Living registry · synced Aug 16, 2026
Find the right model for the job.
Compare context, input, cached-input, and output rates in one place. Every row links to the pricing page it came from, so you can re-check before production billing.
How to read this tableRates are USD per 1M tokens. “—” means the provider does not publish a separate cached-input rate. Regional, batch, long-context, and search-request surcharges can change the final bill.
| Model | Provider | Context | Input / 1M | Cached / 1M | Output / 1M | Fit & source |
|---|---|---|---|---|---|---|
| Claude Fable 5Fable 5 · claude-fable-5Always uses extended thinking; thinking tokens are billed as output. | Anthropic | 1,000,000 | $10 | $1 | $50 | FlagshipOfficial · 2026-08-16Provider source ↗ |
| Claude Mythos 5Mythos 5 · claude-mythos-5Limited availability; priced identically to Fable 5. | Anthropic | 1,000,000 | $10 | $1 | $50 | FlagshipOfficial · 2026-08-16Provider source ↗ |
| Claude Opus 5Opus 5 · claude-opus-5Fast mode doubles rates to $10/$50. | Anthropic | 1,000,000 | $5 | $0.5 | $25 | FlagshipOfficial · 2026-08-16Provider source ↗ |
| Claude Sonnet 5Sonnet 5 · claude-sonnet-5The $2/$10 launch price is now standard; the previously announced Sept 2026 increase to $3/$15 was cancelled. | Anthropic | 1,000,000 | $2 | $0.2 | $10 | BalancedOfficial · 2026-08-16Provider source ↗ |
| Claude Haiku 4.5Haiku 4.5 · claude-haiku-4-5 | Anthropic | 200,000 | $1 | $0.1 | $5 | FastOfficial · 2026-08-16Provider source ↗ |
| GPT-5.6 SolGPT-5.6 · gpt-5.6-solShort-context price; requests over 272K input pay doubled input and 1.5x output on the entire request. | OpenAI | 1,050,000 | $5 | $0.5 | $30 | FlagshipOfficial · 2026-08-16Provider source ↗ |
| GPT-5.6 TerraGPT-5.6 · gpt-5.6-terraShort-context price; requests over 272K input pay doubled input and 1.5x output on the entire request. | OpenAI | 1,050,000 | $2 | $0.2 | $12 | BalancedOfficial · 2026-08-16Provider source ↗ |
| GPT-5.6 LunaGPT-5.6 · gpt-5.6-lunaShort-context price; requests over 272K input pay doubled input and 1.5x output on the entire request. | OpenAI | 1,050,000 | $0.2 | $0.02 | $1.2 | FastOfficial · 2026-08-16Provider source ↗ |
| Gemini 3.1 Pro PreviewGemini 3.1 · gemini-3.1-pro-previewStandard price for prompts up to 200K; larger prompts are tiered at $4/$18. Explicit cache storage costs $4.50 per 1M tokens per hour. | 1,048,576 | $2 | $0.2 | $12 | FlagshipOfficial · 2026-08-16Provider source ↗ | |
| Gemini 3.6 FlashGemini 3.6 · gemini-3.6-flashIntroductory pricing through Dec 31, 2026; standard price becomes $1.50/$7.50 with $0.15 cached. Explicit cache storage $0.50 per 1M tokens per hour during intro. | 1,048,576 | $0.75 | $0.075 | $3.75 | FastOfficial · 2026-08-16Provider source ↗ | |
| Gemini 3.7 FlashGemini 3.7 · gemini-3.7-flashReleased Aug 13, 2026. Introductory pricing through Dec 31, 2026; standard price becomes $1.50/$7.50 with $0.15 cached. | 1,048,576 | $0.75 | $0.075 | $3.75 | FastOfficial · 2026-08-16Provider source ↗ | |
| Gemini 3.5 Flash-LiteGemini 3.5 · gemini-3.5-flash-liteExplicit cache storage costs $1.00 per 1M tokens per hour. | 1,048,576 | $0.3 | $0.03 | $2.5 | FastOfficial · 2026-08-16Provider source ↗ | |
| Grok 4.6Grok 4.6 · grok-4.6Current flagship (Aug 2026). Prompts of 200K+ tokens pay doubled rates on the entire request. | xAI | 500,000 | $2 | $0.5 | $6 | FlagshipOfficial · 2026-08-16Provider source ↗ |
| Grok 4.5Grok 4.5 · grok-4.5Prompts of 200K+ tokens pay doubled rates on the entire request. | xAI | 500,000 | $2 | $0.3 | $6 | BalancedOfficial · 2026-08-16Provider source ↗ |
| Kimi K3Kimi K3 · kimi-k3 | Moonshot | 1,048,576 | $3 | $0.3 | $15 | BalancedOfficial · 2026-08-16Provider source ↗ |
| Kimi K2.7 CodeKimi K2.7 · kimi-k2.7-codeHighspeed variant costs $1.90 input / $8.00 output per 1M. | Moonshot | 262,144 | $0.95 | $0.19 | $4 | FastOfficial · 2026-08-16Provider source ↗ |
| DeepSeek V4 ProDeepSeek V4 · deepseek-v4-proAutomatic disk caching makes cache hits roughly 1/120th of the fresh input price. | DeepSeek | 1,048,576 | $0.435 | $0.003625 | $0.87 | BalancedOfficial · 2026-08-16Provider source ↗ |
| DeepSeek V4 FlashDeepSeek V4 · deepseek-v4-flash | DeepSeek | 1,048,576 | $0.14 | $0.0028 | $0.28 | FastOfficial · 2026-08-16Provider source ↗ |
| Qwen 3.8 MaxQwen 3.8 · qwen3.8-maxReference rate via OpenRouter; Alibaba list pricing varies by region and often runs promotional discounts. | Alibaba | 1,000,000 | $2 | $0.25 | $6 | FlagshipProvider-specific · 2026-08-16Provider source ↗ |
| Llama 4 MaverickLlama 4 · meta-llama/llama-4-maverickHosted-provider reference; self-hosted Llama has no universal per-token price. | Meta | 1,000,000 | $0.2 | — | $0.8 | OpenProvider-specific · 2026-08-16Provider source ↗ |
| Mistral Large 3Mistral Large · mistral-large-latest | Mistral | 262,144 | $0.5 | $0.05 | $1.5 | OpenOfficial · 2026-08-16Provider source ↗ |
| Command ACommand A · command-a-plus-05-2026Current generation (05-2026 refresh) at the same $2.50/$10 price. | Cohere | 256,000 | $2.5 | — | $10 | BalancedOfficial · 2026-08-16Provider source ↗ |
| Sonar ProSonar · sonar-proToken price excludes Sonar search-context request fees ($6/$10/$14 per 1K requests). The Sonar API sunsets Sep 27, 2026 in favor of the Router API. | Perplexity | 200,000 | $3 | — | $15 | BalancedOfficial · 2026-08-16Provider source ↗ |