Living registry · synced July 31, 2026
Find the right model for the job.
Compare context, input, cached-input, and output rates in one place. Every row links to the pricing page it came from, so you can re-check before production billing.
How to read this tableRates are USD per 1M tokens. “—” means the provider does not publish a separate cached-input rate. Regional, batch, long-context, and search-request surcharges can change the final bill.
| Model | Provider | Context | Input / 1M | Cached / 1M | Output / 1M | Fit & source |
|---|---|---|---|---|---|---|
| Claude Fable 5Fable 5 · claude-fable-5 | Anthropic | 1,000,000 | $10 | $1 | $50 | FlagshipOfficial · 2026-07-31Provider source ↗ |
| Claude Opus 5Opus 5 · claude-opus-5 | Anthropic | 1,000,000 | $5 | $0.5 | $25 | FlagshipOfficial · 2026-07-31Provider source ↗ |
| Claude Sonnet 5Sonnet 5 · claude-sonnet-5Introductory pricing through Aug 31, 2026; standard price becomes $3/$15. | Anthropic | 1,000,000 | $2 | $0.2 | $10 | BalancedOfficial · 2026-07-31Provider source ↗ |
| Claude Haiku 4.5Haiku 4.5 · claude-haiku-4-5 | Anthropic | 200,000 | $1 | $0.1 | $5 | FastOfficial · 2026-07-31Provider source ↗ |
| GPT-5.6 SolGPT-5.6 · gpt-5.6-solShort-context price; prompts over 272K use a higher long-context tier. | OpenAI | 1,050,000 | $5 | $0.5 | $30 | FlagshipOfficial · 2026-07-31Provider source ↗ |
| GPT-5.6 TerraGPT-5.6 · gpt-5.6-terraShort-context price; prompts over 272K use a higher long-context tier. | OpenAI | 1,050,000 | $2 | $0.2 | $12 | BalancedOfficial · 2026-07-31Provider source ↗ |
| GPT-5.6 LunaGPT-5.6 · gpt-5.6-lunaShort-context price; prompts over 272K use a higher long-context tier. | OpenAI | 1,050,000 | $0.2 | $0.02 | $1.2 | FastOfficial · 2026-07-31Provider source ↗ |
| Gemini 3.1 Pro PreviewGemini 3.1 · gemini-3.1-pro-previewStandard price for prompts up to 200K; larger prompts are tiered at $4/$18. | 1,000,000 | $2 | $0.2 | $12 | FlagshipOfficial · 2026-07-31Provider source ↗ | |
| Gemini 3.6 FlashGemini 3.6 · gemini-3.6-flash | 1,000,000 | $1.5 | $0.15 | $7.5 | FastOfficial · 2026-07-31Provider source ↗ | |
| Gemini 3.5 Flash-LiteGemini 3.5 · gemini-3.5-flash-lite | 1,000,000 | $0.3 | $0.03 | $2.5 | FastOfficial · 2026-07-31Provider source ↗ | |
| Grok 4.5Grok 4.5 · grok-4.5xAI publishes input/output pricing; no cached-input rate listed. | xAI | 500,000 | $2 | — | $6 | BalancedOfficial · 2026-07-31Provider source ↗ |
| Kimi K3Kimi K3 · kimi-k3 | Moonshot | 1,048,576 | $3 | $0.3 | $15 | BalancedOfficial · 2026-07-31Provider source ↗ |
| DeepSeek V4 ProDeepSeek V4 · deepseek-v4-pro | DeepSeek | 1,000,000 | $0.435 | $0.003625 | $0.87 | BalancedOfficial · 2026-07-31Provider source ↗ |
| DeepSeek V4 FlashDeepSeek V4 · deepseek-v4-flash | DeepSeek | 1,000,000 | $0.14 | $0.0028 | $0.28 | FastOfficial · 2026-07-31Provider source ↗ |
| Qwen 3 235BQwen 3 · qwen3-235b-a22bInternational standard price; Qwen pricing varies by deployment region and mode. | Alibaba | 262,000 | $0.7 | — | $2.8 | OpenOfficial · 2026-07-31Provider source ↗ |
| Llama 4 MaverickLlama 4 · meta-llama/llama-4-maverickHosted-provider reference; self-hosted Llama has no universal per-token price. | Meta | 1,000,000 | $0.2 | — | $0.8 | OpenProvider-specific · 2026-07-31Provider source ↗ |
| Mistral Large 3Mistral Large · mistral-large-latest | Mistral | 256,000 | $0.5 | — | $1.5 | OpenOfficial · 2026-07-31Provider source ↗ |
| Command ACommand A · command-a-03-2025 | Cohere | 256,000 | $2.5 | — | $10 | BalancedOfficial · 2026-07-31Provider source ↗ |
| Sonar ProSonar · sonar-proToken price excludes Sonar search-context request fees ($6/$10/$14 per 1K requests). | Perplexity | 200,000 | $3 | — | $15 | BalancedOfficial · 2026-07-31Provider source ↗ |