Living registry · synced Aug 16, 2026

Find the right model for the job.

Compare context, input, cached-input, and output rates in one place. Every row links to the pricing page it came from, so you can re-check before production billing.

How to read this tableRates are USD per 1M tokens. “—” means the provider does not publish a separate cached-input rate. Regional, batch, long-context, and search-request surcharges can change the final bill.
ModelProviderContextInput / 1MCached / 1MOutput / 1MFit & source
Claude Fable 5Fable 5 · claude-fable-5Always uses extended thinking; thinking tokens are billed as output.Anthropic1,000,000$10$1$50FlagshipOfficial · 2026-08-16Provider source ↗
Claude Mythos 5Mythos 5 · claude-mythos-5Limited availability; priced identically to Fable 5.Anthropic1,000,000$10$1$50FlagshipOfficial · 2026-08-16Provider source ↗
Claude Opus 5Opus 5 · claude-opus-5Fast mode doubles rates to $10/$50.Anthropic1,000,000$5$0.5$25FlagshipOfficial · 2026-08-16Provider source ↗
Claude Sonnet 5Sonnet 5 · claude-sonnet-5The $2/$10 launch price is now standard; the previously announced Sept 2026 increase to $3/$15 was cancelled.Anthropic1,000,000$2$0.2$10BalancedOfficial · 2026-08-16Provider source ↗
Claude Haiku 4.5Haiku 4.5 · claude-haiku-4-5Anthropic200,000$1$0.1$5FastOfficial · 2026-08-16Provider source ↗
GPT-5.6 SolGPT-5.6 · gpt-5.6-solShort-context price; requests over 272K input pay doubled input and 1.5x output on the entire request.OpenAI1,050,000$5$0.5$30FlagshipOfficial · 2026-08-16Provider source ↗
GPT-5.6 TerraGPT-5.6 · gpt-5.6-terraShort-context price; requests over 272K input pay doubled input and 1.5x output on the entire request.OpenAI1,050,000$2$0.2$12BalancedOfficial · 2026-08-16Provider source ↗
GPT-5.6 LunaGPT-5.6 · gpt-5.6-lunaShort-context price; requests over 272K input pay doubled input and 1.5x output on the entire request.OpenAI1,050,000$0.2$0.02$1.2FastOfficial · 2026-08-16Provider source ↗
Gemini 3.1 Pro PreviewGemini 3.1 · gemini-3.1-pro-previewStandard price for prompts up to 200K; larger prompts are tiered at $4/$18. Explicit cache storage costs $4.50 per 1M tokens per hour.Google1,048,576$2$0.2$12FlagshipOfficial · 2026-08-16Provider source ↗
Gemini 3.6 FlashGemini 3.6 · gemini-3.6-flashIntroductory pricing through Dec 31, 2026; standard price becomes $1.50/$7.50 with $0.15 cached. Explicit cache storage $0.50 per 1M tokens per hour during intro.Google1,048,576$0.75$0.075$3.75FastOfficial · 2026-08-16Provider source ↗
Gemini 3.7 FlashGemini 3.7 · gemini-3.7-flashReleased Aug 13, 2026. Introductory pricing through Dec 31, 2026; standard price becomes $1.50/$7.50 with $0.15 cached.Google1,048,576$0.75$0.075$3.75FastOfficial · 2026-08-16Provider source ↗
Gemini 3.5 Flash-LiteGemini 3.5 · gemini-3.5-flash-liteExplicit cache storage costs $1.00 per 1M tokens per hour.Google1,048,576$0.3$0.03$2.5FastOfficial · 2026-08-16Provider source ↗
Grok 4.6Grok 4.6 · grok-4.6Current flagship (Aug 2026). Prompts of 200K+ tokens pay doubled rates on the entire request.xAI500,000$2$0.5$6FlagshipOfficial · 2026-08-16Provider source ↗
Grok 4.5Grok 4.5 · grok-4.5Prompts of 200K+ tokens pay doubled rates on the entire request.xAI500,000$2$0.3$6BalancedOfficial · 2026-08-16Provider source ↗
Kimi K3Kimi K3 · kimi-k3Moonshot1,048,576$3$0.3$15BalancedOfficial · 2026-08-16Provider source ↗
Kimi K2.7 CodeKimi K2.7 · kimi-k2.7-codeHighspeed variant costs $1.90 input / $8.00 output per 1M.Moonshot262,144$0.95$0.19$4FastOfficial · 2026-08-16Provider source ↗
DeepSeek V4 ProDeepSeek V4 · deepseek-v4-proAutomatic disk caching makes cache hits roughly 1/120th of the fresh input price.DeepSeek1,048,576$0.435$0.003625$0.87BalancedOfficial · 2026-08-16Provider source ↗
DeepSeek V4 FlashDeepSeek V4 · deepseek-v4-flashDeepSeek1,048,576$0.14$0.0028$0.28FastOfficial · 2026-08-16Provider source ↗
Qwen 3.8 MaxQwen 3.8 · qwen3.8-maxReference rate via OpenRouter; Alibaba list pricing varies by region and often runs promotional discounts.Alibaba1,000,000$2$0.25$6FlagshipProvider-specific · 2026-08-16Provider source ↗
Llama 4 MaverickLlama 4 · meta-llama/llama-4-maverickHosted-provider reference; self-hosted Llama has no universal per-token price.Meta1,000,000$0.2$0.8OpenProvider-specific · 2026-08-16Provider source ↗
Mistral Large 3Mistral Large · mistral-large-latestMistral262,144$0.5$0.05$1.5OpenOfficial · 2026-08-16Provider source ↗
Command ACommand A · command-a-plus-05-2026Current generation (05-2026 refresh) at the same $2.50/$10 price.Cohere256,000$2.5$10BalancedOfficial · 2026-08-16Provider source ↗
Sonar ProSonar · sonar-proToken price excludes Sonar search-context request fees ($6/$10/$14 per 1K requests). The Sonar API sunsets Sep 27, 2026 in favor of the Router API.Perplexity200,000$3$15BalancedOfficial · 2026-08-16Provider source ↗