Model the same workload
Keep token counts and request volume constant when comparing routes. Otherwise the cheapest-looking result may come from different assumptions rather than a better price.
Enter input, output, monthly volume, and cache-hit assumptions. PromptCostLab applies the same workload to every audited model so you can build a realistic shortlist.
| Model | Per request | Per month | Cost mix | Cache |
|---|---|---|---|---|
| DeepSeek V4 FlashDeepSeek · Fast | $0.0002 | $0.72Lowest estimate | Input $0.43Output $0.29 | 15% applied |
| Llama 4 MaverickMeta · Open | $0.0005 | $1.56 | Input $0.72Output $0.84 | No separate rate |
| GPT-5.6 LunaOpenAI · Fast | $0.0006 | $1.88 | Input $0.62Output $1.26 | 15% applied |
| DeepSeek V4 ProDeepSeek · Balanced | $0.0007 | $2.25 | Input $1.33Output $0.91 | 15% applied |
| Mistral Large 3Mistral · Open | $0.0011 | $3.37 | Input $1.80Output $1.57 | No separate rate |
| Gemini 3.5 Flash-LiteGoogle · Fast | $0.0012 | $3.56 | Input $0.93Output $2.63 | 15% applied |
| Qwen 3 235BAlibaba · Open | $0.0018 | $5.46 | Input $2.52Output $2.94 | No separate rate |
| Claude Haiku 4.5Anthropic · Fast | $0.0028 | $8.36 | Input $3.11Output $5.25 | 15% applied |
Estimate only. Excludes taxes, batch discounts, long-context tiers, tool/search fees, retries, and negotiated pricing. A cache share is applied only when the registry has a published cached-input rate.
Keep token counts and request volume constant when comparing routes. Otherwise the cheapest-looking result may come from different assumptions rather than a better price.
Repeated system prompts and stable context can be cheaper when a provider publishes a cached-input rate. Use a cache share you can actually measure.
A low list price can lose if a model needs more retries or fails your quality bar. Benchmark the shortlist on real requests before routing production traffic.
Multiply fresh input, cached input, and output tokens by their respective per-million-token rates, add them together, then multiply by request volume. PromptCostLab does this independently for each model.
No. Retries, longer outputs, tool fees, search charges, latency, and failure rates can outweigh a lower token price. Treat the table as a shortlist for evaluation.
No. The comparison uses the registry's standard planning rates. Provider-specific batch discounts, long-context tiers, taxes, and negotiated pricing are excluded.