LLM API cost calculator

Compare the cost of one workload across every model.

Enter input, output, monthly volume, and cache-hit assumptions. PromptCostLab applies the same workload to every audited model so you can build a realistic shortlist.

Free · no sign-upCalculations run in your browser.Pricing registry reviewed Jul 31, 2026
START WITH A WORKLOAD
Cost comparison

19 model routes, cheapest first

ModelPer requestPer monthCost mixCache
DeepSeek V4 FlashDeepSeek · Fast$0.0002$0.72Lowest estimateInput $0.43Output $0.2915% applied
Llama 4 MaverickMeta · Open$0.0005$1.56Input $0.72Output $0.84No separate rate
GPT-5.6 LunaOpenAI · Fast$0.0006$1.88Input $0.62Output $1.2615% applied
DeepSeek V4 ProDeepSeek · Balanced$0.0007$2.25Input $1.33Output $0.9115% applied
Mistral Large 3Mistral · Open$0.0011$3.37Input $1.80Output $1.57No separate rate
Gemini 3.5 Flash-LiteGoogle · Fast$0.0012$3.56Input $0.93Output $2.6315% applied
Qwen 3 235BAlibaba · Open$0.0018$5.46Input $2.52Output $2.94No separate rate
Claude Haiku 4.5Anthropic · Fast$0.0028$8.36Input $3.11Output $5.2515% applied

Estimate only. Excludes taxes, batch discounts, long-context tiers, tool/search fees, retries, and negotiated pricing. A cache share is applied only when the registry has a published cached-input rate.

01

Model the same workload

Keep token counts and request volume constant when comparing routes. Otherwise the cheapest-looking result may come from different assumptions rather than a better price.

02

Separate fresh and cached input

Repeated system prompts and stable context can be cheaper when a provider publishes a cached-input rate. Use a cache share you can actually measure.

03

Validate cost per accepted result

A low list price can lose if a model needs more retries or fails your quality bar. Benchmark the shortlist on real requests before routing production traffic.

Plain-English answers

Questions people ask before using the result

How is LLM API cost calculated?

Multiply fresh input, cached input, and output tokens by their respective per-million-token rates, add them together, then multiply by request volume. PromptCostLab does this independently for each model.

Does the cheapest model always cost less in production?

No. Retries, longer outputs, tool fees, search charges, latency, and failure rates can outweigh a lower token price. Treat the table as a shortlist for evaluation.

Are batch and long-context discounts included?

No. The comparison uses the registry's standard planning rates. Provider-specific batch discounts, long-context tiers, taxes, and negotiated pricing are excluded.