LLM API cost calculator

Compare the cost of one workload across every model.

Enter input, output, monthly volume, and cache-hit assumptions. PromptCostLab applies the same workload to every audited model so you can build a realistic shortlist.

Free · no sign-upCalculations run in your browser.Pricing registry reviewed Aug 16, 2026
START WITH A WORKLOAD
Cost comparison

23 model routes, cheapest first

ModelPer requestPer monthCost mixCache
DeepSeek V4 FlashDeepSeek · Fast$0.0002$0.7239Lowest estimateInput $0.4299Output $0.294015% applied
Llama 4 MaverickMeta · Open$0.0005$1.56Input $0.7200Output $0.8400No separate rate
GPT-5.6 LunaOpenAI · Fast$0.0006$1.88Input $0.6228Output $1.2615% applied
DeepSeek V4 ProDeepSeek · Balanced$0.0007$2.25Input $1.33Output $0.913515% applied
Mistral Large 3Mistral · Open$0.0010$3.13Input $1.56Output $1.5715% applied
Gemini 3.5 Flash-LiteGoogle · Fast$0.0012$3.56Input $0.9342Output $2.6315% applied
Gemini 3.6 FlashGoogle · Fast$0.0021$6.27Input $2.34Output $3.9415% applied
Gemini 3.7 FlashGoogle · Fast$0.0021$6.27Input $2.34Output $3.9415% applied

Estimate only — long-context tiers are applied automatically where the registry records them. Excludes taxes, batch discounts, tool/search fees, retries, and negotiated pricing. A cache share applies only when the model has a published cached-input rate.

01

Model the same workload

Keep token counts and request volume constant when comparing routes. Otherwise the cheapest-looking result may come from different assumptions rather than a better price.

02

Separate fresh and cached input

Repeated system prompts and stable context can be cheaper when a provider publishes a cached-input rate. Use a cache share you can actually measure.

03

Validate cost per accepted result

A low list price can lose if a model needs more retries or fails your quality bar. Benchmark the shortlist on real requests before routing production traffic.

Plain-English answers

Questions people ask before using the result

How is LLM API cost calculated?

Multiply fresh input, cached input, and output tokens by their respective per-million-token rates, add them together, then multiply by request volume. PromptCostLab does this independently for each model.

Does the cheapest model always cost less in production?

No. Retries, longer outputs, tool fees, search charges, latency, and failure rates can outweigh a lower token price. Treat the table as a shortlist for evaluation.

Are batch and long-context discounts included?

No. The comparison uses the registry's standard planning rates. Provider-specific batch discounts, long-context tiers, taxes, and negotiated pricing are excluded.