Context window calculator

See where your complete request fits safely.

Paste text, reserve output, and choose a safety margin. PromptCostLab compares the total request budget against every model in the registry.

Free · no sign-upCalculations run in your browser.Pricing registry reviewed Jul 31, 2026
Paste prompt, history, or document text
Model fit

Where this request fits safely

Claude Haiku 4.5Anthropic
1.6%176,852 safe tokens left
Sonar ProPerplexity
1.6%176,852 safe tokens left
Mistral Large 3Mistral
1.2%227,252 safe tokens left
Command ACohere
1.2%227,252 safe tokens left
Qwen 3 235BAlibaba
1.2%232,652 safe tokens left
Grok 4.5xAI
0.63%446,852 safe tokens left
Claude Fable 5Anthropic
0.31%896,852 safe tokens left
Claude Opus 5Anthropic
0.31%896,852 safe tokens left
Claude Sonnet 5Anthropic
0.31%896,852 safe tokens left
Gemini 3.1 Pro PreviewGoogle
0.31%896,852 safe tokens left
Gemini 3.6 FlashGoogle
0.31%896,852 safe tokens left
Gemini 3.5 Flash-LiteGoogle
0.31%896,852 safe tokens left
DeepSeek V4 ProDeepSeek
0.31%896,852 safe tokens left
DeepSeek V4 FlashDeepSeek
0.31%896,852 safe tokens left
Llama 4 MaverickMeta
0.31%896,852 safe tokens left
Kimi K3Moonshot
0.30%940,570 safe tokens left
GPT-5.6 SolOpenAI
0.30%941,852 safe tokens left
GPT-5.6 TerraOpenAI
0.30%941,852 safe tokens left
GPT-5.6 LunaOpenAI
0.30%941,852 safe tokens left

Plain-text estimate only. Chat roles, tool schemas, images, files, hidden formatting, and model-specific tokenizers can change the real count. Verify strict limits with the provider.

01

Reserve output before checking fit

The response shares the request budget on many APIs and can also have a separate cap. Plan output first so a prompt that technically fits still has room to answer.

02

Keep operational headroom

SDK formatting, chat roles, tool schemas, retrieved evidence, and longer user inputs add tokens. A 10–20% margin is a practical starting point for production planning.

03

Improve retrieval before upgrading

A larger context window can hide weak retrieval. Select fewer, more relevant passages before paying for a bigger route or accepting slower requests.

Plain-English answers

Questions people ask before using the result

What is included in an LLM context window?

System instructions, user messages, conversation history, retrieved documents, tool definitions and results, hidden request formatting, and the generated response can all consume context.

Why should I leave safety headroom?

Your planning text may not include chat wrappers, tools, files, images, or unusual user inputs. Headroom reduces truncation and request-limit failures when the real payload is larger.

Is the displayed token count exact?

No. It is a private plain-text estimate. Use provider token-counting APIs or returned usage fields for strict production limits and billing.