What is a context window — and how big is yours?
A context window is the maximum number of tokens an AI model can process in a single request. It includes your instructions, the conversation, retrieved documents, tool results, and the model’s reply. Current models range from 200,000 to about 1,050,000 tokens — roughly one book to ten.
What actually fits in 200K to 1M tokens
| Context size | English words | What that is |
|---|---|---|
| 2,000 tokens | ≈ 1,500 words | a long email thread or meeting notes |
| 8,192 tokens | ≈ 6,100 words | a lengthy technical guide |
| 32,000 tokens | ≈ 24,000 words | a whitepaper with appendices |
| 128K tokens (131,072) | ≈ 98,000 words | a full novel |
| 200K tokens | ≈ 150,000 words | a 300–500 page book — Claude Haiku/Sonnet-class window |
| 500K tokens | ≈ 375,000 words | a large codebase review — Grok-class window |
| 1M tokens | ≈ 750,000 words | about ten novels — current flagship windows |
Word equivalents use the standard 1 token ≈ 0.75 words planning ratio. The window is shared: instructions, history, evidence, and the answer all draw from the same budget.
Every model’s context window, Aug 16, 2026
| Model | Provider | Context window | Max output | Long-context pricing |
|---|---|---|---|---|
| GPT-5.6 SolGPT-5.6 | OpenAI | 1,050,000 (1.05M) | 128,000 | Tiers above 272K |
| GPT-5.6 TerraGPT-5.6 | OpenAI | 1,050,000 (1.05M) | 128,000 | Tiers above 272K |
| GPT-5.6 LunaGPT-5.6 | OpenAI | 1,050,000 (1.05M) | 128,000 | Tiers above 272K |
| Gemini 3.1 Pro PreviewGemini 3.1 | 1,048,576 (1.05M) | 65,536 | Tiers above 200K | |
| Gemini 3.6 FlashGemini 3.6 | 1,048,576 (1.05M) | 65,536 | Standard rates | |
| Gemini 3.7 FlashGemini 3.7 | 1,048,576 (1.05M) | 65,536 | Standard rates | |
| Gemini 3.5 Flash-LiteGemini 3.5 | 1,048,576 (1.05M) | 65,536 | Standard rates | |
| Kimi K3Kimi K3 | Moonshot | 1,048,576 (1.05M) | — | Standard rates |
| DeepSeek V4 ProDeepSeek V4 | DeepSeek | 1,048,576 (1.05M) | 384,000 | Standard rates |
| DeepSeek V4 FlashDeepSeek V4 | DeepSeek | 1,048,576 (1.05M) | 390,000 | Standard rates |
| Claude Fable 5Fable 5 | Anthropic | 1,000,000 (1M) | 128,000 | Standard rates |
| Claude Mythos 5Mythos 5 | Anthropic | 1,000,000 (1M) | 128,000 | Standard rates |
| Claude Opus 5Opus 5 | Anthropic | 1,000,000 (1M) | 128,000 | Standard rates |
| Claude Sonnet 5Sonnet 5 | Anthropic | 1,000,000 (1M) | 128,000 | Standard rates |
| Qwen 3.8 MaxQwen 3.8 | Alibaba | 1,000,000 (1M) | 131,000 | Standard rates |
| Llama 4 MaverickLlama 4 | Meta | 1,000,000 (1M) | — | Standard rates |
| Grok 4.6Grok 4.6 | xAI | 500,000 (500K) | 128,000 | Tiers above 200K |
| Grok 4.5Grok 4.5 | xAI | 500,000 (500K) | 128,000 | Tiers above 200K |
| Kimi K2.7 CodeKimi K2.7 | Moonshot | 262,144 (262K) | — | Standard rates |
| Mistral Large 3Mistral Large | Mistral | 262,144 (262K) | — | Standard rates |
| Command ACommand A | Cohere | 256,000 (256K) | 64,000 | Standard rates |
| Claude Haiku 4.5Haiku 4.5 | Anthropic | 200,000 (200K) | 64,000 | Standard rates |
| Sonar ProSonar | Perplexity | 200,000 (200K) | — | Standard rates |
Windows verified from provider documentation. Long-context tiers apply higher per-token rates above the listed input size — the context-window calculator applies them automatically.
Context window questions
What is a context window?
A context window is the maximum number of tokens a model can process in a single request. It covers system instructions, conversation history, retrieved documents, tool definitions, and the space for the model’s reply. Mainstream models currently offer between 200,000 and about 1,050,000 tokens.
What fits in a 200K context window?
Roughly 150,000 English words — a 300-to-500-page book, a mid-sized codebase, or about 400 pages of extracted PDF text. The window is shared with the model's answer, so effective document capacity is somewhat less.
Do all models have a 1 million token context?
No. In August 2026, Anthropic's Fable/Mythos/Opus/Sonnet, OpenAI's GPT-5.6 family, Google's Gemini lineup, Kimi K3, and DeepSeek V4 all offer roughly 1M-token windows. Grok 4.6/4.5 top out at 500K, Kimi K2.7 Code at 262K, Cohere Command A at 256K, and several models at 200K.
Does a bigger context window cost more?
Often, yes — beyond the obvious per-token input cost. Several providers apply long-context tiers: OpenAI doubles input above 272K tokens, xAI doubles rates at 200K+ prompts, and Gemini 3.1 Pro charges $4/$18 above 200K instead of $2/$12.