AI glossary
Context Window
A context window is the total amount of text a large language model can take into account in a single request, measured in tokens. It includes everything the model is given (system instructions, the conversation so far, attached documents, tool definitions) and everything it generates in reply. Anthropic’s documentation defines it as “all the text a language model can reference when generating a response, including the response itself.”
The context window works like the model’s working memory for one request. It is separate from what the model learned during training: training knowledge is fixed in the weights, while the context window holds only what is passed in now. Anything that does not fit, or is not passed in, the model cannot see.
How a context window is measured
Context windows are counted in tokens, the chunks of text a model actually processes, not in words or characters. OpenAI’s rule of thumb for English is that “1 token is approximately 4 characters” and “approximately three-quarters of a word”, so 100 tokens are about 75 words. On that basis:
| Context window | Roughly, in English words | Roughly, in pages |
|---|---|---|
| 8,192 tokens | about 6,000 | about 12 |
| 128,000 tokens | about 96,000 | a long novel |
| 200,000 tokens | about 150,000 | over 500 (Anthropic’s own estimate) |
| 1,000,000 tokens | about 750,000 | several books |
Other languages and code usually take more tokens per word, so the same window holds less text.
What counts against the window
Input and output share the same budget. If a model has a 200,000-token window and the request already uses 190,000, only 10,000 tokens remain for the answer. Three things are easy to miss:
- The whole conversation is resent. Chat models have no memory between calls; each new message is sent along with the history, so long chats fill the window turn by turn.
- Reasoning tokens count. For reasoning models, the hidden thinking is part of the output. Anthropic states that “all input and output tokens, including thinking tokens, count toward the context window limit”, and OpenAI advises leaving room for reasoning tokens as well as the visible answer.
- Output has its own, smaller cap. Most models limit how many tokens they can generate in one response, well below the full window. Our best AI model for writing guide compares those output caps.
How context windows grew
| Model | Context window | Date |
|---|---|---|
| GPT-3 | 2,048 tokens | May 2020 |
| GPT-4 | 8,192 tokens (32,768 in limited access) | March 2023 |
| Claude | 100,000 tokens (up from 9,000) | May 2023 |
| Claude 2.1 | 200,000 tokens | November 2023 |
| Gemini 1.5 Pro | 128,000 standard; 1 million in preview, 10 million tested in research | February 2024 |
| Gemini 1.5 Pro | 2 million for developers | June 2024 |
Today a million tokens is common at the top end. Of the seven models this site tracks, five offer roughly one million tokens or more: GPT-5.6 Sol (1,050,000), Gemini 3.1 Pro and DeepSeek V4 Pro (1,048,576 each), and Claude Sonnet 5 and Qwen3.8 Max (1,000,000 each). Grok 4.6 offers 500,000 and Llama 4 Maverick 128,000. Current figures are on our model comparison pages.
Why bigger is not automatically better
A larger window lets a model read a whole codebase, contract or book in one go. But size has three costs.
Compute. In a standard transformer, self-attention compares every token with every other token. The original “Attention Is All You Need” paper lists its cost per layer as O(n²·d), quadratic in sequence length n, so doubling the input roughly quadruples that part of the work. The model also keeps a KV cache entry for every token held in context, so memory use rises with every token added.
Price. You pay for every input token on every call. Some providers charge more for very long prompts: Google’s Gemini API pricing doubles the input rate for Gemini 3.1 Pro Preview above 200,000 tokens ($2.00 to $4.00 per million). Others do not; Anthropic says its Claude 4.6 and later models include the full 1M window “at standard pricing”. Prompt caching is the main way to cut the cost of resending the same long context.
Accuracy. Models do not use long context evenly:
- Lost in the middle. Liu et al. (TACL) found that performance “is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades” when the key information sits in the middle.
- Effective vs advertised length. NVIDIA’s RULER benchmark (2024) tested long-context models that “all claim context sizes of 32K tokens or greater”, and found “only half of them can maintain satisfactory performance at the length of 32K.”
- Context rot. Anthropic’s own documentation notes that “as token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
This is why the practical skill is not filling the window but choosing what goes into it, a discipline now called context engineering.
Context window vs related ideas
- Context window vs training data. Training data shapes the model’s weights once; the context window is what it sees in a single request. A model can only cite your document if the document is in the window.
- Context window vs RAG. Retrieval-augmented generation does not enlarge the window. It searches an external store and puts only the relevant passages into it.
- Context window vs memory features. Chat apps that “remember” you store notes outside the model and insert them into the context window when needed. The underlying model itself still starts each request with an empty window.
FAQ
What is a context window in AI?
The maximum amount of text, measured in tokens, that a language model can consider in one request, counting both the input it receives and the output it writes.
How many words is a 128K context window?
About 96,000 English words, using OpenAI’s rule of thumb that a token is roughly three-quarters of a word. Code and many non-English languages use more tokens per word.
What happens when you exceed the context window?
The API rejects the request or the application drops or summarizes older content to make room. In chat apps this is why a model can “forget” the start of a very long conversation.
Does a bigger context window make a model better?
Not by itself. It allows longer inputs, but research shows accuracy drops as inputs grow and when key information is buried in the middle. Longer prompts also cost more and run slower.
Do output tokens count toward the context window?
Yes. Input and output share the window, and for reasoning models the hidden thinking tokens count as well.
Related terms
- Tokens
- Tokenization
- Context engineering
- KV cache
- Attention mechanism
- Retrieval-augmented generation (RAG)
- Large language model
For how attention works under the hood, see How Transformer Attention Actually Works.