Context windows
The context window is the total text a model can hold at once, and exceeding it does not warn you — it either errors, silently drops your oldest messages, or cuts the reply off mid-sentence.
- 21 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
The context window is the total amount of text a model can hold in front of it at one time.
The analogy you have already lived
Picture the whiteboard in a classroom. The teacher works through a problem, filling the board line by line.
The board fills up. To carry on, they wipe the top few lines and write in the space. The earlier working is not saved anywhere. It is gone.
Now the class asks about a step from ten minutes ago. The teacher cannot look it up, because that part of the board no longer exists.
A language model works on exactly this whiteboard. The context window is how big the board is.
Why it exists at all
Every piece of text you send gets chopped into tokens. A token is a small chunk of text, often part of a word. In English one token averages about three-quarters of a word. The tokenization lesson covers the chopping.
The model compares every token against every other token, so that each one can borrow meaning from the rest. That is attention, and it is the reason these models work.
The catch is in the word "every". Double the text and the number of comparisons goes up roughly four times, not two. Memory goes up too. So there has to be a ceiling, and the ceiling is the context window.
The thing almost everyone gets wrong
The model has no memory between messages. None.
When you send your fortieth message in a chat, the app does not send only that message. It sends the entire conversation again, from the first line, every single time.
you type: the app actually sends:
"and what about the tax?" [system instructions]
[your message 1]
[reply 1]
[your message 2]
...
[your message 39]
[reply 39]
"and what about the tax?"The model reads that whole stack fresh, answers, and forgets everything the instant it finishes.
So "it remembers what I told it earlier" is not what happens. Your app is re-reading the transcript to it, every turn, like reminding a stranger of the whole conversation before each question.
This is confusing the first time. Read it once more, because everything else on this page follows from it.
What actually happens when you go over
This is the part nobody explains, so here it is plainly. Three different things can happen, and which one you get depends on the tool.
One: it refuses. If you call a model directly through an API, you get an error. The message names the model's limit and how far past it you went. Nothing is generated. You are billed nothing, and your code either handles it or crashes.
Two: it quietly forgets the beginning. Most chat apps never show you an error. They drop the oldest messages to make room and carry on as if nothing happened. This is the whiteboard being wiped.
The consequences are strange to experience. The assistant starts ignoring an instruction you gave at the start. It asks for a detail you already provided. It contradicts something it said an hour ago. Nothing broke, and nothing warned you.
Three: the reply gets cut off. This one surprises people most. The window is shared between what you send and what comes back.
Suppose the window holds eight thousand tokens. Your input uses seven thousand nine hundred. That leaves a hundred for the answer. You get a reply that stops in the middle of a sentence. The model did not lose its train of thought. It ran out of board.
|<--------------- the whole window ---------------->|
| |
[ system instructions ][ conversation ][ your ask ][ reply ]
^^^^^^^
whatever is left overWhere you have already seen this
- A long chat where the assistant "forgets" your instruction. You told it to answer in Hindi at message three. By message sixty, message three has been wiped off the board.
- A coding assistant losing the file you pasted. The file was in the window. Then it was not.
- A summariser refusing a long PDF. That is behaviour one, an outright refusal.
- An answer that stops mid-sentence. That is behaviour three, the shared budget.
Bigger windows are not the answer
Windows have grown enormously. Some models now accept the text of several novels at once. It is tempting to conclude the problem is solved.
Two reasons it is not.
You pay for every token, every turn. Re-sending a huge conversation forty times means paying for it forty times. The developer section shows the arithmetic, and it is worse than most people expect.
Models use the middle of a long window poorly. Put an important fact at the start or the end of a long document and the model finds it reliably. Bury it in the middle and it is measurably more likely to be missed. This is a documented, repeatable effect, not a rumour.
So a big window lets you fit more. It does not guarantee the model uses it. Those are two different achievements, and only the first one gets advertised.
Remember this
- The window covers everything: instructions, the whole conversation, and the reply that has not been written yet.
- The model has no memory between turns. Your app resends the transcript each time.
- Going over gives you one of three things: an error, silently dropped history, or a truncated reply.
What to learn next
- What is RAG? — fetching only the relevant text instead of sending everything.
- Tokenization — what a token actually is and why counting is not counting words.
- Temperature and sampling — the other setting people get wrong.
Developer — Code and libraries.
Three short programs. The first shows what truncation does to a real conversation. The second shows why long chats get expensive faster than anyone expects. The third is a budget check you should put in front of every model call you ship.
Everything here is plain Python. No dependencies, no network.
Counting tokens
You need a token count before you can manage a budget. The correct way is to use the model's own tokenizer:
pip install tiktoken # OpenAI modelsimport tiktoken
enc = tiktoken.get_encoding("cl100k_base")
print(len(enc.encode("your text here")))No output block here. tiktoken downloads its encoding file on first use, and the
count depends on which encoding your model uses — cl100k_base and o200k_base
disagree on the same string. Run it against your own model's encoding rather than
trusting a number printed on this page.
For the examples below we use a rough stand-in, so the code runs anywhere with nothing installed:
def rough_tokens(text):
return max(1, round(len(text) / 4))One token per four characters is the common rule of thumb for English prose. It is wrong in predictable directions: code splits into more tokens per character, numbers split badly, and Hindi or Tamil written in their own scripts can cost several times more tokens per word than English. Never bill a customer on this estimate. Use it for illustration and use the real tokenizer in production.
What truncation actually deletes
def rough_tokens(text):
return max(1, round(len(text) / 4))
WINDOW = 60 # tiny on purpose, so the wall arrives on screen
RESERVED_FOR_REPLY = 20
conversation = [
("system", "You are a support agent for a bike rental shop in Pune."),
("user", "My booking id is BK-4471 and the bike never arrived."),
("assistant", "I am sorry to hear that. Let me look into it."),
("user", "It was supposed to reach me at 9am today."),
("assistant", "The rider is delayed by traffic near Shivajinagar."),
("user", "Fine. Remind me what my booking id was."),
]
def fit(messages, window, reserved, pin_system=False):
budget = window - reserved
pinned = [m for m in messages if pin_system and m[0] == "system"]
used = sum(rough_tokens(t) for _, t in pinned)
rest = [m for m in messages if m not in pinned]
kept, dropped = [], []
for role, text in reversed(rest): # newest first: the oldest falls off
n = rough_tokens(text)
if used + n <= budget:
kept.append((role, text)); used += n
else:
dropped.append((role, text))
return pinned + list(reversed(kept)), list(reversed(dropped)), used
for label, pin in [("NAIVE: drop from the front", False),
("FIXED: pin the system message", True)]:
kept, dropped, used = fit(conversation, WINDOW, RESERVED_FOR_REPLY, pin)
print(f"--- {label} --- {used}/{WINDOW - RESERVED_FOR_REPLY} input tokens used")
for role, text in kept:
print(f" sent [{role:9}] {text}")
for role, text in dropped:
print(f" DROPPED [{role:9}] {text}")
print()--- NAIVE: drop from the front --- 32/40 input tokens used sent [user ] It was supposed to reach me at 9am today. sent [assistant] The rider is delayed by traffic near Shivajinagar. sent [user ] Fine. Remind me what my booking id was. DROPPED [system ] You are a support agent for a bike rental shop in Pune. DROPPED [user ] My booking id is BK-4471 and the bike never arrived. DROPPED [assistant] I am sorry to hear that. Let me look into it. --- FIXED: pin the system message --- 36/40 input tokens used sent [system ] You are a support agent for a bike rental shop in Pune. sent [assistant] The rider is delayed by traffic near Shivajinagar. sent [user ] Fine. Remind me what my booking id was. DROPPED [user ] My booking id is BK-4471 and the bike never arrived. DROPPED [assistant] I am sorry to hear that. Let me look into it. DROPPED [user ] It was supposed to reach me at 9am today.
Read the first block carefully. The user asks for their booking id, and the message containing the booking id has been deleted. So has the system prompt, which means the assistant no longer knows it works for a bike rental shop.
The model will answer anyway. It will apologise, or invent a plausible id, or ask a question it already asked. From the outside this looks like the model being stupid. It is not. You deleted the answer before asking the question.
The second block fixes half of it. Pinning the system message keeps the persona alive. The booking id still falls off. Pinning is necessary and it is not sufficient — that is the honest result, and it is why the real fix is further down this page.
Why long chats get expensive so fast
Every turn re-sends the entire history. That makes cost grow with the square of the conversation length, not with what you typed.
SYSTEM = 500 # tokens of instructions resent on every single turn
PER_USER = 150
PER_REPLY = 350
history = SYSTEM
total_input = 0
print(f"{'turn':>4} {'input tokens sent':>18} {'running input total':>20}")
for turn in range(1, 21):
history += PER_USER
total_input += history # the WHOLE history is re-sent, every turn
history += PER_REPLY # the reply becomes part of the next turn's input
if turn in (1, 2, 5, 10, 20):
print(f"{turn:>4} {history - PER_REPLY:>18,} {total_input:>20,}")
print()
print(f"turn 20 alone costs {(history - PER_REPLY) / (SYSTEM + PER_USER):.1f}x turn 1")
print(f"cumulative input tokens for a 20-turn chat: {total_input:,}")
print(f"the user only typed about {20 * PER_USER:,} tokens")turn input tokens sent running input total 1 650 650 2 1,150 1,800 5 2,650 8,250 10 5,150 29,000 20 10,150 108,000 turn 20 alone costs 15.6x turn 1 cumulative input tokens for a 20-turn chat: 108,000 the user only typed about 3,000 tokens
Three thousand tokens typed. A hundred and eight thousand tokens billed. That factor of thirty-six is the single most common surprise on a first cloud bill, and it is entirely predictable from the code above.
Two levers follow directly. A shorter system prompt is multiplied by every turn, so trimming 200 tokens from it saves 4,000 tokens over twenty turns. And summarising old turns caps history instead of letting it grow.
Prompt caching changes this arithmetic where it is available. Providers that cache a repeated prefix charge substantially less for the cached portion. It rewards putting the stable content — system prompt, few-shot examples, retrieved documents — at the front, and the volatile content at the back. Structure your prompt that way whether or not you have caching today.
The check to put in front of every call
def rough_tokens(text):
return max(1, round(len(text) / 4))
class ContextBudget:
def __init__(self, window, max_output, safety=0.05):
self.window = window
self.max_output = max_output
self.limit = int((window - max_output) * (1 - safety))
def check(self, messages):
used = sum(rough_tokens(t) for _, t in messages)
return used, self.limit, used <= self.limit
budget = ContextBudget(window=8192, max_output=1000)
short = [("system", "You are helpful."), ("user", "What is UPI?")]
long = [("system", "You are helpful."), ("user", "x " * 20000)]
for name, msgs in [("short chat", short), ("pasted a long document", long)]:
used, limit, ok = budget.check(msgs)
print(f"{name:24} {used:>6} / {limit:>5} tokens {'OK' if ok else 'OVER BUDGET'}")short chat 7 / 6832 tokens OK pasted a long document 10004 / 6832 tokens OVER BUDGET
The safety margin exists because your token estimate is approximate and because chat formats add hidden tokens per message — role markers and separators, typically a few tokens each. Budgeting to exactly the limit fails in production on inputs that passed in testing.
What the three failure modes look like in code
Refusal. A direct API call over the limit returns an HTTP 400. The message names the model's maximum, your request size and the overage. Exact wording differs by provider, so match on the status code and the error type field, never on the English text of the message.
Silent truncation. Nothing raises. Your quality metrics drift down over long sessions and no log line explains it. The fix is to log tokens_sent, messages_dropped and oldest_message_kept on every call. If you are not logging what you dropped, you cannot debug this class of bug at all.
Truncated output. The response arrives with a finish reason of length rather than stop. Check that field on every response. Treating a length-truncated reply as a complete answer is how half-written JSON reaches your parser.
Four ways to stay inside the window
Cap the output. Set max_tokens explicitly. Leaving it unset lets a runaway answer consume the entire remaining budget.
Pin what must survive. System prompt, safety rules, the user's stated constraints. Shown above, and it is the minimum.
Summarise the middle. When history crosses a threshold, replace the oldest turns with a short generated summary. You lose detail and you keep continuity. Log the summary — when a user reports that the assistant contradicted itself, the summary is usually where the fact was lost.
Retrieve instead of resending. Store the conversation and the documents outside the window. Fetch only the passages relevant to the current question. This is RAG, and for anything longer than a short chat it beats a bigger window on cost, latency and accuracy at the same time.
Common mistakes
Believing the model remembers. It does not. Every stateful behaviour you see is your application resending text.
Forgetting the output shares the budget. Input plus output must fit. Reserve the output before you fill the input.
Counting words instead of tokens. For English prose, tokens run roughly one and a third times word count. For code, JSON or Indic scripts the ratio is far worse. Use the real tokenizer.
Assuming a bigger window fixes quality. Accuracy on facts placed in the middle of a long context is measurably lower than at the edges. Putting the important material at the start or the end of your prompt is free and it works.
Not logging what was dropped. Truncation is invisible by design. If you do not record it, you will spend a week blaming the model.
Try it yourself
In truncate.py, raise WINDOW to 100 and re-run. Watch the booking id survive. Then lower it to 40 and watch the system message fall off even with pinning on, because the pinned content alone now exceeds the budget. Decide what your code should do at that point — silently proceed, or raise. There is a correct answer for your application and it is worth deciding on purpose.
Then modify fit() to keep the first two messages and the last four, dropping only the middle. Compare it against the naive version. That shape, first-and-last, is what most production chat systems actually do, and now you know why.
What to learn next
- What is RAG? — retrieving the right text instead of carrying all of it.
- Temperature and sampling — the next setting that silently changes your output.
- Tokenization — accurate counting, and why non-English text costs more.
Researcher — Mathematics and papers.
The mechanics of why attention is quadratic, how the KV cache is sized, and what FlashAttention and grouped-query attention do about it are covered in attention and how LLMs actually work. This section covers what sits on top of that: how windows get extended, why the advertised number is not the usable number, and what the cost model actually looks like in serving.
Extending a trained window
A model trained at length L_train does not generalise to L > L_train without help, because position information falls outside the range seen during training. With RoPE, the standard interventions all manipulate the rotation frequencies.
Position interpolation (Chen et al., 2023, arXiv:2306.15595). Rescale position indices to fit the original range:
m' = m * (L_train / L_target)mis the absolute token position.L_trainis the pretraining context length,L_targetthe desired one.
Interpolation keeps positions inside the trained region rather than extrapolating beyond it. A few hundred to a thousand fine-tuning steps recover quality. The cost is resolution: adjacent positions are squeezed closer together, so fine-grained local ordering degrades.
NTK-aware scaling and YaRN (Peng et al., 2023, arXiv:2309.00071). Rather than scaling all frequencies uniformly, scale high-frequency dimensions less and low-frequency dimensions more. High frequencies carry local ordering, which interpolation damages; low frequencies carry long-range position, which is what needs stretching. YaRN adds a temperature adjustment to attention logits and reaches target lengths with roughly ten times fewer fine-tuning tokens than position interpolation.
Adjusted base frequency (Xiong et al., 2023, arXiv:2309.16039). Increase RoPE's base from 10,000 to a much larger value and continue pretraining at the longer length. Structurally the simplest option and the one most large open-weight releases have used.
The common thread: all of these need continued training at the target length. Applying the transform at inference time alone produces a model that accepts long inputs and degrades badly on them. Reported "context extended to N" without stated continued-training tokens should be read sceptically.
Advertised length is not effective length
This is the most important practical result in the area.
RULER (Hsieh et al., 2024, arXiv:2404.06654) constructs synthetic tasks with controllable length and complexity: multi-key and multi-value retrieval, variable tracking through chains of assignments, and frequent-word extraction. It defines effective context length as the longest length at which a model still beats a strong short-context baseline.
The finding, reproduced across many models: effective length is frequently a small fraction of advertised length. Models claiming 32K commonly hold up only to a few thousand tokens on multi-hop variants, while scoring near-perfectly on single-needle retrieval at full length.
That gap explains why needle-in-a-haystack — inserting one sentence into filler and asking for it back — is a weak benchmark. It tests one lookup with no distractors and no aggregation. Passing it is necessary and nowhere near sufficient. Any long-context claim resting only on a green needle heatmap is under-evidenced.
Complementary benchmarks: LongBench (Bai et al., 2023, arXiv:2308.14508) for realistic bilingual tasks, and ∞Bench (Zhang et al., 2024, arXiv:2402.13718) for beyond-100K evaluation.
Layered on top is the position effect from Liu et al. (2023), Lost in the Middle: a U-shaped accuracy curve over the position of the relevant span. Ordering retrieved passages so the strongest sit at the edges is a free intervention that measurably helps.
The serving cost model
Inference splits into two phases with completely different characteristics, and conflating them causes most capacity-planning errors.
| Prefill | Decode | |
|---|---|---|
| Processes | all n input tokens at once | one token at a time |
| Bottleneck | compute-bound | memory-bandwidth-bound |
| Attention cost | O(n^2 d) | O(n d) per token |
| Determines | time to first token | tokens per second |
| Batching | large batches saturate quickly | large batches help a lot |
Practical consequences:
- Time to first token scales superlinearly with input length. A 100K-token prompt has a prefill cost that dominates the entire request, regardless of how short the answer is.
- KV cache size, not parameter count, caps concurrency. A long-context request holds cache proportional to its length for its entire lifetime.
- A single long request degrades everyone. In a naive scheduler, one large prefill blocks the decode steps of every other in-flight request, producing latency spikes unrelated to load.
PagedAttention (Kwon et al., 2023, arXiv:2309.06180) allocates KV cache in fixed-size blocks with a page table rather than contiguously, cutting internal fragmentation from tens of percent to near zero and enabling copy-on-write sharing of a common prefix across requests. This is the mechanism behind prefix caching.
Chunked prefill (Agrawal et al., Sarathi-Serve, arXiv:2403.02310) splits a long prefill into pieces and interleaves them with decode steps from other requests. It converts a latency spike into slightly slower steady-state throughput, which is almost always the trade you want.
Long context against retrieval
The choice is an economic one, and it is worth writing down explicitly.
Loading a document of n tokens into every request costs prefill proportional to n^2 per request and KV cache proportional to n for the request's lifetime. Retrieval costs one embedding forward pass plus an approximate nearest-neighbour lookup — effectively constant — and then a prompt of k chunks where k * chunk_size is far smaller than n.
At n in the hundreds of thousands and any meaningful request volume, retrieval wins on cost by orders of magnitude. It also wins on time to first token.
Long context wins where retrieval fails structurally: tasks needing global aggregation ("how many times does the contract mention indemnity"), tasks where relevance cannot be judged before reading, and one-off analyses where engineering a retrieval pipeline costs more than the compute saved.
Prompt caching sits between the two. Where a large prefix is genuinely stable across requests, caching amortises the prefill and shifts the crossover point substantially in favour of long context.
Key references
- Chen, S. et al. (2023). Extending Context Window of Large Language Models via Positional Interpolation. arXiv:2306.15595
- Peng, B. et al. (2023). YaRN: Efficient Context Window Extension of Large Language Models. arXiv:2309.00071
- Xiong, W. et al. (2023). Effective Long-Context Scaling of Foundation Models. arXiv:2309.16039
- Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172
- Kwon, W. et al. (2023). Efficient Memory Management for LLM Serving with PagedAttention. arXiv:2309.06180
- Bai, Y. et al. (2023). LongBench. arXiv:2308.14508
- Zhang, X. et al. (2024). ∞Bench: Extending Long Context Evaluation Beyond 100K Tokens. arXiv:2402.13718
- Hsieh, C.-P. et al. (2024). RULER: What's the Real Context Size of Your Long-Context Language Models? arXiv:2404.06654
- Agrawal, A. et al. (2024). Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. arXiv:2403.02310
Open problems
No accepted definition of effective context length. RULER offers one operational definition, tied to its own task suite. Different suites give different answers for the same model, and there is no theory predicting where a given model's usable length should fall.
Position generalisation remains empirical. Interpolation, NTK scaling and base adjustment all work, and none has a derivation explaining why one beats another at a given target length. Choices are made by sweeping.
Aggregation, not retrieval, is the unsolved half. Locating one fact in a long context is largely handled. Counting occurrences, comparing distant sections, or reasoning over the whole document degrades sharply with length, and this is the gap that separates "supports 1M tokens" from "useful at 1M tokens".
What to learn next
- What is RAG? — the retrieval side of this trade-off.
- Attention — the quadratic cost and the efficiency literature around it.
- vLLM — PagedAttention and continuous batching in a real serving stack.