AI glossary

Context window

In one sentence The context window is the maximum amount of text, counted in tokens, that a model can hold in view at one time.

By Updated

The context window is the maximum amount of text a model can hold in view at once, measured in tokens.

Picture a small desk. You can spread out as many papers as fit on it, and you can read any of them instantly. Push one more sheet on and something slides off the edge. The model has no memory beyond what is on that desk at the moment you press send — not from yesterday, not from earlier in the conversation unless it is still on the desk.

The important part, and the part that surprises people, is what counts toward the limit. All of this shares the same desk:

[ system prompt ] [ chat history ] [ retrieved documents ] [ tool output ] [ the reply ]
└──────────────────────── all of it counts ────────────────────────────────────────────┘

The reply counts too. If a model has a 128,000-token window and your prompt uses 127,000, there is room for a thousand tokens of answer and nothing more. Budget backwards: window minus the longest answer you want equals the biggest prompt you may send.

For a rough feel, 1,000 tokens is about 750 English words, so a 128,000-token window is roughly 95,000 words — a decent-sized novel. Hindi, Marathi and other non-Latin scripts use noticeably more tokens per word, so the same passage fills the window faster.

Two practical consequences. Long chats grow until they overflow and produce a context length exceeded error, which is why real chat apps trim or summarise old turns. And a large window is not free — attention cost grows sharply with length, so a stuffed window means slower, pricier answers. Retrieving five relevant chunks with RAG usually beats pasting a whole handbook.

Where to go next