Error database

This model's maximum context length is N tokens (context_length_exceeded)

Your prompt plus the space reserved for the reply is larger than the model's context window. Count tokens before sending and trim history or retrieved context.

The message you saw
This model's maximum context length is N tokens (context_length_exceeded)

By Updated

The error

Output
openai.BadRequestError: Error code: 400 - {'error': {'message': "This model's maximum context length is 128000 tokens. However, your messages resulted in 131204 tokens. Please reduce the length of the messages.", 'type': 'invalid_request_error', 'param': 'messages', 'code': 'context_length_exceeded'}}
Output
anthropic.BadRequestError: Error code: 400 - {'type': 'error', 'error': {'type': 'invalid_request_error', 'message': 'prompt is too long: 208531 tokens > 200000 maximum'}}

Running a model locally gives a different wording for the same wall:

Output
ValueError: Input length of input_ids is 4096, but `max_length` is set to 4096. This can lead to unexpected behavior. You should consider increasing `max_length` or, better yet, setting `max_new_tokens`.

What it means

Every model has a context window — a hard limit on how much text it can hold at once, counted in tokens. You sent more than fits.

The number in the message is exact and it is not negotiable. There is no setting that stretches the window; it is a property of how the model was trained.

Why it happens

The window holds more than people expect. Everything below shares the same budget:

[ system prompt ] [ full chat history ] [ retrieved documents ] [ tool results ] [ the reply ]
└──────────────────────────── all counted against one limit ────────────────────────────────┘

Growing chat history is the slow version. Each turn appends the user message and the model's reply, so a conversation that was 2,000 tokens on turn three is 40,000 on turn thirty. The code that worked all week fails on a long conversation.

RAG with too many chunks is the sudden version. Twenty retrieved chunks at 1,000 tokens each is 20,000 tokens before the question is even added, and one unusually long document can double it.

Tool output is the sneaky version. An agent that fetches a web page and puts the raw HTML into the conversation can add 50,000 tokens in a single step, and the failure happens on the step after that.

Two details catch people out. The reply counts too — if the window is 128,000 and your prompt is 127,500, there is no room for a useful answer even though the request is technically under the limit. And token counts vary by language: Hindi, Marathi, Tamil and other non-Latin scripts commonly use two to three times more tokens than English for the same content, so a prompt that fits in one language overflows in another.

How to fix it

1. Count tokens before sending. Guessing by character count is what led here.

python
import tiktoken

enc = tiktoken.encoding_for_model("gpt-4o-mini")

def count(messages) -> int:
    # a few tokens of per-message overhead, close enough for budgeting
    return sum(len(enc.encode(m["content"])) + 4 for m in messages)

print(count(messages))

For Hugging Face models the equivalent is len(tokenizer.encode(text)). For Anthropic models, the API exposes a token-counting endpoint so you do not have to approximate.

2. Budget backwards from the window. Decide how much answer you need, and give the prompt whatever is left.

python
WINDOW = 128_000
RESERVE_FOR_REPLY = 2_000
SAFETY = 500
max_prompt_tokens = WINDOW - RESERVE_FOR_REPLY - SAFETY     # 125,500

Then enforce that number rather than hoping.

3. Trim the conversation history. Keep the system message, keep the most recent turns, drop the middle.

python
def trim(messages, budget, enc):
    system = [m for m in messages if m["role"] == "system"]
    rest   = [m for m in messages if m["role"] != "system"]

    used = sum(len(enc.encode(m["content"])) + 4 for m in system)
    kept = []
    for m in reversed(rest):                  # newest first
        cost = len(enc.encode(m["content"])) + 4
        if used + cost > budget:
            break
        kept.append(m)
        used += cost
    return system + list(reversed(kept))

Cut at whole messages, not mid-message. Half a reply confuses the model more than a missing one.

4. Summarise instead of dropping, when the old turns matter. Ask a cheap model to compress the oldest turns into a short paragraph, keep that as a system note, and discard the originals. A hundred-token summary can carry what two thousand tokens of transcript conveyed.

5. Retrieve less, and re-rank what you retrieve. In a RAG pipeline, more chunks is not better — beyond a handful, extra chunks add noise as well as tokens.

python
chunks = vector_store.similarity_search(question, k=20)      # cast a wide net
chunks = reranker.rank(question, chunks)[:4]                 # keep only what earns its place
context = "\n\n---\n\n".join(c.page_content for c in chunks)

Four well-ranked chunks routinely beat twenty unranked ones on both accuracy and cost.

6. Truncate tool output before it enters the conversation. Anything an agent fetches should pass through a cap.

python
MAX_TOOL_CHARS = 8_000

def add_tool_result(text: str) -> str:
    if len(text) <= MAX_TOOL_CHARS:
        return text
    return text[:MAX_TOOL_CHARS] + f"\n\n[truncated, {len(text) - MAX_TOOL_CHARS} characters omitted]"

Extracting readable text from HTML before storing it usually cuts the size by an order of magnitude on its own.

7. For a long document, work in pieces. Split it, process each piece, then combine the results — the map-reduce pattern. Summarise each chapter, then summarise the summaries. This works on documents of any size and does not depend on finding a bigger model.

8. Set max_new_tokens, not max_length, when running locally. max_length counts prompt plus output together, which is what produces the third error message above. max_new_tokens counts only the generated part and behaves the way people expect.

python
outputs = model.generate(**inputs, max_new_tokens=256)

9. Move to a larger-window model last, not first. It is a genuine option, and it is also the one that hides the underlying problem. A prompt that grows without limit will overflow a 200,000-token window as surely as a 4,000-token one, and it will be slower and more expensive on the way there.

How to prevent it

Treat the context window as a budget you actively manage, not a limit you discover. Count tokens on the way in, subtract what you reserve for the reply, and refuse or trim before you call the API rather than catching a 400 afterwards.

Cap each source separately — history, retrieved context, tool output — so no single one can consume the whole window. That way a very long web page degrades one tool result instead of breaking the request.

Log the token count of every call. The graph tells you immediately whether you are near the edge, and it makes the cost of a prompt change visible before it reaches production. And test with a deliberately long conversation and a deliberately large document, because both are ordinary user behaviour and neither shows up in a happy-path test.

The lessons behind this error.

  • Generative AI

    Context windows

    The context window is the total text a model can hold at once, and exceeding it does not warn you — it either errors, silently drops your oldest messages, or cuts the reply off mid-sentence.

Back to all errors