Chunking and Long Documents

Long context or retrieval?

A model with a huge context window can read a whole document at once, but that does not always beat retrieval on cost, speed, or even accuracy — this lesson lays out the real trade-off.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A long context window lets a model read an entire document at once, instead of only a retrieved chunk. Reading everything is not always better.

Think about moving house. You could carry every box yourself, in one long, exhausting trip. Or you could send someone ahead to fetch only the boxes you actually need today, and leave the rest in storage. Both get the job done. They cost very different amounts of effort.

Why it exists

This whole section has been about chunking — cutting a document down so only the relevant part gets read. That exists because models used to have small context windows, the maximum amount of text a model can read at once.

Newer models have much larger windows, some large enough to hold an entire book. That raises an honest question: if the model can read the whole thing anyway, why chunk and search at all?

The answer is not simple, and both sides matter. Reading everything is more thorough. It also costs more and runs slower. A large window is not the same thing as reading it well.

How it works

Long context:                     Retrieval:

  [entire 200-page document]  ->    [search] -> [best 3 chunks]
         |                                |
         v                                v
   fed to the model whole          fed to the model, small
   slow, expensive                 fast, cheap
   nothing can be missed           only what search finds
   ...unless buried in the middle  gets read

Where you have already seen it

  • "Summarise this whole PDF" features. These often use a long context window, reading the entire file at once.
  • Customer support bots on a big knowledge base. These almost always use retrieval — reading every article for every question would be too slow and costly.
  • Coding assistants. Reading one open file uses long context. Searching an entire codebase for a relevant function uses retrieval.

Remember this

  • A bigger context window means a model can read more text at once, not that it uses all of it equally well.
  • Long context is simpler and more thorough. Retrieval is cheaper and faster, at large scale.
  • Real systems often use both — retrieval to narrow things down, long context to read what gets found in full.

What to learn next

Developer — Code and libraries.

This counts tokens in a full document versus a retrieved chunk, and estimates the cost difference using an illustrative price.

Setup

bash
pip install transformers

Counting the real cost of "read everything"

long_context_cost.py
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("gpt2")

# A stand-in for a real 40-page policy handbook. Repeating one paragraph
# keeps this example runnable with no file to download.
paragraph = (
    "Employees may carry forward up to 10 unused leave days into the next "
    "financial year, subject to manager approval and departmental workload. "
)
full_handbook = paragraph * 220     # roughly a 40-page handbook
retrieved_chunk = paragraph * 3     # what a retriever would hand back

tokens_full = len(tok.encode(full_handbook))
tokens_chunk = len(tok.encode(retrieved_chunk))

print("full handbook stuffed into context:", tokens_full, "tokens")
print("retrieved chunk only:               ", tokens_chunk, "tokens")
print(f"retrieval uses {tokens_chunk / tokens_full:.1%} of the tokens")

# Illustrative pricing only. Real per-token prices vary by provider and
# change often -- this number is not any specific vendor's real price.
price_per_1k_tokens = 0.003
cost_full = tokens_full / 1000 * price_per_1k_tokens
cost_chunk = tokens_chunk / 1000 * price_per_1k_tokens
print(f"illustrative cost, full handbook:   ${cost_full:.4f}")
print(f"illustrative cost, retrieved chunk: ${cost_chunk:.4f}")
Output
full handbook stuffed into context: 5502 tokens
retrieved chunk only:                77 tokens
retrieval uses 1.4% of the tokens
illustrative cost, full handbook:   $0.0165
illustrative cost, retrieved chunk: $0.0002

Running this prints a warning first — Token indices sequence length is longer than the specified maximum sequence length for this model (5502 > 1024). That is expected: GPT-2's own 1024-token limit is being deliberately exceeded here only to count tokens, not to run the model itself.

Line by line

AutoTokenizer.from_pretrained("gpt2") is used as a stand-in tokenizer. No two models tokenize identically. This gives an honest approximation of token count, not an exact count for whatever model you actually deploy — always measure with your real model's tokenizer for a production estimate.

paragraph * 220 builds a large document from repetition. This keeps the example runnable with no file download, at the cost of being an unrealistically repetitive document. Real handbooks vary far more per paragraph, but the token-count comparison scales the same way regardless.

The 98.6% reduction is the entire argument for retrieval at scale. A single question rarely needs the whole document. Paying to process all of it, every time, adds up fast across many users and many questions.

Common mistakes

Assuming a bigger context window makes retrieval unnecessary. Cost still scales with tokens sent, regardless of window size. A 1-million-token window that gets filled with 1 million tokens on every question is enormously more expensive than retrieving 2,000 relevant tokens.

Ignoring "lost in the middle" effects. Liu et al. (2023) showed models are measurably worse at using information placed in the middle of a long context, compared to the start or end. A larger window does not guarantee even attention across it — see the researcher block below.

Treating this as an either/or decision. Many production systems retrieve first, then feed the retrieved chunks into a still-sizeable context window, combining both techniques rather than picking one exclusively.

Forgetting latency. Processing more input tokens takes more time, not only more money. A long-context call over an entire document is measurably slower to first response than a short, retrieved-context call, which matters for anything interactive.

Try it yourself

Change price_per_1k_tokens to a value ten times higher, and imagine this at 10,000 questions a day instead of one.

At the original price, the daily cost difference between the two approaches is already about $165 versus $2 for 10,000 questions. Multiplying the price by 10 scales both linearly — the ratio between them stays the same, since the cost difference comes entirely from token count, not from price.

What to learn next

  • Context window — the underlying limit this lesson pushes against.
  • Vector databases — the infrastructure retrieval depends on.
  • FAISS — a real tool for searching at the scale where retrieval starts to matter.

Researcher — Mathematics and papers.

The trade-off, formally

Let n be the number of tokens in a document, k be the number of tokens a retriever selects, and c(x) be the (roughly linear, in most current pricing models) cost of processing x input tokens. Long-context processing costs c(n) per query; retrieval-then-generate costs c(k) + r, where r is the retrieval system's own cost — typically small and largely fixed, dominated by an embedding lookup rather than a model call. Retrieval wins on cost whenever k << n, which holds for most real question-answering workloads over large corpora, since a single question rarely requires the entire corpus to answer.

Cost is not the only axis. Xu, Ping, Wu et al. (2023), Retrieval meets Long Context Large Language Models (arXiv:2310.03025), directly compares both approaches empirically and finds that retrieval-augmented generation with a moderate context window can match or exceed pure long-context performance on several benchmarks, while using substantially fewer tokens — and that combining both, retrieving into a still-large context window, performs best of all in their experiments.

Lost in the middle

Liu et al. (2023), Lost in the Middle: How Language Models Use Long Contexts (arXiv:2307.03172), measured retrieval accuracy as a function of where the correct answer sits within a long context. Accuracy is highest when the answer is near the start or end of the input, and measurably lower when it sits in the middle — a U-shaped curve, consistent across the several model families they tested. This means a longer context window does not guarantee uniform attention across its full length. A document's most important fact, placed by chance in the middle of a long-context prompt, can be attended to less reliably than the same fact would be if retrieval had placed it first in a short, focused context.

Why this happens

No settled mechanistic explanation exists for the exact shape of the lost-in-the-middle curve. Proposed contributing factors include training data distributions that under-represent long-range middle-of-document dependencies, and positional encoding schemes whose effective resolution degrades at extreme relative distances. Liu et al. (2023) treat the effect as empirically robust across models and note it persists even in models explicitly trained with long-context objectives, which argues against training data alone being the full explanation.

Practical synthesis

Neither approach dominates unconditionally. Retrieval wins decisively at corpus scale — a customer support knowledge base with thousands of articles cannot fit in any context window, long or short. Long context wins where the whole document is genuinely relevant, and retrieval would risk dropping something needed — a single contract being reviewed in full, for instance. Hybrid architectures, retrieving a generous set of candidate chunks into a moderately large context window rather than either one document or three chunks, are the pattern most production systems converge on as of this writing.

Key references

  • Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172
  • Xu, P., Ping, W., Wu, X. et al. (2023). Retrieval meets Long Context Large Language Models. arXiv:2310.03025
  • Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997

Current state and open problems

Context windows have grown faster than the field's understanding of how reliably models use them. Benchmarks that test only retrieval at the start or end of a context ("needle in a haystack" tests) can overstate real-world long-context capability, since they do not probe the harder, middle-of-document case Liu et al. specifically identify. A rigorous evaluation of any long-context or retrieval system needs to test recall across the full range of positions, not only the easiest ones, before drawing conclusions about which approach to trust for a given task.

What to learn next