Summarising a document longer than the context window
When a document is too long for a model to read in one pass, summarise it in chapters first, then summarise the chapter summaries.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
When a document is too long to read in one pass, summarise it in pieces first, then summarise those summaries.
Picture writing exam notes for a 400-page textbook. You do not read the whole book in one sitting and write one set of notes. You take notes chapter by chapter, then write a shorter summary of your notes.
That two-step process is exactly how hierarchical summarisation works. A model reads and summarises small pieces first, then summarises the summaries.
Why it exists
Every model has a context window. This is a hard limit on how much text it can read at once. It is measured in tokens — small chunks of text. Feed it more, and it either refuses, or silently cuts off the rest.
A single news article fits easily. A 200-page report, a full book, or a year of meeting notes does not. Something has to give.
The fix is not a bigger model. It is a different strategy. Break the document into pieces the model can read. Summarise each piece, then summarise the summaries. This pattern is often called map-reduce, borrowed from data processing. Map a function over many pieces, then reduce the results into one.
How it works
Chapter 1 -> summary 1 \
Chapter 2 -> summary 2 }-> combine -> summary of summaries -> final summary
Chapter 3 -> summary 3 /Each chapter summary is short enough that all of them together still fit in one pass. That combined text becomes the input to one final summarisation step.
For a truly huge document, you can repeat this more than once. Summaries of summaries of summaries, narrowing down with each round.
Where you have already seen it
- Audiobook and podcast apps that summarise a multi-hour episode. No model reads three hours of audio transcript in one shot.
- Legal document review tools. A 500-page contract gets summarised section by section, then combined.
- "Summarise this whole channel" features in chat apps. Weeks of messages, chunked and combined.
- Research tools that summarise an entire paper, including all its sections. Each section is handled on its own first.
Remember this
- Every model has a context window: a hard limit on how much text it can read at once.
- Hierarchical summarisation solves this by summarising small pieces first, then summarising the summaries.
- This is the same "map-reduce" idea used across computing, not something invented specifically for text.
What to learn next
- Context windows — exactly what the limit is, and why it exists.
- Multi-document summarisation — a close cousin: many separate documents instead of one long one.
- Chunking strategies compared — better ways to split a document than by raw length.
Developer — Code and libraries.
Below, a fictional company report has three long "chapters". Each one alone would already push against a small model's limits if there were several more of them. We summarise each chapter, then summarise the combined chapter summaries.
Setup
pip install transformers torchTwo-step summarisation
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("t5-small")
model = AutoModelForSeq2SeqLM.from_pretrained("t5-small")
def summarize(text, max_length=40):
inputs = tokenizer("summarize: " + text, return_tensors="pt",
truncation=True, max_length=512)
ids = model.generate(**inputs, max_length=max_length, min_length=8, num_beams=4)
return tokenizer.decode(ids[0], skip_special_tokens=True)
chapters = [
"Chapter 1. The company opened three new factories in Pune, Chennai and Nagpur this "
"year. Together they created around four thousand jobs. Production capacity rose by "
"sixty percent compared to last year, and the firm expects demand to keep growing.",
"Chapter 2. Revenue grew twenty two percent, driven mostly by exports to Southeast "
"Asia. Costs also rose, largely due to higher shipping prices. Profit margin narrowed "
"slightly even though total profit increased in absolute terms.",
"Chapter 3. The company plans to open two more factories next year, one in Vietnam "
"and one in Bangladesh. It is also investing in solar power for its existing plants "
"to cut long-term energy costs.",
]
# Step 1: summarise each chapter on its own.
chapter_summaries = [summarize(c) for c in chapters]
for i, s in enumerate(chapter_summaries, 1):
print(f"Chapter {i} summary: {s}")
# Step 2: summarise the summaries, now that they fit in one pass.
combined = " ".join(chapter_summaries)
final_summary = summarize(combined, max_length=45)
print("\nFinal summary:", final_summary)Chapter 1 summary: the company opened three new factories in Pune, Chennai and Nagpur this year. production capacity rose by sixty percent compared to last year. Chapter 2 summary: revenue grew twenty two percent, driven mostly by exports to Southeast Asia. costs also rose, partly due to higher shipping prices. Chapter 3 summary: the company plans to open two more factories next year, one in Vietnam and one in Bangladesh. it is also investing in solar power for its existing plants to cut long-term energy costs Final summary: production capacity rose by sixty percent compared to last year. revenue grew twenty two percent, driven mostly by exports to Southeast Asia. costs also rose, partly due to higher shipping prices.
Line by line
Why summarise chapters separately at all. Each chapter is summarised on its own, with its own model.generate() call. This keeps every single call well inside the model's limit, no matter how many chapters exist.
combined = " ".join(chapter_summaries) is the "reduce" step. Three short summaries joined together are far shorter than the three original chapters. That combined text is what finally gets summarised.
Look closely at the wording change. The source said "largely due to higher shipping prices". Chapter 2's summary says "partly due to higher shipping prices". A single word quietly shifted meaning — a small, real example of the risk covered in faithfulness in summarisation.
The final summary lost Chapter 3 entirely. With only 45 tokens to work with, the model dropped the expansion plans altogether. This is the real cost of hierarchical summarisation. Information gets compressed twice. The second pass has no way to know what got trimmed at the first.
Common mistakes
Chunking by raw character count, ignoring sentence boundaries. Cutting a chapter mid-sentence hands the model a broken fragment to summarise, and broken fragments summarise badly.
Assuming the final summary is unbiased. Whatever survives the first round of summarising is all the second round can ever see. An important detail cut in round one is gone for good, with no way to recover it later.
Not tracking word budget per chapter. Say you have ten chapters instead of three, and the combined summaries still exceed the model's limit. You need another round of the same process, not a bigger max_length on the final call.
Try it yourself
Give Chapter 3 a bigger role: raise its max_length to 60 while keeping Chapters 1 and 2 at 40. A longer, more detailed Chapter 3 summary is more likely to survive being compressed a second time.
What to learn next
- Context windows — the exact limit this whole lesson works around.
- Long context or retrieval? — when to reach for a bigger model instead of chunking.
- Faithfulness in summarisation — catching the kind of drift seen in Chapter 2 above.
Researcher — Mathematics and papers.
The problem in one line
A transformer's self-attention cost scales as O(n^2) in sequence length n. Its trained context window is fixed at pre-training time. A document of length N > window cannot be summarised in a single forward pass. This holds regardless of available compute at inference time.
Map-reduce summarisation, formally
Partition the document into chunks c_1, ..., c_k, each satisfying len(c_i) <= window. Summarise each independently:
s_i = f(c_i), for i = 1..kfis the base summariser, a single-pass abstractive or extractive model.
If sum len(s_i) <= window, combine directly:
S = f(concat(s_1, ..., s_k))If not, recurse: treat s_1, ..., s_k as a new set of chunks and repeat. This gives a summarisation tree of depth O(log_b(N / window)), where b is the compression ratio achieved per call.
Where this loses information
Each application of f is lossy and irreversible. A fact present in c_i but omitted from s_i has no path back into S. This compounds across tree depth. Total information retained falls roughly as r^d, for per-step retention rate r and tree depth d.
The Chapter 2 example in the developer block illustrates a second, distinct failure. "Largely" silently became "partly" — not omission, but drift. Nothing enforces that f(c_i) is a strict logical subset of c_i's claims. This is the same faithfulness gap covered in depth in faithfulness in summarisation, compounding once per level of the tree.
Better splits than raw length
Naive fixed-length chunking cuts mid-sentence and mid-argument. Structure-aware chunking, covered in chunking by document structure, splits along document structure instead: headings, paragraphs, speaker turns. This keeps each chunk semantically self-contained, which materially improves per-chunk summary quality.
Alternatives to the map-reduce pattern
Retrieve, then summarise only the relevant parts. Say the final use case is a question, not a full overview. Query-focused summarisation retrieves and summarises only the passages relevant to that question. It sidesteps the need to compress the entire document uniformly.
Long-context models. Architectures such as Longformer (Beltagy et al., 2020) use sparse attention patterns. They reach far longer effective windows at sub-quadratic cost, detailed in Longformer and long-context encoders. Modern LLMs now offer windows in the hundreds of thousands of tokens. That shrinks the class of documents that need chunking at all. But a long context does not guarantee even use of it — see Lost in the Middle.
Incremental summarisation. Rather than a balanced tree, maintain one running summary. Update it as each new chunk arrives: s_i = f(s_{i-1}, c_i). This suits streaming inputs, such as a live meeting transcript. The full document does not exist until the very end.
Key references
- Beltagy, I., Peters, M. & Cohan, A. (2020). Longformer. arXiv:2004.05150
- Liu, P. et al. (2018). Generating Wikipedia by Summarizing Long Sequences. arXiv:1801.10198 — an early large-scale hierarchical approach.
- Liu, N. et al. (2023). Lost in the Middle. arXiv:2307.03172
Current state and open problems
Hierarchical summarisation has not disappeared as context windows have grown; it has moved to a different scale. A million-token context window still cannot fit an entire archive of company documents, only a very long single document.
The open problem is evaluation. ROUGE, scored against one reference summary, does not measure what a hierarchical pipeline loses at each level. No widely agreed metric answers "how much of the original document survived to the final summary" — only proxies.
What to learn next
- Chunking strategies compared — the full menu of ways to split a document before summarising it.
- Long context or retrieval? — deciding between a bigger model and a smarter pipeline.
- Query-focused summarisation — summarising only what is relevant, instead of everything.