Summarisation

Summarising a document longer than the context window

When a document is too long for a model to read in one pass, summarise it in chapters first, then summarise the chapter summaries.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

When a document is too long to read in one pass, summarise it in pieces first, then summarise those summaries.

Picture writing exam notes for a 400-page textbook. You do not read the whole book in one sitting and write one set of notes. You take notes chapter by chapter, then write a shorter summary of your notes.

That two-step process is exactly how hierarchical summarisation works. A model reads and summarises small pieces first, then summarises the summaries.

Why it exists

Every model has a context window. This is a hard limit on how much text it can read at once. It is measured in tokens — small chunks of text. Feed it more, and it either refuses, or silently cuts off the rest.

A single news article fits easily. A 200-page report, a full book, or a year of meeting notes does not. Something has to give.

The fix is not a bigger model. It is a different strategy. Break the document into pieces the model can read. Summarise each piece, then summarise the summaries. This pattern is often called map-reduce, borrowed from data processing. Map a function over many pieces, then reduce the results into one.

How it works

Chapter 1  ->  summary 1   \
Chapter 2  ->  summary 2    }->  combine  ->  summary of summaries  ->  final summary
Chapter 3  ->  summary 3   /

Each chapter summary is short enough that all of them together still fit in one pass. That combined text becomes the input to one final summarisation step.

For a truly huge document, you can repeat this more than once. Summaries of summaries of summaries, narrowing down with each round.

Where you have already seen it

  • Audiobook and podcast apps that summarise a multi-hour episode. No model reads three hours of audio transcript in one shot.
  • Legal document review tools. A 500-page contract gets summarised section by section, then combined.
  • "Summarise this whole channel" features in chat apps. Weeks of messages, chunked and combined.
  • Research tools that summarise an entire paper, including all its sections. Each section is handled on its own first.

Remember this

  • Every model has a context window: a hard limit on how much text it can read at once.
  • Hierarchical summarisation solves this by summarising small pieces first, then summarising the summaries.
  • This is the same "map-reduce" idea used across computing, not something invented specifically for text.

What to learn next

Developer — Code and libraries.

Below, a fictional company report has three long "chapters". Each one alone would already push against a small model's limits if there were several more of them. We summarise each chapter, then summarise the combined chapter summaries.

Setup

bash
pip install transformers torch

Two-step summarisation

hierarchical.py
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tokenizer = AutoTokenizer.from_pretrained("t5-small")
model = AutoModelForSeq2SeqLM.from_pretrained("t5-small")

def summarize(text, max_length=40):
    inputs = tokenizer("summarize: " + text, return_tensors="pt",
                        truncation=True, max_length=512)
    ids = model.generate(**inputs, max_length=max_length, min_length=8, num_beams=4)
    return tokenizer.decode(ids[0], skip_special_tokens=True)

chapters = [
    "Chapter 1. The company opened three new factories in Pune, Chennai and Nagpur this "
    "year. Together they created around four thousand jobs. Production capacity rose by "
    "sixty percent compared to last year, and the firm expects demand to keep growing.",

    "Chapter 2. Revenue grew twenty two percent, driven mostly by exports to Southeast "
    "Asia. Costs also rose, largely due to higher shipping prices. Profit margin narrowed "
    "slightly even though total profit increased in absolute terms.",

    "Chapter 3. The company plans to open two more factories next year, one in Vietnam "
    "and one in Bangladesh. It is also investing in solar power for its existing plants "
    "to cut long-term energy costs.",
]

# Step 1: summarise each chapter on its own.
chapter_summaries = [summarize(c) for c in chapters]
for i, s in enumerate(chapter_summaries, 1):
    print(f"Chapter {i} summary: {s}")

# Step 2: summarise the summaries, now that they fit in one pass.
combined = " ".join(chapter_summaries)
final_summary = summarize(combined, max_length=45)
print("\nFinal summary:", final_summary)
Output
Chapter 1 summary: the company opened three new factories in Pune, Chennai and Nagpur this year. production capacity rose by sixty percent compared to last year.
Chapter 2 summary: revenue grew twenty two percent, driven mostly by exports to Southeast Asia. costs also rose, partly due to higher shipping prices.
Chapter 3 summary: the company plans to open two more factories next year, one in Vietnam and one in Bangladesh. it is also investing in solar power for its existing plants to cut long-term energy costs 

Final summary: production capacity rose by sixty percent compared to last year. revenue grew twenty two percent, driven mostly by exports to Southeast Asia. costs also rose, partly due to higher shipping prices.

Line by line

Why summarise chapters separately at all. Each chapter is summarised on its own, with its own model.generate() call. This keeps every single call well inside the model's limit, no matter how many chapters exist.

combined = " ".join(chapter_summaries) is the "reduce" step. Three short summaries joined together are far shorter than the three original chapters. That combined text is what finally gets summarised.

Look closely at the wording change. The source said "largely due to higher shipping prices". Chapter 2's summary says "partly due to higher shipping prices". A single word quietly shifted meaning — a small, real example of the risk covered in faithfulness in summarisation.

The final summary lost Chapter 3 entirely. With only 45 tokens to work with, the model dropped the expansion plans altogether. This is the real cost of hierarchical summarisation. Information gets compressed twice. The second pass has no way to know what got trimmed at the first.

Common mistakes

Chunking by raw character count, ignoring sentence boundaries. Cutting a chapter mid-sentence hands the model a broken fragment to summarise, and broken fragments summarise badly.

Assuming the final summary is unbiased. Whatever survives the first round of summarising is all the second round can ever see. An important detail cut in round one is gone for good, with no way to recover it later.

Not tracking word budget per chapter. Say you have ten chapters instead of three, and the combined summaries still exceed the model's limit. You need another round of the same process, not a bigger max_length on the final call.

Try it yourself

Give Chapter 3 a bigger role: raise its max_length to 60 while keeping Chapters 1 and 2 at 40. A longer, more detailed Chapter 3 summary is more likely to survive being compressed a second time.

What to learn next

Researcher — Mathematics and papers.

The problem in one line

A transformer's self-attention cost scales as O(n^2) in sequence length n. Its trained context window is fixed at pre-training time. A document of length N > window cannot be summarised in a single forward pass. This holds regardless of available compute at inference time.

Map-reduce summarisation, formally

Partition the document into chunks c_1, ..., c_k, each satisfying len(c_i) <= window. Summarise each independently:

text
s_i = f(c_i),   for i = 1..k
  • f is the base summariser, a single-pass abstractive or extractive model.

If sum len(s_i) <= window, combine directly:

text
S = f(concat(s_1, ..., s_k))

If not, recurse: treat s_1, ..., s_k as a new set of chunks and repeat. This gives a summarisation tree of depth O(log_b(N / window)), where b is the compression ratio achieved per call.

Where this loses information

Each application of f is lossy and irreversible. A fact present in c_i but omitted from s_i has no path back into S. This compounds across tree depth. Total information retained falls roughly as r^d, for per-step retention rate r and tree depth d.

The Chapter 2 example in the developer block illustrates a second, distinct failure. "Largely" silently became "partly" — not omission, but drift. Nothing enforces that f(c_i) is a strict logical subset of c_i's claims. This is the same faithfulness gap covered in depth in faithfulness in summarisation, compounding once per level of the tree.

Better splits than raw length

Naive fixed-length chunking cuts mid-sentence and mid-argument. Structure-aware chunking, covered in chunking by document structure, splits along document structure instead: headings, paragraphs, speaker turns. This keeps each chunk semantically self-contained, which materially improves per-chunk summary quality.

Alternatives to the map-reduce pattern

Retrieve, then summarise only the relevant parts. Say the final use case is a question, not a full overview. Query-focused summarisation retrieves and summarises only the passages relevant to that question. It sidesteps the need to compress the entire document uniformly.

Long-context models. Architectures such as Longformer (Beltagy et al., 2020) use sparse attention patterns. They reach far longer effective windows at sub-quadratic cost, detailed in Longformer and long-context encoders. Modern LLMs now offer windows in the hundreds of thousands of tokens. That shrinks the class of documents that need chunking at all. But a long context does not guarantee even use of it — see Lost in the Middle.

Incremental summarisation. Rather than a balanced tree, maintain one running summary. Update it as each new chunk arrives: s_i = f(s_{i-1}, c_i). This suits streaming inputs, such as a live meeting transcript. The full document does not exist until the very end.

Key references

  • Beltagy, I., Peters, M. & Cohan, A. (2020). Longformer. arXiv:2004.05150
  • Liu, P. et al. (2018). Generating Wikipedia by Summarizing Long Sequences. arXiv:1801.10198 — an early large-scale hierarchical approach.
  • Liu, N. et al. (2023). Lost in the Middle. arXiv:2307.03172

Current state and open problems

Hierarchical summarisation has not disappeared as context windows have grown; it has moved to a different scale. A million-token context window still cannot fit an entire archive of company documents, only a very long single document.

The open problem is evaluation. ROUGE, scored against one reference summary, does not measure what a hierarchical pipeline loses at each level. No widely agreed metric answers "how much of the original document survived to the final summary" — only proxies.

What to learn next