Abstractive summarisation
Abstractive summarisation writes brand-new sentences to capture a document's meaning, instead of copying sentences from it word for word.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Abstractive summarisation writes brand-new sentences that capture a document's meaning, instead of copying sentences from it.
Think about telling a friend what a movie was about. You do not quote the script line by line. You use your own words, and you compress two hours into thirty seconds.
That is abstractive summarisation. The model reads a document, then writes a fresh summary in its own words. Nothing in the output has to appear in the source at all.
Why it exists
Extractive summarisation has a real limit. It can only rearrange sentences that already exist. Say the best summary needs to merge two half-sentences from opposite ends of a document. Extractive methods cannot do that.
Human-written summaries rarely quote the source directly. A news editor writing a headline invents new phrasing. Abstractive summarisation tries to match that skill.
The trade-off is trust. A model that generates free text can also generate a fact that was never in the source. That risk gets its own lesson later: faithfulness in summarisation.
How it works
The model is a sequence-to-sequence system. It reads the whole document, builds an internal understanding, then generates a summary one word at a time.
Full article (200 words)
|
v
Encoder reads it, builds a compressed understanding
|
v
Decoder writes new words, one at a time:
"The" -> "central" -> "bank" -> "held" -> "rates" -> "steady" -> ...Each new word is chosen based on the document's meaning and every word written so far. Every modern chatbot uses this same generate-one-word-at-a-time process, aimed here at compression instead of open conversation.
Where you have already seen it
- Gmail's "Summarize this email" feature. It writes new sentences instead of trimming the email.
- News app "quick read" boxes. Many are written by a model, not copy-pasted from the article.
- ChatGPT, when you paste an article and ask "summarise this". Classic abstractive summarisation, on demand.
- YouTube's auto-generated video descriptions. Increasingly written, not extracted from the transcript.
Remember this
- Abstractive summarisation writes new sentences. It does not copy from the source.
- It can compress and merge ideas in ways extractive methods cannot.
- The risk: it can also state something the source never said. Always check faithfulness before trusting one.
What to learn next
- Extractive summarisation — the copy-based alternative, and when to prefer it.
- Faithfulness in summarisation — checking whether a summary actually stayed true to the source.
- Attention — the mechanism the decoder uses to look back at the document while writing.
Developer — Code and libraries.
t5-small is a genuinely tiny sequence-to-sequence model, small enough to run comfortably on a laptop CPU. It was trained to summarise when its input is prefixed with "summarize: ".
Setup
pip install transformers torchThe first run downloads t5-small, about 244 MB. It is cached after that, so later runs are instant.
Summarising a real paragraph
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("t5-small")
model = AutoModelForSeq2SeqLM.from_pretrained("t5-small")
text = """summarize: India's Chandrayaan-3 mission landed near the Moon's south pole in
August 2023. It was the first spacecraft from any country to land in that region. The
lander carried a rover named Pragyan, which studied the lunar soil for two weeks.
Scientists were excited because the south pole may hold frozen water, which could one
day support a human base on the Moon. The mission cost about 75 million dollars, far
less than similar missions from other countries."""
inputs = tokenizer(text, return_tensors="pt", truncation=True)
output_ids = model.generate(**inputs, max_length=40, min_length=10, num_beams=4)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))the spacecraft landed near the moon's south pole in august 2023. it was the first spacecraft from any country to land in the region.
Line by line
"summarize: " prefix. T5 treats every task as text-to-text. The prefix tells one shared model which task you want. "translate English to German: " and "summarize: " are both valid prefixes for that same model.
tokenizer(text, return_tensors="pt") turns the words into numbers the model can process, packaged as a PyTorch tensor. truncation=True prevents a crash if the input runs past the model's limit.
model.generate(...) does the actual writing, one token at a time. num_beams=4 keeps four candidate summaries alive at once, instead of only the single greediest guess. This usually gives noticeably better wording. max_length and min_length cap the output size — the full mechanics are in controlling summary length.
Lowercase output. Notice the output has no capital letters. t5-small was trained on lowercased text. This is a property of this specific model, not of abstractive summarisation in general.
Common mistakes
Calling pipeline("summarization", ...) and getting a KeyError. Recent versions of transformers no longer register a "summarization" pipeline by default. Load the model and tokenizer directly, as above — it also makes the generation step easier to control.
Forgetting the task prefix. T5 without "summarize: " tries to guess what you want, and usually guesses wrong. Encoder-decoder models like BART do not need a prefix; check the model card before assuming.
Judging a tiny model by a hard example. t5-small is genuinely small. On dense, technical text it can lose nuance a larger model like BART or a modern LLM would keep. Use it to learn the mechanics, then size up for production quality.
Not setting num_beams. The default is greedy decoding, which picks the single most likely next word every time. It is faster but noticeably worse. Beam search almost always improves summary quality for a small speed cost.
Try it yourself
Set num_beams=1 and compare the output to the four-beam version above. Greedy decoding commits to each word instantly, with no chance to reconsider. Watch for repeated or awkward phrasing that beam search avoided.
What to learn next
- Controlling summary length — what
max_lengthandmin_lengthactually do to the model. - Beam search, and why chat models dropped it — the full mechanics behind
num_beams. - Encoder, decoder and encoder-decoder models — where T5's architecture sits among the alternatives.
Researcher — Mathematics and papers.
The task as conditional generation
Abstractive summarisation models the conditional distribution over output sequences given a source document:
P(y | x) = product over t = 1..m of P(y_t | y_1, ..., y_{t-1}, x)xis the source document, a sequence of tokens.y = (y_1, ..., y_m)is the generated summary.- Each
y_tis conditioned on the full source and every token generated so far.
This is the identical factorisation used for machine translation and open-ended text generation. Summarisation differs only in training data and, usually, a stronger compression ratio between input and output length.
Architecture: encoder-decoder
The dominant architecture is encoder-decoder, covered structurally in encoder, decoder and encoder-decoder models. An encoder reads x bidirectionally and produces contextual representations. A decoder generates y autoregressively, attending both to its own prior output and, via cross-attention, to every encoder position.
T5 (Raffel et al., 2020) casts every NLP task, including summarisation, as text-to-text. A task prefix in the input string picks the task. BART (Lewis et al., 2020) instead pre-trains as a denoising autoencoder: corrupt a document, then reconstruct it. That objective transfers unusually well to summarisation, since both tasks boil down to "produce clean text conditioned on a source."
PEGASUS (Zhang et al., 2020) goes further. Its pre-training objective masks out whole sentences that resemble a document's own summary. The task is called Gap Sentence Generation, chosen to match summarisation's structure directly.
Training objective and its known failure mode
Standard training minimises token-level cross-entropy against a single reference summary. It uses teacher forcing: the decoder sees the true previous token at every step, never its own prediction.
This creates exposure bias. At training time the decoder never sees its own mistakes. At inference time, an early error can compound, since the model was never trained to recover from one. This is one root cause of hallucination, alongside the data-level causes explored in faithfulness in summarisation.
Decoding
Generation requires search over an exponentially large space of token sequences. Beam search keeps k partial hypotheses alive at each step, instead of only one. The full mechanics are in beam search, and why chat models dropped it. For summarisation specifically, beam search reliably outperforms greedy decoding, unlike open-ended chat generation, where it can produce dull, repetitive text.
Length penalty corrects a real bias. Unnormalised beam search favours shorter sequences, because every extra token multiplies the sequence probability by another value under one. Practical implementations divide log-probability by length^alpha, with alpha around 0.6 to 1.0, to counteract this.
Evaluation
ROUGE-N and ROUGE-L remain the standard automatic metrics, detailed in ROUGE and what it misses. Abstractive summaries score more variably on ROUGE than extractive ones. ROUGE rewards n-gram overlap, and a faithful paraphrase can share very few n-grams with its reference.
BERTScore compares meaning rather than exact wording. It correlates better with human judgement on abstractive output.
Key references
- Rush, A., Chopra, S. & Weston, J. (2015). A Neural Attention Model for Abstractive Sentence Summarization. arXiv:1509.00685 — the paper that started the neural abstractive line.
- See, A., Liu, P. & Manning, C. (2017). Get To The Point: Summarization with Pointer-Generator Networks. arXiv:1704.04368
- Lewis, M. et al. (2020). BART. arXiv:1910.13461
- Raffel, C. et al. (2020). T5: Exploring the Limits of Transfer Learning. arXiv:1910.10683
- Zhang, J. et al. (2020). PEGASUS. arXiv:1912.08777
Current state and open problems
Large general-purpose LLMs, prompted rather than fine-tuned, now match or beat dedicated summarisation models on many benchmarks. That trade-off is the subject of prompt an LLM or fine-tune a small model?.
The open problem has not moved since See et al. (2017): fluent output and faithful output are two different objectives, and optimising fluency alone can quietly erode faithfulness. No training objective in wide use directly penalises an unsupported claim. Faithfulness is still checked after the fact, not trained in from the start.
What to learn next
- Faithfulness in summarisation — measuring the gap between fluent and true.
- T5 and the text-to-text idea — the architecture used in this lesson's code, in full depth.
- Hierarchical summarisation — what to do when the document does not fit in the model at all.