Summarisation

Extractive summarisation

Extractive summarisation builds a summary by picking the most important sentences straight out of the original text, without writing a single new word.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Extractive summarisation builds a summary by picking the most important sentences straight out of the original text.

Picture revising for an exam with a highlighter pen. You do not rewrite the chapter in your own words. You drag the pen over the lines that matter and skip the rest.

Extractive summarisation is a computer doing exactly that, one sentence at a time. Nothing is invented. Every line in the summary was already sitting somewhere in the original document.

Why it exists

Some jobs cannot afford a summary that puts words in someone's mouth. A court transcript, a doctor's notes, a contract. If the summary adds something the original never said, that is a real problem.

Extractive summarisation solves this by never generating new text at all. It only selects and reorders. Whatever appears in the summary is a verbatim quote. It can always be traced back to exactly where it came from.

It is also far cheaper to build. No writing model is needed, only a way to rank sentences by importance. That made it the default choice for decades, long before today's language models existed.

How it works

The core idea is scoring, not writing.

Document (8 sentences)
     |
     v
Score every sentence by how closely it echoes the rest
     |
     v
Keep only the highest-scoring sentences, in their original order
     |
     v
Summary: sentence 2, then sentence 5 -- copied exactly, nothing rewritten

A sentence scores high when it shares a lot of meaning with many other sentences in the document. That sounds strange at first. The "most important" sentence is the one that overlaps most with the rest. Think of the person at a party whose opinions echo what most guests already believe. A sentence that mentions the document's main topic in several different ways tends to win. A one-off aside about something nobody else mentions tends to lose.

Where you have already seen it

  • News apps that show a 60-word version of a story. Several popular Indian news apps work this way. They pull the sentences that carry the most weight, then show them as bullet points.
  • Google's featured snippet. The highlighted paragraph above search results is often lifted directly from a webpage, not rewritten.
  • "Jump to relevant part" tools in legal and research software. They highlight sentences worth reading rather than rewriting the document.
  • Email thread previews. Some inboxes show you the one sentence from a long thread most likely to matter, pulled verbatim.

Remember this

  • Extractive summarisation copies sentences from the source. It never writes new ones.
  • A sentence scores high when it overlaps in meaning with many other sentences in the document.
  • It is cheap, fast and always traceable back to the source. The trade-off: it can read a little choppy, because the sentences were never written to sit next to each other.

What to learn next

  • Abstractive summarisation — the alternative: writing brand-new sentences instead of picking old ones.
  • Embeddings — a more modern way to measure how similar two sentences are.
  • What is NLP? — the bigger picture this technique sits inside.

Developer — Code and libraries.

TextRank is the classic extractive algorithm, and it needs no training at all. It only needs the document itself. Below, sumy builds the sentence graph and does the ranking for you.

Setup

bash
pip install sumy nltk

sumy needs one small NLTK data file to split text into sentences. The code below downloads it automatically the first time, quietly.

Extracting the two most important sentences

extractive.py
import nltk
nltk.download("punkt_tab", quiet=True)

from sumy.parsers.plaintext import PlaintextParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.text_rank import TextRankSummarizer

text = """
India's Chandrayaan-3 mission landed near the Moon's south pole in August 2023.
It was the first spacecraft from any country to land in that region.
The lander carried a rover named Pragyan, which studied the lunar soil.
Scientists were excited because the south pole may hold frozen water.
Frozen water could one day support a human base on the Moon.
The mission cost about 75 million dollars, far less than similar missions.
Its low cost made engineers around the world take notice.
ISRO plans more Moon and Mars missions in the coming years.
"""

parser = PlaintextParser.from_string(text, Tokenizer("english"))
summary = TextRankSummarizer()(parser.document, sentences_count=2)

for sentence in summary:
    print(sentence)
Output
India's Chandrayaan-3 mission landed near the Moon's south pole in August 2023.
The lander carried a rover named Pragyan, which studied the lunar soil.

Line by line

Tokenizer("english") splits the raw text into sentences and words. This step matters more than it looks. Get sentence splitting wrong, and every later score sits on the wrong units.

TextRankSummarizer() builds a graph where every sentence is a node. An edge between two sentences is weighted by how much their wording overlaps. The algorithm then runs a PageRank-style pass over this graph. PageRank is the same family of algorithm that once ranked web pages by counting links from other important pages. Here, a sentence "votes" for every other sentence it resembles.

sentences_count=2 is the only knob you pulled. There is no model to train, no labelled data to gather. That is the whole appeal of extractive summarisation. It works on any document, in any domain, from day one.

Notice the two chosen sentences sat 15 words apart in the original document. The algorithm still found both, using nothing but word overlap.

Common mistakes

Treating a high score as "most true" rather than "most central". TextRank rewards sentences that overlap with many others, not sentences that are objectively the most important fact. A one-off number buried in paragraph six can matter enormously. It can still score low, because nothing else in the document repeats it.

Forgetting that extractive summaries can read choppily. The chosen sentences were never written to sit next to each other. Pulling two sentences from far apart in a document sometimes loses the connecting logic between them.

Picking sentences_count without checking document length. Two sentences from an eight-sentence document is a fair compression. Two sentences from an eighty-sentence report throws away far more.

Using the default English pipeline on Hindi, Tamil or mixed-language text. Tokenizer("english") assumes English sentence-boundary rules. Feed it Hindi text and sentence splitting quietly breaks, which corrupts every score built on top of it.

Try it yourself

Change sentences_count=2 to sentences_count=4 and compare. More sentences make the summary safer but less compressed. There is no free lunch here, only a trade-off you choose.

Then try swapping TextRankSummarizer for LsaSummarizer (same import path pattern: from sumy.summarizers.lsa import LsaSummarizer). LSA scores sentences using the hidden topics in the document instead of a similarity graph. On short documents the two often agree; on longer, multi-topic documents they can disagree in interesting ways.

What to learn next

  • Embeddings — a way to measure sentence similarity that understands meaning, not only shared words.
  • TF-IDF — the weighting scheme underneath LSA-style extractive summarisers.
  • Abstractive summarisation — what to reach for once copying sentences is not enough.

Researcher — Mathematics and papers.

The task, formally

Given a document D split into sentences s_1, ..., s_n, extractive summarisation selects a subset S of {s_1, ..., s_n}. A length budget k = |S| bounds the selection. The goal: maximise coverage of D's content while minimising redundancy inside S.

Nothing here requires a generative model. The output space is a subset of the input. That is what makes the method both cheap and inherently faithful. Every claim in S is a claim already present in D.

Luhn's method, the origin point

Luhn (1958) is the first published automatic summariser. It scores a sentence by the density of "significant" words it contains. Significance here comes from raw corpus frequency, after removing stopwords. A sentence packed with frequent content words, close together, scores highest. No linear algebra, no graph — only a frequency count and a windowing rule.

Graph-based ranking: TextRank and LexRank

TextRank (Mihalcea & Tarau, 2004) and LexRank (Erkan & Radev, 2004) independently apply the same idea. Build an undirected graph G = (V, E), where V is the set of sentences. Edge weight w(i, j) is a similarity between sentence i and sentence j. TextRank uses normalised word overlap; LexRank uses cosine similarity between TF-IDF sentence vectors.

Sentence importance is then the stationary distribution of a random walk over G. This is the same fixed point PageRank (Page et al., 1998) computes for the web graph:

text
score(s_i) = (1 - d) / n + d * sum over s_j in adj(s_i) of [ w(i,j) / sum over s_k in adj(s_j) of w(j,k) ] * score(s_j)
  • d is a damping factor, conventionally 0.85, matching the original PageRank paper.
  • adj(s_i) is the set of sentences connected to s_i, in practice all other sentences above a similarity threshold.
  • The system is solved by iterating to convergence, or by eigenvector decomposition of the transition matrix directly.

Complexity is O(n^2) to build the similarity matrix, then O(n^2) per power-iteration step. For n in the hundreds, this cost is trivial. For n in the tens of thousands — a full book — the pairwise matrix becomes the bottleneck. That is the practical reason extractive methods are usually applied per chunk, rather than over one entire long document.

Redundancy: Maximal Marginal Relevance

A pure top-k selection by score often picks several near-duplicate sentences. Sentences that overlap with the whole document also tend to overlap with each other. Carbonell & Goldstein (1998) address this directly with MMR:

text
MMR = argmax over s_i in D \ S of [ lambda * sim(s_i, query) - (1 - lambda) * max over s_j in S of sim(s_i, s_j) ]
  • lambda in [0, 1] trades relevance against redundancy; lambda = 1 reduces to plain top-k.
  • sim(s_i, query) measures relevance to the document, or to an explicit query. This is the mechanism query-focused summarisation builds on.
  • The second term penalises a candidate for resembling anything already chosen, greedily, one sentence at a time.

The supervised, neural version

Modern extractive summarisation reframes the task as binary sentence classification. Label every sentence in a document as "keep" or "drop". Then train a classifier on those labels. The labels are rarely human-annotated at the sentence level. Instead, a greedy oracle constructs them. It picks the subset of source sentences that maximises ROUGE against a human-written abstractive summary (Nallapati et al., 2017).

BERTSUM (Liu, 2019) fine-tunes a BERT encoder with a [CLS] token inserted before every sentence. It then trains a lightweight classifier on each [CLS] vector. This gives contextual sentence representations that account for the whole document, not only local word overlap. The cost: you need training data and a GPU.

Evaluation

Extractive and abstractive summaries are scored the same way. ROUGE-N and ROUGE-L compare them against reference summaries. The full method is in ROUGE and what it misses. Extractive methods have a structural advantage on ROUGE. ROUGE rewards n-gram overlap, and extractive output is n-grams lifted verbatim. The reference summary itself was often written from that same document.

Key references

  • Luhn, H. P. (1958). The Automatic Creation of Literature Abstracts. IBM Journal of Research and Development.
  • Page, L., Brin, S., Motwani, R. & Winograd, T. (1998). The PageRank Citation Ranking. Stanford InfoLab.
  • Carbonell, J. & Goldstein, J. (1998). The Use of MMR, Diversity-Based Reranking. SIGIR.
  • Mihalcea, R. & Tarau, P. (2004). TextRank: Bringing Order into Texts. EMNLP.
  • Erkan, G. & Radev, D. (2004). LexRank: Graph-based Lexical Centrality. Journal of Artificial Intelligence Research.
  • Nallapati, R., Zhai, F. & Zhou, B. (2017). SummaRuNNer. AAAI.
  • Liu, Y. (2019). Fine-tune BERT for Extractive Summarization. arXiv:1903.10318

Current state and open problems

Extractive summarisation has not disappeared in the LLM era — it has moved earlier in the pipeline. Retrieve-then-generate systems commonly run a cheap extractive pass first. It cuts a long document down to a size a generative model can read faithfully. Only then does an abstractive model write the final summary. Hierarchical summarisation and chunking strategies both lean on this pattern.

The open problem is the one Luhn faced in 1958. Word or sentence overlap is a proxy for importance, not importance itself. A single sentence can state a number that contradicts the rest of the document. It can be the most important sentence present, while scoring near zero on every centrality measure above.

What to learn next