Summarisation

Summarising many documents at once

Multi-document summarisation combines several documents about the same topic into one summary, while dropping the parts they all repeat.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Multi-document summarisation combines several documents about the same event into one summary, without repeating what they all agree on.

Picture an earthquake. Three different newspapers cover it. Each one repeats the basic facts, but each also adds one detail the others missed.

You do not want three separate summaries that all say "a 6.1 earthquake struck". You want one summary: the shared facts once, plus every extra detail each source added. That is multi-document summarisation.

Why it exists

A single-document summariser reads one thing at a time. Point it at three related articles. It either summarises them separately, or gets confused about which fact came from where.

News aggregation, research reviews and customer feedback all share a problem. Many sources say almost the same thing. A useful summary must notice the overlap and not repeat it three times.

The extra challenge here is redundancy: the same fact showing up in almost identical words across documents. Picking the "most important" sentences alone, without checking for repeats, produces a summary that repeats itself.

How it works

Article A: "6.1 earthquake struck Tuesday. Minor damage. No deaths."
Article B: "Tuesday's 6.1 quake felt in three districts. Few injuries."
Article C: "6.1 quake near coast. Power cut to two towns."
                |
                v
   Combine all sentences from all articles
                |
                v
   Score each by how well it represents the whole group
                |
                v
   Pick the best sentences, skipping ones too similar to an already-picked one
                |
                v
Summary: earthquake happened + injuries + power cut (no repeats)

Where you have already seen it

  • Google News "full coverage" clusters. Multiple outlets on one story, condensed to the key facts without triplicate repetition.
  • Review summary features. "Most reviewers praised the battery life" pulls the same sentiment from hundreds of separate reviews into one line.
  • Research literature review tools. Combining findings from many papers on the same question, without stating the shared background fact ten times.
  • Meeting notes rolled up from several team channels. The same announcement made in three group chats becomes one line in the summary.

Remember this

  • Multi-document summarisation combines several related documents into one summary.
  • The hard part is not finding important sentences. It is avoiding repeats.
  • A good summary states each shared fact once, then adds what makes each source different.

What to learn next

Developer — Code and libraries.

The strategy: embed every sentence from every article. Find the "average meaning" of the whole set. Pick sentences close to that average, while skipping any sentence too similar to one already picked.

Setup

bash
pip install transformers torch

Combining three articles about the same event

multidoc.py
import torch
from transformers import AutoTokenizer, AutoModel

tokenizer = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")

def embed(texts):
    tokens = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
    with torch.no_grad():
        output = model(**tokens)
    mask = tokens["attention_mask"].unsqueeze(-1)
    summed = (output.last_hidden_state * mask).sum(1)
    counts = mask.sum(1).clamp(min=1e-9)
    return summed / counts

# Three short "articles" about the same event, from three different outlets.
articles = [
    "A magnitude 6.1 earthquake struck the coastal district early Tuesday. "
    "Officials reported minor damage to older buildings. No deaths were confirmed.",

    "Tuesday's earthquake, measured at 6.1, was felt across three districts. "
    "Local hospitals said they treated a handful of minor injuries overnight.",

    "The 6.1 quake near the coast on Tuesday briefly cut power to two towns. "
    "Engineers are inspecting a coastal bridge for cracks before reopening it.",
]

# Split into sentences.
sentences = []
for article in articles:
    for s in article.split(". "):
        s = s.strip().rstrip(".")
        if s:
            sentences.append(s + ".")

vecs = embed(sentences)
centroid = vecs.mean(dim=0, keepdim=True)   # the "average meaning" of all sentences

scores = torch.nn.functional.cosine_similarity(vecs, centroid)
ranked = torch.argsort(scores, descending=True)

# Take the top sentences, but skip near-duplicates already picked.
chosen = []
for idx in ranked.tolist():
    candidate = vecs[idx]
    is_new = all(
        torch.nn.functional.cosine_similarity(candidate.unsqueeze(0), vecs[j].unsqueeze(0)).item() < 0.9
        for j in chosen
    )
    if is_new:
        chosen.append(idx)
    if len(chosen) == 3:
        break

print("Multi-document summary:")
for idx in chosen:
    print(f"  ({scores[idx]:.3f}) {sentences[idx]}")
Output
Multi-document summary:
  (0.638) A magnitude 6.1 earthquake struck the coastal district early Tuesday.
  (0.629) Officials reported minor damage to older buildings.
  (0.563) Tuesday's earthquake, measured at 6.1, was felt across three districts.

Line by line

centroid = vecs.mean(dim=0) finds the average vector across every sentence from every article. A sentence close to this average captures the theme most sources agree on. This is the redundancy signal doing its job.

The redundancy check. Before accepting a candidate sentence, the code checks its similarity against every sentence already chosen. A similarity above 0.9 means "basically the same fact, reworded". It gets skipped, even if it scored well on its own.

Only two of the three articles made the final cut. Article C's power-cut sentence never appeared. Its wording sat further from the group's overall centroid, so it lost out to two more "central" facts. This is a real limitation, worth naming honestly. Distinctive information from one source can be crowded out by facts several sources agree on.

Common mistakes

Setting the similarity threshold too low. A threshold like 0.5 would reject almost anything, since even unrelated sentences share some similarity by chance. Start around 0.85 to 0.9 and check the results by eye.

Treating "close to centroid" as "most important". The centroid captures what most sources agree on. A rare but critical detail, mentioned in only one source, sits far from it by definition. It can be missed entirely.

Not weighting by source count. If five near-identical articles get mixed with one different one, the centroid tilts hard toward the five. Deduplicating near-identical articles before embedding often helps more than tuning the threshold.

Try it yourself

Lower the redundancy threshold from 0.9 to 0.75 and rerun. More sentences will get rejected as duplicates, likely producing a shorter, more diverse-but-riskier summary.

What to learn next

Researcher — Mathematics and papers.

The task, formally

Given document set D = {d_1, ..., d_k}, split into sentences, multi-document summarisation selects S from their union. The goal: maximise coverage of that union's content, under a length budget, while minimising pairwise redundancy within S.

This is the MMR formulation from extractive summarisation, applied over a sentence pool drawn from several documents instead of one. The redundancy term matters far more here. Single-document summarisation rarely repeats a fact verbatim. Multi-document input frequently does, by construction.

Centroid-based selection

The developer block implements a simplified version of the centroid method (Radev et al., 2004). The centroid is the mean embedding of every sentence across the document set. It approximates the topic vector: the set's shared, central meaning.

text
centroid = (1/n) * sum over i = 1..n of embed(s_i)
score(s_i) = cos_sim(embed(s_i), centroid)

Selection then proceeds greedily, in descending score order. It rejects any candidate whose similarity to an already-chosen sentence exceeds a threshold tau. This is a redundancy-only simplification of MMR. It skips the tunable relevance-versus-redundancy trade-off, lambda, in favour of a hard cutoff.

Cross-document phenomena that single-document methods never face

Corroboration. A fact repeated across independent sources is more likely to be true. Multi-document summarisation can exploit this directly. It can weight a sentence's score by how many sources support a similar claim. No single-document method can measure that at all.

Contradiction. Sources can disagree on a number, a name, or an outcome. Centroid-based selection does not detect this. Two contradictory sentences about the same entity can both score high on centrality. They stay topically similar even while factually opposed. Detecting and surfacing this is its own problem, covered in handling conflicting sources.

Timeline drift. In an ongoing story, later articles supersede earlier ones — an initial death toll gets revised, for instance. A centroid computed over articles from different times can average together facts that were never simultaneously true.

Neural multi-document summarisation

Modern systems, such as PRIMERA (Xiao et al., 2022), pre-train specifically for the multi-document setting. Its pre-training objective masks entire sentences salient across a cluster of related documents, not only within one. The model learns cross-document salience directly, instead of relying on centroid heuristics at inference time.

An older, widely-used dataset, Multi-News (Fabbri et al., 2019), pairs clusters of news articles with a human-written summary. It remains a standard benchmark.

Evaluation

Standard ROUGE applies, computed against a human-written multi-document reference summary. A metric specific to this setting, though less commonly reported, checks source attribution. Can each claim in the summary be traced to an input document, and is that document correctly identified? This connects directly to faithfulness in summarisation. Here, faithfulness means "true to at least one source", not "true to a single fixed document".

Key references

  • Radev, D. et al. (2004). Centroid-based Summarization of Multiple Documents. Information Processing & Management.
  • Fabbri, A. et al. (2019). Multi-News. arXiv:1906.01749
  • Xiao, W. et al. (2022). PRIMERA. arXiv:2110.08499

Current state and open problems

LLMs with long context windows can now take many full articles as input directly, sidestepping explicit sentence-level selection. Redundancy and contradiction detection do not disappear under this approach. They move from an explicit algorithmic step to an implicit one, buried inside generation. That makes them harder to inspect or debug.

Contradiction handling remains largely unsolved. No widely-adopted production system reliably flags "sources disagree here" inside a generated summary. Most silently pick one version, or blend both into something neither source actually said.

What to learn next