Summarising many documents at once
Multi-document summarisation combines several documents about the same topic into one summary, while dropping the parts they all repeat.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Multi-document summarisation combines several documents about the same event into one summary, without repeating what they all agree on.
Picture an earthquake. Three different newspapers cover it. Each one repeats the basic facts, but each also adds one detail the others missed.
You do not want three separate summaries that all say "a 6.1 earthquake struck". You want one summary: the shared facts once, plus every extra detail each source added. That is multi-document summarisation.
Why it exists
A single-document summariser reads one thing at a time. Point it at three related articles. It either summarises them separately, or gets confused about which fact came from where.
News aggregation, research reviews and customer feedback all share a problem. Many sources say almost the same thing. A useful summary must notice the overlap and not repeat it three times.
The extra challenge here is redundancy: the same fact showing up in almost identical words across documents. Picking the "most important" sentences alone, without checking for repeats, produces a summary that repeats itself.
How it works
Article A: "6.1 earthquake struck Tuesday. Minor damage. No deaths."
Article B: "Tuesday's 6.1 quake felt in three districts. Few injuries."
Article C: "6.1 quake near coast. Power cut to two towns."
|
v
Combine all sentences from all articles
|
v
Score each by how well it represents the whole group
|
v
Pick the best sentences, skipping ones too similar to an already-picked one
|
v
Summary: earthquake happened + injuries + power cut (no repeats)Where you have already seen it
- Google News "full coverage" clusters. Multiple outlets on one story, condensed to the key facts without triplicate repetition.
- Review summary features. "Most reviewers praised the battery life" pulls the same sentiment from hundreds of separate reviews into one line.
- Research literature review tools. Combining findings from many papers on the same question, without stating the shared background fact ten times.
- Meeting notes rolled up from several team channels. The same announcement made in three group chats becomes one line in the summary.
Remember this
- Multi-document summarisation combines several related documents into one summary.
- The hard part is not finding important sentences. It is avoiding repeats.
- A good summary states each shared fact once, then adds what makes each source different.
What to learn next
- Query-focused summarisation — a close cousin, filtering by question instead of by source count.
- Embeddings — the technique used here to measure how similar two sentences are.
- Clustering documents with embeddings — grouping many documents by topic, a related but different problem.
Developer — Code and libraries.
The strategy: embed every sentence from every article. Find the "average meaning" of the whole set. Pick sentences close to that average, while skipping any sentence too similar to one already picked.
Setup
pip install transformers torchCombining three articles about the same event
import torch
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
def embed(texts):
tokens = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
output = model(**tokens)
mask = tokens["attention_mask"].unsqueeze(-1)
summed = (output.last_hidden_state * mask).sum(1)
counts = mask.sum(1).clamp(min=1e-9)
return summed / counts
# Three short "articles" about the same event, from three different outlets.
articles = [
"A magnitude 6.1 earthquake struck the coastal district early Tuesday. "
"Officials reported minor damage to older buildings. No deaths were confirmed.",
"Tuesday's earthquake, measured at 6.1, was felt across three districts. "
"Local hospitals said they treated a handful of minor injuries overnight.",
"The 6.1 quake near the coast on Tuesday briefly cut power to two towns. "
"Engineers are inspecting a coastal bridge for cracks before reopening it.",
]
# Split into sentences.
sentences = []
for article in articles:
for s in article.split(". "):
s = s.strip().rstrip(".")
if s:
sentences.append(s + ".")
vecs = embed(sentences)
centroid = vecs.mean(dim=0, keepdim=True) # the "average meaning" of all sentences
scores = torch.nn.functional.cosine_similarity(vecs, centroid)
ranked = torch.argsort(scores, descending=True)
# Take the top sentences, but skip near-duplicates already picked.
chosen = []
for idx in ranked.tolist():
candidate = vecs[idx]
is_new = all(
torch.nn.functional.cosine_similarity(candidate.unsqueeze(0), vecs[j].unsqueeze(0)).item() < 0.9
for j in chosen
)
if is_new:
chosen.append(idx)
if len(chosen) == 3:
break
print("Multi-document summary:")
for idx in chosen:
print(f" ({scores[idx]:.3f}) {sentences[idx]}")Multi-document summary: (0.638) A magnitude 6.1 earthquake struck the coastal district early Tuesday. (0.629) Officials reported minor damage to older buildings. (0.563) Tuesday's earthquake, measured at 6.1, was felt across three districts.
Line by line
centroid = vecs.mean(dim=0) finds the average vector across every sentence from every article. A sentence close to this average captures the theme most sources agree on. This is the redundancy signal doing its job.
The redundancy check. Before accepting a candidate sentence, the code checks its similarity against every sentence already chosen. A similarity above 0.9 means "basically the same fact, reworded". It gets skipped, even if it scored well on its own.
Only two of the three articles made the final cut. Article C's power-cut sentence never appeared. Its wording sat further from the group's overall centroid, so it lost out to two more "central" facts. This is a real limitation, worth naming honestly. Distinctive information from one source can be crowded out by facts several sources agree on.
Common mistakes
Setting the similarity threshold too low. A threshold like 0.5 would reject almost anything, since even unrelated sentences share some similarity by chance. Start around 0.85 to 0.9 and check the results by eye.
Treating "close to centroid" as "most important". The centroid captures what most sources agree on. A rare but critical detail, mentioned in only one source, sits far from it by definition. It can be missed entirely.
Not weighting by source count. If five near-identical articles get mixed with one different one, the centroid tilts hard toward the five. Deduplicating near-identical articles before embedding often helps more than tuning the threshold.
Try it yourself
Lower the redundancy threshold from 0.9 to 0.75 and rerun. More sentences will get rejected as duplicates, likely producing a shorter, more diverse-but-riskier summary.
What to learn next
- Document clustering with embeddings — grouping documents by topic before you try to summarise them together.
- Handling conflicting sources — what to do when two documents disagree, rather than only add detail.
- BERTScore and embedding-based metrics — measuring summary quality using the same embedding idea used here.
Researcher — Mathematics and papers.
The task, formally
Given document set D = {d_1, ..., d_k}, split into sentences, multi-document summarisation selects S from their union. The goal: maximise coverage of that union's content, under a length budget, while minimising pairwise redundancy within S.
This is the MMR formulation from extractive summarisation, applied over a sentence pool drawn from several documents instead of one. The redundancy term matters far more here. Single-document summarisation rarely repeats a fact verbatim. Multi-document input frequently does, by construction.
Centroid-based selection
The developer block implements a simplified version of the centroid method (Radev et al., 2004). The centroid is the mean embedding of every sentence across the document set. It approximates the topic vector: the set's shared, central meaning.
centroid = (1/n) * sum over i = 1..n of embed(s_i)
score(s_i) = cos_sim(embed(s_i), centroid)Selection then proceeds greedily, in descending score order. It rejects any candidate whose similarity to an already-chosen sentence exceeds a threshold tau. This is a redundancy-only simplification of MMR. It skips the tunable relevance-versus-redundancy trade-off, lambda, in favour of a hard cutoff.
Cross-document phenomena that single-document methods never face
Corroboration. A fact repeated across independent sources is more likely to be true. Multi-document summarisation can exploit this directly. It can weight a sentence's score by how many sources support a similar claim. No single-document method can measure that at all.
Contradiction. Sources can disagree on a number, a name, or an outcome. Centroid-based selection does not detect this. Two contradictory sentences about the same entity can both score high on centrality. They stay topically similar even while factually opposed. Detecting and surfacing this is its own problem, covered in handling conflicting sources.
Timeline drift. In an ongoing story, later articles supersede earlier ones — an initial death toll gets revised, for instance. A centroid computed over articles from different times can average together facts that were never simultaneously true.
Neural multi-document summarisation
Modern systems, such as PRIMERA (Xiao et al., 2022), pre-train specifically for the multi-document setting. Its pre-training objective masks entire sentences salient across a cluster of related documents, not only within one. The model learns cross-document salience directly, instead of relying on centroid heuristics at inference time.
An older, widely-used dataset, Multi-News (Fabbri et al., 2019), pairs clusters of news articles with a human-written summary. It remains a standard benchmark.
Evaluation
Standard ROUGE applies, computed against a human-written multi-document reference summary. A metric specific to this setting, though less commonly reported, checks source attribution. Can each claim in the summary be traced to an input document, and is that document correctly identified? This connects directly to faithfulness in summarisation. Here, faithfulness means "true to at least one source", not "true to a single fixed document".
Key references
- Radev, D. et al. (2004). Centroid-based Summarization of Multiple Documents. Information Processing & Management.
- Fabbri, A. et al. (2019). Multi-News. arXiv:1906.01749
- Xiao, W. et al. (2022). PRIMERA. arXiv:2110.08499
Current state and open problems
LLMs with long context windows can now take many full articles as input directly, sidestepping explicit sentence-level selection. Redundancy and contradiction detection do not disappear under this approach. They move from an explicit algorithmic step to an implicit one, buried inside generation. That makes them harder to inspect or debug.
Contradiction handling remains largely unsolved. No widely-adopted production system reliably flags "sources disagree here" inside a generated summary. Most silently pick one version, or blend both into something neither source actually said.
What to learn next
- Handling conflicting sources — the contradiction problem, in full.
- Document clustering with embeddings — grouping the input documents before summarising each cluster.
- Faithfulness in summarisation — attribution and hallucination, extended to a multi-source setting.