ROUGE and what it misses
ROUGE scores a summary by checking how much of the reference's content it kept, which makes it easy to game with padding.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
ROUGE checks a summary by counting how much of a reference's content made it in. It works the way you check a packing list before a trip.
Picture packing for a trip with a checklist. You go through the list, ticking off items you actually packed. Did you bring the passport? The charger? The medicine?
That checklist check is about recall — how much of what mattered got included. ROUGE, short for Recall-Oriented Understudy for Gisting Evaluation, scores a summary the same way.
Why it exists
BLEU, the metric used for translation, cares mostly about precision — of the words you wrote, how many were correct. That fits translation, where every word in the reference matters.
Summarisation is different. A summary is allowed to use its own words. What matters is whether it kept the important content from the original. Missing a key fact is a bigger sin than phrasing things differently.
ROUGE was built for exactly that shift in priority: reward keeping the content, and only lightly penalise extra length.
How it works
Reference summary: "the fast fox jumped over the lazy dog near the river"
Your summary: "the fox jumped over the dog"
ROUGE checks: how many words from the reference
made it into your summary?
fox -> yes
jumped -> yes
over -> yes
dog -> yes
fast -> missing
lazy -> missing
near -> missing
river -> missing
6 out of 11 reference words recovered -> decent recallROUGE comes in a few common flavours. ROUGE-1 counts single matching words. ROUGE-2 counts matching pairs of words in a row. ROUGE-L finds the longest matching sequence of words in the same order, even with gaps between them.
Where you have already seen it
- News summary tools. Apps that condense articles into three bullet points get graded against human-written summaries using ROUGE.
- Meeting notes generators. Comparing an AI-written meeting summary against what a human note-taker wrote down.
- Research leaderboards. Nearly every summarisation paper reports ROUGE-1, ROUGE-2 and ROUGE-L side by side.
Remember this
- ROUGE checks how much of a reference's content a summary kept, closer to recall than precision.
- ROUGE-1, ROUGE-2 and ROUGE-L check single words, word pairs and longest matching order.
- ROUGE only counts overlapping words. It has no idea whether the summary actually makes sense.
What to learn next
- BLEU, chrF and COMET — the precision-focused cousin of this metric.
- BERTScore — matching meaning instead of exact words.
- Text classification — a different way models handle short text, worth contrasting.
Developer — Code and libraries.
The point here is not only running ROUGE. It is watching it get gamed by a lazy trick, so you know to check for it in your own pipelines.
Setup
pip install rouge-scoreA good summary versus a padded one
from rouge_score import rouge_scorer
scorer = rouge_scorer.RougeScorer(["rouge1"], use_stemmer=True)
reference = "the fast fox jumped over the lazy dog near the river"
good_summary = "the fox jumped over the dog"
padded_summary = "the the the the the fox jumped over the dog the the the"
for name, cand in [("good summary", good_summary), ("padded summary", padded_summary)]:
r1 = scorer.score(reference, cand)["rouge1"]
print(f"{name:14} precision={r1.precision:.3f} recall={r1.recall:.3f} f1={r1.fmeasure:.3f}")good summary precision=1.000 recall=0.545 f1=0.706 padded summary precision=0.538 recall=0.636 f1=0.583
The padded summary is nonsense. It still gets a higher recall than the honest one, 0.636 against 0.545. It stuffed in extra copies of "the", which happens to appear five times in the reference too.
The overall F1 score catches this. Precision drops hard, from 1.000 to 0.538, and the combined score ends up lower. This is exactly why F1 exists: recall alone rewards padding.
Line by line
use_stemmer=True reduces words to their root before comparing, so "jumping" matches "jumped". Without it, ROUGE would miss credit for grammatical variants that mean the same thing.
.precision, .recall, .fmeasure are the three numbers rouge_score returns for every metric. Always look at all three. A single F1 number hides which direction a summary is failing in.
Common mistakes
Reporting recall alone. As shown above, recall can be inflated by padding. Report F1, or precision and recall together, never recall by itself.
Comparing ROUGE scores computed with different settings. Stemming on versus off, and which ROUGE variant, both shift the number. State your settings whenever you report a score.
Trusting a high ROUGE score as proof of a good summary. ROUGE cannot detect a fluent-sounding summary that is factually wrong, if it happens to reuse the reference's words. Word overlap is not the same thing as correctness.
Try it yourself
Write a summary that uses completely different words but says the same thing as the reference, for example "a quick fox leapt across a sleepy hound by the water". Rerun the script.
ROUGE gives it a low score, close to zero, despite it being an accurate summary. This is the same blind spot BLEU has. It is what motivates embedding-based metrics like BERTScore, next.
What to learn next
- BERTScore — scoring meaning with embeddings instead of exact word matches.
- Using an LLM to grade text — a further step past overlap counting.
- Hallucination — why a fluent summary can still say something false.
Researcher — Mathematics and papers.
The three common variants
ROUGE-N is n-gram recall between candidate C and reference R:
ROUGE-N = ( sum over n-grams in R of min(count_C(gram), count_R(gram)) ) / ( sum over n-grams in R of count_R(gram) )count_C(gram),count_R(gram)are the number of times an n-gram appears in the candidate and reference.- The
minclips credit, the same clipping BLEU uses, so repeating a reference n-gram cannot inflate the score. Nis typically 1 or 2 in practice; higher orders get sparse fast on short summaries.
ROUGE-L uses the length of the longest common subsequence (LCS) instead of contiguous n-grams:
R_lcs = LCS(C, R) / len(R)
P_lcs = LCS(C, R) / len(C)
F_lcs = (1 + beta^2) * P_lcs * R_lcs / (R_lcs + beta^2 * P_lcs)LCS(C, R)is the length of the longest subsequence common to both, in order but allowing gaps.betais usually set high, favouring recall, matching ROUGE's original design goal.
LCS-based matching tolerates word reordering and insertion better than fixed n-grams, at the cost of losing sensitivity to exact local phrasing.
Where the original paper set the defaults
Lin (2004), ROUGE: A Package for Automatic Evaluation of Summaries, introduced ROUGE-N, ROUGE-L, ROUGE-W (weighted LCS, favouring consecutive matches) and ROUGE-S (skip-bigram co-occurrence). ROUGE-1, ROUGE-2 and ROUGE-L are the three still commonly reported.
The paper's own validation compared ROUGE scores against human-assigned summary quality on DUC datasets, finding strong correlation for multi-document summarisation specifically. That correlation is weaker outside the conditions Lin tested, which later work returned to repeatedly.
Known correlation problems
Cohan & Goharian (2016) and later work on long-document summarisation found ROUGE correlates poorly with human judgement for abstractive summaries. The gap grows once a summary moves away from copying sentences, toward genuinely rewriting them.
The core mechanical issue: ROUGE cannot reward a paraphrase, and abstractive summarisation exists specifically to produce paraphrases. A system optimised directly against ROUGE (for instance with reinforcement learning) can learn to copy reference phrasing more than it learns to summarise well, a known failure mode called ROUGE hacking.
Complexity
ROUGE-N computation is O(length of C + length of R) with hash-based n-gram counting. ROUGE-L needs an LCS instead, at O(len(C) * len(R)) with the standard dynamic-programming algorithm. That cost is negligible at summary lengths, but worth knowing before applying it to full-document pairs.
Key references
- Lin, C.-Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. ACL Workshop.
- Cohan, A. & Goharian, N. (2016). Revisiting Summarization Evaluation for Scientific Articles. LREC.
- Paulus, R., Xiong, C. & Socher, R. (2018). A Deep Reinforced Model for Abstractive Summarization. arXiv:1705.04304 — an early example of ROUGE-hacking behaviour under RL training.
Current state and open problems
ROUGE remains standard practice for reporting summarisation results, largely for continuity with older literature rather than because it is the best available option. Most current papers report it alongside BERTScore or an LLM-judge score, similar to how translation moved from BLEU-alone to BLEU-plus-COMET.
The unresolved problem is the same one facing every overlap metric here: none of them can verify factual correctness. A summary can score well on ROUGE, BERTScore and even an LLM judge, while quietly asserting something the source document never said. Dedicated factual-consistency metrics exist for this specific failure, and remain an active area rather than a solved one.
What to learn next
- BERTScore — the embedding-based successor most papers report alongside ROUGE now.
- Hallucination — the factual-correctness failure ROUGE cannot see.
- Using an LLM to grade text — evaluation that can, imperfectly, judge meaning.