Evaluating Text Systems

ROUGE and what it misses

ROUGE scores a summary by checking how much of the reference's content it kept, which makes it easy to game with padding.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

ROUGE checks a summary by counting how much of a reference's content made it in. It works the way you check a packing list before a trip.

Picture packing for a trip with a checklist. You go through the list, ticking off items you actually packed. Did you bring the passport? The charger? The medicine?

That checklist check is about recall — how much of what mattered got included. ROUGE, short for Recall-Oriented Understudy for Gisting Evaluation, scores a summary the same way.

Why it exists

BLEU, the metric used for translation, cares mostly about precision — of the words you wrote, how many were correct. That fits translation, where every word in the reference matters.

Summarisation is different. A summary is allowed to use its own words. What matters is whether it kept the important content from the original. Missing a key fact is a bigger sin than phrasing things differently.

ROUGE was built for exactly that shift in priority: reward keeping the content, and only lightly penalise extra length.

How it works

  Reference summary:  "the fast fox jumped over the lazy dog near the river"
  Your summary:        "the fox jumped over the dog"

  ROUGE checks: how many words from the reference
  made it into your summary?

  fox        -> yes
  jumped     -> yes
  over       -> yes
  dog        -> yes
  fast       -> missing
  lazy       -> missing
  near       -> missing
  river      -> missing

  6 out of 11 reference words recovered -> decent recall

ROUGE comes in a few common flavours. ROUGE-1 counts single matching words. ROUGE-2 counts matching pairs of words in a row. ROUGE-L finds the longest matching sequence of words in the same order, even with gaps between them.

Where you have already seen it

  • News summary tools. Apps that condense articles into three bullet points get graded against human-written summaries using ROUGE.
  • Meeting notes generators. Comparing an AI-written meeting summary against what a human note-taker wrote down.
  • Research leaderboards. Nearly every summarisation paper reports ROUGE-1, ROUGE-2 and ROUGE-L side by side.

Remember this

  • ROUGE checks how much of a reference's content a summary kept, closer to recall than precision.
  • ROUGE-1, ROUGE-2 and ROUGE-L check single words, word pairs and longest matching order.
  • ROUGE only counts overlapping words. It has no idea whether the summary actually makes sense.

What to learn next

Developer — Code and libraries.

The point here is not only running ROUGE. It is watching it get gamed by a lazy trick, so you know to check for it in your own pipelines.

Setup

bash
pip install rouge-score

A good summary versus a padded one

rouge_demo.py
from rouge_score import rouge_scorer

scorer = rouge_scorer.RougeScorer(["rouge1"], use_stemmer=True)

reference = "the fast fox jumped over the lazy dog near the river"
good_summary = "the fox jumped over the dog"
padded_summary = "the the the the the fox jumped over the dog the the the"

for name, cand in [("good summary", good_summary), ("padded summary", padded_summary)]:
    r1 = scorer.score(reference, cand)["rouge1"]
    print(f"{name:14} precision={r1.precision:.3f}  recall={r1.recall:.3f}  f1={r1.fmeasure:.3f}")
Output
good summary   precision=1.000  recall=0.545  f1=0.706
padded summary precision=0.538  recall=0.636  f1=0.583

The padded summary is nonsense. It still gets a higher recall than the honest one, 0.636 against 0.545. It stuffed in extra copies of "the", which happens to appear five times in the reference too.

The overall F1 score catches this. Precision drops hard, from 1.000 to 0.538, and the combined score ends up lower. This is exactly why F1 exists: recall alone rewards padding.

Line by line

use_stemmer=True reduces words to their root before comparing, so "jumping" matches "jumped". Without it, ROUGE would miss credit for grammatical variants that mean the same thing.

.precision, .recall, .fmeasure are the three numbers rouge_score returns for every metric. Always look at all three. A single F1 number hides which direction a summary is failing in.

Common mistakes

Reporting recall alone. As shown above, recall can be inflated by padding. Report F1, or precision and recall together, never recall by itself.

Comparing ROUGE scores computed with different settings. Stemming on versus off, and which ROUGE variant, both shift the number. State your settings whenever you report a score.

Trusting a high ROUGE score as proof of a good summary. ROUGE cannot detect a fluent-sounding summary that is factually wrong, if it happens to reuse the reference's words. Word overlap is not the same thing as correctness.

Try it yourself

Write a summary that uses completely different words but says the same thing as the reference, for example "a quick fox leapt across a sleepy hound by the water". Rerun the script.

ROUGE gives it a low score, close to zero, despite it being an accurate summary. This is the same blind spot BLEU has. It is what motivates embedding-based metrics like BERTScore, next.

What to learn next

Researcher — Mathematics and papers.

The three common variants

ROUGE-N is n-gram recall between candidate C and reference R:

text
ROUGE-N = ( sum over n-grams in R of min(count_C(gram), count_R(gram)) ) / ( sum over n-grams in R of count_R(gram) )
  • count_C(gram), count_R(gram) are the number of times an n-gram appears in the candidate and reference.
  • The min clips credit, the same clipping BLEU uses, so repeating a reference n-gram cannot inflate the score.
  • N is typically 1 or 2 in practice; higher orders get sparse fast on short summaries.

ROUGE-L uses the length of the longest common subsequence (LCS) instead of contiguous n-grams:

text
R_lcs = LCS(C, R) / len(R)
P_lcs = LCS(C, R) / len(C)
F_lcs = (1 + beta^2) * P_lcs * R_lcs / (R_lcs + beta^2 * P_lcs)
  • LCS(C, R) is the length of the longest subsequence common to both, in order but allowing gaps.
  • beta is usually set high, favouring recall, matching ROUGE's original design goal.

LCS-based matching tolerates word reordering and insertion better than fixed n-grams, at the cost of losing sensitivity to exact local phrasing.

Where the original paper set the defaults

Lin (2004), ROUGE: A Package for Automatic Evaluation of Summaries, introduced ROUGE-N, ROUGE-L, ROUGE-W (weighted LCS, favouring consecutive matches) and ROUGE-S (skip-bigram co-occurrence). ROUGE-1, ROUGE-2 and ROUGE-L are the three still commonly reported.

The paper's own validation compared ROUGE scores against human-assigned summary quality on DUC datasets, finding strong correlation for multi-document summarisation specifically. That correlation is weaker outside the conditions Lin tested, which later work returned to repeatedly.

Known correlation problems

Cohan & Goharian (2016) and later work on long-document summarisation found ROUGE correlates poorly with human judgement for abstractive summaries. The gap grows once a summary moves away from copying sentences, toward genuinely rewriting them.

The core mechanical issue: ROUGE cannot reward a paraphrase, and abstractive summarisation exists specifically to produce paraphrases. A system optimised directly against ROUGE (for instance with reinforcement learning) can learn to copy reference phrasing more than it learns to summarise well, a known failure mode called ROUGE hacking.

Complexity

ROUGE-N computation is O(length of C + length of R) with hash-based n-gram counting. ROUGE-L needs an LCS instead, at O(len(C) * len(R)) with the standard dynamic-programming algorithm. That cost is negligible at summary lengths, but worth knowing before applying it to full-document pairs.

Key references

  • Lin, C.-Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. ACL Workshop.
  • Cohan, A. & Goharian, N. (2016). Revisiting Summarization Evaluation for Scientific Articles. LREC.
  • Paulus, R., Xiong, C. & Socher, R. (2018). A Deep Reinforced Model for Abstractive Summarization. arXiv:1705.04304 — an early example of ROUGE-hacking behaviour under RL training.

Current state and open problems

ROUGE remains standard practice for reporting summarisation results, largely for continuity with older literature rather than because it is the best available option. Most current papers report it alongside BERTScore or an LLM-judge score, similar to how translation moved from BLEU-alone to BLEU-plus-COMET.

The unresolved problem is the same one facing every overlap metric here: none of them can verify factual correctness. A summary can score well on ROUGE, BERTScore and even an LLM judge, while quietly asserting something the source document never said. Dedicated factual-consistency metrics exist for this specific failure, and remain an active area rather than a solved one.

What to learn next