BLEU, chrF and COMET
BLEU scores a translation by counting matching word chunks against a reference, which is fast but blind to correct paraphrases.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
BLEU checks a translation by counting matching word chunks against a reference. It works the way a teacher checks homework against an answer key.
Picture a teacher marking a translation exercise. She has the official answer key. She counts how many words and phrases the student's answer shares with it. More matches, higher marks.
That is BLEU, short for Bilingual Evaluation Understudy. It compares machine output against one or more reference translations, word chunk by word chunk.
Why it exists
Before BLEU, checking translation quality meant paying a human to read every sentence. That does not scale. A company shipping updates daily cannot wait days for human review of each one.
BLEU gave researchers a number they could compute in seconds. Change your model, rerun BLEU, see if the number went up. No human needed for every test.
The catch: BLEU only counts matching words. It has no idea what the sentence means.
How it works
Reference: "the cat sat on the mat"
Candidate: "the cat sat on the mat" -> huge overlap, high score
Reference: "the cat sat on the mat"
Candidate: "a cat was resting on the rug" -> correct meaning,
almost no shared words,
low scoreBLEU breaks both sentences into overlapping chunks of one, two, three and four words. It counts how many chunks in the candidate also appear in the reference. It also penalises answers that are suspiciously short, since a one-word answer can accidentally match perfectly.
That second example is the whole problem with BLEU in one line. A fluent, accurate translation that happens to use different words still scores badly.
chrF fixes part of this by matching chunks of letters instead of whole words. "resting" and "rug" still share no words with "sat" and "mat". But chrF gives partial credit for close spellings and shared word endings. That helps languages where one word changes shape with grammar, such as Hindi or Tamil.
COMET goes further still. It uses a trained neural network instead of counting overlapping text. That network learned what a good translation "feels like", from thousands of human judgements. It can reward a correct paraphrase that shares no words with the reference at all. The tradeoff: COMET needs a model to run. That costs more compute and takes longer than counting words.
Where you have already seen it
- Google Translate's internal testing. Every improvement to a translation model gets a BLEU number before and after, long before a human reviews it.
- Research papers. Almost every machine translation paper for two decades reported a BLEU score, for better or worse.
- Leaderboards. Translation competitions such as WMT rank systems by BLEU, chrF and increasingly COMET, side by side.
Remember this
- BLEU counts matching word chunks against a reference. It is fast, but blind to correct paraphrases.
- chrF matches letter chunks instead of whole words, which helps with rich grammar and small spelling differences.
- COMET uses a trained model to judge meaning. It catches good paraphrases that BLEU and chrF miss, at a higher compute cost.
What to learn next
- ROUGE and what it misses — the same idea, used for summaries instead of translations.
- BERTScore — matching meaning with embeddings instead of counting letters.
- Tokenization — the chopping step every one of these metrics depends on.
Developer — Code and libraries.
Two things are worth running here. First, watch BLEU fail on a correct paraphrase. Second, build a tiny chrF from scratch, so it stops being a black box.
Both run on plain Python, using only nltk for the standard BLEU implementation.
Setup
pip install nltkBLEU on an exact match versus a correct paraphrase
from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction
reference = ["the cat sat on the mat".split()]
exact = "the cat sat on the mat".split()
paraphrase = "a cat was resting on the rug".split()
smooth = SmoothingFunction().method1
for name, cand in [("exact match", exact), ("good paraphrase", paraphrase)]:
score = sentence_bleu(reference, cand, smoothing_function=smooth)
print(f"{name:16} BLEU = {score:.4f}")exact match BLEU = 1.0000 good paraphrase BLEU = 0.0773
The paraphrase is a completely correct translation of the reference's meaning. BLEU still gives it almost nothing, because it shares almost no word chunks with the reference. This is the central limitation of BLEU, reproduced in four lines.
A tiny chrF, built from scratch
from collections import Counter
def char_ngrams(text, n):
text = text.replace(" ", "_")
return [text[i:i + n] for i in range(len(text) - n + 1)]
def chrf(reference, candidate, n=4, beta=2):
ref_grams = Counter(char_ngrams(reference, n))
cand_grams = Counter(char_ngrams(candidate, n))
overlap = sum((ref_grams & cand_grams).values())
precision = overlap / max(sum(cand_grams.values()), 1)
recall = overlap / max(sum(ref_grams.values()), 1)
if precision + recall == 0:
return 0.0
return (1 + beta**2) * precision * recall / (beta**2 * precision + recall)
reference = "the cat sat on the mat"
exact = "the cat sat on the mat"
paraphrase = "a cat was resting on the rug"
for name, cand in [("exact match", exact), ("good paraphrase", paraphrase)]:
print(f"{name:16} chrF(n=4) = {chrf(reference, cand):.4f}")exact match chrF(n=4) = 1.0000 good paraphrase chrF(n=4) = 0.3465
chrF gives the paraphrase roughly four and a half times the score BLEU did. Still far from perfect, since chrF only sees letters, not meaning, but noticeably more forgiving.
Line by line
char_ngrams replaces spaces with _ first. This keeps a chunk like at_m distinguishable from one that spans no word boundary. Without it, chrF would silently merge unrelated words.
ref_grams & cand_grams on two Counter objects returns the overlap — the minimum count of each shared chunk in both. This is standard n-gram precision and recall counting, the same idea BLEU uses at the word level.
beta=2 weights recall twice as heavily as precision. This is chrF's standard setting. Missing content in a translation is usually judged worse than adding a little extra.
Common mistakes
Comparing BLEU scores across different tokenizers. BLEU counts word chunks, so a different word-splitting rule produces a different score on identical text. Always report which tokenizer produced your BLEU number.
Running BLEU on a single sentence and trusting it. BLEU was designed as a corpus-level metric, averaged over many sentences. A single-sentence BLEU score is noisy and can mislead. The score above shows one bad edge case looking catastrophic in isolation.
Assuming a higher BLEU always means a better translation. A model can learn to copy input words to inflate its BLEU score without producing fluent output. Always spot-check translations a metric scores highly.
Try it yourself
Change paraphrase to a bad translation that shares BLEU's favourite word chunks with the reference. Try "the cat sat on the rug", which differs only in the last word. Rerun both scripts.
BLEU stays high, since almost every chunk still matches. chrF also stays reasonably high, since only four letters changed. Neither metric notices the ending changed the meaning entirely, from a floor mat to a different object. That judgement needs a human, or a metric like COMET, trained on human judgements.
What to learn next
- ROUGE and what it misses — the summarisation cousin of this problem.
- BERTScore — scoring meaning instead of letters or words.
- Byte pair encoding — how the "words" being counted are actually produced.
Researcher — Mathematics and papers.
BLEU
BLEU = BP * exp( sum_{n=1}^{N} w_n * log(p_n) )p_nis the modified n-gram precision at ordern: matching n-gram count divided by total candidate n-gram count. Matches are clipped so a chunk cannot count more times than it appears in the reference.w_nare weights, conventionally1/NforN = 4(uniform across unigrams through 4-grams).BPis the brevity penalty:1if candidate lengthcexceeds reference lengthr, elseexp(1 - r/c). This exists because precision alone rewards short, safe outputs.
Papineni et al. (2002) proposed BLEU as a corpus-level metric: p_n is computed by summing counts across the entire test set before taking the ratio, not by averaging per-sentence scores. Per-sentence BLEU, as shown in the developer block, is a common but non-standard simplification, and is far noisier because a short sentence has few n-grams to match.
Clipping is the detail most reimplementations get wrong: a candidate that repeats a common reference word many times must not be credited more than the reference actually contains that word.
chrF
chrF_beta = (1 + beta^2) * chrP * chrR / (beta^2 * chrP + chrR)chrP,chrRare precision and recall over character n-grams (typically n up to 6), averaged over orders.betais set to 2 by convention (Popovic, 2015), weighting recall twice as heavily as precision.
Because chrF operates below the word level, it needs no tokenizer at all. This removes an entire class of BLEU's reproducibility problems. It also makes chrF more robust for morphologically rich and low-resource languages, where BLEU's word-level matching is especially brittle.
COMET
Rei et al. (2020) frame evaluation as a learned regression problem rather than a counting problem. A pretrained cross-lingual encoder (XLM-RoBERTa) embeds the source, the candidate and the reference. A small feed-forward head then predicts a quality score directly, trained on the WMT Metrics human-judgement datasets:
score = f( embed(source), embed(candidate), embed(reference) )fis a learned regressor, trained to match human direct-assessment scores.
Because it is trained on human judgements rather than surface overlap, COMET correlates with human ratings substantially better than BLEU or chrF at the segment level. The cost is a forward pass through a multi-hundred-million-parameter model for every sentence scored.
Complexity and cost
| Metric | Per-sentence cost | Needs a GPU | Tokenizer-sensitive |
|---|---|---|---|
| BLEU | O(sentence length) counting | No | Yes |
| chrF | O(sentence length) counting | No | No |
| COMET | One transformer forward pass | Preferred, not required | No |
Known failure modes
Callison-Burch, Osborne & Koehn (2006), Re-evaluating the Role of BLEU in Machine Translation Research, showed BLEU rankings can disagree with human rankings. This was most visible when comparing systems built on different paradigms — rule-based versus statistical, at the time.
Mathur, Baldwin & Cohn (2020), Tangled up in BLEU, found that metric reliability varies by domain and by how close the compared systems are in quality. No single metric, including COMET, should be trusted without checking it against the specific comparison being made.
Key references
- Papineni, K., Roukos, S., Ward, T. & Zhu, W.-J. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation. ACL.
- Popovic, M. (2015). chrF: character n-gram F-score for automatic MT evaluation. WMT.
- Rei, R. et al. (2020). COMET: A Neural Framework for MT Evaluation. arXiv:2009.09025
- Callison-Burch, C., Osborne, M. & Koehn, P. (2006). Re-evaluating the Role of BLEU in Machine Translation Research. EACL.
- Mathur, N., Baldwin, T. & Cohn, T. (2020). Tangled up in BLEU. arXiv:2006.06264
Current state and open problems
Modern MT leaderboards report BLEU alongside chrF and COMET rather than BLEU alone, because the three metrics disagree often enough that reporting only one hides real quality differences.
Newer learned metrics, such as COMET-QE and reference-free variants, score translations without needing any human-written reference at all. This helps languages where references are scarce, though these estimates remain less reliable than reference-based COMET.
The open problem is the same one that motivated COMET. Overlap-based metrics are cheap but shallow. Learned metrics are accurate, but need to be trusted, calibrated and periodically re-validated against fresh human judgement data, as quality improves and old benchmarks saturate.
What to learn next
- ROUGE and what it misses — the same overlap-counting idea, applied to summarisation.
- BERTScore — a lighter-weight embedding-based alternative to COMET.
- Embeddings — the representations that make meaning-aware metrics possible at all.