Evaluating Text Systems

BERTScore and embedding-based metrics

BERTScore compares meaning instead of exact words, by matching each word to its closest counterpart in embedding space.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

BERTScore checks if two pieces of text mean the same thing, even when they use different words.

Picture two friends describing a movie they both watched. One says "the hero saved the city". The other says "the main character stopped the disaster". Different words, same story.

You know they mean the same thing because you understand the words, not because the sentences match letter for letter. BERTScore gives a computer a version of that understanding.

Why it exists

BLEU and ROUGE both count matching words. A correct paraphrase, one that uses different words for the same meaning, scores badly on both. That is a real gap, since good writing rarely repeats the reference word for word.

The fix needs a way to compare meaning, not spelling. An embedding is a list of numbers standing in for a word's meaning. Similar-meaning words get similar numbers.

How it works

  Reference:  "the cat sat on the mat"
  Candidate:  "a cat was resting on the rug"

  Each word becomes a list of numbers (an embedding).
  Instead of checking for identical words, BERTScore checks
  which words are CLOSEST in meaning:

  cat  <-> cat        (same word, perfect match)
  sat  <-> resting     (different words, close meaning)
  mat  <-> rug          (different words, close meaning)

  Result: a high score, because the closest matches
  are all genuinely similar in meaning.

Every word in the candidate gets matched to its closest-meaning word in the reference. The overall score is built from how close those best matches are, averaged across the sentence.

Where you have already seen it

  • Paraphrase detection. Tools that flag two differently-worded sentences as saying the same thing use this exact matching idea.
  • Modern research papers. Most summarisation and translation papers now report BERTScore alongside ROUGE or BLEU, precisely to catch what those miss.
  • Plagiarism checkers. Meaning-aware checkers that catch a copied idea even after every word has been swapped for a synonym.

Remember this

  • BERTScore matches words by meaning, using embeddings, not by exact spelling.
  • It gives real credit for correct paraphrases, which BLEU and ROUGE cannot do.
  • It still needs a model to run, so it costs more compute than counting words.

What to learn next

Developer — Code and libraries.

BERTScore needs a real embedding model. This example uses distilbert-base-uncased, a smaller cousin of BERT, so the download stays around 270MB and it runs on CPU in seconds.

Setup

bash
pip install bert-score

Scoring a paraphrase against an unrelated sentence

bertscore_demo.py
import logging
logging.getLogger("transformers").setLevel(logging.ERROR)
from bert_score import score

references = ["the cat sat on the mat", "the cat sat on the mat"]
candidates = ["a cat was resting on the rug",       # good paraphrase, different words
              "the stock market fell sharply today"]  # unrelated sentence

P, R, F1 = score(candidates, references, model_type="distilbert-base-uncased", verbose=False)

for text, f in zip(["good paraphrase   ", "unrelated sentence"], F1):
    print(f"{text} BERTScore F1 = {f.item():.4f}")

The first run downloads the model (about 270MB) and prints a short loading report from the transformers library. Only the two score lines below matter.

Output
good paraphrase    BERTScore F1 = 0.9040
unrelated sentence BERTScore F1 = 0.6903

The paraphrase scores 0.9040, close to a perfect match, despite sharing almost no exact words with the reference. Recall BLEU gave a similar paraphrase only 0.0773. This is the entire point of BERTScore.

Notice the unrelated sentence still scores 0.6903, not close to zero. Embedding similarity between any two ordinary English sentences rarely drops to zero, since common words like "the" always partly match. Read BERTScore as a comparison between candidates, not as an absolute 0-to-1 correctness scale.

Line by line

model_type="distilbert-base-uncased" picks which embedding model does the matching. The original BERTScore paper defaults to a much larger RoBERTa model. A smaller model runs faster, at some cost to how well it captures subtle meaning differences.

P, R, F1 are precision, recall and F1, the same idea as ROUGE. Precision checks whether the candidate's words are backed by the reference. Recall checks whether the reference's words got covered.

Common mistakes

Comparing BERTScore numbers computed with different model_type settings. A different embedding model produces a different scale of scores entirely. Always report which model produced your BERTScore.

Treating BERTScore as an absolute quality percentage. As the unrelated-sentence example shows, scores rarely go near zero. Compare scores against each other, not against an assumed threshold like 0.9 meaning "correct".

Running it on a GPU-less machine expecting BLEU-like speed. Every sentence needs a forward pass through a real model. For thousands of sentences, this is meaningfully slower than counting words, and worth budgeting for.

Try it yourself

Add a third candidate with related-sounding words that change the meaning: "the cat sat on the roof". Compare its score against the two above.

It should score high. "Roof" sits embedding-close to "mat" and "rug", and the sentence structure matches closely. This is a real limitation. BERTScore rewards surface and topical similarity. It can still miss a single word that flips the actual meaning.

What to learn next

  • Embeddings — the representations this metric depends on entirely.
  • BERT — the model family BERTScore is named after and built on.
  • Natural language inference — a metric built specifically to catch meaning flips like the one above.

Researcher — Mathematics and papers.

The matching procedure

Given a reference R = (r_1, ..., r_k) and candidate C = (c_1, ..., c_l), both tokenized and embedded by a pretrained contextual model, BERTScore computes:

text
Recall    = (1/k) * sum_{i=1}^{k} max_{j} cos(r_i, c_j)
Precision = (1/l) * sum_{j=1}^{l} max_{i} cos(c_j, r_i)
F1        = 2 * Precision * Recall / (Precision + Recall)
  • cos(a, b) is cosine similarity between two token embeddings.
  • Each reference token greedily matches its single closest candidate token by cosine similarity, and vice versa for precision; matches are not required to be mutual or one-to-one.
  • Embeddings are contextual: the same word gets a different vector depending on surrounding words, unlike static embeddings such as Word2Vec.

Zhang et al. (2020) additionally introduce importance weighting using inverse document frequency, so that matching a rare, informative word counts for more than matching "the". This is optional in most implementations, off by default in the example above.

Why contextual embeddings matter here

A static embedding gives "bank" one vector regardless of context, unable to distinguish "river bank" from "bank account". A contextual model, using self-attention over the full sentence, produces a different vector for "bank" in each case.

This is precisely why BERTScore uses BERT-family models rather than Word2Vec or GloVe: matching needs to reflect meaning-in-context, not only word identity.

Correlation with human judgement

Zhang et al. (2020) report BERTScore correlating with human quality judgements substantially better than BLEU and ROUGE, across machine translation and image captioning benchmarks, using WMT and other human-annotated datasets as ground truth.

The gain is largest on correct paraphrases, exactly where overlap-based metrics fail. It is smallest where the reference and candidate already share most surface words, where BLEU already does reasonably well.

Complexity and cost

StepCost
Embedding both textsOne transformer forward pass per sentence pair
Pairwise cosine matrixO(k * l) per sentence pair
Greedy matchingO(k * l), same matrix reused for precision and recall

The dominant cost by far is the transformer forward pass, not the matching arithmetic, which is why model choice (distilbert versus roberta-large) is the main speed lever available.

Known limitations

BERTScore inherits whatever blind spots its underlying embedding model has. Sinha et al. (2021) found embedding-based metrics can miss negation and numeric changes. "The price rose 5%" and "the price rose 50%" embed as near-identical sentences, despite an enormous factual difference.

This connects directly to the model-internals work on probing: a metric is only as good at detecting a distinction as the representation it is built on is at encoding that distinction.

Key references

  • Zhang, T., Kishore, V., Wu, F., Weinberger, K. & Artzi, Y. (2020). BERTScore: Evaluating Text Generation with BERT. arXiv:1904.09675
  • Sinha, K. et al. (2021). Unnatural Language Inference. ACL — related evidence on what contextual embeddings fail to encode.

Current state and open problems

BERTScore is now reported routinely alongside BLEU and ROUGE, though rarely alone, since reviewers expect the overlap-based numbers for continuity with older results.

Newer metrics such as BLEURT (Sellam, Das & Parikh, 2020) fine-tune a model directly on human-rating data, rather than relying only on generic sentence embeddings, the way COMET improved on BERTScore's approach for translation.

The open question is the same across this whole family of metrics. An embedding-based score is only as trustworthy as the training data and objective that shaped the embedding space. That makes it harder to audit than counting words, not easier.

What to learn next