Exact match and F1 for question answering
Exact match checks if an answer matches word for word after cleanup, while F1 gives partial credit for overlapping words.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Exact match checks if a model's answer matches the correct answer exactly. F1 gives partial credit when it is close.
Picture a teacher marking a fill-in-the-blank test. "What is the capital of France?" One student writes "Paris". Another writes "It's Paris, I think". Both got it right, in spirit.
A strict teacher only accepts "Paris" and marks the second student wrong. A fair teacher notices "Paris" is in there and gives some credit. Question-answering evaluation faces the exact same choice.
Why it exists
A model answering questions rarely writes the exact same words as the correct answer. It might add "the answer is", use different capitalisation, or include an unnecessary word like "the".
Checking only for an identical string punishes harmless differences. That is too strict, and it hides real progress.
Never checking exactness has the opposite problem. A model that rambles around the right answer then looks the same as one that nails it. That is too loose.
Exact match (EM) and F1 are the two-metric compromise question-answering settled on, ever since the SQuAD benchmark introduced them.
How it works
Gold answer: "the Eiffel Tower"
Model answer: "The Eiffel Tower"
Step 1: clean up both (lowercase, drop punctuation, drop "a/an/the")
"eiffel tower" vs "eiffel tower" -> now they match!
Step 2: Exact match -> yes, identical after cleanup
F1 -> perfect, every word overlapsTake a wordier answer, like "a very tall tower in Paris". Exact match now fails outright, since the words genuinely differ. F1 still gives partial credit for the real overlap that remains, though that credit drops sharply once an answer strays this far from the gold text.
Where you have already seen it
- Chatbot benchmarks. Leaderboards for question-answering models report EM and F1 side by side, the standard since the SQuAD benchmark.
- Customer support bots. Testing whether a support bot's answer contains the correct order number or policy detail.
- Search engine "featured snippets". Checking if the short answer box above search results actually answers the question asked.
Remember this
- Exact match is strict: the cleaned-up answer must match word for word.
- F1 is forgiving: it rewards partial word overlap between the two answers.
- Both metrics need a normalisation step first, or harmless differences get punished.
What to learn next
- BERTScore — grading by meaning instead of exact words, for longer answers.
- Named entity recognition — finding the exact span an answer needs to come from.
- Natural language inference — checking if two statements agree or contradict.
Developer — Code and libraries.
This is the exact scoring function used by the SQuAD benchmark, rebuilt from scratch. No libraries beyond the Python standard library are needed.
Setup
Nothing to install. This uses only re, string and collections from the standard library.
The SQuAD-style scoring function
import re, string
from collections import Counter
def normalize(text):
text = text.lower()
text = "".join(ch for ch in text if ch not in string.punctuation)
text = re.sub(r"\b(a|an|the)\b", " ", text)
return " ".join(text.split())
def exact_match(pred, gold):
return int(normalize(pred) == normalize(gold))
def f1(pred, gold):
pred_tokens = normalize(pred).split()
gold_tokens = normalize(gold).split()
common = Counter(pred_tokens) & Counter(gold_tokens)
overlap = sum(common.values())
if overlap == 0:
return 0.0
precision = overlap / len(pred_tokens)
recall = overlap / len(gold_tokens)
return 2 * precision * recall / (precision + recall)
gold = "the Eiffel Tower"
predictions = ["The Eiffel Tower", "Eiffel Tower", "a very tall tower in Paris"]
for pred in predictions:
print(f"pred={pred!r:32} EM={exact_match(pred, gold)} F1={f1(pred, gold):.3f}")pred='The Eiffel Tower' EM=1 F1=1.000 pred='Eiffel Tower' EM=1 F1=1.000 pred='a very tall tower in Paris' EM=0 F1=0.286
The second prediction dropped "the" entirely, yet still scores EM=1. That is normalize at work: it strips articles before comparing, so "Eiffel Tower" and "the Eiffel Tower" become identical.
Line by line
normalize runs three cleanup steps in order: lowercase, strip punctuation, then drop the articles "a", "an" and "the". This exact recipe comes from the original SQuAD evaluation script, and most QA benchmarks since have copied it.
Counter(pred_tokens) & Counter(gold_tokens) returns the overlap between the two bags of words, counting each shared word only as many times as it appears in both. This is the same clipping idea used in BLEU and ROUGE.
F1 returns 0.0 early when overlap is 0, avoiding a divide-by-zero later in the precision and recall calculation.
Common mistakes
Forgetting to normalise before comparing. Without it, "Paris." and "paris" count as completely different strings, and a correct answer gets marked wrong for punctuation alone.
Using only exact match for open-ended answers. EM is strict by design. For any answer longer than a few words, F1 is the more informative number, and EM alone will look unfairly low.
Assuming F1 needs the words in the same order. This implementation uses bag-of-words overlap, not sequence order. "Tower Eiffel the" scores identically to "the Eiffel Tower" here, which is a real limitation worth knowing.
Try it yourself
Add a fourth prediction, "Big Ben", a real landmark but the wrong one. Rerun the script.
Both EM and F1 land at 0, since it shares no words with "the Eiffel Tower". This is the case both metrics handle well: a confident, fluent, completely wrong answer.
What to learn next
- Named entity recognition — extracting the kind of short spans this metric grades.
- BERTScore — a meaning-aware alternative for longer, less span-like answers.
- Text classification — a different shape of question-answering task, picking from fixed options.
Researcher — Mathematics and papers.
The SQuAD evaluation protocol
EM(p, g) = 1[ normalize(p) = normalize(g) ]
F1(p, g) = 2 * Precision * Recall / (Precision + Recall)pis the predicted answer span,ga single gold answer.normalizelowercases, strips punctuation and articles, and collapses whitespace, exactly as in the developer block.Precision = |overlap| / |p_tokens|,Recall = |overlap| / |g_tokens|, withoverlapcomputed via clipped multiset intersection.
Rajpurkar et al. (2016) introduced this protocol for SQuAD, with one addition not shown above: multiple gold answers per question, each collected from a different human annotator. The reported score for a question is the maximum EM and F1 across all gold answers, since any one of them is an acceptable correct answer.
EM_final = max over gold answers g of EM(p, g)
F1_final = max over gold answers g of F1(p, g)Skipping the max-over-golds step is a common reimplementation bug: it silently lower-bounds every score by treating one arbitrary gold answer as the only correct one.
The unanswerable-question extension
SQuAD 2.0 (Rajpurkar, Jia & Liang, 2018) added questions with no answer in the passage at all, requiring the model to predict an explicit "no answer" token. The metric extends naturally. Predicting "no answer" when the gold label agrees counts as a correct exact match. A wrong "no answer" guess, or a wrong span when one did exist, both count as a miss.
This single addition meaningfully changed which systems performed best. Models that reliably extracted a plausible-looking span, whether or not the passage actually supported one, dropped in rank relative to models that could also recognise absence.
Why this metric family does not transfer to generative QA
Both EM and F1 assume the answer is a short span, with a small, well-defined space of acceptable phrasings. Free-form generative answers, from a chatbot rather than an extractive reader, break that assumption. A fully correct answer can be phrased in dozens of ways no bag-of-words metric will catch, which is what motivates BERTScore and LLM-judge approaches instead.
Complexity
Both EM and F1 are O(n) in the number of tokens per answer, dominated by the string normalisation step rather than the comparison itself. This is negligible compared to model inference cost, and is one reason EM and F1 remain standard even in eras where compute-heavy metrics are otherwise preferred.
Key references
- Rajpurkar, P., Zhang, J., Lopyrev, K. & Liang, P. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. arXiv:1606.05250
- Rajpurkar, P., Jia, R. & Liang, P. (2018). Know What You Don't Know: Unanswerable Questions for SQuAD. arXiv:1806.03822
Current state and open problems
EM and F1 remain the default for extractive question answering, largely because the task itself, pick a span from a given passage, keeps the space of correct answers naturally small.
The open problem sits at the boundary this metric family was never built for. In retrieval-augmented and generative QA, an answer is synthesised rather than extracted, and a fully correct response may share no exact span with any gold answer at all. Most systems in that setting report EM and F1 only against an extractive baseline, and lean on LLM-as-judge or citation-based metrics otherwise.
What to learn next
- Using an LLM to grade text — the metric family generative QA has moved toward.
- Named entity recognition — the span-extraction task this evaluation protocol was originally built around.
- What is RAG — the generative-QA setting where this metric family stops applying cleanly.