Question Answering

What SQuAD and Natural Questions actually test

A strong score on a QA benchmark proves a system is good at that benchmark's specific style of question, not that it will handle real, differently-shaped questions equally well.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A QA benchmark tests one specific kind of question, and a high score proves skill at that kind, not at every kind.

Think about a driving test that only checks parallel parking. Passing it proves someone can parallel park. It does not prove they can handle a busy highway, a hill start, or driving at night. None of those get tested.

Why it exists

Building a question-answering system needs some way to measure whether one version is better than another. Benchmarks like SQuAD and Natural Questions exist to give researchers a shared, repeatable test everyone can compare scores on.

That is genuinely useful. It is also easy to over-trust. A benchmark comes from a specific process — often, questions written by people looking at Wikipedia articles. That process shapes exactly what gets tested, and what does not.

A model can score very well on a benchmark's own style of question. It can still struggle on real users' messier, differently-phrased, differently-sourced questions. The benchmark score describes the benchmark, not automatically the real world.

How it works

SQuAD-style question:
  Written by someone reading a Wikipedia paragraph, answer is a
  short span, always present somewhere in that same paragraph.

Real user question:
  Typed by someone who has not read any paragraph, might be
  vague, might have no answer anywhere in your documents,
  might need combining facts from three different pages.

A model tuned only on the first kind can struggle badly on the second,
even after scoring 90%+ on the benchmark it was tuned against.

Where you have already seen it

  • Self-driving car demos versus real roads. A demo route rehearsed many times looks far more impressive than driving is in general.
  • School exam scores versus real-world skill. Acing a specific exam format does not guarantee the broader skill it was meant to represent.
  • "Works great in the demo" software, that behaves very differently once real, unpredictable users start typing into it.

Remember this

  • A benchmark score measures performance on that benchmark's specific question style, not question-answering in general.
  • Real user questions are messier than benchmark questions — more ambiguous, less certain to have an answer at all.
  • A high score is a reason for interest, not proof a system is ready for your specific use case.

What to learn next

Developer — Code and libraries.

This implements SQuAD's own scoring method by hand, then shows exactly where it penalises a correct, differently-worded answer.

Setup

bash
python --version   # 3.9 or newer, only the standard library is used

Exact match and F1, the way SQuAD actually scores

squad_metrics.py
import re
import string
from collections import Counter

def normalize(text):
    text = text.lower()
    text = "".join(ch for ch in text if ch not in string.punctuation)
    text = re.sub(r"\b(a|an|the)\b", " ", text)
    return " ".join(text.split())

def exact_match(pred, gold):
    return int(normalize(pred) == normalize(gold))

def f1(pred, gold):
    pred_tokens = normalize(pred).split()
    gold_tokens = normalize(gold).split()
    common = Counter(pred_tokens) & Counter(gold_tokens)
    num_same = sum(common.values())
    if num_same == 0:
        return 0.0
    precision = num_same / len(pred_tokens)
    recall = num_same / len(gold_tokens)
    return 2 * precision * recall / (precision + recall)

gold = "the Reserve Bank of India"
predictions = [
    "the Reserve Bank of India",   # exact
    "Reserve Bank of India",       # article dropped, fine after normalising
    "RBI",                         # correct in meaning, zero word overlap
    "India's central bank",        # correct in meaning, different words
]

for p in predictions:
    print(f"pred={p!r:30} EM={exact_match(p, gold)}  F1={f1(p, gold):.2f}")
Output
pred='the Reserve Bank of India'    EM=1  F1=1.00
pred='Reserve Bank of India'        EM=1  F1=1.00
pred='RBI'                          EM=0  F1=0.00
pred="India's central bank"         EM=0  F1=0.29

Line by line

normalize() lowercases, strips punctuation, and drops "a," "an," "the." This is SQuAD's actual normalisation, designed to stop trivial differences like capitalisation from counting as wrong answers.

exact_match is strict, all-or-nothing — 1 only if the normalised strings match exactly. f1 is more forgiving, based on shared word overlap between prediction and gold answer.

"RBI" scores 0 on both metrics, despite being a completely correct answer. It shares zero words with "Reserve Bank of India," because it is an abbreviation, not a paraphrase built from the same words. This is the exact limitation the beginner block warns about, made concrete with real numbers.

"India's central bank" scores a nonzero but low F1, 0.29. It shares "bank" with the gold answer, contributing a little overlap, but the rest of the words are unique to each side. A human grading this by hand would very likely call it correct.

Common mistakes

Treating F1 as a general similarity score for any two pieces of text. It is specifically a word-overlap score, blind to meaning. A synonym-heavy correct paraphrase scores low; a word-salad of gold-answer terms in the wrong order can score deceptively high.

Assuming a single gold answer is enough. Real SQuAD evaluation stores multiple acceptable gold answers per question exactly to reduce this problem, and reports the best score against any of them — a detail worth remembering before comparing scores computed against only one gold answer.

Comparing scores computed with different normalisation rules. A slightly different punctuation-stripping or article-removal rule changes scores measurably. Comparing your own system's score against a published SQuAD leaderboard number is only fair if the exact same scoring code was used.

Reading a small F1 gap as a meaningful quality difference. Two systems scoring 88.1 and 88.4 F1 are not necessarily different in any real-world useful sense. Benchmark score differences need statistical context, not bare comparison.

Try it yourself

Add a fifth prediction, "India's national bank" — a small, real-sounding factual error, "national" instead of "central."

Compare its F1 score against "RBI"'s score. The factually wrong answer likely scores higher than the correct abbreviation, purely from sharing more surface words with the gold answer — a genuinely useful, uncomfortable thing to see directly in your own output.

What to learn next

Researcher — Mathematics and papers.

What SQuAD and Natural Questions each actually measure

Rajpurkar, Zhang, Lopyrev & Liang (2016), SQuAD: 100,000+ Questions for Machine Comprehension of Text (arXiv:1606.05250), sourced questions from crowdworkers who read a specific Wikipedia paragraph and wrote a question they already knew the answer to, drawn from that same paragraph. This process guarantees answerability by construction in the original release, and produces questions with unusually high lexical overlap with their source passage, since the question-writer had the answer text in view while writing the question.

Kwiatkowski, T. et al. (2019), Natural Questions: A Benchmark for Question Answering Research, Transactions of the Association for Computational Linguistics, took a different approach: questions are real, anonymised queries issued to a search engine, and annotators then searched for a Wikipedia page and passage that could answer each one. This produces a meaningfully different question distribution — genuine information-seeking questions, asked before any answer was known, rather than questions reverse-engineered from a known answer. Natural Questions also permits a "no answer" label when no Wikipedia page adequately answers the query, closer to real search traffic than SQuAD's original always-answerable design.

Why this construction difference matters

A model can learn to exploit SQuAD's construction artefact — high lexical overlap between question and answer-bearing sentence — as a shortcut, performing well without necessarily building the general reading-comprehension capability the benchmark is meant to proxy for. Jia & Liang (2017), Adversarial Examples for Evaluating Reading Comprehension Systems, demonstrated this directly: inserting a single adversarial, lexically-similar-but-irrelevant sentence into a SQuAD passage caused a substantial accuracy drop across the systems they tested, showing those systems were partly relying on surface pattern matching rather than robust comprehension.

Benchmark score as a proxy, not a target

Goodhart's law — a measure that becomes a target stops being a good measure — applies directly here. Optimising a model specifically against SQuAD's exact-match and F1 metrics, through architecture or training choices tuned to this benchmark, risks improving the score without improving general question-answering capability in the way the benchmark was originally meant to indicate. This is a structural concern with any fixed benchmark used for extended periods as the primary target of active optimisation, not a criticism specific to SQuAD or Natural Questions.

Key references

  • Rajpurkar, P., Zhang, J., Lopyrev, K. & Liang, P. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. arXiv:1606.05250
  • Rajpurkar, P., Jia, R. & Liang, P. (2018). Know What You Don't Know: Unanswerable Questions for SQuAD. arXiv:1806.03822
  • Kwiatkowski, T. et al. (2019). Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics 7.
  • Jia, R. & Liang, P. (2017). Adversarial Examples for Evaluating Reading Comprehension Systems. A direct demonstration of shortcut-learning in SQuAD-trained models.

Current state and open problems

Both benchmarks are now comfortably saturated by modern large language models on their original leaderboards, which itself signals their limited remaining power to discriminate between current strong and weak systems. This has driven a broader shift toward evaluating QA systems on task-specific, custom evaluation sets built from a deployment's actual question distribution, rather than relying on general-purpose public benchmarks as evidence of production readiness — the same conclusion the beginner block reaches informally, and the direct motivation for Building a retrieval eval set. A benchmark score remains useful for comparing methods under controlled, shared conditions. It is not, on its own, evidence that a system is ready for a specific real-world deployment.

What to learn next