Question Answering

Teaching a model to say 'not in the document'

A trustworthy question-answering system needs to recognise when a passage does not contain the answer, and say so, instead of confidently pointing at the nearest wrong span.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Answerability is a model's ability to recognise when a passage does not contain the answer, and say so.

Think about a shopkeeper asked for a product they do not stock. A good shopkeeper says "we don't have that," and points you elsewhere. A bad one waves vaguely at the nearest shelf, hoping you will not check too closely.

Why it exists

The last lesson's extractive model always returns some span, even when nothing in the passage actually answers the question. Ask it something the passage never covers, and it still picks its best guess — confidently, with no warning attached.

That is a serious problem in practice. A user cannot tell a real answer from a confident wrong guess without checking the source. Checking every answer defeats the point of asking at all.

Answerability training fixes this directly. The model trains on examples with no real answer, and learns a specific "no answer" signal, instead of forcing a guess.

How it works

Passage: "The RBI raised the repo rate to 6.5 percent in
          February 2023, aiming to control inflation."

Q1: "What was the new repo rate?"
    -> "6.5 percent"                (real answer, found in passage)

Q2: "Who is the Governor of the RBI?"
    -> NOT IN THE DOCUMENT           (correctly refuses to guess)

Both answers come from the same model, reading the same passage. The difference is whether the passage actually contains what was asked.

Where you have already seen it

  • "I don't have information about that" replies from a chatbot, instead of a made-up answer.
  • Search engines showing "no results" honestly, rather than the closest unrelated match.
  • Customer support bots that escalate to a human when a question falls outside their knowledge base, instead of guessing.

Remember this

  • A useful question-answering system needs a way to say "I don't know," not only a way to answer.
  • Without it, a wrong guess and a real answer look identical to the user.
  • This is trained behaviour, learned from examples where the correct output is "no answer."

What to learn next

Developer — Code and libraries.

This uses a model trained on SQuAD 2.0, which includes unanswerable questions, and shows how it decides between answering and abstaining.

Setup

bash
pip install transformers torch

The first run downloads deepset/minilm-uncased-squad2, about 130 MB — the same model used in the previous lesson.

Answer, or abstain

no_answer.py
import torch
from transformers import AutoTokenizer, AutoModelForQuestionAnswering

name = "deepset/minilm-uncased-squad2"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForQuestionAnswering.from_pretrained(name)
model.eval()

context = (
    "The Reserve Bank of India raised the repo rate to 6.5 percent in "
    "February 2023. The move was aimed at controlling inflation."
)

def answer_or_abstain(question, context):
    inputs = tok(question, context, return_tensors="pt")
    with torch.no_grad():
        out = model(**inputs)
    start_logits, end_logits = out.start_logits[0], out.end_logits[0]

    # Position 0 is [CLS]. SQuAD 2.0 training points start AND end at [CLS]
    # when the passage has no answer, so its score is a learned "I don't know."
    no_answer_score = (start_logits[0] + end_logits[0]).item()

    best_start = torch.argmax(start_logits[1:]) + 1
    best_end = torch.argmax(end_logits[1:]) + 1
    best_span_score = (start_logits[best_start] + end_logits[best_end]).item()

    if no_answer_score > best_span_score:
        return None
    return tok.decode(inputs["input_ids"][0][best_start:best_end + 1]).strip()

questions = [
    "What was the new repo rate?",
    "Why did the RBI raise the rate?",
    "Who is the Governor of the RBI?",
    "What was India's GDP growth that year?",
]

for q in questions:
    ans = answer_or_abstain(q, context)
    print(f"Q: {q}")
    print(f"A: {ans if ans else 'NOT IN THE DOCUMENT'}")
Output
Q: What was the new repo rate?
A: 6. 5 percent
Q: Why did the RBI raise the rate?
A: controlling inflation
Q: Who is the Governor of the RBI?
A: NOT IN THE DOCUMENT
Q: What was India's GDP growth that year?
A: NOT IN THE DOCUMENT

Line by line

no_answer_score reads the model's score for pointing both start and end at [CLS], position 0. SQuAD 2.0's training data specifically labels unanswerable questions this way, so the model learns this position as a genuine "no answer" output, not an arbitrary default.

best_span_score is computed by excluding position 0, so it represents the strongest real candidate span in the passage, separate from the no-answer option.

Comparing the two scores decides the outcome. If the model is more confident in "no answer" than in its best real span, answer_or_abstain returns None. This is the entire mechanism — no separate classifier, no extra model, only one more candidate position competing in the same span-scoring system.

Note the small formatting artefact, "6. 5 percent." The tokenizer splits "6.5" into subword pieces, and naive decoding adds a space back between them. Production systems apply extra text-cleanup rules after decoding — this is a real, common rough edge, not a bug specific to this code.

Common mistakes

Treating the gap between the two scores as a calibrated confidence percentage. In practice, that gap is often tiny — sometimes a difference of 0.0001 — even on decisions that turn out correct. Use the comparison to decide yes-or-no. Do not report the gap itself as "N% confident" without separately calibrating it on your own data.

Testing only with questions that are easy calls — plainly answerable, or plainly not. Real user questions are messier — partially related, ambiguously worded, or answerable only with a stretch. Evaluate on genuinely hard borderline cases, not only easy ones.

Assuming every model exposes a no-answer score this way. This SQuAD 2.0-style mechanism is specific to models trained on unanswerable examples. A model trained only on SQuAD 1.1, which has no unanswerable questions, has no such signal and will always guess.

Deploying without any fallback message design. Returning None is a good start. What the user sees instead — "not found," an escalation link, a suggestion to rephrase — is a real product decision, not an afterthought.

Try it yourself

Add a question about something adjacent to the passage but not stated in it, such as "What was the previous repo rate before this increase?"

The passage says the rate was "raised," implying a previous value existed, but never states what it was. Predict whether the model abstains before running it, then check.

What to learn next

Researcher — Mathematics and papers.

Formal extension of the span model

SQuAD 2.0 (Rajpurkar, Jia & Liang, 2018) extends the span-prediction formulation from the previous lesson by adding a null option. The prediction is:

text
answer = span(i*, j*)   if   score(i*, j*) > score(0, 0)
       = NULL             otherwise
  • score(i, j) = p_start[i] + p_end[j], as before.
  • score(0, 0) is the model's learned score for the [CLS] position doubling as both start and end, trained to be high exactly when the passage contains no answer.

This reduces answerability detection to a single additional comparison within the existing span-scoring framework, rather than requiring a separate binary classifier — an elegant design choice, since it lets the same training signal shape both span selection and the answer/no-answer decision jointly.

Threshold calibration in practice

The raw comparison score(i*, j*) > score(0, 0) is what the developer block implements, and it is what the original paper reports results against. Production systems frequently add a tunable margin, score(i*, j*) > score(0, 0) + tau, adjusting tau to trade recall (answering more questions) against precision (avoiding wrong answers), based on a validation set matching the deployment domain rather than SQuAD's own distribution.

Why calibration transfers poorly across domains

A model fine-tuned on SQuAD 2.0's Wikipedia-derived unanswerable questions learns what "looks unanswerable" for that specific data distribution — often, superficial lexical mismatch between question and passage. A domain with different surface patterns, such as legal or medical text, can produce a systematically miscalibrated no-answer score, confidently answering when it should abstain, or the reverse. This is a known, general limitation of models fine-tuned on one benchmark and deployed on a different domain's distribution, not specific to this task.

Key references

  • Rajpurkar, P., Jia, R. & Liang, P. (2018). Know What You Don't Know: Unanswerable Questions for SQuAD. arXiv:1806.03822 — introduces SQuAD 2.0 and this null-answer formulation.
  • Rajpurkar, P., Zhang, J., Lopyrev, K. & Liang, P. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. arXiv:1606.05250 — the original, always-answerable version this extends.

Current state and open problems

Answerability remains a narrower problem than the general hallucination question facing generative systems. Span-based abstention only detects "this passage does not contain a matching span" — it cannot detect a passage that is misleading, outdated, or subtly wrong while still containing a lexically matching span. Generative RAG systems built on top of retrieval face a strictly harder version of this problem. They must judge not only "is there a matching span" but "does the retrieved context actually support a truthful answer" — closer to natural language inference than to span prediction. See Checking that an answer really came from the source for that harder, more general case.

What to learn next