Teaching a model to say 'not in the document'
A trustworthy question-answering system needs to recognise when a passage does not contain the answer, and say so, instead of confidently pointing at the nearest wrong span.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Answerability is a model's ability to recognise when a passage does not contain the answer, and say so.
Think about a shopkeeper asked for a product they do not stock. A good shopkeeper says "we don't have that," and points you elsewhere. A bad one waves vaguely at the nearest shelf, hoping you will not check too closely.
Why it exists
The last lesson's extractive model always returns some span, even when nothing in the passage actually answers the question. Ask it something the passage never covers, and it still picks its best guess — confidently, with no warning attached.
That is a serious problem in practice. A user cannot tell a real answer from a confident wrong guess without checking the source. Checking every answer defeats the point of asking at all.
Answerability training fixes this directly. The model trains on examples with no real answer, and learns a specific "no answer" signal, instead of forcing a guess.
How it works
Passage: "The RBI raised the repo rate to 6.5 percent in
February 2023, aiming to control inflation."
Q1: "What was the new repo rate?"
-> "6.5 percent" (real answer, found in passage)
Q2: "Who is the Governor of the RBI?"
-> NOT IN THE DOCUMENT (correctly refuses to guess)Both answers come from the same model, reading the same passage. The difference is whether the passage actually contains what was asked.
Where you have already seen it
- "I don't have information about that" replies from a chatbot, instead of a made-up answer.
- Search engines showing "no results" honestly, rather than the closest unrelated match.
- Customer support bots that escalate to a human when a question falls outside their knowledge base, instead of guessing.
Remember this
- A useful question-answering system needs a way to say "I don't know," not only a way to answer.
- Without it, a wrong guess and a real answer look identical to the user.
- This is trained behaviour, learned from examples where the correct output is "no answer."
What to learn next
- Extractive question answering — the base mechanism this lesson extends.
- Hallucination — the broader problem answerability training helps guard against.
- Diagnosing a wrong answer — what to check when a system answers when it should not.
Developer — Code and libraries.
This uses a model trained on SQuAD 2.0, which includes unanswerable questions, and shows how it decides between answering and abstaining.
Setup
pip install transformers torchThe first run downloads deepset/minilm-uncased-squad2, about 130 MB — the same model used in the previous lesson.
Answer, or abstain
import torch
from transformers import AutoTokenizer, AutoModelForQuestionAnswering
name = "deepset/minilm-uncased-squad2"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForQuestionAnswering.from_pretrained(name)
model.eval()
context = (
"The Reserve Bank of India raised the repo rate to 6.5 percent in "
"February 2023. The move was aimed at controlling inflation."
)
def answer_or_abstain(question, context):
inputs = tok(question, context, return_tensors="pt")
with torch.no_grad():
out = model(**inputs)
start_logits, end_logits = out.start_logits[0], out.end_logits[0]
# Position 0 is [CLS]. SQuAD 2.0 training points start AND end at [CLS]
# when the passage has no answer, so its score is a learned "I don't know."
no_answer_score = (start_logits[0] + end_logits[0]).item()
best_start = torch.argmax(start_logits[1:]) + 1
best_end = torch.argmax(end_logits[1:]) + 1
best_span_score = (start_logits[best_start] + end_logits[best_end]).item()
if no_answer_score > best_span_score:
return None
return tok.decode(inputs["input_ids"][0][best_start:best_end + 1]).strip()
questions = [
"What was the new repo rate?",
"Why did the RBI raise the rate?",
"Who is the Governor of the RBI?",
"What was India's GDP growth that year?",
]
for q in questions:
ans = answer_or_abstain(q, context)
print(f"Q: {q}")
print(f"A: {ans if ans else 'NOT IN THE DOCUMENT'}")Q: What was the new repo rate? A: 6. 5 percent Q: Why did the RBI raise the rate? A: controlling inflation Q: Who is the Governor of the RBI? A: NOT IN THE DOCUMENT Q: What was India's GDP growth that year? A: NOT IN THE DOCUMENT
Line by line
no_answer_score reads the model's score for pointing both start and end at [CLS], position 0. SQuAD 2.0's training data specifically labels unanswerable questions this way, so the model learns this position as a genuine "no answer" output, not an arbitrary default.
best_span_score is computed by excluding position 0, so it represents the strongest real candidate span in the passage, separate from the no-answer option.
Comparing the two scores decides the outcome. If the model is more confident in "no answer" than in its best real span, answer_or_abstain returns None. This is the entire mechanism — no separate classifier, no extra model, only one more candidate position competing in the same span-scoring system.
Note the small formatting artefact, "6. 5 percent." The tokenizer splits "6.5" into subword pieces, and naive decoding adds a space back between them. Production systems apply extra text-cleanup rules after decoding — this is a real, common rough edge, not a bug specific to this code.
Common mistakes
Treating the gap between the two scores as a calibrated confidence percentage. In practice, that gap is often tiny — sometimes a difference of 0.0001 — even on decisions that turn out correct. Use the comparison to decide yes-or-no. Do not report the gap itself as "N% confident" without separately calibrating it on your own data.
Testing only with questions that are easy calls — plainly answerable, or plainly not. Real user questions are messier — partially related, ambiguously worded, or answerable only with a stretch. Evaluate on genuinely hard borderline cases, not only easy ones.
Assuming every model exposes a no-answer score this way. This SQuAD 2.0-style mechanism is specific to models trained on unanswerable examples. A model trained only on SQuAD 1.1, which has no unanswerable questions, has no such signal and will always guess.
Deploying without any fallback message design. Returning None is a good start. What the user sees instead — "not found," an escalation link, a suggestion to rephrase — is a real product decision, not an afterthought.
Try it yourself
Add a question about something adjacent to the passage but not stated in it, such as "What was the previous repo rate before this increase?"
The passage says the rate was "raised," implying a previous value existed, but never states what it was. Predict whether the model abstains before running it, then check.
What to learn next
- Extractive question answering — the underlying span-prediction mechanism.
- Checking that an answer really came from the source — a complementary check, run after an answer is given.
- Exact match and F1 for QA — how SQuAD 2.0 scores no-answer predictions.
Researcher — Mathematics and papers.
Formal extension of the span model
SQuAD 2.0 (Rajpurkar, Jia & Liang, 2018) extends the span-prediction formulation from the previous lesson by adding a null option. The prediction is:
answer = span(i*, j*) if score(i*, j*) > score(0, 0)
= NULL otherwisescore(i, j) = p_start[i] + p_end[j], as before.score(0, 0)is the model's learned score for the[CLS]position doubling as both start and end, trained to be high exactly when the passage contains no answer.
This reduces answerability detection to a single additional comparison within the existing span-scoring framework, rather than requiring a separate binary classifier — an elegant design choice, since it lets the same training signal shape both span selection and the answer/no-answer decision jointly.
Threshold calibration in practice
The raw comparison score(i*, j*) > score(0, 0) is what the developer block implements, and it is what the original paper reports results against. Production systems frequently add a tunable margin, score(i*, j*) > score(0, 0) + tau, adjusting tau to trade recall (answering more questions) against precision (avoiding wrong answers), based on a validation set matching the deployment domain rather than SQuAD's own distribution.
Why calibration transfers poorly across domains
A model fine-tuned on SQuAD 2.0's Wikipedia-derived unanswerable questions learns what "looks unanswerable" for that specific data distribution — often, superficial lexical mismatch between question and passage. A domain with different surface patterns, such as legal or medical text, can produce a systematically miscalibrated no-answer score, confidently answering when it should abstain, or the reverse. This is a known, general limitation of models fine-tuned on one benchmark and deployed on a different domain's distribution, not specific to this task.
Key references
- Rajpurkar, P., Jia, R. & Liang, P. (2018). Know What You Don't Know: Unanswerable Questions for SQuAD. arXiv:1806.03822 — introduces SQuAD 2.0 and this null-answer formulation.
- Rajpurkar, P., Zhang, J., Lopyrev, K. & Liang, P. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. arXiv:1606.05250 — the original, always-answerable version this extends.
Current state and open problems
Answerability remains a narrower problem than the general hallucination question facing generative systems. Span-based abstention only detects "this passage does not contain a matching span" — it cannot detect a passage that is misleading, outdated, or subtly wrong while still containing a lexically matching span. Generative RAG systems built on top of retrieval face a strictly harder version of this problem. They must judge not only "is there a matching span" but "does the retrieved context actually support a truthful answer" — closer to natural language inference than to span prediction. See Checking that an answer really came from the source for that harder, more general case.
What to learn next
- Extractive question answering — the span-prediction base this extends.
- Natural language inference — the more general "does this support that" problem.
- Checking that an answer really came from the source — answerability's harder, generative-era cousin.