Question Answering

Extractive question answering

Extractive QA finds the exact words that answer a question inside a given passage, instead of writing a new sentence from scratch.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Extractive question answering finds the exact words that answer a question, inside a passage it was already given.

Think about an open-book exam. The question asks for a date, and you underline that exact date on the page in front of you. You are not writing your own sentence. You are pointing at words already sitting on the page.

Why it exists

Early question-answering systems generated a fresh sentence for every answer, word by word. That is powerful. It is also how a model invents facts that sound right and are not — see Hallucination.

Extractive QA sidesteps that risk in one specific setting: when the answer is known to sit somewhere inside a given passage. Rather than generating text, the model only has to point at a start position and an end position within that passage.

Because every answer is a real quote from the source, extractive QA cannot invent a fact that is not on the page. It can still pick the wrong span, or miss the right one, but it cannot make something up out of thin air.

How it works

Question: "Where is the Eiffel Tower located?"
Passage:  "The Eiffel Tower is located in Paris, France. It was
           built in 1889."

Model finds a start position and an end position:

  "The Eiffel Tower is located in [ Paris, France ]. It was built..."
                                    ^start          ^end

Answer: "Paris, France"

The answer is a direct quote, lifted from an exact span of the passage.

Where you have already seen it

  • "People also ask" boxes on Google. A short, quoted answer under a question, pulled from a specific web page.
  • Document search tools that highlight the answer. Some PDF search tools jump straight to, and highlight, the exact sentence answering your query.
  • Voice assistants answering from a single web result, quoting a specific fact rather than composing a new sentence.

Remember this

  • Extractive QA finds a span of text inside a passage. It does not write new sentences.
  • The model needs both a question and a passage that actually contains the answer.
  • Because the answer is a direct quote, it cannot invent a fact that is not in the source.

What to learn next

Developer — Code and libraries.

This loads a small BERT-family model already fine-tuned for extractive QA, and finds the answer span directly from its start and end logits.

Setup

bash
pip install transformers torch

The first run downloads deepset/minilm-uncased-squad2, about 130 MB.

Finding the answer span

extractive_qa.py
import torch
from transformers import AutoTokenizer, AutoModelForQuestionAnswering

name = "deepset/minilm-uncased-squad2"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForQuestionAnswering.from_pretrained(name)
model.eval()

context = (
    "The Eiffel Tower is located in Paris, France. It was built in 1889 "
    "and stands 330 metres tall."
)
question = "Where is the Eiffel Tower located?"

inputs = tok(question, context, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

start = torch.argmax(outputs.start_logits)
end = torch.argmax(outputs.end_logits) + 1
answer = tok.decode(inputs["input_ids"][0][start:end]).strip()

print("question:", question)
print("answer:  ", answer)
print("start logit score:", round(outputs.start_logits[0, start].item(), 2))
print("end logit score:  ", round(outputs.end_logits[0, end - 1].item(), 2))
Output
question: Where is the Eiffel Tower located?
answer:   paris, france
start logit score: 6.75
end logit score:   6.51

The model is uncased, so the answer prints in lowercase. Logit scores are model-specific numbers, not probabilities — they matter only relative to each other, not as a percentage.

Line by line

tok(question, context, return_tensors="pt") packs both texts into one input, following the format the model was trained on: question first, then context, separated by a special token the tokenizer inserts automatically.

outputs.start_logits and outputs.end_logits are two full-length vectors, one score per input token, representing how likely that token is to start, or end, the answer.

torch.argmax(...) picks the single highest-scoring position for each. The + 1 on end accounts for Python slicing being exclusive at the end — without it, the last token of the answer would be cut off.

tok.decode(...) turns the token IDs back into text. This is the extractive step made concrete: the answer is a literal slice of inputs["input_ids"], not new text generated by the model.

Common mistakes

Passing a question with no matching context. If the passage genuinely does not contain the answer, the model still returns some span — usually a low-confidence, wrong one. Handling this properly is the entire subject of the next lesson.

Forgetting the passage has a length limit. This model, like most BERT-family models, truncates input past a fixed number of tokens — commonly 512. A long passage needs chunking, tying this lesson directly back to the earlier section on chunking long documents.

Mixing up start_logits and end_logits argmax without checking start <= end. On adversarial or unusual inputs, the raw argmax of each can occasionally produce end before start. Production code checks for this and falls back to the best valid span.

Assuming the printed logit score is a confidence percentage. It is a raw, unnormalised number. Comparing it against another logit from the same model, on the same input, is meaningful. Treating it as "68% confident" is not — softmax normalisation would be needed for that, and even then, it would not calibrate cleanly to true accuracy without deliberate calibration.

Try it yourself

Change the question to "How tall is the Eiffel Tower?" and re-run.

The model should now point at "330 metres" instead of "paris, france" — a different span, extracted from the same passage, because the question changed which fact is being asked for. The passage was never re-read differently; only the target span moved.

What to learn next

Researcher — Mathematics and papers.

Formulation

Extractive QA is framed as span prediction over a tokenized sequence. Given a question Q and context C, tokenized jointly as X = [CLS] Q [SEP] C [SEP] of length n, the model produces two vectors p_start, p_end in R^n, typically from two independent linear heads over the final hidden states. The predicted span is:

text
(i*, j*) = argmax over valid (i, j), i <= j   of   p_start[i] + p_end[j]
  • p_start[i] is the start logit at position i.
  • p_end[j] is the end logit at position j.
  • The constraint i <= j restricts to non-degenerate spans; production implementations also cap j - i to a maximum answer length.

Training objective

Training minimises cross-entropy loss independently on the gold start position and gold end position, treating span prediction as two separate token-classification problems over the same sequence, not a single joint distribution over (i, j) pairs. This independence assumption is a modelling simplification: the true joint distribution over spans is not the product of independent start and end marginals, but the approximation trains efficiently and works well in practice.

Why BERT-family, encoder-only models

Extractive QA needs bidirectional context — the correct start position can depend on words that appear after it in the passage. Encoder-only models, trained with a masked language modelling objective rather than left-to-right generation, provide this natively; see Encoder vs decoder models. This is why extractive QA baselines are almost universally BERT-family, not decoder-only generative models, even in an era where generative models dominate most other NLP tasks.

Complexity

For a sequence of length n and model width d, a forward pass costs the same as any transformer encoder pass: dominated by attention's O(n^2 * d) and the feedforward blocks' O(n * d^2), as detailed in Attention. The span-selection step itself is O(n^2) in the worst case, if scoring every valid (i, j) pair, but is typically implemented as two independent O(n) argmax operations, as in the developer block, trading joint optimality for speed.

Key references

  • Rajpurkar, P., Zhang, J., Lopyrev, K. & Liang, P. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. arXiv:1606.05250 — the benchmark that established this task formulation at scale.
  • Devlin, J. et al. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. The original architecture this task is built on.

Current state and open problems

Extractive QA performs strongly on benchmark-style, single-passage, single-hop questions, and degrades sharply outside that setting — multi-document reasoning, arithmetic over extracted numbers, and questions with no single contiguous answer span all fall outside what the span-prediction formulation can represent. Generative QA, where a model composes a free-text answer rather than extracting a span, has become the more common production pattern for exactly these harder cases, at the cost of reintroducing the hallucination risk extractive QA was designed to avoid. Most modern RAG systems use extractive QA as one signal among several, not as the final answer format shown to a user.

What to learn next