Handling 'what about last year?'
A follow-up question like "what about last year?" means nothing to a search system on its own, so it needs rewriting into a standalone question before retrieval can work.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Conversational question rewriting turns a follow-up like "what about last year?" into a standalone question a search system can use.
Think about a phone call. A friend asks, "and what about tomorrow?" You only understand because you remember what was asked a moment earlier. Anyone listening in from the middle of the call would be lost.
Why it exists
A search system has no memory of the conversation by default. Every question gets treated as if it arrived alone, with nothing said before it.
"What was the revenue in 2023?" searches fine on its own. "What about last year?," asked right after, searches badly. The words "last year" carry no topic of their own. A search system matches them against whatever loosely mentions a year, not necessarily the right one.
The fix rewrites the follow-up into a complete, standalone question, before it reaches search. Words like "last year," "it," and "that" get replaced by the actual thing they refer to, pulled from earlier turns.
How it works
Turn 1 (user): "What was the revenue in 2023?"
Turn 2 (user): "What about last year?"
Rewritten before search:
"What about last year?"
|
v (using turn 1 for context)
"What was the revenue in 2022?"
|
v
search runs on the REWRITTEN question, not the raw oneWhere you have already seen it
- Voice assistants handling follow-ups. "Set a timer for 10 minutes." "Make it 15 instead." The assistant resolves "it" using the previous command.
- Customer support chat. "Can I return my order?" then "What about a different item?" A good bot tracks which item is now being discussed.
- Search engines with conversational mode. Typing a follow-up question after a first search increasingly gets rewritten before the second search runs.
Remember this
- A follow-up question often makes no sense without the conversation before it.
- Rewriting it into a full, standalone question has to happen before search, not after.
- Words like "it," "that," and "last year" are the usual signs a rewrite is needed.
What to learn next
- Extractive question answering — what runs after a question has been made standalone.
- Keyword vs semantic search — the search step this rewriting step protects.
- Matching a user question to an existing FAQ — another place a badly-worded question causes a bad match.
Developer — Code and libraries.
This shows a raw follow-up question retrieving the wrong document, then shows a rewritten version retrieving the right one.
Setup
pip install scikit-learnSearch before and after rewriting
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
docs = [
"Our revenue in 2023 was 65 crore rupees, up 60 percent year on year.",
"Our revenue in 2022 was 40 crore rupees, up 33 percent year on year.",
"Our office moved to a new campus in Bengaluru in March 2023.",
]
vec = TfidfVectorizer(stop_words="english").fit(docs)
doc_vecs = vec.transform(docs)
def retrieve(query, k=1):
qv = vec.transform([query])
sims = cosine_similarity(qv, doc_vecs)[0]
ranked = sims.argsort()[::-1][:k]
return [(docs[i], round(float(sims[i]), 3)) for i in ranked]
turn1 = "What was the revenue in 2023?"
turn2_raw = "What about last year?" # no context of its own
turn2_rewritten = "What was the revenue in 2022?" # rewritten using turn 1's topic
print("turn 1: ", retrieve(turn1))
print("turn 2, raw follow-up:", retrieve(turn2_raw))
print("turn 2, rewritten: ", retrieve(turn2_rewritten))turn 1: [('Our revenue in 2023 was 65 crore rupees, up 60 percent year on year.', 0.401)]
turn 2, raw follow-up: [('Our revenue in 2023 was 65 crore rupees, up 60 percent year on year.', 0.567)]
turn 2, rewritten: [('Our revenue in 2022 was 40 crore rupees, up 33 percent year on year.', 0.455)]Scores are TF-IDF cosine similarities over this exact three-document corpus. They will differ on a different or larger corpus.
Line by line
The raw follow-up retrieves the wrong document, and does so with an even higher score (0.567) than turn 1's own correct match (0.401). "What about last year?" shares no distinctive words with the 2022 document — "year" appears in all three documents — so it defaults to the most generally revenue-related match, which happens to be the 2023 one again.
The rewritten query fixes this completely, retrieving the correct 2022 document, because it now explicitly contains "2022" — a word the 2022 document actually has, and the others do not.
Nothing about the retrieval function changed between the two calls. retrieve() is identical. The entire difference in outcome comes from the text of the query itself — proof that this is a query-quality problem, not a search-algorithm problem.
Common mistakes
Trying to fix this by making the retriever "smarter" instead of fixing the query. A more powerful embedding model still cannot resolve "last year" without knowing what year the conversation was already discussing — that information is not present in the follow-up text at all, no matter how good the search algorithm is.
Rewriting only the most recent turn, ignoring longer context. "What about 2022?" then "And profit?" needs both prior turns — the year and the original topic — to become "What was the profit in 2022?" A rewrite that looks back only one turn misses this.
Rewriting after retrieval instead of before. The whole benefit comes from giving the search step a query it can actually match against. Rewriting the display text after an answer is already wrong does not help retrieval find the right source.
Assuming every follow-up needs rewriting. "What was the revenue in 2023?" followed by "What was the profit in 2023?" is already standalone — rewriting an already-complete question risks introducing an error where none existed.
Try it yourself
Add a third turn, "And the year before that?", following turn 2. Write the rewritten version by hand, using both turn 1 and turn 2's topic, then test it with retrieve().
There is no 2021 document in this small corpus, so a correctly rewritten "What was the revenue in 2021?" should retrieve weakly against all three documents — a useful, realistic case where rewriting succeeds and retrieval still correctly finds nothing strong to return.
What to learn next
- Multi-hop questions — a related but different chaining problem, across facts rather than conversational turns.
- Query expansion — a related retrieval-time technique, not a conversation-time one.
- AI agents — systems that carry conversational state across many turns like this.
Researcher — Mathematics and papers.
Formal task definition
Given a conversation history H = (q_1, a_1, ..., q_{n-1}, a_{n-1}) and a new, possibly context-dependent question q_n, conversational question rewriting learns a function rewrite(q_n, H) -> q_n', where q_n' is a standalone question, interpretable and retrievable with no access to H. This decouples conversational understanding from retrieval: the retriever only ever sees standalone questions, and all conversational reasoning is concentrated in one component.
Why this decoupling is architecturally attractive
Retrieval systems — BM25, dense bi-encoders, hybrid search — are built and evaluated against standalone queries almost universally, including every retrieval method covered earlier in this section and in the semantic search section. Rewriting lets a conversational product reuse those systems unmodified, rather than requiring every downstream retrieval component to be conversation-aware. The cost is a single point of failure: if rewrite() fails, every downstream component fails with it, since none of them has access to the original conversational context to recover.
Rule-based versus learned rewriting
The developer block's hand-written rewrite is a stand-in for what production systems generate with a language model, prompted with the conversation history and asked to produce a self-contained question. This is a well-suited task for a general-purpose LLM: it requires broad language understanding, coreference resolution, and ellipsis resolution, but produces a short, checkable output. Rule-based approaches — pattern-matching "what about X" and similar templates — cover a narrow set of common cases cheaply, and fail on any phrasing outside the templates. In practice, most production conversational search systems now use an LLM call for this step, since the accuracy gap over rule-based rewriting is substantial for anything beyond simple pronoun substitution.
Relation to coreference resolution and ellipsis
Two distinct linguistic phenomena drive this problem. Coreference is a pronoun or noun phrase referring to something mentioned earlier — "it," "that," "the company." Ellipsis is an omission the listener must fill in — "what about last year?" omits "was the revenue" entirely, not only its referent. A rewriting system needs to handle both, and ellipsis is generally the harder case, since there is no explicit word in the follow-up to anchor a coreference resolver to in the first place.
Key references
There is no single foundational paper this specific technique traces to in the way BPE traces to Sennrich et al. or attention traces to Vaswani et al. Conversational question rewriting draws on decades of coreference resolution research combined with the practical retrieval-augmented generation pattern described in Lewis et al. (2020), Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv:2005.11401), and is largely treated as an applied engineering problem, addressed with general-purpose LLMs, rather than a distinct research subfield with its own dedicated benchmark literature at the scale of SQuAD or HotpotQA.
Current state and open problems
Evaluating a rewrite in isolation is difficult: a rewrite can look reasonable to a human and still retrieve badly, or look slightly odd and retrieve perfectly, since what matters operationally is retrieval quality downstream, not surface fluency. Most teams evaluate rewriting indirectly, by measuring end-to-end retrieval or answer accuracy with and without the rewriting step, rather than scoring the rewritten text against a reference rewrite directly. There is no consensus benchmark for this task independent of the retrieval system it feeds into.
What to learn next
- Query expansion — a related technique operating at retrieval time rather than conversation time.
- Natural language inference — a building block for judging whether a rewrite preserved meaning.
- AI agents — where conversational state management is handled at a larger scale.