Question Answering

Matching a user question to an existing FAQ

FAQ matching finds the pre-written question closest in meaning to what a user actually typed, so a trusted, already-approved answer can be reused instead of generated fresh.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

FAQ matching finds the existing, pre-written question closest in meaning to what a user typed.

Think about calling a helpline and pressing through recorded questions until one matches yours. You are not getting a fresh answer made up on the spot. You are being routed to the closest existing one.

Why it exists

Many businesses already have a list of frequently asked questions, each with a carefully written, approved answer. A user rarely types a question in exactly those words.

"I forgot my login password" and "How do I reset my password?" mean the same thing, in different words. Matching on exact text would miss this pairing entirely, since not one word overlaps between "forgot," "login," and "reset," "password."

FAQ matching solves this with meaning, not exact wording. It compares the user's question to every FAQ question by meaning, and returns the closest one's already-written answer.

How it works

FAQ list:
  Q: "How do I reset my password?"       -> A: "Go to Settings > Security..."
  Q: "What is your refund policy?"       -> A: "Refunds are processed..."
  Q: "How long does shipping take?"      -> A: "Standard shipping takes..."

User types: "I forgot my login password"

  Compare by meaning against every FAQ question:
    vs "reset my password"     -> very close match
    vs "refund policy"         -> not close
    vs "shipping take"         -> not close

  Best match: "How do I reset my password?"
  Return its existing, approved answer.

Where you have already seen it

  • Help centre search boxes. Typing a question into a support site's search often surfaces the closest matching help article, worded nothing like your search.
  • Chatbot widgets on e-commerce sites. A quick question gets routed to a pre-written answer, not a freshly generated one.
  • In-app help search, where typing "can't log in" correctly surfaces an article titled "Trouble Signing In."

Remember this

  • FAQ matching compares meaning, not exact words, between a user's question and existing FAQ questions.
  • The reused answer is pre-written and pre-approved — nothing is generated fresh.
  • This works only as well as the FAQ list covers what users actually ask.

What to learn next

Developer — Code and libraries.

This matches a user question to the closest FAQ entry using sentence embeddings, and returns that entry's pre-written answer.

Setup

bash
pip install transformers torch

The first run downloads sentence-transformers/all-MiniLM-L6-v2, about 90 MB.

Matching by meaning, not by keyword

faq_match.py
import torch
from transformers import AutoTokenizer, AutoModel

tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()

def embed(texts):
    enc = tok(texts, padding=True, truncation=True, return_tensors="pt")
    with torch.no_grad():
        out = model(**enc)
    mask = enc["attention_mask"].unsqueeze(-1).float()
    pooled = (out.last_hidden_state * mask).sum(1) / mask.sum(1)
    return torch.nn.functional.normalize(pooled, dim=1)

faqs = [
    "How do I reset my password?",
    "What is your refund policy?",
    "How long does shipping take?",
    "Can I change my delivery address after ordering?",
]
answers = [
    "Go to Settings > Security and click 'Reset password'.",
    "Refunds are processed within 7 business days of the return.",
    "Standard shipping takes 3-5 business days within India.",
    "Yes, within 1 hour of placing the order, from your Orders page.",
]

faq_vecs = embed(faqs)

def match(user_question, k=1):
    qv = embed([user_question])
    sims = (qv @ faq_vecs.T)[0]
    best = torch.argsort(sims, descending=True)[:k]
    return [(faqs[i], answers[i], round(sims[i].item(), 3)) for i in best]

for q in [
    "I forgot my login password",
    "Can I get my money back?",
    "My order hasn't shipped, when will it arrive?",
]:
    print("user asked:", q)
    for faq, ans, score in match(q):
        print(f"  matched FAQ ({score}): {faq}")
        print(f"  answer: {ans}")
Output
user asked: I forgot my login password
  matched FAQ (0.784): How do I reset my password?
  answer: Go to Settings > Security and click 'Reset password'.
user asked: Can I get my money back?
  matched FAQ (0.518): What is your refund policy?
  answer: Refunds are processed within 7 business days of the return.
user asked: My order hasn't shipped, when will it arrive?
  matched FAQ (0.745): How long does shipping take?
  answer: Standard shipping takes 3-5 business days within India.

Line by line

embed() reuses the same mean-pooling pattern seen throughout this batch of lessons — one normalised vector per text, so qv @ faq_vecs.T computes cosine similarity against every FAQ question in a single matrix multiply.

None of the three test questions share more than one or two exact words with their correct FAQ match. "Money back" matched "refund policy" with no word overlap at all. This is the entire value of embedding-based matching over exact or keyword search.

The refund match scored notably lower, 0.518, than the other two. "Money back" and "refund policy" are related but phrased quite differently — a reasonable case where a production system might show this as a suggestion rather than an automatic answer, using a score threshold. See the next lesson for what to check once an answer is given.

Common mistakes

Returning a match with no minimum score threshold. If nothing in the FAQ list is genuinely relevant, the function still returns whatever scores highest, even at a low, unreliable score. Set a minimum threshold, below which the system says "I could not find a matching answer" instead of guessing.

Growing the FAQ list without checking for near-duplicate questions. Two very similarly worded FAQ entries compete for the same user questions, splitting the score between them, and a small rewording can flip which one wins unpredictably.

Re-embedding the FAQ list on every request. faq_vecs should be computed once and cached, or stored in a vector database, not recomputed every time a user asks something. This code recomputes it once, up front, which is correct for a small demo but worth calling out.

Assuming the top match is always right. For ambiguous questions, the top match and the second-best match can be close in score. Showing the top 2 or 3 matches, not only the single best one, is a common and often better UX for uncertain cases.

Try it yourself

Add a fifth FAQ entry that overlaps closely with an existing one — for example, "How do I change my password?" alongside the existing "How do I reset my password?" — and re-run the password-related query.

Watch which one wins, and by how much. Near-duplicate FAQ entries are a genuinely common maintenance problem in real FAQ systems, and this shows exactly why.

What to learn next

Researcher — Mathematics and papers.

FAQ matching as retrieval with a closed answer set

FAQ matching is a constrained special case of the retrieval problem covered throughout the semantic search and question-answering sections: the "documents" being searched are a small, closed, curated set of question-answer pairs, rather than an open corpus of arbitrary passages. This constraint is what makes the problem tractable with a simple bi-encoder similarity search, as in the developer block, rather than needing the full retrieve-then-read machinery of open-domain QA.

Bi-encoder choice and its limits

A bi-encoder embeds the user question and each FAQ question independently, then compares via a cheap dot product — the pattern used throughout this lesson set. This scales to thousands of FAQ entries with sub-millisecond query latency, since FAQ embeddings are precomputed once. See Bi-encoders vs cross-encoders for the accuracy cost of this speed: a cross-encoder, scoring the user question and each FAQ question jointly, generally ranks more accurately but must run once per FAQ entry per query, making it impractical at large FAQ-list scale without a bi-encoder pre-filtering step first.

Score calibration across question pairs

The similarity scores in the developer block's output are not comparable across different user questions in any absolute sense — a 0.518 score for one query and a 0.518 score for a different query do not represent equal underlying match quality, since embedding-space geometry is not uniformly calibrated across arbitrary sentence pairs. A minimum-score threshold, as recommended in the developer block's common mistakes, should be tuned empirically against labelled examples from your specific FAQ list and expected query distribution, not set from a general rule of thumb.

Relation to intent classification

FAQ matching overlaps substantially with intent classification, a longstanding task in dialogue systems research: both map a free-text user utterance to one of a fixed set of known categories. The distinction is largely architectural. Intent classification is typically trained as a supervised classifier over a fixed label set, requiring labelled training examples per intent, while FAQ matching via embedding similarity requires no training data beyond the FAQ questions themselves and generalises to a newly added FAQ entry with zero additional training — a meaningful operational advantage when a FAQ list changes frequently.

Key references

FAQ matching, as an embedding-similarity application, builds directly on the sentence embedding literature covered in Sentence transformers and the dense retrieval work of Karpukhin et al. (2020), Dense Passage Retrieval for Open-Domain Question Answering (arXiv:2004.04906). It is not typically treated as a distinct research subfield with its own benchmark literature, since it is a direct, practitioner-facing application of general-purpose semantic search.

Current state and open problems

The dominant practical failure mode is FAQ list drift: new products, policies, or edge cases generate user questions with no genuinely close FAQ match, and a naive system still returns its closest available entry rather than surfacing the gap. Detecting "this question has no good match in our FAQ list" reliably is the same answerability problem covered earlier in this section, applied to a closed answer set instead of an open passage. Monitoring low-confidence matches in production, and routing them for human review to either answer directly or seed a new FAQ entry, remains a manual, operational process rather than a solved technical one.

What to learn next