Monitoring Models in Production
Monitoring a vector index
A vector index needs its own kind of monitoring, watching whether new questions still have a close match inside it, separate from monitoring the model that built it.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A vector index needs its own monitoring, checking whether new questions still have a genuinely close match stored inside it.
An embedding is a list of numbers standing in for the meaning of a piece of text, so similar meanings become similar-looking lists.
The analogy you have already lived
A library shelves books by topic, so that similar books sit near each other. Someone wanting a cookbook walks straight to the cooking shelf, not the whole building.
That shelving only works for topics the library actually stocked. A brand-new subject the library never bought a single book on has no shelf to walk to, near or far.
A vector index is that library: a large, searchable collection of embeddings, built so a new question can quickly find its closest matches. New questions outside what it stocked have nothing close to find.
Why it exists
Semantic search and question-answering tools store a large set of documents as embeddings, once. A new question also becomes an embedding, and the system looks up whichever stored documents sit closest to it.
That stored set does not update itself. New products launch. New questions get asked. The index keeps offering only what it held on the day it was built, unless someone deliberately refreshes it.
How it works
week 1: questions about billing, shipping
-> the index HAS billing and shipping docs
-> every question finds a close match
week 6: a new feature launched, questions shift toward it
-> the index does NOT have docs about the new feature
-> those questions find nothing close, silentlyNothing crashes. The system returns its best available match every time, even when its best available match is genuinely far away.
A real example you have seen
A help-chat bot trained on a product's older documentation keeps answering fluently after a big new feature ships. It quietly stretches old, unrelated answers to cover new questions, because it was never told there was nothing good to offer.
The honest part
A vector index failing does not look like an error message. It looks like a slightly-off answer, delivered with the same confidence as a good one.
That is the hardest part to accept. The system has no built-in way to say "I have nothing close to this". Watching match quality yourself is the only way to know.
Remember this
- An embedding turns meaning into numbers; a vector index stores many of them for fast lookup.
- An index only covers what it was built from, and never updates itself.
- A poor match is returned exactly as confidently as a good one, with no warning.
What to learn next
- Feedback loops in production — how a stale index can reinforce its own blind spots.
- Tracing an LLM application — watching a retrieval step like this one inside a larger LLM pipeline.
- Measuring data drift — the same distributional thinking, applied to raw features instead of embeddings.
Developer — Code and libraries.
Setup
pip install numpyWatching match quality drop as new topics appear
A brute-force cosine-similarity index over sixteen-dimensional vectors, small enough to run entirely in memory. No external vector database needed to see the effect.
import numpy as np
rng = np.random.RandomState(0)
dim = 16
def unit_vectors(center, n, noise=0.3):
v = center + rng.normal(0, noise, (n, dim))
return v / np.linalg.norm(v, axis=1, keepdims=True)
center_a = rng.normal(0, 1, dim) # topic: "billing questions"
center_b = rng.normal(0, 1, dim) # topic: "shipping questions"
center_c = rng.normal(0, 1, dim) # topic: "a brand-new feature launched last week"
# The index was built when the product only had billing and shipping docs.
index_vectors = np.vstack([unit_vectors(center_a, 150), unit_vectors(center_b, 150)])
def best_match_similarity(queries, index):
sims = queries @ index.T # cosine similarity, since all vectors are unit length
return sims.max(axis=1)
# Week 1: support questions still match the topics the index was built for.
queries_week1 = np.vstack([unit_vectors(center_a, 40), unit_vectors(center_b, 40)])
# Week 6: the new feature launched, and a growing share of questions are
# about it -- a topic that was never added to the index.
queries_week6 = np.vstack([
unit_vectors(center_a, 15),
unit_vectors(center_b, 15),
unit_vectors(center_c, 30),
])
sim1 = best_match_similarity(queries_week1, index_vectors)
sim6 = best_match_similarity(queries_week6, index_vectors)
print(f"week 1 -- mean best-match similarity: {sim1.mean():.3f} min: {sim1.min():.3f}")
print(f"week 6 -- mean best-match similarity: {sim6.mean():.3f} min: {sim6.min():.3f}")
weak = (sim6 < 0.3).mean()
print(f"week 6 -- fraction of queries with NO good match (similarity < 0.3): {weak:.1%}")week 1 -- mean best-match similarity: 0.979 min: 0.952 week 6 -- mean best-match similarity: 0.563 min: 0.030 week 6 -- fraction of queries with NO good match (similarity < 0.3): 50.0%
Exact output from this seeded script. Half of week 6's queries had no genuinely close match in the index, and the index returned its "best" answer for every one of them regardless.
Line-by-line walkthrough
unit_vectors scatters points around a centre and normalises them, standing in for real text embeddings, which are also typically compared by cosine similarity.
index_vectors deliberately only contains topics A and B. Topic C never went in, standing in for a feature that launched after the index was last built.
best_match_similarity computes, for every query, the similarity to its single closest indexed vector. It never returns "nothing found" — there is always a best match, however weak.
The weak fraction is the number worth alerting on. It directly counts questions the index had nothing genuinely relevant for.
Common mistakes
Monitoring only whether the search returns results. It always will. Monitor the actual similarity score of the top result, not whether a result exists.
Never re-embedding the index after documents change. If the embedding model itself is upgraded but the stored vectors are not recomputed, queries and documents are compared in two different, incompatible number spaces.
Treating a low similarity score as acceptable because the system still answered. A confident-sounding answer built on a weak match is often worse than a system that says it does not know. See redacting personal data from LLM logs and tracing an LLM application for how this connects to a full retrieval-augmented pipeline.
Sizing an alert threshold from a handful of manual checks. Log the top-match similarity for every real query, the way prediction logging does for a model, and set the threshold from that real distribution.
Try it yourself
Add index_vectors = np.vstack([index_vectors, unit_vectors(center_c, 150)]) after week 6's queries are generated, simulating a refreshed index. Recompute sim6 against the new index_vectors, and check how much the mean similarity recovers.
What to learn next
- Feedback loops in production — how a stale index can reinforce its own blind spots.
- Tracing an LLM application — watching a retrieval step like this one inside a larger LLM pipeline.
- Measuring data drift — the same distributional thinking, applied to raw features instead of embeddings.
Researcher — Mathematics and papers.
What "index staleness" actually covers
Three distinct failure modes hide under one name, and they need different fixes:
- Coverage gaps — the topic was never indexed at all, demonstrated above. Fixed by adding documents.
- Embedding-space drift — the embedding model itself was upgraded or fine-tuned, and old vectors were never recomputed under the new model, so old and new vectors are no longer comparable by distance at all. Fixed only by full re-embedding.
- Approximate-search degradation — for large indexes, exact search is replaced by an approximate nearest-neighbour structure (HNSW, IVF), whose recall against true nearest neighbours can degrade as the index grows past the parameters it was tuned for.
Measuring approximate-search recall directly
For a sample of queries, compute recall@k against exact brute-force search as ground truth:
$$\text{recall@}k = \frac{1}{|Q|} \sum_{q \in Q} \frac{|\text{ANN}_k(q) \cap \text{Exact}_k(q)|}{k}$$
Where $\text{ANN}_k(q)$ is the top-$k$ result set from the approximate index for query $q$, and $\text{Exact}_k(q)$ is the true top-$k$ by brute force over the same vectors. This isolates approximation error from coverage error — a healthy index can still show low recall@k if its HNSW parameters (ef_search, graph connectivity) are under-tuned for its current size.
Detecting coverage gaps at scale
The developer example computes best-match similarity per query, which is exactly the score used by out-of-distribution query detection in production retrieval systems. At scale, this is tracked as a running histogram of top-1 similarity scores, with alerting on the histogram's low tail growing — directly analogous to the PSI-based drift monitoring from measuring data drift, applied to similarity scores instead of raw features.
Complexity and cost
Brute-force search is $O(nd)$ per query for $n$ indexed vectors of dimension $d$, as used in the developer example. HNSW search is $O(\log n)$ expected per query after $O(n \log n)$ index construction (Malkov and Yashunin, 2018), the standard trade made once brute force stops being fast enough — typically past a few hundred thousand vectors on a single machine.
Papers and systems
- Malkov and Yashunin, Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs, TPAMI 2018 — arxiv.org/abs/1603.09320
- Johnson, Douze and Jégou, Billion-Scale Similarity Search with GPUs, IEEE Big Data 2019 — the FAISS paper.
- Reimers and Gurevych, Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, EMNLP 2019 — the widely-used approach behind most modern text embeddings this lesson's vectors stand in for.
What to learn next
- Feedback loops in production — how a stale index can reinforce its own blind spots.
- Tracing an LLM application — watching a retrieval step like this one inside a larger LLM pipeline.
- Measuring data drift — the same distributional thinking, applied to raw features instead of embeddings.