Feature and Data Pipelines in Production

Re-embedding and reindexing

Re-embedding means recomputing every stored vector with a new model, because vectors from two different embedding models are not comparable, even for the exact same text.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Re-embedding means recomputing every stored vector with a new model, because vectors from two different models are not comparable.

The analogy you have already lived

Imagine your school upgrades its ID card camera. Every new photo is taken with the new camera, in a new format. Old photo IDs, taken with the old camera, cannot be swapped one-for-one into the new system — they need to be reshot.

An embedding is like a photo: a fixed-size list of numbers a model produces for a piece of text or an image. A new embedding model is a new camera. Its numbers are not compatible with the old camera's numbers, even for the exact same thing.

Why it exists

Search and recommendation systems compare embeddings using distance. How close two vectors sit decides how similar the system thinks two things are. That only works if every vector was produced by the same model.

When a better embedding model comes out, it is tempting to switch the search system over right away. For new queries only. But every document already stored was embedded with the OLD model. Comparing a new-model query against old-model documents compares two things measured in different, incompatible units.

How it works

model v1 embeds "python tutorial"  -> [0.2, -0.5, 0.8, ...]
model v2 embeds "python tutorial"  -> [-0.3, 0.9, 0.1, ...]     <- totally different numbers

same sentence, different model  =  different, incomparable vectors

query embedded with v2  vs  documents still stored as v1
   -> comparing numbers that were never meant to be compared
   -> search results become meaningless, even for an exact match

The fix: re-embed every stored document with the new model, so queries and documents are measured in the same units again.

A real example you have seen

A photo search app that upgrades its image-recognition model. Every photo you ever uploaded needs to be re-processed with the new model. Only then do new-style searches ("find photos with a dog") work correctly on your old photos.

Remember this

  • An embedding is only comparable to other embeddings from the same model version.
  • Switching embedding models without re-embedding stored data makes search results quietly meaningless, not visibly broken.
  • Re-embedding means running every stored item through the new model, not only new items going forward.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Proving the incompatibility, with identical text

reembed_bug.py
import hashlib
import numpy as np

# Two "embedding models" stand in as two fixed random projections into a
# small space. This is not a real embedding model -- it exists only to make
# one fact concrete: v1 and v2 place the same sentence in different,
# unrelated coordinates.
def fake_embed(text, model_seed, dim=16):
    # A hash of the bytes, not Python's built-in hash() -- that one is
    # randomised per process on purpose, which would make this "deterministic"
    # function give a different answer on every run.
    key = f"{text}|{model_seed}".encode()
    h = int(hashlib.sha256(key).hexdigest(), 16) % (10**6)
    v = np.random.RandomState(h).normal(size=dim)
    return v / np.linalg.norm(v)

def cosine(a, b):
    return float(np.dot(a, b))

MODEL_V1, MODEL_V2 = 1, 2
doc_text = "python programming tutorial"

# The document was embedded and stored back when the service ran model v1.
doc_vector_v1 = fake_embed(doc_text, MODEL_V1)

# The service was upgraded to model v2. A query arrives with the EXACT SAME
# TEXT as the stored document -- the best possible match that could exist.
query_vector_v2 = fake_embed(doc_text, MODEL_V2)

# What a correct system does: re-embed the document with v2 too, so query
# and document live in the same space again.
doc_vector_v2 = fake_embed(doc_text, MODEL_V2)

mismatched = cosine(query_vector_v2, doc_vector_v1)   # v2 query vs OLD v1 doc
matched    = cosine(query_vector_v2, doc_vector_v2)   # v2 query vs RE-EMBEDDED doc

print("query text and document text are IDENTICAL:", repr(doc_text))
print(f"similarity, v2 query vs the OLD v1 vector       : {mismatched:.3f}  (should be 1.0 -- it is not)")
print(f"similarity, v2 query vs a RE-EMBEDDED v2 vector : {matched:.3f}")
Output
query text and document text are IDENTICAL: 'python programming tutorial'
similarity, v2 query vs the OLD v1 vector       : 0.052  (should be 1.0 -- it is not)
similarity, v2 query vs a RE-EMBEDDED v2 vector : 1.000

fake_embed is not a real embedding model — it is a deliberately simple stand-in so the point stays visible: identical text still needs to go through the same model to be comparable. A real model (Sentence-BERT, OpenAI's embedding API, a HuggingFace model from the HuggingFace stack) has the identical property, for the identical reason.

Line-by-line walkthrough

fake_embed is deterministic per (text, model_seed) pair — the same text with the same model always gives the same vector, but a different model gives an unrelated one. That mirrors a real embedding model's behaviour exactly, minus the actual language understanding.

The result is the entire lesson in two numbers: 0.052 where a perfect match should score 1.000. The query and the document say the exact same thing, and the mismatched-model comparison cannot tell.

Common mistakes

Shipping a new embedding model for queries only, "to start seeing improvements sooner". This is the bug demonstrated above, deployed. Every existing document becomes silently incomparable to every new query.

Re-embedding without re-indexing. New vectors are useless until the search index built on top of them (see rebuilding an index with no downtime) is also rebuilt to reflect them.

Assuming two versions of "the same" model are compatible. A fine-tuned checkpoint, a different pooling strategy, even a different normalisation step can break comparability, not only a full model swap. Treat any change to the embedding pipeline as a new, incompatible version.

No plan for how long re-embedding will take. For millions of documents through a large model, this can take hours to days. Plan the cutover, do not discover the timeline mid-incident.

Try it yourself

Add a second document with different text and confirm its similarity to the query is also affected by which model produced it — the bug is not specific to the identical-text case, that case is only the clearest way to see it.

What to learn next

Researcher — Mathematics and papers.

Why embedding spaces are not aligned

Two independently trained embedding models learn different coordinate systems, related by no fixed transformation in general. Even models trained on similar objectives can differ by an arbitrary rotation, reflection or nonlinear warp of the space — there is no guarantee, and usually no truth, that $\text{embed}_1(x) \approx \text{embed}_2(x)$ for any $x$.

Some techniques attempt to align two embedding spaces after the fact — Procrustes alignment for a linear rotation-based mapping, or a learned linear probe — but these are approximate and add error; full re-embedding remains the only exact fix.

Cost of a full re-embed

For $n$ stored documents through a model with per-item cost $c$, a full re-embed costs $O(n \times c)$, typically dominated by $c$ when $c$ involves a forward pass through a neural network. For a corpus of $10^7$ documents through a mid-sized transformer, this is commonly hours on a GPU cluster — enough that the operation is planned and batched like a training job, not run casually.

Staged migration strategies

  • Dual-write, dual-index — write both v1 and v2 embeddings during the transition, serve from v1 until v2's index is fully built and validated, then cut over. Doubles storage temporarily, costs nothing in risk.
  • Shadow re-embedding with offline evaluation — build the v2 index fully offline, run retrieval-quality metrics (recall@k against a labelled set) against both indexes before switching any live traffic.
  • Partial re-embedding by priority — re-embed the most-queried or most-recent documents first if a full re-embed cannot complete before a deadline, accepting degraded recall on the long tail during the transition.

The cutover itself is a specific instance of rebuilding an index with no downtime: the new, fully-built v2 index is swapped in atomically, never built in place over the live one.

Papers and systems

  • Reimers and Gurevych, Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, EMNLP 2019 — the model family most production text-embedding pipelines are built on or descended from.
  • Schönemann, A Generalized Solution of the Orthogonal Procrustes Problem, Psychometrika 1966 — the classical alignment technique referenced above, still used as an approximate stopgap.
  • Johnson, Douze and Jégou, Billion-scale similarity search with GPUs (FAISS), IEEE Big Data 2019 — the indexing layer that has to be rebuilt alongside the vectors themselves.

What to learn next