Feature and Data Pipelines in Production
Rebuilding an index with no downtime
A search or vector index can be rebuilt from scratch in the background and swapped in with one atomic pointer change, so readers never see a half-built or missing index.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A safe index rebuild builds the new index fully in the background, then swaps it in with one instant change. It is never rebuilt in place over the live one.
The analogy you have already lived
A shop repainting its signboard does not take the old sign down first. It does not paint the new one in its place while customers watch a blank wall. It paints the new sign somewhere else, finished and ready, then swaps it in overnight.
Rebuilding a search index the naive way is like painting over the live sign. Readers see a half-finished, confusing result the entire time.
Why it exists
An index — the structure that makes search or lookup fast — sometimes has to be rebuilt entirely. This happens after re-embedding, after a schema change, or on a routine schedule to clear out staleness.
Rebuilding it in place means the index is incomplete for however long the rebuild takes. That can be minutes for a small one, hours for a large one. Every read during that window either fails, or returns a wrong, partial answer.
How it works
step 1: keep serving reads from index_v1 (untouched, fully working)
step 2: build index_v2 in the background, fully, taking as long as it needs
step 3: the ONLY visible moment: flip a pointer from "v1" to "v2"
step 4: old index_v1 can now be safely deletedThe entire risk is compressed into one tiny, instant step: changing what the pointer says. Everything slow and error-prone happens before that point is ever reached, invisible to readers.
A real example you have seen
A city repainting lane markings on a busy highway usually paints the new lines beside the old ones first. It closes the road only for the few minutes it takes to switch which lines are active. It never redraws lanes underneath moving traffic.
Remember this
- Never rebuild an index in place while it is still being read from.
- Build the new version fully, in the background, alongside the old one.
- The cutover should be one small, instant, atomic step — a pointer change, not a rebuild.
What to learn next
- Document ingestion pipelines for RAG — the pipeline that keeps feeding an index that gets rebuilt this way.
- Re-embedding and reindexing — the most common reason a vector index needs this exact kind of rebuild.
- Model serving — the same blue-green pattern, one layer up, for the model itself.
Developer — Code and libraries.
Setup
No installation needed — this uses only Python's standard library sqlite3 and threading.
Reads that never see a broken index
import sqlite3
import threading
import time
import random
conn = sqlite3.connect(":memory:", check_same_thread=False)
lock = threading.Lock()
conn.execute("CREATE TABLE index_v1 (doc_id INTEGER, text TEXT)")
conn.executemany("INSERT INTO index_v1 VALUES (?,?)",
[(i, f"doc-{i}-old") for i in range(5)])
# a one-row table that says which index reads should use right now
conn.execute("CREATE TABLE active_pointer (active TEXT)")
conn.execute("INSERT INTO active_pointer VALUES ('index_v1')")
conn.commit()
def read_active_doc(doc_id):
with lock:
active = conn.execute("SELECT active FROM active_pointer").fetchone()[0]
row = conn.execute(
f"SELECT text FROM {active} WHERE doc_id = ?", (doc_id,)
).fetchone()
return row[0] if row else None
stop = False
errors = []
reads = []
def reader_loop():
while not stop:
try:
reads.append(read_active_doc(random.randint(0, 4)))
except Exception as e:
errors.append(str(e))
time.sleep(0.001)
reader_thread = threading.Thread(target=reader_loop)
reader_thread.start()
# Meanwhile, build the NEW index in the background. Readers above keep
# hitting index_v1 the whole time this runs -- they never see a half-built table.
time.sleep(0.05)
with lock:
conn.execute("CREATE TABLE index_v2 (doc_id INTEGER, text TEXT)")
for i in range(5):
time.sleep(0.01) # simulate real build work, one document at a time
with lock:
conn.execute("INSERT INTO index_v2 VALUES (?,?)", (i, f"doc-{i}-new"))
conn.commit()
# The swap: one tiny atomic update. This is the only moment readers move
# from the old index to the new one.
with lock:
conn.execute("UPDATE active_pointer SET active = 'index_v2'")
conn.commit()
time.sleep(0.05)
stop = True
reader_thread.join()
print(f"reads served during the whole rebuild+swap : {len(reads)}")
print(f"read errors (should be 0) : {len(errors)}")
print(f"first few reads : {reads[:3]}")
print(f"last few reads : {reads[-3:]}")reads served during the whole rebuild+swap : 85 read errors (should be 0) : 0 first few reads : ['doc-4-old', 'doc-1-old', 'doc-1-old'] last few reads : ['doc-1-new', 'doc-1-new', 'doc-0-new']
The exact read count and the exact sample values depend on timing and will differ between runs — this uses real threads with real sleeps. What is not timing-dependent: zero errors, and every single read returned a complete, consistent answer, either fully old or fully new.
Line-by-line walkthrough
active_pointer is the entire mechanism. Readers never ask "is the index finished building" — they only ever ask "which table is active right now", which is always a complete answer.
The lock around both reads and the pointer swap makes the swap itself atomic from a reader's point of view. Without it, a reader could read the pointer as index_v2 while index_v2 is still mid-build — the exact bug this whole pattern exists to prevent.
Common mistakes
DELETE FROM index followed by re-inserting rows. This is rebuilding in place with extra steps — the table is empty or partial for the entire rebuild, and any concurrent read sees it.
Forgetting to delete the old index after the swap. Not a correctness bug, but a slow leak of storage and cost across repeated rebuilds. Clean it up once you are confident the new one is good.
No verification step before swapping. Building index_v2 and swapping to it blindly ships any bug in the rebuild straight to production. Compare row counts, spot-check known queries, before flipping the pointer — this is the shadow-deployment idea from model serving's researcher section, applied to data instead of a model.
Treating this as unique to search. The identical pattern — this pattern is a form of blue-green deployment — is exactly how model serving recommends shipping a new model version without downtime.
Try it yourself
Add a deliberate bug to the index_v2 build loop — skip inserting doc_id=3. Add a verification step after the build that checks index_v2's row count matches index_v1's before allowing the swap, and confirm it catches the bug.
What to learn next
- Document ingestion pipelines for RAG — the pipeline that keeps feeding an index that gets rebuilt this way.
- Re-embedding and reindexing — the most common reason a vector index needs this exact kind of rebuild.
- Model serving — the same blue-green pattern, one layer up, for the model itself.
Researcher — Mathematics and papers.
This is blue-green deployment, applied to data
The pattern here is exactly model serving's blue-green deployment strategy, with "model version" replaced by "index version". Both share the same core guarantee: at every instant, 100% of traffic is served by one complete, consistent version — never a mix, never a partial one.
The cost is also the same: roughly double the storage or compute during the transition window, in exchange for zero read-path risk. For most systems this trade is favourable, since storage is cheap relative to a production outage or a window of wrong answers.
Making the swap actually atomic
A single-row UPDATE under a lock, as in the demo, is sufficient for a single-process system. Distributed systems need a stronger primitive:
- A versioned alias, as used by Elasticsearch and OpenSearch — an alias name points at a concrete index, and repointing an alias is a single atomic cluster-state operation, visible to all nodes at once.
- Compare-and-swap on a pointer key in a coordination service (etcd, ZooKeeper) or a strongly consistent key-value store, so concurrent swap attempts cannot race.
- A new deployment behind a load balancer, for index formats that live as whole running services rather than tables — the container-level equivalent of the same alias-flip idea.
The correctness property required in every case is the same: readers must never be able to observe a state where the pointer says "new" but the new structure is not yet fully built and durable.
Validation before cutover
A rebuilt index should be validated against the live one before the swap, not after. Standard checks: row/vector count parity within an expected tolerance, a fixed set of canary queries with known expected results (the same "frozen exam paper" idea from CI/CD for ML), and a sample-based comparison of top-k results between old and new for a batch of real recent queries, flagging large unexplained divergence.
Rollback
Because the old index is not deleted immediately after the swap, rollback is symmetric to the forward swap: flip the pointer back. Keeping the old version alive for a defined grace period — commonly the length of one full monitoring cycle — is standard practice specifically so this option exists without needing to rebuild from scratch under incident pressure.
Papers and systems
- Elasticsearch's index alias documentation is the clearest production reference implementation of atomic pointer-swap reindexing at scale.
- Kleppmann, Designing Data-Intensive Applications, Chapter 9 — the broader treatment of atomic commit and consensus this pattern relies on when distributed.
- Humble and Farley, Continuous Delivery, 2010 — the original popularisation of blue-green deployment, of which this is a data-layer instance.
What to learn next
- Document ingestion pipelines for RAG — the pipeline that keeps feeding an index that gets rebuilt this way.
- Re-embedding and reindexing — the most common reason a vector index needs this exact kind of rebuild.
- Model serving — the same blue-green pattern, one layer up, for the model itself.