Swapping weights with no downtime
A zero-downtime weight swap replaces a model's numbers inside an already-running server, the way a relay runner hands off the baton at full speed — the new runner is carrying it before the old one lets go, so the race never actually stops.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A zero-downtime weight swap replaces a model's numbers inside a server that is already running, without restarting it or dropping any request.
The analogy you have already lived
You have watched a relay race. The baton passes from one runner to the next while both are at full speed. There is no moment where the baton is on the ground, and no moment where the race actually stops. The handoff happens, and the race continues without ever stopping, faster hands than eyes.
Swapping a model's weights with no downtime is that same handoff. The server keeps running, keeps answering requests, the entire time — it is only the numbers inside it, the model's learned weights, that change hands.
Why it exists
Blue-green deployment solves a similar problem by running two entirely separate environments and switching between them. That is the right tool when you genuinely need two full, independent copies of everything.
Often you do not. Sometimes all that changed is the model file itself — the server code, the request handling, the validation, all of it stays exactly the same. Spinning up a whole second environment for that is real, avoidable cost. A weight swap gets the same "no downtime, no dropped requests" guarantee, inside the single process you already have running.
How it works
server is running, answering requests with model version A
|
v
load model version B fully into memory, OFF TO THE SIDE
(the server keeps answering with A this whole time)
|
v
ONE atomic switch: "the live model is now B"
|
v
every request from this instant on uses B
every request already in flight finishes safely with whichever
version it started withThe loading is slow and happens quietly, before anyone is affected. The switch itself is a single, instant action.
A real example you have seen
A GPS navigation app updating its traffic model mid-journey does not restart the app or interrupt your current route. The underlying model behind its ETA predictions can change between one screen refresh and the next, with the app itself never stopping.
Remember this
- A weight swap changes the model inside a running process, without a restart — a lighter-weight tool than a full blue-green deploy.
- Load the new version fully, before switching — the switch itself should be the fast, easy part.
- The switch must be one atomic action — anything done in two steps reopens the exact race condition atomicity is meant to close.
What to learn next
- Versioning a model API — keeping callers correctly informed about which weights actually answered their request, across any of these swap mechanisms.
- Blue-green model deployments — the heavier-weight sibling of this pattern, for when the change is more than the model's weights alone.
- Model warm-up and cold starts — the general problem of loading a model without making the first requests after pay for it.
Developer — Code and libraries.
Setup
pip install scikit-learn pandas numpy joblib pytestThe safe pattern: atomic file write, then atomic in-memory swap
Two separate risks need handling here. The first: writing a new model file to disk while something might be reading it. The second: swapping the in-memory reference while requests are actively being served. Both get the same fix — do the slow part off to the side, and make the actual switch a single, instant step.
"""An in-process model host: swaps the live model's weights inside a single
running server, with no restart and no dropped requests.
"""
import os
import tempfile
import joblib
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
COLUMNS = ["income", "years", "age"]
def make_training_data(n=400, seed=0):
rng = np.random.RandomState(seed)
X = pd.DataFrame({
"income": rng.uniform(5, 80, n).round(1),
"years": rng.uniform(0, 10, n).round(1),
"age": rng.randint(21, 65, n).astype(float),
})
score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
y = (score + rng.normal(0, 0.8, n) > 0).astype(int)
return X, y
def train_and_save(path: str, seed: int):
"""Writes a model to disk SAFELY: write to a temp file, then atomically
rename it into place. A reader can never observe a half-written file.
"""
X, y = make_training_data(seed=seed)
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X, y)
directory = os.path.dirname(path) or "."
fd, tmp_path = tempfile.mkstemp(dir=directory)
os.close(fd)
joblib.dump(model, tmp_path)
os.replace(tmp_path, path) # atomic on POSIX and on Windows (same volume)
class ModelHost:
def __init__(self, path: str):
self.path = path
self._model = joblib.load(path)
def reload(self):
"""Loads the NEW model fully, off to the side, then swaps in one step."""
new_model = joblib.load(self.path) # slow work happens BEFORE the swap
self._model = new_model # single attribute write -- atomic
def predict(self, applicant: dict) -> int:
row = pd.DataFrame([applicant], columns=COLUMNS)
return int(self._model.predict(row)[0])
train_and_save("model_v1.joblib", seed=0)
host = ModelHost("model_v1.joblib")
print("initial prediction:", host.predict({"income": 60.0, "years": 7.0, "age": 34.0}))
train_and_save("model_v1.joblib", seed=1) # a fresh version, same filename, written atomically
host.reload()
print("prediction after hot reload:", host.predict({"income": 60.0, "years": 7.0, "age": 34.0}))initial prediction: 1 prediction after hot reload: 1
(Both answers happen to agree for this particular applicant — the point being demonstrated here is that the reload runs cleanly, not that the answer necessarily changes.)
Proving zero downtime, under real concurrent load
import threading
import time
from host import ModelHost, train_and_save
train_and_save("model_v2.joblib", seed=0)
host = ModelHost("model_v2.joblib")
applicant = {"income": 60.0, "years": 7.0, "age": 34.0}
results, errors = [], []
stop = threading.Event()
def hammer():
while not stop.is_set():
try:
results.append(host.predict(applicant))
except Exception as e:
errors.append(repr(e))
threads = [threading.Thread(target=hammer) for _ in range(8)]
for t in threads:
t.start()
time.sleep(0.05)
train_and_save("model_v2.joblib", seed=1) # write the new version to disk, atomically
host.reload() # load it fully, then swap in one step
time.sleep(0.05)
stop.set()
for t in threads:
t.join()
print(f"total predictions during the reload: {len(results)}")
print(f"errors: {len(errors)}")total predictions during the reload: 412 errors: 0
Eight threads made 412 real predictions while a full reload — writing a new file to disk and swapping it into a live host — happened in the middle of that traffic. Zero errors. The exact count of 412 will differ run to run; the zero should not.
Now break something on purpose: writing straight over the live file
def unsafe_writer():
"""The unsafe version: writes straight over the live file, no temp file, no rename."""
X2, y2 = make_training_data(seed=1)
new_model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X2, y2)
joblib.dump(new_model, PATH) # overwrites the file readers are actively readingSix reader threads repeatedly call joblib.load(PATH) — standing in for other worker processes polling the same file — while the writer overwrites it directly, twenty times, with no temp file and no atomic rename:
successful reads: 830 errors while reading a file being overwritten mid-write: 33 example errors: ['EOFError()', 'EOFError()', 'EOFError()']
Thirty-three real failures, out of 863 total attempted reads. A reader that happened to open the file partway through the writer's joblib.dump call read a truncated, half-written file and hit a genuine EOFError. This is exactly the failure the temp-file-then-rename pattern in train_and_save prevents: os.replace is atomic at the filesystem level, so any reader either sees the complete old file or the complete new one — there is no instant in between where a partial file is visible to anyone.
Testing the reload behaviour
import glob
from host import ModelHost, train_and_save
def test_reload_actually_changes_the_served_predictions_where_they_differ():
train_and_save("model_test.joblib", seed=0)
host = ModelHost("model_test.joblib")
applicant = {"income": 7.6, "years": 7.6, "age": 53.0} # a real point where the two models disagree
before = host.predict(applicant)
train_and_save("model_test.joblib", seed=1)
host.reload()
after = host.predict(applicant)
assert (before, after) == (0, 1)
def test_atomic_save_leaves_no_temp_file_behind():
before = set(glob.glob("tmp*"))
train_and_save("model_test3.joblib", seed=0)
after = set(glob.glob("tmp*"))
assert after == before # the temp file was renamed away, not left sitting aroundpytest test_host.py -q.. [100%] 2 passed in 0.86s
Common mistakes
Writing directly to the live file path. The unsafe_write.py demonstration above is a real, common bug — it looks harmless in a quick manual test and fails under exactly the concurrent load a production service actually has.
Loading the new model on the same thread that is serving requests. If reload() blocks the only thread handling traffic while it reads a large file from disk, you have not achieved zero downtime — you have achieved a very short, very real downtime. Run the load on a background thread or worker, and only touch the shared reference from the main path, briefly, for the swap itself.
Two-step swaps. Setting self._model = None and then self._model = new_model a moment later reopens the exact race demonstrated in blue-green model deployments — any request landing in that gap sees a model that is not there.
Forgetting the memory cost. For a brief window, both the old and new model are fully loaded in memory at once — the old one still being referenced by in-flight requests, the new one already loaded and ready. For a small model this is nothing; for a multi-gigabyte model, plan for genuinely needing roughly double the memory during a swap, not only during a full blue-green deploy.
No way to confirm the reload actually happened. Log the model's version or a content hash on every reload, and expose it on a health or status endpoint — otherwise "did the reload actually take?" is a question you can only answer by guessing from behaviour.
Try it yourself
Add a version field to the object train_and_save writes to disk (a simple string, alongside the model), and have ModelHost expose host.current_version(). Then write a test that reloads twice in a row and confirms the version reported changes each time, in order — proof the host is not accidentally still serving a stale in-memory copy.
What to learn next
- Versioning a model API — keeping callers correctly informed about which weights actually answered their request, across any of these swap mechanisms.
- Blue-green model deployments — the heavier-weight sibling of this pattern, for when the change is more than the model's weights alone.
- Model warm-up and cold starts — the general problem of loading a model without making the first requests after pay for it.
Researcher — Mathematics and papers.
Where the atomicity guarantees actually come from
Two distinct atomicity properties are combined in this lesson, from two different layers:
In-memory swap. As in blue-green model deployments, a single attribute assignment is atomic under CPython's Global Interpreter Lock for simple reference reassignment. This guarantee is process-local and CPython-specific; it says nothing about coordination across separate OS processes, which is the common real deployment shape (multiple worker processes behind a load balancer, each with its own ModelHost) — each process needs its own reload triggered independently, or a shared mechanism (a file-watch, a signal, a message) to coordinate them.
On-disk atomic rename. os.replace (and the POSIX rename(2) syscall it wraps) is guaranteed atomic by the filesystem for a rename within the same filesystem/volume: a concurrent reader observes either the complete old file or the complete new one, never a partial write. This guarantee does not hold across filesystem boundaries (renaming across mounted volumes falls back to a non-atomic copy-then-delete on many systems) and does not, by itself, guarantee the data is durably on disk after a crash — that additionally requires an fsync before the rename, which joblib.dump does not perform for you.
Why the unsafe demonstration reliably fails
joblib.dump internally performs multiple write syscalls for anything beyond a genuinely small object (pickling a scikit-learn pipeline touches numpy array buffers written in several chunks). A reader's open() racing against the writer's open(path, "wb") — which truncates the file to zero length immediately, before any new content is written — can observe the file at any point along that multi-write sequence, including the truncated, empty state, which is what produces the EOFError captured above. The failure rate observed (33 of 863) is a function of the relative timing of reader and writer syscalls on one machine under one load pattern, not a fixed probability — a different machine, filesystem, or I/O load could show a higher or lower rate, though the failure mode itself is structural, not incidental.
Reload coordination at scale
Beyond a single process, production hot-reload commonly uses one of: a shared file-modification-time or content-hash check polled periodically by each worker; a pub/sub message (Redis, a message queue) broadcasting "reload now" to every worker simultaneously; or an orchestrator-level rolling restart, which trades true zero-downtime-within-a-process for the operational simplicity of only restarting workers one at a time behind a load balancer that only routes to ready ones — effectively a blue-green pattern applied at the granularity of individual worker processes rather than whole environments.
Papers and prior art
- Lamport, On Interprocess Communication, 1986 — foundational treatment of atomicity and the guarantees (and non-guarantees) of shared mutable state under concurrency, the general principle both swaps in this lesson rely on.
- POSIX.1-2017,
rename()specification — the formal atomicity contractos.replacedepends on.
What to learn next
- Versioning a model API — keeping callers correctly informed about which weights actually answered their request, across any of these swap mechanisms.
- Blue-green model deployments — the heavier-weight sibling of this pattern, for when the change is more than the model's weights alone.
- Model warm-up and cold starts — the general problem of loading a model without making the first requests after pay for it.