Shadow deployment
A shadow deployment lets a new model answer every real request in secret, alongside the model actually serving users, so you can compare them honestly without ever risking a wrong answer reaching anyone.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A shadow deployment runs a new model on real traffic, in secret, without ever showing its answers to a real user.
The analogy you have already lived
You have seen a trainee doctor shadow a senior one — standing beside them during a real consultation, quietly forming their own diagnosis, comparing notes afterward. The trainee's opinion is never what the patient hears. Only the senior doctor's decision reaches the patient, every single time, while the trainee is learning from the exact same real cases.
A shadow deployment is that same arrangement for a model. The new model watches every real request and forms its own answer. Only the model already trusted in production ever replies to the user.
Why it exists
Testing a new model offline, on old data, only tells you how it would have done on cases that already happened. Real traffic today is not identical to a saved test set from last month. The only way to know how a new model handles today's traffic is to actually show it today's traffic.
Sending real users the new model's real answers is the most direct way to find out. It is also the riskiest: if the new model is wrong in a way nobody caught, real people get hurt by it before anyone notices. Shadow deployment gets that same information — how does the new model behave on live, current traffic — with none of that risk. Its answers never leave the log file they were written to.
How it works
a real request arrives
|
+----------------------+
| |
v v
OLD model (live) NEW model (shadow)
| |
v v
answer sent to the answer written to a
real user, every time log file, and ONLY thereThe two models see the identical request, at the same time. Only one of them is ever allowed to speak to the user.
A real example you have seen
Before a new fraud-detection model goes live at a bank, it is common to run it in shadow first. It scores every real transaction quietly, alongside the model that is actually blocking or allowing payments. If the new model would have blocked a transaction the old one allowed, someone reviews that case before the new model ever gets a vote.
Remember this
- A shadow model sees every real request, but its answer is only ever logged, never returned.
- It is the lowest-risk way to compare a new model against real, current traffic — nothing a shadow model says can go wrong for a user.
- The cost is doubled compute — you are running two models for the price of one answer — and shadow testing alone cannot tell you how users would react to the new answers, only whether the two models agree.
What to learn next
- Canary releases for models — the next step, giving the new model a small amount of real influence instead of none.
- Champion and challenger models — a longer-running version of the same old-versus-new comparison built here.
- Integration testing an inference server — making sure the serving code itself is solid before you trust it to run two models per request.
Developer — Code and libraries.
Setup
pip install scikit-learn pandas numpy pytestThe shadow scorer
Two versions of the loan-approval model from model serving — the current one, and a candidate trained with different regularisation. Every request goes to both. Only the current model's answer is returned.
"""A shadow deployment: the OLD model answers every real request.
The NEW model scores the same request in the background, and its
answer is only logged -- it is never returned to a caller.
"""
import time
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
COLUMNS = ["income", "years", "age"]
def make_training_data(n=400, seed=0):
rng = np.random.RandomState(seed)
X = pd.DataFrame({
"income": rng.uniform(5, 80, n).round(1),
"years": rng.uniform(0, 10, n).round(1),
"age": rng.randint(21, 65, n).astype(float),
})
score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
y = (score + rng.normal(0, 0.8, n) > 0).astype(int)
return X, y
X, y = make_training_data()
OLD_MODEL = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000, C=1.0)).fit(X, y)
# "NEW": retrained with different regularisation -- a realistic candidate improvement
NEW_MODEL = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000, C=0.1)).fit(X, y)
SHADOW_LOG = []
def score_request(applicant: dict) -> dict:
row = pd.DataFrame([applicant], columns=COLUMNS)
t0 = time.perf_counter()
old_answer = int(OLD_MODEL.predict(row)[0])
old_ms = (time.perf_counter() - t0) * 1000
t0 = time.perf_counter()
new_answer = int(NEW_MODEL.predict(row)[0])
new_ms = (time.perf_counter() - t0) * 1000
SHADOW_LOG.append({
"applicant": applicant,
"old_answer": old_answer,
"new_answer": new_answer,
"agree": old_answer == new_answer,
"old_ms": round(old_ms, 4),
"new_ms": round(new_ms, 4),
})
return {"repaid": old_answer} # ONLY the old model's answer ever leaves this functionRunning it against 200 simulated requests
from shadow import *
rng = np.random.RandomState(1)
for _ in range(200):
applicant = {
"income": round(rng.uniform(5, 80), 1),
"years": round(rng.uniform(0, 10), 1),
"age": float(rng.randint(21, 65)),
}
score_request(applicant)
agree_rate = sum(r["agree"] for r in SHADOW_LOG) / len(SHADOW_LOG)
avg_old_ms = sum(r["old_ms"] for r in SHADOW_LOG) / len(SHADOW_LOG)
avg_new_ms = sum(r["new_ms"] for r in SHADOW_LOG) / len(SHADOW_LOG)
print(f"requests scored: {len(SHADOW_LOG)}")
print(f"agreement rate between old and new: {agree_rate:.3f}")
print(f"avg old-model time per request: {avg_old_ms:.4f} ms")
print(f"avg new-model time per request (shadow, never shown to user): {avg_new_ms:.4f} ms")requests scored: 200 agreement rate between old and new: 0.980 avg old-model time per request: 0.4239 ms avg new-model time per request (shadow, never shown to user): 0.4025 ms
Those exact numbers are one real run on one machine — the timings especially will vary run to run and machine to machine. The agreement rate (98% on this batch) is the number worth watching for a real decision, and it will shift somewhat with a different random batch of requests too.
Line-by-line: reading a disagreement
from shadow import *
rng = np.random.RandomState(1)
for _ in range(200):
applicant = {
"income": round(rng.uniform(5, 80), 1),
"years": round(rng.uniform(0, 10), 1),
"age": float(rng.randint(21, 65)),
}
score_request(applicant)
disagreements = [r for r in SHADOW_LOG if not r["agree"]]
print(f"disagreements: {len(disagreements)} -- first one: {disagreements[0]}")disagreements: 4 -- first one: {'applicant': {'income': 34.7, 'years': 3.9, 'age': 27.0}, 'old_answer': 0, 'new_answer': 1, 'agree': False, 'old_ms': 0.3785, 'new_ms': 0.3564}This is the actual value of shadow deployment: not a single aggregate agreement number, but the specific cases where the two models disagree, saved with the exact input that produced the split. A human reviewing this list can look at the four disagreements out of two hundred and judge whether the new model's opinion is an improvement, before it ever gets to act on a real applicant's behalf.
Testing the shadow property itself
The one property that matters most is also the easiest to accidentally break — a refactor that starts returning the new model's answer by mistake. Test it directly.
from shadow import OLD_MODEL, SHADOW_LOG, score_request
def test_the_returned_answer_always_matches_the_old_model():
SHADOW_LOG.clear()
applicant = {"income": 60.0, "years": 7.0, "age": 34.0}
result = score_request(applicant)
import pandas as pd
expected = int(OLD_MODEL.predict(pd.DataFrame([applicant], columns=["income", "years", "age"]))[0])
assert result["repaid"] == expected
def test_the_new_model_is_logged_but_never_returned():
SHADOW_LOG.clear()
result = score_request({"income": 60.0, "years": 7.0, "age": 34.0})
assert "new_answer" not in result # the response never carries it
assert "new_answer" in SHADOW_LOG[-1] # but the log always does
def test_every_scored_request_gets_a_log_entry():
SHADOW_LOG.clear()
for i in range(5):
score_request({"income": 30.0 + i, "years": 2.0, "age": 25.0})
assert len(SHADOW_LOG) == 5pytest test_shadow.py -q... [100%] 3 passed in 0.87s
Common mistakes
Letting the shadow call block the real response. If NEW_MODEL.predict(...) runs synchronously in the request path and is slow, users pay for the shadow model's latency even though they never see its answer. In a real service, run the shadow call in the background (a queue, a separate worker, asyncio.create_task) so a slow or crashing shadow model can never affect a real response.
Forgetting the shadow model can crash. A shadow model throwing an exception should never take down the request that is actually serving a user. Wrap the shadow call in its own error handling, and let a shadow failure produce nothing worse than a gap in the log.
Comparing only the final label. Two models can agree on the yes/no answer while disagreeing sharply on confidence — one at 0.51, the other at 0.98. Log the full probability, not only the rounded decision, so a review can see near-misses as well as outright disagreements.
Treating a high agreement rate as proof the new model is better. A 98% agreement rate says the two models are similar. It says nothing about which one is right. Shadow deployment tells you where models disagree; deciding which one is correct still needs labelled outcomes, a canary release, or human review of the disagreements.
Running shadow forever, on everything. Shadow deployment doubles compute cost for as long as it runs. It is a phase with a start and an end, not a permanent architecture — once you trust the new model, move to a canary and retire the shadow.
Try it yourself
Extend SHADOW_LOG to also record each model's predicted probability, not only the rounded 0/1 answer. Then find the disagreement with the smallest gap between the two probabilities, and the one with the largest — the smallest gap is a genuine toss-up; the largest is the case most worth a human looking at first.
What to learn next
- Canary releases for models — the next step, giving the new model a small amount of real influence instead of none.
- Champion and challenger models — a longer-running version of the same old-versus-new comparison built here.
- Integration testing an inference server — making sure the serving code itself is solid before you trust it to run two models per request.
Researcher — Mathematics and papers.
What shadow deployment can and cannot measure
Shadow deployment estimates $P(\hat{y}{\text{new}} \ne \hat{y}{\text{old}} \mid x)$ over the live input distribution $x$ — the disagreement rate — without any risk exposure, because $\hat{y}{\text{new}}$ is never acted upon. It cannot estimate which of $\hat{y}{\text{new}}$ or $\hat{y}_{\text{old}}$ is closer to the true outcome $y$, because $y$ is frequently unobserved or only observed with a delay (a loan default, a fraud confirmation). This is the structural reason shadow deployment is a necessary but not sufficient step before a live rollout: it derisks the comparison, not the decision of which model is actually better.
Selection bias when the shadow model would have changed behaviour
A subtler limitation applies to systems where the served decision changes the state the next request is evaluated against — a ranking system, a recommender, a pricing engine. If the new model would have shown a different item, at a different price, the resulting user action (click, purchase) under shadow is never observed for the new model's counterfactual choice, only for the old model's actual choice. Shadow deployment is unbiased for models that classify or score a request without materially altering downstream state; it under-informs for models whose decision changes what happens next. This is a specific case of the offline-online evaluation gap discussed in CI/CD for machine learning, and the reason interleaving and true canary exposure remain necessary for ranking systems even after a clean shadow run.
Cost and infrastructure
Shadow deployment's compute cost is, at minimum, the cost of one extra full inference pass per request — for an LLM-scale model this is not a rounding error, and teams commonly shadow only a sampled fraction of traffic (1-10%) rather than 100%, trading statistical power in the disagreement estimate for cost. The sampling itself should be a fixed hash of a request identifier rather than independent per-request random sampling, so that any downstream analysis joining shadow results to other logged features remains consistent.
Where this sits among release strategies
Crankshaw et al. (2017), Clipper, and the broader model-serving systems literature treat shadowing as the zero-risk end of a spectrum that continues through canary (small real exposure), then blue-green (full exposure, instant switch), each trading risk against the fidelity of the signal obtained — a spectrum this section covers lesson by lesson.
Papers
- Crankshaw et al., Clipper: A Low-Latency Online Prediction Serving System, NSDI 2017 — arxiv.org/abs/1612.03079
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — the feedback-loop risk of a decision-changing shadow deployment traces to the same "hidden feedback loops" pattern named here.
What to learn next
- Canary releases for models — the next step, giving the new model a small amount of real influence instead of none.
- Champion and challenger models — a longer-running version of the same old-versus-new comparison built here.
- Integration testing an inference server — making sure the serving code itself is solid before you trust it to run two models per request.