Releasing Models Safely

Shadow deployment

A shadow deployment lets a new model answer every real request in secret, alongside the model actually serving users, so you can compare them honestly without ever risking a wrong answer reaching anyone.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A shadow deployment runs a new model on real traffic, in secret, without ever showing its answers to a real user.

The analogy you have already lived

You have seen a trainee doctor shadow a senior one — standing beside them during a real consultation, quietly forming their own diagnosis, comparing notes afterward. The trainee's opinion is never what the patient hears. Only the senior doctor's decision reaches the patient, every single time, while the trainee is learning from the exact same real cases.

A shadow deployment is that same arrangement for a model. The new model watches every real request and forms its own answer. Only the model already trusted in production ever replies to the user.

Why it exists

Testing a new model offline, on old data, only tells you how it would have done on cases that already happened. Real traffic today is not identical to a saved test set from last month. The only way to know how a new model handles today's traffic is to actually show it today's traffic.

Sending real users the new model's real answers is the most direct way to find out. It is also the riskiest: if the new model is wrong in a way nobody caught, real people get hurt by it before anyone notices. Shadow deployment gets that same information — how does the new model behave on live, current traffic — with none of that risk. Its answers never leave the log file they were written to.

How it works

   a real request arrives
            |
            +----------------------+
            |                      |
            v                      v
      OLD model (live)      NEW model (shadow)
            |                      |
            v                      v
     answer sent to the      answer written to a
     real user, every time   log file, and ONLY there

The two models see the identical request, at the same time. Only one of them is ever allowed to speak to the user.

A real example you have seen

Before a new fraud-detection model goes live at a bank, it is common to run it in shadow first. It scores every real transaction quietly, alongside the model that is actually blocking or allowing payments. If the new model would have blocked a transaction the old one allowed, someone reviews that case before the new model ever gets a vote.

Remember this

  • A shadow model sees every real request, but its answer is only ever logged, never returned.
  • It is the lowest-risk way to compare a new model against real, current traffic — nothing a shadow model says can go wrong for a user.
  • The cost is doubled compute — you are running two models for the price of one answer — and shadow testing alone cannot tell you how users would react to the new answers, only whether the two models agree.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn pandas numpy pytest

The shadow scorer

Two versions of the loan-approval model from model serving — the current one, and a candidate trained with different regularisation. Every request goes to both. Only the current model's answer is returned.

shadow.py
"""A shadow deployment: the OLD model answers every real request.
The NEW model scores the same request in the background, and its
answer is only logged -- it is never returned to a caller.
"""
import time

import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

COLUMNS = ["income", "years", "age"]


def make_training_data(n=400, seed=0):
    rng = np.random.RandomState(seed)
    X = pd.DataFrame({
        "income": rng.uniform(5, 80, n).round(1),
        "years": rng.uniform(0, 10, n).round(1),
        "age": rng.randint(21, 65, n).astype(float),
    })
    score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
    y = (score + rng.normal(0, 0.8, n) > 0).astype(int)
    return X, y


X, y = make_training_data()
OLD_MODEL = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000, C=1.0)).fit(X, y)
# "NEW": retrained with different regularisation -- a realistic candidate improvement
NEW_MODEL = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000, C=0.1)).fit(X, y)

SHADOW_LOG = []


def score_request(applicant: dict) -> dict:
    row = pd.DataFrame([applicant], columns=COLUMNS)

    t0 = time.perf_counter()
    old_answer = int(OLD_MODEL.predict(row)[0])
    old_ms = (time.perf_counter() - t0) * 1000

    t0 = time.perf_counter()
    new_answer = int(NEW_MODEL.predict(row)[0])
    new_ms = (time.perf_counter() - t0) * 1000

    SHADOW_LOG.append({
        "applicant": applicant,
        "old_answer": old_answer,
        "new_answer": new_answer,
        "agree": old_answer == new_answer,
        "old_ms": round(old_ms, 4),
        "new_ms": round(new_ms, 4),
    })

    return {"repaid": old_answer}  # ONLY the old model's answer ever leaves this function

Running it against 200 simulated requests

python
from shadow import *

rng = np.random.RandomState(1)
for _ in range(200):
    applicant = {
        "income": round(rng.uniform(5, 80), 1),
        "years": round(rng.uniform(0, 10), 1),
        "age": float(rng.randint(21, 65)),
    }
    score_request(applicant)

agree_rate = sum(r["agree"] for r in SHADOW_LOG) / len(SHADOW_LOG)
avg_old_ms = sum(r["old_ms"] for r in SHADOW_LOG) / len(SHADOW_LOG)
avg_new_ms = sum(r["new_ms"] for r in SHADOW_LOG) / len(SHADOW_LOG)
print(f"requests scored: {len(SHADOW_LOG)}")
print(f"agreement rate between old and new: {agree_rate:.3f}")
print(f"avg old-model time per request: {avg_old_ms:.4f} ms")
print(f"avg new-model time per request (shadow, never shown to user): {avg_new_ms:.4f} ms")
Output
requests scored: 200
agreement rate between old and new: 0.980
avg old-model time per request: 0.4239 ms
avg new-model time per request (shadow, never shown to user): 0.4025 ms

Those exact numbers are one real run on one machine — the timings especially will vary run to run and machine to machine. The agreement rate (98% on this batch) is the number worth watching for a real decision, and it will shift somewhat with a different random batch of requests too.

Line-by-line: reading a disagreement

python
from shadow import *

rng = np.random.RandomState(1)
for _ in range(200):
    applicant = {
        "income": round(rng.uniform(5, 80), 1),
        "years": round(rng.uniform(0, 10), 1),
        "age": float(rng.randint(21, 65)),
    }
    score_request(applicant)

disagreements = [r for r in SHADOW_LOG if not r["agree"]]
print(f"disagreements: {len(disagreements)} -- first one: {disagreements[0]}")
Output
disagreements: 4 -- first one: {'applicant': {'income': 34.7, 'years': 3.9, 'age': 27.0}, 'old_answer': 0, 'new_answer': 1, 'agree': False, 'old_ms': 0.3785, 'new_ms': 0.3564}

This is the actual value of shadow deployment: not a single aggregate agreement number, but the specific cases where the two models disagree, saved with the exact input that produced the split. A human reviewing this list can look at the four disagreements out of two hundred and judge whether the new model's opinion is an improvement, before it ever gets to act on a real applicant's behalf.

Testing the shadow property itself

The one property that matters most is also the easiest to accidentally break — a refactor that starts returning the new model's answer by mistake. Test it directly.

test_shadow.py
from shadow import OLD_MODEL, SHADOW_LOG, score_request


def test_the_returned_answer_always_matches_the_old_model():
    SHADOW_LOG.clear()
    applicant = {"income": 60.0, "years": 7.0, "age": 34.0}
    result = score_request(applicant)

    import pandas as pd
    expected = int(OLD_MODEL.predict(pd.DataFrame([applicant], columns=["income", "years", "age"]))[0])
    assert result["repaid"] == expected


def test_the_new_model_is_logged_but_never_returned():
    SHADOW_LOG.clear()
    result = score_request({"income": 60.0, "years": 7.0, "age": 34.0})

    assert "new_answer" not in result          # the response never carries it
    assert "new_answer" in SHADOW_LOG[-1]        # but the log always does


def test_every_scored_request_gets_a_log_entry():
    SHADOW_LOG.clear()
    for i in range(5):
        score_request({"income": 30.0 + i, "years": 2.0, "age": 25.0})
    assert len(SHADOW_LOG) == 5
bash
pytest test_shadow.py -q
Output
...                                                                      [100%]
3 passed in 0.87s

Common mistakes

Letting the shadow call block the real response. If NEW_MODEL.predict(...) runs synchronously in the request path and is slow, users pay for the shadow model's latency even though they never see its answer. In a real service, run the shadow call in the background (a queue, a separate worker, asyncio.create_task) so a slow or crashing shadow model can never affect a real response.

Forgetting the shadow model can crash. A shadow model throwing an exception should never take down the request that is actually serving a user. Wrap the shadow call in its own error handling, and let a shadow failure produce nothing worse than a gap in the log.

Comparing only the final label. Two models can agree on the yes/no answer while disagreeing sharply on confidence — one at 0.51, the other at 0.98. Log the full probability, not only the rounded decision, so a review can see near-misses as well as outright disagreements.

Treating a high agreement rate as proof the new model is better. A 98% agreement rate says the two models are similar. It says nothing about which one is right. Shadow deployment tells you where models disagree; deciding which one is correct still needs labelled outcomes, a canary release, or human review of the disagreements.

Running shadow forever, on everything. Shadow deployment doubles compute cost for as long as it runs. It is a phase with a start and an end, not a permanent architecture — once you trust the new model, move to a canary and retire the shadow.

Try it yourself

Extend SHADOW_LOG to also record each model's predicted probability, not only the rounded 0/1 answer. Then find the disagreement with the smallest gap between the two probabilities, and the one with the largest — the smallest gap is a genuine toss-up; the largest is the case most worth a human looking at first.

What to learn next

Researcher — Mathematics and papers.

What shadow deployment can and cannot measure

Shadow deployment estimates $P(\hat{y}{\text{new}} \ne \hat{y}{\text{old}} \mid x)$ over the live input distribution $x$ — the disagreement rate — without any risk exposure, because $\hat{y}{\text{new}}$ is never acted upon. It cannot estimate which of $\hat{y}{\text{new}}$ or $\hat{y}_{\text{old}}$ is closer to the true outcome $y$, because $y$ is frequently unobserved or only observed with a delay (a loan default, a fraud confirmation). This is the structural reason shadow deployment is a necessary but not sufficient step before a live rollout: it derisks the comparison, not the decision of which model is actually better.

Selection bias when the shadow model would have changed behaviour

A subtler limitation applies to systems where the served decision changes the state the next request is evaluated against — a ranking system, a recommender, a pricing engine. If the new model would have shown a different item, at a different price, the resulting user action (click, purchase) under shadow is never observed for the new model's counterfactual choice, only for the old model's actual choice. Shadow deployment is unbiased for models that classify or score a request without materially altering downstream state; it under-informs for models whose decision changes what happens next. This is a specific case of the offline-online evaluation gap discussed in CI/CD for machine learning, and the reason interleaving and true canary exposure remain necessary for ranking systems even after a clean shadow run.

Cost and infrastructure

Shadow deployment's compute cost is, at minimum, the cost of one extra full inference pass per request — for an LLM-scale model this is not a rounding error, and teams commonly shadow only a sampled fraction of traffic (1-10%) rather than 100%, trading statistical power in the disagreement estimate for cost. The sampling itself should be a fixed hash of a request identifier rather than independent per-request random sampling, so that any downstream analysis joining shadow results to other logged features remains consistent.

Where this sits among release strategies

Crankshaw et al. (2017), Clipper, and the broader model-serving systems literature treat shadowing as the zero-risk end of a spectrum that continues through canary (small real exposure), then blue-green (full exposure, instant switch), each trading risk against the fidelity of the signal obtained — a spectrum this section covers lesson by lesson.

Papers

  • Crankshaw et al., Clipper: A Low-Latency Online Prediction Serving System, NSDI 2017 — arxiv.org/abs/1612.03079
  • Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — the feedback-loop risk of a decision-changing shadow deployment traces to the same "hidden feedback loops" pattern named here.

What to learn next

What to learn next

These follow on from what you just read.

  • Releasing Models Safely

    Canary releases for models

    A canary release gives a new model a small, real slice of traffic and watches it closely, the way miners once carried a caged canary underground — small warning, before anyone bigger got hurt.

  • Releasing Models Safely

    Blue-green model deployments

    A blue-green deploy keeps two complete, fully warmed-up environments and flips every request between them with one instant switch, the way a home changeover switch moves the whole house from grid power to a generator, never a blend of both.

  • Releasing Models Safely

    Rolling back a model

    Rolling back means switching straight back to the last version you trusted, without stopping to diagnose the new one first — the same instinct as putting your old shoes back on the moment new ones start to hurt.