Canary releases for models
A canary release gives a new model a small, real slice of traffic and watches it closely, the way miners once carried a caged canary underground — small warning, before anyone bigger got hurt.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A canary release sends a small, real slice of traffic to a new model, and watches it closely before sending any more.
The analogy you have already lived
You may know where the phrase comes from, even if you have not seen it done. Coal miners once carried a caged canary underground. The bird's small body reacted to dangerous gas long before a person would notice anything wrong. If the canary showed distress, the miners left immediately, before it became a danger to them.
A canary release is that same idea, applied to a model. A small slice of real traffic — say, five in every hundred requests — goes to the new model first. If something is wrong, only that small slice is affected, and it shows up before the new model ever reaches everyone.
Why it exists
Shadow deployment is the safest possible test — the new model never actually decides anything for a real user. But it also cannot tell you everything. It cannot measure how real users respond to the new model's actual decisions, because those decisions are never shown to anyone.
At some point, the new model has to be allowed to actually act. A canary release lets that happen for a small, controlled slice of traffic first, instead of everyone at once. If the new model turns out to have a real problem, only a small number of real people were affected, not all of them. You find out before it spreads any further.
How it works
100 real requests arrive
|
+---- 95 go to the OLD model (control) -- answers as always
|
+---- 5 go to the NEW model (canary) -- answers for real, watched closely
|
v
metrics on the canary slice look healthy?
/ \
yes no
| |
raise the canary's roll back to 0% immediately;
share a little only a small slice was ever affected
(5% -> 25% -> 100%)The canary's share grows step by step, only after each step looks healthy, not all at once.
A real example you have seen
A ride-hailing or food-delivery app rolling out a new pricing model rarely flips it on for every city at once. A handful of cities, or a small percentage of trips, gets the new pricing first. The company watches whether people still book rides at the same rate before turning it on everywhere.
Remember this
- A canary release gives a new model a small, real share of traffic, not everyone and not no one.
- The share grows gradually, only after the current step looks healthy.
- The whole method depends on the split being trustworthy and the comparison being statistically honest — both covered in the developer section below.
What to learn next
- Blue-green model deployments — the alternative when you want an instant, all-at-once switch instead of a gradual ramp.
- Sample ratio mismatch — a deeper look at diagnosing exactly why a split went wrong.
- Rolling back a model — what actually happens the moment a canary's guardrail metric fails.
Developer — Code and libraries.
Setup
pip install scipy numpy pytestRouting traffic: sticky, not random
The same user hitting the service twice must land in the same bucket both times — a user who sees the new model once and the old model the next time has an inconsistent, confusing experience, and any comparison between the two groups becomes unreliable. Hashing the user ID gives a stable, deterministic assignment without needing to store anything.
"""A sticky, hash-based canary router: the same user always lands in the
same bucket, and the split is close to the requested percentage on average.
"""
import hashlib
def assign_bucket(user_id: str, canary_percent: float) -> str:
"""Deterministic: hashing the same user_id always gives the same bucket."""
digest = hashlib.md5(user_id.encode()).hexdigest()
bucket_value = int(digest[:8], 16) % 10_000 # 0..9999
return "canary" if bucket_value < canary_percent * 100 else "control"
users = [f"user-{i}" for i in range(20_000)]
buckets = [assign_bucket(u, canary_percent=5.0) for u in users]
canary_count = buckets.count("canary")
print(f"requested canary: 5.0% measured canary: {canary_count/len(users)*100:.2f}% ({canary_count}/{len(users)})")
repeats = {assign_bucket("user-42", 5.0) for _ in range(5)}
print("user-42 bucket across 5 calls:", repeats)requested canary: 5.0% measured canary: 5.01% (1003/20000)
user-42 bucket across 5 calls: {'control'}Twenty thousand distinct users, hashed once each, landed within 0.01 percentage points of the requested 5% — real measured output from hashlib.md5, not a guaranteed exact value (a different set of user IDs will land close to, but not exactly on, 5.00%). The repeated calls for user-42 all agree, confirming the routing is genuinely sticky.
Before trusting any numbers: is the split actually what you asked for?
A sample ratio mismatch (SRM) is a bug in the split itself — the canary is quietly getting the wrong share of traffic, often because of something unrelated to the model, like older app versions being unable to reach the new routing path at all. If the split is broken, every comparison built on it is meaningless, no matter how careful the rest of the analysis is.
"""Two checks every canary needs before you trust its numbers at all."""
import numpy as np
from scipy import stats
def sample_ratio_mismatch_check(canary_n: int, control_n: int, expected_canary_frac: float):
"""Chi-square goodness-of-fit against the requested split."""
total = canary_n + control_n
expected = [total * expected_canary_frac, total * (1 - expected_canary_frac)]
observed = [canary_n, control_n]
chi2, p_value = stats.chisquare(observed, expected)
return chi2, p_value
print("healthy 5% split (1003 canary / 18997 control):")
chi2, p = sample_ratio_mismatch_check(1003, 18997, 0.05)
print(f" chi2={chi2:.3f} p={p:.3f}")
print("broken split (a routing bug sends only 2200 to canary out of 20000):")
chi2, p = sample_ratio_mismatch_check(2200, 17800, 0.05)
print(f" chi2={chi2:.3f} p={p:.6f}")healthy 5% split (1003 canary / 18997 control): chi2=0.009 p=0.922 broken split (a routing bug sends only 2200 to canary out of 20000): chi2=1515.789 p=0.000000
A p-value near 1 says the split is statistically indistinguishable from what was requested — the healthy case. A p-value essentially at zero, as in the broken case, says the observed split could not plausibly have come from the intended 5/95 rule — something in the routing itself is broken, and no metric comparison downstream of it should be trusted until it is fixed.
The actual comparison: is the canary worse on something that matters?
def error_rate_guardrail(canary_errors, canary_n, control_errors, control_n, alpha=0.05):
"""A two-proportion z-test: is the canary's error rate significantly worse?"""
p_canary = canary_errors / canary_n
p_control = control_errors / control_n
p_pooled = (canary_errors + control_errors) / (canary_n + control_n)
se = np.sqrt(p_pooled * (1 - p_pooled) * (1 / canary_n + 1 / control_n))
z = (p_canary - p_control) / se
p_value = 1 - stats.norm.cdf(z) # one-sided: is canary WORSE?
return p_canary, p_control, z, p_value, p_value < alpha
p_c, p_ctrl, z, p_val, worse = error_rate_guardrail(21, 1000, 190, 9500)
print(f"healthy canary: canary={p_c:.4f} control={p_ctrl:.4f} z={z:.3f} p={p_val:.3f} worse={bool(worse)}")
p_c, p_ctrl, z, p_val, worse = error_rate_guardrail(42, 1000, 190, 9500)
print(f"unhealthy canary: canary={p_c:.4f} control={p_ctrl:.4f} z={z:.3f} p={p_val:.5f} worse={bool(worse)}")healthy canary: canary=0.0210 control=0.0200 z=0.214 p=0.415 worse=False unhealthy canary: canary=0.0420 control=0.0200 z=4.502 p=0.00000 worse=True
The healthy canary's error rate (2.1%) is barely different from control (2.0%), and the test correctly refuses to call that a real difference — p=0.415 is nowhere near significant. The unhealthy canary's error rate is roughly double control's, and the test flags it without ambiguity. This is the automatic version of the check from evaluation gates in CI, applied to live traffic instead of an offline report.
Testing the routing and the checks
from canary import assign_bucket
from guardrail import error_rate_guardrail, sample_ratio_mismatch_check
def test_the_same_user_always_lands_in_the_same_bucket():
results = {assign_bucket("user-99", 10.0) for _ in range(20)}
assert len(results) == 1
def test_the_measured_split_is_close_to_the_requested_percentage():
users = [f"user-{i}" for i in range(20_000)]
canary_count = sum(assign_bucket(u, 5.0) == "canary" for u in users)
measured_pct = canary_count / len(users) * 100
assert 4.0 <= measured_pct <= 6.0
def test_srm_check_flags_a_badly_broken_split():
_, p = sample_ratio_mismatch_check(2200, 17800, 0.05)
assert p < 0.01
def test_guardrail_flags_a_genuinely_worse_canary():
*_, worse = error_rate_guardrail(42, 1000, 190, 9500)
assert worse # scipy returns a numpy bool here, not Python's True -- see belowpytest test_canary.py -q.... [100%] 4 passed in 0.44s
Notice the last test asserts worse, not worse is True. scipy and numpy comparisons return numpy.bool_, a distinct type from Python's built-in bool — numpy.bool_(True) is True is actually False, because is compares identity, not value. This has genuinely broken real test suites; assert on truthiness, not identity, whenever a value came from numpy or scipy.
Common mistakes
Random per-request assignment instead of sticky routing. Without hashing on a stable ID, the same user can bounce between old and new behaviour on every request, which is confusing for the user and statistically invalid for the comparison — a single user's outcome should not be split across both groups.
Reading the guardrail metric before there is enough traffic. A z-test on 30 canary requests has almost no power to detect a real problem — see statistical power and sample size. Ramping too fast, before the current step has accumulated enough traffic to say anything, defeats the entire purpose.
Never checking for sample ratio mismatch. A broken split silently invalidates every comparison built on it. Run the SRM check first, every time, before looking at any other metric — it is common enough in real systems to have its own name and its own literature.
Watching only the metric you hope improves. A pricing model canary that only tracks conversion, and never tracks complaints or refunds, can look like a clear win while quietly making things worse somewhere nobody was watching. Guardrail metrics exist specifically to catch harm in places you were not focused on.
Ramping to 100% the moment the first check passes. One green result at 5% is a start, not a conclusion. Increase gradually — 5% to 25% to 100% is a common shape — re-checking at each step, because a problem that only shows up under real scale is exactly what a single small step cannot reveal.
Try it yourself
Change error_rate_guardrail into a two-sided test (is the canary different in either direction, not only worse), and re-run it against a canary that is meaningfully better than control. Decide for yourself whether a two-sided or one-sided test is the right choice for a guardrail metric — there is a real, defensible argument for each.
What to learn next
- Blue-green model deployments — the alternative when you want an instant, all-at-once switch instead of a gradual ramp.
- Sample ratio mismatch — a deeper look at diagnosing exactly why a split went wrong.
- Rolling back a model — what actually happens the moment a canary's guardrail metric fails.
Researcher — Mathematics and papers.
The statistics underneath a guardrail check
The two-proportion z-test implemented above assumes both groups' outcomes are independent Bernoulli trials, and relies on the normal approximation to the binomial being adequate — reasonable once $n \cdot p$ and $n \cdot (1-p)$ are both comfortably above about 5 for each group, which a small early canary slice can violate. For small samples, an exact test (Fisher's exact test, scipy.stats.fisher_exact) avoids the approximation error at some cost in analytical simplicity.
Sample ratio mismatch as a diagnostic, not only a symptom
Fabijan et al. (2019), among others, document SRM as one of the most common silent failure modes in real online experimentation platforms, typically traced to one of a small number of root causes: assignment logic that runs after a filtering step correlated with the treatment (e.g., bot filtering that behaves differently for the two code paths), caching that serves a stale bucket assignment inconsistently, or client-side routing bugs affecting specific app versions or browsers disproportionately. The chi-square test in this lesson detects that something is wrong; root-causing it always requires segmenting the mismatch by platform, version, and geography rather than stopping at the aggregate p-value.
Peeking and the multiple-looks problem
A canary is, by construction, checked repeatedly as its traffic share ramps — a form of sequential testing rather than a single fixed-sample test. Evaluating a fixed-alpha significance test at every ramp step inflates the true false-positive rate well above the nominal $\alpha$, because each additional look is another chance to cross the threshold by chance alone. Peeking and sequential testing covers the corrections — group-sequential boundaries, always-valid p-values — that a production canary pipeline checked at multiple ramp stages should actually use instead of the naive single-test approach shown for clarity above.
Guardrail metric selection
A guardrail metric should satisfy three properties: it should be sensitive to the specific harms the release could plausibly cause, it should have low enough baseline variance to reach significance on realistic canary traffic volumes, and — critically — it should be pre-registered before the release, not chosen after looking at the data. Choosing a guardrail after seeing which metric moved is a direct instance of the multiple-comparisons problem covered in multiple testing correction.
Papers
- Fabijan et al., Diagnosing Sample Ratio Mismatch in Online Controlled Experiments, KDD 2019 — doi.org/10.1145/3292500.3330722
- Kohavi et al., Trustworthy Online Controlled Experiments, Cambridge University Press, 2020 — the standard reference for guardrail metrics, SRM, and sequential peeking together.
- Johari et al., Always Valid Inference: Continuous Monitoring of A/B Tests, Operations Research 2022 — the formal treatment of the multiple-looks problem this lesson's ramp steps run into.
What to learn next
- Blue-green model deployments — the alternative when you want an instant, all-at-once switch instead of a gradual ramp.
- Sample ratio mismatch — a deeper look at diagnosing exactly why a split went wrong.
- Rolling back a model — what actually happens the moment a canary's guardrail metric fails.