Experiment Design and A/B Testing
Sample ratio mismatch
Sample ratio mismatch is when your 50/50 split arrives as something else — a tiny imbalance in user counts that signals the experiment is silently broken.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Sample ratio mismatch (SRM) is when the number of users in each group does not match the split you configured. It is a sign that something is quietly eating your data.
Imagine dealing a deck of cards to two players, one card each, alternating. At the end, one player holds 60 cards and the other 40. You do not shrug at that — a deal like that means cards were dropped, stuck together, or palmed. The deal itself was broken, so no game played with it can be fair.
An experiment configured 50/50 should produce nearly equal groups. "Nearly" — coin tosses wobble. But past a point, the wobble is too large for luck, and the alarm must ring.
Why it exists
The dangerous part is why the counts drift apart. Users rarely vanish at random. The new page crashes on old phones — so old-phone users silently drop out of treatment. A redirect loses impatient users. A bot filter eats one arm more than the other.
Each cause removes a specific kind of user from one side. The groups stop being twins. Every comparison after that is polluted, even if the remaining data looks plentiful. A small SRM regularly hides a large bias.
How it works
configured: 50.0% | 50.0%
observed: 50.8% | 49.2% from 100,000 users
is a 0.8% tilt luck? → statistics says: almost never
→ ALARM: find the leak before reading any metricThe check is a routine statistical test on the two counts. It answers one question: could a fair coin produce a tilt this size, this often? With big numbers, even tiny tilts become impossible to blame on luck.
A real example you have seen
Microsoft reported that around 6% of their experiments trigger SRM alarms — and nearly all turned out to be genuine bugs, not luck. One famous case: a new feature loaded slower, timing-out exactly the users on weak connections, who then never got logged. The metric looked great because struggling users had been deleted from one side.
Remember this
- SRM means the group sizes disagree with the configured split beyond luck.
- The missing users are almost never random — so the surviving comparison is biased.
- Check SRM before looking at any metric; a failed check voids the experiment.
What to learn next
- Colliders and selection bias — the causal structure behind why lost users poison comparisons.
- A/A tests — the broader health check SRM belongs to.
- Guardrail metrics — the other checks that run before you celebrate a win.
Developer — Code and libraries.
Setup
pip install scipyVerified with scipy 1.14.1.
The chi-square check
A chi-square test compares observed counts against expected counts and returns the probability that luck alone explains the gap.
from scipy import stats
# two experiments, both configured for a 50/50 split
splits = [("experiment 1", 50_211, 49_789),
("experiment 2", 50_800, 49_200)]
for name, control, treatment in splits:
total = control + treatment
res = stats.chisquare([control, treatment], f_exp=[total / 2, total / 2])
print(f"{name}: control={control:,} treatment={treatment:,} "
f"p={res.pvalue:.7f}")experiment 1: control=50,211 treatment=49,789 p=0.1820462 experiment 2: control=50,800 treatment=49,200 p=0.0000004
The walkthrough
Experiment 1 is healthy. A 422-user gap in 100,000 happens by luck about 18% of the time. No alarm.
Experiment 2 is broken — and would fool your eye. 50.8% versus 49.2% looks close enough. The test says a fair split produces this less than once in a million runs. Someone is losing users, and you must find out who before touching any metric.
The convention: alarm when the SRM p-value drops below 0.001. That threshold is deliberately strict — this check runs on every experiment every day, and you want alarms you always investigate. Note the reversal from normal testing: here a small p-value is bad news.
Where to hunt once it fires:
- Assignment — is the hash actually uniform? Was the split changed mid-flight? (Changing allocation mid-experiment causes SRM against the new ratio.)
- Execution — does one arm crash, redirect, or time out more? Slow variants lose impatient users before logging fires.
- Logging — do both arms log at the same point in the flow? A treatment that logs after a new interstitial loses everyone who bounces at it.
- Filtering — bot removal, deduplication and "activity" filters applied after assignment can eat arms unevenly.
Common mistakes
Reading metrics from an experiment with SRM anyway. The lost users are selected, not random — the bias can dwarf any plausible treatment effect. Fix the leak and rerun; do not "adjust".
Checking counts only at the end. SRM from a deploy bug is findable on day one. Automate the check into the daily dashboard — unlike metric peeking, count-checking causes no statistical harm (peeking applies to effects, not sanity checks).
Testing the ratio with the wrong expected split. A 90/10 experiment checked against 50/50 fires forever. The expected counts must reflect the configured allocation.
Averaging SRM away across segments. Counts can balance overall while badly mismatched inside a segment (say, iOS only) — a segment-level bug with an overall alibi. Check the ratio within any segment you plan to report.
Try it yourself
Find the smallest imbalance the test can catch: keep the total at 100,000 and move users from control to treatment until the p-value crosses 0.001. Then repeat with a total of 10,000. What you learn: SRM detection sharpens as traffic grows — big experiments cannot hide even tiny leaks.
What to learn next
- Colliders and selection bias — the causal structure behind why lost users poison comparisons.
- A/A tests — the broader health check SRM belongs to.
- Guardrail metrics — the other checks that run before you celebrate a win.
Researcher — Mathematics and papers.
The test
Under $H_0$ (assignment follows the configured ratio), the observed counts $(n_1, n_2)$ follow a binomial split of $N = n_1 + n_2$. The Pearson statistic
$$ X^2 = \sum_{i=1}^{2} \frac{(n_i - E_i)^2}{E_i} \;\sim\; \chi^2_1 $$
- $n_i$ — observed count in arm $i$.
- $E_i$ — expected count, $N \cdot p_i$ for configured proportion $p_i$.
- $\chi^2_1$ — chi-square distribution with one degree of freedom.
An exact binomial test is preferable at small $N$; at experimentation scale the chi-square approximation is indistinguishable. Detectable imbalance shrinks as $O(1/\sqrt{N})$, and the threshold bites fast: at $N = 10^5$ a 0.5-point deviation from 50.0 gives $X^2 = 10$, $p \approx 1.6 \times 10^{-3}$ — under suspicion but short of the 0.001 alarm — while the developer block's 0.8-point deviation gives $X^2 = 25.6$, $p \approx 4 \times 10^{-7}$. Quadrupling $N$ halves the deviation the alarm can catch.
Why SRM implies bias, formally
Let $S \in {0,1}$ indicate surviving into the logged sample. SRM means $\Pr(S = 1 \mid T = 1) \ne \Pr(S = 1 \mid T = 0)$. The analysed contrast is
$$ \mathbb{E}[Y \mid T=1, S=1] - \mathbb{E}[Y \mid T=0, S=1] $$
Conditioning on $S$ — a variable affected by treatment — breaks the exchangeability randomisation bought. This is precisely collider/selection bias: treatment and user traits both cause survival, so conditioning on survival links them. The bias magnitude is unbounded by the SRM magnitude — a 0.1% count gap can carry a large metric bias if the lost users are extreme.
Taxonomy and practice
Fabijan et al. (2019), Diagnosing Sample Ratio Mismatch in Online Controlled Experiments (KDD), catalogue root causes from thousands of Microsoft experiments across five stages: assignment, execution, log processing, analysis, and interference — with observed SRM rates near 6% of experiments. Kohavi, Tang and Xu (2020), chapter 21, treat SRM as the first of the "trust checks" gating any metric read-out.
Related guards worth automating alongside SRM:
- Pre-period balance: arms should match on pre-experiment metrics; imbalance despite matching counts flags assignment correlation bugs.
- Segment SRM: run the chi-square within key segments (platform, country, new/returning) — Simpson-style cancellation across segments occurs in practice.
- Triggered-analysis SRM: when analysing only users who triggered the feature, the trigger condition must be computable identically in both arms (counterfactual triggering); otherwise trigger-stage SRM appears.
What to learn next
- Colliders and selection bias — the causal structure behind why lost users poison comparisons.
- A/A tests — the broader health check SRM belongs to.
- Guardrail metrics — the other checks that run before you celebrate a win.