Experiment Design and A/B Testing

A/A tests

An A/A test shows the same version to both groups on purpose — if your experiment system finds a "winner" anyway, the system itself is broken.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An A/A test splits users into two groups and shows both groups the exact same thing — to check that your testing machinery finds no difference.

Before weighing gold, a jeweller puts identical weights on both pans of the balance. If the needle tilts, the balance is faulty — and every weighing after that would lie. An A/A test is the same trick for experiments. Both groups get identical treatment, so any "difference" the system reports is a fault in the system.

Why it exists

An A/B testing setup has many parts: the splitter, the logging, the metric pipeline, the statistics. Any one of them can be quietly broken. Maybe the splitter favours heavy users on one side. Maybe one arm's events get dropped by a buggy logger.

You cannot see these faults by running A/B tests, because a real difference and a broken pipeline look identical from outside. The A/A test removes the real difference on purpose. Whatever remains is the machinery talking.

How it works

all users → coin toss → group 1 → same old page → measure
                      → group 2 → same old page → measure

healthy system:  the two numbers differ only by luck
broken system:   a "significant winner" appears from nowhere

One catch: even a perfect system cries wolf occasionally. Statistics works with a small allowed rate of false alarms. So one alarming A/A result means little — the pattern across many A/A runs is what tells the story.

A real example you have seen

Microsoft's experimentation team famously runs continuous A/A tests on Bing. Several large companies found real bugs this way — logging delays that hit one arm first, bots piling into one bucket, caches serving one group faster. The balance-pan check catches what code review missed.

Remember this

  • An A/A test gives both groups the same experience — on purpose.
  • Its job is to test the testing system, not the product.
  • A healthy system still fires a few false alarms — the pattern over many runs is the real signal.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Verified with numpy 1.26.4 and scipy 1.14.1. Exact counts below depend on the seed only — they repeat as printed on any machine with these versions.

Two thousand A/A tests in a few seconds

A p-value is the probability of seeing a gap this large when no real difference exists. Under a true null, p-values should land anywhere between 0 and 1 with equal chance — a flat histogram. That flatness is the health certificate.

aa_simulation.py
import numpy as np
from scipy import stats

rng = np.random.default_rng(42)
runs, hits = 2000, 0
bins = [0] * 10                          # p-values sorted into ten 0.1-wide bins
for _ in range(runs):
    a = rng.normal(250, 60, 1000)        # both arms drawn from the SAME world
    b = rng.normal(250, 60, 1000)        # so any "difference" is pure luck
    p = stats.ttest_ind(a, b).pvalue
    hits += p < 0.05
    bins[min(int(p * 10), 9)] += 1

print(f"significant at 0.05: {hits} of {runs} ({hits/runs:.1%})")
print("p-value histogram, 0.0 to 1.0:", bins)
Output
significant at 0.05: 100 of 2000 (5.0%)
p-value histogram, 0.0 to 1.0: [184, 182, 200, 187, 231, 200, 193, 192, 202, 229]

The walkthrough

5.0% of runs were "significant" — and that is correct behaviour. Testing at the 0.05 level means accepting a 5% false alarm rate. An A/A suite that never fires is as suspicious as one that fires constantly.

The histogram is roughly flat — around 200 per bin. This is the deeper check. A healthy pipeline produces uniform p-values under the null. A bump near zero means too many false wins: peeking, dependent data, or a broken variance formula. A bump near one usually means duplicated or correlated observations.

Both arms come from one distribution. In production you do this with real traffic: run the splitter, serve everyone the identical page, log and analyse as usual. The simulation shows the expected shape; the production A/A shows whether your real pipeline matches it.

What real A/A failures look like

  • Too many wins at 0.05 — often the analysis assumes independent users while the data has repeated sessions per user. The fix is analysing at the randomisation unit.
  • A lopsided user count — the splitter or logging is dropping one arm's traffic. That specific failure has its own alarm: sample ratio mismatch.
  • One metric always wins in A/A — that metric's pipeline computes the arms differently, perhaps a timezone or caching asymmetry.

Common mistakes

Running one A/A test and declaring victory. One run tells you almost nothing — it passes 95% of the time even when subtly broken. Run hundreds (or simulate from logged data) and look at the p-value distribution.

Panicking at a single significant A/A. One in twenty is the contract you signed at the 0.05 level. Investigate patterns, not single fires.

Using A/A data to "calibrate away" a bias. If A/A shows a systematic tilt, find the bug. Subtracting the tilt hides a fault that will grow back in a different shape.

Try it yourself

Break the simulation on purpose: draw arm b with standard deviation 90 while keeping the same mean, and re-run. Watch what the t-test does to the histogram. Then instead give arm b a mean of 253 and see how many of the 2000 runs now "win" — that number is your power against a small effect, for free.

What to learn next

Researcher — Mathematics and papers.

Uniformity of the null p-value distribution

For a continuous test statistic $S$ with null CDF $F$, the p-value $P = 1 - F(S)$ satisfies $P \sim \text{Uniform}(0,1)$ under $H_0$ — the probability integral transform.

  • $S$ — the test statistic (here, Welch's or Student's $t$).
  • $F$ — its cumulative distribution under the null hypothesis.
  • $P$ — the p-value as a random variable.

Every deviation from uniformity is diagnostic information. Kolmogorov–Smirnov or Anderson–Darling tests against Uniform(0,1) turn an A/A suite into a formal goodness-of-fit test of the whole pipeline.

What A/A tests estimate beyond validity

  1. The realised false positive rate at level $\alpha$: proportion of A/A runs with $p < \alpha$. Deviations flag violated test assumptions — unequal variances, heavy tails at small $n$, dependence.
  2. Metric variance for power calculations. The empirical variance of the A/A difference is the honest input to sample-size formulas, capturing clustering and mixed populations that textbook formulas miss.
  3. Pre-experiment covariate quality. The correlation between pre-period and in-period metrics measured on A/A data determines how much CUPED will help.

Dependence: the usual killer

With session-level observations nested in users, the variance of a user-mean metric is inflated by intra-user correlation. Analysing sessions as independent understates the standard error by roughly $\sqrt{1 + (m-1)\rho}$ (with $m$ sessions per user and intra-user correlation $\rho$), producing excess A/A failures. The delta method or user-level aggregation restores calibration — Deng, Knoblich and Lu (2018), Applying the Delta Method in Metric Analytics, KDD.

Reading

  • Kohavi, Tang and Xu (2020), Trustworthy Online Controlled Experiments, chapter 19 — A/A tests as continuous infrastructure monitoring.
  • Kohavi, Longbotham, Sommerfield and Henne (2009), Controlled experiments on the web: survey and practical guide, Data Mining and Knowledge Discovery 18 — the early practical account, including A/A failures found at Microsoft.
  • Fabijan et al. (2019), Diagnosing Sample Ratio Mismatch in Online Controlled Experiments, KDD — what to check when the A/A alarm is a count imbalance.

What to learn next