Experiment Design and A/B Testing
CUPED and variance reduction
CUPED uses each user's own pre-experiment behaviour to cancel predictable noise, shrinking error bars and letting the same experiment detect smaller effects faster.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
CUPED makes experiments sharper by comparing each user against their own past, instead of comparing raw totals.
Think of judging a tuition centre. Compare the final exam marks of its students against everyone else, and the comparison is dominated by who walked in. Toppers score high with or without tuition. But compare each student's marks against their own last year's marks, and the noise of "who walked in" cancels. What remains is the improvement — the thing you actually wanted.
CUPED (Controlled-experiment Using Pre-Experiment Data) does exactly this with product metrics. A user who spent a lot last month will likely spend a lot this month. That predictable part is noise for your experiment, and it can be subtracted.
Why it exists
The biggest cost in experimentation is waiting. Big companies can detect a 0.1% change in days because they have oceans of users. Everyone else waits weeks — the wobble from user-to-user differences drowns small effects.
But much of that wobble is predictable before the experiment even starts. Heavy users stay heavy; occasional users stay occasional. CUPED removes the part of the wobble that last month already told you about. Same users, same duration — noticeably narrower error bars.
How it works
raw comparison: this month's spend → huge spread per person
CUPED comparison: this month's spend
- what last month predicted for that person
─────────────────────────────────────────
= the surprise → much smaller spread
groups differ only in the surprise → the effect stands out soonerOne promise makes this safe: the past data comes from before the experiment began. The treatment cannot have touched it, so subtracting it cannot smuggle in bias. It can shrink the noise — never bend the answer.
A real example you have seen
Microsoft introduced CUPED for Bing and reported cutting required experiment duration roughly in half for key metrics. Netflix, Airbnb and Booking.com all run versions of it as the default. When a streaming app measures whether a new player raises watch time, last month's watch time is doing quiet work in the background.
Remember this
- CUPED subtracts the part of each user's result that their pre-experiment behaviour already predicted.
- Same data, same users — smaller error bars, faster answers.
- Pre-experiment data is untouched by treatment, so the adjustment is safe by construction.
What to learn next
- Novelty and primacy effects — why the first week of an experiment can mislead even with tight error bars.
- Double machine learning — the same residual-on-residual idea, grown into a full causal method.
- A/A tests — where to measure your covariate correlations safely.
Developer — Code and libraries.
Setup
pip install numpyVerified with numpy 1.26.4.
CUPED in fifteen lines
The recipe: measure how strongly pre-experiment spend predicts in-experiment spend (a number called theta), subtract that predictable part from every user, then run the ordinary comparison on what remains.
import numpy as np
rng = np.random.default_rng(3)
n = 5000
pre = rng.gamma(2, 150, 2 * n) # last month's spend, per user
spend = 100 + 0.6 * pre + rng.normal(0, 80, 2 * n) # habit carries over
spend[:n] += 12 # true treatment effect: +12 rupees
t, c = spend[:n], spend[n:]
naive = t.mean() - c.mean()
se_naive = np.sqrt(t.var() / n + c.var() / n)
theta = np.cov(spend, pre)[0, 1] / pre.var() # how strongly pre-spend predicts spend
adj = spend - theta * (pre - pre.mean()) # subtract the predictable part
cuped = adj[:n].mean() - adj[n:].mean()
se_cuped = np.sqrt(adj[:n].var() / n + adj[n:].var() / n)
print(f"naive: {naive:+6.2f} (standard error {se_naive:.2f})")
print(f"CUPED: {cuped:+6.2f} (standard error {se_cuped:.2f})")
print(f"variance removed: {1 - adj.var() / spend.var():.1%}")naive: +11.62 (standard error 3.01) CUPED: +13.34 (standard error 1.58) variance removed: 72.3%
The walkthrough
Both estimates are near the truth (+12) — CUPED's error bar is half the size. The standard error fell from 3.01 to 1.58. Since required sample size scales with the square of the standard error, this experiment now needs roughly a quarter of the users for the same certainty.
theta is a regression slope. It is computed by pooling both arms — that keeps it identical for treatment and control, which is part of why no bias can enter. Its value here is close to 0.6, the strength we built into the simulation.
pre.mean() is subtracted inside the adjustment so the adjusted metric keeps the same average as the original. The estimate stays in rupees, directly comparable to the naive one.
72.3% of variance removed matches the theory: the reduction equals the squared correlation between pre and post. Correlation 0.85 → about 72% gone. Your gain depends entirely on how predictive your pre-period metric is.
Common mistakes
Using pre-period data collected after assignment. If treatment can touch the covariate, subtracting it does bias the estimate — this is the same poison as adjusting for a post-treatment variable in causal DAGs. "Pre" must mean strictly before the experiment start.
Expecting magic on new users. New users have no pre-period. Their theta-adjustment does nothing (use the mean for missing values, contributing zero reduction). Products dominated by new users gain little from CUPED.
Computing theta separately per arm. Slightly different thetas across arms reintroduce a bias term. Pool the arms — one theta for everyone.
Choosing a pre-metric different from the outcome metric without checking correlation. CUPED with a weakly correlated covariate is harmless but useless. Measure the correlation on A/A data first; below about 0.3 the gain is under 10%.
Try it yourself
Weaken the habit line to 0.2 * pre and rerun — watch the variance reduction collapse, and connect it to the correlation. Then break the rules on purpose: recompute pre as pre + 0.5 * (np.arange(2 * n) < n), mimicking a covariate leaked from the experiment period, and watch the CUPED estimate drift off the truth.
What to learn next
- Novelty and primacy effects — why the first week of an experiment can mislead even with tight error bars.
- Double machine learning — the same residual-on-residual idea, grown into a full causal method.
- A/A tests — where to measure your covariate correlations safely.
Researcher — Mathematics and papers.
The estimator
For outcome $Y$ and pre-experiment covariate $X$ (unaffected by treatment), define the adjusted metric
$$ \tilde{Y} = Y - \theta (X - \bar{X}), \qquad \theta = \frac{\operatorname{Cov}(Y, X)}{\operatorname{Var}(X)} $$
- $Y$ — the in-experiment metric; $X$ — the pre-experiment covariate.
- $\theta$ — the pooled regression slope of $Y$ on $X$.
- $\bar{X}$ — the grand mean of $X$ (centring keeps $\mathbb{E}[\tilde{Y}] = \mathbb{E}[Y]$).
Because randomisation gives $\mathbb{E}[X \mid T=1] = \mathbb{E}[X \mid T=0]$, the adjustment term has equal expectation in both arms, so $\mathbb{E}[\tilde{Y}_t - \tilde{Y}_c] = \mathbb{E}[Y_t - Y_c]$: unbiasedness is inherited, not assumed. The variance becomes
$$ \operatorname{Var}(\tilde{Y}) = \operatorname{Var}(Y)\,(1 - \rho^2) $$
with $\rho$ the correlation between $Y$ and $X$. The optimal single-covariate coefficient is exactly the OLS slope; with several covariates, $\theta$ becomes the vector $\operatorname{Var}(X)^{-1}\operatorname{Cov}(X, Y)$ and $\rho^2$ becomes the multiple $R^2$.
Equivalences
CUPED with a pre-period covariate is ANCOVA — regression of $Y$ on treatment and $X$ — restricted to randomised data; it is also the regression-adjustment estimator of Lin (2013), Agnostic notes on regression adjustments to experimental data, Annals of Applied Statistics, which shows arm-specific slopes with interaction terms are never asymptotically worse. The original industrial formulation is Deng, Xu, Kohavi and Walker (2013), Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data, WSDM.
Extensions
- CUPAC / ML-adjusted CUPED: replace $\theta X$ with $\hat{f}(X)$ from any ML model trained on pre-experiment data to predict $Y$ (DoorDash's CUPAC; also "MLRATE", Guo et al. 2021). Validity needs the prediction built from pre-treatment inputs only; cross-fitting avoids overfitting bias — the same trick as in double machine learning.
- Binary and ratio metrics: variance reduction applies through the delta method; conversion metrics typically see 10–30% reductions versus 40–70% for continuous engagement metrics.
- Triggered analyses: dilution correction plus CUPED compound — both shrink the denominator of detectable effect.
Practical numbers
Reported reductions: Bing ~50% variance on query share (Deng et al., 2013); Netflix 30–60% on engagement; effects equivalent to doubling-to-quadrupling traffic. The binding constraint is covariate availability — new-user-heavy surfaces and cold-start products see the least gain.
What to learn next
- Novelty and primacy effects — why the first week of an experiment can mislead even with tight error bars.
- Double machine learning — the same residual-on-residual idea, grown into a full causal method.
- A/A tests — where to measure your covariate correlations safely.