Recommender Systems

A/B testing a recommender

Offline scores rank ideas but cannot settle them, so a new recommender is decided by showing it to a random half of real users and comparing what happens — with enough traffic, fixed metrics and no peeking.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why offline scores are not enough
  4. How it is run
  5. The three things that surprise people
  6. Choose the metric before you start
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An A/B test shows the new recommender to a random half of your users, and keeps the old one for the other half. You then compare what the two halves actually do.

The analogy you have already lived

You cook dal two ways for a family dinner. Half the table gets the old recipe, half gets the new one, and nobody is told which is which.

Then you look at the plates. The honest answer is in what came back empty, not in what people said politely.

Two details make it a real test rather than a guess. The two halves were chosen at random, so it is not that all the children got one version. And both versions were served on the same evening, so the weather, the day and everyone's mood were the same for both.

That is an A/B test. Randomised, simultaneous, and judged on behaviour.

Why offline scores are not enough

Ranking metrics are computed against what people did while a different system was choosing what to show them.

So an offline score answers a question shaped like "how well does the new model agree with the old model's choices?" That is useful for catching regressions. It is useful for cutting a hundred ideas down to three. It cannot tell you what people will do when the new model is the one choosing.

There are effects an offline score cannot see at all. A slower model that shows better items may still lose, because people leave while the page loads. A more accurate model may reduce time on the app, because people find what they wanted faster. Neither of those appears in nDCG.

How it is run

   every arriving user
          |
     hash their id  ->  a stable random half
          |
     ---------------------------
     |                         |
   group A                   group B
   old recommender           new recommender
     |                         |
     -----------  compare  ------
                   |
        one metric, decided in advance

Two things in that picture are load-bearing.

The split is by user, not by visit. A person must land in the same group every time they return. Otherwise a single person sees both systems, and their behaviour belongs to neither.

The metric was chosen before the test started. Run a test, look at forty numbers afterwards, and something will look impressive by chance. Choosing afterwards turns an experiment into a search for a good-looking coincidence.

The three things that surprise people

One. You need far more traffic than you expect.

Small improvements need enormous samples to prove. The developer block computes it: with a 10 percent click rate, spotting a 1 percent relative improvement takes over 1.4 million users per group.

Most teams do not have that traffic. The honest response is to accept that small gains are not measurable for you, and to stop pretending otherwise.

Two. New things look good for a while, then stop.

Change anything visible and some people click it because it is different. This is the novelty effect, and it fades over one to three weeks. A test stopped on day three measures curiosity, not value.

The mirror image also happens. Loyal users are annoyed by change and behave worse at first, then recover. Both effects push you towards running tests for at least two weeks.

Three. Checking early ruins the answer.

This is the one that catches good engineers. Watch the numbers every day, and stop the moment they look significant, and you will declare winners that are not there.

The developer block runs two thousand experiments where both groups are identical. Waiting until the end wrongly declares a winner 5 percent of the time, which is exactly what the statistics promise. Stopping at the first good-looking day does it 21.6 percent of the time.

One in five of your "wins" would then be nothing at all.

Choose the metric before you start

A recommender can win on clicks and lose on everything that matters. Pick one primary metric, plus a short list of things that must not get worse.

  • Primary: something close to real value. Completed purchases, sessions in the following week, a satisfying watch rather than a click.
  • Guardrails: page load time, complaint rate, unsubscribes, catalogue coverage, and how much attention the top few items are taking.

Write them down before the test. A guardrail added afterwards is not a guardrail.

The honest part

Some things cannot be A/B tested cleanly, and pretending otherwise is worse than admitting it.

If your platform is a marketplace, promoting an item to group B removes stock or attention from group A. The two groups are not independent, so the comparison is contaminated. This is called interference, and it needs cluster-based designs rather than user-level ones.

Long-term effects are worse. Whether a change makes users healthier or the catalogue richer over two years cannot be read from a two-week test. Every organisation ends up making that call partly on judgement. The ones that admit it make better calls than the ones hiding behind a p-value.

Remember this

  • Randomise by user, run both versions at the same time, and fix the metric in advance.
  • Small lifts need enormous samples; check the number before running the test.
  • Do not stop early. Peeking turned a 5 percent error rate into 21.6 percent.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

The statistics use math.erf from the standard library, so there is no SciPy dependency. The whole file runs in a couple of seconds.

Power, one experiment, and the cost of peeking

ab_test.py
import math
import numpy as np

Z_ALPHA, Z_POWER = 1.959964, 0.841621          # two-sided 5% significance, 80% power


def two_proportion_test(clicks_a, n_a, clicks_b, n_b):
    """Returns the two rates, the absolute difference and a two-sided p-value."""
    p_a, p_b = clicks_a / n_a, clicks_b / n_b
    pooled = (clicks_a + clicks_b) / (n_a + n_b)
    se = math.sqrt(pooled * (1 - pooled) * (1 / n_a + 1 / n_b))
    z = (p_b - p_a) / se if se > 0 else 0.0
    p_value = 2 * (1 - 0.5 * (1 + math.erf(abs(z) / math.sqrt(2))))
    return p_a, p_b, p_b - p_a, z, p_value


def users_per_arm(baseline, relative_lift):
    p1 = baseline
    p2 = baseline * (1 + relative_lift)
    return math.ceil((Z_ALPHA + Z_POWER) ** 2 *
                     (p1 * (1 - p1) + p2 * (1 - p2)) / (p2 - p1) ** 2)


print("how many users you need before you can see a lift at all")
print(f"{'baseline CTR':>13}{'lift you hope for':>19}{'users per arm':>16}")
for base in (0.10, 0.02):
    for lift in (0.20, 0.05, 0.01):
        print(f"{base:>13.2%}{lift:>19.0%}{users_per_arm(base, lift):>16,}")

print("\none experiment, 20,000 users per arm, a real 5% lift underneath")
rng = np.random.default_rng(7)
N = 20_000
clicks_a = int(rng.binomial(N, 0.100))
clicks_b = int(rng.binomial(N, 0.105))
pa, pb, diff, z, p = two_proportion_test(clicks_a, N, clicks_b, N)
print(f"  control    {clicks_a:>6} clicks   {pa:.4f}")
print(f"  treatment  {clicks_b:>6} clicks   {pb:.4f}")
print(f"  difference {diff:+.4f}   relative {diff / pa:+.1%}   z {z:.2f}   p {p:.3f}")
print(f"  needed for 80% power at this size: {users_per_arm(0.10, 0.05):,} per arm")

print("\nwhat happens when you check the result every day (2,000 A/A tests, no real difference)")
TRIALS, CHECKS, PER_CHECK = 2000, 14, 2_000
rng = np.random.default_rng(11)
a = rng.binomial(PER_CHECK, 0.10, size=(TRIALS, CHECKS)).cumsum(axis=1)
b = rng.binomial(PER_CHECK, 0.10, size=(TRIALS, CHECKS)).cumsum(axis=1)
n = np.arange(1, CHECKS + 1) * PER_CHECK
peeked = 0
final = 0
for t in range(TRIALS):
    hits = [two_proportion_test(int(a[t, c]), int(n[c]), int(b[t, c]), int(n[c]))[4] < 0.05
            for c in range(CHECKS)]
    peeked += any(hits)
    final += hits[-1]
print(f"  stopping only at the end        : {final / TRIALS:.1%} called a winner")
print(f"  stopping the first time p < 0.05: {peeked / TRIALS:.1%} called a winner")
Output
how many users you need before you can see a lift at all
 baseline CTR  lift you hope for   users per arm
       10.00%                20%           3,839
       10.00%                 5%          57,760
       10.00%                 1%       1,419,070
        2.00%                20%          21,106
        2.00%                 5%         315,204
        2.00%                 1%       7,729,568

one experiment, 20,000 users per arm, a real 5% lift underneath
  control      1996 clicks   0.0998
  treatment    2039 clicks   0.1019
  difference +0.0021   relative +2.2%   z 0.71   p 0.475
  needed for 80% power at this size: 57,760 per arm

what happens when you check the result every day (2,000 A/A tests, no real difference)
  stopping only at the end        : 5.0% called a winner
  stopping the first time p < 0.05: 21.6% called a winner

The simulated numbers come from fixed seeds. NumPy does not promise an identical random stream across major versions, so if your digits differ slightly, read the pattern rather than the decimals.

The first table is a reality check, and you should run it on your own numbers

Detecting a 20 percent relative lift on a 10 percent baseline takes about 3,800 users per arm. Any product can do that.

Detecting a 1 percent relative lift on the same baseline takes 1,419,070 per arm. Drop the baseline to 2 percent, which is a realistic purchase rate, and the same 1 percent lift needs 7.7 million per arm.

Sit with those two numbers. Most recommender improvements are worth low single-digit percentages. Most products do not have millions of users per arm per fortnight. The conclusion is uncomfortable and correct: a great many shipped recommender changes were never actually measurable.

The honest responses are to aim for bigger changes, to use a more sensitive metric than a raw rate, or to reduce variance with the techniques in the researcher block. Running an underpowered test and reading the p-value is not on the list.

The middle block shows what underpowered looks like from the inside

The simulation was generated with a genuine 5 percent lift. Ten percent against ten-point-five percent, built into the data.

The result: p = 0.475. Nothing. The measured relative difference is +2.2%, less than half of what is really there, and it is indistinguishable from noise at this sample size.

Nobody running this test would know a real improvement had been thrown away. There is no warning in the output. This is why the power calculation comes before the experiment: it is the only step that tells you an experiment is not worth running.

The last block is the one to show your team

Both groups are identical, drawn from the same distribution, with no effect whatsoever. Every "winner" here is false by construction.

Stopping only at the end declares a winner 5.0% of the time, which is precisely the 5 percent the test promises.

Stopping the first time the dashboard shows p < 0.05 declares a winner 21.6% of the time. More than one in five.

Nothing was cheated. Nobody manipulated the data. Somebody looked at a live dashboard fourteen times and acted on the best-looking moment. That is the entire mechanism, and it is the most common way a recommender team ships an improvement that does not exist.

The fix is boring: decide the sample size in advance, and do not look at the primary metric until you reach it. If you must monitor continuously, use a method built for it — see the researcher block.

Line by line, for the parts that are not obvious

pooled = (clicks_a + clicks_b) / (n_a + n_b) — under the null hypothesis both groups share one rate, so the standard error is computed from the combined rate. Using each group's own rate gives a slightly different, less standard test.

0.5 * (1 + math.erf(abs(z) / math.sqrt(2))) — the standard normal cumulative distribution, from the error function. This is why no SciPy import is needed.

(Z_ALPHA + Z_POWER) ** 2 — the two constants combine because you need the difference to clear the noise threshold and to do so reliably. Dropping Z_POWER gives a sample size that detects the effect about half the time, which is the most common error in home-made power calculators.

.cumsum(axis=1) — each checkpoint contains all earlier data, which is what a real dashboard shows. The checks are heavily correlated, and the peeking inflation comes from taking a maximum over correlated draws.

hits[-1] versus any(hits) — one line apart, and the difference between a 5 percent error rate and a 21.6 percent one.

Common mistakes

Randomising by session instead of by user. A returning user lands in a different group and their behaviour is split across both. Hash a stable user or device id, and keep the hash function fixed forever.

Skipping the A/A test. Before trusting your platform, run an experiment where both groups get the identical system. If it reports winners more than about 5 percent of the time, your splitting or logging is broken, and every result you have ever shipped is suspect.

Ignoring a sample-ratio mismatch. You asked for a 50/50 split and got 50.4/49.6 on a million users. That is not rounding; it is a bug in assignment, filtering or logging. Test the ratio itself, and treat a failure as a stop-the-experiment event.

Measuring clicks alone. Clicks are the easiest metric to improve and the least connected to value. Pair every click metric with a downstream one: completion, purchase, or a return visit next week.

Running fifty variants and reporting the best. With fifty comparisons at 5 percent significance you expect two or three false winners. Correct for multiple comparisons or pre-register a single variant.

Forgetting that the control is also changing. Other teams ship during your test. Compare against a concurrent control, never against last month.

Try it yourself

Change PER_CHECK from 2_000 to 500 and CHECKS from 14 to 56, keeping the total sample identical. Peek more often over the same data and watch the false-winner rate climb further.

Then set the true treatment rate in the single experiment to 0.100 — no effect at all — and run it a few times with different seeds. Count how often it looks like a winner. That exercise does more for a team's statistical instincts than any amount of reading.

What to learn next

Researcher — Mathematics and papers.

The estimator

For a binary outcome with $n_A, n_B$ users per arm and observed rates $\hat{p}_A, \hat{p}_B$, the pooled two-proportion $z$ statistic is

$$ z = \frac{\hat{p}_B - \hat{p}_A}{\sqrt{\hat{p}\,(1-\hat{p})\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}}, \qquad \hat{p} = \frac{c_A + c_B}{n_A + n_B} $$

with $c_A, c_B$ the click counts. Required sample size per arm for a two-sided test at level $\alpha$ with power $1-\beta$:

$$ n = \frac{\left(z_{1-\alpha/2} + z_{1-\beta}\right)^2 \left[p_1(1-p_1) + p_2(1-p_2)\right]}{(p_2 - p_1)^2} $$

The inverse-square dependence on the effect size is the governing fact of the whole discipline. Halving the detectable effect quadruples the cost.

The ratio metric caveat. Click-through rate is a ratio of two random quantities, and the randomisation unit is the user while the denominator counts impressions. Treating each impression as independent understates the variance badly, because impressions from one user are correlated. Use the delta method for the variance of a ratio, or bootstrap over users. Reporting a naive binomial standard error on a per-impression CTR is one of the most common errors in industrial experiment reports.

Variance reduction

CUPED (Deng, Xu, Kohavi and Walker, 2013), Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data:

$$ Y^{\text{cuped}} = Y - \theta (X - \mathbb{E}[X]), \qquad \theta = \frac{\operatorname{Cov}(Y, X)}{\operatorname{Var}(X)} $$

$X$ is a pre-experiment covariate — most usefully the same metric measured on the same users before the test began. The adjusted estimator is unbiased because $X$ is independent of assignment, and its variance is reduced by a factor $(1 - \rho^2)$ with $\rho = \operatorname{Corr}(Y, X)$.

For engagement metrics $\rho$ of 0.5 to 0.8 is common, which cuts required sample size by 25 to 64 percent. It is the highest-return change available to most experimentation platforms, and it costs one extra column in the data.

Triggered analysis. Restrict the analysis to users who could have been affected — those who actually saw a recommendation slot. Diluting the effect across users who never reached the surface reduces power for no gain. The trigger condition must be computable identically in both arms, or the comparison is broken.

Sequential testing

The peeking inflation in the developer block is the classical multiple-testing problem under correlated looks. Two principled remedies:

Group sequential designs (O'Brien-Fleming, Pocock boundaries) fix the number of interim analyses in advance and spend the type-I error budget across them via an alpha-spending function.

Always-valid inference gives a $p$-value that is valid at every sample size. Johari, Koomen, Pekelis and Walsh (2017), Peeking at A/B Tests, construct a mixture sequential probability ratio test whose $p$-value process satisfies

$$ P\left( \exists\, t : p_t \le \alpha \mid H_0 \right) \le \alpha $$

so continuous monitoring is legitimate. The cost is a larger expected sample size when the effect is small — the guarantee is bought with power, not created from nothing.

Interleaving

For ranking changes specifically, interleaving is dramatically more sensitive than an A/B test. Team-draft interleaving (Radlinski, Kurup and Joachims, 2008) merges the two rankings into one list shown to each user, attributing each click to the ranker that contributed the item. Each user is their own control, so between-user variance vanishes.

Chapelle et al. (2012), Large-Scale Validation and Analysis of Interleaved Search Evaluation, report sensitivity gains of one to two orders of magnitude in required traffic relative to A/B testing, with rankings that agree with A/B conclusions.

The limits are real: interleaving compares two rankings over the same candidate pool and cannot measure changes to the interface, to latency, or to anything session-level. Use it to choose between rankers, then A/B test the winner against the incumbent for the value metrics.

Interference

User-level randomisation assumes the stable unit treatment value assumption: one user's outcome does not depend on another user's assignment. Recommenders violate it routinely.

  • Marketplaces: finite inventory means promoting an item to treatment users removes it from control users.
  • Two-sided platforms: creator-side effects propagate to consumers in both arms.
  • Shared model state: online learning updated by both arms couples them directly.

Standard remedies are cluster randomisation (by region, by supply pool, by time block) and switchback designs alternating the whole system between arms across time intervals. Both raise variance considerably, since the effective sample size drops to the number of clusters or intervals. Estimating the bias from ignoring interference is generally harder than accepting the variance cost.

The long-term problem

Two-week experiments measure two-week outcomes. Hohnhold, O'Brien and Tang (2015), Focusing on the Long-term, describe long-term holdback groups kept on the old system for months, and cohort-based analyses that estimate the accumulated effect of a sequence of shipped changes.

This is the only rigorous answer to the diversity trade-off from diversity and filter bubbles. A diversity change that costs 1 percent of clicks this fortnight and improves retention over six months is undetectable in the short experiment and visible in the holdback. Organisations without a long-term holdback are structurally unable to make that trade correctly, however good their statistics are.

Papers

  • Radlinski, Kurup and Joachims (2008), How Does Clickthrough Data Reflect Retrieval Quality?, CIKM — team-draft interleaving.
  • Chapelle et al. (2012), Large-Scale Validation and Analysis of Interleaved Search Evaluation, ACM TOIS.
  • Deng, Xu, Kohavi and Walker (2013), Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data (CUPED), WSDM.
  • Kohavi, Deng, Frasca, Walker, Xu and Pohlmann (2013), Online Controlled Experiments at Large Scale, KDD.
  • Hohnhold, O'Brien and Tang (2015), Focusing on the Long-term: It's Good for Users and Business, KDD.
  • Johari, Koomen, Pekelis and Walsh (2017), Peeking at A/B Tests: Why it matters, and what to do about it, KDD.
  • Kohavi, Tang and Xu (2020), Trustworthy Online Controlled Experiments — the standard book-length treatment.

What to learn next