Experiment Design and A/B Testing

Bayesian A/B testing

Bayesian A/B testing answers the question teams actually ask — "what is the probability B is better, and how much do we lose if we are wrong?" — using beliefs updated by data.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Bayesian A/B testing treats each version's true quality as something you hold a belief about, and lets data sharpen that belief. It ends with a sentence like "there is a 94% chance B is better".

Think of tasting dal while it cooks. Before the first taste you have a rough belief about the salt. Each spoonful updates it. After three tastes you are fairly sure; after ten, confident. You never receive a certificate that the salt is correct. You hold a belief that strengthens with each taste. At some point you are sure enough to serve.

The classical approach from hypothesis testing answers a stranger question — "how often would luck produce a gap this big?". The Bayesian approach answers the one your team actually asked: how likely is it that B beats A?

Why it exists

P-values confuse almost everyone. A p-value of 0.04 does not mean "96% chance the change works" — yet that is how meetings retell it. Bayesian results say the natural sentence directly, because the machinery is built to produce it.

It also reframes the shipping decision. Instead of "significant or not", it offers expected loss: if we ship B and B is actually worse, how much do we expect to lose? When that number falls below what you would happily ignore, ship.

How it works

belief about A's rate:      ▁▂▅█▅▂▁        (centred near 5.0%)
belief about B's rate:        ▁▂▅█▅▂▁      (centred near 6.0%)

overlap small  →  "94% chance B is better"

more data  →  both humps narrow  →  the answer sharpens

Each version's conversion rate gets a "belief curve" — wide when data is scarce, narrowing as clicks arrive. The headline numbers fall out by comparing the two curves.

A real example you have seen

Most commercial testing tools — VWO, parts of Optimizely, many in-house platforms — report results in Bayesian language: "chance to beat control: 94%". Product teams adopted it because that sentence survives a meeting without being mangled, while "we reject the null at the 5% level" rarely does.

Remember this

  • Bayesian testing outputs the sentence people want: "probability B is better = 94%".
  • Beliefs start wide and narrow as data arrives.
  • Decisions use expected loss — ship when being wrong would cost less than you care about.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Verified with numpy 1.26.4 and scipy 1.14.1.

The whole method in twenty lines

For click/no-click data, beliefs about a rate are captured by a Beta distribution — a curve over possible rates. Starting from a flat prior, the update rule is nothing more than addition: successes and failures become the curve's two parameters.

bayes_ab.py
import numpy as np
from scipy import stats

rng = np.random.default_rng(1)
# the data so far: A got 120 clicks from 2400 views, B got 145 from 2400
post_a = stats.beta(1 + 120, 1 + 2400 - 120)   # Beta(1,1) prior + observed counts
post_b = stats.beta(1 + 145, 1 + 2400 - 145)

draws_a = post_a.rvs(200_000, random_state=rng)
draws_b = post_b.rvs(200_000, random_state=rng)

p_b_wins = (draws_b > draws_a).mean()
loss_if_ship_b = np.maximum(draws_a - draws_b, 0).mean()

print(f"P(B beats A) = {p_b_wins:.3f}")
print(f"expected loss if we ship B = {loss_if_ship_b:.5f} click-rate points")
print(f"B's 95% credible interval: {post_b.ppf(0.025):.4f} to {post_b.ppf(0.975):.4f}")
Output
P(B beats A) = 0.943
expected loss if we ship B = 0.00016 click-rate points
B's 95% credible interval: 0.0516 to 0.0707

The walkthrough

stats.beta(1 + wins, 1 + losses) is the entire model. The 1 + is the flat prior — believing nothing in particular before data. The posterior (the updated belief curve) needs no fitting, no iteration: conjugacy makes the update pure arithmetic.

The Monte Carlo step answers awkward questions directly. "What is the chance B beats A" has no tidy formula, but sampling 200,000 pairs and counting settles it to three decimals in milliseconds. Any question about the two rates — "chance B is at least 10% better", "distribution of the ratio" — is one line of the same shape.

Expected loss is the shipping criterion. 0.00016 points means: in the worlds where B is secretly worse, the average damage is a hundredth of a percent of click rate. If that is beneath your caring threshold, ship B — even at 94% rather than 99% certainty. This converts statistics into a business decision with units.

The credible interval reads the way people want. "The true rate lies in 5.2%–7.1% with 95% probability" is the correct reading of a credible interval — the reading that is famously wrong for classical confidence intervals.

Common mistakes

Believing Bayes makes peeking free. Checking the posterior daily and stopping when "P(B wins)" crosses 95% still inflates how often you ship duds — the posterior is honest about beliefs, but repeated decisions against a threshold still select lucky moments. Fixed horizons, or decision rules based on expected loss with a preset budget, keep you honest.

Flat priors on tiny samples. With 30 views per arm, the flat prior lets the posterior claim wild rates like 20% for a product whose historical rate never left 4–6%. A weak prior centred on history — say Beta(5, 95) — buys stability and is honest about what you knew.

Using the posterior of the difference to claim "no difference exists". A posterior straddling zero means uncertainty, not equality — the same trap as reading a big p-value as proof of no effect.

Forgetting the guardrails. A Bayesian goal metric changes nothing about guardrail discipline — damage elsewhere still vetoes a ship.

Try it yourself

Rerun with ten times the data (1,200 and 1,450 clicks from 24,000 views each) and watch P(B wins) and the credible interval sharpen. Then encode a sceptical prior — Beta(10, 190), believing rates near 5% — and see how much data B needs before the scepticism washes out.

What to learn next

Researcher — Mathematics and papers.

The conjugate model

For arm $j$ with true rate $\theta_j$, prior $\theta_j \sim \text{Beta}(\alpha_0, \beta_0)$ and likelihood $x_j \sim \text{Binomial}(n_j, \theta_j)$ give the posterior

$$ \theta_j \mid x_j \;\sim\; \text{Beta}(\alpha_0 + x_j,\; \beta_0 + n_j - x_j) $$

  • $\theta_j$ — the unknown conversion rate of arm $j$.
  • $x_j, n_j$ — observed successes and trials.
  • $\alpha_0, \beta_0$ — prior pseudo-counts of successes and failures.

The comparison quantities are functionals of the joint posterior (independent across arms):

$$ \Pr(\theta_B > \theta_A \mid \text{data}), \qquad \mathbb{E}!\left[\max(\theta_A - \theta_B, 0)\right] $$

The first has a closed form as a finite sum (Cook, 2005, Exact calculation of beta inequalities); Monte Carlo is standard practice. The second is the expected loss of choosing B; the decision rule "ship when expected loss < $\varepsilon$" is the industry-standard formulation from Stucchio (2015), Bayesian A/B Testing at VWO.

Continuous metrics and hierarchy

Revenue-per-user style metrics use Normal-Inverse-Gamma or, for zero-inflated spend, a compound model (Bernoulli purchase × LogNormal basket) with conjugate or MCMC updates. Hierarchical priors across many experiments (empirical Bayes) shrink per-experiment estimates toward the platform's base-rate distribution of effects — directly attacking the winner's curse in large experiment programmes; Azevedo et al. (2020), A/B Testing with Fat Tails (JPE), derive the optimal experimentation portfolio when true effects are heavy-tailed.

Frequentist properties of Bayesian rules

A Bayesian threshold rule evaluated repeatedly is a stopping rule, and its type-I error is governed by martingale arguments, not by the posterior's calibration. Under optional stopping with a flat prior, $\Pr(\text{ever claim } \theta_B > \theta_A \text{ at } 95\%)$ under the null exceeds 5% substantially. Fixes with guarantees:

  • Expected-loss stopping with caps — Stucchio's rule bounds the loss rather than the error rate; simulation calibration is the honest accompaniment.
  • Bayes factors / e-values with optional stopping validity — under specific mixtures, the safe-testing framework (Grünwald, de Heide and Koolen, 2024, Safe Testing, JRSS-B) gives anytime-valid guarantees, converging with the sequential testing literature.
  • Thompson sampling — using the posterior to allocate traffic rather than test, trading estimation precision for regret minimisation; see multi-armed bandits.

Reading

  • Stucchio (2015), Bayesian A/B Testing at VWO — the applied blueprint, including expected loss.
  • Gelman et al., Bayesian Data Analysis (3rd ed., 2013) — chapters 2 and 5 for conjugacy and hierarchy.
  • Deng, Lu and Chen (2016), Continuous monitoring of A/B tests without pain (WSDM) — the peeking problem analysed for Bayesian dashboards.

What to learn next