Experiment Design and A/B Testing

Randomisation units and assignment

Randomisation decides who sees which version of your product, and the unit you randomise — person, session or city — quietly decides what your experiment can prove.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Randomisation means letting a coin toss — not a person — decide who gets the new version and who gets the old one.

Think of dividing a class into two cricket teams. If the two captains pick players, the teams end up lopsided — every good batsman on one side. If you pull names from a bowl instead, luck spreads the strong and weak players evenly. An A/B test — showing version A to some users and version B to others, then comparing — works the same way. The coin toss builds two groups that are alike in every way except the thing you changed.

Why it exists

Suppose you launch a new checkout page and let users opt in. The people who opt in are the curious, adventurous ones. They were always going to buy more. Now you cannot tell whether the page helped or whether eager people gathered on one side.

Randomisation removes the choosing. When chance builds the groups, they match on everything — age, city, phone type, mood — even things you never measured. Any gap that appears afterwards must come from your change.

How it works

user id  →  scrambler  →  heads: version A
                          tails: version B

same user tomorrow  →  same scrambler  →  same side, always

The scrambler is not a fresh coin each visit. It is a fixed recipe, so one user always lands on the same side. Imagine a user seeing the new checkout in the morning and the old one after lunch. They would be confused, and your data would be soup.

There is one more choice hiding here: the randomisation unit — the thing that gets its own coin toss. It can be a person, a single visit, or even a whole city. Toss per person, and a person's visits all stay together. Toss per visit, and the same person flips sides between visits. The unit decides who must stay together.

A real example you have seen

Google, Amazon and Netflix each run thousands of experiments a year. Right now, your YouTube homepage is almost certainly part of several. A friend's app can look slightly different from yours — different button colour, different ranking — because a scrambler put you on different sides.

Remember this

  • Randomisation lets chance build the groups, so they match on everything you never measured.
  • The same user must land on the same side every time.
  • Choose the randomisation unit — person, session, city — before anything else.

What to learn next

Developer — Code and libraries.

Setup

No installs needed — this uses only Python's standard library.

Deterministic assignment with a hash

Production systems do not store a giant table of who-got-what. They recompute the answer from a hash — a function that scrambles text into a large, stable number.

assign.py
import hashlib

def assign(user_id: str, experiment: str) -> str:
    # hash of experiment + user: stable for a user, fresh across experiments
    key = f"{experiment}:{user_id}".encode()
    h = int(hashlib.md5(key).hexdigest(), 16)
    return "treatment" if h % 2 else "control"

print(assign("user_1042", "new-checkout"))
print(assign("user_1042", "new-checkout"))   # same user -> same arm, every time
print(assign("user_1042", "new-ranking"))    # new experiment -> a fresh coin toss

counts = {"control": 0, "treatment": 0}
for i in range(100_000):
    counts[assign(f"user_{i}", "new-checkout")] += 1
print(counts)
Output
treatment
treatment
control
{'control': 50268, 'treatment': 49732}

The walkthrough

The experiment name goes into the hash. Hash only the user id, and every experiment splits users identically — the same people always in treatment. Your experiments stop being independent, and a leftover effect from one contaminates the next. Salting the hash with the experiment name reshuffles everyone each time.

h % 2 maps the huge number to two arms. For a 90/10 split, use h % 100 < 10. Hash output spreads evenly, so proportions come out as configured.

The split is 50,268 to 49,732 — not exactly 50/50. That wobble is normal coin-toss noise. A/A tests and sample ratio checks tell you when a wobble is too big to be luck.

No database of assignments exists. Any server, any day, recomputes the same answer from the same string. That is what makes assignment stateless and cheap.

Common mistakes

Calling random.random() at page load. The user flips arms on every visit. Their experience turns inconsistent, and each person's metrics become a blend of both versions. The fix: hash a stable id, never draw fresh randomness.

Randomising by session while measuring per user. A user with nine sessions appears nine times, sometimes on both sides. Per-user metrics like "did this person subscribe this month" become undefined. Match the unit to the metric, or aggregate up before analysing.

Ignoring shared accounts and devices. A family tablet is one "user" to your hash but four people to reality. This blurs the treatment and weakens measured effects. Know what your id actually identifies.

Letting each arm trigger assignment at a different point. If the new flow's entry button is more visible, more users "enter" treatment. The groups are no longer comparable at the door. Assign at a point both arms share.

Try it yourself

Change the modulus line to give treatment 10% of users, and verify the counts. Then remove the experiment name from the key and assign the same 100,000 users in two different experiments. Count how many users land on the same arm in both — and explain the number you see.

What to learn next

Researcher — Mathematics and papers.

Why randomisation identifies the causal effect

In potential-outcomes notation, each unit $i$ has $Y_i(1)$ and $Y_i(0)$, and the estimand is the average treatment effect:

$$ \tau = \mathbb{E}[Y(1) - Y(0)] $$

  • $Y(1), Y(0)$ — the outcomes unit $i$ would have under treatment and under control.
  • $\tau$ — the average treatment effect (ATE) over the population.

Randomisation makes treatment $T$ independent of both potential outcomes: $T \perp (Y(0), Y(1))$. Then

$$ \mathbb{E}[Y \mid T=1] - \mathbb{E}[Y \mid T=0] = \mathbb{E}[Y(1)] - \mathbb{E}[Y(0)] = \tau $$

so the plain difference in group means is unbiased for the ATE. No model, no adjustment, no functional-form assumptions. This is the entire epistemic advantage of experiments over observational data, developed further in potential outcomes.

The unit determines the variance, not only the semantics

With $n$ units per arm and outcome variance $\sigma^2$, the difference-in-means estimator has variance $2\sigma^2 / n$. Randomise clusters (cities, stores, classrooms) instead of individuals, and effective sample size shrinks by the design effect:

$$ \text{deff} = 1 + (m - 1)\,\rho $$

  • $m$ — average units per cluster.
  • $\rho$ — the intraclass correlation (ICC): the fraction of outcome variance attributable to between-cluster differences.

With $m = 1000$ users per city and a modest $\rho = 0.01$, deff $\approx 11$: you need eleven times the sample for the same power. This is the price of cluster randomisation, paid willingly when interference between users would otherwise bias the estimate. The formula is due to Kish (1965), Survey Sampling.

Assignment mechanics at scale

  • Hash-based bucketing (MD5 or SipHash of salt:unit_id) approximates i.i.d. assignment with zero storage; uniformity of the hash makes bucket proportions exact in expectation.
  • Layered designs (Tang et al., 2010, Overlapping Experiment Infrastructure, KDD — Google's system) run many experiments concurrently: layers partition parameters, and per-layer salts make assignments independent across layers.
  • Stratified or blocked randomisation pre-balances known covariates; at online sample sizes the gain is usually small next to CUPED-style post-adjustment.

Reading

  • Fisher (1935), The Design of Experiments — where randomised assignment as an inferential principle begins.
  • Kohavi, Tang and Xu (2020), Trustworthy Online Controlled Experiments — the standard industry reference; randomisation units are chapter 14.
  • Imbens and Rubin (2015), Causal Inference for Statistics, Social, and Biomedical Sciences — the formal treatment.

What to learn next

What to learn next

These follow on from what you just read.

  • Experiment Design and A/B Testing

    A/A tests

    An A/A test shows the same version to both groups on purpose — if your experiment system finds a "winner" anyway, the system itself is broken.

  • Experiment Design and A/B Testing

    Peeking and sequential testing

    Checking your experiment repeatedly and stopping the moment it looks significant quietly multiplies your false alarms — sequential testing is how to look early without lying.

  • Experiment Design and A/B Testing

    Sample ratio mismatch

    Sample ratio mismatch is when your 50/50 split arrives as something else — a tiny imbalance in user counts that signals the experiment is silently broken.