Experiment Design and A/B Testing
Interference and network effects
When treated users affect control users — through friendships, shared markets or shared drivers — the A/B comparison measures something the launch will never deliver.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Interference means one user's treatment leaks into another user's outcome — the two groups stop being separate worlds, and the comparison stops meaning what you think.
Vaccinate half a village and the unvaccinated half also falls sick less — fewer sick neighbours to catch anything from. Comparing the two halves now understates the vaccine, because the control group quietly received some of its benefit. The experiment's two groups were supposed to live in separate worlds. In a village, nobody does.
Apps are villages. Users message each other, compete for the same rickshaws, bid on the same auctions.
Why it exists
The classic A/B test rests on a hidden promise: my treatment touches my outcome only, never yours. Statisticians call this SUTVA — the assumption that no unit's outcome depends on anyone else's assignment.
Social products break it through connections: give half of one friend group a new video-call feature and their control friends receive more calls. Marketplaces break it through scarcity: make treated buyers keener, and they buy up inventory control buyers would have bought. In both cases, treatment helps one group at the expense of — or as a gift to — the other.
How it works
what you want to compare:
world A: everyone has the feature world B: nobody has it
what the A/B test gives you:
one shared world, both groups inside it,
treated users eating from the same plate as control users
→ control is contaminated → the difference is distortedThe distortion has a direction. When treatment users take from a shared pool (rides, stock, ad slots), control looks worse than a real no-launch world — the lift is overstated. When treatment spills benefits onto control (network features, vaccines), the lift is understated.
A real example you have seen
Food delivery apps learned this expensively. A feature that makes treated customers order faster grabs delivery riders sooner. Control customers wait longer than they would in a world without the feature. Not because the feature is good, but because the experiment made them compete with it. The A/B dashboard shows a big win; the citywide launch shows nothing.
Remember this
- A/B testing assumes separate worlds; friends, markets and shared fleets connect them.
- Taking from a shared pool inflates measured lift; spillover of benefits deflates it.
- When users interact, redesign the experiment — switchbacks and cluster splits exist for this.
What to learn next
- Switchback experiments — the time-slicing design built for shared markets.
- Randomisation units and assignment — cluster randomisation and its variance cost.
- Potential outcomes and counterfactuals — the formal language SUTVA lives in.
Developer — Code and libraries.
Setup
pip install numpyVerified with numpy 1.26.4.
Simulating a marketplace experiment that lies
Two thousand buyers want 900 items. Treatment makes buyers keener. The A/B test splits buyers half and half — inside one shared market.
import numpy as np
rng = np.random.default_rng(5)
n_users, stock = 2000, 900 # more interested buyers than items
def bookings(treated):
p = np.where(treated, 0.60, 0.50) # the new page makes users keener to book
wants = rng.random(n_users) < p
ids = np.flatnonzero(wants)
rng.shuffle(ids) # shared stock: first come, first served
served = np.zeros(n_users, bool)
served[ids[:stock]] = True
return served
# the A/B test: half of users treated, all fighting over the same stock
mask = np.zeros(n_users, bool)
mask[:n_users // 2] = True
s = bookings(mask)
ab_lift = s[mask].mean() - s[~mask].mean()
# the truth: run the whole market one way, then the other
all_treated = bookings(np.ones(n_users, bool)).mean()
all_control = bookings(np.zeros(n_users, bool)).mean()
print(f"the A/B test claims: {ab_lift:+.3f} bookings per user")
print(f"launching to everyone: {all_treated - all_control:+.3f}")the A/B test claims: +0.078 bookings per user launching to everyone: +0.000
The walkthrough
The A/B test reports a 7.8-point lift. The launch delivers zero. Not a small bias — the entire measured effect is an artifact. Demand already exceeded the 900 items, so making buyers keener changes who gets an item, never how many are sold.
Where the fake lift comes from: treated users want items more often, so they claim more of the fixed stock — directly out of control users' hands. Control's rate drops below what a no-launch world would show. The comparison measures cannibalisation, not creation.
The bookings function is the whole trick. Both arms flow into one stock counter. Comparing bookings(all ones) against bookings(all zeros) simulates the two separate worlds the A/B test pretends to have — a luxury reality never grants, which is exactly why this bias is so persistent.
Real severity depends on supply tightness. With stock = 2000 (nobody constrained), rerun and watch the A/B estimate become honest. Interference bias scales with how contested the shared resource is.
Designs that survive interference
- Cluster randomisation: assign whole friend-groups, cities or markets to one arm, so leakage stays inside a cluster. Costs power — see the design effect in randomisation units.
- Switchback experiments: alternate the entire market between arms over time slices — the standard for ride-hailing and delivery; the next lesson builds one.
- Budget-split / shadow auctions for ads: duplicate the auction with separated budgets rather than separated bidders.
- Ego-cluster designs for social features: treat a user and their neighbourhood, measure only the centre.
Common mistakes
Assuming your product has no interference. Search ranking (shared index updates), pricing (competitive response), notifications (attention is shared), two-sided markets — most valuable surfaces have some. The question is magnitude, not existence.
Detecting interference with an SRM-style check. Counts stay perfectly balanced under interference — the sample ratio check passes while the estimate is badly wrong. Detection needs designed contrasts: vary the treated fraction across markets and see whether per-user lift depends on it.
Cluster-randomising but analysing at user level. The clustering must enter the standard errors; user-level analysis of a city-level split understates uncertainty severely.
Try it yourself
Add a dose-response check: run the same market at 10%, 50% and 90% treated fractions and print the A/B lift at each. Under no interference the lift is flat in the fraction; here it will move — that movement is the standard production diagnostic.
What to learn next
- Switchback experiments — the time-slicing design built for shared markets.
- Randomisation units and assignment — cluster randomisation and its variance cost.
- Potential outcomes and counterfactuals — the formal language SUTVA lives in.
Researcher — Mathematics and papers.
SUTVA and exposure mappings
The Stable Unit Treatment Value Assumption (Rubin, 1980) requires $Y_i(\mathbf{T}) = Y_i(T_i)$ — unit $i$'s outcome depends only on its own assignment, not the full assignment vector $\mathbf{T} \in {0,1}^n$. Under interference there are up to $2^n$ potential outcomes per unit, and "the treatment effect" is undefined without further structure.
The standard repair is an exposure mapping (Aronow and Samii, 2017, Estimating average causal effects under general interference, Annals of Applied Statistics): a function $g(\mathbf{T}, i)$ compressing the assignment vector into a small exposure set — e.g. (own treatment, fraction of neighbours treated). Estimands are then contrasts between exposures, estimable when the design gives each exposure positive probability.
Estimands worth separating
- Total treatment effect (TTE / GATE): all-treated world versus all-control world — the launch decision quantity, which the naive A/B estimator does not target under interference.
- Direct effect: own treatment, holding others' exposure fixed.
- Spillover (indirect) effect: others' treatment, holding own fixed.
Hudgens and Halloran (2008), Toward causal inference with interference (JASA), formalise these under partial interference — spillovers confined within known groups — using two-stage randomisation: first the group's treated fraction, then individuals within groups.
Design and inference results
- Graph cluster randomisation: Ugander, Karrer, Backstrom and Kleinberg (2013), KDD — cluster the social graph, randomise clusters; bias falls with cluster quality while variance grows via the design effect.
- Ego-cluster designs: Saint-Jacques et al. (LinkedIn, 2019) — one-hop neighbourhoods treated as units for network-feature spillover measurement.
- Marketplace bias magnitude: Blake and Coey (2014), Why marketplace experimentation is harder than it seems (EC) — eBay evidence that cannibalisation inflates naive estimates; Fradkin (Airbnb) similarly via structural simulation.
- Dose-response detection: randomising the treated fraction across markets identifies the interference slope; a flat dose-response is evidence (not proof) of SUTVA holding.
Variance under cluster designs follows the Kish design effect $1 + (m-1)\rho$; graph-cluster designs trade this against exposure fidelity — smaller clusters mean better power and worse containment.
Reading
- Aronow and Samii (2017) — the general framework.
- Karrer et al. (2021), Network experimentation at scale (Facebook, KDD) — engineering and analysis of cluster experiments in production.
- Johari, Li, Liskovich and Weintraub (2022), Experimental design in two-sided platforms: an analysis of bias (Management Science) — when listing-side versus customer-side randomisation each fail.
What to learn next
- Switchback experiments — the time-slicing design built for shared markets.
- Randomisation units and assignment — cluster randomisation and its variance cost.
- Potential outcomes and counterfactuals — the formal language SUTVA lives in.