Experiment Design and A/B Testing
Switchback experiments
A switchback experiment flips the whole system between old and new versions in randomised time blocks — the standard design when users share one market that cannot be split.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A switchback experiment gives the new version to everyone — but only during randomly chosen time blocks — and compares those blocks against the old-version blocks.
A family kitchen cannot cook two dinners at once to compare recipes. So you cook the new recipe on some nights and the old one on others, and compare across nights. One kitchen, one family, alternating in time. That is a switchback: the whole system flips between versions, block by block.
The order of nights must come from a coin, not convenience. Cook the new recipe only on relaxed Sundays and you have measured Sundays, not the recipe.
Why it exists
The previous lesson showed how shared markets poison user-splitting. Treated users grab rides, stock or riders out of control users' hands, and the comparison measures the grabbing.
The switchback removes the sharing problem by never splitting the market. During a treatment block, everyone gets the new dispatch algorithm — there is no control user to rob. The price: you now compare across time, and time is moody. Lunch rush is not 3 pm. The randomised block order is what keeps moody time from taking sides.
How it works
hour: 9 10 11 12 13 14 15 16 17
version: B A A B A B B A B ← coin per hour
compare: average of B-hours vs average of A-hours
(the coin scattered rush hours fairly across both)One wrinkle needs care: the system remembers. If the new algorithm sends drivers to the wrong side of town, the next block inherits that mess. Blocks must be long enough for the system to settle — or the analysis must expect the leakage between neighbouring blocks.
A real example you have seen
The surge price you see on a ride app at 6 pm was almost certainly tuned by switchbacks. DoorDash, Uber, Lyft and Swiggy-style platforms all run them for dispatch, pricing and matching changes. One city's drivers are a single shared pool, and no user-split can divide that pool fairly.
Remember this
- A switchback flips the whole system between versions in randomised time blocks.
- It exists because shared markets make user-splitting lie.
- Blocks must be long enough to settle, and their order must come from a coin.
What to learn next
- Guardrail metrics — what to watch while the market flips.
- Interference and network effects — the disease this design treats.
- Difference-in-differences — the observational cousin of comparing across time.
Developer — Code and libraries.
Setup
pip install numpyVerified with numpy 1.26.4.
A week of switchbacks, and the design it beats
The simulation models a dispatch change that truly saves 3 minutes of wait time, on top of a strong daily demand cycle. Compare a lazy day/night split against a randomised switchback.
import numpy as np
rng = np.random.default_rng(9)
hours = 168 # one week of one-hour blocks
rush = 20 + 10 * np.sin(np.arange(hours) / 24 * 2 * np.pi) # daily demand cycle
# bad idea: new dispatcher by day, old one by night
day = np.arange(hours) % 24 < 12
wait_day_split = rush - 3 * day + rng.normal(0, 2, hours) # true effect: -3 min
biased = wait_day_split[day].mean() - wait_day_split[~day].mean()
# switchback: a coin toss decides each hour block independently
flip = np.zeros(hours, bool)
flip[rng.permutation(hours)[: hours // 2]] = True
wait_switch = rush - 3 * flip + rng.normal(0, 2, hours)
honest = wait_switch[flip].mean() - wait_switch[~flip].mean()
print(f"day-vs-night split: {biased:+.2f} minutes")
print(f"switchback estimate: {honest:+.2f} minutes (truth: -3.00)")day-vs-night split: +9.07 minutes switchback estimate: -2.40 minutes (truth: -3.00)
The walkthrough
The day/night split gets the sign wrong. The new dispatcher saves 3 minutes, yet the split reports waits 9 minutes longer — because daytime demand is higher, and demand got bundled into the comparison. Any schedule chosen for convenience bundles time's moods into the answer.
The switchback lands at -2.40 against a truth of -3.00. The gap is honest sampling noise from 168 blocks; more weeks tighten it. The coin scattered rush hours evenly, so the demand cycle cancels.
The unit of analysis is the block, not the ride. 168 blocks means 168 observations. Analysing millions of individual rides as independent would shrink error bars dishonestly — rides inside one block share conditions. This mirrors the cluster-versus-user issue in randomisation units.
Carryover is the design's standing weakness. This simulation has none: each hour's wait depends only on that hour's version. Real dispatch systems drag state — driver positions, queued orders — across the boundary, biasing estimates toward zero. Longer blocks dilute carryover; discarding the first minutes of each block ("burn-in") is the common patch.
Common mistakes
Alternating A-B-A-B on a fixed schedule. A deterministic alternation lines up with any 2-hour periodicity in the world, and confounds with everything on that rhythm. Randomise the block assignments.
Blocks shorter than the system's memory. If rebalancing drivers takes 40 minutes, 30-minute blocks measure mostly transition states. Pick block length from the system's settling time, not from impatience for sample size.
Counting per-ride observations in the standard error. Cluster by block (or use block means, as above). The honest error bar is wider; it is also the true one.
Forgetting user-facing consistency. Switchbacks show every user both versions. For invisible backend changes (dispatch, pricing logic, ranking weights) this is fine. For visible UI, flipping hourly would be absurd — switchbacks are a backend tool.
Try it yourself
Add carryover: make each hour's wait depend 30% on the previous hour's version. Watch the estimate shrink toward zero. Then discard the first "20 minutes" of each block (weight boundary hours by 2/3) and see how much of the truth returns. Finally, try 4-hour blocks — 42 of them — and compare the error bar against the hourly design.
What to learn next
- Guardrail metrics — what to watch while the market flips.
- Interference and network effects — the disease this design treats.
- Difference-in-differences — the observational cousin of comparing across time.
Researcher — Mathematics and papers.
The estimator and its variance
Partition time into $K$ blocks; assign $W_k \in {0,1}$ i.i.d. Bernoulli(½). With block-mean outcomes $\bar{Y}_k$, the difference-in-means estimator is
$$ \hat{\tau} = \frac{\sum_k W_k \bar{Y}_k}{\sum_k W_k} - \frac{\sum_k (1 - W_k) \bar{Y}_k}{\sum_k (1 - W_k)} $$
- $W_k$ — the arm serving during block $k$.
- $\bar{Y}_k$ — the metric averaged within block $k$.
Serial correlation across blocks (demand cycles, weather) inflates the variance of $\hat{\tau}$ relative to i.i.d. sampling; inference should use block-level randomisation tests or HAC/cluster-robust variance at the block level. Within-block observation counts add nothing beyond reducing within-block measurement noise.
Carryover bias
Let the outcome in block $k$ respond to current and lagged assignment: $\mathbb{E}[\bar{Y}_k] = \mu_k + \tau W_k + \gamma W_{k-1}$. Under i.i.d. Bernoulli assignment $W_{k-1}$ is independent of $W_k$, so the lagged term contributes $\gamma/2$ to both arms and cancels: $\hat{\tau} \to \tau$. That looks like good news and is not. The quantity a launch decision needs is the all-treated versus all-control contrast, which is $\tau + \gamma$, so the naive estimator understates the shipping effect by exactly the carryover — silently, with nothing in the block-level data to flag it. Designs that alternate deterministically are worse: there $W_{k-1} = 1 - W_k$, and the estimator converges to $\tau - \gamma$.
Bojinov, Simchi-Levi and Zhao (2023), Design and analysis of switchback experiments (Management Science), formalise this with $m$-block carryover: the optimal design lengthens blocks until carryover is negligible relative to block length, and they derive minimax-optimal switch frequencies plus randomisation-based inference that remains valid under misspecified memory. Burn-in deletion is the practical approximation: drop the first $b$ periods of each block, trading data for bias.
Design refinements
- Blocked / paired switchbacks: pair adjacent same-length windows (Tue 2–4 pm this week vs next week), randomising within pairs — removes low-frequency drift; analysis by paired differences.
- Crossover designs with washout are the clinical-trial ancestor (Jones and Kenward, Design and Analysis of Cross-Over Trials, 3rd ed., 2014); "washout" is the medical name for burn-in.
- Regional switchbacks: randomise (region × time-block) cells; regions add replication and allow spatial spillover checks — the DoorDash production design (Kastelman and Ramesh, 2018, DoorDash engineering, Switchback Tests and Randomized Experimentation Under Network Effects).
- Synthetic-control hybrids: for very few units (one city), pair switchbacks with time-series counterfactual models; the frontier for low-replication settings connects to difference-in-differences.
Reading
- Bojinov, Simchi-Levi and Zhao (2023) — the definitive design/analysis treatment.
- Hu and Wager (2022), Switchback experiments under geometric mixing (arXiv:2209.00197) — bias-variance trade-offs when the system state mixes at a known rate.
- Glynn, Johari and Rasouli (2020), Adaptive experimental design with temporal interference (NeurIPS) — Markov-chain framing of switchbacks.
What to learn next
- Guardrail metrics — what to watch while the market flips.
- Interference and network effects — the disease this design treats.
- Difference-in-differences — the observational cousin of comparing across time.