Experiment Design and A/B Testing
Novelty and primacy effects
New features get curiosity clicks that fade, and changed habits drag old users down before they adapt — both effects make short experiments lie about the long run.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A novelty effect is a burst of attention a change gets only because it is new. A primacy effect is the opposite: users stumble at first because their habits broke.
A new restaurant opens on your street and the first week it is packed. Half the crowd came out of curiosity, not hunger. By week six, you can see whether the food actually holds customers. Now picture the opposite: your neighbourhood shop rearranges its shelves. For two weeks, regulars grumble and take longer to find things — then they adapt, and maybe the new layout is genuinely better.
Experiments suffer both. Measure during the curiosity burst and you overrate the change. Measure during the stumbling weeks and you underrate it.
Why it exists
An experiment is a sample of time, and early time is not typical time. A shiny new button gets clicked because it is shiny and new. A rearranged menu gets fumbled because fingers memorised the old one.
Both distortions wear off. The danger is that experiments are expensive, teams are impatient, and week-one results arrive first — looking decisive in either direction.
How it works
novelty: effect starts high, decays toward the truth
+9 +5 +4 +3 +1 +1 +1 ... real effect: +1
primacy: effect starts low, climbs toward the truth
-4 -2 -1 0 +1 +2 +2 ... real effect: +2
week-one reading: wrong in both cases, in opposite directionsThe defences are patience and comparison. Run long enough to see the curve flatten. Compare new users (no habits to break, no novelty to chase — the feature is their normal) against long-time users. If the two groups disagree strongly, time effects are in play.
A real example you have seen
Every app redesign you have hated follows the primacy script. When Instagram or WhatsApp moves a button, the internet complains for two weeks — then adapts. If those companies trusted week-one metrics alone, no redesign would ever ship. Facebook and Microsoft both publish experiment guidelines requiring at least two full weeks partly for this reason.
Remember this
- Novelty: curiosity inflates early results, then fades.
- Primacy: broken habits deflate early results, then recover.
- Trust the trend across weeks and the new-user segment, not the day-one splash.
What to learn next
- Interference and network effects — the next way an honest-looking experiment misleads.
- Peeking and sequential testing — the stopping discipline that novelty makes doubly necessary.
- Guardrail metrics — watching more than the metric you hope will move.
Developer — Code and libraries.
Setup
pip install numpyVerified with numpy 1.26.4.
Watching a novelty effect decay
The simulation gives treatment a lift that decays over two weeks — the curiosity wearing off — plus daily measurement noise.
import numpy as np
rng = np.random.default_rng(11)
days = np.arange(1, 15)
true_lift = 10 * np.exp(-days / 4) # curiosity clicks that fade away
measured = []
for lift in true_lift:
control = rng.normal(100, 0.8) # daily clicks per 1000 users
treatment = rng.normal(100 + lift, 0.8)
measured.append(treatment - control)
measured = np.array(measured)
for d, m in zip(days, measured):
bar = "#" * max(int(m * 3), 0)
print(f"day {d:>2}: lift {m:+6.2f} {bar}")
print(f"week 1 average: {measured[:7].mean():+.2f}")
print(f"week 2 average: {measured[7:].mean():+.2f}")day 1: lift +8.85 ########################## day 2: lift +4.68 ############## day 3: lift +4.54 ############# day 4: lift +3.18 ######### day 5: lift +0.79 ## day 6: lift +0.90 ## day 7: lift +1.08 ### day 8: lift +2.03 ###### day 9: lift +0.23 day 10: lift +1.49 #### day 11: lift +0.12 day 12: lift -0.35 day 13: lift +1.27 ### day 14: lift -0.28 week 1 average: +3.43 week 2 average: +0.64
The walkthrough
A week-one stop reports +3.43; the settled effect is under +1. A team stopping on day three would celebrate a lift five times larger than reality delivers. The yearly numbers would then mysteriously refuse to grow — a pattern experimentation teams know painfully well.
The daily printout is the diagnostic. A single end-of-experiment average hides the decay; the day-by-day series shows it. Any decent experiment dashboard plots the treatment effect over time exactly for this reason.
Noise complicates the reading. Days 9 through 14 wobble around a small positive value. Distinguishing "still decaying" from "settled plus noise" needs either more days or a fitted decay curve — the researcher block sketches the model.
Segmenting is the second diagnostic. Users who joined after the experiment started never saw the old version. For them there is no novelty and no broken habit. If their lift is +0.6 while veterans show +3.4, time effects are confirmed, and the new-user number is closer to the long-run truth.
Common mistakes
Stopping the moment significance appears. Novelty inflates early effects, and peeking rewards exactly those inflated days. The two failures compound each other.
Running exactly seven days. One week captures each weekday once but gives no second week to compare against. Two weeks minimum lets you split the data and check stability — and catches weekday-weekend cycles twice.
Treating all fading effects as novelty. Effects also fade when a marketing push ends, when a bug fix lands mid-experiment, or when heavy users saturate a promotion. Read the timeline alongside the deploy log before naming the cause.
Believing week two is always the truth. Some effects need months to surface — subscription renewals, habit formation. For those, a two-week experiment measures a leading indicator at best. Long-term holdouts (below) are the honest tool.
Try it yourself
Flip the simulation into a primacy story: make true_lift start at -6 and climb to +2 with the same exponential shape. Confirm week one would kill a feature that week two vindicates. Then add a fitted curve: np.polyfit on log-transformed lifts recovers the decay rate.
What to learn next
- Interference and network effects — the next way an honest-looking experiment misleads.
- Peeking and sequential testing — the stopping discipline that novelty makes doubly necessary.
- Guardrail metrics — watching more than the metric you hope will move.
Researcher — Mathematics and papers.
A time-varying treatment effect model
Let the treatment effect at experiment day $t$ be
$$ \tau(t) = \tau_\infty + (\tau_0 - \tau_\infty)\, e^{-t/\lambda} $$
- $\tau_0$ — the initial effect (inflated by novelty, or deflated by primacy).
- $\tau_\infty$ — the asymptotic (long-run) effect: the estimand that matters for shipping.
- $\lambda$ — the adaptation time constant, in days.
Fitting this on daily effect estimates gives both a settled-effect estimate and a data-driven answer to "have we run long enough" ($t \gg \lambda$). The naive experiment-average estimates $\frac{1}{T}\int_0^T \tau(t)\,dt$, which overweights the transient by construction.
Diagnostics with identification power
- Cohort-by-exposure analysis: stratify the effect by user tenure at first exposure. Users first exposed in week 2 replicate the week-1 curve of week-1 users if the dynamics are exposure-driven (novelty/primacy), but not if they are calendar-driven (seasonality, marketing) — a clean separation test.
- New-user segment as a proxy for $\tau_\infty$: for new users, treatment is the baseline. Their effect estimate skips the transient entirely, at the price of a population shift (new users are not veterans).
- Dosage curves: plot effect against the number of exposures per user. Novelty predicts declining effect in exposure count within users, not only in calendar time.
Long-term measurement designs
- Long-term holdouts: keep a small control group (1–5%) unexposed for months after launch. This is the gold standard for $\tau_\infty$; the cost is product inconsistency and maintenance burden.
- Surrogate indices: Athey, Chetty, Imbens and Kang (2019), The Surrogate Index (NBER w26463) — combine many short-term metrics into a statistically validated proxy for the long-term outcome, trained on historical experiments with long follow-ups.
- Hohnhold, O'Brien and Tang (2015), Focusing on the Long-term: It's Good for Users and Business (KDD) — Google's ads work distinguishing user learning effects from instantaneous effects, measuring ads blindness developing over weeks and its reversal.
Reading
- Kohavi, Tang and Xu (2020), Trustworthy Online Controlled Experiments, chapter 23 — measuring long-term treatment effects, and how long to run.
- Sadeghi et al. (2021), Novelty and Primacy: A Long-Term Estimator for Online Experiments (arXiv:2102.12893) — formal estimators for the dynamics above.
What to learn next
- Interference and network effects — the next way an honest-looking experiment misleads.
- Peeking and sequential testing — the stopping discipline that novelty makes doubly necessary.
- Guardrail metrics — watching more than the metric you hope will move.