Experiment Design and A/B Testing

Novelty and primacy effects

New features get curiosity clicks that fade, and changed habits drag old users down before they adapt — both effects make short experiments lie about the long run.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A novelty effect is a burst of attention a change gets only because it is new. A primacy effect is the opposite: users stumble at first because their habits broke.

A new restaurant opens on your street and the first week it is packed. Half the crowd came out of curiosity, not hunger. By week six, you can see whether the food actually holds customers. Now picture the opposite: your neighbourhood shop rearranges its shelves. For two weeks, regulars grumble and take longer to find things — then they adapt, and maybe the new layout is genuinely better.

Experiments suffer both. Measure during the curiosity burst and you overrate the change. Measure during the stumbling weeks and you underrate it.

Why it exists

An experiment is a sample of time, and early time is not typical time. A shiny new button gets clicked because it is shiny and new. A rearranged menu gets fumbled because fingers memorised the old one.

Both distortions wear off. The danger is that experiments are expensive, teams are impatient, and week-one results arrive first — looking decisive in either direction.

How it works

novelty:   effect starts high, decays toward the truth
           +9  +5  +4  +3  +1  +1  +1  ...   real effect: +1

primacy:   effect starts low, climbs toward the truth
           -4  -2  -1   0  +1  +2  +2  ...   real effect: +2

week-one reading: wrong in both cases, in opposite directions

The defences are patience and comparison. Run long enough to see the curve flatten. Compare new users (no habits to break, no novelty to chase — the feature is their normal) against long-time users. If the two groups disagree strongly, time effects are in play.

A real example you have seen

Every app redesign you have hated follows the primacy script. When Instagram or WhatsApp moves a button, the internet complains for two weeks — then adapts. If those companies trusted week-one metrics alone, no redesign would ever ship. Facebook and Microsoft both publish experiment guidelines requiring at least two full weeks partly for this reason.

Remember this

  • Novelty: curiosity inflates early results, then fades.
  • Primacy: broken habits deflate early results, then recover.
  • Trust the trend across weeks and the new-user segment, not the day-one splash.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Verified with numpy 1.26.4.

Watching a novelty effect decay

The simulation gives treatment a lift that decays over two weeks — the curiosity wearing off — plus daily measurement noise.

novelty.py
import numpy as np

rng = np.random.default_rng(11)
days = np.arange(1, 15)
true_lift = 10 * np.exp(-days / 4)        # curiosity clicks that fade away

measured = []
for lift in true_lift:
    control = rng.normal(100, 0.8)        # daily clicks per 1000 users
    treatment = rng.normal(100 + lift, 0.8)
    measured.append(treatment - control)
measured = np.array(measured)

for d, m in zip(days, measured):
    bar = "#" * max(int(m * 3), 0)
    print(f"day {d:>2}: lift {m:+6.2f}  {bar}")
print(f"week 1 average: {measured[:7].mean():+.2f}")
print(f"week 2 average: {measured[7:].mean():+.2f}")
Output
day  1: lift  +8.85  ##########################
day  2: lift  +4.68  ##############
day  3: lift  +4.54  #############
day  4: lift  +3.18  #########
day  5: lift  +0.79  ##
day  6: lift  +0.90  ##
day  7: lift  +1.08  ###
day  8: lift  +2.03  ######
day  9: lift  +0.23  
day 10: lift  +1.49  ####
day 11: lift  +0.12  
day 12: lift  -0.35  
day 13: lift  +1.27  ###
day 14: lift  -0.28  
week 1 average: +3.43
week 2 average: +0.64

The walkthrough

A week-one stop reports +3.43; the settled effect is under +1. A team stopping on day three would celebrate a lift five times larger than reality delivers. The yearly numbers would then mysteriously refuse to grow — a pattern experimentation teams know painfully well.

The daily printout is the diagnostic. A single end-of-experiment average hides the decay; the day-by-day series shows it. Any decent experiment dashboard plots the treatment effect over time exactly for this reason.

Noise complicates the reading. Days 9 through 14 wobble around a small positive value. Distinguishing "still decaying" from "settled plus noise" needs either more days or a fitted decay curve — the researcher block sketches the model.

Segmenting is the second diagnostic. Users who joined after the experiment started never saw the old version. For them there is no novelty and no broken habit. If their lift is +0.6 while veterans show +3.4, time effects are confirmed, and the new-user number is closer to the long-run truth.

Common mistakes

Stopping the moment significance appears. Novelty inflates early effects, and peeking rewards exactly those inflated days. The two failures compound each other.

Running exactly seven days. One week captures each weekday once but gives no second week to compare against. Two weeks minimum lets you split the data and check stability — and catches weekday-weekend cycles twice.

Treating all fading effects as novelty. Effects also fade when a marketing push ends, when a bug fix lands mid-experiment, or when heavy users saturate a promotion. Read the timeline alongside the deploy log before naming the cause.

Believing week two is always the truth. Some effects need months to surface — subscription renewals, habit formation. For those, a two-week experiment measures a leading indicator at best. Long-term holdouts (below) are the honest tool.

Try it yourself

Flip the simulation into a primacy story: make true_lift start at -6 and climb to +2 with the same exponential shape. Confirm week one would kill a feature that week two vindicates. Then add a fitted curve: np.polyfit on log-transformed lifts recovers the decay rate.

What to learn next

Researcher — Mathematics and papers.

A time-varying treatment effect model

Let the treatment effect at experiment day $t$ be

$$ \tau(t) = \tau_\infty + (\tau_0 - \tau_\infty)\, e^{-t/\lambda} $$

  • $\tau_0$ — the initial effect (inflated by novelty, or deflated by primacy).
  • $\tau_\infty$ — the asymptotic (long-run) effect: the estimand that matters for shipping.
  • $\lambda$ — the adaptation time constant, in days.

Fitting this on daily effect estimates gives both a settled-effect estimate and a data-driven answer to "have we run long enough" ($t \gg \lambda$). The naive experiment-average estimates $\frac{1}{T}\int_0^T \tau(t)\,dt$, which overweights the transient by construction.

Diagnostics with identification power

  • Cohort-by-exposure analysis: stratify the effect by user tenure at first exposure. Users first exposed in week 2 replicate the week-1 curve of week-1 users if the dynamics are exposure-driven (novelty/primacy), but not if they are calendar-driven (seasonality, marketing) — a clean separation test.
  • New-user segment as a proxy for $\tau_\infty$: for new users, treatment is the baseline. Their effect estimate skips the transient entirely, at the price of a population shift (new users are not veterans).
  • Dosage curves: plot effect against the number of exposures per user. Novelty predicts declining effect in exposure count within users, not only in calendar time.

Long-term measurement designs

  • Long-term holdouts: keep a small control group (1–5%) unexposed for months after launch. This is the gold standard for $\tau_\infty$; the cost is product inconsistency and maintenance burden.
  • Surrogate indices: Athey, Chetty, Imbens and Kang (2019), The Surrogate Index (NBER w26463) — combine many short-term metrics into a statistically validated proxy for the long-term outcome, trained on historical experiments with long follow-ups.
  • Hohnhold, O'Brien and Tang (2015), Focusing on the Long-term: It's Good for Users and Business (KDD) — Google's ads work distinguishing user learning effects from instantaneous effects, measuring ads blindness developing over weeks and its reversal.

Reading

  • Kohavi, Tang and Xu (2020), Trustworthy Online Controlled Experiments, chapter 23 — measuring long-term treatment effects, and how long to run.
  • Sadeghi et al. (2021), Novelty and Primacy: A Long-Term Estimator for Online Experiments (arXiv:2102.12893) — formal estimators for the dynamics above.

What to learn next