Experiment Design and A/B Testing

Guardrail metrics

Guardrail metrics are the dials you watch to make sure a "winning" change is not quietly breaking speed, trust or revenue somewhere you were not looking.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A guardrail metric is a measurement you are not trying to improve — you are promising not to damage it.

Driving to the station faster is the goal, but the driver still glances at the fuel gauge, the engine temperature and the mirrors. Nobody wins a race by ignoring an overheating engine. In an experiment, the metric you hope to move is the speedometer. The guardrails are every other dial that must stay healthy while you speed up.

Why it exists

Almost any metric can be improved by quietly damaging another. Autoplaying videos raises watch time and burns through mobile data. A pushier signup popup lifts registrations and teaches visitors to avoid the site. A heavier page adds a feature and slows every load.

Teams are rewarded for moving their own metric. Without guardrails, each team's win exports a small injury to someone else's metric, and the product decays while every dashboard is green.

How it works

experiment: new product page

goal metric:      purchases        → hoping it goes UP
guardrails:       page load time   → must not go up
                  crashes          → must not go up
                  unsubscribes     → must not go up
                  revenue/user     → must not go down

decision: ship only if the goal wins AND every guardrail holds

There is a subtle twist. For a guardrail, the question flips from "did it improve?" to "can we rule out meaningful damage?". Absence of a visible drop is not proof of safety — small experiments cannot see small damage. Good guardrail practice asks for evidence that harm, if any, is smaller than an agreed limit.

A real example you have seen

Amazon and Google both published versions of the same finding: an extra 100 milliseconds of page delay measurably cuts sales and searches. That is why load time is a near-universal guardrail. A feature that wins its own metric while adding 200 ms is, at these companies, dead on arrival.

Remember this

  • Guardrails are metrics you promise not to harm, alongside the one you hope to move.
  • The common ones: speed, crashes, revenue, retention, complaints.
  • For guardrails, ask "can we rule out damage?" — silence is not safety.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Verified with numpy 1.26.4 and scipy 1.14.1.

A win with a broken guardrail

A new product page converts better — and, being heavier, loads slower. Both facts are in the data, waiting to be checked.

guardrails.py
import numpy as np
from scipy import stats

rng = np.random.default_rng(21)
n = 20_000
buy_c = rng.random(n) < 0.100          # control conversion: 10.0%
buy_t = rng.random(n) < 0.112          # the new page converts better...
ms_c = rng.gamma(4, 90, n)             # ...but it is heavier, so it loads slower
ms_t = rng.gamma(4, 99, n)

for name, c, t in [("purchase rate", buy_c, buy_t),
                   ("load time, ms", ms_c, ms_t)]:
    p = stats.ttest_ind(t.astype(float), c.astype(float)).pvalue
    print(f"{name}: control {c.mean():7.3f}  treatment {t.mean():7.3f}  p={p:.4f}")
Output
purchase rate: control   0.097  treatment   0.111  p=0.0000
load time, ms: control 357.284  treatment 397.305  p=0.0000

The walkthrough

Both rows are significant — and they disagree about shipping. Conversion rose from 9.7% to 11.1%. Load time rose 40 ms. A dashboard showing only the first row says ship; the second row says the win was partly bought with speed, and speed debts compound across teams.

The decision is a policy, not a p-value. A sane policy, agreed before launch: ship if the goal metric wins and no guardrail shows damage beyond its margin — say, 20 ms for load time. Here the measured 40 ms breaches the margin, so the feature goes back for a diet, even though its own metric won.

Choose few, sensitive guardrails. Standard set: latency percentiles (p50/p95), crash and error rates, revenue per user, retention or return rate, unsubscribe and complaint rates, plus the sample ratio check, which is a guardrail on the experiment itself.

Watch the multiple-testing arithmetic. Ten guardrails, each read at 0.05, fire falsely about 40% of the time somewhere. Either apply a correction, or treat single marginal guardrail alarms as triggers for a closer look rather than automatic vetoes. What you must not do is scan ten dials and quote the worst one as if it were the only test run.

Common mistakes

Reading "no significant change" as "safe". An underpowered experiment cannot distinguish "no damage" from "damage too small for this sample". For guardrails that matter, state a margin and require the confidence interval to exclude damage beyond it — a non-inferiority test, not a significance test.

Adding twenty guardrails to look rigorous. Each extra dial adds false alarms and dilutes attention. Five well-chosen guardrails that always trigger investigation beat twenty that everyone learns to ignore.

Letting the goal team define its own guardrails. The team's incentive is a green dashboard. Guardrail definitions and margins belong to a central experimentation function, agreed before the experiment starts.

Averaging away tail damage. A mean load time can hold steady while the p95 doubles — the slowest users absorbing all the pain. Guardrail latency at percentiles, not means.

Try it yourself

Implement the non-inferiority version: with a 20 ms margin, compute the 95% confidence interval for the load-time difference (stats.ttest_ind(...).confidence_interval()) and check whether it lies entirely below +20 ms. Then shrink n to 2,000 and watch the interval widen until safety can no longer be demonstrated — the honest cost of small experiments.

What to learn next

Researcher — Mathematics and papers.

Non-inferiority formulation

For a guardrail where increase means harm, fix margin $\delta > 0$ and test

$$ H_0: \mu_t - \mu_c \ge \delta \quad \text{vs} \quad H_1: \mu_t - \mu_c < \delta $$

  • $\mu_t, \mu_c$ — treatment and control means of the guardrail metric.
  • $\delta$ — the largest harm considered tolerable, fixed in advance.

Practice: conclude safety iff the upper limit of the $(1 - 2\alpha)$ CI for $\mu_t - \mu_c$ falls below $\delta$. This inverts the burden of proof — the default is "harmful until shown otherwise", the correct default for trust and performance metrics. Power against $\mu_t = \mu_c$ requires $n \propto 2\sigma^2 (z_{1-\alpha} + z_{1-\beta})^2 / \delta^2$ per arm: tight margins on noisy metrics are expensive, which is why guardrail margins must be negotiated, not wished.

The OEC and metric hierarchies

Kohavi's overall evaluation criterion (OEC) frames the ideal: a single scalar combining goal and countervailing metrics with explicit weights, optimised jointly — degrading a component becomes visible in the number being maximised. Reality usually keeps them separate:

  • Goal metrics — the hypothesis, tested for superiority.
  • Guardrails — tested for non-inferiority against margins.
  • Data-quality checks — SRM, logging health; failure voids the run rather than informing it.

Deng and Shi (2016), Data-driven metric development for online controlled experiments (KDD), cover sensitivity/directionality trade-offs: guardrails should be sensitive (move when harmed) and directional (their movement has an agreed reading). Dmitriev et al. (2017), A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls (KDD), catalogue guardrail misreads from Microsoft practice.

Multiplicity across guardrails

With $G$ guardrails at level $\alpha$, family-wise false alarm probability approaches $1 - (1-\alpha)^G$ under independence. Options:

  • Bonferroni / Holm on the guardrail family — conservative, simple, appropriate for veto-power dials.
  • False discovery rate (Benjamini–Hochberg) — appropriate when guardrail alarms trigger investigation rather than automatic vetoes.
  • Hierarchical gatekeeping — test guardrails only when the goal metric wins, preserving error rates through the fixed sequence.

The latency-revenue exchange rate deserves its own estimate: dedicated slowdown experiments (deliberately injecting delay, as in Google/Bing's 2009-2012 studies reported by Schurman and Brutlag, and Kohavi et al. 2013 KDD) price the guardrail margin in revenue terms, converting "20 ms" into money both sides of a shipping debate understand.

What to learn next