Hypothesis Testing and Inference

Multiple testing correction

Run twenty tests on pure noise and one will look significant — corrections like Bonferroni and Benjamini-Hochberg keep mass testing honest.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Multiple testing correction stops you fooling yourself when you run many tests at once.

Buy one lottery ticket and winning would astonish you. Buy ten thousand and a win astonishes nobody — somebody's number had to come up. The win means luck, not destiny.

Every statistical test you run is a lottery ticket for a false alarm. Run enough tests and a "significant" result is guaranteed, even when nothing anywhere is real.

Why it exists

Modern data work runs tests in bulk. Compare twenty metrics between two app versions. Screen ten thousand genes for a disease link. Try fifty model variants against a baseline.

At the usual significance threshold, each innocent test has a one-in-twenty chance of crying wolf. Twenty innocent tests give you about a 64% chance of at least one false discovery. Someone always "finds something" — and it is noise wearing a medal.

A famous demonstration made this vivid: researchers put a dead salmon in a brain scanner, ran thousands of uncorrected tests, and found "brain activity". The correction, not the salmon, was the missing ingredient.

How it works

Two families of fixes, with different personalities:

Bonferroni:  make each test stricter.
             20 tests → each must be 20x more convincing.
             almost no false alarms — but real effects get missed.

Benjamini-Hochberg (BH): control the false discovery rate.
             accept that ~5% of your DISCOVERIES may be false,
             in exchange for catching far more real ones.

Bonferroni suits verdicts you must not get wrong — a medicine's headline claim. BH suits screening — finding candidate genes or features worth a closer look, where a few duds among the discoveries are a fair price.

A real example you have seen

"Scientists find chocolate linked to happiness" headlines often come from studies testing dozens of foods against dozens of moods. Hundreds of tests, one shiny survivor, no correction — then a press release. You now own the tool to ask the right question: how many tests did they run?

Remember this

  • Many tests = many lottery tickets for false alarms.
  • Bonferroni: strict, protects against any false alarm, misses real effects.
  • Benjamini–Hochberg: tolerates a small share of false discoveries, finds more truth.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Outputs verified with numpy 1.26 and scipy 1.14, CPU. The p-values are simulated; other numpy versions may shift the exact digits, not the lesson.

Twenty tests on pure noise

Twenty A/B tests where nothing differs in any of them — both groups drawn from the same distribution.

corrections.py
import numpy as np
from scipy import stats

rng = np.random.default_rng(3)
pvals = np.array([
    stats.ttest_ind(rng.normal(0, 1, 50), rng.normal(0, 1, 50)).pvalue
    for _ in range(20)
])
print(f"raw p-values below 0.05:    {(pvals < 0.05).sum()} of 20  (every one a false alarm)")
print(f"smallest raw p-value:       {pvals.min():.4f}")

print(f"after Bonferroni (0.05/20): {(pvals < 0.05 / 20).sum()} of 20")

adjusted = stats.false_discovery_control(pvals)   # Benjamini-Hochberg
print(f"after BH adjustment:        {(adjusted < 0.05).sum()} of 20")
Output
raw p-values below 0.05:    1 of 20  (every one a false alarm)
smallest raw p-value:       0.0354
after Bonferroni (0.05/20): 0 of 20
after BH adjustment:        0 of 20

The walkthrough

One of twenty innocent tests fired anyway. p = 0.0354 looks publishable. It is noise — we built this world, so we know. On average one such imposter appears per twenty null tests; some runs produce two or three.

Bonferroni divides the threshold by the test count. Each test now needs p < 0.0025. Brutal but transparent, and it guards the family-wise error rate: the chance of even one false alarm across the whole batch stays at 5%.

false_discovery_control implements Benjamini–Hochberg. It returns adjusted p-values you compare to 0.05 as usual. The procedure ranks the p-values and asks whether the small ones cluster more than uniform noise would allow. Here they do not — correct verdict, zero discoveries.

When real effects exist, the two fixes part ways. Mix 5 real effects into 100 tests: Bonferroni typically keeps two or three; BH keeps most of the five while letting the occasional imposter through. That trade is the entire choice.

Where hidden multiplicity lives in ML: trying many random seeds and reporting the best, tuning hyperparameters on the test set, comparing every pair of models in a leaderboard. Each is multiple testing without the name — see validation strategies for the honest workflow.

Common mistakes

Correcting only the tests you reported. If you ran 200 and show 3, the correction needs 200. The family is everything you ran (and, for strict honesty, everything you would have run).

Bonferroni across thousands of screens. At 10,000 genes, each test needs p < 0.000005 — real biology gets thrown out wholesale. Screening is FDR territory; use BH.

Treating a BH survivor as a proven fact. BH promises that about 95% of survivors are real — collectively. Any individual survivor still deserves replication.

Running tests until something clears the bar, then stopping. No correction rescues a stopping rule chosen by the results. Pre-register the family of tests, or use sequential methods designed for peeking.

Try it yourself

Add five real effects: draw those groups from rng.normal(0.5, 1, 50) versus rng.normal(0, 1, 50). Rerun with 100 total tests and compare how many of the five each method recovers.

What to learn next

Researcher — Mathematics and papers.

Error rates, precisely

For $m$ tests with $V$ false rejections and $R$ total rejections:

  • FWER $= \Pr(V \geq 1)$ — family-wise error rate.
  • FDR $= \mathbb{E}!\left[\frac{V}{\max(R, 1)}\right]$ — expected fraction of rejections that are false (Benjamini and Hochberg, 1995, Controlling the false discovery rate).

FWER control implies FDR control; the converse fails. The regimes differ by what a single error costs.

Bonferroni and its refinements

Testing each hypothesis at $\alpha/m$ gives $\text{FWER} \leq \sum_i \Pr(p_i \leq \alpha/m \mid H_0) = \alpha$ by the union bound — no independence needed. Šidák's $1 - (1-\alpha)^{1/m}$ is exact under independence, marginally less conservative.

Holm's step-down procedure (1979) uniformly improves Bonferroni at no assumption cost: order $p_{(1)} \leq \dots \leq p_{(m)}$; reject while $p_{(k)} \leq \frac{\alpha}{m - k + 1}$; stop at the first failure. There is no reason to use plain Bonferroni when Holm is available; plain Bonferroni survives because one sentence states it.

Benjamini–Hochberg

Find the largest $k$ with

$$ p_{(k)} \leq \frac{k}{m}\,\alpha $$

and reject hypotheses $1, \dots, k$ (after ordering). Under independence, $\text{FDR} \leq \frac{m_0}{m}\alpha \leq \alpha$, where $m_0$ is the number of true nulls. The geometric picture: plot ordered p-values against rank; reject everything below the line through the origin with slope $\alpha/m$.

  • Dependence: BH remains valid under positive regression dependence (PRDS — Benjamini and Yekutieli, 2001); the same paper's BY procedure (inflate by $\sum_{i=1}^m 1/i \approx \ln m$) covers arbitrary dependence, at real power cost.
  • Adaptive versions estimate $m_0$ (Storey's q-value, 2002) and recycle the slack, tightening control to $\approx \alpha$ exactly.
  • Adjusted p-values: $\tilde{p}{(k)} = \min{j \geq k} \frac{m\, p_{(j)}}{j}$ — what scipy.stats.false_discovery_control returns.

The silent multiplicity problem

Corrections assume the family is declared. The deeper failure mode is the garden of forking paths (Gelman and Loken, 2014): analysis choices made after seeing data create multiplicity with $m$ effectively unknown and uncounted. Formal countermeasures: pre-registration; sample splitting; selective inference, which computes p-values conditional on the selection event (Taylor and Tibshirani, 2015, Statistical learning and selective inference); and knockoff filters for FDR-controlled feature selection in regression (Barber and Candès, 2015).

Complexity and scale

All procedures here are $O(m \log m)$ — sorting dominates. Genomics routinely runs $m \sim 10^6$ (GWAS uses a Bonferroni-derived genome-wide threshold of $5 \times 10^{-8}$); neuroimaging corrects over correlated voxel maps with random-field theory or permutation-based maxT (Westfall and Young, 1993), which respects arbitrary dependence exactly. The dead-salmon study is Bennett et al. (2009), Neural correlates of interspecies perspective taking in the post-mortem Atlantic salmon — an Ig Nobel-winning argument for correction.

What to learn next