Resampling, Likelihood and Bayes
The bootstrap
The bootstrap measures the uncertainty of any statistic — median, correlation, anything — by resampling your own data instead of trusting a formula.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The bootstrap measures how much your estimate would wobble, by resampling your own data over and over.
The honest way to know your estimate's wobble is to repeat the whole survey many times. Survey 15 customers, get an answer; survey 15 fresh ones, get a slightly different answer. The spread across surveys is the uncertainty.
Nobody can afford a thousand surveys. The bootstrap's trick: treat the one sample you have as a stand-in for the whole population, and draw your thousand surveys from it.
Why it exists
For the mean, statisticians have a textbook wobble formula. But for the median? The 90th percentile? A correlation? A model's F1 score? The formulas get ugly, exist only under assumptions, or do not exist at all.
Bradley Efron's 1979 insight: the computer can replace the formula. Resampling costs electricity, and electricity got cheap. One idea now covers every statistic you can compute.
The name comes from "pulling yourself up by your own bootstraps" — the data lifts itself, no outside formula needed.
How it works
Draw a fake sample from your real sample, the same size, allowing repeats. Compute your statistic on it. Do this thousands of times.
real sample (15 values)
│ draw 15 WITH REPEATS → fake sample 1 → median = 102
│ draw 15 WITH REPEATS → fake sample 2 → median = 99
│ ... 10,000 times ...
▼
the 10,000 medians pile up → their spread = your uncertaintyAllowing repeats is the entire trick. Each fake sample leaves some values out and doubles others — mimicking how a fresh survey would catch some customers and miss others. Without repeats you would copy the sample exactly, and learn nothing.
A real example you have seen
Election forecasters and sports analytics sites publish ranges like "projected: 48 to 52 seats". Under the hood, many of them resample polling data thousands of times and read off the spread. Machine-learning papers that report a score as "0.87, give or take 0.02" have often measured that give-or-take by bootstrapping the test set.
Remember this
- Resample your own data with repeats, recompute, watch the pile of answers.
- The pile's spread estimates your statistic's wobble — no formula needed.
- Works for medians, percentiles, correlations, model scores — anything computable.
What to learn next
- Permutation tests — resampling's sibling for yes/no hypothesis questions.
- Confidence intervals — the classical formulas the bootstrap generalises.
- Monte Carlo simulation — the simulate-and-count principle underneath it all.
Developer — Code and libraries.
Setup
pip install numpy scipyOutputs verified with numpy 1.26 and scipy 1.14, CPU. Bootstrap digits reproduce with these versions; other versions may shift them by a millisecond or two.
A confidence interval for a median
Response times for 15 requests, including one timeout horror. The median is the sane summary — and the median has no friendly textbook formula. Perfect bootstrap territory.
import numpy as np
rng = np.random.default_rng(42)
data = np.array([102, 98, 115, 87, 2200, 110, 95, 104, 122, 99,
91, 108, 130, 96, 101])
print(f"observed median: {np.median(data):.0f} ms")
# Resample WITH replacement 10,000 times; take the median each time.
medians = np.array([
np.median(rng.choice(data, size=len(data), replace=True))
for _ in range(10_000)
])
lo, hi = np.percentile(medians, [2.5, 97.5])
print(f"95% bootstrap CI for the median: ({lo:.0f}, {hi:.0f}) ms")observed median: 102 ms 95% bootstrap CI for the median: (96, 115) ms
The walkthrough
replace=True is the load-bearing argument. Each fake sample of 15 typically omits about a third of the original values and repeats others — the resampling fingerprint.
np.percentile(medians, [2.5, 97.5]) cuts off the middle 95% of the pile. This is the percentile interval: the plainest of several bootstrap interval recipes.
Notice what the outlier did and did not do. The 2200 ms timeout sits in the data, and the interval (96, 115) ignores it politely — because the median ignores it. Bootstrap the mean of this data instead and the interval becomes wide and lopsided. The bootstrap reports the uncertainty of whatever statistic you feed it, honestly either way.
scipy has a batteries-included version with a better interval recipe (BCa) as its default:
from scipy import stats
res = stats.bootstrap((data,), np.median, confidence_level=0.95,
random_state=rng)
print(res.confidence_interval)ConfidenceInterval(low=98.0, high=115.0)
Nearly the same answer — BCa nudged the lower end up by correcting for the pile's slight skew. On strongly skewed statistics that correction earns its keep.
Common mistakes
Resampling without replacement. Every "fake sample" is the original in shuffled order; every median identical; the interval collapses to a point. If your interval has zero width, check this first.
Bootstrapping tiny samples and trusting the result. With 5 values there are few distinct resamples; the interval is noise wearing a suit. Below roughly 15–20 values, treat bootstrap intervals as sketches.
Breaking dependence. Resampling individual days of a time series shreds the ordering that makes it a time series. Dependent data needs block bootstraps — resample chunks, not points. Same trap for grouped data: resample customers, not individual clicks of customers.
Bootstrapping the test set to compare two models, resampling each separately. Their scores share the same examples, so resample paired rows and bootstrap the score difference.
Try it yourself
Swap np.median for np.mean and rerun. Predict first: which direction will the interval stretch, and why will it not be symmetric around the observed mean?
What to learn next
- Permutation tests — resampling's sibling for yes/no hypothesis questions.
- Confidence intervals — the classical formulas the bootstrap generalises.
- Monte Carlo simulation — the simulate-and-count principle underneath it all.
Researcher — Mathematics and papers.
The plug-in principle
Let $\hat\theta = t(F_n)$ be a statistic written as a functional of the empirical distribution $F_n$ (which puts mass $\frac{1}{n}$ on each observation). The unknown sampling distribution of $\hat\theta$ under the true $F$ is approximated by the computable sampling distribution under $F_n$:
$$ \operatorname{Law}_F\left(t(F_n)\right) \;\approx\; \operatorname{Law}_{F_n}\left(t(F_n^*)\right) $$
Where:
- $F_n^$ — the empirical distribution of a resample $X_1^, \dots, X_n^*$ drawn i.i.d. from $F_n$.
- $t(\cdot)$ — the statistical functional (median, correlation, …).
Monte Carlo with $B$ resamples approximates the right-hand side; the statistical error of the bootstrap is in the substitution $F \to F_n$, not in $B$. Standard choices: $B \approx 1000$ for standard errors, $B \geq 10{,}000$ for tail quantiles.
Consistency and failure cases
The bootstrap is consistent when $t$ is suitably smooth (Hadamard-differentiable) at $F$ — covering means, quantiles of continuous distributions, correlations, regression coefficients. Known failures:
- Extremes: $t = \max$; the resample maximum equals the sample maximum with probability $1 - (1 - \tfrac1n)^n \to 1 - e^{-1}$, so the bootstrap distribution degenerates.
- Non-smooth points: parameters on a boundary; $|\mu|$ at $\mu = 0$.
- Heavy tails: infinite-variance means.
- Remedy in several such cases: the m-out-of-n bootstrap (resample $m \ll n$) restores consistency at an efficiency price (Bickel, Götze and van Zwet, 1997).
Interval constructions and their accuracy
For a $1-2\alpha$ interval, with $\hat\theta^*_{(q)}$ the $q$-quantile of the bootstrap replicates:
| method | interval | coverage error |
|---|---|---|
| percentile | $[\hat\theta^{(\alpha)},\, \hat\theta^{(1-\alpha)}]$ | $O(n^{-1/2})$ |
| basic (reflected) | $[2\hat\theta - \hat\theta^{(1-\alpha)},\, 2\hat\theta - \hat\theta^{(\alpha)}]$ | $O(n^{-1/2})$ |
| bootstrap-t | studentise each replicate | $O(n^{-1})$ |
| BCa | percentile with bias ($z_0$) and acceleration ($a$) corrections | $O(n^{-1})$ |
BCa (Efron, 1987, Better bootstrap confidence intervals) estimates $z_0$ from the fraction of replicates below $\hat\theta$ and $a$ from a jackknife skewness estimate; it is transformation-respecting and second-order accurate, hence scipy's default. Theory via Edgeworth expansions: Hall (1992), The Bootstrap and Edgeworth Expansion.
Dependent and structured data
- Block bootstrap for stationary series: moving blocks (Künsch, 1989) or the stationary bootstrap with geometric block lengths (Politis and Romano, 1994); block length trades bias against variance, typically $\ell \sim n^{1/3}$.
- Cluster bootstrap: resample groups, preserving within-group dependence.
- Residual and pairs bootstraps for regression; the wild bootstrap (Wu, 1986) for heteroscedastic errors, multiplying residuals by random signs.
- Bayesian bootstrap (Rubin, 1981): Dirichlet weights instead of multinomial counts — a smoothed sibling with a posterior interpretation.
Cost
$O(B \cdot c(n))$ where $c(n)$ is one statistic evaluation — embarrassingly parallel. The origin paper: Efron (1979), Bootstrap methods: another look at the jackknife, which reframed Quenouille–Tukey jackknifing as a special case of a general simulation principle.
What to learn next
- Permutation tests — resampling's sibling for yes/no hypothesis questions.
- Confidence intervals — the classical formulas the bootstrap generalises.
- Monte Carlo simulation — the simulate-and-count principle underneath it all.