Hypothesis Testing and Inference
The central limit theorem
Averages of enough samples form a bell curve, no matter how ugly the individual data looks — the quiet result that makes most of statistics work.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The central limit theorem says: average enough independent samples, and the averages form a bell curve — even when the raw data does not.
Individual auto-rickshaw rides through the city are chaotic. One catches every green light; another hits a procession and takes an hour. The pattern of single rides is lopsided and wild.
But track your monthly average commute, month after month. Those averages cluster tamely around a middle — a few low, a few high, most in between. Averaging washed the chaos into a bell shape.
Why it matters
Look back at this section. The t-test trusted averages to behave. Confidence intervals stretched "about two wobbles" either side, assuming a bell. Why is a bell curve hiding inside every recipe?
This theorem is the answer. It is the load-bearing wall of the whole building. It explains why tools built on bell-curve assumptions keep working on data that looks nothing like a bell.
How it works
Extreme averages require conspiracies. For your monthly average to be terrible, many rides must go wrong together — and independent bad luck rarely coordinates.
single rides: ▂▇█▅▃▂▁▁▁▁▁ ← lopsided, long slow tail
averages of 10: ▁▃▇█▇▃▁ ← pulling toward a bell
averages of 50: ▁▆█▆▁ ← narrow, symmetric bellTwo things happen as you average more samples. The shape becomes a bell, whatever shape you started from. And the bell gets narrower — averages of many samples barely wobble.
The theorem needs the samples to be independent and the extremes not too wild. Averaging a few days of stock-market chaos, where one crash dwarfs everything, can defeat it.
A real example you have seen
Exit polls interview a few thousand voters out of millions, then report a percentage with a small margin. That confidence comes from here: averages (and percentages are averages) settle into a predictable bell with a known width.
Same for quality control — a factory tastes a handful of biscuits per batch, not the batch.
Remember this
- Averages of independent samples approach a bell curve, regardless of the raw shape.
- More samples per average → a narrower bell.
- This theorem is why t-tests and confidence intervals work on non-bell data.
What to learn next
- Statistical power and sample size — using the square-root law to plan experiments.
- The bootstrap — what to do when your statistic has no CLT of its own.
- Monte Carlo simulation — the simulate-and-average method this theorem underwrites.
Developer — Code and libraries.
Setup
pip install numpy scipyOutputs verified with numpy 1.26 and scipy 1.14, CPU. Simulated digits reproduce with these versions; other versions may differ in the last decimal.
Watching the theorem happen
Delivery times drawn from a heavily skewed distribution — most fast, some awful. We average groups of 2, 10 and 50 and watch the skew die.
import numpy as np
from scipy import stats
rng = np.random.default_rng(0)
population = rng.exponential(scale=30, size=100_000)
print(f"one delivery at a time: mean {population.mean():.1f} min, "
f"skewness {stats.skew(population):.2f}")
for n in [2, 10, 50]:
means = rng.exponential(scale=30, size=(10_000, n)).mean(axis=1)
print(f"averages of {n:2d}: mean {means.mean():5.1f}, "
f"spread {means.std():5.2f}, skewness {stats.skew(means):5.2f}")one delivery at a time: mean 29.9 min, skewness 2.01 averages of 2: mean 29.8, spread 21.35, skewness 1.36 averages of 10: mean 30.1, spread 9.63, skewness 0.64 averages of 50: mean 30.0, spread 4.25, skewness 0.30
The walkthrough
skewness measures lopsidedness: 0 is symmetric, 2.01 is strongly right-tailed. Watch the column melt: 2.01 → 1.36 → 0.64 → 0.30. Averaging is actively destroying the skew.
The spread column follows a square-root law. Raw deliveries wobble by 30. Averages of 10 wobble by 9.63 — close to 30 divided by the square root of 10 (9.49). Averages of 50: 4.25, near 30 over root-50 (4.24). This is the exact law behind "four times the data, half the error bar".
The mean column never moves. Averaging changes the shape and spread of the wobble, never the centre. Estimates stay honest; they get precise.
How many samples is "enough"? Folklore says 30. Honest answer: it depends on the skew. Mild skew — a dozen may do. This exponential still shows skew 0.30 at n = 50. Heavier tails need more; a few distributions (infinite-variance ones) never comply.
Common mistakes
Believing the CLT normalises your data. It never does. The raw deliveries stay skewed forever; only the averages turn bell-shaped. If a method needs normal raw data, the CLT is no rescue.
Applying it to dependent samples. Server response times in a traffic spike move together; effective sample size collapses and the bell arrives late or not at all. The independence clause has teeth.
Trusting it for medians, maxima or ratios. The theorem is about sums and averages. Maxima follow different laws entirely; for medians and ratios, use the bootstrap.
Trusting it in the far tails. The bell approximation is best in the middle. "Five-sigma" events in skewed data occur far more often than the bell predicts — a lesson finance learned expensively in 2008.
Try it yourself
Replace rng.exponential(scale=30, ...) with rng.standard_cauchy(...) (a distribution with wild, heavy tails) and rerun. Watch the spread column refuse to shrink — you have met a distribution the theorem cannot tame.
What to learn next
- Statistical power and sample size — using the square-root law to plan experiments.
- The bootstrap — what to do when your statistic has no CLT of its own.
- Monte Carlo simulation — the simulate-and-average method this theorem underwrites.
Researcher — Mathematics and papers.
Statement
Let $X_1, X_2, \dots$ be i.i.d. with mean $\mu$ and finite variance $\sigma^2$. Then:
$$ \frac{\bar{X}_n - \mu}{\sigma / \sqrt{n}} \;\xrightarrow{d}\; \mathcal{N}(0, 1) \quad \text{as } n \to \infty $$
Where:
- $\bar{X}n = \frac{1}{n}\sum{i=1}^n X_i$ — the sample mean.
- $\mu, \sigma^2$ — population mean and variance; finiteness of $\sigma^2$ is the binding condition.
- $\xrightarrow{d}$ — convergence in distribution.
Equivalently $\bar{X}_n \approx \mathcal{N}(\mu, \sigma^2/n)$ for large $n$ — the source of the $\sqrt{n}$ in every standard error.
Proof sketch and rate
The characteristic function of the standardised sum is $\left[\varphi!\left(\tfrac{t}{\sqrt{n}}\right)\right]^n$ with $\varphi(s) = 1 - \tfrac{s^2}{2} + o(s^2)$ after centring and scaling; the limit is $e^{-t^2/2}$, the standard normal, by Lévy's continuity theorem.
The Berry–Esseen theorem (1941/1942) quantifies the speed: with finite third absolute moment $\rho = \mathbb{E}|X - \mu|^3$,
$$ \sup_x \left| F_n(x) - \Phi(x) \right| \leq \frac{C \rho}{\sigma^3 \sqrt{n}} $$
with $C \leq 0.4748$ (Shevtsova, 2011). The error decays like $n^{-1/2}$ and scales with the skew — the formal version of "how many samples depends on how ugly the data is". Edgeworth expansions refine this: the leading error term is proportional to skewness, the next to kurtosis.
Generalisations and failure modes
- Lindeberg–Feller CLT: independent, non-identical summands, under the Lindeberg condition (no single term dominates the variance).
- Dependent data: CLTs exist for martingales and mixing sequences (weak, decaying dependence); strong dependence changes the norming or the limit.
- Infinite variance: with tail index $\alpha < 2$, sums converge to $\alpha$-stable laws instead (generalised CLT, Gnedenko and Kolmogorov, 1954). The Cauchy case in the exercise is $\alpha = 1$: the mean of $n$ Cauchy draws is Cauchy again — averaging accomplishes nothing.
- Extremes: maxima converge to Gumbel/Fréchet/Weibull families (Fisher–Tippett–Gnedenko), not the normal.
The delta method
Smooth functions of averages inherit asymptotic normality:
$$ \sqrt{n}\left(g(\bar{X}_n) - g(\mu)\right) \xrightarrow{d} \mathcal{N}!\left(0, \; g'(\mu)^2 \sigma^2\right) $$
for differentiable $g$ with $g'(\mu) \neq 0$ — the workhorse for standard errors of ratios, log-odds and transformed metrics.
Historical line
De Moivre (1733) for binomial sums; Laplace (1810) in general form; Lyapunov (1901) with a rigorous proof under moment conditions; Lindeberg (1922) with the modern sufficient condition. Pólya coined "central" (zentraler Grenzwertsatz, 1920) for its central role in probability, not for the centre of the distribution.
What to learn next
- Statistical power and sample size — using the square-root law to plan experiments.
- The bootstrap — what to do when your statistic has no CLT of its own.
- Monte Carlo simulation — the simulate-and-average method this theorem underwrites.