Hypothesis Testing and Inference
Confidence intervals
A confidence interval replaces a single guess with an honest range — and its 95% refers to the method's success rate, not to your one interval.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A confidence interval is a range that says "the true value is probably somewhere in here" — with a stated success rate for the method.
Ask a fisherman when he will return. He never says "6:14 pm". He says "between 6 and 7" — a range wide enough that he is usually right, narrow enough to be useful.
A single number pretends to a precision you do not have. A range tells the truth about how much you know. Statistics formalises the fisherman's habit.
Why it exists
Your test set says the model is 84.0% accurate. On different test data, it would have said 83.1%, or 85.2%. The single number hides that wobble.
Jerzy Neyman introduced confidence intervals in 1937 to carry the wobble along with the estimate. "84% accurate, give or take 3" and "84% accurate, give or take 0.2" justify completely different decisions. A plain 84% hides which world you are in.
How it works
Collect a sample and compute your estimate. Then compute how much such estimates typically wobble, and stretch a range around yours.
one sample → estimate 84% → range: 81% ─────●───── 87%
the promise (for a 95% interval):
repeat the whole study 100 times, build 100 ranges →
about 95 of them will contain the true valueRead that promise carefully. The 95% belongs to the method, not to your one range. Any single range either contains the truth or misses it — you never learn which. What you know is the method's track record.
This distinction is genuinely slippery. Nearly everyone gets it wrong the first time, and that is normal.
Wider range or higher confidence — pick one. More data is the only way to get both.
A real example you have seen
Election polls: "42%, plus or minus 3%". Weather: "31 to 34 degrees expected". Delivery apps: "arriving in 25 to 35 minutes". Each is an interval doing the fisherman's job — honest about uncertainty, still useful for decisions.
Remember this
- An interval reports the estimate and its wobble, in one object.
- The 95% is the method's long-run success rate, not your interval's probability.
- Narrower needs more data; more confident needs wider. No free lunch.
What to learn next
- The central limit theorem — why the mean's wobble is bell-shaped at all.
- The bootstrap — intervals for medians, correlations, anything, by resampling.
- Bayesian vs frequentist thinking — the interval that does mean "95% probability the truth is here".
Developer — Code and libraries.
Setup
pip install numpy scipyOutputs verified with numpy 1.26 and scipy 1.14, CPU. The simulation line reproduces with these versions; other versions may shift it slightly.
Three intervals, then the promise, checked
Delivery times in minutes for 12 orders, then an on-time proportion, then a simulation that puts the "95%" claim on trial.
import numpy as np
from scipy import stats
times = np.array([31, 28, 35, 30, 42, 29, 33, 38, 27, 36, 31, 34])
mean = times.mean()
sem = stats.sem(times) # standard error of the mean
lo, hi = stats.t.interval(0.95, df=len(times) - 1, loc=mean, scale=sem)
print(f"mean {mean:.1f} min, 95% CI ({lo:.1f}, {hi:.1f})")
# A proportion: 43 of 50 deliveries arrived on time.
ci = stats.binomtest(43, 50).proportion_ci(confidence_level=0.95)
print(f"on-time rate 0.86, 95% CI ({ci.low:.3f}, {ci.high:.3f})")
# The promise, tested: simulate 1000 studies where the truth IS 33.
rng = np.random.default_rng(7)
hits = 0
for _ in range(1000):
s = rng.normal(33, 4, size=12)
lo, hi = stats.t.interval(0.95, 11, loc=s.mean(), scale=stats.sem(s))
hits += (lo <= 33 <= hi)
print(f"intervals that caught the true mean: {hits} / 1000")mean 32.8 min, 95% CI (30.0, 35.6) on-time rate 0.86, 95% CI (0.733, 0.942) intervals that caught the true mean: 945 / 1000
The walkthrough
stats.sem is the standard error — the typical wobble of the mean across repeat samples. It shrinks with the square root of the sample size: four times the data halves the error bar.
stats.t.interval stretches about two standard errors either side, using the t-distribution to be honest about small samples. With 12 points, it stretches slightly more than two.
The proportion interval is lopsided on purpose. 0.86 sits closer to the upper end (0.942) than the lower (0.733). Proportions near 0 or 1 cannot wobble symmetrically — the exact method knows that; the naive "plus-minus two standard errors" formula does not.
945 out of 1000. The method promised roughly 95% and delivered — that is what "confidence" operationally means. No single interval knew whether it was one of the 945.
Intervals and tests are the same machinery. A 95% interval for a difference that excludes zero is a two-sided test rejecting at 0.05. The interval reports strictly more: direction, size, and precision.
Common mistakes
Saying "there is a 95% chance the truth is in (30.0, 35.6)". In the frequentist framework, the truth is fixed; only intervals vary. If you want a statement about the parameter's probability, that is a credible interval — see bayesian-vs-frequentist.
Using ±2 standard errors for accuracy near 100%. It can produce intervals beyond 1.0. Use proportion_ci or a Wilson interval for benchmark accuracies.
Reporting model accuracy with no interval at all. 84.0% from 200 test examples carries a 95% interval running roughly 78% to 89% — about ten points wide. A model that scores two points higher has not been shown to be better. See model-evaluation.
Comparing two models by eyeballing overlapping intervals. Two intervals can overlap while the difference is significant. Build the interval for the difference itself instead.
Try it yourself
Change the sample size in the simulation from 12 to 3 and rerun. Coverage stays near 950 — but print one interval's width and see the real cost of tiny samples.
What to learn next
- The central limit theorem — why the mean's wobble is bell-shaped at all.
- The bootstrap — intervals for medians, correlations, anything, by resampling.
- Bayesian vs frequentist thinking — the interval that does mean "95% probability the truth is here".
Researcher — Mathematics and papers.
Formal definition
A $1-\alpha$ confidence procedure for parameter $\theta$ is a data-dependent set $C(X)$ with
$$ \Pr_\theta\left(\theta \in C(X)\right) \geq 1 - \alpha \quad \text{for every } \theta $$
Where:
- $\theta$ — the fixed, unknown parameter.
- $C(X)$ — the random interval, a function of the data $X$.
- $1 - \alpha$ — the coverage probability, a property of the procedure across repetitions.
The probability statement is over $X$, not $\theta$ — the source of every misinterpretation.
Classical constructions
Pivots. A pivot is a function of data and parameter whose distribution is parameter-free. For normal samples, $T = \frac{\bar{X} - \mu}{S/\sqrt{n}} \sim t_{n-1}$ exactly; inverting $|T| \leq t_{1-\alpha/2,\,n-1}$ gives
$$ \bar{X} \pm t_{1-\alpha/2,\,n-1}\, \frac{S}{\sqrt{n}} $$
with $S$ the sample standard deviation.
Test inversion. $C(X) = {\theta_0 : \text{test of } H_0!: \theta = \theta_0 \text{ does not reject}}$. Every test family yields an interval family and conversely — the duality used throughout applied statistics.
Asymptotic (Wald) intervals. $\hat\theta \pm z_{1-\alpha/2}\,\widehat{\operatorname{se}}(\hat\theta)$, justified by the estimator's asymptotic normality. Wald intervals for binomial proportions are notoriously poor: coverage oscillates far below nominal even at large $n$ (Brown, Cai and DasGupta, 2001, Interval estimation for a binomial proportion). Preferred: Wilson (score inversion, 1927) or Clopper–Pearson (exact, conservative, 1934 — what proportion_ci computes by default).
Coverage versus precision
Among valid procedures, shorter is better. The Neyman–Pearson duality carries over: uniformly most accurate intervals correspond to uniformly most powerful tests. Exact methods (Clopper–Pearson) buy guaranteed coverage at the price of over-coverage (real coverage above $1-\alpha$), i.e. wider intervals; score-based intervals trade a little guarantee for shorter length and near-nominal average coverage.
Beyond the classical setups
- Bootstrap intervals (percentile, BCa) replace distributional assumptions with resampling — see the bootstrap lesson; BCa achieves second-order accuracy, $O(n^{-1})$ coverage error versus $O(n^{-1/2})$ for the percentile method (Efron, 1987).
- Profile likelihood intervals for multiparameter models invert the likelihood-ratio test using Wilks' theorem.
- Simultaneous intervals (Bonferroni, Scheffé, Tukey) hold jointly over many parameters; per-comparison intervals do not.
- Conformal prediction gives distribution-free prediction intervals for individual outcomes — a different target from parameter intervals (Vovk, Gammerman and Shafer, 2005, Algorithmic Learning in a Random World).
Origin: Neyman (1937), Outline of a theory of statistical estimation based on the classical theory of probability.
What to learn next
- The central limit theorem — why the mean's wobble is bell-shaped at all.
- The bootstrap — intervals for medians, correlations, anything, by resampling.
- Bayesian vs frequentist thinking — the interval that does mean "95% probability the truth is here".