Statistics
Statistics is how you judge a whole pot from one spoonful. Your test set is that spoonful, which is why a benchmark number needs an error bar.
- 17 min read
- 3 reading levels
- Published
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Statistics is how you say something true about a whole crowd after looking at only a few of them.
Think about a cook with a large pot of dal. She stirs it, lifts one spoonful, tastes it, and adds salt.
She never tastes the whole pot. One spoon told her enough about all of it. That move — judging the whole from a small taste — is the entire subject.
Why you should care
Every number you will ever report about an AI model is a spoonful.
You test on a few hundred examples and get an accuracy. The world contains millions of examples you did not test on. Your number is a taste, and the honest question is how much the taste can be trusted.
Teams ship models on differences that are pure noise. They see one number beat another and celebrate. Statistics is what stops you doing that.
Question one: what is typical
The average is the fair-share number. Add everything up, split it equally among all of them.
It is useful, and on its own it lies more often than people expect.
Picture a small street where most people earn an ordinary wage. One film star moves in. The average income on that street now suggests everyone is wealthy. Nobody's actual life changed.
So there is a second number: the middle value. Line everyone up from poorest to richest and pick the person standing in the middle. The film star moves the average enormously and the middle value hardly at all.
When those two numbers disagree, something lopsided is going on. That gap is a signal, and it is free to check.
Question two: how spread out
Two students both average sixty marks across a year.
six tests each. a longer bar means higher marks.
Asha ###### ##### ###### ###### #####
Ravi ######### ## ######## ### #########The same average describes two completely different students. You would prepare them for an exam in different ways.
The number that captures this is the spread — how far, on average, things sit from the middle. Small spread means dependable. Large spread means the average is hiding something.
Spread is the single most under-reported number in AI. A model that averages high accuracy but collapses on some kinds of input is Ravi.
The taste has to be a fair one
Back to the pot. The cook stirs before she tastes, and that is not a habit — it is the whole method.
Salt sinks. Spices settle. A spoon lifted straight off the top tells you about the top, not the pot.
stirred pot → one spoon → a fair picture of the whole
unstirred pot → one spoon → a confident, wrong answerData settles in exactly the same way. Files arrive sorted by date, or by category, or by whoever collected them. Take the first thousand rows and you have tasted the top of the pot.
This is why you shuffle your data before splitting it. It is also why a bigger unstirred spoon does not help you. More of the wrong thing is still the wrong thing.
Small spoon, noisy answer
A tiny taste can mislead even from a well-stirred pot. That is not carelessness; it is unavoidable.
The fix is a habit rather than a formula. Never report a single number on its own. Report how much that number would wobble if you had tasted a different spoonful.
With a few hundred test examples, that wobble is larger than most people guess. It is often larger than the improvement being celebrated.
The link to a model that memorises
Here is where spread turns into the most famous failure in machine learning.
A student who memorises last year's question paper scores well on last year's paper. Give them this year's paper and they collapse. They learned the spoonful, not the dal.
A model does this readily. It can fit every wrinkle of your training data, including the accidents. Those wrinkles were noise in your particular sample, and they will not repeat.
The warning sign is always the same shape. Performance on data it has seen keeps improving. Performance on data it has not seen starts getting worse. There is a whole lesson on this: overfitting and underfitting.
Where you have already seen it
- Exit polls, which ask a few thousand voters and predict a whole state.
- Quality checks at a factory, where a handful of pieces from each batch get tested.
- A doctor's blood test, run on a few drops rather than all your blood.
- App updates rolled out to a slice of users first, to see whether anything breaks.
The honest part
Two hard truths, and both cost people real money.
A biased sample cannot be rescued by size. If your spoon only ever comes from the top of the pot, taking a bigger spoon changes nothing. Collect differently, or admit the limit out loud.
Most reported improvements are noise. Run the same training twice with different starting randomness and you get different numbers. If a new method beats the old one by less than that gap, nothing has been shown.
This is uncomfortable and it is standard. Good teams rerun with several starting points and report the spread.
Remember this
- Every number you measure comes from a sample, so it comes with a wobble.
- The average alone hides things. Always look at the spread as well.
- Stir before you taste: shuffle your data, or your split is measuring the wrong thing.
What to learn next
- Overfitting and underfitting — memorising the spoonful instead of learning the dal.
- Train, test and validation splits — how to take a fair taste.
- Model evaluation — which numbers to report and which to distrust.
Developer — Code and libraries.
Setup
pip install numpyThree programs, each answering a question you will face on a real project. What is typical? Is this difference real? Am I overfitting?
The average alone is not enough
import numpy as np
model_a = np.array([88.0, 87.0, 89.0, 88.0, 88.0]) # accuracy on five different splits
model_b = np.array([95.0, 78.0, 92.0, 80.0, 95.0])
for name, s in (("model A", model_a), ("model B", model_b)):
print(f"{name}: mean {s.mean():5.1f} spread {s.std(ddof=1):5.2f} worst {s.min():5.1f}")
salaries = np.array([25.0, 28.0, 30.0, 32.0, 35.0, 400.0]) # thousands per month
print("salary mean :", round(float(salaries.mean()), 2))
print("salary median:", round(float(np.median(salaries)), 2))model A: mean 88.0 spread 0.71 worst 87.0 model B: mean 88.0 spread 8.34 worst 78.0 salary mean : 91.67 salary median: 31.0
Identical means, and the two models are nothing alike. Model B would have failed a user badly on one of those five splits. A results table showing only the mean makes them look like the same model.
The salary rows show the other half. One large value drags the mean to 91.67, a figure that describes nobody in the list. The median stays at 31.0, which describes the group honestly.
ddof=1 matters. NumPy's default is ddof=0, which divides by n. That is right when you hold the whole population. On a sample it is biased low. For sample data — which is all data you will ever handle — use ddof=1. On five values the difference is over 10%.
Is that difference real?
You compare two models on a 100-example test set. One gets 91%, the other 89%. Ship the winner?
n = 100
for name, acc in (("A", 0.91), ("B", 0.89)):
se = (acc * (1 - acc) / n) ** 0.5
lo, hi = acc - 1.96 * se, acc + 1.96 * se
print(f"model {name}: {acc:.2f} on {n} examples -> 95% interval [{lo:.3f}, {hi:.3f}]")
print()
for n in (100, 1_000, 10_000, 100_000):
half = 1.96 * (0.9 * 0.1 / n) ** 0.5
print(f"test set of {n:>6}: accuracy known to +/- {half*100:.2f} points")model A: 0.91 on 100 examples -> 95% interval [0.854, 0.966] model B: 0.89 on 100 examples -> 95% interval [0.829, 0.951] test set of 100: accuracy known to +/- 5.88 points test set of 1000: accuracy known to +/- 1.86 points test set of 10000: accuracy known to +/- 0.59 points test set of 100000: accuracy known to +/- 0.19 points
The two intervals overlap almost completely. That 2-point win is inside the noise, and the ranking could flip on a different test set.
The second table is the one to memorise. On 100 examples you cannot resolve anything smaller than about 6 points. On 1,000 you get to about 2 points. Precision improves with the square root of n, so a hundredfold larger test set buys only a tenfold tighter estimate.
This is the arithmetic behind a great deal of published nonsense. A leaderboard gap of 0.4 points on a small evaluation set is not a result.
Two honest caveats on that formula. It is the normal approximation to a binomial interval. Near 0 and 1 it misbehaves, and a Wilson interval is the fix.
It also assumes independent test examples. Sentences from one document are not independent. Neither are frames from one video. The true interval is then wider than this.
Watching a model overfit
import numpy as np
rng = np.random.default_rng(0) # fixed seed, so your numbers match this page
x = np.linspace(0.0, 1.0, 12)
y = np.sin(2 * np.pi * x) + rng.normal(0.0, 0.25, size=x.size)
train = np.arange(0, 12, 2) # six points to learn from
test = np.arange(1, 12, 2) # six unseen points to be judged on
for deg in (1, 3, 5):
coef = np.polyfit(x[train], y[train], deg)
tr = float(np.mean((np.polyval(coef, x[train]) - y[train]) ** 2))
te = float(np.mean((np.polyval(coef, x[test]) - y[test]) ** 2))
print(f"degree {deg}: train error {tr:7.4f} test error {te:8.3f}")degree 1: train error 0.2786 test error 0.331 degree 3: train error 0.0199 test error 0.023 degree 5: train error 0.0000 test error 1.370
Read the two columns together, because neither one alone tells you anything.
Degree 1 is a straight line trying to follow a curve. Both errors are high and similar. That is underfitting: the model is not flexible enough.
Degree 3 is the sweet spot here. Low error on both, and the two numbers agree with each other.
Degree 5 fits six training points with six coefficients, so it passes through every one of them exactly. Training error is essentially zero. Test error is 60 times worse than degree 3.
Nothing about degree 5 is broken. It fit the noise, and the noise does not repeat on new points. This is the whole of overfitting in nine lines of code, and you can watch it happen.
That last row also explains why a training loss of zero is a warning and not an achievement.
Common mistakes
1. Reporting one number from one run. Train the same model three times with different seeds. The spread across those runs is your resolution limit. Any claimed improvement smaller than it is unproven.
2. Splitting before shuffling. Data arrives sorted more often than you expect — by date, by class, by collection batch. train_test_split shuffles by default, df[:800] does not.
3. Leaking test data into training. Scaling the whole dataset before splitting means your training set has seen the test set's mean. Fit the scaler on training data alone, then apply it to test data.
4. Tuning on the test set. Every time you look at test accuracy and change something, you overfit to the test set a little. Keep a validation set for decisions and touch the test set once.
5. Averaging accuracies across imbalanced groups. Score 99% on a common class and 20% on a rare one. The overall number still looks fine. Report per-group results.
6. Using .std() without ddof=1. Silently biased low on small samples, and nothing warns you.
Try it yourself
In overfit.py, change the noise from 0.25 to 0.0. Rerun and watch degree 5 stop overfitting entirely. Overfitting is a response to noise, not to model size on its own.
Then change the seed from 0 to 1 and rerun. The exact numbers move. The pattern across the three degrees does not, and that stability is what makes the conclusion trustworthy.
What to learn next
- Train, test and validation splits — building an evaluation you can defend.
- Model evaluation — metrics beyond accuracy.
- Overfitting and underfitting — diagnosing and fixing the gap above.
Researcher — Mathematics and papers.
Estimators and their error
An estimator theta_hat is a function of the sample used to approximate a population quantity theta. Its quality decomposes exactly:
MSE(theta_hat) = E[ (theta_hat - theta)**2 ]
= Bias(theta_hat)**2 + Var(theta_hat)
Bias(theta_hat) = E[theta_hat] - thetaE[.]is expectation over repeated draws of the sample, not over the data points in one sample.Varmeasures how much the estimate moves between samples.
The sample variance with divisor n - 1 (Bessel's correction) is unbiased for the population variance; the divisor n version is biased low by a factor (n-1)/n. That is ddof=1 in the Developer block. Note that the corrected variance is unbiased. Its square root, the standard deviation, is not. Jensen's inequality guarantees a downward bias there.
The bias-variance decomposition for prediction
For squared loss at a point x, with y = f(x) + eps and Var[eps] = sigma**2:
E[ (y - g_D(x))**2 ] = sigma**2
+ ( E_D[g_D(x)] - f(x) )**2 the squared bias
+ E_D[ ( g_D(x) - E_D[g_D(x)] )**2 ] the varianceg_D is the model fitted on training set D, and E_D averages over draws of D. The first term is irreducible. The polynomial experiment above is this decomposition made visible: degree 1 is bias-dominated, degree 5 variance-dominated.
Two honest limitations. The decomposition is specific to squared loss. The 0-1 loss analogue (Domingos, 2000) is messier and not additive in general.
The classical U-shaped picture is also incomplete for overparameterised models. Double descent (Belkin et al., 2019; Nakkiran et al., 2020) shows test error falling again past the interpolation threshold. The bias-variance frame still fits the underparameterised regime, where most tabular and small-model work lives.
Confidence intervals for a proportion
Test accuracy on n independent examples is a binomial proportion. The Wald interval used in the Developer block is
p_hat +/- z * sqrt( p_hat (1 - p_hat) / n )with z = 1.96 for 95% nominal coverage. Its actual coverage is poor for small n or extreme p_hat, sometimes far below nominal (Brown, Cai and DasGupta, 2001). Prefer the Wilson score interval:
p_hat + z**2/(2n) +/- z * sqrt( p_hat(1-p_hat)/n + z**2/(4 n**2) )
CI = ------------------------------------------------------------------
1 + z**2/nIt stays inside [0, 1] and holds coverage far better at the edges. scipy.stats.binomtest(...).proportion_ci(method="wilson") gives it directly.
Independence is the load-bearing assumption, and it is routinely violated. Evaluation data comes in clusters: utterances per speaker, frames per video, questions per document. Those clusters correlate.
Use a cluster bootstrap that resamples whole clusters. Without it the interval is optimistic by a wide margin.
Comparing two models properly
Two models evaluated on the same test set produce paired, dependent measurements. A two-sample test discards that pairing and loses power.
McNemar's test is the right default for classifier comparison on a shared test set. Let b be the count where model A is right and B wrong, and c the reverse. Under the null of equal error rates:
chi2 = ( |b - c| - 1 )**2 / ( b + c ) ~ chi-squared with 1 dfOnly the disagreements carry information; the examples both models get right are uninformative. Dietterich (1998) evaluated five tests for this setting and recommended McNemar and 5x2-fold cross-validated t for the two common scenarios.
Bootstrap is the general-purpose alternative, and it needs no distributional assumption. Resample the test set with replacement B times. Recompute the metric difference on each resample. Take the empirical 2.5th and 97.5th percentiles.
B = 10,000 is cheap and standard. It works for metrics with no closed-form variance: BLEU, F1, nDCG, calibration error. That covers most of the interesting ones.
The multiplicity problem
Evaluating m variants at level alpha gives a family-wise error rate up to 1 - (1 - alpha)**m. At alpha = 0.05 and m = 20, that is 64%. A hyperparameter sweep is exactly this situation.
Bonferroni (alpha / m) controls family-wise error and is conservative. Benjamini-Hochberg controls the false discovery rate. It is the better choice when several genuine effects are expected.
Neither repairs the deeper problem. Pick the best configuration on a validation set, then report that configuration's validation score. The number is upward-biased by construction. Report a held-out score from data used for nothing else.
Related, and worth naming: reporting the maximum over k seeds rather than the mean and spread. Dodge et al. (2019) showed results depend strongly on the search budget spent. They proposed reporting expected validation performance as a function of that budget.
Effect size, not only significance
With a large enough test set every difference becomes statistically significant. Significance answers "is it non-zero", not "does it matter". Report the difference with its interval, and state the smallest difference that would change a decision, before running the test.
Power analysis inverts the same arithmetic. To detect a difference delta in accuracy at power 1 - beta and level alpha, the required n scales as (z_{alpha/2} + z_beta)**2 * sigma**2 / delta**2. Resolving a 1-point difference needs roughly a hundred times the data of a 10-point one.
Sampling and validity
Selection effects dominate sampling error in practice, and no amount of n fixes them.
- Covariate shift: the input distribution moves,
P(y|x)holds. Importance weighting can correct it when supports overlap. - Label shift: class priors move. Correctable with a confusion matrix from the source domain (Lipton et al., 2018).
- Concept drift:
P(y|x)itself changes. Nothing corrects this without new labels. See monitoring and model drift. - Benchmark contamination: evaluation data present in pretraining corpora. This makes reported scores meaningless rather than mildly biased. Assume contamination on any public benchmark until a decontamination check is reported.
Stratified sampling reduces variance when strata differ and you know the strata proportions. Cross-validation reduces variance in the estimate, at the cost of k times the compute. One warning about it is widely ignored. The standard error across CV folds is not a valid standard error for the generalisation estimate, because the folds share training data. Bengio and Grandvalet (2004) proved no unbiased estimator of the variance of k-fold CV exists.
References
- Wasserman, L. All of Statistics. Springer, 2004. Compact and aimed at exactly this audience.
- Efron, B., Tibshirani, R. An Introduction to the Bootstrap. Chapman and Hall, 1993.
- Brown, L. D., Cai, T. T., DasGupta, A. "Interval Estimation for a Binomial Proportion." Statistical Science 16(2), 101–133, 2001.
- Dietterich, T. G. "Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms." Neural Computation 10(7), 1895–1923, 1998.
- Domingos, P. "A Unified Bias-Variance Decomposition." ICML, 2000.
- Bengio, Y., Grandvalet, Y. "No Unbiased Estimator of the Variance of K-Fold Cross-Validation." JMLR 5, 1089–1105, 2004.
- Belkin, M., Hsu, D., Ma, S., Mandal, S. "Reconciling modern machine-learning practice and the classical bias-variance trade-off." PNAS 116(32), 2019.
- Nakkiran, P., et al. "Deep Double Descent." ICLR, 2020. arXiv:1912.02292
- Dodge, J., Gururangan, S., Card, D., Schwartz, R., Smith, N. A. "Show Your Work: Improved Reporting of Experimental Results." EMNLP, 2019. arXiv:1909.03004
- Lipton, Z. C., Wang, Y.-X., Smola, A. "Detecting and Correcting for Label Shift with Black Box Predictors." ICML, 2018.
What to learn next
- Model evaluation — metric choice and threshold selection.
- Calculus — the other half of the maths, and how a model improves.
- Monitoring and model drift — detecting when the distribution moves under you.