Hypothesis Testing and Inference
ANOVA
ANOVA compares three or more group averages in one shot, by asking whether the groups differ more between themselves than within themselves.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
ANOVA tests whether three or more group averages differ by more than noise — all in a single test.
Three cooks make the same dal, and you want to know if any of them stands out. Every pot varies day to day, even from the same cook. The real question is whether the cooks differ more from each other than each cook differs from herself.
That comparison — between-group spread against within-group spread — is the whole of ANOVA. The name means analysis of variance: studying averages by studying spreads.
Why it exists
With three groups you could run three t-tests: A vs B, A vs C, B vs C. With six groups, fifteen tests. Every extra test is another lottery ticket for a false alarm.
Ronald Fisher, working at an agricultural research station in the 1920s, needed to compare many crop treatments fairly. ANOVA asks one combined question first: "is anything going on here at all?"
How it works
cook A pots: ●●●●● within-group spread: how much one
cook B pots: ●●●●● cook varies pot to pot
cook C pots: ●●●●●
between-group spread: how far the
A B C cooks' averages sit apart
─┴────┴────┴─
between spread ÷ within spread = the F-value
near 1 → groups blend into the noise
large → at least one group stands apartOne detail trips everyone: a significant result says at least one group differs. It does not say which. A follow-up comparison answers that — you will run one in the developer tab.
A real example you have seen
Farm trials comparing three fertilisers on identical plots. Schools comparing exam results across four teaching methods. An ML team comparing model accuracy across three prompt styles, with several runs each.
Any time you hear "we compared five variants", a version of this test decided whether the comparison meant anything.
Remember this
- ANOVA compares many group averages with one test, avoiding a pile-up of t-tests.
- The trick: between-group spread divided by within-group spread.
- A significant result means "something differs" — a follow-up test finds which.
What to learn next
- Non-parametric tests — the rank-based fallback when spreads and outliers misbehave.
- Multiple testing correction — the false-alarm budget problem ANOVA partially solves.
- Confidence intervals — reporting gaps with uncertainty instead of verdicts.
Developer — Code and libraries.
Setup
pip install numpy scipyOutputs verified with numpy 1.26 and scipy 1.14, CPU.
One F-test, then the follow-up
Crop yield in quintals per acre under three fertilisers, five plots each.
import numpy as np
from scipy import stats
a = np.array([21.5, 22.8, 20.9, 23.1, 22.0])
b = np.array([24.2, 25.1, 23.8, 24.9, 25.5])
c = np.array([22.1, 21.8, 23.0, 22.5, 21.4])
f, p = stats.f_oneway(a, b, c)
print(f"ANOVA: F = {f:.2f}, p = {p:.5f}")
# ANOVA says "at least one differs". Tukey's test says which pairs.
res = stats.tukey_hsd(a, b, c)
print(res)ANOVA: F = 19.98, p = 0.00015 Tukey's HSD Pairwise Group Comparisons (95.0% Confidence Interval) Comparison Statistic p-value Lower CI Upper CI (0 - 1) -2.640 0.000 -3.903 -1.377 (0 - 2) -0.100 0.976 -1.363 1.163 (1 - 0) 2.640 0.000 1.377 3.903 (1 - 2) 2.540 0.000 1.277 3.803 (2 - 0) 0.100 0.976 -1.163 1.363 (2 - 1) -2.540 0.000 -3.803 -1.277
The walkthrough
F = 19.98 — the between-group spread is nearly twenty times the within-group spread. Under "all fertilisers equal", an F this large appears about 15 times in 100,000 experiments.
tukey_hsd is the "which one?" answer. Group indices follow argument order: 0 = a, 1 = b, 2 = c. Rows (0 - 1) and (1 - 2) are decisive: fertiliser b beats both others by about 2.5 quintals. Row (0 - 2) shows a and c are indistinguishable — the interval spans zero.
Tukey's intervals are already corrected for making three comparisons at once. That is the reason to use it instead of three raw t-tests after the ANOVA.
Why not skip the ANOVA and go straight to pairwise tests? With three groups you could. With ten groups and 45 pairs, the single gatekeeper question keeps the false-alarm budget under control. See multiple-testing-correction for the general problem.
Common mistakes
Running ANOVA on two groups. It works — F is then t squared — but a t-test reports direction more readably.
Ignoring wildly unequal spreads. Classic ANOVA assumes similar within-group spread. f_oneway has no Welch option, but scipy ships an unequal-variance alternative: stats.alexandergovern(a, b, c). Reach for it when one group is far noisier than the rest.
Believing a significant F identifies the best group. It does not even say one group is best — a spread-out middle can trigger it. Read the pairwise intervals before making claims.
Forgetting that plots must be independent. Five yields from the same plot re-measured are not five samples. That structure needs repeated-measures ANOVA, a different tool.
Try it yourself
Shift every value in c up by 2.5 and rerun. Predict first: which Tukey rows flip to significant, and what happens to F?
What to learn next
- Non-parametric tests — the rank-based fallback when spreads and outliers misbehave.
- Multiple testing correction — the false-alarm budget problem ANOVA partially solves.
- Confidence intervals — reporting gaps with uncertainty instead of verdicts.
Researcher — Mathematics and papers.
The decomposition
For $k$ groups with sizes $n_j$, means $\bar{y}_j$, grand mean $\bar{y}$, and $N = \sum_j n_j$ observations total:
$$ \underbrace{\sum_{j=1}^{k}\sum_{i=1}^{n_j} (y_{ij} - \bar{y})^2}{SS{total}} = \underbrace{\sum_{j=1}^{k} n_j (\bar{y}j - \bar{y})^2}{SS_{between}} + \underbrace{\sum_{j=1}^{k}\sum_{i=1}^{n_j} (y_{ij} - \bar{y}j)^2}{SS_{within}} $$
Where:
- $y_{ij}$ — observation $i$ in group $j$.
- $SS_{between}$ — variation of group means around the grand mean.
- $SS_{within}$ — variation of observations around their own group mean.
The test statistic divides each sum by its degrees of freedom:
$$ F = \frac{SS_{between} / (k-1)}{SS_{within} / (N-k)} $$
Under $H_0$ (all group means equal) with normal errors of common variance $\sigma^2$, $F \sim F_{k-1,\,N-k}$. Both numerator and denominator estimate $\sigma^2$ under $H_0$, so $F$ concentrates near 1; under the alternative the numerator inflates by $\frac{1}{k-1}\sum_j n_j(\mu_j - \bar\mu)^2$.
ANOVA as a linear model
One-way ANOVA is the linear model $y_{ij} = \mu + \tau_j + \varepsilon_{ij}$ with $\varepsilon_{ij} \sim \mathcal{N}(0, \sigma^2)$, tested against the intercept-only model. The F-test is the general nested-model comparison; regression F-tests, two-way ANOVA with interactions, and ANCOVA are the same machinery with different design matrices. This is the cleanest way to remember the assumptions: independent errors, homoscedasticity, normality — in decreasing order of importance.
Assumption failure modes
- Heteroscedasticity with unequal $n_j$ biases the level in either direction (liberal when big groups are quiet, conservative when big groups are noisy). Welch's ANOVA (Welch, 1951) and the Alexander–Govern test replace the pooled denominator; simulation evidence favours them as defaults.
- Non-normality: the F level is fairly robust for moderate $N$ by the CLT, but power suffers under heavy tails; the rank-based Kruskal–Wallis test (next lesson) is the standard fallback.
- Dependence (clustered or repeated measurements) is the assumption that actually destroys the test; mixed models or repeated-measures ANOVA are required.
Post-hoc theory
Tukey's HSD (Tukey, 1949) controls the family-wise error rate over all $\binom{k}{2}$ pairwise comparisons using the studentised range distribution $q_{k, N-k}$; each interval is $\bar{y}_i - \bar{y}j \pm \frac{q{1-\alpha}}{\sqrt{2}} \, s_p \sqrt{\frac{1}{n_i} + \frac{1}{n_j}}$. Scheffé's method covers all contrasts (any weighted combination of means) at extra width; Dunnett's compares every group to one control at less width.
Effect size
$\eta^2 = SS_{between}/SS_{total}$ is the variance fraction explained by group membership; $\omega^2$ corrects its upward bias:
$$ \omega^2 = \frac{SS_{between} - (k-1)\,MS_{within}}{SS_{total} + MS_{within}} $$
Fisher introduced the framework in Statistical Methods for Research Workers (1925) and the agricultural design context in The Design of Experiments (1935).
What to learn next
- Non-parametric tests — the rank-based fallback when spreads and outliers misbehave.
- Multiple testing correction — the false-alarm budget problem ANOVA partially solves.
- Confidence intervals — reporting gaps with uncertainty instead of verdicts.