Hypothesis Testing and Inference
Non-parametric tests
Rank-based tests compare who finished ahead of whom instead of averaging raw values — so one wild outlier cannot hijack the verdict.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Non-parametric tests compare groups by who ranks ahead of whom, so extreme values cannot hijack the answer.
Think of a school race. You could report each runner's exact seconds, or you could report finishing positions: first, second, third. If one runner trips and takes ten minutes, the positions of everyone else do not change.
That is the whole idea. Work with positions — ranks — instead of raw numbers, and one disaster stops dominating the story.
Why it exists
The t-test leans on averages, and averages are fragile. One billionaire walks into a tea stall and the "average customer wealth" becomes a joke.
Real data is full of such wreckers: one server timeout among fast responses, one viral post among ordinary ones. Rank-based tests were built for exactly this data — and for data where the assumptions behind the t-test feel shaky.
The name is unhelpful, so decode it once: parametric tests assume the data follows a known shape, like the bell curve. Non-parametric tests skip that assumption.
How it works
raw values: old server: 110 125 98 3200 ...
new server: 88 95 102 2900 ...
pool them, rank them: 1st, 2nd, 3rd, ... 20th
question: do one group's ranks pile up at the front?If the two groups were identical, their ranks would interleave like well-shuffled cards. If one group's ranks cluster near the front, that group tends to be smaller — regardless of how extreme the extremes are.
The price of this safety: with clean, bell-shaped data, rank tests are slightly less sharp than the t-test. You pay a little sensitivity for a lot of robustness.
A real example you have seen
Hospital pain scores — "rate your pain 1 to 10" — are ranks by nature; averaging them is dubious, ranking them is honest. Website response times, house prices, income comparisons: all outlier-riddled, all natural territory for rank tests.
Remember this
- Rank tests compare positions, not raw values, so outliers lose their veto.
- Use them for skewed data, ordered scores, and small suspicious samples.
- The trade: a small loss of sharpness on clean data buys robustness on ugly data.
What to learn next
- Pearson, Spearman and Kendall — the same rank idea applied to correlation.
- Permutation tests — build an exact test for any statistic by shuffling.
- The bootstrap — robust effect sizes to pair with rank-test verdicts.
Developer — Code and libraries.
Setup
pip install numpy scipyOutputs verified with numpy 1.26 and scipy 1.14, CPU.
The rank tests versus the t-test, on ugly data
API response times in milliseconds. Each group has one horror story.
import numpy as np
from scipy import stats
old = np.array([110, 125, 98, 104, 3200, 115, 108, 122, 96, 130])
new = np.array([88, 95, 102, 79, 91, 85, 99, 93, 2900, 87])
u, p = stats.mannwhitneyu(old, new, alternative="two-sided")
print(f"Mann-Whitney U: U = {u:.0f}, p = {p:.4f}")
t, p_t = stats.ttest_ind(old, new, equal_var=False)
print(f"t-test on the same data: p = {p_t:.4f}")
# Paired version: ten endpoints measured before and after adding a cache.
before = np.array([210, 180, 195, 220, 205, 190, 215, 185, 200, 225])
after = np.array([170, 165, 180, 175, 168, 182, 160, 172, 178, 169])
w, p_w = stats.wilcoxon(before, after)
print(f"Wilcoxon: W = {w:.0f}, p = {p_w:.4f}")Mann-Whitney U: U = 87, p = 0.0058 t-test on the same data: p = 0.9081 Wilcoxon: W = 0, p = 0.0020
The walkthrough
The headline is the middle line. The new server wins 87 of the 100 possible old-versus-new comparisons — and the t-test reports p = 0.91, total blindness. Two outliers inflated both means and the noise estimate until the real pattern drowned. The Mann–Whitney U test (two independent groups, ranks) sees through it: p = 0.0058.
U = 87 counts, across all 100 old-new pairings, how often the old value ranks higher. The maximum here is 100; 87 is a strong skew toward "old is slower". U near half the maximum means thorough interleaving — no difference.
The Wilcoxon signed-rank test is the paired sibling. It ranks the differences per endpoint. W = 0 means not one endpoint got slower after the cache — the strongest possible paired result at this sample size.
Three or more groups? stats.kruskal(a, b, c) is the rank-based ANOVA. Same recipe: pool, rank, compare rank piles.
Common mistakes
Describing Mann–Whitney as "a test of medians". It tests whether one group tends to rank higher — a stochastic-dominance question. With very different group shapes, medians can be equal while the test fires. Say "tends to be larger", not "has a larger median".
Using wilcoxon on independent groups. It silently pairs first-with-first, second-with-second. If rows are not the same unit measured twice, the pairing is fiction. Fix: mannwhitneyu for independent groups.
Throwing away a real effect size. Rank tests give verdicts, not magnitudes. Pair them with a robust size estimate: the median difference, or a bootstrap interval on it.
Defaulting to rank tests out of fear. With roughly bell-shaped data and no outliers, the t-test sees smaller effects at the same sample size. Look at your data first; choose second.
Try it yourself
Delete the two outliers (3200 and 2900) and rerun both tests. Predict first: which p-value moves more, and in which direction?
What to learn next
- Pearson, Spearman and Kendall — the same rank idea applied to correlation.
- Permutation tests — build an exact test for any statistic by shuffling.
- The bootstrap — robust effect sizes to pair with rank-test verdicts.
Researcher — Mathematics and papers.
Mann–Whitney U
For samples $X_1, \dots, X_m$ and $Y_1, \dots, Y_n$:
$$ U = \sum_{i=1}^{m} \sum_{j=1}^{n} \mathbf{1}[X_i > Y_j] + \tfrac{1}{2}\,\mathbf{1}[X_i = Y_j] $$
Where:
- $\mathbf{1}[\cdot]$ — the indicator function, 1 when the condition holds, else 0.
- $m, n$ — the two sample sizes; $U \in [0, mn]$.
$U/mn$ estimates $\Pr(X > Y) + \tfrac12 \Pr(X = Y)$ — the common-language effect size, directly interpretable and worth reporting. Under $H_0: \Pr(X > Y) = \tfrac12$,
$$ \mathbb{E}[U] = \frac{mn}{2}, \qquad \operatorname{Var}(U) = \frac{mn(m+n+1)}{12} $$
and the standardised $U$ is asymptotically normal; scipy uses the exact distribution for small samples without ties. The test is equivalent to Wilcoxon's rank-sum (Wilcoxon, 1945; Mann and Whitney, 1947, On a test of whether one of two random variables is stochastically larger than the other).
What the test actually assumes
The general null is $F_X = F_Y$ (identical distributions). Rejection means stochastic ordering evidence, not automatically a median difference. Under the location-shift model $F_Y(t) = F_X(t - \Delta)$ it becomes a valid test for $\Delta \neq 0$, and inverting it yields the Hodges–Lehmann estimator — the median of all $mn$ pairwise differences — with a distribution-free confidence interval.
Efficiency
The asymptotic relative efficiency (ARE) of Mann–Whitney against the t-test:
- Normal data: $3/\pi \approx 0.955$ — you lose under 5% efficiency.
- Any distribution: never below $0.864$ (Hodges and Lehmann, 1956).
- Heavy tails (e.g. double exponential: 1.5; $t_3$: about 1.9): the rank test wins, sometimes by a lot.
This asymmetry — bounded small loss, unbounded gain — is the practical argument for rank tests on suspect data.
Wilcoxon signed-rank and Kruskal–Wallis
Signed-rank: rank $|d_i|$ for nonzero paired differences $d_i$, sum the ranks of positive differences to get $W^+$; under symmetry of $d$ about zero, $\mathbb{E}[W^+] = \frac{n(n+1)}{4}$, $\operatorname{Var}(W^+) = \frac{n(n+1)(2n+1)}{24}$. Note the extra assumption: symmetry of the differences, not of the raw data.
Kruskal–Wallis (1952) generalises to $k$ groups: with pooled ranks $R_{ij}$ and group rank-means $\bar{R}_j$,
$$ H = \frac{12}{N(N+1)} \sum_{j=1}^{k} n_j \left(\bar{R}j - \frac{N+1}{2}\right)^2 \;\xrightarrow{d}\; \chi^2{k-1} $$
Post-hoc pairwise comparisons use Dunn's test (1964) with a multiplicity correction.
Complexity and modern context
Ranking dominates: $O(N \log N)$ per test. Ties require midranks and a variance correction (scipy applies both). Modern alternatives occupying the same niche: permutation tests (exact, any statistic), trimmed-mean t-tests (Yuen, 1974), and the Brunner–Munzel test, which drops the equal-variance-under-null requirement that Mann–Whitney retains.
What to learn next
- Pearson, Spearman and Kendall — the same rank idea applied to correlation.
- Permutation tests — build an exact test for any statistic by shuffling.
- The bootstrap — robust effect sizes to pair with rank-test verdicts.