Hypothesis Testing and Inference

Non-parametric tests

Rank-based tests compare who finished ahead of whom instead of averaging raw values — so one wild outlier cannot hijack the verdict.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Non-parametric tests compare groups by who ranks ahead of whom, so extreme values cannot hijack the answer.

Think of a school race. You could report each runner's exact seconds, or you could report finishing positions: first, second, third. If one runner trips and takes ten minutes, the positions of everyone else do not change.

That is the whole idea. Work with positions — ranks — instead of raw numbers, and one disaster stops dominating the story.

Why it exists

The t-test leans on averages, and averages are fragile. One billionaire walks into a tea stall and the "average customer wealth" becomes a joke.

Real data is full of such wreckers: one server timeout among fast responses, one viral post among ordinary ones. Rank-based tests were built for exactly this data — and for data where the assumptions behind the t-test feel shaky.

The name is unhelpful, so decode it once: parametric tests assume the data follows a known shape, like the bell curve. Non-parametric tests skip that assumption.

How it works

raw values:   old server: 110  125   98  3200 ...
              new server:  88   95  102  2900 ...

pool them, rank them:      1st, 2nd, 3rd, ... 20th

question: do one group's ranks pile up at the front?

If the two groups were identical, their ranks would interleave like well-shuffled cards. If one group's ranks cluster near the front, that group tends to be smaller — regardless of how extreme the extremes are.

The price of this safety: with clean, bell-shaped data, rank tests are slightly less sharp than the t-test. You pay a little sensitivity for a lot of robustness.

A real example you have seen

Hospital pain scores — "rate your pain 1 to 10" — are ranks by nature; averaging them is dubious, ranking them is honest. Website response times, house prices, income comparisons: all outlier-riddled, all natural territory for rank tests.

Remember this

  • Rank tests compare positions, not raw values, so outliers lose their veto.
  • Use them for skewed data, ordered scores, and small suspicious samples.
  • The trade: a small loss of sharpness on clean data buys robustness on ugly data.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Outputs verified with numpy 1.26 and scipy 1.14, CPU.

The rank tests versus the t-test, on ugly data

API response times in milliseconds. Each group has one horror story.

rank_tests.py
import numpy as np
from scipy import stats

old = np.array([110, 125, 98, 104, 3200, 115, 108, 122, 96, 130])
new = np.array([88, 95, 102, 79, 91, 85, 99, 93, 2900, 87])

u, p = stats.mannwhitneyu(old, new, alternative="two-sided")
print(f"Mann-Whitney U: U = {u:.0f}, p = {p:.4f}")

t, p_t = stats.ttest_ind(old, new, equal_var=False)
print(f"t-test on the same data: p = {p_t:.4f}")

# Paired version: ten endpoints measured before and after adding a cache.
before = np.array([210, 180, 195, 220, 205, 190, 215, 185, 200, 225])
after  = np.array([170, 165, 180, 175, 168, 182, 160, 172, 178, 169])
w, p_w = stats.wilcoxon(before, after)
print(f"Wilcoxon: W = {w:.0f}, p = {p_w:.4f}")
Output
Mann-Whitney U: U = 87, p = 0.0058
t-test on the same data: p = 0.9081
Wilcoxon: W = 0, p = 0.0020

The walkthrough

The headline is the middle line. The new server wins 87 of the 100 possible old-versus-new comparisons — and the t-test reports p = 0.91, total blindness. Two outliers inflated both means and the noise estimate until the real pattern drowned. The Mann–Whitney U test (two independent groups, ranks) sees through it: p = 0.0058.

U = 87 counts, across all 100 old-new pairings, how often the old value ranks higher. The maximum here is 100; 87 is a strong skew toward "old is slower". U near half the maximum means thorough interleaving — no difference.

The Wilcoxon signed-rank test is the paired sibling. It ranks the differences per endpoint. W = 0 means not one endpoint got slower after the cache — the strongest possible paired result at this sample size.

Three or more groups? stats.kruskal(a, b, c) is the rank-based ANOVA. Same recipe: pool, rank, compare rank piles.

Common mistakes

Describing Mann–Whitney as "a test of medians". It tests whether one group tends to rank higher — a stochastic-dominance question. With very different group shapes, medians can be equal while the test fires. Say "tends to be larger", not "has a larger median".

Using wilcoxon on independent groups. It silently pairs first-with-first, second-with-second. If rows are not the same unit measured twice, the pairing is fiction. Fix: mannwhitneyu for independent groups.

Throwing away a real effect size. Rank tests give verdicts, not magnitudes. Pair them with a robust size estimate: the median difference, or a bootstrap interval on it.

Defaulting to rank tests out of fear. With roughly bell-shaped data and no outliers, the t-test sees smaller effects at the same sample size. Look at your data first; choose second.

Try it yourself

Delete the two outliers (3200 and 2900) and rerun both tests. Predict first: which p-value moves more, and in which direction?

What to learn next

Researcher — Mathematics and papers.

Mann–Whitney U

For samples $X_1, \dots, X_m$ and $Y_1, \dots, Y_n$:

$$ U = \sum_{i=1}^{m} \sum_{j=1}^{n} \mathbf{1}[X_i > Y_j] + \tfrac{1}{2}\,\mathbf{1}[X_i = Y_j] $$

Where:

  • $\mathbf{1}[\cdot]$ — the indicator function, 1 when the condition holds, else 0.
  • $m, n$ — the two sample sizes; $U \in [0, mn]$.

$U/mn$ estimates $\Pr(X > Y) + \tfrac12 \Pr(X = Y)$ — the common-language effect size, directly interpretable and worth reporting. Under $H_0: \Pr(X > Y) = \tfrac12$,

$$ \mathbb{E}[U] = \frac{mn}{2}, \qquad \operatorname{Var}(U) = \frac{mn(m+n+1)}{12} $$

and the standardised $U$ is asymptotically normal; scipy uses the exact distribution for small samples without ties. The test is equivalent to Wilcoxon's rank-sum (Wilcoxon, 1945; Mann and Whitney, 1947, On a test of whether one of two random variables is stochastically larger than the other).

What the test actually assumes

The general null is $F_X = F_Y$ (identical distributions). Rejection means stochastic ordering evidence, not automatically a median difference. Under the location-shift model $F_Y(t) = F_X(t - \Delta)$ it becomes a valid test for $\Delta \neq 0$, and inverting it yields the Hodges–Lehmann estimator — the median of all $mn$ pairwise differences — with a distribution-free confidence interval.

Efficiency

The asymptotic relative efficiency (ARE) of Mann–Whitney against the t-test:

  • Normal data: $3/\pi \approx 0.955$ — you lose under 5% efficiency.
  • Any distribution: never below $0.864$ (Hodges and Lehmann, 1956).
  • Heavy tails (e.g. double exponential: 1.5; $t_3$: about 1.9): the rank test wins, sometimes by a lot.

This asymmetry — bounded small loss, unbounded gain — is the practical argument for rank tests on suspect data.

Wilcoxon signed-rank and Kruskal–Wallis

Signed-rank: rank $|d_i|$ for nonzero paired differences $d_i$, sum the ranks of positive differences to get $W^+$; under symmetry of $d$ about zero, $\mathbb{E}[W^+] = \frac{n(n+1)}{4}$, $\operatorname{Var}(W^+) = \frac{n(n+1)(2n+1)}{24}$. Note the extra assumption: symmetry of the differences, not of the raw data.

Kruskal–Wallis (1952) generalises to $k$ groups: with pooled ranks $R_{ij}$ and group rank-means $\bar{R}_j$,

$$ H = \frac{12}{N(N+1)} \sum_{j=1}^{k} n_j \left(\bar{R}j - \frac{N+1}{2}\right)^2 \;\xrightarrow{d}\; \chi^2{k-1} $$

Post-hoc pairwise comparisons use Dunn's test (1964) with a multiplicity correction.

Complexity and modern context

Ranking dominates: $O(N \log N)$ per test. Ties require midranks and a variance correction (scipy applies both). Modern alternatives occupying the same niche: permutation tests (exact, any statistic), trimmed-mean t-tests (Yuen, 1974), and the Brunner–Munzel test, which drops the equal-variance-under-null requirement that Mann–Whitney retains.

What to learn next