Hypothesis Testing and Inference

Q-Q plots and normality checks

A Q-Q plot lines your data up against the bell curve's predictions, point by point, and shows exactly where and how the data misbehaves.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A Q-Q plot checks whether your data follows the bell curve, by lining both up side by side.

Tailors check a shirt against a paper pattern by laying one on top of the other. Wherever fabric bulges past the pattern, the mismatch is visible — and where it bulges tells you what went wrong: sleeves, collar, or length.

A Q-Q plot lays your data over the bell curve's pattern the same way. Points on the line: good fit. Points peeling away: a mismatch, located.

Why it exists

Several tools in this section assume data is roughly bell-shaped. The t-test leans on it for small samples. Confidence-interval recipes assume it. Get it badly wrong and their answers mislead quietly.

So you need a habit: check the shape before trusting the tool. A histogram gives a rough look, but its shape shifts with how you slice the bins. The Q-Q plot is sharper, and it puts its finger on where the mismatch lives.

How it works

The name means quantile-quantile. A quantile is a checkpoint: the value below which a given share of the data falls — the halfway point, the 90% point, and so on.

Line up your data's checkpoints against the bell curve's checkpoints, pair by pair:

bell curve says      your data says
   1%  point  -2.3      -5.2     ← too extreme! heavy lower tail
  25%  point  -0.7      -0.8     ← fine
  50%  point   0.0      -0.1     ← fine
  99%  point  +2.3      +4.6     ← too extreme! heavy upper tail

plotted:   middle points hug the line,
           tail points bend away = "heavy tails"

The bend's shape is a diagnosis. Both ends bending outward: extreme values are too common — heavy tails. Only one end bending: lopsided data. An S-shape: squeezed middle.

A real example you have seen

Stock returns are the textbook case. Most days the market moves modestly, matching the bell — then a crash lands far beyond what the bell allows. On a Q-Q plot, decades of returns hug the line in the middle and peel off violently at both ends. Risk models that skipped this check helped make 2008 worse.

Remember this

  • A Q-Q plot pairs your data's checkpoints with the bell curve's, one by one.
  • Points on the line = good fit; the shape of the bend names the problem.
  • Check the shape before trusting shape-assuming tools.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy matplotlib

Outputs verified with numpy 1.26 and scipy 1.14, CPU. Simulated digits reproduce with these versions.

Numbers first, then the picture

Two datasets: one genuinely bell-shaped, one with heavy tails (Student-t with 3 degrees of freedom — a stand-in for stock returns).

qq_check.py
import numpy as np
from scipy import stats

rng = np.random.default_rng(2)
normal = rng.normal(size=200)
heavy = rng.standard_t(df=3, size=200)

qs = [0.01, 0.25, 0.50, 0.75, 0.99]
theory = stats.norm.ppf(qs)
print("quantile   theory   normal-data   heavy-tailed")
for q, t in zip(qs, theory):
    print(f"{q:8.2f}   {t:6.2f}   {np.quantile(normal, q):11.2f}   {np.quantile(heavy, q):12.2f}")

print(f"Shapiro-Wilk, normal data: p = {stats.shapiro(normal).pvalue:.3f}")
print(f"Shapiro-Wilk, heavy tails: p = {stats.shapiro(heavy).pvalue:.1e}")
Output
quantile   theory   normal-data   heavy-tailed
    0.01    -2.33         -2.12          -4.69
    0.25    -0.67         -0.72          -0.96
    0.50     0.00         -0.03          -0.16
    0.75     0.67          0.67           0.52
    0.99     2.33          2.08           4.61
Shapiro-Wilk, normal data: p = 0.899
Shapiro-Wilk, heavy tails: p = 1.3e-13

The actual plot is two lines. They print nothing to the console — the result is a PNG file on disk:

python
import matplotlib.pyplot as plt
stats.probplot(heavy, dist="norm", plot=plt)
plt.savefig("qq.png")

Open qq.png (640×480, written into your working directory): the middle hugs the red line; both ends peel away steeply — the heavy-tail signature drawn for you.

The walkthrough

Read the table like a Q-Q plot in text form. The middle checkpoints (25%, 50%, 75%) of both datasets sit near theory. The tails split the story: the normal data's 1% point is -2.12 against a predicted -2.33 — fine. The heavy-tailed data's is -4.69 — more than double the prediction. Same at the top: 4.61 versus 2.33.

stats.norm.ppf is the bell curve's checkpoint calculator — the inverse of the CDF. ppf(0.01) answers "below what value does 1% of a standard bell fall?".

The Shapiro–Wilk test turns the picture into a p-value. Null hypothesis: the data is normal. p = 0.899 gives no reason to doubt the first dataset; p = 1.3e-13 demolishes the second. It is the most powerful of the common normality tests at small-to-medium sizes.

But the plot outranks the test — in both directions. At n = 20, real heavy tails often survive the test (low power). At n = 20,000, a harmless rounding artefact fails it (overwhelming power against irrelevant flaws). The plot shows what kind and how much; the test only says "detectably different".

Common mistakes

Testing normality of the raw data when the assumption is about something else. The t-test cares about near-normal sample means — which the central limit theorem delivers for moderate samples anyway. For regression, the assumption is about residuals, never the raw target. Check the right thing.

Dropping "outliers" to pass the check. Heavy tails are the distribution, not contamination. Deleting real tail points forges the certificate. Model them instead: rank tests, robust methods, or a t-distribution likelihood.

Reacting to a failed test at huge n. With enough data, the test flags cosmetic deviations while the CLT has already rescued your t-test. Judge the plot's bend size, not the p-value's zeros.

Checking normality with a histogram. Bin width changes the verdict; tails hide in short bars. The Q-Q plot needs no binning and puts the tails where you cannot miss them.

Try it yourself

Generate rng.exponential(size=200) and print the same table. Predict the pattern first: which checkpoints will overshoot, which will undershoot, and why does only one side misbehave?

What to learn next

Researcher — Mathematics and papers.

Construction

Order the sample $x_{(1)} \leq \dots \leq x_{(n)}$ and plot $x_{(i)}$ against theoretical quantiles:

$$ q_i = F^{-1}!\left(\frac{i - 0.5}{n}\right) $$

Where:

  • $F^{-1}$ — the quantile function of the reference distribution.
  • $\frac{i - 0.5}{n}$ — a plotting position; alternatives include $\frac{i - 3/8}{n + 1/4}$ (Blom, near-optimal for the normal) and Filliben's medians, which scipy's probplot uses.

If the data is $F$ up to location and scale, the points fall near a straight line with intercept $\mu$ and slope $\sigma$ — which is why a normal Q-Q plot needs no parameter estimates to be read.

Sampling variability is largest in the tails: $\operatorname{Var}(x_{(i)}) \approx \frac{p_i(1-p_i)}{n f(q_i)^2}$ with $f$ the reference density — tail order statistics wobble enormously, so pointwise confidence bands (or the tighter Aldor-Noiman et al. TS bands, 2013) belong on any Q-Q plot used for decisions.

Shapiro–Wilk

The statistic measures agreement between order statistics and their normal expectations:

$$ W = \frac{\left(\sum_{i=1}^{n} a_i x_{(i)}\right)^2}{\sum_{i=1}^{n} (x_i - \bar{x})^2} $$

Where $a_i$ are constants from the expected values $m_i$ and covariance matrix $V$ of standard normal order statistics: $a = m^\top V^{-1} / |V^{-1} m|$. $W \in (0, 1]$ is, in effect, the squared correlation of the Q-Q plot with its best-fit line; small $W$ means curvature. Original paper: Shapiro and Wilk (1965), An analysis of variance test for normality; Royston (1992) extended the p-value approximation to large $n$ (scipy follows Royston, accurate to $n = 5000$).

Comparative power (Razali and Wah, 2011): Shapiro–Wilk ≥ Anderson–Darling > Lilliefors > Kolmogorov–Smirnov across most alternatives. Anderson–Darling weights the tail discrepancies explicitly:

$$ A^2 = -n - \frac{1}{n}\sum_{i=1}^{n} (2i - 1)\left[\ln F(x_{(i)}) + \ln!\left(1 - F(x_{(n+1-i)})\right)\right] $$

Kolmogorov–Smirnov against a fitted normal requires the Lilliefors correction — using the plain KS null distribution after estimating $\mu, \sigma$ from the same data is a standard, silent error.

Moment-based diagnostics

Sample skewness $g_1 = m_3 / m_2^{3/2}$ and excess kurtosis $g_2 = m_4/m_2^2 - 3$ (with $m_k$ the $k$-th central sample moment) feed D'Agostino–Pearson's $K^2$ omnibus test (scipy.stats.normaltest). The $t_3$ distribution above has infinite fourth moment — sample kurtosis then grows without bound as $n$ increases, a useful red flag on streaming data.

Decision-relevance, not ritual

The consequential question is never "is the data exactly normal" (it is not) but "does the deviation break my next step". Guidance: t-procedure levels are robust to moderate non-normality at $n \gtrsim 30$ per group (CLT), but power degrades under heavy tails — switch to rank or permutation methods. Variance-based statistics (F-tests for variances, Levene's original form) are fragile to kurtosis at any $n$. Prediction intervals and tail-risk estimates depend on the tails precisely where the approximation is worst — model tails explicitly (extreme value theory) rather than certifying normality.

What to learn next

What to learn next

These follow on from what you just read.

  • Resampling, Likelihood and Bayes

    The bootstrap

    The bootstrap measures the uncertainty of any statistic — median, correlation, anything — by resampling your own data instead of trusting a formula.

  • Resampling, Likelihood and Bayes

    Permutation tests

    A permutation test builds the null distribution by shuffling group labels — an exact test for any statistic, with almost no assumptions to break.

  • Resampling, Likelihood and Bayes

    Monte Carlo simulation

    When the formula is too hard, roll the dice instead — simulate the process thousands of times and count what happens.