Hypothesis Testing and Inference

T-tests

The t-test decides whether two averages differ by more than their noise — the most-used statistical test in the world, in its three everyday forms.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A t-test decides whether two averages are truly different, or different only by noise.

Two street vendors both claim the hotter chai. You cannot taste every cup they ever pour. You try a few cups from each and compare — but every cup varies a little on its own.

The t-test weighs the gap between the two averages against how much individual cups wobble. A big gap with a small wobble is convincing. A big gap with a huge wobble is not.

Why it exists

In 1908, William Gosset was a brewer at Guinness, checking barley quality with tiny samples. Existing methods assumed huge samples, and his batches were four or five measurements.

He worked out how averages behave in small samples and published under the pen name "Student". That is why you will see the name Student's t-test — same test, historical alias.

How it works

The test squeezes your data into one number, then asks how rare that number would be if both groups were secretly identical.

gap between the two averages
────────────────────────────────   →  the t-value
wobble expected from noise alone

small t (near 0)  →  noise explains everything
large t           →  the gap outgrew the noise

The wobble estimate shrinks as you collect more data. So the same gap becomes more convincing with more samples — which matches common sense.

There are three everyday versions. Two separate groups (two-sample). One group against a fixed target (one-sample). The same individuals measured twice, before and after (paired).

A real example you have seen

Every "clinically proven" toothpaste or shampoo ad hides a t-test. One group used the product, one group did not, and someone compared the average outcomes.

App teams do it too. When YouTube trials a new video player with some users, comparing average watch time between groups is t-test territory.

Remember this

  • A t-test compares an average gap against the noise level.
  • More data shrinks the noise estimate, making real gaps easier to detect.
  • Three forms: two groups, one group versus a target, and paired before/after.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Outputs verified with numpy 1.26 and scipy 1.14, CPU.

All three t-tests on one dataset

Page load times in seconds, eight visits each, old server versus new.

three_t_tests.py
import numpy as np
from scipy import stats

old = np.array([3.1, 2.8, 3.4, 3.0, 3.3, 2.9, 3.2, 3.5])
new = np.array([2.6, 2.9, 2.5, 2.8, 2.7, 3.0, 2.4, 2.6])

t, p = stats.ttest_ind(old, new, equal_var=False)   # Welch's t-test
print(f"two-sample: t = {t:.2f}, p = {p:.4f}")

# Same eight pages timed on both servers -> paired test.
t, p = stats.ttest_rel(old, new)
print(f"paired:     t = {t:.2f}, p = {p:.4f}")

# One sample against a promise: does the new server load under 3 seconds?
t, p = stats.ttest_1samp(new, 3.0, alternative="less")
print(f"one-sample: t = {t:.2f}, p = {p:.4f}")
Output
two-sample: t = 4.11, p = 0.0011
paired:     t = 3.14, p = 0.0165
one-sample: t = -4.35, p = 0.0017

The walkthrough

equal_var=False is not optional decoration. It selects Welch's t-test, which drops the assumption that both groups have equal spread. The classic Student version breaks quietly when spreads differ and group sizes are unequal. Welch costs nothing when spreads happen to match, so make it your default.

ttest_rel pairs the measurements up. It tests the eight differences, one per page. Pairing removes the page-to-page variation from the noise, which is why paired designs need far fewer samples for the same conclusion.

The sign of t carries direction. t = -4.35 means the new server's average sits below the 3-second target. The magnitude says by how many noise-units.

alternative="less" asks the one-sided question "is the average under 3.0?". Commit to the direction before running the test.

Common mistakes

Using the two-sample test on paired data. It still runs, but it ignores the pairing and wastes your design's power. If the same items appear in both columns, use ttest_rel.

Trusting the test on heavily skewed, tiny samples. With 8–10 values and a wild outlier, the mean itself becomes a poor summary. Switch to a rank-based test — see nonparametric-tests — or a permutation test.

Reporting p and hiding the size of the gap. p = 0.0011 sounds impressive; readers still need "0.45 seconds faster". Report the difference and its interval alongside — see confidence-intervals.

Running the test repeatedly as data trickles in, stopping when p dips below 0.05. This inflates false alarms badly. Fix the sample size in advance, using statistical-power-and-sample-size.

Try it yourself

Add one slow outlier, 9.9, to new and rerun the two-sample line. Watch a "significant" result evaporate, then explain where the extra noise entered the formula.

What to learn next

Researcher — Mathematics and papers.

The statistic

For samples $x_1, \dots, x_{n_1}$ and $y_1, \dots, y_{n_2}$, Welch's statistic is:

$$ t = \frac{\bar{x} - \bar{y}}{\sqrt{\dfrac{s_1^2}{n_1} + \dfrac{s_2^2}{n_2}}} $$

Where:

  • $\bar{x}, \bar{y}$ — the sample means.
  • $s_1^2, s_2^2$ — the unbiased sample variances (divisor $n-1$).
  • $n_1, n_2$ — the sample sizes.

Under the null of equal population means, with normal populations, $t$ follows approximately a Student's t-distribution with the Welch–Satterthwaite degrees of freedom:

$$ \nu \approx \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{(s_1^2/n_1)^2}{n_1 - 1} + \frac{(s_2^2/n_2)^2}{n_2 - 1}} $$

$\nu$ is generally non-integer; it interpolates between $\min(n_1, n_2) - 1$ and $n_1 + n_2 - 2$. The t-distribution has heavier tails than the normal, precisely accounting for the extra uncertainty of estimating variance from the same small sample.

Why t and not z

If $\sigma$ were known, $(\bar{x} - \mu)/(\sigma/\sqrt{n})$ would be exactly standard normal. Replacing $\sigma$ with $s$ injects estimation noise; Gosset ("Student", 1908, The probable error of a mean) derived the resulting distribution. As $\nu \to \infty$, $t_\nu \to \mathcal{N}(0,1)$; at $n = 5$ the tails differ enough to double small-sample false-alarm rates if you use z.

Robustness, stated honestly

  • Normality of the population matters little for moderate $n$: the CLT makes $\bar{x}$ near-normal regardless. For $n \gtrsim 30$ per group and mild skew, the level holds well.
  • Heavy tails and strong skew at small $n$ break the level and, worse, the mean stops being the estimand you care about.
  • Unequal variances with unequal $n$ destroy the pooled (Student) version — the true Type I rate can hit several times the nominal $\alpha$ — while Welch stays close to nominal. Delacre, Lakens and Leys (2017) argue Welch should be the unconditional default; pre-testing variances with Levene's test and then choosing is known to distort the level.

The paired test is a one-sample t-test on differences $d_i = x_i - y_i$, so it requires only the differences (not the raw scores) to be near-normal.

Effect size

Cohen's $d = (\bar{x} - \bar{y}) / s_p$ (with $s_p$ the pooled standard deviation) expresses the gap in noise units, independent of $n$. Hedges' $g$ applies a small-sample bias correction factor $\left(1 - \frac{3}{4\nu - 1}\right)$. Report one of them with the p-value; $t$ alone conflates effect and sample size since $t \approx d\sqrt{n/2}$.

Cost and variants

Computation is $O(n)$; nothing here strains hardware. Variants:

  • Trimmed-means / Yuen's t-test (Yuen, 1974) for heavy tails.
  • Mann–Whitney U as the rank-based alternative (next lessons).
  • Bayesian t-test via Bayes factors (Rouder et al., 2009, Bayesian t tests for accepting and rejecting the null hypothesis).

What to learn next

What to learn next

These follow on from what you just read.

  • Hypothesis Testing and Inference

    The chi-squared test

    The chi-squared test works on counts — clicks, votes, defects — and asks whether the numbers in your table drifted too far from what chance would deal.

  • Hypothesis Testing and Inference

    ANOVA

    ANOVA compares three or more group averages in one shot, by asking whether the groups differ more between themselves than within themselves.

  • Hypothesis Testing and Inference

    Non-parametric tests

    Rank-based tests compare who finished ahead of whom instead of averaging raw values — so one wild outlier cannot hijack the verdict.