Hypothesis Testing and Inference
T-tests
The t-test decides whether two averages differ by more than their noise — the most-used statistical test in the world, in its three everyday forms.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A t-test decides whether two averages are truly different, or different only by noise.
Two street vendors both claim the hotter chai. You cannot taste every cup they ever pour. You try a few cups from each and compare — but every cup varies a little on its own.
The t-test weighs the gap between the two averages against how much individual cups wobble. A big gap with a small wobble is convincing. A big gap with a huge wobble is not.
Why it exists
In 1908, William Gosset was a brewer at Guinness, checking barley quality with tiny samples. Existing methods assumed huge samples, and his batches were four or five measurements.
He worked out how averages behave in small samples and published under the pen name "Student". That is why you will see the name Student's t-test — same test, historical alias.
How it works
The test squeezes your data into one number, then asks how rare that number would be if both groups were secretly identical.
gap between the two averages
──────────────────────────────── → the t-value
wobble expected from noise alone
small t (near 0) → noise explains everything
large t → the gap outgrew the noiseThe wobble estimate shrinks as you collect more data. So the same gap becomes more convincing with more samples — which matches common sense.
There are three everyday versions. Two separate groups (two-sample). One group against a fixed target (one-sample). The same individuals measured twice, before and after (paired).
A real example you have seen
Every "clinically proven" toothpaste or shampoo ad hides a t-test. One group used the product, one group did not, and someone compared the average outcomes.
App teams do it too. When YouTube trials a new video player with some users, comparing average watch time between groups is t-test territory.
Remember this
- A t-test compares an average gap against the noise level.
- More data shrinks the noise estimate, making real gaps easier to detect.
- Three forms: two groups, one group versus a target, and paired before/after.
What to learn next
- The chi-squared test — the equivalent tool when your data is counts, not averages.
- Non-parametric tests — what to reach for when outliers wreck the mean.
- Statistical power and sample size — how many samples a t-test needs before it can see anything.
Developer — Code and libraries.
Setup
pip install numpy scipyOutputs verified with numpy 1.26 and scipy 1.14, CPU.
All three t-tests on one dataset
Page load times in seconds, eight visits each, old server versus new.
import numpy as np
from scipy import stats
old = np.array([3.1, 2.8, 3.4, 3.0, 3.3, 2.9, 3.2, 3.5])
new = np.array([2.6, 2.9, 2.5, 2.8, 2.7, 3.0, 2.4, 2.6])
t, p = stats.ttest_ind(old, new, equal_var=False) # Welch's t-test
print(f"two-sample: t = {t:.2f}, p = {p:.4f}")
# Same eight pages timed on both servers -> paired test.
t, p = stats.ttest_rel(old, new)
print(f"paired: t = {t:.2f}, p = {p:.4f}")
# One sample against a promise: does the new server load under 3 seconds?
t, p = stats.ttest_1samp(new, 3.0, alternative="less")
print(f"one-sample: t = {t:.2f}, p = {p:.4f}")two-sample: t = 4.11, p = 0.0011 paired: t = 3.14, p = 0.0165 one-sample: t = -4.35, p = 0.0017
The walkthrough
equal_var=False is not optional decoration. It selects Welch's t-test, which drops the assumption that both groups have equal spread. The classic Student version breaks quietly when spreads differ and group sizes are unequal. Welch costs nothing when spreads happen to match, so make it your default.
ttest_rel pairs the measurements up. It tests the eight differences, one per page. Pairing removes the page-to-page variation from the noise, which is why paired designs need far fewer samples for the same conclusion.
The sign of t carries direction. t = -4.35 means the new server's average sits below the 3-second target. The magnitude says by how many noise-units.
alternative="less" asks the one-sided question "is the average under 3.0?". Commit to the direction before running the test.
Common mistakes
Using the two-sample test on paired data. It still runs, but it ignores the pairing and wastes your design's power. If the same items appear in both columns, use ttest_rel.
Trusting the test on heavily skewed, tiny samples. With 8–10 values and a wild outlier, the mean itself becomes a poor summary. Switch to a rank-based test — see nonparametric-tests — or a permutation test.
Reporting p and hiding the size of the gap. p = 0.0011 sounds impressive; readers still need "0.45 seconds faster". Report the difference and its interval alongside — see confidence-intervals.
Running the test repeatedly as data trickles in, stopping when p dips below 0.05. This inflates false alarms badly. Fix the sample size in advance, using statistical-power-and-sample-size.
Try it yourself
Add one slow outlier, 9.9, to new and rerun the two-sample line. Watch a "significant" result evaporate, then explain where the extra noise entered the formula.
What to learn next
- The chi-squared test — the equivalent tool when your data is counts, not averages.
- Non-parametric tests — what to reach for when outliers wreck the mean.
- Statistical power and sample size — how many samples a t-test needs before it can see anything.
Researcher — Mathematics and papers.
The statistic
For samples $x_1, \dots, x_{n_1}$ and $y_1, \dots, y_{n_2}$, Welch's statistic is:
$$ t = \frac{\bar{x} - \bar{y}}{\sqrt{\dfrac{s_1^2}{n_1} + \dfrac{s_2^2}{n_2}}} $$
Where:
- $\bar{x}, \bar{y}$ — the sample means.
- $s_1^2, s_2^2$ — the unbiased sample variances (divisor $n-1$).
- $n_1, n_2$ — the sample sizes.
Under the null of equal population means, with normal populations, $t$ follows approximately a Student's t-distribution with the Welch–Satterthwaite degrees of freedom:
$$ \nu \approx \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{(s_1^2/n_1)^2}{n_1 - 1} + \frac{(s_2^2/n_2)^2}{n_2 - 1}} $$
$\nu$ is generally non-integer; it interpolates between $\min(n_1, n_2) - 1$ and $n_1 + n_2 - 2$. The t-distribution has heavier tails than the normal, precisely accounting for the extra uncertainty of estimating variance from the same small sample.
Why t and not z
If $\sigma$ were known, $(\bar{x} - \mu)/(\sigma/\sqrt{n})$ would be exactly standard normal. Replacing $\sigma$ with $s$ injects estimation noise; Gosset ("Student", 1908, The probable error of a mean) derived the resulting distribution. As $\nu \to \infty$, $t_\nu \to \mathcal{N}(0,1)$; at $n = 5$ the tails differ enough to double small-sample false-alarm rates if you use z.
Robustness, stated honestly
- Normality of the population matters little for moderate $n$: the CLT makes $\bar{x}$ near-normal regardless. For $n \gtrsim 30$ per group and mild skew, the level holds well.
- Heavy tails and strong skew at small $n$ break the level and, worse, the mean stops being the estimand you care about.
- Unequal variances with unequal $n$ destroy the pooled (Student) version — the true Type I rate can hit several times the nominal $\alpha$ — while Welch stays close to nominal. Delacre, Lakens and Leys (2017) argue Welch should be the unconditional default; pre-testing variances with Levene's test and then choosing is known to distort the level.
The paired test is a one-sample t-test on differences $d_i = x_i - y_i$, so it requires only the differences (not the raw scores) to be near-normal.
Effect size
Cohen's $d = (\bar{x} - \bar{y}) / s_p$ (with $s_p$ the pooled standard deviation) expresses the gap in noise units, independent of $n$. Hedges' $g$ applies a small-sample bias correction factor $\left(1 - \frac{3}{4\nu - 1}\right)$. Report one of them with the p-value; $t$ alone conflates effect and sample size since $t \approx d\sqrt{n/2}$.
Cost and variants
Computation is $O(n)$; nothing here strains hardware. Variants:
- Trimmed-means / Yuen's t-test (Yuen, 1974) for heavy tails.
- Mann–Whitney U as the rank-based alternative (next lessons).
- Bayesian t-test via Bayes factors (Rouder et al., 2009, Bayesian t tests for accepting and rejecting the null hypothesis).
What to learn next
- The chi-squared test — the equivalent tool when your data is counts, not averages.
- Non-parametric tests — what to reach for when outliers wreck the mean.
- Statistical power and sample size — how many samples a t-test needs before it can see anything.