Hypothesis Testing and Inference
The chi-squared test
The chi-squared test works on counts — clicks, votes, defects — and asks whether the numbers in your table drifted too far from what chance would deal.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The chi-squared test checks whether counted categories — clicks, votes, defects — split the way chance would split them.
Deal a pack of cards to four players. Everyone expects 13 cards each. If one player gets 20, you would re-check the shuffle, because the deal drifted too far from fair.
The chi-squared test measures exactly that drift. It compares the counts you got against the counts a fair split would give, across every category at once.
Why it exists
T-tests need measurements — seconds, rupees, kilograms. But much real data is counted, not measured. Clicked or ignored. Passed or failed. Voted A, B or C.
Karl Pearson built this test in 1900 so counted data could be tested too. It became one of the most used tools in science, from genetics to market research.
How it works
You build a table of counts. The test then computes what the table would look like if the two things in it had no connection — and measures the gap.
clicked ignored
blue button 34 166 "If colour changed nothing,
green button 51 149 both rows should click at
the same rate. Do they?"
actual counts vs no-connection counts → one drift scoreA small drift score means chance covers the differences. A large one means the rows genuinely behave differently — colour and clicking are connected.
The same trick, run against any expected split, also checks fairness. Sixty dice rolls should give each face about ten appearances. Big drift from that says loaded die.
A real example you have seen
Election exit polls versus actual results get compared this way. So do hospital studies asking "did more vaccinated people avoid infection?" — vaccinated or not, infected or not, a two-by-two table of people.
Every "which version got more clicks" experiment inside apps like Flipkart is a table of counts waiting for this test.
Remember this
- Chi-squared is for counts in categories, not for averages of measurements.
- It compares your table against the table chance would produce.
- One score summarises the drift across every cell at once.
What to learn next
- ANOVA — comparing averages across three or more groups.
- Permutation tests — the exact, assumption-free version of this idea.
- Multiple testing correction — what happens when you run many tables at once.
Developer — Code and libraries.
Setup
pip install numpy scipyOutputs verified with numpy 1.26 and scipy 1.14, CPU.
Independence and goodness of fit
Two hundred visitors saw a blue button; two hundred saw green. Did colour change clicking?
import numpy as np
from scipy import stats
# clicked ignored
table = np.array([[34, 166], # blue button
[51, 149]]) # green button
res = stats.chi2_contingency(table)
print(f"chi2 = {res.statistic:.2f}, p = {res.pvalue:.4f}")
print("expected counts if colour made no difference:")
print(res.expected_freq)
# Goodness of fit: one die, rolled 60 times, counts per face.
rolls = np.array([5, 8, 9, 8, 10, 20])
res2 = stats.chisquare(rolls) # default expectation: equal counts
print(f"die: chi2 = {res2.statistic:.2f}, p = {res2.pvalue:.4f}")chi2 = 3.82, p = 0.0505 expected counts if colour made no difference: [[ 42.5 157.5] [ 42.5 157.5]] die: chi2 = 13.40, p = 0.0199
The walkthrough
res.expected_freq is the null hypothesis, printed. Overall, 85 of 400 visitors clicked — 21.25%. Applied to each row of 200, chance predicts 42.5 clicks per colour. The statistic totals the squared gaps between actual and expected, scaled by expected.
That p-value of 0.0505 is a teaching gift. It sits a hair above the usual 0.05 line. The honest report is "suggestive, not conclusive — collect more data", not a triumphant yes and not a flat no. Thresholds are conventions; data near the line stays ambiguous.
chi2_contingency tests connection; chisquare tests fit. The first invents its own expected table from the margins. The second needs you to know the expectation up front — here, ten per face. The die's six landed 20 times instead of ten, and p = 0.0199 calls that suspicious.
Degrees of freedom come for free. For a table, it is (rows − 1) × (columns − 1); the p-value already accounts for table size.
Common mistakes
Feeding in percentages or rates. The test needs raw counts. Convert 17% of 200 back to 34 before building the table — the sample size is the evidence.
Using it with tiny expected counts. When any expected cell drops below about 5, the approximation frays. Fix: use stats.fisher_exact(table) for 2×2 tables — exact at any size.
Concluding cause from connection. A significant table says the variables move together, not that one drives the other. Ice-cream sales and drownings rise together every summer; the connection is the season, not the cone. Randomised assignment, not the test, is what buys causal claims.
Double-counting people. Each person must land in exactly one cell. If the same visitor saw both colours, the independence assumption is broken; restructure the experiment instead.
Try it yourself
Multiply the whole table by 4 (same rates, 1,600 visitors) and rerun. Watch the same percentages become decisive. Then work out why sample size alone moved the p-value.
What to learn next
- ANOVA — comparing averages across three or more groups.
- Permutation tests — the exact, assumption-free version of this idea.
- Multiple testing correction — what happens when you run many tables at once.
Researcher — Mathematics and papers.
The statistic
For observed counts $O_i$ and expected counts $E_i$ over $k$ cells:
$$ \chi^2 = \sum_{i=1}^{k} \frac{(O_i - E_i)^2}{E_i} $$
Where:
- $O_i$ — the observed count in cell $i$.
- $E_i$ — the count expected under $H_0$.
- $k$ — the number of cells.
Under $H_0$, $\chi^2$ converges in distribution to a chi-squared distribution with $\nu$ degrees of freedom. For an $r \times c$ contingency table with margins estimated from the data, $\nu = (r-1)(c-1)$; for goodness of fit against a fully specified distribution, $\nu = k - 1$ (subtract one more per parameter estimated from the data).
The result is asymptotic — Pearson (1900, On the criterion that a given system of deviations…) — and each term is a squared standardised residual: divide by $E_i$ because a Poisson-like count with mean $E_i$ has variance $\approx E_i$.
Derivation sketch
Cell counts under $H_0$ are multinomial. The vector $(O_i - E_i)/\sqrt{E_i}$ is asymptotically multivariate normal with a rank-$\nu$ covariance projection; the squared norm of such a Gaussian vector is $\chi^2_\nu$. This also frames the test as a score test; the likelihood-ratio counterpart is the G-test:
$$ G = 2 \sum_i O_i \log\frac{O_i}{E_i} $$
$G$ and Pearson's $\chi^2$ agree to second order; $G$ is additive over nested partitions, which some fields prefer.
Small samples and corrections
- Rule of thumb: all $E_i \geq 5$ (Cochran, 1954). Below that, the asymptotic level degrades.
- Fisher's exact test (Fisher, 1935) conditions on both margins and computes hypergeometric probabilities — exact for 2×2, extendable at combinatorial cost.
- Yates' continuity correction subtracts 0.5 from $|O_i - E_i|$ in 2×2 tables; it overcorrects and is now widely discouraged for moderate samples (scipy applies it by default for 2×2 — pass
correction=Falseto disable). - Modern alternative: simulate the null by permutation for any table shape and size.
Effect size
$\chi^2$ grows linearly with $n$ at fixed proportions, so it is not an effect size. Report Cramér's V:
$$ V = \sqrt{\frac{\chi^2}{n \, \min(r-1, c-1)}} $$
with $n$ the grand total; $V \in [0, 1]$ measures association strength independent of sample size.
Complexity and relatives
Computation is $O(rc)$. Related machinery: the chi-squared statistic underlies feature selection for categorical inputs (sklearn.feature_selection.chi2, scikit-learn 1.7), the McNemar test for paired binary outcomes (compare two classifiers on the same test set), and the likelihood-ratio tests of the previous lesson's researcher block.
What to learn next
- ANOVA — comparing averages across three or more groups.
- Permutation tests — the exact, assumption-free version of this idea.
- Multiple testing correction — what happens when you run many tables at once.