Hypothesis Testing and Inference

Statistical power and sample size

Power is a test's chance of catching a real effect — and most disappointing experiments were doomed before the first data point, by a sample too small to see anything.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Power is the probability that your test detects an effect that is genuinely there.

You are looking for your black cat in a dark garden with a weak torch. The cat is out there. Whether you find it depends on your torch's brightness, how dark the night is, and how big the cat is.

Power is the brightness of your statistical torch. A weak study misses real effects — and then someone announces "no effect found", when the honest sentence is "our torch was too dim".

Why it exists

In the 1960s, Jacob Cohen surveyed published psychology studies and found something embarrassing. The typical study had roughly a coin-flip's chance of detecting the effects it hunted.

Half of all real discoveries were being missed by design, before any data arrived. Power analysis exists so you can compute that chance before spending money — and fix it.

How it works

Three things brighten the torch:

bigger effect     easy to find   (a tiger in the garden)
more data         more light     (a stronger torch)
less noise        clearer night  (less fog)

power = chance of finding what is really there
      → too low?  collect more data, or hunt bigger effects

The standard target is 80% power: four chances in five of catching the effect if it exists. Why not 99%? The data cost climbs steeply — each step toward certainty costs more than the last.

There is a nasty twist worth knowing. An underpowered study that does get lucky tends to exaggerate: among the rare wins, only the flukiest, biggest-looking effects cross the finish line. Small studies do worse than miss truths — the "successes" they publish run inflated.

A real example you have seen

Drug trials publish their power analysis before starting — regulators demand proof that the study could detect the benefit it hunts. A/B testing tools like those inside big e-commerce firms show "minimum detectable effect" calculators. That is power analysis with a friendlier name.

Remember this

  • Power = the chance your test catches a real effect.
  • "Not significant" from a weak study means "could not see", never "does not exist".
  • Decide sample size before the experiment, from the smallest effect you care about.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Outputs verified with numpy 1.26 and scipy 1.14, CPU. Simulated power reproduces with these versions; other versions may shift values by about ±0.01.

Measuring power by simulation

The honest, assumption-light way: simulate the experiment many times with the effect you hope to detect, and count how often the test fires.

power_sim.py
import numpy as np
from scipy import stats

rng = np.random.default_rng(1)

def power(n, effect=3.0, sd=8.0, trials=2000):
    # Fraction of experiments reaching p < 0.05 when the effect is real.
    wins = 0
    for _ in range(trials):
        a = rng.normal(100, sd, n)
        b = rng.normal(100 + effect, sd, n)
        wins += stats.ttest_ind(a, b, equal_var=False).pvalue < 0.05
    return wins / trials

for n in [10, 30, 100, 200]:
    print(f"n = {n:3d} per group -> power {power(n):.2f}")
Output
n =  10 per group -> power 0.12
n =  30 per group -> power 0.29
n = 100 per group -> power 0.75
n = 200 per group -> power 0.96

The scenario: a real improvement of 3 units on a metric that wobbles by 8. With 10 users per group you catch it 12% of the time — an experiment-shaped coin toss, rigged against you.

The walkthrough

The function simulates the world where you are right. Group b genuinely sits 3 units higher. Power is how often the t-test notices, at your chosen sample size.

effect=3.0, sd=8.0 is the pair that matters. Their ratio (3/8 ≈ 0.4) is the effect size in noise units. Halve the effect and you need roughly four times the data for the same power — the square-root law from central-limit-theorem working against you.

Where does effect=3.0 come from? Not from wishing. It is the smallest effect worth acting on — the minimum improvement that would pay for shipping the change. Powering for a fantasy effect produces a study that can only confirm fantasies.

Simulation beats formulas when reality is messy. Formulas exist for clean t-tests. But swap in a Mann–Whitney test, unequal group sizes, or a clipped metric, and the same loop keeps working — change the inside and count.

Common mistakes

Computing power after the experiment from the observed effect. "Post-hoc power" is a one-to-one transform of the p-value dressed as new information. Power analysis is a planning tool; run it before.

Reading a low-powered null as evidence of absence. At n = 30 above, "no significant difference" happens 71% of the time with the effect present. To support "no meaningful effect", run an equivalence test against a pre-declared triviality margin.

Peeking at results daily and stopping at the first p < 0.05. Repeated looks multiply false alarms. Fix the sample size in advance, or use a sequential design built for peeking.

Trusting the effect size of a barely-significant small study. Those estimates are inflated by selection (the "winner's curse"). Plan replications using a deflated effect.

Try it yourself

Find the smallest n giving power ≥ 0.80 for effect=3.0 (try values between 100 and 200). Then halve the effect to 1.5 and find it again. Confirm the roughly-four-times rule with your own numbers.

What to learn next

Researcher — Mathematics and papers.

Formal setup

For a test of $H_0: \theta = \theta_0$ at level $\alpha$, the power function is

$$ \pi(\theta) = \Pr_\theta(\text{reject } H_0) $$

Where:

  • $\pi(\theta_0) \leq \alpha$ — the false-alarm constraint at the null.
  • $\pi(\theta_1) = 1 - \beta$ — power at the alternative of interest $\theta_1$.
  • $\beta$ — the Type II error rate.

The two-sample normal case, closed form

For a two-sided level-$\alpha$ z-approximation with common variance $\sigma^2$, group size $n$, and true difference $\delta$:

$$ 1 - \beta \approx \Phi!\left( \frac{\delta}{\sigma\sqrt{2/n}} - z_{1-\alpha/2} \right) $$

Inverting for the required per-group $n$:

$$ n = \frac{2\,(z_{1-\alpha/2} + z_{1-\beta})^2}{d^2}, \qquad d = \frac{\delta}{\sigma} $$

Where:

  • $\Phi$ — the standard normal CDF; $z_q$ — its $q$-quantile.
  • $d$ — Cohen's standardised effect size.

At $\alpha = 0.05$, power $0.80$: $(1.96 + 0.84)^2 \approx 7.85$, so $n \approx 15.7/d^2$. For the developer example, $d = 3/8$: $n \approx 112$ — matching the simulation's crossing near 100–130. Exact t-based calculations replace the normal with a noncentral t-distribution with noncentrality $\lambda = d\sqrt{n/2}$; the noncentral $\chi^2$ and F play the same role for chi-squared and ANOVA power.

Why underpowered literatures self-poison

Let $\pi$ be typical power and $\phi$ the pre-study odds that a tested effect is real. The positive predictive value of a significant claim is

$$ PPV = \frac{\pi \phi}{\pi \phi + \alpha (1 - \phi)} $$

Low $\pi$ shrinks the numerator while $\alpha$ keeps feeding false positives — the arithmetic core of Ioannidis (2005), Why most published research findings are false. Conditioning on significance also biases $|\hat\delta|$ upward (winner's curse); Gelman and Carlin (2014) quantify this as the Type M (magnitude) error, alongside Type S (sign) errors, and both explode as power drops below ~0.5.

Sequential and adaptive designs

Fixed-$n$ designs waste samples when effects are large and peeking invalidates $\alpha$. Remedies:

  • Group-sequential designs: O'Brien–Fleming or Pocock spending functions allocate the $\alpha$ budget across planned interim looks.
  • Always-valid inference: mixture sequential probability ratio tests and e-processes allow continuous monitoring with anytime-valid error control (Howard et al., 2021; Ramdas et al., 2023).
  • Wald's SPRT (1945) is the optimal simple-vs-simple sequential test, reaching decisions with roughly half the average sample size of the fixed design.

References

  • Cohen, J. (1988), Statistical Power Analysis for the Behavioral Sciences, 2nd ed. — the $d$/$f$/$w$ effect-size conventions.
  • Cohen, J. (1962) — the original power survey of published research.
  • Button et al. (2013), Power failure: why small sample size undermines the reliability of neuroscience, Nature Reviews Neuroscience.

What to learn next