Statistics questions, with answers
Twenty real statistics interview questions with worked answers — p-values, A/B tests, the disease-test trap, bootstrap, Simpson's paradox and the deep follow-ups.
- 17 min read
- 3 reading levels
- Published
Read these first
On this page 9
- Why this round exists
- The instinct, in one picture
- Q1. What does a p-value actually mean?
- Q2. What is a confidence interval, in words?
- Q3. Type I versus Type II errors — which is worse?
- Q4. Correlation is not causation — give me a real example
- Q5. Why do bigger samples make us more confident?
- Remember this
- What to learn next
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The statistics round tests one thing: whether numbers can fool you.
You have run this test yourself at a cricket toss. A captain wins the toss eight times in a row, and the other side starts muttering about the coin. The real question everyone is asking: would an honest coin do that often enough that we should calm down? That instinct — "could this be luck?" — is the entire round.
Data scientist loops lean on this round hardest. ML engineer loops still visit it, because A/B tests — showing two versions to two random user groups and comparing — decide whether models ship.
Why this round exists
Companies run experiments constantly: new button colour, new ranking model, new price. Someone must read the results and say "real improvement" or "luck". Reading them wrong burns money in both directions — shipping changes that do nothing, and killing changes that worked.
The five questions below are asked of nearly everyone, and are answered without a single formula. Questions 6–15 are in the developer tab, 16–20 in the researcher tab.
The instinct, in one picture
"the new design got more clicks!"
│
▼
how big is the difference?
│
▼
how much do results wobble
by pure luck at this size?
│
▼
difference >> wobble → probably real
difference ≈ wobble → cannot tell yetQ1. What does a p-value actually mean?
It answers: if the change truly did nothing, how often would luck alone produce a result this impressive? A small p-value says "luck rarely looks this good", so you lean towards the effect being real.
It does not say the probability your change works. That misreading is the trap inside this question, and interviewers listen for it specifically. The coin-toss captain is the whole idea: eight in a row from an honest coin is rare — that rarity is the p-value.
Q2. What is a confidence interval, in words?
A range that says "the true value plausibly lives somewhere in here", with the width telling you how much your data wobbles. "The new design adds between 0.3 and 4 minutes of watch time" is far more useful than "it adds 2.2 minutes", because the width warns you how sure to be.
The strong candidate adds: if the interval comfortably contains "no difference at all", the honest summary is "we cannot tell yet".
Q3. Type I versus Type II errors — which is worse?
A Type I error is a false alarm: declaring an effect that does not exist. A Type II error is a miss: failing to detect an effect that does exist. Neither is universally worse — it depends on the cost of each mistake, and saying so is the answer.
A spam filter's false alarm deletes an important mail (costly); its miss shows you one more spam (cheap). A disease screening's miss can be fatal; its false alarm costs one follow-up test. Strong answers name the costs, then choose.
Q4. Correlation is not causation — give me a real example
Ice cream sales and drowning deaths rise together every year. Ice cream does not cause drowning; summer causes both. A hidden third factor driving two things is called a confounder.
The business version scores higher: "users of our new feature churn less" does not mean the feature prevents churn — the users who chose to try it were the engaged ones already. That one sentence is why companies run experiments instead of mining old data. The full machinery lives in the causal inference section of the statistics track.
Q5. Why do bigger samples make us more confident?
Think of tasting a huge pot of dal. One spoon tells you almost nothing unless the pot is well stirred. Many spoons, from different parts of the pot, and the taste settles towards the truth. Averages calm down as samples grow — small samples swing wildly, large ones stabilise.
The follow-up trap: does more data fix bias? No. If every spoon comes from the top of an unstirred pot, a million spoons stay wrong. Sample size cures wobble, never a crooked sampling method.
Remember this
- Every question here is really: could this be luck?
- A p-value is how good luck would have to be, not the chance you are right.
- Big samples cure wobble, never bias.
What to learn next
- Deep learning questions, with answers — the next question set in the loop.
- Statistics — the foundations behind every answer here.
- Probability — Bayes' theorem and the disease-test arithmetic, taught fully.
Developer — Code and libraries.
Setup
Questions 6–15, with runnable code where code answers best.
pip install numpyRuns on CPU in about a second. Outputs made with NumPy 2.4; identical seeds reproduce them on any recent version.
Q6. Design and analyse an A/B test for me
The design half: randomise users into two groups, decide the metric and the sample size before starting, run to the planned size, then test. The analysis half is beautifully demonstrated by the permutation test — an approach worth knowing because it also explains what a p-value is:
import numpy as np
rng = np.random.default_rng(0)
# minutes watched per user in a tiny A/B test - invented numbers, realistic scale
a = np.array([12.1, 8.4, 15.2, 9.8, 11.5, 7.9, 13.3, 10.6, 9.1, 12.7]) # old design
b = np.array([14.9, 11.2, 16.8, 10.4, 13.1, 12.6, 15.5, 9.9, 14.2, 13.8]) # new design
observed = b.mean() - a.mean()
pooled = np.concatenate([a, b])
# if the design changed nothing, the group labels are meaningless -
# so shuffle the labels many times and see how often chance beats reality
extreme = 0
n_shuffles = 100_000
for _ in range(n_shuffles):
rng.shuffle(pooled)
diff = pooled[10:].mean() - pooled[:10].mean()
if abs(diff) >= abs(observed):
extreme += 1
print(f"observed difference: {observed:.2f} minutes")
print(f"permutation p-value: {extreme / n_shuffles:.4f}")
# bootstrap: resample each group with replacement to see how much the
# difference itself wobbles from sample to sample
diffs = np.array([
rng.choice(b, size=10).mean() - rng.choice(a, size=10).mean()
for _ in range(10_000)
])
lo, hi = np.percentile(diffs, [2.5, 97.5])
print(f"95% bootstrap interval for the difference: [{lo:.2f}, {hi:.2f}] minutes")observed difference: 2.18 minutes permutation p-value: 0.0501 95% bootstrap interval for the difference: [0.28, 4.06] minutes
That p-value landed at 0.0501 — a fraction above the traditional 0.05 line — and that is a gift of an interview moment. The weak candidate says "not significant, discard". The strong candidate says: 0.0501 and 0.0499 describe the same evidence; the difference is 2.2 minutes with an interval excluding zero; with ten users per group the right conclusion is "promising, underpowered — run it bigger". Treating 0.05 as a cliff edge is exactly the misunderstanding this round probes.
Q7. What is the bootstrap, and when would you reach for it?
The bootstrap re-samples your data with replacement many times, recomputing your statistic each time, and reads the spread of results as the uncertainty of that statistic. It is the second half of the code above, and it answers "how much would this number change if I could rerun the experiment?" — without rerunning anything.
Reach for it when no textbook formula exists for your statistic's uncertainty: medians, ratios, a model's AUC, revenue per user with its wild skew. Limitations to volunteer: it struggles with very small samples (resampling ten points only rearranges those ten values) and with extremes like the maximum.
Q8. A disease test is 99% accurate. You test positive. How worried should you be?
The most famous statistics interview question, and the answer shocks everyone once. Suppose 1 person in 1,000 has the disease, the test catches 99% of sick people, and it false-alarms on 5% of healthy people.
Take 100,000 people. About 100 are sick, and the test catches 99 of them. About 99,900 are healthy, and 5% false-alarm — 4,995 people. Your positive result puts you in a room of 99 + 4,995 = 5,094 positives, of whom 99 are sick:
P(sick | positive) = 99 / 5,094 ≈ 1.9%Under 2%, because the disease is rare and false alarms swamp true ones. This is Bayes' theorem in action — updating a belief with evidence, where the base rate (how common the condition is) dominates. The interview transfer: a 99%-accurate fraud model drowning in false positives on rare fraud is the same arithmetic, which connects straight to precision and recall in ml-theory-questions. Foundations: probability.
Q9. What is statistical power, and what increases it?
Power is the probability your test detects an effect that genuinely exists — the complement of the Type II error rate. Underpowered tests are how real improvements get killed: the effect was there, the test could not see it, the feature died.
Four levers: bigger samples, bigger true effects (test bold changes, not button shades), less noisy metrics, and a looser significance threshold. The practice sentence: decide power before the test, by computing the sample size needed to detect the smallest effect worth shipping — never after.
Q10. Why can't I check my A/B test daily and stop when it becomes significant?
Because you are giving luck many chances and keeping its best attempt. Each daily peek is another lottery ticket; across weeks of peeking, the chance of at least one accidental "significant!" balloons far past 5%. The test's guarantee applied to one look at a pre-planned size.
Name the practice: this is the peeking problem, and it produces winners that evaporate after shipping. The fixes: commit to the sample size in advance, or use methods built for continuous monitoring — sequential tests, which spend an error budget across looks. Knowing that such methods exist is the senior part of the answer.
Q11. You measured 20 metrics and one moved significantly. What do you tell your boss?
That this is expected under pure chance: at a 5% threshold, one-in-twenty metrics false-alarms on average even when nothing changed. Finding exactly one mover in twenty is what "no effect anywhere" looks like.
The multiple comparisons problem, by name. Remedies: pick one primary metric before the experiment and let the rest be exploratory; or correct the threshold — Bonferroni divides it by the number of tests (crude, safe), false-discovery-rate methods are gentler at scale. The one-liner that lands: "if you torture the data long enough, it will confess to anything."
Q12. Standard deviation versus standard error — what is the difference?
Standard deviation measures how much individual data points spread around their mean — a fact about the population that more data does not shrink. Standard error measures how much the sample mean itself would wobble across repeated samples — and it shrinks as the sample grows, falling with the square root of the sample size.
The square-root part is worth saying: four times the data halves the standard error. Precision gets expensive fast — a tenth of the error costs a hundred times the sample. Interviewers probe this with: "your mean is stable but users vary hugely — contradiction?" No: small standard error, large standard deviation, at the same time.
Q13. What is Simpson's paradox? Construct an example
A trend that holds in every subgroup can reverse when the subgroups are pooled. The canonical real dataset — kidney-stone treatments, Charig et al. (1986):
| Cure rate | Small stones | Large stones | Pooled |
|---|---|---|---|
| Treatment A | 81/87 (93%) | 192/263 (73%) | 273/350 (78%) |
| Treatment B | 234/270 (87%) | 55/80 (69%) | 289/350 (83%) |
A wins on small stones. A wins on large stones. Pooled, B "wins". The reversal happens because A — the more invasive treatment — was given mostly the hard cases, so its pooled rate is dragged down by case mix, not by quality.
The lesson to state: pooled comparisons silently assume the groups faced the same mix of cases. When a lurking variable decides who gets what, subgroup analysis — or better, randomisation — is the defence.
Q14. When is the mean the wrong summary?
Whenever the distribution is skewed or heavy-tailed, because a few extreme values drag the mean away from the typical case. Nine people earning 30,000 rupees and one earning 3,000,000 have a mean income of 327,000 — describing nobody. The median (50th percentile) says 30,000 and describes almost everyone.
The engineering version scores extra: latency. Mean response time hides disasters; the 95th and 99th percentiles are what users feel. "Report medians and tail percentiles for anything skewed — money, latency, session length" is the practitioner's answer.
Q15. What is regression to the mean? Why do managers believe punishment works?
Extreme results are part skill, part luck, and the luck does not repeat. The best performer this month was both good and lucky; next month the luck resets and they look worse — with no real change. The worst performer improves the same way.
The famous consequence (Kahneman's flight instructors): punish after terrible performance, and things improve anyway, because extremes drift back towards normal. Praise after brilliance, and things decline anyway. Punishment gets false credit; praise gets false blame. The interview transfer: "our worst-performing stores improved after the intervention" is uninterpretable without a control group, because the worst stores were partly unlucky and would have rebounded regardless.
Common mistakes in this round
Reciting mechanics without the trap. Every question above has a famous misreading; naming it is most of the score.
Speaking in formulas. This round is spoken. The candidate who can explain a p-value to a smart non-statistician outscores the one who writes integrals.
Forgetting the business. "Significant" and "worth shipping" are different claims — a real effect can be too small to matter. End answers with the decision, not the statistic.
Try it yourself
Rerun permutation_ab.py after changing one user's number in group b — make 9.9 into 13.9 — and predict the p-value's direction before running. Then answer Q8 aloud with a 1-in-100 base rate instead of 1-in-1,000, doing the people-count arithmetic as you speak.
What to learn next
- Deep learning questions, with answers — the next question set in the loop.
- Statistics — the foundations behind every answer here.
- Probability — Bayes' theorem and the disease-test arithmetic, taught fully.
Researcher — Mathematics and papers.
Questions 16–20: the follow-ups that separate "took a course" from "understands the machinery".
Q16. State the Central Limit Theorem precisely — and what it does not say
For i.i.d. random variables $X_1, \dots, X_n$ with mean $\mu$ and finite variance $\sigma^2$:
$$ \sqrt{n}\,\frac{\bar{X}_n - \mu}{\sigma} \;\xrightarrow{d}\; \mathcal{N}(0, 1) \quad \text{as } n \to \infty $$
Where:
- $\bar{X}n = \frac{1}{n}\sum{i=1}^n X_i$ — the sample mean.
- $\xrightarrow{d}$ — convergence in distribution: the CDF of the left side converges pointwise to the standard normal CDF.
What it does not say, which is where the marks are: it says nothing about any fixed $n$ (the "$n \ge 30$" rule is folklore, not theorem — heavily skewed data can need thousands); it concerns the mean, not the data (the data's distribution never becomes normal); it requires finite variance (Cauchy-tailed data breaks it; stable laws take over); and independence matters (correlated observations converge more slowly or to different limits, which is why clustered A/B assignments need clustered standard errors). The Berry–Esseen theorem bounds the finite-$n$ error at rate $O(1/\sqrt{n})$ with a constant depending on the third moment — skewness literally slows the convergence.
Q17. What exactly does "95% confidence" mean — and what do people think it means?
The parameter $\theta$ is fixed; the interval is random. A 95% procedure satisfies:
$$ P_\theta\big(\theta \in [L(X), U(X)]\big) \ge 0.95 \quad \text{for all } \theta $$
The probability statement is about the procedure across repeated samples: 95% of the intervals it generates would contain the truth. Any one realised interval either contains $\theta$ or does not — assigning it "95% probability" is strictly a Bayesian statement about a credible interval, which requires a prior. The two intervals can coincide numerically (flat priors, nice likelihoods) while meaning different things, and saying that cleanly is the whole answer.
Worth adding: the same logic explains why a p-value is not $P(H_0 \mid \text{data})$ — frequentist machinery never assigns probabilities to hypotheses. The ASA's 2016 statement on p-values (Wasserstein and Lazar) is the citable reference for the misuse catalogue.
Q18. Derive a maximum likelihood estimator, and tell me when MLE misbehaves
MLE picks the parameter making the observed data most probable: $\hat{\theta} = \arg\max_\theta \sum_i \log p(x_i \mid \theta)$, logs taken because sums are differentiable and numerically sane where products underflow.
The interview-standard derivation, Bernoulli: with $k$ successes in $n$ trials, $\ell(p) = k \log p + (n-k)\log(1-p)$. Setting $\ell'(p) = k/p - (n-k)/(1-p) = 0$ gives $\hat{p} = k/n$ — the sample proportion, reassuringly.
The misbehaviour catalogue is the researcher part. Small samples: MLE is biased (variance MLE divides by $n$, not $n-1$). Boundary cases: zero successes gives $\hat{p} = 0$, an absurd certainty — the cure is smoothing, which is a prior in disguise (MAP estimation). Singularities: Gaussian mixtures where one component's variance shrinks onto a point send the likelihood to infinity. And the deep connection: maximising likelihood is minimising KL divergence from the empirical distribution to the model family — which is why cross-entropy training of neural networks is MLE, tying this answer to loss-functions.
Q19. Bayesian versus frequentist — what changes in practice, not philosophy?
The operational differences, which is what a hiring panel wants. Bayesian inference produces a posterior $p(\theta \mid x) \propto p(x \mid \theta)\,p(\theta)$, so uncertainty statements are direct probabilities over parameters; small-sample inference is coherent given a prior; and sequential updating is natural — no peeking problem, because the posterior is valid whenever you look (though decision rules on top of it can still inflate errors, a caveat worth voicing).
The costs: priors must come from somewhere and can be contested; computation historically needed MCMC, though variational methods changed the economics.
The bridge that earns the senior nod: regularisation is priors. Ridge regression is MAP with a Gaussian prior; lasso is MAP with a Laplace prior — so most practitioners are quietly Bayesian already, and the L1/L2 question in ml-theory-questions is this question wearing different clothes.
Q20. Why 0.05? Defend or attack the threshold
History, not mathematics. Fisher suggested one-in-twenty as a convenient benchmark in the 1920s, tables were printed around it, and journals fossilised it. No optimality theorem exists for 0.05.
The Neyman–Pearson framing makes the threshold a decision variable: choose $\alpha$ by trading the cost of false positives against false negatives at your sample size. Particle physics demands five sigma ($\alpha \approx 3 \times 10^{-7}$) because a false discovery embarrasses a field; a low-cost website tweak might rationally ship at $\alpha = 0.2$. The modern critique to cite: Benjamin et al. (2018) proposed 0.005 for discovery claims; the deeper point is that any bright-line threshold invites the gaming (p-hacking, optional stopping — Q10) that pre-registration and holdout discipline exist to prevent. Ending with "report the estimate, the interval and the cost of both errors; the threshold is a decision, not a truth" is the strongest close available.
Sources
- Charig, C. R. et al. (1986), Comparison of treatment of renal calculi..., BMJ 292 — the Simpson's paradox kidney-stone data.
- Wasserstein, R. and Lazar, N. (2016), The ASA statement on p-values, The American Statistician 70(2).
- Benjamin, D. et al. (2018), Redefine statistical significance, Nature Human Behaviour 2.
- Efron, B. (1979), Bootstrap methods: another look at the jackknife, Annals of Statistics 7(1).
- Kahneman, D. (2011), Thinking, Fast and Slow — regression to the mean and the flight-instructor story.
- Wasserman, L., All of Statistics — the compact reference for Q16–Q19.
What to learn next
- Deep learning questions, with answers — the next question set in the loop.
- Statistics — the foundations behind every answer here.
- Probability — Bayes' theorem and the disease-test arithmetic, taught fully.