Resampling, Likelihood and Bayes
Bayesian vs frequentist thinking
Two working definitions of probability — long-run frequency versus degree of belief — and how each answers the questions the other cannot ask.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Frequentists define probability as long-run frequency; Bayesians define it as degree of belief. Each definition unlocks different questions.
Ask "what is the chance this coin lands heads?" and both camps answer comfortably: flip it many times, count. Frequency works because flipping repeats.
Now ask "what is the chance it rains during tomorrow's match?" Tomorrow happens once. There is no long run to count. Yet your brain still holds a number — a strength of belief, updated every time you glance at the clouds.
Why the split exists
Frequentist tools ruled the twentieth century: p-values, confidence intervals, all built on "imagine repeating the study". They ask no opinions, which feels objective — and they answer only repeatable-style questions.
The Bayesian view is older — Thomas Bayes, 1763 — and it treats unknown things themselves as uncertain, describable by belief. It fell from favour partly because its computations were hopeless by hand. Computers fixed that, and the approach returned in force.
How the two think
frequentist: the true value is FIXED but unknown.
randomness lives in the DATA.
"If I repeated this survey 100 times,
95 of my intervals would catch the truth."
bayesian: beliefs come as distributions.
start with a PRIOR belief,
let data reshape it into a POSTERIOR belief.
"Given this data, I am 95% sure the truth
is between here and here."The prior is what you believed before the data. The posterior is what you believe after. Data converts one into the other.
Note the trade at the heart of it. Bayesians get to answer the question everyone actually asks: "how sure are we, given this data?" The price is stating a starting belief before the data arrives. Frequentists avoid stating beliefs, and in exchange answer a more roundabout question.
With plenty of data, the two camps nearly always agree. The philosophy matters most when data is scarce.
A real example you have seen
Election forecasts saying "72% chance candidate X wins" are Bayesian statements — one election, no long run. Spam filters literally run on Bayes' rule. Weather "chance of rain" mixes both traditions. And every drug's approval file is written mostly in frequentist language, because regulators want procedures with guaranteed error rates.
Remember this
- Frequentist: probability = long-run frequency; randomness lives in the data.
- Bayesian: probability = belief; data updates prior into posterior.
- Lots of data → both agree. Scarce data → the prior (and the philosophy) matters.
What to learn next
- Priors, posteriors and conjugate updating — the mechanics of belief updating, in closed form.
- MCMC from scratch — posterior computation when closed forms run out.
- Confidence intervals — the frequentist promise, stated precisely.
Developer — Code and libraries.
Setup
pip install scipyOutputs verified with scipy 1.14, CPU.
One dataset, both interviews
A new feature succeeded in 9 of 12 demos. Question it both ways.
from scipy import stats
k, n = 9, 12
# Frequentist: a confidence interval for the long-run success rate.
ci = stats.binomtest(k, n).proportion_ci(confidence_level=0.95)
print(f"frequentist 95% CI: ({ci.low:.2f}, {ci.high:.2f})")
# Bayesian: flat prior + data -> posterior. Read off beliefs directly.
posterior = stats.beta(1 + k, 1 + n - k)
lo, hi = posterior.ppf([0.025, 0.975])
print(f"Bayesian 95% credible interval: ({lo:.2f}, {hi:.2f})")
print(f"P(success rate > 0.5) = {1 - posterior.cdf(0.5):.3f}")frequentist 95% CI: (0.43, 0.95) Bayesian 95% credible interval: (0.46, 0.91) P(success rate > 0.5) = 0.954
The walkthrough
The last line is the whole argument for Bayes. "There is a 95.4% probability the feature succeeds more often than not" — a direct answer to the business question. No frequentist quantity can say that sentence; a p-value talks about data under a null, not about the parameter.
stats.beta(1 + k, 1 + n - k) is Bayes' rule solved in closed form. Start from Beta(1, 1) — a flat prior, every success rate equally believable — observe 9 successes and 3 failures, and the posterior is Beta(10, 4). Why this pairing works so neatly is the next lesson, conjugate priors.
The two intervals nearly coincide — by design of the example. Twelve data points already outweigh a flat prior. Shrink the data to 3 successes in 4 demos and the intervals visibly part ways; the prior's voice grows as the data's shrinks.
A credible interval and a confidence interval are different promises. The credible interval: "given this data and prior, 95% of my belief sits here". The confidence interval: "the recipe that made me catches the truth in 95% of repeat studies". The confidence-intervals lesson covers why the second is so often misread as the first.
Common mistakes
Calling a flat prior "no assumptions". Flat on the rate is not flat on, say, the log-odds — flatness depends on the ruler. A flat prior is an assumption wearing camouflage; state it and move on.
Using a Bayesian posterior but reporting it as a p-value (or vice versa). Mixed language produces claims neither framework supports. Pick the question first: "guaranteed error rate" → frequentist; "belief given this data" → Bayesian.
Fighting about philosophy when n is large. Rerun the code with 900 of 1200: the intervals agree to two decimals. Save the debate for small-data problems, where it earns its keep.
Choosing a prior after seeing the data. That double-counts the data and manufactures confidence. Priors come from before: past experiments, physical limits, or deliberate vagueness.
Try it yourself
Replace the flat prior with a sceptical one, stats.beta(2 + k, 8 + n - k) — a prior leaning toward low success rates. Recompute P(success rate > 0.5) and watch scarce data lose an argument with a strong opinion.
What to learn next
- Priors, posteriors and conjugate updating — the mechanics of belief updating, in closed form.
- MCMC from scratch — posterior computation when closed forms run out.
- Confidence intervals — the frequentist promise, stated precisely.
Researcher — Mathematics and papers.
The formal divide
Both camps accept Bayes' theorem as algebra:
$$ p(\theta \mid x) = \frac{p(x \mid \theta)\, p(\theta)}{p(x)} $$
Where:
- $p(\theta)$ — the prior over parameter $\theta$.
- $p(x \mid \theta)$ — the likelihood.
- $p(x) = \int p(x \mid \theta) p(\theta)\, d\theta$ — the evidence (marginal likelihood).
- $p(\theta \mid x)$ — the posterior.
The divide is whether $p(\theta)$ is meaningful. Frequentism: $\theta$ is a fixed constant; only sampling statements $\Pr_\theta(X \in A)$ are licensed; optimality means risk guarantees uniform over $\theta$. Bayesianism: probability is coherent degree of belief (de Finetti, 1937; Savage, 1954); all inference flows from the posterior; decisions minimise posterior expected loss.
Exchangeability, the quiet bridge
De Finetti's representation theorem: an infinite exchangeable binary sequence is distributionally a mixture of i.i.d. Bernoulli sequences —
$$ p(x_1, \dots, x_n) = \int_0^1 \prod_{i=1}^n \theta^{x_i}(1-\theta)^{1-x_i} \; d\mu(\theta) $$
so "a parameter with a prior" emerges as a theorem about symmetric beliefs rather than a metaphysical posit. This is the cleanest philosophical foundation on offer.
Asymptotic reconciliation
The Bernstein–von Mises theorem: under regularity, the posterior converges to $\mathcal{N}!\left(\hat\theta_{MLE},\, I(\theta_0)^{-1}/n\right)$ in total variation — priors wash out, and credible intervals acquire asymptotically correct frequentist coverage. Caveats: fails for infinite-dimensional parameters in general (Freedman, 1999), under misspecification, and in the small-$n$ regime where the debate actually lives.
Complete-class results run the other direction: admissible frequentist procedures are (limits of) Bayes rules (Wald, 1950) — even a committed frequentist is, in effect, choosing among priors.
Where each pays rent
- Frequentist strengths: error-rate guarantees without belief elicitation; multiple-testing control; distribution-free methods (permutation, conformal); regulatory auditability.
- Bayesian strengths: coherent uncertainty for one-shot events; principled pooling via hierarchical models (partial pooling beats both no-pooling and complete-pooling — Gelman et al., Bayesian Data Analysis, 3rd ed., 2013); sequential updating without stopping-rule penalties (the likelihood principle); decision theory with explicit losses; small-$n$ regularisation via priors.
- ML translations: ridge = MAP with Gaussian prior; dropout and ensembling as approximate posterior inference; Thompson sampling for bandits is Bayesian decision theory in production at every large tech firm.
Modern practice is a portfolio
Empirical Bayes estimates the prior from data (Efron, Large-Scale Inference, 2010 — the workhorse of genomics). Calibrated Bayes (Little, 2006) uses Bayesian machinery, then checks frequentist operating characteristics by simulation. The pragmatic rule: choose the framework per question, and validate whichever you use.
Key texts: Bayes (1763, published posthumously by Price); Fisher (1925); Neyman (1937); Jaynes, Probability Theory: The Logic of Science (2003) for the maximal Bayesian case; Efron and Hastie, Computer Age Statistical Inference (2016) for the truce.
What to learn next
- Priors, posteriors and conjugate updating — the mechanics of belief updating, in closed form.
- MCMC from scratch — posterior computation when closed forms run out.
- Confidence intervals — the frequentist promise, stated precisely.