Resampling, Likelihood and Bayes

Priors, posteriors and conjugate updating

With the right prior family, Bayesian updating collapses to addition — Beta plus coin flips stays Beta — giving exact posteriors at zero computational cost.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A conjugate prior is a starting belief shaped so that updating it with data keeps the same shape — turning Bayes' rule into plain addition.

Think of a ledger you keep on a new employee: one column of "times they delivered", one of "times they slipped". Every week you add this week's counts to the running totals. Your opinion at any moment is the ledger.

Conjugate updating is exactly this. Belief lives as a pair of counts; data arrives as counts; update means add. Nothing else.

Why it exists

Bayes' rule in general demands ugly integrals — for two centuries, that made Bayesian analysis mostly a philosophy. Raiffa and Schlaifer formalised the workaround in the 1960s. Choose the prior from a family matched to your data type, and the integral solves itself, forever.

Matched pairs exist for the common data types. Yes/no data pairs with the Beta family. Count data pairs with the Gamma family. Bell-curve data pairs with the normal family. Learn one pair well and the pattern carries.

How it works

The Beta family describes beliefs about an unknown rate — a click rate, a success rate. It has two dials, and they read as imaginary prior counts: successes seen, failures seen.

prior:   Beta(3, 7)      "as if I'd already seen 3 opens, 7 ignores"
data:    4 opens, 6 ignores

update:  add successes to the first dial, failures to the second
posterior: Beta(3+4, 7+6) = Beta(7, 13)

Everything you want falls out of the updated pair. Your best estimate is roughly the success fraction of the combined counts. Your uncertainty shrinks as the totals grow — big ledgers are hard to sway.

And you can update daily, weekly, or all at once: the ledger ends up identical. That stream-friendliness is why this machinery loves live systems.

A real example you have seen

Think of "new seller" badges versus ratings on established sellers. One bad review moves a newcomer's score a lot, and a veteran's barely. Thousands of prior counts anchor the veteran. Ad systems and recommendation engines keep exactly these running-count beliefs per ad or item, updating with every impression. That is conjugate updating in production, billions of times a day.

Remember this

  • Conjugate prior + matching data = posterior in the same family; update = add counts.
  • Beta's two dials read as imaginary successes and failures already seen.
  • Streaming updates and one big update agree — order does not matter.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scipy

Outputs verified with scipy 1.14, CPU.

A newsletter's open rate, learned as data arrives

Prior belief: open rates around 30%, held loosely — worth about ten imaginary emails. Then three batches of real data arrive.

beta_updating.py
from scipy import stats

a, b = 3, 7                # Beta(3, 7): mean 0.3, weight of ~10 observations

batches = [(4, 6), (11, 19), (37, 63)]      # (opened, ignored) per batch
for opened, ignored in batches:
    a, b = a + opened, b + ignored          # the entire update rule
    post = stats.beta(a, b)
    lo, hi = post.ppf([0.025, 0.975])
    print(f"after +{opened:>2} opens / {opened + ignored:>3} mails -> Beta({a},{b}), "
          f"mean {post.mean():.3f}, 95% ({lo:.3f}, {hi:.3f})")
Output
after + 4 opens /  10 mails -> Beta(7,13), mean 0.350, 95% (0.163, 0.566)
after +11 opens /  30 mails -> Beta(18,32), mean 0.360, 95% (0.234, 0.496)
after +37 opens / 100 mails -> Beta(55,95), mean 0.367, 95% (0.292, 0.445)

The walkthrough

a, b = a + opened, b + ignored — the promised addition. Bayes' rule, integral and all, compressed into one line because the shapes match.

Watch the interval tighten: width 0.40 → 0.26 → 0.15 as evidence piles up. The mean drifts toward the data's rate (about 0.37) and away from the prior's 0.30. Both movements are automatic.

The posterior mean has a readable formula: total successes over total counts, 55 / 150 = 0.367 — prior imaginary counts included. A conjugate posterior is a weighted compromise between prior and data, with weights equal to their counts.

Choosing the prior dials. Mean first: a / (a + b) should match your honest expectation. Strength second: a + b is how many real observations your opinion is worth. Beta(3, 7) says "30%, weakly held". Beta(30, 70) says the same mean, stubbornly. Beta(1, 1) is the flat prior from bayesian-vs-frequentist.

Answer real questions straight off the posterior:

python
print(f"P(open rate > 0.40) = {1 - post.cdf(0.40):.3f}")
Output
P(open rate > 0.40) = 0.197

Common mistakes

Making the prior secretly enormous. Beta(300, 700) needs a thousand real observations to argue with. If the posterior refuses to move, print a + b and check who is outvoting whom.

Updating with data that is not yes/no per trial. The Beta pairs with independent success/failure trials. Opens per user with repeat sends, or correlated batches, violate the model; aggregate to the right unit first.

Reporting the posterior mean without the interval. mean 0.350 after ten emails hides an interval spanning 0.16 to 0.57 — a decision made on that mean alone is a coin toss in disguise.

Recomputing from scratch each batch out of distrust. The ledger property is exact: Beta(3+52, 7+88) in one step equals the three-step result. Test it once in a REPL and then trust the algebra.

Try it yourself

Run the same three batches against a stubborn prior, a, b = 30, 70. Compare the final means and intervals. Then compute how many opens-in-a-row it would take to drag the stubborn posterior's mean above 0.40.

What to learn next

Researcher — Mathematics and papers.

Definition

A family $\mathcal{P} = {p(\theta \mid \eta)}$ is conjugate to a likelihood $p(x \mid \theta)$ when the posterior remains in $\mathcal{P}$ for every prior in $\mathcal{P}$ and every dataset. For the Bernoulli likelihood and Beta prior:

$$ p(\theta \mid x) \propto \underbrace{\theta^{k}(1-\theta)^{n-k}}{\text{likelihood}} \; \underbrace{\theta^{\alpha-1}(1-\theta)^{\beta-1}}{\text{Beta}(\alpha,\beta)} = \theta^{\alpha+k-1}(1-\theta)^{\beta+n-k-1} $$

Where:

  • $k$ — successes in $n$ trials; $\alpha, \beta$ — prior shape parameters.
  • The result is $\text{Beta}(\alpha + k,\ \beta + n - k)$ by inspection — no integral computed.

Posterior mean $\frac{\alpha + k}{\alpha + \beta + n}$ interpolates prior mean and MLE with weights $\alpha + \beta$ and $n$; posterior variance contracts at rate $O(1/n)$.

The exponential-family engine

Every regular exponential family likelihood

$$ p(x \mid \theta) = h(x) \exp!\left(\eta(\theta)^\top T(x) - A(\theta)\right) $$

admits a conjugate prior $p(\theta) \propto \exp!\left(\eta(\theta)^\top \tau - \nu A(\theta)\right)$, with hyperparameters $(\tau, \nu)$ acting as pseudo-sufficient-statistics and pseudo-count. Updating: $\tau \mathrel{+}= \sum_i T(x_i)$, $\nu \mathrel{+}= n$. The classical table falls out:

likelihoodconjugate priorposterior update
Bernoulli/BinomialBeta$(\alpha, \beta)$$\alpha{+}k,\ \beta{+}n{-}k$
PoissonGamma$(\alpha, \beta)$$\alpha{+}\sum x_i,\ \beta{+}n$
ExponentialGamma$(\alpha, \beta)$$\alpha{+}n,\ \beta{+}\sum x_i$
Normal (known $\sigma^2$)Normal$(\mu_0, \tau_0^2)$precision-weighted mean
Normal (unknown both)Normal-Inverse-Gammafour-parameter update
MultinomialDirichlet$(\boldsymbol\alpha)$$\alpha_j {+} c_j$
Multivariate normal cov.Inverse-Wishartscatter-matrix update

Diaconis and Ylvisaker (1979, Conjugate priors for exponential families) characterise these priors and prove the linearity of posterior expectation in the data — the "weighted compromise" made precise.

Marginals and the evidence

Conjugacy also yields the prior predictive in closed form. For Beta-Bernoulli, the marginal likelihood of $k$ successes in $n$ trials is the Beta-Binomial:

$$ p(k \mid n) = \binom{n}{k} \frac{B(\alpha + k, \beta + n - k)}{B(\alpha, \beta)} $$

with $B$ the Beta function — closed-form evidence enables exact Bayes factors and empirical-Bayes fitting of $(\alpha, \beta)$ by maximising this marginal across many units (the classic batting-average analysis: Efron and Morris, 1975).

Uses and limits

Conjugacy powers Thompson sampling (sample $\theta \sim \text{Beta}(\alpha_i, \beta_i)$ per arm, pull the argmax — Thompson, 1933; Chapelle and Li, 2011 for the modern revival), naive Bayes smoothing (Dirichlet pseudo-counts = Laplace smoothing), Kalman filtering (Gaussian conjugacy through time), and the collapsed Gibbs samplers of topic models.

Limits: conjugate families are single-peaked and thin-tailed — they cannot express "either near 0.1 or near 0.9"; mixtures of conjugate priors remain conjugate (with weight updates) and fix some of this. Hierarchical models break clean conjugacy at the top level, and non-exponential-family likelihoods (logistic, neural) have none — the opening for MCMC and variational inference.

Reference text: Raiffa and Schlaifer, Applied Statistical Decision Theory (1961), which coined "conjugate"; Gelman et al., Bayesian Data Analysis (3rd ed., 2013), ch. 2.

What to learn next