Calibration and Uncertainty

Proper scoring rules

A proper scoring rule is a grading scheme under which honest probabilities earn the best expected grade — hedging and bluffing both lose, which accuracy cannot promise.

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A proper scoring rule grades predictions so that telling your true belief is the winning strategy.

Bluff or hedge, and the grading takes points off you in the long run.

Imagine running a cricket-prediction contest in your office. Everyone states how confident they are that India wins, and you award points after the match. Design the grading badly and people stop reporting beliefs: they learn that shouting "100%!" gains them points, or that mumbling "50-50" avoids all risk. The contest now measures gaming skill.

Design the grading well and something lovely happens: the way to win is to say exactly what you believe. Such a grading scheme is called a proper scoring rule.

Why this matters

Grading predictions by accuracy alone is a badly designed contest. Accuracy checks only which side of the halfway mark you stood on. A timid 51% and a ringing 99% earn identical credit when right. They earn identical blame when wrong. Under that grading, confidence numbers are decoration, and a model (or a forecaster) has no reason to make them honest.

The moment you use the confidence — pricing decisions by cost, triaging patients, sizing bets — dishonest confidence costs real money. So the grading must reward honesty. Proper scoring rules are exactly the gradings that do.

How it works

Two proper gradings dominate practice, and they have different tempers.

The Brier score is the even-tempered judge. Bad calls cost you steadily. No single miss can wipe out a good year.

Log-loss is the harsh judge. It shrugs when you were roughly right. It turns savage as a wrong claim approaches certainty.

what you said  ->  it happened  ->  Brier penalty     log-loss penalty

    90%               yes            tiny                tiny
    55%               yes            small               small
     1%               yes            huge                ENORMOUS

The character difference matters. Brier forgives a rare disaster; log-loss never forgets one. A single "0.1% chance" claim on something that happens can outweigh a thousand good predictions under log-loss. Choose the rule whose temperament matches the price of overconfidence in your world.

A real example you have seen

Weather services grade their forecasters on the Brier score — it was invented for rain forecasts. That grading is why app forecasts feel trustworthy. A forecaster who hedges everything, or who bluffs certainty, loses points year after year. The colleagues reporting honest chances win.

Remember this

  • Proper rule = honesty is the winning strategy; hedging and bluffing both lose.
  • Brier: even-tempered. Log-loss: brutal on confident wrongness.
  • Accuracy ignores confidence entirely — never grade probabilities with it alone.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2.

Three forecasters, one truth

Same eight matches, three prediction styles. All three land the same side of 50% every time — identical accuracy — and the proper rules see through them:

three_forecasters.py
import numpy as np
from sklearn.metrics import accuracy_score, brier_score_loss, log_loss

y = np.array([1, 0, 1, 1, 0, 1, 0, 0])
honest = np.array([0.8, 0.2, 0.7, 0.9, 0.3, 0.8, 0.2, 0.6])
hedger = np.array([0.6, 0.4, 0.55, 0.6, 0.45, 0.6, 0.4, 0.55])
cocky = np.array([0.999, 0.001, 0.999, 0.999, 0.001, 0.999, 0.001, 0.999])

for name, p in [("honest", honest), ("hedger", hedger), ("cocky", cocky)]:
    acc = accuracy_score(y, p >= 0.5)
    print(f"{name:7s} accuracy={acc:.3f}  "
          f"Brier={brier_score_loss(y, p):.3f}  log-loss={log_loss(y, p):.3f}")
Output
honest  accuracy=0.875  Brier=0.089  log-loss=0.328
hedger  accuracy=0.875  Brier=0.188  log-loss=0.569
cocky   accuracy=0.875  Brier=0.125  log-loss=0.864

The walkthrough

Accuracy: a three-way tie at 0.875. All three called the same 7 of 8 matches. If accuracy is your only judge, these forecasters are interchangeable — which is exactly the blindness this lesson exists to cure.

The honest forecaster wins both proper scores. Brier 0.089, log-loss 0.328 — comfortably ahead. Confident when the data justified it, moderate when not, and both rules paid out.

The hedger's timidity costs steadily. Never below 0.4, never above 0.6, so mistakes stay cheap — but every correct call earns feeble credit too. Brier 0.188 is the worst of the three: constant mild dishonesty, priced in full.

The cocky forecaster splits the two judges — and this is the subtlest line of the output. On Brier, 0.125: seven near-perfect calls hedge one disaster (the match 8 miss at 99.9% confidence), landing better than the hedger. On log-loss, 0.864: worst by far, because one certain-and-wrong claim carries a penalty of about 6.9 alone against near-zero elsewhere. Same predictions — the rules disagree by design about how unforgivable overconfidence is.

Both are proper; neither is wrong. Choose log-loss when a confident falsehood is catastrophic (medical triage, LLM training uses it as cross-entropy). Choose Brier when you want steadier gradients and bounded penalties (forecast verification, model comparison dashboards).

Common mistakes

Reporting accuracy for a probability model. The output above is the full argument: three very different probability qualities, one accuracy.

Clipping predictions to dodge log-loss infinities, silently. log_loss on an exact 0 for an event that happens is infinite by definition; sklearn clips internally at float precision. If you clip more aggressively for stability, disclose it — clipping changes the score you claim to report.

Comparing proper scores across different datasets. Both rules depend on how hard the problem is. Brier 0.089 on easy matches is no better than 0.15 on coin-flip matches. Compare models on the same test set, or report skill scores relative to a baseline.

Optimising a proper score, expecting calibration alone. Proper scores reward calibration and discrimination together. A model can improve its Brier by ranking better while staying miscalibrated — the decomposition separates the two.

Try it yourself

Add a fourth forecaster: contrarian = 1 - honest. Predict its accuracy and both scores before running. Then soften cocky's claims to 0.99 / 0.01 and watch which judge's opinion changes more — that sensitivity is the temperament difference.

What to learn next

Researcher — Mathematics and papers.

Definition and the two canonical rules

A scoring rule $S(p, y)$ assigns a penalty to forecast $p$ against outcome $y$. It is proper if truth-telling minimises expected penalty: for all $q$,

$$ \mathbb{E}{y \sim \mathrm{Bern}(q)}\big[S(q, y)\big] \;\le\; \mathbb{E}{y \sim \mathrm{Bern}(q)}\big[S(p, y)\big] \quad \forall p $$

and strictly proper if equality holds only at $p = q$.

Where:

  • $q$ — the forecaster's true belief; $p$ — the reported probability.
  • $\mathrm{Bern}(q)$ — outcomes drawn with true probability $q$.

The Brier score $S(p,y) = (p - y)^2$ (Brier, 1950) and the logarithmic score $S(p,y) = -\log p_y$ (Good, 1952) are both strictly proper. Accuracy — the 0/1 score thresholded at 0.5 — is proper but not strictly: any $p$ on the correct side is optimal, which formalises the three-way tie in the developer block.

Structure theory

The Savage representation (Savage, 1971; Gneiting and Raftery, 2007, the field's definitive survey): every proper scoring rule corresponds to a convex function $G$ (the generalised entropy), with

$$ S(p, y) = -G(p) - G'(p)\,(y - p) $$

Brier arises from $G(p) = p(1-p)$ scaled; log-loss from Shannon entropy. Consequences: proper rules form a convex family; each embodies a different weighting over decision thresholds. Schervish (1989) makes the weighting exact — every proper rule is a mixture of elementary cost-weighted decision losses, tying this lesson to cost-based thresholds: Brier weights all cost ratios uniformly; log-loss up-weights extreme cost ratios, hence its severity at confident errors.

The Murphy decomposition (1973) splits the Brier score into reliability minus resolution plus uncertainty — calibration and discrimination as orthogonal components, the quantitative bridge to reliability diagrams.

Selection consequences and elicitation

  • Optimising a strictly proper loss during training (log-loss is the standard classification objective) makes calibrated probabilities a stationary target — though modern over-parameterised optimisation still lands miscalibrated (Guo et al., 2017), which is why post-hoc calibration persists.
  • Improper objectives invite gaming: optimising F1 or accuracy at a threshold pushes probability mass to the threshold's sides, destroying probability meaning.
  • Elicitation theory generalises to other functionals: means are elicitable (squared error), quantiles are elicitable (pinball loss — the engine of prediction intervals), variance alone is not (Gneiting, 2011, Making and evaluating point forecasts).
  • For full distributions over continuous outcomes, the CRPS (continuous ranked probability score) integrates Brier over all thresholds and is strictly proper; it anchors probabilistic forecast evaluation in meteorology and, increasingly, ML forecasting.

Prediction markets and forecasting tournaments (Tetlock's Good Judgment Project) run on proper rules for the same reason your office contest should: the rule is the incentive design.

What to learn next