Calibration and Uncertainty

Reliability diagrams and calibration error

A model is calibrated when its 70% predictions come true about 70% of the time — reliability diagrams draw that promise-versus-reality check, and ECE compresses it to one number.

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A model is calibrated when its confidence numbers mean what they say.

Of all the times it claims "70% sure", the thing should happen about 70% of the time.

Judge a TV weather forecaster the fair way. Collect every day she said "80% chance of rain" — say forty days over a year. Now count how many of those days actually got rain. Around 32 of 40? Her 80% means 80%. Rain on only 15? Her "80%" was bluster, whatever her overall accuracy.

This promise-versus-reality audit is called calibration, and models deserve it as much as forecasters.

Why this matters

Downstream decisions consume the probability, not the label. The cost-based threshold formula, a doctor weighing treatment against a 12% risk figure, an insurer pricing a policy — all take the number literally. If the model's "12%" really means 30%, every one of those decisions is quietly wrong, even when the model ranks patients perfectly.

Accuracy and ranking scores cannot see this failure. A model can order every case correctly while overstating each confidence — great ROC-AUC, useless probabilities.

How it works

Group predictions into bins by claimed confidence, then compare each bin's claim with what actually happened:

claimed      what happened     verdict

~10%            9%             honest
~30%           14%             overconfident here
~50%           52%             honest
~90%           96%             underconfident here

Plot claimed against happened and you get a reliability diagram. A truthful model draws the diagonal — claimed equals happened everywhere. Bulges below the diagonal mean overconfidence; bulges above mean underconfidence.

One number summarises the picture: the expected calibration error, or ECE — the average promise-versus-reality gap, with busier bins counting for more. Zero is perfect; 0.05 means claims are off by five percentage points on a typical prediction.

A real example you have seen

Weather apps live this daily, and the good ones are genuinely calibrated — of all their "60% rain" forecasts, close to 60% get rain. That is why the number feels trustworthy enough to plan a wedding around. Nobody would use an app whose "90%" delivered rain half the time.

Remember this

  • Calibration asks whether confidence numbers are honest, separate from accuracy.
  • Reliability diagram: claimed versus happened, bin by bin; the diagonal is truth.
  • ECE: the average gap, one number, weighted by how often each claim is made.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU; last-digit drift elsewhere is normal.

Auditing a random forest's honesty

Random forests rank well but their probabilities drift from the diagonal — a perfect audit subject:

reliability_audit.py
import numpy as np
from sklearn.calibration import calibration_curve
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=6000, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)
p = RandomForestClassifier(random_state=0).fit(Xtr, ytr).predict_proba(Xte)[:, 1]

happened, said = calibration_curve(yte, p, n_bins=10)
for s, h in zip(said, happened):
    print(f"model said {s:.2f}  ->  it happened {h:.2f}")

def ece(probs):
    """Average promise-vs-reality gap, weighted by how full each bin is."""
    bins = np.clip((probs * 10).astype(int), 0, 9)
    return sum((bins == b).mean() * abs(yte[bins == b].mean() - probs[bins == b].mean())
               for b in range(10) if (bins == b).any())

logistic = LogisticRegression(max_iter=1000).fit(Xtr, ytr).predict_proba(Xte)[:, 1]
print(f"forest   expected calibration error: {ece(p):.3f}")
print(f"logistic expected calibration error: {ece(logistic):.3f}")
Output
model said 0.03  ->  it happened 0.02
model said 0.14  ->  it happened 0.08
model said 0.25  ->  it happened 0.09
model said 0.36  ->  it happened 0.18
model said 0.45  ->  it happened 0.54
model said 0.56  ->  it happened 0.56
model said 0.67  ->  it happened 0.72
model said 0.76  ->  it happened 0.89
model said 0.87  ->  it happened 0.92
model said 0.95  ->  it happened 0.98
forest   expected calibration error: 0.035
logistic expected calibration error: 0.020

The walkthrough

Read the table like the forecaster audit. When this forest says 25%, reality delivers 9% — it overstates modest risks by nearly 3x. When it says 76%, reality delivers 89% — it understates strong cases. The pattern (stretched away from the extremes toward the middle) is the classic forest signature: averaging many trees pulls probabilities toward the centre.

A decision built on these numbers inherits the lies. Price an offer for customers above "25% churn risk" and you target a group whose true risk is 9% — the threshold formula computes nonsense from dishonest inputs. Meanwhile this model's ranking is excellent; accuracy dashboards show nothing wrong.

calibration_curve bins predictions and returns both sides of the audit. Ten equal-width bins here; strategy="quantile" instead gives equal-population bins — better when predictions crowd one region.

The ece helper makes the definition concrete: per bin, weight (fraction of predictions in the bin) times gap (mean outcome minus mean claimed). The forest's 0.035 means a typical confidence is off by about 3.5 points. There is no universal "good" ECE. The comparison is what informs you: the same audit gives 0.020 for a plain logistic model on this data, so the forest is roughly twice as dishonest. Judge that against your decision's sensitivity.

Common mistakes

Auditing on training data. Models are most overconfident about data they memorised. Calibration is measured on held-out data, full stop.

Trusting one binning. With 10 bins and 3,000 points, sparse bins wobble. Recompute with 15 and with quantile bins; conclusions that survive rebinning are real. Small audit sets need wider error margins — a bin of 40 predictions cannot pin its rate to two decimals.

Concluding the model is good because ECE is low. A model that predicts the base rate for everyone is perfectly calibrated and perfectly useless. Calibration and discrimination are separate virtues; check ranking quality alongside.

Fixing miscalibration by retraining bigger. The cure is cheaper: a small mapping learned on validation data straightens the curve — Platt scaling next, isotonic after.

Try it yourself

Run the same audit on LogisticRegression and compare tables — near-diagonal, ECE about a third of the forest's. Then re-audit the forest with n_bins=20 and strategy="quantile" and check the 0.035 holds up.

What to learn next

Researcher — Mathematics and papers.

Definitions

Perfect calibration for a binary probabilistic predictor $\hat{p}(x)$:

$$ P\big(y = 1 \mid \hat{p}(x) = q\big) = q \quad \text{for all } q \in [0, 1] $$

The binned estimator of expected calibration error over bins ${B_b}$:

$$ \widehat{ECE} = \sum_{b} \frac{|B_b|}{n} \, \big| \bar{y}(B_b) - \bar{\hat{p}}(B_b) \big| $$

Where:

  • $\bar{y}(B_b)$ — the empirical outcome rate in bin $b$.
  • $\bar{\hat{p}}(B_b)$ — the mean claimed probability in bin $b$.
  • $|B_b|/n$ — the bin's share of predictions.

The estimator is biased: binning smooths true deviations (downward bias) while finite samples inflate per-bin gaps (upward bias); the balance depends on bin count. Debiased and adaptive-bin estimators exist (Nixon et al., 2019; Kumar, Liang, Ma, 2019 — the latter shows binned ECE underestimates the calibration error of continuous-output models, and proposes scaling-binning to make it measurable).

Decompositions and history

The Brier score admits the Murphy (1973) decomposition:

$$ \text{Brier} = \underbrace{\text{reliability}}{\text{calibration gap}} - \underbrace{\text{resolution}}{\text{discrimination}} + \underbrace{\text{uncertainty}}_{\bar{y}(1-\bar{y})} $$

— calibration and discrimination as separable components of one proper score. Reliability diagrams themselves date to the forecasting-verification literature (Murphy and Winkler, 1977); DeGroot and Fienberg (1983) formalised calibration versus refinement. Meteorology solved cultural calibration decades before ML: operational precipitation forecasts are among the best-calibrated probabilistic predictions in any field.

Modern findings

  • Niculescu-Mizil and Caruana (2005), Predicting good probabilities with supervised learning: the classic map of which learners miscalibrate in which direction — boosting and SVMs sigmoid-distorted, bagged trees centre-pulled (the developer block's pattern), logistic regression near-honest.
  • Guo et al. (2017), On calibration of modern neural networks: depth, width and less regularisation worsened calibration even as accuracy improved; temperature scaling — one scalar — repairs most of it. The paper that revived the field.
  • Minderer et al. (2021) re-ran the audit on newer architectures: modern vision transformers are better calibrated than the ResNet era, cautioning against "deep nets are miscalibrated" as a timeless law.
  • Multiclass ECE variants (top-label, classwise, KDE-based) differ materially; Vaicenavicius et al. (2019) treat the estimation problem rigorously.
  • Calibration under distribution shift degrades before accuracy does (Ovadia et al., 2019) — monitoring ECE in production catches drift that accuracy metrics miss.

Proper-score-based checks (log-loss and Brier on held-out data) complement diagram-based audits: they are unbinned and decision-relevant, but conflate calibration with discrimination — which is why both views belong in an evaluation report.

What to learn next

What to learn next

These follow on from what you just read.

  • Calibration and Uncertainty

    Platt scaling

    Platt scaling fixes a model's dishonest confidence by fitting a tiny two-number translator on held-out data — the model's ranking stays untouched while its probabilities start meaning something.

  • Calibration and Uncertainty

    Isotonic calibration

    Isotonic calibration redraws a model's confidence scale as a free-form rising staircase — more flexible than Platt's fixed S-curve, and hungrier for data because of it.

  • Calibration and Uncertainty

    Proper scoring rules

    A proper scoring rule is a grading scheme under which honest probabilities earn the best expected grade — hedging and bluffing both lose, which accuracy cannot promise.