Calibration and Uncertainty

Isotonic calibration

Isotonic calibration redraws a model's confidence scale as a free-form rising staircase — more flexible than Platt's fixed S-curve, and hungrier for data because of it.

Read these first

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Isotonic calibration corrects a model's probabilities using a staircase of any shape.

It promises only one thing: a higher raw score never turns into a lower probability.

A strict examiner marks a class, and his marks are bent in a strange way. Brutal in the middle range, generous at the top, odd in patches. You trust one thing about him — a student he marked higher genuinely did better. To convert his marks into fair percentages, you would not force one fixed formula onto them. You would redraw the whole scale piece by piece. Each stretch of marks finds its own honest level. A higher mark is never allowed to land below a lower one.

That order-respecting redraw is isotonic calibration. "Isotonic" means order-preserving — the one promise the staircase keeps.

Why this matters

Platt scaling buys safety with rigidity: two numbers, one S-shape. When a model's dishonesty follows that shape, perfect. When the reliability diagram shows a lopsided bend, a zigzag, or honesty in one region with delusion in another, the S physically cannot follow the damage.

Isotonic calibration drops the fixed shape. It fits a staircase: flat stretches joined by upward steps, placed wherever the data demands. Any rising pattern of dishonesty can be traced exactly.

How it works

honest probability
   |                            ____
   |                      _____|        <- staircase: rises in
   |               ______|                 free-form steps,
   |          ____|                        never dips down
   |      ___|
   |_____|
   +--------------------------------
        raw model score ->

Flexibility has a price, and the price is data. A two-number S-curve can be pinned down by a few hundred examples. A staircase with dozens of free steps can memorise a few hundred examples instead — each little accident of the calibration set becomes a step. Fed on scraps, isotonic produces a jagged staircase that fits the past and lies about the future.

The working rule of thumb: below roughly a thousand calibration examples, prefer the S-curve. Above, the staircase usually wins, because real dishonesty rarely follows one textbook shape.

A real example you have seen

Ad systems predict click chances at enormous scale, and tiny probability biases turn directly into money misallocated across millions of auctions. With billions of past examples, data hunger is no obstacle — flexible staircase-style calibration is standard machinery in that world.

Remember this

  • Isotonic fits a free-form rising staircase; only the ordering is sacred.
  • It fixes bends that Platt's fixed S-curve cannot reach.
  • It needs roughly 1,000+ examples; on less, it memorises noise. Prefer Platt there.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU; last-digit drift elsewhere is normal.

Platt versus isotonic, starved and well-fed

Same overconfident Naive Bayes as the previous lesson, calibrated on 300 examples, then on 20,000. Brier score judges (lower is better):

platt_vs_isotonic.py
import numpy as np
from sklearn.calibration import CalibratedClassifierCV
from sklearn.datasets import make_classification
from sklearn.metrics import brier_score_loss
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import GaussianNB

X, y = make_classification(n_samples=23000, n_informative=8, random_state=0)
Xte, Xpool, yte, ypool = train_test_split(X, y, train_size=3000, random_state=0)

for n in (300, 20000):
    Xtr, ytr = Xpool[:n], ypool[:n]
    for method in ("sigmoid", "isotonic"):
        cal = CalibratedClassifierCV(GaussianNB(), method=method, cv=5).fit(Xtr, ytr)
        b = brier_score_loss(yte, cal.predict_proba(Xte)[:, 1])
        print(f"n={n:5d}  {method:8s}  Brier = {b:.4f}")
Output
n=  300  sigmoid   Brier = 0.0572
n=  300  isotonic  Brier = 0.0570
n=20000  sigmoid   Brier = 0.0538
n=20000  isotonic  Brier = 0.0529

The walkthrough

At n=20,000 isotonic wins outright: 0.0529 against sigmoid's 0.0538. Naive Bayes' distortion is not a clean S, and with ample data the staircase traces the true shape the sigmoid cannot.

At n=300 they tie — 0.0570 versus 0.0572, inside noise. With cv=5, each staircase is fitted on only 240 points, and its overfitting eats the flexibility advantage. The textbook prediction is that isotonic falls well behind here; on this easy, low-noise data the tie is as far as starvation goes. The reliable pattern is the n=20,000 row: flexibility pays exactly when data is plentiful.

method="isotonic" is the entire code change. Under the hood, scikit-learn's IsotonicRegression fits the staircase to (score, outcome) pairs from each cross-validation fold, exactly as Platt fitting did with the S-curve.

Watch the edges of the staircase. Isotonic maps everything below the lowest calibration score to one constant, and everything above the highest to another — frequently exactly 0.0 or 1.0. Downstream code taking log(p) will meet infinities. Clip outputs into [0.001, 0.999] before logging.

Common mistakes

Reaching for isotonic because it is "more powerful". On a few hundred calibration points the power turns against you. Check your calibration-set size first; the previous lesson's two-parameter map is the safe default.

Fitting the staircase on training predictions. Same leakage as always in calibration: the model's training scores flatter it, and the staircase learns the flattery. CalibratedClassifierCV exists to prevent exactly this; keep cv ≥ 2.

Believing a probability of exactly 1.0. That is a staircase edge, not evidence of certainty. The clip-before-log advice above is not cosmetic — one production log(0) pays for the whole lesson.

Assuming the model's dishonesty is monotone. Isotonic preserves ranking by design. If the raw scores genuinely rank backwards in some region — probability 0.3 cases outperforming probability 0.5 cases — no order-preserving repair can fix it, and the staircase will flatten that whole region instead. That is a modelling bug upstream, visible as a non-monotone reliability diagram.

Try it yourself

Extract one fitted staircase and look at it directly: fit IsotonicRegression(out_of_bounds="clip") on (p_raw, y) from a held-out split, then plot its predictions over a grid of 100 scores. Count the steps at n=300 versus n=20,000 — you will see the memorisation shrink as data grows.

What to learn next

Researcher — Mathematics and papers.

The optimisation problem

Given calibration pairs $(s_i, y_i)$ sorted by score, isotonic regression solves:

$$ \min_{g} \sum_{i=1}^{n} \big(y_i - g(s_i)\big)^2 \quad \text{s.t.} \quad g(s_i) \le g(s_j) \text{ whenever } s_i \le s_j $$

Where:

  • $s_i$ — the raw model scores; $y_i \in {0, 1}$ the outcomes.
  • $g$ — the fitted calibration map, constrained monotone non-decreasing.

The solution is unique and piecewise-constant — the staircase — computed by the pool adjacent violators algorithm (PAV): walk the sorted sequence, merging each adjacent pair violating monotonicity into a pooled block whose value is the block average, repeating until clean. Runs in $O(n)$ after the $O(n \log n)$ sort (Ayer et al., 1955; Barlow et al., 1972, for the order-restricted-inference foundations).

A useful identity: the PAV solution equals the left derivative of the greatest convex minorant of the cumulative sum diagram — which is why the fitted values are also the calibrated conditional means given the ordering constraint, regardless of the loss being squared or Bernoulli likelihood (they coincide under monotonicity).

Statistical properties

  • Consistency: if the true regression $\mathbb{E}[y \mid s]$ is monotone, isotonic converges at rate $n^{-1/3}$ (pointwise, Chernoff-type limits) — slower than parametric $n^{-1/2}$, the formal version of "data-hungry".
  • The number of steps grows like $O(n^{1/3})$ in probability, quantifying how memorisation shrinks with data.
  • Zadrozny and Elkan (2002), Transforming classifier scores into accurate multiclass probability estimates, introduced PAV calibration to ML and the one-vs-rest-then-normalise multiclass extension.
  • Niculescu-Mizil and Caruana (2005) supply the empirical selection rule quoted everywhere: sigmoid below ~1,000 calibration points, isotonic above — with the caveat that the crossover depends on how non-sigmoid the distortion is.

Refinements

  • Smoothing: the raw staircase's discontinuities and 0/1 edge values motivate smoothed isotonic variants and Platt-on-isotonic hybrids; IRoVa (isotonic regression over Venn predictors) and Venn–Abers predictors (Vovk and Petej, 2014) produce interval-valued, provably calibrated outputs from PAV machinery, connecting to conformal prediction.
  • Scaling–binning (Kumar, Liang, Ma, 2019) combines a parametric fit with binning to get measurable calibration error with isotonic-like flexibility.
  • For multiclass beyond one-vs-rest: Dirichlet calibration (Kull et al., 2019) is the parametric competitor; per-class isotonic with normalisation remains the strong nonparametric baseline.
  • Under distribution shift, any calibrator fitted in-distribution decays; label-shift corrections re-weight the calibration pairs (Alexandari et al., 2020) rather than refitting the shape.

What to learn next