Calibration and Uncertainty

Platt scaling

Platt scaling fixes a model's dishonest confidence by fitting a tiny two-number translator on held-out data — the model's ranking stays untouched while its probabilities start meaning something.

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Platt scaling learns a small conversion curve for a model's exaggerated confidence.

It is like learning a friend's exaggeration factor and correcting for it.

Everyone has that friend. "Five minutes away" means twenty. "Almost done" means half done. After enough dinners, you carry a private conversion chart in your head. His messages become perfectly useful. Not because he changed, but because you learned his exact style of exaggeration.

Platt scaling builds that conversion chart for a model.

Why this matters

The reliability audit often convicts a model of consistent dishonesty: its "90%" events happen 99% of the time, its "20%" events almost never. The ranking is fine — higher scores do mean more likely — but the numbers themselves are bent.

Retraining a bigger model to fix this is expensive and often fails. The cheap observation: if the dishonesty is consistent, a fixed translation repairs it. Feed the model's raw score through a learned correction curve, and out comes an honest probability. The model itself is never touched.

How it works

raw score from model  ->  conversion curve  ->  honest probability

        0.95         ->        ~~ S ~~       ->      0.83
        0.50         ->      learned from    ->      0.55
        0.10         ->      held-out data   ->      0.24

Platt scaling chooses one specific shape for the curve: a smooth S-shape controlled by two dials. One dial squashes or stretches confidence overall. The other shifts it up or down. Fitting means finding the two numbers that make the translated probabilities match reality on a set of examples the model has never seen.

Why held-out examples? The model is most deluded about its own training data — it saw the answers. A translator fitted there would learn to trust the delusion. Fresh data shows the model's honest error rates, which is exactly what the curve must learn.

Two dials is a tiny budget, and that is the method's character. It can fix a consistently bent confidence scale beautifully. It cannot fix weirder, lumpier dishonesty. For that, the next lesson drops the fixed shape.

A real example you have seen

Spam filters were the classic customer. The strongest early filters produced scores with no probability meaning, yet mail systems wanted "99% spam" to justify auto-deleting. A Platt layer on top turned raw scores into percentages honest enough to act on. That is the exact recipe still used behind many "confidence" numbers you see in products.

Remember this

  • Platt scaling is a two-number S-curve translator on top of a frozen model.
  • It is fitted on held-out data, where the model's true error rates show.
  • Ranking never changes — only the honesty of the numbers.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU; last-digit drift elsewhere is normal.

Repairing Naive Bayes, the classic overconfident model

Naive Bayes assumes features do not interact, and pays with wildly overconfident probabilities. Watch the repair:

platt_repair.py
import numpy as np
from sklearn.calibration import CalibratedClassifierCV
from sklearn.datasets import make_classification
from sklearn.metrics import brier_score_loss
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import GaussianNB

X, y = make_classification(n_samples=6000, n_informative=8, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)

raw = GaussianNB().fit(Xtr, ytr)
p_raw = raw.predict_proba(Xte)[:, 1]
print("Brier, raw Naive Bayes:", round(brier_score_loss(yte, p_raw), 4))
print("predictions above 0.99 or below 0.01:", round(((p_raw > 0.99) | (p_raw < 0.01)).mean(), 2))

platt = CalibratedClassifierCV(GaussianNB(), method="sigmoid", cv=5).fit(Xtr, ytr)
p_cal = platt.predict_proba(Xte)[:, 1]
print("Brier, after Platt:    ", round(brier_score_loss(yte, p_cal), 4))
print("accuracy unchanged:", (raw.predict(Xte) == platt.predict(Xte)).mean())
Output
Brier, raw Naive Bayes: 0.0617
predictions above 0.99 or below 0.01: 0.42
Brier, after Platt:     0.0613
accuracy unchanged: 0.984

The walkthrough

42% of raw predictions claim near-certainty. Naive Bayes multiplies per-feature evidence as if each feature were fresh news; correlated features get counted repeatedly, and confidence compounds toward 0 or 1. That 0.42 is the smoking gun — real uncertainty rarely permits certainty on 4 predictions in 10.

The Brier score improved modestly, from 0.0617 to 0.0613. The Brier score is the mean squared gap between claimed probability and outcome — a metric that punishes dishonest confidence, covered properly in proper scoring rules. The gain looks small because Brier mixes calibration with ranking quality, and ranking already dominated. The calibration component itself improves far more — rerun the reliability audit on both and compare tables.

method="sigmoid" is Platt scaling inside CalibratedClassifierCV. With cv=5, the training set is split five ways: each fold's model scores data it never saw, those honest scores fit the S-curve, and the five results are averaged. Nothing touches the test set.

The last line shows 98.4% of hard predictions agree — not 100%, because the S-curve's crossing point sits slightly off the raw 0.5 boundary. Ranking within each fold's model is preserved exactly; the tiny label disagreement is the threshold shifting, which is a calibration feature.

scikit-learn version note. The wrapped-model argument is named estimator since 1.2 (base_estimator was removed in 1.4); tutorials older than that break on 1.7. To calibrate an already-trained frozen model, wrap it in FrozenEstimator (added in 1.6) instead of the removed cv="prefit" pattern.

Common mistakes

Calibrating on training data. The model's training scores are deluded, so the curve learns to endorse delusion. Always cross-validation folds or a dedicated calibration split.

Expecting Platt to fix any curve shape. Two parameters draw one S-shape. If the reliability diagram shows a zigzag or one-sided bend, the S cannot follow it — that is isotonic territory.

Judging the repair by accuracy. Accuracy compares against one threshold and barely moves. Judge with Brier, log-loss, or ECE — the metrics that see probability quality.

Calibrating a model whose ranking is broken. Translation cannot reorder. If low-risk cases score above high-risk ones, no monotone curve helps; fix the model first.

Try it yourself

Compute ECE (the loop from the reliability lesson) for p_raw and p_cal — the improvement is far more dramatic than the Brier deltas suggest. Then print both models' probabilities for the same ten test rows and watch near-certain claims relax toward honesty.

What to learn next

Researcher — Mathematics and papers.

The method

Platt (1999), Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Given a frozen scorer $f(x)$, fit:

$$ P(y = 1 \mid x) = \frac{1}{1 + \exp(A f(x) + B)} $$

Where:

  • $f(x)$ — the raw decision score (SVM margin in the original; any monotone score in general).
  • $A < 0$ — slope: how sharply score maps to probability (the "exaggeration factor").
  • $B$ — offset: where the 50% point sits.

$A$ and $B$ are fitted by maximum likelihood (logistic loss) on held-out scores — the procedure is exactly a one-feature logistic regression taking $f(x)$ as input. Platt's paper adds two details often skipped: out-of-sample scores via cross-validation (the cv=5 machinery), and target smoothing — replacing labels ${0,1}$ with $\frac{N_-+2}{...}$-style soft targets — a Bayesian shrinkage that prevents the fitted map from saturating; scikit-learn implements it.

When the sigmoid is the correct shape

Platt scaling is well-specified when class-conditional score distributions are Gaussian with equal variance: then the true posterior is exactly a sigmoid in the score. Kull, Silva Filho and Flach (2017) generalise: for scores already living in $[0,1]$ with Beta-distributed class-conditionals, the right family is beta calibration — three parameters, able to fix S-shaped and inverse-S miscalibration, containing Platt as a special case. Their empirical result: beta calibration dominates sigmoid on most classifiers whose outputs are probabilities rather than margins.

Temperature scaling, the deep-learning descendant

Guo et al. (2017) restrict Platt's map for multiclass networks: divide logits by a single learned temperature $T$ before softmax ($A = 1/T$, $B = 0$, applied per-logit). One parameter cannot change the argmax, so accuracy is exactly preserved — and it repairs most of modern networks' overconfidence. Limits documented since: it calibrates marginally, not per-class or per-group (Kull et al., 2019 propose Dirichlet calibration for multiclass; Ovadia et al., 2019 show temperature fitted in-distribution fails under shift).

Sample complexity and comparisons

Two parameters need little data — Platt behaves well from a few hundred calibration points, versus roughly a thousand-plus for isotonic regression; this is the practical selection rule, confirmed across learners in Niculescu-Mizil and Caruana (2005). Their taxonomy of which models need which repair remains the field's working map: boosted models and SVMs show sigmoid-shaped distortion (Platt's home turf), Naive Bayes shows heavier-tailed distortion that isotonic handles better given data, bagged trees sit between.

The verification-theory footnote: any post-hoc calibrator trades a little sharpness (concentration of predictions near 0 and 1) for calibration; the aim, per Gneiting, Balabdaoui and Raftery (2007), is maximising sharpness subject to calibration — sharpening beyond honesty is exactly the disease being cured.

What to learn next