Calibration and Uncertainty

Conformal prediction

Conformal prediction wraps any trained model and returns prediction sets that contain the true answer a promised fraction of the time — a guarantee bought with one held-out calibration set.

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Conformal prediction turns a model's single guesses into sets of plausible answers.

The sets are sized so the true answer falls inside them a promised share of the time. That holds no matter how flawed the model is.

A tailor stitches school uniforms from measurements taken months ago, so his cutting always runs a little off. Before promising anything, he checks his last hundred deliveries and measures how far off each one ran. Ninth-tenths of them erred within two centimetres. From now on he promises: "the fit will be within two centimetres." He is right about 90% of the time, whatever his cutting habits are. The promise was measured from his own track record.

Conformal prediction is that track-record trick, applied to any model.

Why this matters

Models hand you one answer with no honest error bars. Calibration repairs the probability numbers, but its promises are checked in bulk and lean on the repair holding up. Conformal prediction makes a stronger, stranger promise. Pick a target, say 90%. For each new case it then produces a prediction set: several possible labels, or a range of values. That set contains the truth 90% of the time. The guarantee holds even when the underlying model is mediocre, because it never trusts the model — only the measured track record.

The price of the promise is paid in set size. A weak model keeps its 90% pledge by returning big, vague sets. A strong model keeps it with sharp ones. Honesty is guaranteed; usefulness still depends on the model.

How it works

Set aside a calibration set — labelled examples the model never trained on. That is the track record.

1. score each calibration example:      how badly did the model
                                        miss the true answer?
2. find the miss size that covers       "90% of misses were
   90% of the track record         ->    smaller than THIS bar"
3. for a new case: include every        the prediction set
   answer within that bar of miss

For a new case, include every candidate answer the model would miss by less than the bar. An easy case yields a set of one. A confusing case yields three or four candidates — and that widening is the model admitting confusion, case by case.

One honest caveat: the promise is an average over many predictions. It is not "this particular set is 90% safe". Slow drift in your data also erodes it. The tailor's track record stops binding if he changes his scissors.

A real example you have seen

Medical AI triage increasingly reports exactly this way: instead of "diagnosis: X", the tool shows "consistent with X or Y — 90% coverage". Doctors rule out the rest and focus. A set of one means the machine is effectively sure; a set of four tells the radiologist this scan deserves their full attention.

Remember this

  • Conformal wraps any model and promises coverage — truth inside the set a chosen % of the time.
  • The promise comes from a held-out track record, not from trusting the model.
  • Weak models pay in bigger sets, never in broken promises.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU; last-digit drift elsewhere is normal.

Split conformal for a 3-class problem, from scratch

The whole method is a quantile of held-out scores — few enough lines to own outright:

conformal_sets.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=6000, n_classes=3, n_informative=6, random_state=0)
Xtr, Xrest, ytr, yrest = train_test_split(X, y, test_size=0.4, random_state=0)
Xcal, Xte, ycal, yte = train_test_split(Xrest, yrest, test_size=0.5, random_state=0)

model = LogisticRegression(max_iter=1000).fit(Xtr, ytr)

# Score = how much probability the model FAILED to give the true class.
cal_p = model.predict_proba(Xcal)
scores = 1 - cal_p[np.arange(len(ycal)), ycal]

n = len(scores)
q = np.quantile(scores, np.ceil(0.9 * (n + 1)) / n, method="higher")
print(f"cut-off score: {q:.3f}")

test_p = model.predict_proba(Xte)
sets = test_p >= 1 - q                    # keep every class that clears the bar
covered = sets[np.arange(len(yte)), yte].mean()
print(f"true class inside the set: {covered:.1%} (promised: 90%)")
print(f"average set size: {sets.sum(axis=1).mean():.2f} classes")
print("sizes seen:", dict(zip(*np.unique(sets.sum(axis=1), return_counts=True))))
Output
cut-off score: 0.802
true class inside the set: 89.8% (promised: 90%)
average set size: 1.76 classes
sizes seen: {1: 393, 2: 701, 3: 106}

The walkthrough

The promise held: 89.8% against a 90% target. No tuning, no luck — the guarantee is arithmetic. The model's probabilities were never assumed calibrated; only their ranking of misses on the calibration set mattered.

Three splits, three jobs. Training data fits the model. Calibration data — 1,200 points the model never saw — builds the track record. Test data audits the promise. Reusing training data as calibration silently voids the guarantee, the same leakage sin as always.

The miss score 1 - p(true class) is the simplest sensible choice: 0 when the model gave the truth full probability, near 1 when it starved the truth. Any miss measure works — the guarantee never depends on the choice, only the set sizes do.

The odd-looking quantile level ceil(0.9 * (n+1)) / n with method="higher" is the finite-sample correction — at n=1,200 it takes the quantile a touch above 0.9. Skip it and coverage lands slightly under target on average.

Read the size histogram as an uncertainty report. 393 easy cases got a single class; 106 hard ones got all three — the model formally shrugging. Per-case honesty like this is what plain point predictions cannot express.

Common mistakes

Calibrating on training data. The model's misses on data it memorised are unrealistically small, the bar comes out too low, and real coverage falls short of the promise. The calibration set must be untouched.

Reading 90% as per-prediction. The guarantee is marginal — across many predictions on exchangeable data. Any single set either contains the truth or does not, and subgroups (one class, one region) can run below 90% while the average holds.

Using a tiny calibration set. The quantile from 50 points is noisy, and realised coverage wobbles several points around target. Aim for 500 to 1,000; the guarantee is exact in expectation at any size, but the variance is yours to manage.

Expecting the promise to survive drift. Exchangeability between calibration and future data is the entire foundation. Under drift, recalibrate on recent data — cheap, since the model itself never refits.

Try it yourself

Change the target to 99% and rerun: watch the cut-off rise and the size histogram lurch toward 3-class sets — certainty priced in vagueness. Then swap the model for a RandomForestClassifier and check the promise still holds with sharper sets.

What to learn next

Researcher — Mathematics and papers.

The split conformal theorem

Given exchangeable pairs $(X_i, Y_i)$, a fitted model, and nonconformity scores $s_i = s(X_i, Y_i)$ on a calibration set of size $n$, define $\hat{q}$ as the $\lceil (n+1)(1-\alpha) \rceil / n$ empirical quantile of the scores. The prediction set for a new $X_{n+1}$:

$$ C(X_{n+1}) = {\, y : s(X_{n+1}, y) \le \hat{q} \,} \qquad \Rightarrow \qquad P\big(Y_{n+1} \in C(X_{n+1})\big) \ge 1 - \alpha $$

Where:

  • $\alpha$ — the tolerated miss rate (0.1 for 90% coverage).
  • $s(\cdot,\cdot)$ — any measurable nonconformity function; the guarantee is assumption-free about the model and score.
  • Exchangeability — the only distributional assumption: calibration and test pairs are jointly permutation-invariant (i.i.d. suffices).

Proof sketch: under exchangeability, the rank of $s_{n+1}$ among all $n+1$ scores is uniform; the truth escapes the set only when its score ranks in the top $\alpha$ fraction, corrected for finiteness. With continuous scores there is also an upper bound: coverage $\le 1 - \alpha + \frac{1}{n+1}$ — the method is not conservative in any meaningful way.

Origins: Vovk, Gammerman and Shafer, Algorithmic Learning in a Random World (2005) — full (transductive) conformal prediction, which refits per candidate label and needs no calibration split at the price of massive compute. Split (inductive) conformal (Papadopoulos et al., 2002) is the practical form above. Angelopoulos and Bates (2021), A gentle introduction to conformal prediction, is the standard modern tutorial.

Better scores, better sets

The score determines efficiency, never validity:

  • APS (Romano, Sesia, Candès, 2020): score by cumulative probability mass down the sorted class list — sets adapt to the full probability profile, improving conditional behaviour.
  • RAPS (Angelopoulos et al., 2021): APS plus a rank penalty, shrinking sets on ImageNet-scale problems.
  • CQR (Romano et al., 2019) conformalises quantile regression for regression intervals — the subject of the next lesson.
  • Class-conditional (Mondrian) conformal runs the calibration per group, buying per-class validity at the cost of per-group calibration data (Vovk, 2012).

Limits and frontier

  • Conditional coverage is impossible in full generality: distribution-free $P(Y \in C(X) \mid X = x) \ge 1-\alpha$ for all $x$ forces uninformative, infinitely wide sets (Vovk, 2012; Lei and Wasserman, 2014; Barber et al., 2020). Everything practical is an approximation between marginal and conditional.
  • Beyond exchangeability: weighted conformal handles covariate shift with known likelihood ratios (Tibshirani et al., 2019); nonexchangeable conformal (Barber et al., 2023) degrades gracefully under drift with quantified slack; adaptive conformal inference (Gibbs and Candès, 2021) tracks online distribution shift.
  • Risk control generalisations: conformal risk control (Angelopoulos et al., 2022) extends the quantile trick from miscoverage to any monotone loss — false-negative-rate control in medical screening being the flagship use.
  • Full conformal costs $O(|\mathcal{Y}|)$ model refits per prediction; split conformal costs one quantile — the trade underneath every deployment decision. Jackknife+ and CV+ (Barber et al., 2021) sit between, recycling cross-validation residuals with $1 - 2\alpha$ worst-case guarantees.

What to learn next