Calibration and Uncertainty

Choosing a threshold from costs

The 0.5 cut-off is a default nobody chose on purpose — when the two mistakes have different prices, the cheapest threshold follows from those prices, and it is rarely 0.5.

On this page 5
  1. Why this matters
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A classifier's threshold is a business decision in disguise.

The right cut-off depends on the price of each kind of mistake. The default of 0.5 is almost never that cut-off.

Think of a smoke alarm's sensitivity knob. Turned high, it screams at burnt toast — annoying, cheap. Turned low, it sleeps through a real fire — quiet, catastrophic. The right setting depends on one comparison: how bad is a false scream versus a missed fire? Nobody sane sets it "in the middle" for symmetry.

Your model's threshold is that knob.

Why this matters

Models output a score between 0 and 1, and someone converts it into an action: flag or pass, call or ignore, treat or wait. The lazy conversion — act when the score crosses 0.5 — encodes a silent assumption: both mistakes cost the same. They almost never do.

A missed fraud costs the bank thousands; a false alarm costs one verification SMS. A missed cancer costs a life; a false alarm costs a follow-up scan. Whenever mistake prices differ by 10x or 100x, the mid-point threshold quietly burns money — or worse.

How it works

Write down the price of each mistake. Then, for any candidate threshold, count what your model's mistakes would cost at that setting, and pick the cheapest.

threshold  false alarms  missed cases   total cost
   high         few          many       heavy  (misses are pricey)
   0.5         some          some       still heavy
   low         many          few        cheapest!   <- when misses cost 24x
   too low     floods        none       alarm-flood costs take over

Here is the pattern that surprises people. When misses are far pricier than false alarms, the best threshold drops low. The model should alarm even on weak suspicion. In the opposite case (alarms expensive, misses tolerable), it climbs high. The threshold follows the prices, mechanically.

The two prices themselves come from the business, not the data. Estimating them roughly is fine. Being 30% wrong about a price moves the best threshold a little. Ignoring prices entirely means using a threshold that answers no question at all.

A real example you have seen

Your bank blocks a suspicious card payment and sends "was this you?" over SMS. That system alarms on faint suspicion. Its threshold is set low on purpose. A fraud slipping through costs the bank real money; a false block costs one moment of your patience.

Remember this

  • Threshold 0.5 assumes equal mistake prices — a claim, not a neutral default.
  • Price the two mistakes, then pick the threshold with the lowest total cost.
  • Expensive misses push the threshold down; expensive alarms push it up.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU; last-digit drift elsewhere is normal.

Pricing the mistakes on a churn model

A telecom model predicts which customers will leave. A retention call costs ₹50; a lost customer costs ₹1,200:

cost_threshold.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=4000, weights=[0.9], random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, stratify=y, random_state=0)
p = LogisticRegression().fit(Xtr, ytr).predict_proba(Xte)[:, 1]

COST_FP, COST_FN = 50, 1200      # rupees: one retention call vs one lost customer

def cost_at(t):
    pred = (p >= t).astype(int)
    fp = int(((pred == 1) & (yte == 0)).sum())
    fn = int(((pred == 0) & (yte == 1)).sum())
    return COST_FP * fp + COST_FN * fn

print("cost at threshold 0.50:", cost_at(0.50))
thresholds = np.round(np.arange(0.05, 1.0, 0.05), 2)
best = min(thresholds, key=cost_at)
print("cheapest threshold:", best, "with cost:", cost_at(best))
print("theory says:", round(COST_FP / (COST_FP + COST_FN), 3))
Output
cost at threshold 0.50: 146100
cheapest threshold: 0.05 with cost: 47800
theory says: 0.04

The walkthrough

The default threshold cost three times the right one. ₹1,46,100 versus ₹47,800 on the same model with the same predictions. Nothing about the model changed — only the conversion from score to action. This is the cheapest model improvement that exists.

The best threshold landed at 0.05, not near 0.5. With misses 24 times pricier than calls, the model should call anyone with even 5% churn risk. Counterintuitive until you reprice it: a ₹50 call that prevents a 5% chance of losing ₹1,200 is a good trade — expected saving ₹60.

The theory line is a one-line shortcut. For a model whose probabilities are honest, the optimal threshold is the false-alarm price divided by the sum of both prices: 50 / 1250 = 0.04. Our swept 0.05 matches to within the grid. That formula only works when the probabilities mean what they say — which is exactly why calibration is the next lesson.

The sweep is still worth running even with the formula in hand: it reveals how flat or sharp the cost curve is near the optimum. A flat valley means the exact threshold barely matters; a sharp one means revisit it whenever costs or data drift.

Common mistakes

Tuning the threshold on the test set. The threshold is a fitted parameter. Choose it on a validation split, then report cost on untouched test data — standard train/test discipline.

Optimising accuracy or F1, then wondering about money. Every metric implies some price ratio (accuracy implies equal prices; F1 implies a particular prevalence-dependent ratio nobody chose). If real prices exist, use them directly.

Forgetting capacity constraints. The cost-optimal threshold might imply 3,000 calls a day from a 10-person team. When capacity binds, choose the top-k by score instead — the precision-at-k view.

Setting prices once and never revisiting. The cost of a missed churn changes with margins and competition. Log the prices with the model version, and re-derive the threshold whenever either changes.

Try it yourself

Set COST_FP = 400 — pushy calls now annoy customers into leaving — and predict the new best threshold with the formula before rerunning the sweep. Then plot cost_at(t) for all thresholds and find where the curve is flat enough that a ±0.05 error is harmless.

What to learn next

Researcher — Mathematics and papers.

The Bayes-optimal threshold

With calibrated posterior $\eta(x) = P(y = 1 \mid x)$, costs $c_{FP}$ and $c_{FN}$ (and zero cost for correct decisions), the expected-cost-minimising rule is: predict positive iff

$$ \eta(x) \ge t^* = \frac{c_{FP}}{c_{FP} + c_{FN}} $$

Derivation in one line: alarming costs $(1 - \eta)\,c_{FP}$ in expectation, staying silent costs $\eta\, c_{FN}$; alarm when the first is smaller. With non-zero correct-decision utilities the same algebra yields $t^*$ from the four-entry utility matrix (Elkan, 2001, The foundations of cost-sensitive learning — which also shows class rebalancing is equivalent to threshold shifting for calibrated models).

Where:

  • $\eta(x)$ — the true conditional positive probability; a calibrated model's output estimates it.
  • $t^*$ — the optimal operating threshold, independent of prevalence given $\eta(x)$.

The empirical sweep in the developer block agrees with $t^*$ only because the logistic model is near-calibrated. Miscalibrated scores shift the empirical optimum away from the formula — the practical argument for calibrating first, then thresholding by formula, rather than sweeping per deployment.

Cost curves and dominance

Drummond and Holte (2006) replace ROC space with cost space: x-axis the probability-cost function $PC(+) = \pi c_{FN} / (\pi c_{FN} + (1-\pi) c_{FP})$, y-axis normalised expected cost. Each classifier-threshold pair is a line; the lower envelope over thresholds shows expected cost across all operating conditions at once. Two properties make it superior to ROC for deployment: dominance is visible as one envelope under another, and the regions of operating conditions where model A beats model B are read directly off the axis.

Hernández-Orallo, Flach and Ferri (2012) unify this with metrics: ROC-AUC equals expected cost under a particular distribution over thresholds, and the Brier score equals expected cost under a uniform distribution over cost ratios when thresholding by $t^*$ — a bridge between this lesson and proper scoring rules.

Beyond fixed, known costs

  • Example-dependent costs (Bahnsen et al., 2015): each transaction has its own $c_{FN}$ (the transaction amount). The rule becomes $\eta(x) \ge c_{FP}(x) / (c_{FP}(x) + c_{FN}(x))$ pointwise; fraud systems implement exactly this.
  • Uncertain costs: when prices are estimates, the flatness of the cost curve near $t^*$ (visible in the sweep) is the robustness margin; minimax-regret thresholds hedge across a cost interval.
  • Deferral / reject options (Chow, 1970): with an abstain action costing $c_r < \min(c_{FP}, c_{FN}) \cdot \text{(error prob)}$, the optimal policy thresholds twice, deferring the uncertain middle band to humans — the theoretical basis of human-in-the-loop triage.
  • Capacity-constrained deployment replaces thresholds with top-k selection; the Neyman–Pearson view (fix an FPR budget, maximise TPR) is the statistical twin.

What to learn next