Choosing the metric that matches the decision
A metric is a promise about what the business will feel, and the wrong one lets a useless model score brilliantly.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Choose the number that measures success only after writing down the decision the model will drive — the decision picks the metric, never the reverse.
Think of an exam with negative marking. The moment wrong answers cost a mark, you stop guessing blindly — same student, same knowledge, different behaviour. The scoring rule changed what "playing well" means.
Models are the same. Train and select them by one score, and you get a model shaped by that score. Pick a score that ignores what the business loses, and you get a model that ignores it too.
Why it exists
Accuracy — the fraction of answers that are right — sounds like the natural score. But when only 3 transactions in 100 are fraud, a model that says "all fine, always" is 97% accurate and 100% useless.
The failure is not the model. The failure is the metric. It priced a missed fraud and a false alarm as equal. One costs thousands; the other costs a phone call.
How it works
Work backwards from the decision:
decision: which claims get auto-approved?
mistakes: approve a bad claim → costs a payout
hold a good claim → costs a review + an annoyed customer
metric: total cost of mistakes, in rupees, per 1000 claimsTwo named mistake-rates matter constantly. Recall asks: of the things we should have caught, how many did we catch? Precision asks: of the things we flagged, how many deserved it? Every alarm system trades one against the other — and the right trade depends on the price of each mistake.
When mistakes have prices, use the prices. A metric in rupees needs no interpretation meeting.
A real example you have seen
A smoke alarm is tuned for recall: it must not miss a real fire, so it accepts screaming at burnt toast. Face-unlock on your phone is tuned for precision: letting a stranger in is far worse than asking you to try again. Same kind of system, opposite metric — because the decisions differ.
Remember this
- The decision and the price of each mistake choose the metric. Write both down first.
- Accuracy misleads whenever one class is rare or one mistake is pricier.
- A metric in business units (rupees, hours, missed patients) ends arguments that percentages start.
What to learn next
- Designing for a human reviewer — when the decision includes a person, the metric must too.
- Model evaluation — precision, recall, and their relatives, in full.
- Evaluating by slice, not by average — one metric can hide a failing subgroup.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7 and numpy 1.26.
The accurate model that loses money
One fraud model, three operating thresholds. Accuracy prefers one; the cost sheet prefers another.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import confusion_matrix
from sklearn.model_selection import train_test_split
# 3% of transactions are fraud, like a real payment stream.
X, y = make_classification(n_samples=20000, n_features=8, n_informative=6,
weights=[0.97], random_state=7)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)
probs = LogisticRegression(max_iter=1000).fit(X_tr, y_tr).predict_proba(X_te)[:, 1]
print(f"'never flag anything' accuracy: {(y_te == 0).mean():.3f}\n")
print("threshold accuracy missed reviews cost")
for t in [0.5, 0.2, 0.05]:
pred = (probs >= t).astype(int)
tn, fp, fn, tp = confusion_matrix(y_te, pred).ravel()
# a missed fraud costs 4000 rupees, a human review costs 50
cost = fn * 4000 + (tp + fp) * 50
acc = (pred == y_te).mean()
print(f"{t:9.2f} {acc:8.3f} {fn:6d} {tp+fp:7d} {cost:6d}")'never flag anything' accuracy: 0.965
threshold accuracy missed reviews cost
0.50 0.965 175 0 700000
0.20 0.965 170 12 680600
0.05 0.808 99 938 442900The walkthrough
At threshold 0.50 the model flags nothing. Zero reviews. Its accuracy exactly equals "never flag anything" — the model is deployed and doing no work. Accuracy cannot see the difference.
At 0.05 accuracy collapses to 0.808 — and cost falls by a third. The model now catches 76 more frauds at the price of 938 reviews. In rupees, the "worst" threshold by accuracy is the best decision on the table.
The threshold is a business dial, not a modelling detail. The default 0.5 in predict() encodes the assumption that both mistakes cost the same. Almost no real decision satisfies it. Score with predict_proba, then choose the threshold where the cost curve bottoms out.
Confusion-matrix names, since they appear everywhere: fn (false negatives) are the missed frauds; fp (false positives) are innocent transactions flagged; tp (true positives) are frauds caught.
Common mistakes
Reporting accuracy on imbalanced data. The stakeholder hears 96.5% and celebrates the do-nothing model. Report recall, precision, and cost at the operating threshold instead — model evaluation walks through each.
Optimising a metric no one can act on. A beautiful AUC (a threshold-free ranking score) means little if the business runs at one fixed threshold with a fixed review budget. Evaluate at the budget you will actually have.
One metric when you need a metric plus a guardrail. "Minimise cost, subject to: no more than 1,000 reviews per day, and recall above 50%." Real deployments are constrained optimisation; state the constraints or the model will spend them.
Changing the metric mid-project. Every metric change silently invalidates earlier comparisons. Log the metric definition with the same care as the code version.
Try it yourself
Re-price the mistakes: missed fraud ₹500, review ₹50 — a low-stakes product. Find the new best threshold. Then answer in one sentence: why does the best threshold rise when missed fraud gets cheaper?
What to learn next
- Designing for a human reviewer — when the decision includes a person, the metric must too.
- Model evaluation — precision, recall, and their relatives, in full.
- Evaluating by slice, not by average — one metric can hide a failing subgroup.
Researcher — Mathematics and papers.
Decision-theoretic foundation
Given action set $A$, outcomes $Y$, and a loss $L(a, y)$, the optimal decision rule given calibrated posteriors is:
$$ a^*(x) = \arg\min_{a \in A} \; \mathbb{E}_{y \sim P(Y \mid X=x)}\left[L(a, y)\right] $$
Where:
- $a^*(x)$ — the cost-minimising action for input $x$.
- $P(Y \mid X = x)$ — the model's predictive distribution, assumed calibrated.
- $L(a, y)$ — the loss of taking action $a$ when the truth is $y$.
For binary classification with false-positive cost $c_{FP}$ and false-negative cost $c_{FN}$, this reduces to thresholding the posterior at:
$$ t^* = \frac{c_{FP}}{c_{FP} + c_{FN}} $$
Elkan (2001), The Foundations of Cost-Sensitive Learning (IJCAI), derives this and its corollary: cost-sensitivity can live entirely in the threshold when probabilities are honest. The catch is the premise — miscalibrated scores break the formula, which is why calibration (reliability diagrams, Platt scaling, isotonic regression) is a scoping concern, not a finishing touch.
Proper scoring rules
Selecting models by a proper scoring rule — one minimised in expectation by the true probabilities, such as log loss or the Brier score — preserves the option to change thresholds later. Improper targets (accuracy at a fixed threshold) entangle the model with today's costs. Gneiting and Raftery (2007), Strictly Proper Scoring Rules, Prediction, and Estimation (JASA), is the canonical treatment.
Metrics as incentives: Goodhart effects
Once a metric selects models (or people), optimisation pressure exploits its gaps — Goodhart's law. ML-specific analyses: Manheim and Garrabrant (2018), Categorizing Variants of Goodhart's Law; reward hacking in RL (Amodei et al., 2016, Concrete Problems in AI Safety). The mitigation pattern is a small metric portfolio: one optimised metric, plus guardrail metrics that must not regress — the structure Ng formalises as optimising versus satisficing metrics (Machine Learning Yearning, 2018).
Offline metric versus online outcome
The offline metric is itself a proxy for a business outcome measured later by experiment. Correlation between offline gains and A/B results is imperfect and domain-specific; recommender-system studies report notable disagreement rates (e.g., Garcin et al., 2014; the broader debate in Jannach and Bauer, 2020). The scoping implication: pre-register the offline metric you expect to transfer, then verify the transfer on the first deployment rather than assuming it.
What to learn next
- Designing for a human reviewer — when the decision includes a person, the metric must too.
- Model evaluation — precision, recall, and their relatives, in full.
- Evaluating by slice, not by average — one metric can hide a failing subgroup.