Scoping an ML Project

Choosing the metric that matches the decision

A metric is a promise about what the business will feel, and the wrong one lets a useless model score brilliantly.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Choose the number that measures success only after writing down the decision the model will drive — the decision picks the metric, never the reverse.

Think of an exam with negative marking. The moment wrong answers cost a mark, you stop guessing blindly — same student, same knowledge, different behaviour. The scoring rule changed what "playing well" means.

Models are the same. Train and select them by one score, and you get a model shaped by that score. Pick a score that ignores what the business loses, and you get a model that ignores it too.

Why it exists

Accuracy — the fraction of answers that are right — sounds like the natural score. But when only 3 transactions in 100 are fraud, a model that says "all fine, always" is 97% accurate and 100% useless.

The failure is not the model. The failure is the metric. It priced a missed fraud and a false alarm as equal. One costs thousands; the other costs a phone call.

How it works

Work backwards from the decision:

decision:  which claims get auto-approved?
mistakes:  approve a bad claim   → costs a payout
           hold a good claim     → costs a review + an annoyed customer
metric:    total cost of mistakes, in rupees, per 1000 claims

Two named mistake-rates matter constantly. Recall asks: of the things we should have caught, how many did we catch? Precision asks: of the things we flagged, how many deserved it? Every alarm system trades one against the other — and the right trade depends on the price of each mistake.

When mistakes have prices, use the prices. A metric in rupees needs no interpretation meeting.

A real example you have seen

A smoke alarm is tuned for recall: it must not miss a real fire, so it accepts screaming at burnt toast. Face-unlock on your phone is tuned for precision: letting a stranger in is far worse than asking you to try again. Same kind of system, opposite metric — because the decisions differ.

Remember this

  • The decision and the price of each mistake choose the metric. Write both down first.
  • Accuracy misleads whenever one class is rare or one mistake is pricier.
  • A metric in business units (rupees, hours, missed patients) ends arguments that percentages start.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7 and numpy 1.26.

The accurate model that loses money

One fraud model, three operating thresholds. Accuracy prefers one; the cost sheet prefers another.

metric_vs_decision.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import confusion_matrix
from sklearn.model_selection import train_test_split

# 3% of transactions are fraud, like a real payment stream.
X, y = make_classification(n_samples=20000, n_features=8, n_informative=6,
                           weights=[0.97], random_state=7)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)

probs = LogisticRegression(max_iter=1000).fit(X_tr, y_tr).predict_proba(X_te)[:, 1]

print(f"'never flag anything' accuracy: {(y_te == 0).mean():.3f}\n")
print("threshold  accuracy   missed  reviews    cost")
for t in [0.5, 0.2, 0.05]:
    pred = (probs >= t).astype(int)
    tn, fp, fn, tp = confusion_matrix(y_te, pred).ravel()
    # a missed fraud costs 4000 rupees, a human review costs 50
    cost = fn * 4000 + (tp + fp) * 50
    acc = (pred == y_te).mean()
    print(f"{t:9.2f}  {acc:8.3f}  {fn:6d}  {tp+fp:7d}  {cost:6d}")
Output
'never flag anything' accuracy: 0.965

threshold  accuracy   missed  reviews    cost
     0.50     0.965     175        0  700000
     0.20     0.965     170       12  680600
     0.05     0.808      99      938  442900

The walkthrough

At threshold 0.50 the model flags nothing. Zero reviews. Its accuracy exactly equals "never flag anything" — the model is deployed and doing no work. Accuracy cannot see the difference.

At 0.05 accuracy collapses to 0.808 — and cost falls by a third. The model now catches 76 more frauds at the price of 938 reviews. In rupees, the "worst" threshold by accuracy is the best decision on the table.

The threshold is a business dial, not a modelling detail. The default 0.5 in predict() encodes the assumption that both mistakes cost the same. Almost no real decision satisfies it. Score with predict_proba, then choose the threshold where the cost curve bottoms out.

Confusion-matrix names, since they appear everywhere: fn (false negatives) are the missed frauds; fp (false positives) are innocent transactions flagged; tp (true positives) are frauds caught.

Common mistakes

Reporting accuracy on imbalanced data. The stakeholder hears 96.5% and celebrates the do-nothing model. Report recall, precision, and cost at the operating threshold instead — model evaluation walks through each.

Optimising a metric no one can act on. A beautiful AUC (a threshold-free ranking score) means little if the business runs at one fixed threshold with a fixed review budget. Evaluate at the budget you will actually have.

One metric when you need a metric plus a guardrail. "Minimise cost, subject to: no more than 1,000 reviews per day, and recall above 50%." Real deployments are constrained optimisation; state the constraints or the model will spend them.

Changing the metric mid-project. Every metric change silently invalidates earlier comparisons. Log the metric definition with the same care as the code version.

Try it yourself

Re-price the mistakes: missed fraud ₹500, review ₹50 — a low-stakes product. Find the new best threshold. Then answer in one sentence: why does the best threshold rise when missed fraud gets cheaper?

What to learn next

Researcher — Mathematics and papers.

Decision-theoretic foundation

Given action set $A$, outcomes $Y$, and a loss $L(a, y)$, the optimal decision rule given calibrated posteriors is:

$$ a^*(x) = \arg\min_{a \in A} \; \mathbb{E}_{y \sim P(Y \mid X=x)}\left[L(a, y)\right] $$

Where:

  • $a^*(x)$ — the cost-minimising action for input $x$.
  • $P(Y \mid X = x)$ — the model's predictive distribution, assumed calibrated.
  • $L(a, y)$ — the loss of taking action $a$ when the truth is $y$.

For binary classification with false-positive cost $c_{FP}$ and false-negative cost $c_{FN}$, this reduces to thresholding the posterior at:

$$ t^* = \frac{c_{FP}}{c_{FP} + c_{FN}} $$

Elkan (2001), The Foundations of Cost-Sensitive Learning (IJCAI), derives this and its corollary: cost-sensitivity can live entirely in the threshold when probabilities are honest. The catch is the premise — miscalibrated scores break the formula, which is why calibration (reliability diagrams, Platt scaling, isotonic regression) is a scoping concern, not a finishing touch.

Proper scoring rules

Selecting models by a proper scoring rule — one minimised in expectation by the true probabilities, such as log loss or the Brier score — preserves the option to change thresholds later. Improper targets (accuracy at a fixed threshold) entangle the model with today's costs. Gneiting and Raftery (2007), Strictly Proper Scoring Rules, Prediction, and Estimation (JASA), is the canonical treatment.

Metrics as incentives: Goodhart effects

Once a metric selects models (or people), optimisation pressure exploits its gaps — Goodhart's law. ML-specific analyses: Manheim and Garrabrant (2018), Categorizing Variants of Goodhart's Law; reward hacking in RL (Amodei et al., 2016, Concrete Problems in AI Safety). The mitigation pattern is a small metric portfolio: one optimised metric, plus guardrail metrics that must not regress — the structure Ng formalises as optimising versus satisficing metrics (Machine Learning Yearning, 2018).

Offline metric versus online outcome

The offline metric is itself a proxy for a business outcome measured later by experiment. Correlation between offline gains and A/B results is imperfect and domain-specific; recommender-system studies report notable disagreement rates (e.g., Garcin et al., 2014; the broader debate in Jannach and Bauer, 2020). The scoping implication: pre-register the offline metric you expect to transfer, then verify the transfer on the first deployment rather than assuming it.

What to learn next