Machine Learning

Model evaluation

Model evaluation is measuring whether a trained model is actually any good, using scores that reveal its real mistakes instead of hiding them.

Read these first

On this page 8
  1. Why one number is not enough
  2. The two ways to be wrong
  3. You cannot have both
  4. How it works
  5. Judging a number-predicting model
  6. Where you have seen this
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Model evaluation is measuring how good your model really is. A good measurement shows the model's mistakes instead of hiding them.

Think about the smoke alarm in a kitchen. It can fail in two completely different ways. It can shriek every time you fry an onion, when there is no fire. Or it can stay silent while the kitchen actually burns.

Both are failures. They are not equally bad, and no single number describes both. That is the whole of this lesson.

Why one number is not enough

Most people reach for accuracy — the share of predictions that were correct. It feels like the obvious measure. It is also the one that misleads most often.

Here is the trap.

Imagine a disease that four people in a hundred actually have. I will now give you a model. My model ignores its input entirely and answers "healthy" every single time.

That model is correct ninety-six times out of a hundred. Ninety-six percent accuracy.

It has also never once found a sick person. It is completely worthless, and its score sounds excellent. This is not a made-up edge case. It is the normal situation in fraud detection, disease screening, and defect spotting. The thing you are looking for is rare.

The two ways to be wrong

Give the mistakes names and everything gets clearer.

  • A false alarm is shouting when nothing is there. The alarm goes off while you fry onions. A healthy person is told they are ill.
  • A miss is staying silent when something is there. The kitchen burns and the alarm sleeps. A sick person is sent home.

Now the two useful questions.

When it shouts, how often is it right? This is called precision. Low precision means a noisy alarm nobody trusts. Eventually somebody takes the battery out.

Of all the real fires, how many did it catch? This is called recall. Low recall means a quiet alarm that lets the house burn.

You cannot have both

This is the part worth slowing down for.

Make the alarm more sensitive and it catches every fire — and screams at every onion. Recall goes up, precision goes down.

Make it less sensitive and it stops screaming at onions — and starts sleeping through small fires. Precision goes up, recall goes down.

There is no setting that maximises both. You have to choose, based on which mistake costs more.

  • Screening for a treatable cancer? Chase recall. A false alarm means an extra test. A miss means a death.
  • Filtering spam? Chase precision. A little spam in the inbox is annoying. A job offer in the spam folder is a disaster.

That decision is not a technical one. It belongs to whoever will live with the consequences.

How it works

Every prediction lands in one of four boxes.

                    WHAT THE MODEL SAID
                    "sick"          "healthy"
                 +---------------+---------------+
   W    actually |   caught it   |   MISSED it   |
   H       sick  |               |               |
   A             +---------------+---------------+
   T    actually |  FALSE ALARM  |  correctly     |
        healthy  |               |  left alone    |
   IS            +---------------+---------------+


   Precision = of everything in the "sick" column, how much was truly sick
   Recall    = of everything in the "actually sick" row, how much was caught

This grid is called a confusion matrix. Look at it before you look at any score. It shows you the actual mistakes, and scores are only summaries of it.

Judging a number-predicting model

If your model predicts a number rather than a bucket, precision and recall do not apply. You measure the size of the misses instead.

  • Average miss. On average, how far off is a prediction? If a price model is off by 3 lakh on average, that sentence means something to anybody.
  • Average miss, punishing large errors more. This is the usual default. It cares much more about being wildly wrong once than slightly wrong often.

Which one you want depends on whether one huge mistake is worse than many small ones. For predicting delivery times, one two-hour error is far worse than twenty one-minute errors.

Where you have seen this

  • Spam filters tuned for precision, which is why some spam still reaches you.
  • Bank fraud blocks tuned for recall, which is why your genuine payment sometimes gets stopped.
  • Airport security tuned heavily for recall, which is why so many bags get flagged.
  • Medical screening, where a positive result usually means "take a better test", not "you are ill".

Remember this

  • Accuracy hides the truth whenever the thing you are looking for is rare.
  • Precision asks how often an alarm is right. Recall asks how many real cases were caught.
  • Raising one lowers the other, so the choice depends on which mistake hurts more.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

No training in this example. We hand-write the true labels and the predictions, so every number can be checked by counting. That makes it a good file to keep and reread.

The rare-disease trap, in code

evaluate.py
from sklearn.metrics import confusion_matrix, classification_report
from sklearn.metrics import accuracy_score, recall_score

# 20 patients. 1 means the illness is present, 0 means it is absent.
y_true = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,1,1,1]
# A lazy model that answers "healthy" for everybody
y_lazy = [0] * 20
# A model that actually looked at the data
y_pred = [0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,1,1,0,1]

print("lazy model  -> accuracy", accuracy_score(y_true, y_lazy),
      "| recall on sick patients", recall_score(y_true, y_lazy, zero_division=0))
print("real model  -> accuracy", accuracy_score(y_true, y_pred),
      "| recall on sick patients", recall_score(y_true, y_pred))
print()
print("confusion matrix for the real model (rows = truth, columns = prediction):")
print(confusion_matrix(y_true, y_pred))
print()
print(classification_report(y_true, y_pred, target_names=["healthy", "sick"], zero_division=0))
Output
lazy model  -> accuracy 0.8 | recall on sick patients 0.0
real model  -> accuracy 0.9 | recall on sick patients 0.75

confusion matrix for the real model (rows = truth, columns = prediction):
[[15  1]
 [ 1  3]]

              precision    recall  f1-score   support

     healthy       0.94      0.94      0.94        16
        sick       0.75      0.75      0.75         4

    accuracy                           0.90        20
   macro avg       0.84      0.84      0.84        20
weighted avg       0.90      0.90      0.90        20

Read the first line again

The lazy model scores 80 percent accuracy while finding zero sick patients. Its recall is exactly 0.0.

If accuracy were your only metric, you would have shipped a model that is a constant. This is why 80 percent, on its own, tells you nothing at all.

Reading the confusion matrix

scikit-learn orders rows and columns by sorted label value, so with labels 0 and 1 the layout is:

              predicted 0   predicted 1
   actual 0  [    15            1      ]   <- 15 correct, 1 false alarm
   actual 1  [     1            3      ]   <-  1 missed,  3 caught

Check the numbers by hand, because it builds the intuition permanently:

  • Precision for "sick" = caught / everything called sick = 3 / (3 + 1) = 0.75
  • Recall for "sick" = caught / everyone actually sick = 3 / (3 + 1) = 0.75
  • Accuracy = everything correct / everything = (15 + 3) / 20 = 0.90

Both denominators happen to be 4 here, which is a coincidence of this data. They are different quantities and will normally differ.

Which line of the report to read

The classification_report gives four summary rows, and picking the wrong one is a common error.

  • sick row — the metrics for the class you care about. Usually the row that matters.
  • accuracy — overall correctness. Unreliable when classes are imbalanced, as here.
  • macro avg — plain average across classes, treating both equally. Use it when the rare class matters as much as the common one.
  • weighted avg — average weighted by class size. It is dominated by the majority class, so it reproduces the same blind spot as accuracy. Here it reads 0.90, which flatters the model.

f1-score is the harmonic mean of precision and recall — a single number that stays low unless both are decent. It is a reasonable default when you have no reason to prefer one over the other. It is the wrong choice when you do have a reason.

Moving the threshold

Precision and recall are not fixed properties of a model. They depend on the cut-off you apply to its probability output.

threshold.py
from sklearn.metrics import precision_score, recall_score

y_true = [0, 0, 0, 0, 0, 0, 1, 1, 1, 1]
# The probabilities a model assigned to each patient being sick
probs = [0.02, 0.10, 0.18, 0.31, 0.44, 0.55, 0.40, 0.62, 0.78, 0.91]

for cut in (0.30, 0.50, 0.70):
    pred = [1 if p >= cut else 0 for p in probs]
    p = precision_score(y_true, pred, zero_division=0)
    r = recall_score(y_true, pred, zero_division=0)
    print(f"threshold {cut:.2f} -> precision {p:.2f}  recall {r:.2f}")
Output
threshold 0.30 -> precision 0.57  recall 1.00
threshold 0.50 -> precision 0.75  recall 0.75
threshold 0.70 -> precision 1.00  recall 0.50

One model. One set of predictions. Three completely different characters.

At 0.30 it catches every sick patient and raises three false alarms. At 0.70 it never raises a false alarm and misses half the sick patients. The trade-off from the Beginner section is right there in the numbers.

The default threshold of 0.5 is a convention, not a recommendation. Choosing it deliberately is part of building the model.

Common mistakes

Reporting accuracy on imbalanced data. The headline error of this lesson. Report per-class precision and recall, or use average='macro'.

Evaluating on the training set. Any score computed on rows the model trained on measures memorisation. See train, test and validation splits.

Leaving the threshold at 0.5 without thinking. Shown above. Pick it from the cost of each mistake, not from the default.

Comparing F1 across datasets with different class balance. F1 depends on the positive-class rate. An F1 of 0.7 on one dataset is not comparable to 0.7 on another.

Using ROC-AUC on heavily imbalanced data. ROC curves use the false-positive rate, whose denominator is the large negative class. That makes the curve look flattering when positives are rare. Use a precision-recall curve and average_precision_score instead.

Trusting predict_proba as a real probability. Many models, including boosted trees and most neural networks, produce systematically overconfident scores. If you need a genuine probability, calibrate with CalibratedClassifierCV.

Try it yourself

In threshold.py, add 0.01 and 0.90 to the loop to see both extremes:

Output
threshold 0.01 -> precision 0.40  recall 1.00
threshold 0.90 -> precision 1.00  recall 0.25

At 0.01 the model flags all ten patients. Recall reaches 1.00, and precision falls to 0.40 — which is exactly the share of patients who are sick. Flagging everybody always gives you precision equal to the base rate.

At 0.90 it flags a single patient and gets that one right. Precision is 1.00 and recall drops to 0.25.

Those two rows are the boundaries of what any threshold on this model can achieve. Everything useful lives between them.

Now try 0.20 and predict the result before running it. Many people expect it to flag everybody. It does not. The lowest two probabilities are 0.02 and 0.10, so seven patients get flagged. That gives precision 0.57 and recall 1.00, identical to the 0.30 row. Thresholds only matter where predictions actually sit.

What to learn next

Researcher — Mathematics and papers.

Definitions from the confusion matrix

For binary classification with positive class 1:

TP = true positives    (predicted 1, actually 1)
FP = false positives   (predicted 1, actually 0)   -- Type I error
FN = false negatives   (predicted 0, actually 1)   -- Type II error
TN = true negatives    (predicted 0, actually 0)
Accuracy    = (TP + TN) / (TP + TN + FP + FN)
Precision   = TP / (TP + FP)                      also called PPV
Recall      = TP / (TP + FN)                      also sensitivity, TPR
Specificity = TN / (TN + FP)                      also TNR
FPR         = FP / (FP + TN)                      = 1 - Specificity
F_beta      = (1 + beta^2) * P * R / (beta^2 * P + R)
  • P, R — precision and recall
  • beta — weight on recall relative to precision. beta = 1 gives F1. beta = 2 weights recall four times as heavily

Note the asymmetry that makes rare-event evaluation hard. Precision's denominator depends on the number of predicted positives. Recall's denominator depends on the number of actual positives. Only precision is affected by class prevalence. That is why precision moves when the base rate moves, and recall does not.

The base-rate problem, formally

By Bayes' theorem, the positive predictive value is:

PPV = ( Sens * Prev ) / ( Sens * Prev + (1 - Spec) * (1 - Prev) )
  • Sens — sensitivity (recall), Spec — specificity, Prev — prevalence of the positive class

Take a test with 99 percent sensitivity and 99 percent specificity applied to a condition with prevalence 0.001. Then PPV = (0.99 * 0.001) / (0.99 * 0.001 + 0.01 * 0.999) ≈ 0.090.

A 99-percent-accurate test yields a positive result that is wrong about 91 percent of the time. This is not a modelling failure; it is arithmetic, and it is why screening programmes are always two-stage.

Threshold-free measures

ROC-AUC is the area under the TPR-versus-FPR curve. It equals the probability that a random positive scores higher than a random negative. That is the Mann-Whitney U statistic. It is invariant to class prevalence, which is both its strength and its trap.

PR-AUC / average precision integrates precision against recall. Its baseline is the positive rate, not 0.5, so it moves with prevalence.

Davis & Goadrich (2006) proved that a curve dominating in ROC space also dominates in PR space, and conversely. The visual impression, however, differs sharply under imbalance. Saito & Rehmsmeier (2015) showed empirically that ROC plots mislead for rare positives. Use PR curves when positives are rare and you care about them.

Average precision = SUM_n ( R_n - R_{n-1} ) * P_n
  • P_n, R_n — precision and recall at the n-th threshold

Prefer average_precision_score over the trapezoidal auc(recall, precision), which interpolates optimistically between operating points.

Matthews correlation coefficient

MCC = ( TP*TN - FP*FN ) / sqrt( (TP+FP)(TP+FN)(TN+FP)(TN+FN) )

MCC is the Pearson correlation between predicted and true binary labels. It ranges from -1 to +1, with 0 meaning chance. Unlike F1 it uses all four cells of the confusion matrix. It goes high only when the model does well on both classes.

Chicco & Jurman (2020) argue MCC should replace F1 and accuracy as the default binary metric. The argument is strong. F1 ignores TN entirely, so a model that handles the negative class badly can still post a good F1.

Calibration

Discrimination and calibration are distinct properties. A model can rank perfectly (AUC 1.0) while its probabilities are badly wrong.

Brier score = (1/n) * SUM_i ( p_i - y_i )^2
  • p_i — predicted probability for example i, y_i — its true binary label

The Brier score is a strictly proper scoring rule, meaning it is uniquely minimised by reporting true probabilities. Log loss is another. Murphy (1973) decomposes Brier into reliability, resolution, and uncertainty.

Guo et al. (2017) showed that modern deep networks are far worse calibrated than the shallower networks of the 1990s, despite higher accuracy. Increased capacity and reduced regularisation push confidence up faster than correctness.

Temperature scaling, a single scalar fitted on validation data, corrects most of it. Platt scaling (1999) and isotonic regression (Zadrozny & Elkan, 2002) are the classical alternatives.

Expected Calibration Error bins predictions by confidence and averages the confidence-accuracy gap. It is sensitive to binning choices, so report the scheme used.

Regression metrics

MAE  = (1/n) * SUM |y_i - y_hat_i|
MSE  = (1/n) * SUM (y_i - y_hat_i)^2
RMSE = sqrt(MSE)
MAPE = (100/n) * SUM | (y_i - y_hat_i) / y_i |

RMSE is minimised by the conditional mean, MAE by the conditional median. That is the substantive difference; the robustness of MAE to outliers follows from it rather than being a separate fact.

MAPE has two defects that disqualify it more often than people realise. It is undefined at y_i = 0. And it penalises over-prediction more heavily than under-prediction, which biases forecasts downward. Prefer MASE (Hyndman & Koehler, 2006) for scale-free forecast comparison.

Uncertainty in the estimate itself

A reported metric is a point estimate from a finite sample and carries variance. For accuracy on n test points, the normal-approximation interval is:

SE = sqrt( acc * (1 - acc) / n )

With n = 1000 and accuracy 0.90, SE ≈ 0.0095, so the 95 percent interval spans roughly 0.881 to 0.919. A competitor scoring 0.905 is not meaningfully better. Use the Wilson interval rather than the normal approximation for small n or extreme rates.

To compare two models on the same test set, use McNemar's test (1947) on the discordant pairs. Dietterich (1998) evaluates the common alternatives and recommends it for the single-test-set case. Cross-validated t-tests violate independence, and Nadeau & Bengio (2003) give the variance correction.

Key references

  • Davis, J. & Goadrich, M. (2006). The Relationship Between Precision-Recall and ROC Curves. ICML.
  • Saito, T. & Rehmsmeier, M. (2015). The Precision-Recall Plot Is More Informative than the ROC Plot on Imbalanced Datasets. PLoS ONE 10(3).
  • Guo, C., Pleiss, G., Sun, Y. & Weinberger, K. (2017). On Calibration of Modern Neural Networks. ICML.
  • Chicco, D. & Jurman, G. (2020). The Advantages of the Matthews Correlation Coefficient over F1 and Accuracy. BMC Genomics 21(6).
  • Dietterich, T. (1998). Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms. Neural Computation 10(7).
  • Hyndman, R. & Koehler, A. (2006). Another Look at Measures of Forecast Accuracy. Int. J. Forecasting 22(4).
  • Murphy, A. (1973). A New Vector Partition of the Probability Score. J. Applied Meteorology 12(4).

Current state

Single-number leaderboard metrics are increasingly recognised as inadequate. Current practice favours:

  • Disaggregated evaluation — report per-subgroup metrics, since aggregate parity can hide large per-group disparities (Buolamwini & Gebru, 2018).
  • Behavioural testing — CheckList (Ribeiro et al., 2020) tests capabilities directly rather than scoring a held-out sample.
  • Distribution shift — WILDS (Koh et al., 2021) evaluates on deliberately shifted test distributions. In-distribution accuracy is a weak predictor of performance there.
  • Documentation — Model Cards (Mitchell et al., 2019) and Datasheets (Gebru et al., 2021). Both standardise reporting of intended use and evaluation conditions.

For generative models the picture is less settled. Perplexity measures fit to a distribution rather than usefulness. Automated LLM-as-judge protocols correlate reasonably with human preference but carry known position and verbosity biases (Zheng et al., 2023). Treat any single generative benchmark number with more scepticism than a classification metric.

What to learn next