AI in Healthcare

Label leakage in clinical models

A clinical model can score close to perfect by learning what a doctor already suspected, not what actually predicts a patient's illness.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Label leakage means a model is secretly handed the answer, through a feature shaped by a doctor's own suspicion.

Think about a student who scores 100% on a test because the answer key was tucked inside it. The student did not learn the material. They read the answers.

A clinical model can do the same thing without anyone noticing. It finds a feature that is really a trace of the diagnosis. That trace was left behind by a doctor's own actions — and the model calls it "prediction."

Why it exists

A doctor who suspects a heart attack orders a troponin blood test. A doctor who does not suspect one, usually does not order it.

Now imagine a model predicting heart attacks, given "was a troponin test ordered" as one input. That feature is almost a copy of the answer. It does not measure the disease — it measures a doctor's belief about the disease, formed before the model ever ran.

This is label leakage. A feature could only exist after the outcome was already known or suspected. That feature gets smuggled into the model's own inputs. It is one of the most common reasons a clinical model looks brilliant in testing. Then it fails the moment it meets a real patient.

How it works

What actually happened, in order:

  Patient shows symptoms  ->  Doctor suspects heart attack  ->  Orders troponin test  ->  Test result confirms it


What the leaky model sees, with no sense of that order:

  "troponin test ordered" (input)   ->   "had a heart attack" (label)
                                              |
                          looks like a strong predictor.
                          is actually the doctor's opinion, restated.

Where you have already seen it

  • A spam filter trained on "was this moved to the Spam folder" would predict spam without learning anything real. The moving is the label — not a clue about it.
  • A loan approval model with "was a background check ordered" leaks the underwriter's own suspicion into the data.

An honest warning

Leakage is not a rare mistake made by careless people. It hides inside ordinary-looking columns with innocent names, and a model happily exploits it without telling anyone.

The giveaway is almost always the same: a result that seems too good. If a healthcare model reports near-perfect accuracy, ask one question first. Not "how did we get so lucky" — instead, "which feature is secretly the answer?"

Remember this

  • Leakage happens when a feature could only exist after the outcome was already known or suspected.
  • Clinical data is full of it, because so many recorded actions are themselves a doctor's diagnosis in disguise.
  • A suspiciously good result is a warning sign, not good news — check for leakage before celebrating.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn pandas

Minimal runnable code

leakage_demo.py
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

rng = np.random.default_rng(0)
n = 2000

# The real signal a model should learn from: age and resting heart rate.
age = rng.normal(55, 15, n)
resting_hr = rng.normal(75, 12, n)
risk_score = 0.04 * age + 0.03 * resting_hr - 4.5
true_prob = 1 / (1 + np.exp(-risk_score))
had_mi = rng.binomial(1, true_prob)

# A troponin test is ordered almost only when a clinician already suspects
# a heart attack. This column is downstream of the diagnosis, not a cause
# of it -- it leaks the answer straight into the features.
troponin_ordered = np.where(had_mi == 1, rng.binomial(1, 0.97, n), rng.binomial(1, 0.04, n))

df = pd.DataFrame({"age": age, "resting_hr": resting_hr,
                    "troponin_ordered": troponin_ordered, "had_mi": had_mi})

def evaluate(features):
    X_train, X_test, y_train, y_test = train_test_split(
        df[features], df["had_mi"], test_size=0.3, random_state=0, stratify=df["had_mi"]
    )
    model = LogisticRegression().fit(X_train, y_train)
    return accuracy_score(y_test, model.predict(X_test))

leaky_accuracy = evaluate(["age", "resting_hr", "troponin_ordered"])
honest_accuracy = evaluate(["age", "resting_hr"])

print(f"with the leaky 'troponin_ordered' feature: {leaky_accuracy:.3f}")
print(f"with only age and resting heart rate:      {honest_accuracy:.3f}")
Output
with the leaky 'troponin_ordered' feature: 0.960
with only age and resting heart rate:      0.630

What actually happened

Adding one leaky column jumped accuracy from 0.630 to 0.960. Nothing about the patients changed between those two runs — only whether the model was allowed to peek at a proxy for the doctor's own suspicion.

0.630 is a believably modest, honest number for a two-feature clinical model. 0.960 is the kind of number that should make an experienced practitioner immediately distrustful, not impressed.

Line by line, the parts that are not obvious:

  • troponin_ordered is generated from had_mi, not the other way round — rng.binomial(1, 0.97, n) when had_mi == 1. In code, that dependency is visible. In a real hospital extract, it is buried in the order in which columns were recorded, and invisible unless someone checks.
  • stratify=y_test keeps the same class balance in train and test, so the accuracy difference is not an artefact of an unlucky split.
  • Both models are evaluated on an identical test set, random_state=0, so this is a fair, apples-to-apples comparison.

Common mistakes

Trusting a feature because it is numeric and structured. troponin_ordered looks like an ordinary binary column. Nothing about its data type reveals that it is derived from the label.

Only checking for leakage when a result looks suspicious. By the time a result looks suspicious, the leaky feature may already have been reported to a stakeholder as a finding.

Removing the leaky feature but keeping others similar to it. A real EHR extract often has several near-copies of the diagnosis-adjacent signal — a follow-up appointment type, a discharge disposition, a specific medication class. Removing one and missing the rest still leaks.

Try it yourself

Add a second leaky feature, cardiology_consult, generated the same way troponin_ordered was — dependent on had_mi. Retrain with all three features and watch accuracy climb even further above 0.960, purely from stacking two leaks on top of each other.

What to learn next

Researcher — Mathematics and papers.

A formal definition

Leakage occurs when a feature x_i is not conditionally independent of the label y given the true causal information available at prediction time t:

x_i is leaky  if  x_i was generated at time  t_gen > t_predict
              and  x_i is a function of information that determined y
  • t_predict — the moment a real deployed model would have to make its prediction
  • t_gen — the moment the feature's value was actually recorded

Any feature with t_gen > t_predict is, by construction, using information from the future relative to a real deployment. The troponin-order example is a specific case: t_gen (when the order was placed) postdates the clinical suspicion that also determines y, even though both may be timestamped identically in a naive extract.

Why cross-validation does not catch this

Standard k-fold cross-validation, discussed generally in k-fold cross-validation, assumes leakage is a property of the split, not the feature. A feature that is leaky with respect to the causal timeline remains leaky in every fold of a random cross-validation, because the leak is baked into the feature's definition, not into which rows land in train versus test. This is why cross-validated accuracy on leaky clinical data can be simultaneously high, stable across folds, and completely uninformative about real-world performance.

Detection: temporal feature auditing

The reliable defence is a feature-availability audit: for every feature, explicitly record the earliest timestamp at which it could legitimately have been known, and reject any feature whose availability timestamp is not strictly before the prediction cutoff. This is standard practice in point-in-time correct feature pipelines, closely related to the temporal discipline covered in using the future to predict the past.

Documented real-world cases

The pattern generalises beyond troponin tests. Caruana et al. (2015) documented a related, though distinct, failure in a pneumonia mortality-risk model: asthmatic patients with pneumonia showed lower predicted risk, because in practice they were routed to more aggressive care (ICU admission), which improved their outcomes and confounded the model's association between "has asthma" and "low risk of dying." That case is a confounding failure rather than a strict leakage failure — the asthma flag was genuinely available before the outcome — but it illustrates the same underlying danger: a feature can carry a hidden signal about the care a patient received, not about their underlying condition, and a black-box model has no way to tell the two apart.

Key references

  • Kaufman, S., Rosset, S. & Perlich, C. (2012). Leakage in Data Mining: Formulation, Detection, and Avoidance. ACM TKDD. The general formal treatment.
  • Caruana, R. et al. (2015). Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission. KDD. The asthma/pneumonia confounding case.
  • Wong, A. et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model. JAMA Internal Medicine. A case where undisclosed feature construction contributed to an optimistic reported performance later not matched in practice.

What to learn next