Predicting patient deterioration
A model that predicts patient decline is an ordinary classifier wearing high stakes, where the usual accuracy number can hide a model that is nearly useless.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Predicting deterioration means guessing, hours ahead, whether a patient is about to get significantly worse.
Think of a teacher who notices a usually-attentive student going quiet and staring at the desk. Nothing dramatic has happened yet. The teacher has learned, from years of small signals, that this often comes before a bigger problem.
A deterioration model is trying to be that teacher, using numbers instead of instinct. Heart rate creeping up, blood pressure drifting down — small changes that alone mean nothing, but together are a warning.
Why it exists
Hospital wards already use simple scoring systems for this. Add up points for abnormal heart rate, breathing rate, and blood pressure, and you get a single early-warning number. Nurses check this score by hand, on a set schedule.
A machine learning model tries to do the same job with more signals, checked automatically. It can check far more often than a person could manage. The goal stays the same as any early-warning system: catch the problem while there is still time to act.
How it works
Vitals + labs over the last few hours
|
v
[ trained model ]
|
v
probability the patient
deteriorates in the next 6h -> 0.83 -> alert a nurseThe model outputs a probability, not a yes-or-no answer. A hospital then has to pick a threshold — how high that probability needs to be before someone gets alerted.
Where you have already seen it
- A car's low-fuel light is a simple threshold system: it warns you before the tank is empty, not after.
- A smoke detector trades false alarms for the guarantee that it almost never misses a real fire.
Hospital deterioration alerts make the exact same trade-off, with far higher stakes on both sides.
An honest warning
In most hospitals, the great majority of patients stay stable. Deterioration is the rare case, not the common one.
That single fact breaks a beginner's instinct to trust "accuracy" as a measure of a good model. A model that never raises an alarm can still score highly accurate. It is right about the common case almost every time — and useless for the one job it exists to do. This part is confusing the first time you see the numbers. Read the developer example below twice if it does not click immediately. That reaction is normal, not a sign you missed something.
Remember this
- The task is ordinary binary classification, made high-stakes by what a wrong answer costs.
- Deterioration is rare, so raw accuracy is a misleading way to judge the model.
- The probability threshold used to trigger an alert is a decision, not a detail. Get it wrong, and the model ends up useless or ignored.
- A model like this needs clinical validation and regulatory sign-off before it runs near a real patient. A nurse or doctor still reviews every alert.
What to learn next
- Classification — the general technique this lesson applies to a specific, high-stakes setting.
- False alarms and alert fatigue — what happens when the threshold above is set carelessly.
- ROC vs precision-recall curves — the right way to evaluate a model like this one.
Developer — Code and libraries.
Setup
pip install scikit-learnMinimal runnable code
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score, average_precision_score, accuracy_score, recall_score
# Stand-in for vitals: heart rate, respiratory rate, systolic BP, temperature,
# age. "1" means the patient deteriorated in the next 6 hours. On a real
# general ward, that outcome is genuinely rare.
X, y = make_classification(
n_samples=2000, n_features=5, n_informative=4, n_redundant=0,
weights=[0.93, 0.07], flip_y=0.02, class_sep=1.1, random_state=0,
)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, stratify=y, random_state=0
)
print(f"deterioration rate in the data: {y.mean():.1%}")
model = LogisticRegression().fit(X_train, y_train)
probs = model.predict_proba(X_test)[:, 1]
preds = model.predict(X_test)
# The trap: a model that never predicts deterioration still scores high on
# plain accuracy, because "stable" is right 93% of the time by default.
always_stable_accuracy = 1 - y_test.mean()
print(f"accuracy of a model that always predicts 'stable': {always_stable_accuracy:.3f}")
print(f"accuracy of the trained model: {accuracy_score(y_test, preds):.3f}")
print()
print(f"ROC-AUC: {roc_auc_score(y_test, probs):.3f}")
print(f"PR-AUC: {average_precision_score(y_test, probs):.3f}")
print(f"recall at the default 0.5 threshold: {recall_score(y_test, preds):.3f}")deterioration rate in the data: 7.6% accuracy of a model that always predicts 'stable': 0.923 accuracy of the trained model: 0.923 ROC-AUC: 0.729 PR-AUC: 0.460 recall at the default 0.5 threshold: 0.000
What actually happened
Read those two accuracy numbers again. They are identical: 0.923. The trained model matched a model that does nothing at all.
Now read the recall: 0.000. At the default 0.5 threshold, this model never once predicted deterioration on the test set — not correctly, not incorrectly, not at all. It never crossed 0.5. And yet the ROC-AUC of 0.729 says the model has real information inside it; it ranks deteriorating patients higher than stable ones far better than chance. The information is there. The default threshold is throwing it away.
Line by line, the parts that are not obvious:
weights=[0.93, 0.07]sets the class imbalance directly — 7% positive class, chosen to resemble a plausible ward-level deterioration rate for this illustration..predict()applies a threshold of exactly 0.5 internally..predict_proba()returns the raw probability, before any threshold is applied.- PR-AUC (
average_precision_score) is the more honest single number here. Unlike ROC-AUC, it does not get inflated by the (huge) number of easy true negatives — see ROC vs precision-recall curves for why.
Common mistakes
Reporting accuracy as the headline metric. On imbalanced clinical data, accuracy rewards a model for agreeing with the majority class, which is exactly the class nobody needs a model to detect.
Leaving the threshold at the library default. 0.5 is an arbitrary number with no clinical meaning. The right threshold depends on how many false alarms a ward can absorb per shift — see choosing a threshold from costs.
Training and testing on a random shuffle of all patients. A patient's own future data can leak into their own past if visits are split randomly instead of by patient — see group leakage.
Try it yourself
Lower the threshold from 0.5 to 0.2: preds_low = (probs >= 0.2).astype(int). Recompute recall and accuracy. Watch recall rise and accuracy fall — that trade-off, not either number alone, is the real design decision a hospital has to make.
What to learn next
- Class weights — a standard way to make a model pay more attention to the rare, important class during training.
- False alarms and alert fatigue — what a poorly chosen threshold costs in practice.
- Choosing a threshold from costs — turning this trade-off into an actual number.
Researcher — Mathematics and papers.
Framing the target
Deterioration prediction is framed as binary classification over a fixed prediction horizon h and a gap g:
y_t = 1 if a deterioration event occurs in (t + g, t + g + h]
y_t = 0 otherwiset— the time of predictiong— a mandatory gap between prediction and the earliest permitted event, deliberately excluding events that were already clinically obvious at prediction timeh— the prediction horizon, e.g. 6 hours
g exists specifically to prevent the leakage failure mode covered in label leakage in clinical models: without it, a model can appear to "predict" an event that clinical staff had already started responding to before t.
Why ROC-AUC misleads under severe imbalance
ROC-AUC is computed from the true-positive rate and false-positive rate, both of which are normalised within their own class:
TPR = TP / (TP + FN)
FPR = FP / (FP + TN)Because TN is enormous under a 7% positive rate, FPR stays small even when the count of false positives is large relative to the count of true positives. Precision, by contrast, is directly sensitive to that imbalance:
Precision = TP / (TP + FP)Davis & Goadrich (2006) formalised the resulting rule: a large gap between ROC-AUC and PR-AUC is expected, not a bug, whenever the positive class is rare — exactly the regime deterioration prediction lives in.
Real-world validation results
The best-known clinical deterioration model is the Epic Sepsis Model, deployed across hundreds of US hospitals. An independent external validation by Wong et al. (2021, JAMA Internal Medicine) on over 27,000 patients found substantially worse discrimination than the vendor's reported figures, and found that the model missed the majority of sepsis cases while generating a large volume of alerts. Readers should consult the primary source for the exact reported numbers rather than trust any single restated figure, including this one — but the qualitative finding is well established and widely cited: a model's reported performance at development time is not a reliable guide to its performance after deployment at a new site.
Cost
For n training examples and d features, a gradient-boosted tree ensemble (the most common production choice for this task) costs O(n d log n) typical training time and O(depth x n_trees) per prediction — cheap enough to run continuously in real time on ordinary hospital IT hardware, which is one reason it dominates deep sequence models in this specific domain.
Key references
- Davis, J. & Goadrich, M. (2006). The Relationship Between Precision-Recall and ROC Curves. ICML.
- Wong, A. et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine.
- Churpek, M. et al. (2016). Multicenter Comparison of Machine Learning Methods and Conventional Regression for Predicting Clinical Deterioration on the Wards. Critical Care Medicine.
- Escobar, G. et al. (2020). Automated Identification of Adults at Risk for In-Hospital Clinical Deterioration. NEJM.
Current state
Gradient-boosted trees and logistic regression remain competitive with, and often outperform, deep learning approaches on this task, largely because clinical tabular datasets are small relative to the feature engineering effort available, and because interpretability requirements favour simpler models. See prospective clinical validation for why strong retrospective numbers on any of these architectures are not sufficient evidence for deployment.
What to learn next
- ROC vs precision-recall curves — the evaluation machinery this task depends on.
- Validating a model before it touches patients — why retrospective AUC is not the end of the story.
- Reliability diagrams and calibration error — whether the predicted probability itself can be trusted.