When medical AI has caused harm
Documented cases of clinical AI causing real harm share a pattern — quietly biased proxies, and confident deployment before real-world testing caught up.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Medical AI has already caused documented, real-world harm — not hypothetical harm, but cases with names, published papers, and patients affected.
Think about a car recall. A defect nobody noticed in testing shows up once thousands of real drivers use the car. The company has to publicly explain what went wrong, and fix it.
Medical AI has had its own version of this, more than once. A model that looked promising in a research paper caused a real problem. That problem only showed up once the model met real patients, at scale. Studying those cases honestly is not pessimism. It is how the field learns.
Why it exists
Every earlier lesson in this section exists because of failures like these — leakage, alert fatigue, validation before deployment. They are not hypothetical warnings invented to sound careful. They are lessons the field learned the hard way.
Two patterns show up again and again in documented failures. First: a model trained on a proxy for what it actually needed to predict — where that proxy was itself unfair. Second: a model deployed, or nearly deployed, before it was properly tested.
How it works
A pattern seen across real documented cases:
Model trained on a proxy -> proxy is biased in a way
for the real target nobody checked for
| |
v v
looks accurate in testing + quietly unfair in deployment
|
v
harm discovered after the fact,
usually by outside researchersA real, well-documented example
In 2019, researchers published a widely cited study (Obermeyer et al., Science, 2019). It examined a healthcare risk-prediction algorithm used across the United States, deciding which patients received extra care. The algorithm predicted future healthcare cost, as a stand-in for future healthcare need. That sounds reasonable — cost is easy to measure, and need is not.
Here is the problem. At the same level of true underlying illness, Black patients had historically generated lower healthcare costs than white patients. Not because they were healthier — because of well-documented, unequal access to care. The algorithm learned that pattern faithfully. It systematically underestimated how much care Black patients needed. Correcting the algorithm to predict health need directly, instead of cost, substantially closed that gap. The paper itself is the authoritative source for the exact figures, and is well worth reading in full.
Where you have already seen this shape of problem
You have not seen this specific case before reading this lesson — and that is worth noticing. Most people never hear about the AI systems that quietly affect their care. These decisions happen behind the scenes of a hospital or insurer's operations, not in a visible app.
An honest warning
None of these documented failures happened because the people building them were careless or unqualified. Healthcare AI is genuinely hard to get right. A model can look accurate by every standard metric. It can still be unfair or unsafe in a way those metrics never measured. That is precisely why every lesson in this section keeps returning to the same message. A toy model, including every example in this section, is for learning the ideas — never for a real clinical decision.
Remember this
- Documented, real-world harms from medical AI exist, and studying them is how the field improves.
- A common pattern is a biased proxy standing in for the thing a model actually needed to predict.
- Standard accuracy metrics can look fine while a model is unfair or unsafe in ways those metrics do not capture.
What to learn next
- Bias in datasets — the general version of the proxy problem in this lesson's case study.
- Deploying responsibly — practical steps that address the pattern described here.
- Validating a model before it touches patients — the safeguard most directly aimed at catching this before deployment.
Developer — Code and libraries.
Setup
pip install numpy pandasMinimal runnable code
This is a synthetic illustration of the mechanism behind the Obermeyer et al. case above — not a reproduction of their data or their numbers. It exists to make the underlying logic concrete.
import numpy as np
import pandas as pd
rng = np.random.default_rng(0)
n = 4000
# SYNTHETIC illustration, not real hospital data. "true_need" is an
# unobserved severity score: how much care a patient actually needs.
# "group" stands in for a demographic split with unequal historical
# access to care.
group = rng.choice(["A", "B"], size=n, p=[0.5, 0.5])
true_need = rng.normal(50, 15, n)
# Group B historically faced more barriers to using the health system --
# cost, transport, trust -- so the SAME level of need produced LESS spend.
access_penalty = np.where(group == "B", 0.55, 1.00)
healthcare_cost = true_need * access_penalty + rng.normal(0, 5, n)
df = pd.DataFrame({"group": group, "true_need": true_need, "healthcare_cost": healthcare_cost})
# A model that ranks patients for an extra-care programme by predicted
# cost -- cheap to measure, and a common real-world shortcut for "need".
flagged_by_cost = df["healthcare_cost"] >= df["healthcare_cost"].quantile(0.80)
# The same top 20%, ranked by the thing that actually matters.
flagged_by_need = df["true_need"] >= df["true_need"].quantile(0.80)
print("share of each group flagged for extra care, by predicted cost:")
print(df.assign(flagged=flagged_by_cost).groupby("group")["flagged"].mean().round(3).to_string())
print()
print("share of each group flagged for extra care, by true need:")
print(df.assign(flagged=flagged_by_need).groupby("group")["flagged"].mean().round(3).to_string())share of each group flagged for extra care, by predicted cost: group A 0.392 B 0.002 share of each group flagged for extra care, by true need: group A 0.198 B 0.202
What actually happened
Ranking by cost, group A gets flagged for extra care 39.2% of the time. Group B gets flagged 0.2% of the time — almost never, despite having the same distribution of true need in this synthetic data. Ranking by true need instead, both groups land close to the intended 20% flag rate: 19.8% and 20.2%.
Nothing in this code makes the model "racist" in any explicit sense. It never sees group as an input. The unfairness enters entirely through the proxy — cost — carrying a historical pattern that has nothing to do with medical need.
Line by line, the parts that are not obvious:
groupis never given to the model as a feature, and the bias appears anyway. Removing a sensitive attribute from the input data does not remove bias that is encoded in a different column correlated with it.access_penaltyis the entire mechanism: identicaltrue_need, differenthealthcare_cost, purely because of an access difference built into this synthetic world.- The
.quantile(0.80)threshold is the same operation a real "top 20% highest risk" clinical programme might use — the code models a realistic decision rule, even though the data is invented.
Common mistakes
Assuming "the model doesn't see race" means the model can't be biased on race. This example demonstrates exactly how that assumption fails, through an ordinary correlated proxy.
Treating fairness as something checked once, after training. Checking whether outcomes differ across relevant groups needs to happen at every stage — feature choice, label choice, and post-deployment monitoring.
Assuming a proxy is safe because it is easy to measure. Cost is easy to measure precisely because it is generated by the existing healthcare system — which means it inherits every inequity already built into that system.
Try it yourself
Change access_penalty for group B from 0.55 to 0.85, simulating a smaller access gap. Rerun and watch how much the flagging gap narrows — the bias in this toy model scales directly with the size of the underlying access disparity you set.
What to learn next
- Bias in datasets — the general pattern this case study demonstrates.
- Fairness metrics — ways to formally check for a gap like the one shown above.
- SHAP and LIME — tools for finding which features are driving a model's decisions.
Researcher — Mathematics and papers.
The Obermeyer et al. (2019) finding
Obermeyer, Powers, Vyas & Mullainathan (Science 366, 2019) audited a commercial risk-prediction algorithm used to identify patients for a hospital's high-risk care management programme, applied to roughly 200 million people annually in the US healthcare system at the time of the study. The algorithm's label was healthcare cost, used as a proxy for healthcare need. The authors found that Black patients, at any given predicted risk score, were substantially sicker than white patients at that same score — because the algorithm had learned the historical relationship between race and healthcare spending, not the relationship between race and underlying illness. Readers should consult the original paper for the precise reported figures; this lesson deliberately does not restate them from memory, in keeping with the standard that a cited statistic should be independently verifiable rather than reconstructed from recall.
Other documented cases
IBM Watson for Oncology. Internal training documents, reported by STAT News in 2018 (Ross & Swetlitz), showed the system had produced "multiple examples of unsafe and incorrect" cancer treatment recommendations during internal testing, reportedly due in part to training on a limited number of synthetic and hypothetical cases rather than sufficient real patient data. IBM discontinued sales of the product in subsequent years. This case is widely cited as a caution against training and validating primarily on synthetic or limited data before clinical deployment.
The Epic Sepsis Model. Covered in predicting patient deterioration and validating a model before it touches patients: an independent external validation (Wong et al., 2021, JAMA Internal Medicine) found substantially worse real-world discrimination than the vendor's reported figures, despite the model being deployed across hundreds of hospitals prior to that independent audit.
A general causal framing
Each case above fits a common structure: a model was trained to predict a proxy Z (cost, a synthetic case outcome, a specific EHR-derived sepsis signal) standing in for a true target Y (health need, correct treatment, actual sepsis onset), where Z and Y diverge systematically for some subgroup G:
E[Z | Y, G] varies with G, even holding Y fixedE[Z | Y, G]— the expected value of the proxy, given the true target and group membership
Whenever this holds, a model optimised purely to predict Z will inherit whatever pattern relates G to the gap between Z and Y, regardless of whether G is an explicit input feature. This is the formal reason "the model doesn't use race as a feature" is not a sufficient fairness argument, demonstrated concretely in the developer block above.
Key references
- Obermeyer, Z., Powers, B., Vyas, V. & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464).
- Ross, C. & Swetlitz, I. (2018). IBM's Watson supercomputer recommended 'unsafe and incorrect' cancer treatments, internal documents show. STAT News.
- Wong, A. et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine.
- Chen, I., Johansson, F. & Sontag, D. (2018). Why Is My Classifier Discriminatory? NeurIPS. A formal treatment of proxy-driven disparity of the kind shown above.
Current state
Regulatory and institutional scrutiny of clinical algorithms has increased substantially since these cases became public, including growing requirements for algorithmic impact assessments before deployment in some jurisdictions. This remains an active, evolving area of policy, not a fully settled one, and readers building anything intended for real clinical use should treat current regulatory guidance, not this lesson, as the authoritative source.
What to learn next
- Bias in datasets — the general framework this section's case studies are instances of.
- Explainability — techniques for surfacing exactly what a deployed model is actually keying on.
- What hospital data actually looks like — returning to the start of this section, now with these failure modes in mind.