Why your model fails at the next hospital
A model trained on one hospital's scanner can quietly learn that scanner's look, not the disease, and fail the moment it meets a different machine.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A model trained on one hospital's scans can fail on a different hospital's scans of the exact same condition.
Think about learning to read a teacher's handwriting all year. You get good at it. A substitute teacher then writes on the board, in a completely different style. Suddenly you are struggling again — even though the words mean the same thing.
Medical scanners have their own "handwriting." Different machines, different settings, different hospitals — the same disease can look subtly different in every one of them.
Why it exists
Every scanner has its own brightness, contrast, and noise pattern. These are shaped by its manufacturer, its exact settings, and how staff at that hospital use it. None of this has anything to do with the disease being diagnosed.
A model trained on images from one hospital can accidentally learn to key on these machine-specific quirks. It might learn them alongside the real medical signal — or worse, instead of it. It looks accurate during testing, because the test images came from the very same machines. Move that model to a new hospital, with different equipment, and the quirks it relied on disappear.
How it works
Hospital A's scanner -> its own particular "look" -> model trained here
|
v
looks accurate
Hospital B's scanner -> a different "look" -> same model, same disease
|
v
accuracy quietly dropsNothing about the disease changed between the two hospitals. Only the machine did.
Where you have already seen it
- A photo filter that shifts colors can make the same food look different, across two different phone cameras.
- Voice recognition trained on one microphone often struggles on a different microphone, even hearing the exact same words.
An honest warning
This failure is dangerous specifically because it is silent. A model does not announce that it has started relying on the wrong signal. It performs worse, in a new setting, with no built-in warning. That is exactly why testing at more than one hospital matters, before real deployment. Regulators and radiologists decide when it is ready to move to a new site — not the model's own reported accuracy.
Remember this
- A model can accidentally learn a scanner's specific look, instead of, or alongside, the real medical signal.
- Strong performance at one hospital does not guarantee strong performance at another.
- This kind of failure gives no warning on its own — it has to be actively tested for.
What to learn next
- Dataset bias and shortcut learning — the general version of this exact failure mode.
- Domain adaptation for vision — techniques built specifically to address it.
- Generating radiology reports — a different task, but one that inherits this same risk.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyMinimal runnable code
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
# Handwritten digits stand in for medical images here -- small, built into
# scikit-learn, and the point (a model tied to one machine's "look") does
# not need a real scanner to demonstrate.
digits = load_digits()
X, y = digits.data, digits.target
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)
# "Hospital A" scanner: images as they come, scaled to [0, 1].
X_train_a = X_train / 16.0
X_test_a = X_test / 16.0
# "Hospital B" scanner: a different make of machine, with different contrast
# and brightness settings. Nothing about the digits themselves changed.
def hospital_b_style(X):
X = X / 16.0
X = X * 0.6 + 0.25 # lower contrast, brighter baseline
return np.clip(X, 0, 1)
X_test_b = hospital_b_style(X_test)
model = RandomForestClassifier(n_estimators=200, random_state=0)
model.fit(X_train_a, y_train)
acc_same_hospital = accuracy_score(y_test, model.predict(X_test_a))
acc_other_hospital = accuracy_score(y_test, model.predict(X_test_b))
print(f"accuracy on the SAME scanner style it was trained on: {acc_same_hospital:.3f}")
print(f"accuracy on a DIFFERENT scanner's images, same patients: {acc_other_hospital:.3f}")
print("nothing about the underlying digits changed -- only brightness and contrast did.")accuracy on the SAME scanner style it was trained on: 0.978 accuracy on a DIFFERENT scanner's images, same patients: 0.756 nothing about the underlying digits changed -- only brightness and contrast did.
What actually happened
Accuracy drops from 97.8% to 75.6% — a real, substantial fall — using the exact same trained model, on the exact same underlying digits. Only the brightness and contrast changed, standing in for a different scanner's settings.
Nothing here involved a harder task, or a different disease, or noisier data in any meaningful sense. The model had not seen this particular "look" before, and part of what it had learned turned out to be tied to the specific appearance of its training images.
Line by line, the parts that are not obvious:
hospital_b_styleis a deliberately simple, global transform — real inter-scanner differences are usually more complex than a linear brightness and contrast shift, which makes this toy example, if anything, an optimistic case.- The model is never retrained or shown any Hospital B images. This is a pure generalisation test, measuring only what the original model already learned.
RandomForestClassifierwas chosen for speed here, but this failure mode is not specific to any one model type — it can affect any model trained on data from a narrow source.
Common mistakes
Validating a medical imaging model only on data from the same source as training. A high validation score under these conditions says nothing about performance at a different hospital, with different equipment.
Assuming a bigger model or a fancier architecture fixes this by itself. Model capacity does not address a training data limitation — a model can only be as source-diverse as the data it saw.
Fixing this with a single global correction, once, at deployment time. Different institutions can differ in more ways than brightness and contrast alone — resolution, positioning conventions, and patient population all vary too, and a single fixed correction rarely covers all of them.
Try it yourself
Make the shift more severe: change X * 0.6 + 0.25 to X * 0.3 + 0.4. Rerun and watch how much further accuracy falls — a rough proxy for how much a training set's source diversity matters.
What to learn next
- Dataset bias and shortcut learning — the general pattern this lesson's example demonstrates.
- Domain adaptation for vision — techniques for closing this exact gap.
- Validating a model before it touches patients — why external validation, at a different site, matters before deployment.
Researcher — Mathematics and papers.
Covariate shift and domain generalization
Formally, this failure is an instance of covariate shift: the input distribution changes between training and deployment, P_train(x) != P_deploy(x), while the true relationship between input and label, P(y|x), is assumed to stay approximately constant. Domain generalization is the broader problem of training a model on data from one or more source domains such that it generalises to an unseen target domain, without access to any target-domain data at training time — the exact setting the developer block's Hospital B represents.
A landmark negative result
Zech et al. (2018, PLOS Medicine) trained a pneumonia-detection CNN on chest X-rays from multiple hospital systems and found that performance varied substantially by hospital of origin — not because pneumonia looked different, but because the model had partly learned to detect hospital-specific image characteristics (including, in some analyses, portable-versus-fixed X-ray equipment markers that correlated with patient severity at a given site) rather than purely the pathology itself. This is among the most widely cited empirical demonstrations that a medical imaging model's validation accuracy at its training site does not reliably predict performance elsewhere.
Harmonisation approaches
For MRI specifically, where scanner and protocol differences are especially pronounced, ComBat (Fortin et al., 2017, adapting an earlier batch-effect correction method from genomics) is a widely used statistical harmonisation technique. It models site-specific additive and multiplicative effects on extracted imaging features and removes them, while explicitly preserving effects of interest such as age or diagnosis:
y_ijv = alpha_v + X_ij*beta_v + gamma_iv + delta_iv*epsilon_ijvy_ijv— the observed featurevfor subjectjat siteiX_ij*beta_v— the effect of biological covariates of interest (preserved, not removed)gamma_iv,delta_iv— the site-specific additive and multiplicative effects (estimated and removed)
Deep-learning-native approaches include domain-adversarial training (Ganin et al., 2016), which trains a feature extractor to simultaneously predict the task label while being unable to predict which site an example came from, encouraging site-invariant representations.
Cost
Domain generalization techniques generally trade some in-distribution accuracy for out-of-distribution robustness — there is no free lunch here either, and a model tuned aggressively for cross-site robustness can underperform a narrowly trained model when deployed back at its original source site. Multi-site external validation, rather than any single technique, remains the primary tool for detecting this trade-off before deployment.
Key references
- Zech, J. et al. (2018). Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs. PLOS Medicine 15(11).
- Fortin, J. et al. (2017). Harmonization of multi-site diffusion tensor imaging data. NeuroImage 161.
- Ganin, Y. et al. (2016). Domain-Adversarial Training of Neural Networks. JMLR 17(1).
- AlBadawy, E., Saha, A. & Mazurowski, M. (2018). Deep learning for segmentation of brain tumors: Impact of cross-institutional training and testing. Medical Physics 45(3). A direct demonstration of cross-institutional performance drop in segmentation.
Current state
Multi-site, multi-scanner external validation is increasingly expected in published medical imaging AI research, and is a standard requirement in regulatory review, though historically it has been inconsistently reported. This remains an active area of methodological standardisation, connected directly to the prospective validation practices covered in validating a model before it touches patients.
What to learn next
- Dataset bias and shortcut learning — the general framework this lesson's failure mode is an instance of.
- Domain adaptation for vision — the general techniques this lesson's citations draw from.
- Validating a model before it touches patients — where external, multi-site validation fits into a real deployment process.