Vision Datasets and Annotation

Dataset bias and shortcut learning

A model takes the easiest route to a correct answer, and if an accidental clue in your data is easier than the real feature, that clue is what it learns.

On this page 10
  1. The short answer
  2. The analogy
  3. Why models do this
  4. Real cases, not made-up ones
  5. The clue that decides everything
  6. How to catch it
  7. The uncomfortable part
  8. Where you have seen this
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A model learns whatever gets the right answer most easily, and that is often not the thing you meant.

The analogy

Think about a child learning to tell cows from horses, from a picture book.

Every cow in the book stands in a green field. Every horse stands on brown sand.

The child scores full marks. Then you show them a horse in a green field, and they say cow.

They never learned about the animals. They learned about the ground. And the ground was easier.

Why models do this

A model has no idea what you meant. It has pictures and answers, and it finds the shortest path between them.

If the background predicts the label, the background is the answer. Backgrounds are large, uniform and easy to see. Animal shapes are complicated.

The model is not being lazy or stupid. It solved the problem you actually posed.

Real cases, not made-up ones

Hospital scanners. A model detecting lung disease from chest images learned to recognise which hospital took the image. One hospital had more sick patients than the other. It scored beautifully, then failed at a new hospital.

Wolves and huskies. A famous demonstration built a classifier that looked at whether there was snow in the picture. Wolves were photographed in snow. Huskies were not.

Rulers in skin images. Skin cancer photographs taken by specialists often include a measuring ruler. Healthy skin photos do not. So the model learned to look for rulers.

None of these are exotic. They are what happens by default.

The clue that decides everything

Here is the part that makes this predictable rather than mysterious.

The model takes the shortcut only when the shortcut is easier than the real thing.

If the real feature is clear and simple, the model learns it and ignores the clue. If the real feature is hard, noisy or small, the model reaches for the clue instead.

You will see this crossover measured in the next section, from clean to noisy, and the switch is sharp.

How to catch it

Test somewhere different. New hospital, new camera, new city, new lighting. A big drop means a shortcut.

Break the clue on purpose. Paint over the corner, blank the background, shuffle the thing you suspect. If accuracy collapses, you found it.

Look at what the model looks at. Tools exist that highlight which pixels drove a decision. If they highlight the background, you have your answer.

The uncomfortable part

Fixing this makes your reported number go down.

The shortcut was inflating your score. Remove it and the honest score is lower, sometimes much lower.

That is a difficult conversation to have with whoever is expecting the higher number. Have it early, because the alternative is having it after deployment.

Where you have seen this

  • A phone feature that works at your house and not at your friend's.
  • A medical study that did not repeat in a different hospital.
  • A tool that works well for some groups of people and badly for others.

Remember this

  • Models take the easiest route to the right answer, not the one you intended.
  • Shortcuts appear when the accidental clue is easier than the real feature.
  • Removing a shortcut makes your reported number fall, and that is the honest number.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch==2.5.1

Making shortcut learning appear and disappear on demand

Two shapes to classify. A bright corner patch agrees with the class 95 percent of the time, which is the shortcut. Noise is added to the shapes but never to the corner patch. So the difficulty of the real feature is a dial, and the shortcut stays easy.

shortcut.py
import torch, torch.nn as nn, time
S = 16

def make(n, cue_agrees, noise, seed):
    """Half squares, half bars. A clean 3x3 corner patch is the SHORTCUT.
    `noise` makes the real feature (the shape) harder to read. The cue stays clean."""
    g = torch.Generator().manual_seed(seed)
    grid = torch.arange(S).float()
    yy, xx = torch.meshgrid(grid, grid, indexing="ij")
    y = torch.randint(0, 2, (n,), generator=g)
    c = torch.randint(6, 11, (n, 2, 1, 1), generator=g).float()
    dx, dy = xx - c[:, 0], yy - c[:, 1]
    img = torch.where(y[:, None, None].bool(),
                      ((dx.abs() < 3) & (dy.abs() < 3)).float(),      # class 1: a filled square
                      ((dx.abs() < 4) & (dy.abs() < 1)).float())      # class 0: a thin bar
    img = img + noise * torch.randn(img.shape, generator=g)
    cue = torch.where(torch.rand(n, generator=g) < cue_agrees, y, 1 - y)
    img[:, :3, :3] = torch.where(cue[:, None, None].bool(), 1.0, 0.0)  # the cue is never noisy
    return img.unsqueeze(1), y

def train(cue_agrees, noise, steps=300):
    torch.manual_seed(0)
    m = nn.Sequential(nn.Conv2d(1, 8, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
                      nn.Conv2d(8, 16, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
                      nn.Flatten(), nn.Linear(16 * 4 * 4, 2))
    opt = torch.optim.Adam(m.parameters(), lr=3e-3)
    for s in range(steps):
        x, y = make(128, cue_agrees, noise, seed=1000 + s)
        loss = nn.functional.cross_entropy(m(x), y)
        opt.zero_grad(); loss.backward(); opt.step()
    return m

@torch.no_grad()
def acc(m, x, y): return (m(x).argmax(1) == y).float().mean().item()

t0 = time.time()
print("How hard is the real feature?   test where the cue     test where the cue")
print("(noise on the shape)            still agrees (95%)     is reversed (5%)")
for noise in (0.0, 0.3, 0.6, 1.0, 1.5):
    m = train(0.95, noise)
    a1 = acc(m, *make(2000, 0.95, noise, 7))
    a2 = acc(m, *make(2000, 0.05, noise, 9))
    tag = "  <- ignored the cue" if a2 > 0.8 else ("  <- rides the cue" if a2 < 0.3 else "")
    print(f"{noise:24.1f} {a1:22.1%} {a2:22.1%}{tag}")

print("\nSame data, but the cue is randomised during training (agrees only 50% of the time):")
for noise in (0.6, 1.0, 1.5):
    m = train(0.50, noise)
    print(f"{noise:24.1f} {acc(m, *make(2000, 0.95, noise, 7)):22.1%} "
          f"{acc(m, *make(2000, 0.05, noise, 9)):22.1%}")
print(f"\nseconds: {time.time()-t0:.0f}")
Output
How hard is the real feature?   test where the cue     test where the cue
(noise on the shape)            still agrees (95%)     is reversed (5%)
                     0.0                 100.0%                 100.0%  <- ignored the cue
                     0.3                 100.0%                  99.9%  <- ignored the cue
                     0.6                  99.8%                  98.4%  <- ignored the cue
                     1.0                  97.4%                  73.2%
                     1.5                  94.5%                  22.8%  <- rides the cue

Same data, but the cue is randomised during training (agrees only 50% of the time):
                     0.6                  99.4%                  99.4%
                     1.0                  92.4%                  94.5%
                     1.5                  81.8%                  85.3%

seconds: 14

Exact figures depend on your PyTorch build. The crossover, and its position, reproduce.

Reading the output

At low noise the model ignores the shortcut completely. At noise 0.0, 0.3 and 0.6, accuracy is essentially identical whether the cue agrees or actively lies. The cue was available, correlated at 95 percent, and free to use. The model did not use it, because the shape was easier to read.

This is the finding that makes shortcut learning tractable. It is not that models prefer shortcuts. It is that models prefer whatever is easier, and a clean, simple real feature wins.

At noise 1.5 the model has switched. 94.5% when the cue agrees, 22.8% when it is reversed. Worse than chance, which means it is actively following the cue against the evidence. The reported headline number, 94.5%, is almost entirely the cue.

Noise 1.0 is the crossover. 97.4% and 73.2%: partly shape, partly cue. Real datasets sit here. The symptom is usually a moderate unexplained drop on new data, not a total collapse.

Randomising the cue during training fixes it, and costs you the headline. At noise 1.5, the honest model scores 81.8% and 85.3%, statistically the same under both test conditions. Compare that against the shortcut model's 94.5%.

Read those two numbers together. Fixing the shortcut moved the reported accuracy from 94.5% down to 81.8%. Nothing got worse. The 94.5% was never real. That conversation is the hard part of this work. It is easier to have when you can show a table.

The four diagnostics

Test on shifted data. Different camera, hospital, city, season. The most informative test you can run, and the one most projects never run.

Ablate the suspect. Blank it, blur it, or randomise it, and re-measure. The output above is exactly this test.

Train on the suspect alone. Crop out only the corner patch and train a classifier on that. If it scores well, the cue is sufficient on its own. Your main model is very likely using it.

Attribution maps. Grad-CAM or similar, to see which pixels drove the decision. Useful for generating hypotheses, and weak as proof. Several saliency methods pass sanity checks poorly. A map highlighting the object does not establish that the object was the cause. See SHAP and LIME for the general caveats.

The four fixes, in order of how well they work

Fix the data. Break the correlation at collection time. Photograph both classes in both settings. This is the only fix that removes the problem rather than managing it.

Augment away the cue. If the cue is background, colour or texture, augmentation that randomises it works well. If the cue is the object's own shape or position, augmentation cannot reach it.

Balance by group. If you know the confounding variable, reweight or resample. Make it independent of the label within each training batch.

Change the objective. Group DRO optimises the worst-performing group rather than the average, which directly targets the failure. It needs group labels at training time, which is a strong requirement.

Sagawa et al. (2020) report a finding worth internalising: strong regularisation is necessary for group DRO to help in over-parameterised networks. Without it, the model achieves near-zero training loss on every group. The worst-group objective then has nothing left to optimise.

Common mistakes

Assuming a high test score means no shortcut. The test set carries the same correlation as the training set. It cannot detect this by construction. Only shifted data can.

Testing on shifted data and concluding the model needs more capacity. More capacity fits the shortcut better.

Trusting a saliency map as proof. It is a hypothesis generator. Confirm by ablation.

Removing a shortcut and quietly comparing against the old number. Report both, and say which is honest.

Treating this as only a fairness problem. It is a generalisation problem that frequently has fairness consequences, because demographic attributes are common accidental cues. Both framings are correct and the technical fix is the same.

Try it yourself

Move the cue from a corner patch into the shape itself: make class 1 squares one pixel larger. Now augmentation cannot remove it and cropping cannot reach it. That is what a genuinely entangled shortcut looks like. It is why data collection is the fix that matters.

What to learn next

Researcher — Mathematics and papers.

The framing

Geirhos et al. (2020) define shortcuts as decision rules that perform well on standard benchmarks. They fail to transfer to more challenging testing conditions, such as real-world scenarios. Their paper is Shortcut Learning in Deep Neural Networks, in Nature Machine Intelligence. Their argument is that many apparently separate failures of deep learning share this single cause. They draw a parallel to comparative psychology, education and linguistics. Shortcut learning may be a general characteristic of learning systems, not a quirk of neural networks.

The mechanism is a mismatch. The intended solution set differs from the set of solutions consistent with the training data. Any decision rule achieving low training loss is admissible. The optimiser selects among them by inductive bias and ease of optimisation, not by semantic correctness. The developer results above make that selection criterion visible. The cue was available and unused when the intended feature was easier to fit. It was adopted when it was not.

Documented instances

DomainShortcutReference
Object recognitionTexture over shapeGeirhos et al., 2019
Chest radiographyHospital and scanner identityZech et al., 2018
COVID-19 imagingPatient position, source dataset, text markersDeGrave et al., 2021
DermatologySurgical skin markings, rulersWinkler et al., 2019
Object detectionContext and co-occurrenceChoi et al., 2012
Natural language inferenceLexical overlap heuristicsMcCoy et al., 2019

Geirhos et al. (2019), ImageNet-trained CNNs are biased towards texture, is the sharpest single result. They use cue-conflict images: a cat's shape with an elephant's texture. Standard ImageNet CNNs classify predominantly by texture, while humans classify predominantly by shape. Training on Stylized-ImageNet, where texture is randomised, shifts the bias toward shape and improves robustness to common corruptions.

DeGrave et al. (2021) applied explanation methods and generative adversarial ablation to published COVID-19 chest radiograph models. They traced performance to laterality markers, patient positioning and dataset source, not to lung pathology. Their paper is AI for radiographic COVID-19 detection selects shortcuts over signal. External-validation performance collapsed accordingly.

Why the optimiser prefers shortcuts

Simplicity bias. Networks trained by gradient descent are biased toward simple functions, and this bias can be extreme. Shah et al. (2020), The Pitfalls of Simplicity Bias in Neural Networks, construct datasets with two perfectly predictive features. One is simple and one is complex. Networks rely exclusively on the simpler one, even when that destroys robustness and margin.

Learning order. Nam et al. (2020) show bias-aligned examples are learned first. They characterise "easy" bias as bias learned early in training. Their debiasing method trains a second network on examples the first gets wrong.

Gradient starvation. Pezeshki et al. (2021) analyse how a dominant feature suppresses the gradient signal for other predictive features. The model then never learns them at all. This is the mechanism that makes shortcut adoption self-reinforcing rather than a matter of degree.

The developer table is a small empirical instance of all three. As the intended feature's signal-to-noise falls, the cue becomes the simpler and earlier-learned solution. The model then commits to it.

Mitigation, and its limits

Data intervention is the only approach that removes the confound rather than managing it. Balanced collection, or targeted collection of counter-examples where the cue and label disagree. Expensive and correct.

Group DRO (Sagawa et al., 2020) minimises worst-group loss:

$$ \min_{\theta} \; \max_{g \in \mathcal{G}} \; \mathbb{E}_{(x,y) \sim P_g}\left[\ell(\theta; x, y)\right] $$

$\mathcal{G}$ is a partition of the data by the confounding attribute. The paper's central practical result concerns regularisation. Strong regularisation, through early stopping or a large weight-decay penalty, is necessary in over-parameterised models. An unregularised network drives training loss to zero on every group, and the worst-group objective becomes vacuous.

Without group labels. JTT (Liu et al., 2021) trains an initial model, upweights the examples it misclassifies, and retrains. LfF (Nam et al., 2020) trains a deliberately biased model with generalised cross entropy. It then reweights by relative difficulty. Both approximate group DRO using the model's own errors as a proxy for the unknown group.

Invariant risk minimisation (Arjovsky et al., 2019) seeks a representation whose optimal classifier is identical across training environments. Theoretically attractive; Rosenfeld et al. (2021) show the practical objective can fail to recover the invariant predictor, except under restrictive conditions. It frequently underperforms plain ERM on benchmarks.

Benchmarks. WILDS (Koh et al., 2021) collects real distribution shifts across domains. Its consistent finding is that no method reliably closes the gap. Treat any method claiming to solve this generally with the corresponding scepticism.

Reporting standard

An honest evaluation of a vision model includes:

  • At least one out-of-distribution test set, drawn from a different site, device, time period or population. Report it alongside the in-distribution number.
  • Metrics disaggregated by every recorded attribute, not only aggregate accuracy. A worst-group number is more informative than a mean.
  • At least one explicit ablation of a hypothesised shortcut, with the accuracy under ablation reported.
  • The in-distribution to out-of-distribution gap stated as a number, since that gap is the quantity a reader actually needs.

The default reporting practice is a single accuracy on a randomly split test set. It cannot detect any of the failures on this page. It is a measurement instrument that is blind to the most common way these systems fail.

References

  • Torralba and Efros, Unbiased Look at Dataset Bias, CVPR 2011
  • Zech et al., Variable Generalization Performance of a Deep Learning Model to Detect Pneumonia in Chest Radiographs, PLOS Medicine, 2018 — arxiv.org/abs/1807.00431
  • Geirhos et al., ImageNet-trained CNNs are biased towards texture, ICLR 2019 — arxiv.org/abs/1811.12231
  • Arjovsky et al., Invariant Risk Minimization, 2019 — arxiv.org/abs/1907.02893
  • Geirhos et al., Shortcut Learning in Deep Neural Networks, Nature Machine Intelligence 2, 665-673, 2020 — arxiv.org/abs/2004.07780
  • Sagawa et al., Distributionally Robust Neural Networks for Group Shifts, ICLR 2020 — arxiv.org/abs/1911.08731
  • Shah et al., The Pitfalls of Simplicity Bias in Neural Networks, NeurIPS 2020 — arxiv.org/abs/2006.07710
  • Nam et al., Learning from Failure, NeurIPS 2020 — arxiv.org/abs/2007.02561
  • DeGrave et al., AI for Radiographic COVID-19 Detection Selects Shortcuts Over Signal, Nature Machine Intelligence, 2021
  • Koh et al., WILDS: A Benchmark of in-the-Wild Distribution Shifts, ICML 2021 — arxiv.org/abs/2012.07421
  • Liu et al., JTT: Improving Group Robustness without Training Group Information, ICML 2021 — arxiv.org/abs/2107.09044
  • Pezeshki et al., Gradient Starvation, NeurIPS 2021 — arxiv.org/abs/2011.09468

What to learn next