Testing robustness to corruptions
Real cameras give you noise, blur, glare and compression, so measure your model against those on purpose before your users do it for you.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 11
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Robustness testing means damaging your test images on purpose, in the ways real cameras do, and measuring what breaks.
The analogy you have lived
Your phone unlocks with your face every morning, indoors, in even light. It works perfectly.
Then you step outside into hard afternoon sun. Half your face is glare, the other half is shadow. It stops recognising you.
The phone did not get worse. The world changed slightly, in a way the phone was never tested against. Your test set was your bedroom.
Why the test set lies to you
Your test images came from the same batch as your training images. Same cameras, same lighting, same photographer, same day, often the same compression settings.
So the test set measures one thing: how well the model handles pictures exactly like the ones it learned from. That is a real question. It is not the question your users are asking.
Users hand you a smudged lens. A dark corridor. A photo through a bus window. A screenshot of a screenshot. None of those appeared in your neat test folder.
The fix, and it is a cheap one
Take your existing test images. Damage each one on purpose, in a specific, controlled way. Then measure accuracy again.
clean test set → accuracy 0.95
|
+-- add camera noise → accuracy ?
+-- blur it slightly → accuracy ?
+-- brighten it → accuracy ?
+-- squeeze it as a JPEG → accuracy ?
+-- shrink and re-enlarge → accuracy ?You need no new labels. The picture of a dog is still a dog after you blur it. That is what makes this test almost free, and it is why every serious vision team runs it.
Doing it at several strengths
One level of damage is not enough. A trace of noise and a blizzard of it are different questions.
The standard approach applies each kind of damage at five increasing strengths, and reports accuracy at each. The shape of that curve is what tells you something.
A model that holds steady across all five and then falls off a cliff has a threshold. A model that slides down gently from the first level has no margin at all.
Reading the results
Do not compare raw accuracy across models. Compare how much worse each model got, relative to its own clean score.
A model starting at 0.98 and dropping to 0.70 lost a lot. A model starting at 0.85 and dropping to 0.80 lost far less. It may be the safer one to deploy. The leaderboard will disagree.
This relative view is the whole point. It separates "how clever is this model" from "how fragile is this model", and those are different questions.
The most useful thing you will find
A large share of robustness failures are not model failures at all. They are preparation failures.
If a model breaks when the picture gets brighter, quite often the model was never told to ignore overall brightness. Fix the preparation step and the fragility disappears, with no retraining.
That is a very good outcome, because it costs an afternoon rather than a fortnight of compute.
Where you have seen this
- Number-plate cameras that work in daylight and fail in rain.
- A document scanner app that reads a crisp printout and gives up on a fax.
- Video calls where the background replacement falls apart in dim light.
- A shop's checkout camera that stops recognising products after the lights are changed.
The honest part
This test measures the damage you thought to apply. It says nothing about the damage you did not think of.
Nobody has a complete list of the ways the world differs from a dataset. A model can pass every corruption test and still fail on a camera model you have never seen.
Treat robustness testing as a floor, never a certificate.
Remember this
- Your test set shares the flaws of your training set. Damage it on purpose to learn more.
- Apply each kind of damage at several strengths, and look at the shape of the fall.
- Compare relative loss between models, not raw accuracy.
What to learn next
- Adversarial attacks — the worst-case cousin of this average-case test.
- Data augmentation — the main lever for improving these numbers.
- Monitoring and drift — catching the corruptions you never thought to test.
Developer — Code and libraries.
Setup
pip install numpy scipy pillowThe example below uses a deliberately simple classifier so the numbers are reproducible and the failure has one obvious cause. The method transfers unchanged to a real network.
Corrupt, measure, then fix the preparation
import io
import numpy as np
from scipy.ndimage import gaussian_filter
from PIL import Image
rng = np.random.default_rng(7)
S = 32
CLASSES = ["disc", "square", "triangle"]
def draw(kind, rng):
"""A bright shape on a dark background, roughly centred, size varying."""
img = np.full((S, S), 0.15)
r = rng.integers(8, 12)
cy, cx = 16 + rng.integers(-2, 3), 16 + rng.integers(-2, 3)
yy, xx = np.mgrid[0:S, 0:S]
if kind == 0:
m = (yy - cy) ** 2 + (xx - cx) ** 2 <= r * r
elif kind == 1:
m = (abs(yy - cy) <= r) & (abs(xx - cx) <= r)
else:
m = (yy - cy >= -r) & (yy - cy <= r) & (abs(xx - cx) <= (yy - cy + r) / 2)
img[m] = 0.85
return np.clip(img + rng.normal(0, 0.02, (S, S)), 0, 1)
def make(n):
y = rng.integers(0, 3, size=n)
return np.stack([draw(k, rng) for k in y]), y
train_x, train_y = make(300)
test_x, test_y = make(300)
def build(x, y, norm):
return np.stack([prep(x[y == c], norm).mean(0) for c in range(3)])
def prep(x, norm):
if not norm:
return x
m = x.mean(axis=(1, 2), keepdims=True)
s = x.std(axis=(1, 2), keepdims=True) + 1e-6
return (x - m) / s # per-image standardisation
def acc(x, y, templates, norm):
z = prep(x, norm)
d = ((z[:, None] - templates[None]) ** 2).sum(axis=(2, 3))
return float((d.argmin(1) == y).mean())
# --- six corruptions, five severities each ---------------------------------
def gaussian_noise(x, s):
return np.clip(x + rng.normal(0, [.04, .08, .12, .18, .26][s-1], x.shape), 0, 1)
def blur(x, s):
return np.stack([gaussian_filter(i, [.6, 1.0, 1.4, 1.9, 2.5][s-1]) for i in x])
def contrast(x, s):
f = [.65, .5, .38, .28, .18][s-1]
return np.clip((x - x.mean()) * f + x.mean(), 0, 1)
def brightness(x, s):
return np.clip(x + [.1, .18, .26, .34, .42][s-1], 0, 1)
def pixelate(x, s):
k = [2, 3, 4, 6, 8][s-1]
out = []
for i in x:
im = Image.fromarray((i * 255).astype(np.uint8))
im = im.resize((S // k, S // k), Image.BOX).resize((S, S), Image.NEAREST)
out.append(np.asarray(im, dtype=np.float64) / 255)
return np.stack(out)
def jpeg(x, s):
q = [40, 25, 15, 10, 7][s-1]
out = []
for i in x:
buf = io.BytesIO()
Image.fromarray((i * 255).astype(np.uint8)).save(buf, "JPEG", quality=q)
out.append(np.asarray(Image.open(buf), dtype=np.float64) / 255)
return np.stack(out)
CORRUPTIONS = {"gaussian_noise": gaussian_noise, "blur": blur,
"contrast": contrast, "brightness": brightness,
"pixelate": pixelate, "jpeg": jpeg}
def report(norm):
templates = build(train_x, train_y, norm)
clean = acc(test_x, test_y, templates, norm)
print(f"clean accuracy {clean:.3f} (error {1-clean:.3f})")
print(f"{'corruption':<16} " + " ".join(f"{'s'+str(s):>6}" for s in range(1, 6))
+ f" {'mean err':>9} {'relative':>9}")
rel = []
for name, fn in CORRUPTIONS.items():
a = [acc(fn(test_x, s), test_y, templates, norm) for s in range(1, 6)]
err = 1 - np.mean(a)
rel.append(err / max(1 - clean, 1e-9))
print(f"{name:<16} " + " ".join(f"{v:>6.3f}" for v in a)
+ f" {err:>9.3f} {rel[-1]:>8.1f}x")
print(f"mean relative corruption error: {np.mean(rel):.1f}x the clean error\n")
print("=== raw pixels ===")
report(norm=False)
print("=== same model, each image standardised first ===")
report(norm=True)=== raw pixels === clean accuracy 0.843 (error 0.157) corruption s1 s2 s3 s4 s5 mean err relative gaussian_noise 0.843 0.843 0.843 0.847 0.847 0.155 1.0x blur 0.843 0.843 0.843 0.843 0.843 0.157 1.0x contrast 0.880 0.827 0.807 0.643 0.407 0.287 1.8x brightness 0.817 0.650 0.397 0.320 0.320 0.499 3.2x pixelate 0.843 0.843 0.843 0.767 0.817 0.177 1.1x jpeg 0.843 0.843 0.843 0.843 0.843 0.157 1.0x mean relative corruption error: 1.5x the clean error === same model, each image standardised first === clean accuracy 0.853 (error 0.147) corruption s1 s2 s3 s4 s5 mean err relative gaussian_noise 0.853 0.847 0.850 0.847 0.840 0.153 1.0x blur 0.853 0.853 0.857 0.840 0.827 0.154 1.0x contrast 0.853 0.853 0.853 0.853 0.853 0.147 1.0x brightness 0.853 0.853 0.853 0.853 0.853 0.147 1.0x pixelate 0.853 0.847 0.847 0.790 0.533 0.226 1.5x jpeg 0.853 0.853 0.847 0.847 0.847 0.151 1.0x mean relative corruption error: 1.1x the clean error
What that table actually says
Brightness costs 3.2 times the clean error, and accuracy hits 0.320 — chance level for three classes. The model is not confused. It is destroyed. Meanwhile noise, blur and JPEG do nothing at all. A single average over corruptions would have reported "1.5x" and hidden a total failure behind five harmless columns. Always print the rows.
The failure curve has a shape worth reading. Brightness goes 0.817, 0.650, 0.397, 0.320, 0.320. It falls fast and then flattens on the floor. Contrast goes 0.880, 0.827, 0.807, 0.643, 0.407: a plateau, then a cliff between severity 3 and 4. Those two curves describe different problems. A flat-then-cliff curve means you have a usable operating range and need to know where its edge is.
Notice contrast at severity 1 scored 0.880, better than clean. Mild contrast reduction pulled the shapes closer to the class templates. Corruption is not monotone, and a single lucky severity level is not evidence of robustness.
The second table is the payoff. The same model, with per-image standardisation added to the preparation step, has brightness and contrast at exactly 1.0x — completely immune. Nothing was retrained. The fragility was never in the model; it was in feeding it absolute pixel values when only relative ones carried the signal.
And there is no free lunch. Pixelation got worse, from 1.1x to 1.5x, and severity 5 fell from 0.817 to 0.533. Standardisation amplifies whatever variation survives, and heavy pixelation leaves blocky variation to amplify. Every robustness fix trades one exposure for another. Measure the whole table after every change, not the one row you were working on.
Applying this to a real model
The synthetic dataset here is a teaching device. The protocol is the real thing, and it is what Hendrycks and Dietterich (ICLR 2019) standardised as ImageNet-C:
pip install imagecorruptionsfrom imagecorruptions import corrupt, get_corruption_names
for name in get_corruption_names(): # 15 standard corruptions
for severity in range(1, 6):
damaged = corrupt(image, corruption_name=name, severity=severity)No output block: it depends on your images and your model. The library implements the exact corruption functions from the paper, so your numbers are comparable to published ones. Rolling your own blur will not be.
The fifteen corruptions are grouped as noise (gaussian, shot, impulse), blur (defocus, glass, motion, zoom), weather (snow, frost, fog, brightness) and digital (contrast, elastic, pixelate, JPEG). Four more — speckle noise, gaussian blur, spatter, saturate — are held out as a validation set, so you can check you did not accidentally tune against the test corruptions.
Common mistakes
Training on the corruptions you test with. ImageNet-C exists to measure generalisation to unseen damage. Augment with the four held-out corruptions if you want to improve robustness; keep the fifteen for measurement only. See data augmentation.
Comparing raw corrupted accuracy across models. A stronger model starts higher and can look more robust while losing more. Divide by each model's own clean error, exactly as the relative column does.
Applying corruption before your production preprocessing instead of in place of the camera. Damage belongs at the point the camera would have introduced it: on the full-resolution image, before resizing. Corrupting the 224-pixel tensor tests something that cannot happen.
Reporting one averaged number. The average is for the leaderboard. The row that fell to chance is for you.
Forgetting compression. JPEG is the corruption every deployed system actually experiences, on every single request, and it is the one most often left out of test suites.
Try it yourself
Add a darkness corruption that subtracts rather than adds. Predict from the raw-pixel table whether it will hurt as much as brightness did. Then add a fixed offset to the training images too and see whether the model learns to ignore brightness on its own, or whether it needs the preparation fix.
What to learn next
- Adversarial attacks — the worst-case cousin of this average-case test.
- Data augmentation — the main lever for improving these numbers.
- Monitoring and drift — catching the corruptions you never thought to test.
Researcher — Mathematics and papers.
The benchmark
Hendrycks and Dietterich (2019), Benchmarking Neural Network Robustness to Common Corruptions and Perturbations, ICLR. ImageNet-C applies 15 corruption functions at 5 severities to the 50,000 ImageNet validation images, with 4 further corruptions held out. CIFAR-10-C, CIFAR-100-C, ImageNet-P (perturbation sequences) and later ImageNet-3DCC follow the same construction.
The headline metric normalises by a reference network, historically AlexNet, so that corruptions of differing intrinsic difficulty contribute comparably:
$$ \mathrm{CE}c^{f} = \frac{\sum{s=1}^{5} E_{s,c}^{f}}{\sum_{s=1}^{5} E_{s,c}^{\text{AlexNet}}}, \qquad \mathrm{mCE}^{f} = \frac{1}{15}\sum_{c} \mathrm{CE}_c^{f} $$
$E_{s,c}^{f}$ is the top-1 error of model $f$ under corruption $c$ at severity $s$. The relative corruption error subtracts the clean error first,
$$ \mathrm{rCE}_c^{f} = \frac{\sum_s \left(E_{s,c}^{f} - E_{\text{clean}}^{f}\right)}{\sum_s \left(E_{s,c}^{\text{AlexNet}} - E_{\text{clean}}^{\text{AlexNet}}\right)} $$
which is the quantity that separates degradation from general accuracy. The runnable table above computes the un-normalised analogue of rCE, dividing by the model's own clean error rather than a reference network's.
Corruption robustness is not adversarial robustness
They are different threat models and they do not transfer.
Adversarial robustness concerns worst-case perturbations inside a small $\ell_p$ ball, found by an optimiser with access to the model. Corruption robustness concerns average-case shifts of large magnitude that are semantically natural. Ford et al. (2019) show the two are related through the model's error surface but empirically decoupled: adversarial training frequently reduces clean accuracy and gives inconsistent corruption gains, and corruption-robust models remain easy to attack. See adversarial attacks for the other half.
What actually improves it
| Method | Mechanism | Note |
|---|---|---|
| AugMix (Hendrycks et al., 2020) | stochastic augmentation chains, mixed, plus a Jensen-Shannon consistency loss | uses only held-out ops |
| DeepAugment | passing images through corrupted image-to-image networks | composes with AugMix |
| PixMix (Hendrycks et al., 2022) | mixing with structurally complex fractal images | improves calibration too |
| Anti-aliased downsampling (Zhang, 2019) | low-pass filter before every stride | fixes shift instability |
| Larger pretraining data | broader natural covariate coverage | the most reliable single factor |
| ViT and ConvNeXt architectures | weaker texture bias, global context | see below |
Geirhos et al. (2019) identified the underlying cause for CNNs: ImageNet-trained networks classify predominantly by texture, not shape, and texture is exactly what noise, blur and compression destroy. Training on stylised ImageNet shifts the bias toward shape and improves corruption robustness directly. This is the most explanatory result in the area.
Bhojanapalli et al. (2021) and Liu et al. (2022) find transformers and modernised CNNs more corruption-robust than classic ResNets at matched compute, though a substantial part of the gap is attributable to training recipe and data scale rather than architecture.
Methodological cautions
Severity levels are not calibrated across corruption types. Severity 3 fog and severity 3 shot noise are not equally hard. Averaging across corruptions at fixed severity is a convention, not a measurement.
The corruptions are synthetic. ImageNet-C's fog is a diamond-square noise field, not photographed fog. Taori et al. (2020), Measuring Robustness to Natural Distribution Shifts, evaluated 200-plus models and found that robustness gains on synthetic benchmarks transfer only weakly to natural shift sets such as ImageNetV2, ObjectNet and ImageNet-R. Synthetic corruption robustness is necessary evidence and nowhere near sufficient.
Compression interacts with the corruption. ImageNet validation images are already JPEG. Applying a corruption and re-encoding compounds two artefacts, and implementations differ in whether they re-encode. This is the same class of problem documented for FID by Parmar et al. (2022).
Accuracy under corruption hides calibration collapse. Confidence is typically far more damaged than accuracy: models stay reasonably accurate while becoming badly overconfident on corrupted inputs. Ovadia et al. (2019), Can You Trust Your Model's Uncertainty?, is the reference result, and it matters more than the accuracy table for any system with a human-review threshold.
Papers
- Hendrycks and Dietterich, Benchmarking Neural Network Robustness to Common Corruptions and Perturbations, ICLR 2019 — arxiv.org/abs/1903.12261
- Geirhos et al., ImageNet-trained CNNs are biased towards texture, ICLR 2019 — arxiv.org/abs/1811.12231
- Zhang, Making Convolutional Networks Shift-Invariant Again, ICML 2019 — arxiv.org/abs/1904.11486
- Hendrycks et al., AugMix, ICLR 2020 — arxiv.org/abs/1912.02781
- Taori et al., Measuring Robustness to Natural Distribution Shifts in Image Classification, NeurIPS 2020 — arxiv.org/abs/2007.00644
- Ovadia et al., Can You Trust Your Model's Uncertainty?, NeurIPS 2019 — arxiv.org/abs/1906.02530
- Hendrycks et al., PixMix, CVPR 2022 — arxiv.org/abs/2112.05135
What to learn next
- Adversarial attacks — the worst-case cousin of this average-case test.
- Data augmentation — the main lever for improving these numbers.
- Monitoring and drift — catching the corruptions you never thought to test.