Adversarial attacks
A change too small for a person to see can flip a model's answer — here is the attack in twenty lines, and the honest state of the defences.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An adversarial attack is a tiny, deliberate change to an input that flips a model's answer, while looking unchanged to a person.
The analogy you have already lived
Think of a scale at a vegetable shop. Press one finger lightly on the corner of the pan while the shopkeeper is looking elsewhere. The vegetables look identical. The reading is now wrong.
You did not break the scale. You found the one place where a small push produces a large change in the number.
A model has thousands of such places. An attacker computes exactly where to push, and how hard.
Why this is different from ordinary noise
Here is the fact that makes this a security problem rather than a quality problem.
Add random specks to a photo and a good model ignores them. Add the same amount of carefully chosen specks and the model can be wrong almost every time.
The experiment below shows exactly this on real handwritten digits. Random noise leaves accuracy at 98%. Noise of identical size, aimed properly, drops it to 23%.
The size of the change is not what matters. The direction is.
Why it happens at all
A model draws a boundary between "this is a 3" and "this is an 8". Real examples sit near that boundary far more often than you would guess. The model reads thousands of input pixels, and in some direction the boundary always passes close by.
The attacker has one advantage that decides everything: they can compute which direction moves the answer fastest. Nudging every pixel a little in that exact direction adds up to a large push, even though no single pixel moved much.
original photo → model → "panda", 58% sure
+
tiny crafted noise (invisible to you)
↓
attacked photo → model → "gibbon", 99% sureThat example is from the 2014 paper that named the problem, and it is still the clearest picture of it.
Where this actually bites
- Stickers placed on a stop sign that make a road-sign classifier read a speed limit.
- Patterns printed on a T-shirt that stop a person being detected by a camera.
- Spam and malware built to slide past a filter, which is the oldest version of this game.
- Audio that sounds like ordinary music to you and issues a command to a voice assistant.
Notice that three of those four are physical. This is not only a digital-file problem.
What is honestly hard here
Ten years of work has not solved this. That statement is not pessimism, it is the state of the field.
Many published defences were later broken by stronger attacks, often within months. The ones that hold up are expensive and cost real accuracy. There is a genuine trade: a more robust model is usually a less accurate model on ordinary inputs.
If someone tells you their model is attack-proof, ask which attacks they tested and how hard they tried.
Remember this
- Small aimed changes beat large random ones, by a huge margin.
- The attacker's advantage is knowing which direction moves the answer.
- Defences are partial, costly, and an active research area, not a settled one.
What to learn next
- Prompt injection — the same idea, aimed at language models.
- Deploying responsibly — the system-level defences that do not need a robust model.
- Image classification — the task these attacks were first built against.
Developer — Code and libraries.
Build the attack, then measure the defence
FGSM — the Fast Gradient Sign Method — is one line of maths and the right starting point. Take the gradient of the loss with respect to the input, take its sign, and step a fixed distance in that direction.
The dataset is scikit-learn's built-in 8x8 digits. It ships inside the library, so nothing downloads and this runs on any laptop.
Setup
pip install numpy scikit-learnThe attack, and the comparison that matters
import numpy as np
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
# 8x8 handwritten digits, ships inside scikit-learn - nothing is downloaded
digits = load_digits()
keep = np.isin(digits.target, [3, 8])
X = digits.data[keep] / 16.0 # pixels scaled to 0..1
y = (digits.target[keep] == 8).astype(int) # 1 = the digit is an 8
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)
clf = LogisticRegression(max_iter=2000).fit(Xtr, ytr)
print(f"clean test accuracy: {clf.score(Xte, yte):.4f} ({len(yte)} images)")
def fgsm(model, X, y, eps):
"""One step along the sign of the loss gradient - Goodfellow et al., 2014."""
w, b = model.coef_[0], model.intercept_[0]
p = 1.0 / (1.0 + np.exp(-(X @ w + b)))
grad = (p - y)[:, None] * w # d(log-loss)/dx for logistic regression
return np.clip(X + eps * np.sign(grad), 0.0, 1.0) # keep pixels in the valid range
print("\nFGSM: every pixel nudged by at most eps, in the worst direction")
print(f"{'eps':>6}{'accuracy':>11}{'max pixel change':>19}{'mean |change|':>16}")
for eps in (0.0, 0.02, 0.05, 0.10, 0.20):
Xadv = fgsm(clf, Xte, yte, eps)
delta = np.abs(Xadv - Xte)
print(f"{eps:>6.2f}{clf.score(Xadv, yte):>11.4f}{delta.max():>19.3f}{delta.mean():>16.4f}")
print("\nRANDOM noise of exactly the same size, for comparison")
rng = np.random.default_rng(0)
for eps in (0.05, 0.10, 0.20):
Xrand = np.clip(Xte + eps * np.sign(rng.normal(size=Xte.shape)), 0.0, 1.0)
print(f" eps={eps:.2f} accuracy = {clf.score(Xrand, yte):.4f}")
# For a linear model the worst-case shift in the score is exactly eps * ||w||_1,
# so shrinking the weights buys robustness - and costs clean accuracy.
print("\nWEIGHT SIZE VS ROBUSTNESS (C is the inverse regularisation strength)")
print(f"{'C':>8}{'||w||_1':>10}{'clean':>9}{'eps=0.10':>10}{'eps=0.20':>10}")
for C in (10.0, 1.0, 0.1, 0.01):
m = LogisticRegression(C=C, max_iter=5000).fit(Xtr, ytr)
print(f"{C:>8}{np.abs(m.coef_[0]).sum():>10.2f}{m.score(Xte, yte):>9.4f}"
f"{m.score(fgsm(m, Xte, yte, 0.10), yte):>10.4f}"
f"{m.score(fgsm(m, Xte, yte, 0.20), yte):>10.4f}")clean test accuracy: 0.9907 (108 images)
FGSM: every pixel nudged by at most eps, in the worst direction
eps accuracy max pixel change mean |change|
0.00 0.9907 0.000 0.0000
0.02 0.9815 0.020 0.0139
0.05 0.9352 0.050 0.0348
0.10 0.8611 0.100 0.0682
0.20 0.2315 0.200 0.1317
RANDOM noise of exactly the same size, for comparison
eps=0.05 accuracy = 0.9907
eps=0.10 accuracy = 0.9907
eps=0.20 accuracy = 0.9815
WEIGHT SIZE VS ROBUSTNESS (C is the inverse regularisation strength)
C ||w||_1 clean eps=0.10 eps=0.20
10.0 53.65 0.9907 0.8241 0.1296
1.0 27.05 0.9907 0.8611 0.2315
0.1 11.94 0.9630 0.8704 0.3426
0.01 3.44 0.9444 0.8796 0.3889The three findings in that output
Aimed noise destroys the model; random noise does not. At eps=0.20, FGSM drops accuracy to 0.2315 — worse than guessing. Random noise of exactly the same per-pixel magnitude leaves it at 0.9815. Same budget, opposite outcome. Any evaluation that tests robustness with random noise has measured nothing.
Pixel changes are small. At eps=0.10 the largest change to any pixel is one tenth of the full brightness range, and the average change is 0.068. On an 8x8 image that is a faint haze.
Robustness is buyable, and it costs accuracy. The last table is the trade written out. Shrinking the weights from ||w||_1 = 53.65 to 3.44 raises accuracy under the strong attack from 0.1296 to 0.3889, and drops clean accuracy from 0.9907 to 0.9444. Neither column can be maximised alone.
For a linear model that trade is exact rather than empirical. The attack shifts the score by at most eps * ||w||_1, so a smaller weight vector means a smaller worst-case shift. Deep networks have no such clean formula, but the same tension appears.
Line by line
grad = (p - y)[:, None] * w is the gradient of the log-loss with respect to the input, not the weights. For logistic regression it has this closed form. In PyTorch you would set x.requires_grad_(True), call loss.backward(), and read x.grad.
np.sign(grad) throws away the magnitude and keeps only the direction, per pixel. This is what makes it an $L_\infty$ attack: every pixel moves by exactly eps, so the largest single change is bounded.
np.clip(..., 0.0, 1.0) keeps the result a valid image. Skipping this produces adversarial examples with negative brightness that cannot be saved to a file, which is how people accidentally report attack success rates that do not survive contact with a real image pipeline.
stratify=y keeps the class balance identical in both splits, so the 108-image test set is not accidentally lopsided.
Common mistakes
Testing with FGSM only. It is a single gradient step and the weakest standard attack. A defence that survives FGSM and fails PGD is broken. Use a maintained suite: torchattacks, foolbox or AutoAttack.
Gradient masking. Adding a rounding step or a randomised input transform makes gradients uninformative, so gradient attacks fail, and the model looks robust. Athalye et al. (2018) broke seven of nine ICLR defences this way. If your defence dramatically beats the attack but a black-box transfer attack still works, you have masked gradients rather than gained robustness.
Reporting robustness without the threat model. "Robust" is meaningless without the norm, the radius, and whether the attacker can see your weights. State all three.
Assuming a private model is safe. Adversarial examples transfer between models trained on similar data. An attacker can craft on their own copy and send it to yours.
Forgetting the cheap non-model defences. Rate limiting, input logging, confidence thresholds and human review on high-stakes decisions all raise the attacker's cost without touching the model.
Try it yourself
Add a second attack that steps three times with eps/3 each time, re-computing the gradient between steps. That is a miniature PGD. Compare its accuracy at a fixed budget against single-step FGSM, and you will see why single-step results overstate robustness.
What to learn next
- Prompt injection — the same idea, aimed at language models.
- Deploying responsibly — the system-level defences that do not need a robust model.
- Image classification — the task these attacks were first built against.
Researcher — Mathematics and papers.
Threat model first
A robustness claim without a threat model is not a claim. Specify:
- Perturbation set $\Delta$. Usually an $\ell_p$ ball ${\delta : |\delta|_p \le \varepsilon}$, with $p \in {0, 1, 2, \infty}$. Also: spatial transformations, colour shifts, physically realisable patches.
- Adversary knowledge. White-box (weights and gradients), black-box score-based (probabilities), black-box decision-based (labels only), or transfer-only.
- Adversary capability. Query budget, ability to modify the physical scene, ability to poison training data.
$\ell_p$ balls are a tractable proxy for "perceptually similar", not a definition of it. Robustness inside an $\ell_\infty$ ball says nothing about robustness to a small rotation.
The objective
Madry et al. (2018) frame robust training as a saddle-point problem:
$$ \min_{\theta} \; \mathbb{E}{(x,y) \sim \mathcal{D}} \Big[ \max{\delta \in \Delta} \; \mathcal{L}\big(f_\theta(x + \delta),\, y\big) \Big] $$
$\theta$ are the model parameters, $\mathcal{D}$ the data distribution, $\mathcal{L}$ the loss, $\Delta$ the perturbation set. The inner maximisation is the attack; the outer minimisation is adversarial training. The inner problem is non-concave, so it is solved approximately, and the quality of that approximation determines whether the outer solution is meaningfully robust.
Attacks
FGSM (Goodfellow et al., 2014): $x' = x + \varepsilon\,\text{sign}(\nabla_x \mathcal{L})$. One step. Exact for linear models — the demonstration in the developer block — and a weak approximation otherwise.
PGD (Madry et al., 2018): iterate $x^{t+1} = \Pi_{\mathcal{B}(x,\varepsilon)}\big(x^t + \alpha\,\text{sign}(\nabla_x \mathcal{L})\big)$, with $\Pi$ the projection back onto the ball and random restarts. Treated as the standard first-order attack.
C&W (Carlini and Wagner, 2017): minimise $|\delta|_2 + c \cdot g(x + \delta)$ with a margin-based $g$ optimised in a change of variables that removes the box constraint. Finds much smaller perturbations than PGD and broke defensive distillation.
AutoAttack (Croce and Hein, 2020): an ensemble of APGD with two losses, FAB and Square Attack, parameter-free. The current default for reporting, and RobustBench standardises on it.
Black-box. Square Attack (Andriushchenko et al., 2020) is score-based and random-search driven. Boundary Attack (Brendel et al., 2018) needs only decision labels. Transfer attacks exploit the shared decision boundaries documented by Papernot et al. (2016).
Why adversarial examples exist
Two accounts, both supported.
Linearity (Goodfellow et al., 2014). For a linear model the logit shift under an $\ell_\infty$ perturbation is exactly $\varepsilon |w|_1$, which grows with input dimension. Modern networks are locally close enough to linear for the argument to bite.
Non-robust features (Ilyas et al., 2019). Adversarial vulnerability arises from genuinely predictive but brittle features present in the data. Their key experiment: a dataset relabelled purely by adversarial perturbations still trains a model that generalises to the clean test set, showing the perturbations carry real signal. This reframes robustness as a choice about which features to use, not as a defect to patch.
Defences, and the evaluation problem
Adversarial training is the only broadly surviving defence. Madry et al. (2018) train on PGD examples. Costs: roughly $k\times$ training compute for $k$ inner steps, and a consistent clean-accuracy drop. Zhang et al. (2019), TRADES, make the accuracy-robustness trade explicit through a decomposed loss.
Certified defences give provable guarantees on a radius. Randomised smoothing (Cohen et al., 2019) constructs $g(x) = \arg\max_c \Pr_{\eta \sim \mathcal{N}(0,\sigma^2 I)}[f(x + \eta) = c]$, certified in $\ell_2$ with radius $\frac{\sigma}{2}(\Phi^{-1}(p_A) - \Phi^{-1}(p_B))$ where $p_A, p_B$ are the top two class probabilities under noise and $\Phi^{-1}$ the inverse Gaussian CDF. Scales to ImageNet; certified radii remain small.
The methodological warning. Athalye, Carlini and Wagner (2018), Obfuscated Gradients Give a False Sense of Security, broke seven of nine defences accepted at ICLR 2018 within weeks. Carlini et al. (2019) publish an evaluation checklist. The diagnostic signals of gradient masking are concrete: unbounded attacks fail to reach 100% success, black-box attacks beat white-box attacks, iterative attacks do not beat single-step attacks, and increasing $\varepsilon$ does not monotonically increase success.
RobustBench (Croce et al., 2021) maintains standardised leaderboards under AutoAttack. Current CIFAR-10 $\ell_\infty$ robustness at $\varepsilon = 8/255$ sits well below clean accuracy, and progress has been incremental for years. Treat any claim far above the leaderboard as an evaluation artefact until independently reproduced.
Physical and non-vision attacks
Eykholt et al. (2018) achieved high misclassification rates on physical stop signs with printed stickers under varied distances and angles. Athalye et al. (2018) produced 3D-printed objects misclassified from most viewpoints. Sharif et al. (2016) used printed eyeglass frames to impersonate identities in face recognition.
Beyond vision, the same framework covers audio (Carlini and Wagner, 2018), malware detection, and — for language models — prompt injection, where the perturbation set is natural-language text and no $\ell_p$ ball applies.
Papers
- Szegedy et al., Intriguing Properties of Neural Networks, 2013 — arxiv.org/abs/1312.6199
- Goodfellow, Shlens and Szegedy, Explaining and Harnessing Adversarial Examples, 2014 — arxiv.org/abs/1412.6572
- Madry et al., Towards Deep Learning Models Resistant to Adversarial Attacks, ICLR 2018 — arxiv.org/abs/1706.06083
- Athalye, Carlini and Wagner, Obfuscated Gradients Give a False Sense of Security, ICML 2018 — arxiv.org/abs/1802.00420
- Ilyas et al., Adversarial Examples Are Not Bugs, They Are Features, NeurIPS 2019 — arxiv.org/abs/1905.02175
- Cohen, Rosenfeld and Kolter, Certified Adversarial Robustness via Randomized Smoothing, ICML 2019 — arxiv.org/abs/1902.02918
- Croce and Hein, Reliable Evaluation of Adversarial Robustness (AutoAttack), ICML 2020 — arxiv.org/abs/2003.01690
What to learn next
- Prompt injection — the same idea, aimed at language models.
- Deploying responsibly — the system-level defences that do not need a robust model.
- Image classification — the task these attacks were first built against.