Data augmentation
Augmentation makes new training examples by changing old ones in ways the real world also changes them, and a transform that does not match reality makes the model worse.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Data augmentation means making new training examples by changing the ones you already have, in the ways the real world changes them too.
Think about learning to recognise a friend's face. You have seen them in bright sun, in a dark room, wearing sunglasses, side-on, laughing.
That is why you would still know them in a badly lit photo. You did not memorise one picture. You saw the same face under many conditions.
Augmentation gives a model the same variety, from a small number of original examples.
Why it exists
Labelled examples are expensive. Someone has to look at each one and write down the answer.
Augmentation is a way to squeeze more learning out of what you already paid for. Take one photo, produce a slightly rotated copy, a slightly darker copy, a slightly cropped copy. Now you have four training examples from one label.
The rule that decides everything
The change must be one the real world also makes, and it must not change the answer.
A photo of a cat, rotated a little, is still a cat. Good augmentation.
A photo of the digit 2, flipped left to right, is not a 2 any more. Bad augmentation, and it will make your model worse than doing nothing at all.
That difference is the whole subject.
How it looks
one photo of a cat
|
+--> slightly rotated -> still a cat [keep]
+--> darker -> still a cat [keep]
+--> cropped a bit -> still a cat [keep]
+--> flipped upside down -> a cat, upside down
do photos of cats
arrive upside down?
if no -> [discard]The last branch is the one to think about. The question is never "is this change allowed". It is "does this change happen in the data my model will meet".
Somewhere you have seen this
Any phone app that reads a document works this way. Real users hold the phone at an angle, in bad light, with a thumb over one corner. The training data had to contain all of that, and most of it was made by augmentation.
Voice assistants too. Recordings get background noise and echo added, because kitchens have both.
What is honestly hard here
Augmentation cannot invent information that was never collected. It stretches what you have.
Say your training photos are all one type of pen on white paper. No amount of rotating produces chalk on a blackboard. Augmentation is a patch on a collection problem, not a cure for it.
And picking transforms takes real thought about your data. Copying a list of augmentations from a blog post is how people quietly make their models worse.
Remember this
- Augmentation makes new examples by changing old ones in realistic ways.
- The change must not alter the correct answer. A flipped digit is a different digit.
- It stretches the data you have. It cannot replace data you never collected.
What to learn next
- Synthetic data — generating rows from scratch, and what that cannot give you.
- Overfitting and underfitting — the problem augmentation is really treating.
- Image classification — where these transforms do the most work.
Developer — Code and libraries.
Here is the rule measured. We use the handwritten digits dataset that ships inside scikit-learn — 1,797 images of 8 by 8 pixels, no download, runs on any laptop.
The scenario is realistic. Training images were all captured neatly centred. Real users do not centre anything, so the test set is shifted by up to one pixel in each direction.
We then compare a transform that matches reality (shifting) against one that does not (mirroring).
Setup
pip install numpy scikit-learnMatched transform versus mismatched transform
import numpy as np
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
digits = load_digits() # 1797 hand-written digits, 8x8 pixels, ships inside scikit-learn
X_train, X_test, y_train, y_test = train_test_split(
digits.images, digits.target, test_size=0.4, random_state=0, stratify=digits.target)
def shift(images, down, right):
"""Slide every image by one pixel and fill the vacated edge with background."""
out = np.zeros_like(images)
out[:, max(down, 0):8 + min(down, 0), max(right, 0):8 + min(right, 0)] = \
images[:, max(-down, 0):8 + min(-down, 0), max(-right, 0):8 + min(-right, 0)]
return out
# Real users do not centre the digit perfectly. Training never saw that; the test set now does.
rng = np.random.default_rng(0)
offsets = rng.integers(-1, 2, (len(X_test), 2))
X_test = np.concatenate([shift(X_test[i:i + 1], int(offsets[i, 0]), int(offsets[i, 1]))
for i in range(len(X_test))])
FOUR_WAYS = [(1, 0), (-1, 0), (0, 1), (0, -1)]
shifted = np.concatenate([X_train] + [shift(X_train, d, r) for d, r in FOUR_WAYS])
mirrored = np.concatenate([X_train, X_train[:, :, ::-1]])
def accuracy(images, labels):
model = LogisticRegression(max_iter=5000).fit(images.reshape(len(images), -1), labels)
return round(accuracy_score(y_test, model.predict(X_test.reshape(len(X_test), -1))), 3)
print("labelled training images:", len(X_train))
print("test images, off-centre :", len(X_test))
print()
print(f"{'training set':40s}{'images':>8}{'accuracy':>10}")
print(f"{'the originals, all neatly centred':40s}{len(X_train):>8}{accuracy(X_train, y_train):>10}")
print(f"{'originals + one-pixel shifted copies':40s}{len(shifted):>8}"
f"{accuracy(shifted, np.concatenate([y_train] * 5)):>10}")
print(f"{'originals + left-right mirrored copies':40s}{len(mirrored):>8}"
f"{accuracy(mirrored, np.concatenate([y_train] * 2)):>10}")labelled training images: 1078 test images, off-centre : 719 training set images accuracy the originals, all neatly centred 1078 0.356 originals + one-pixel shifted copies 5390 0.583 originals + left-right mirrored copies 2156 0.299
Three numbers, three lessons
0.356 is the collection failure. A model trained on perfectly centred digits scores barely above a third when digits arrive off-centre. Nothing is wrong with the model. It was never shown the variation it now faces.
0.583 is augmentation earning its place. Adding shifted copies raised accuracy by 22.7 points, at zero labelling cost. The transform matched the variation in the real world, so the model learned to ignore position.
0.299 is augmentation doing harm. Mirrored copies made the model worse than doing nothing. A mirrored 2 is not a 2, and a mirrored 5 looks something like a 2. We taught the model that two different digits share a label.
And 0.583 is still bad. This is the honest part. Augmentation recovered part of the loss and nowhere near all of it. Off-centre training images, collected properly, would beat this comfortably. Augmentation is a repair, not a substitute for collecting the right data.
Why mirroring is worse than useless
Adding a transform that changes the label injects contradictory training signal. The model sees two identical inputs with different labels and settles on a blurred compromise for both.
That is why the result lands below the no-augmentation baseline, rather than only failing to help. A wrong augmentation is not a wasted augmentation. It is an active harm.
Mirroring is fine for photographs of cats. It is wrong for digits, for text, for medical scans where left and right differ, and for road scenes in a country that drives on one particular side.
A quick tour of transforms that usually match reality
Images. Small rotations, small shifts, small crops, brightness and contrast changes, mild blur, sensor noise. Horizontal flip only when your subject is symmetric in the wild.
Text. Synonym replacement, random word deletion, back-translation through another language. Be careful: negation words and named entities must never be swapped, and word order carries meaning in most languages.
Audio. Time stretching, pitch shifting within a small range, added room noise, added background chatter. SpecAugment masks bands of time and frequency directly on the spectrogram.
Tabular data. The weakest case. Adding Gaussian noise to columns rarely helps and frequently breaks relationships between columns. Prefer synthetic data with explicit constraints, and expect modest results.
Common mistakes
Augmenting the test set. The test set represents reality and must be left untouched. In the code above the shift applied to X_test is the world being messy, not an augmentation, and it is applied before any model is fitted.
Augmenting before the split. Do the split first. Otherwise a shifted copy of a training image lands in the test set, and you are scoring the model on rows it has already seen. This is the leakage described in cleaning data.
Copying a transform list from a different domain. Standard ImageNet augmentation includes horizontal flip. Apply it to X-rays and you will teach the model that a heart on the right side is normal.
Augmenting so hard the label becomes doubtful. If a human cannot classify the augmented example, it is noise with a confident label attached. Look at 20 augmented samples before training. Every time.
Using augmentation to fix class imbalance and stopping there. Ten copies of one rare fraud case is still one fraud case worth of information. See imbalanced data.
Try it yourself
Add a third augmentation: vertical flip, X_train[:, ::-1, :]. Predict the result before running it. It should land near the mirrored case, because up-down flipping also destroys digit identity.
Then reduce the test-set offsets from rng.integers(-1, 2, ...) to all zeros, so the test digits stay centred. Re-run. The shifted augmentation now hurts, because you added variation the test set does not contain. That single change flips the conclusion, and it is the clearest possible demonstration of the rule.
What to learn next
- Synthetic data — generating rows from scratch, and what that cannot give you.
- Overfitting and underfitting — the problem augmentation is really treating.
- Image classification — where these transforms do the most work.
Researcher — Mathematics and papers.
Augmentation as an invariance prior
Let $\mathcal{T}$ be a set of transformations $t: \mathcal{X} \to \mathcal{X}$. Augmentation asserts a label-invariance assumption:
$$ p(y \mid x) \;=\; p\big(y \mid t(x)\big) \qquad \forall\, t \in \mathcal{T} $$
Training on augmented data approximates minimisation of the transformation-averaged risk:
$$ \hat{R}{\mathcal{T}}(f) \;=\; \frac{1}{n}\sum{i=1}^{n} \mathbb{E}_{t \sim \mathcal{T}}\Big[ \ell\big(f(t(x_i)),\, y_i\big) \Big] $$
When the invariance assumption holds, this is a form of regularisation that shrinks the effective hypothesis class to functions approximately invariant under $\mathcal{T}$, reducing variance without adding bias.
When the assumption fails — mirroring a digit — the augmented risk is minimised by a function that cannot be correct on the true distribution. The bias term becomes irreducible, and no amount of data or capacity removes it. The experiment above measures exactly this term.
The variance-reduction view
Chen, Dobriban and Lee (2020), A Group-Theoretic Framework for Data Augmentation, JMLR, formalise augmentation over a group $G$ acting on the input space. Their main result: for the orbit-averaged estimator, augmentation reduces variance and the reduction is expressible through the group's action, while leaving the bias unchanged provided the target function is genuinely $G$-invariant.
Dao et al. (2019), A Kernel Theory of Modern Data Augmentation, ICML, decompose the effect into a kernel term and a variance-regularisation term, giving a first-order approximation of augmentation as an explicit regulariser. Both papers converge on the same practical statement: augmentation encodes a prior, and a wrong prior is a bias you cannot train away.
Learned augmentation policies
- AutoAugment (Cubuk et al., 2019, CVPR) searches a discrete policy space of 16 operations with magnitude and probability, using reinforcement learning against validation accuracy. Effective and extremely expensive — thousands of GPU-hours for the original search.
- RandAugment (Cubuk et al., 2020) removes the search entirely, reducing the space to two integers: $N$ operations sampled uniformly, each at a single shared magnitude $M$. It matches or exceeds AutoAugment at a search cost of a small grid over two parameters. This is the sensible default.
- TrivialAugment (Müller and Hutter, 2021, ICCV) removes even that, sampling one operation and one magnitude uniformly per image, with no tuning. Competitive with both, which is a pointed result about how much of the earlier gain came from the search.
Mixing-based augmentation
mixup (Zhang et al., 2018, ICLR) trains on convex combinations of pairs:
$$ \tilde{x} = \lambda x_i + (1-\lambda) x_j, \qquad \tilde{y} = \lambda y_i + (1-\lambda) y_j, \qquad \lambda \sim \text{Beta}(\alpha, \alpha) $$
Note this breaks the label-invariance framing above: the label changes deliberately. It is better read as enforcing linear behaviour between training examples, and it measurably improves calibration and robustness to label noise.
CutMix (Yun et al., 2019, ICCV) pastes a rectangular patch from one image into another and mixes labels in proportion to patch area. Cutout (DeVries and Taylor, 2017) masks a random square, forcing reliance on distributed evidence rather than one discriminative region.
For tabular data, none of these transfer cleanly. mixup between two customers produces a customer who does not exist, and the interpolated categorical columns are meaningless.
Domain-specific notes
Audio. SpecAugment (Park et al., 2019, Interspeech) applies time warping, frequency masking and time masking directly to the log-mel spectrogram. It produced substantial word-error-rate improvements on LibriSpeech without any change to the model, and it is now standard. See waveforms and spectrograms.
Text. EDA (Wei and Zou, 2019, EMNLP) — synonym replacement, random insertion, swap and deletion — helps most in low-resource settings and its benefit largely vanishes above a few thousand labelled examples. Back-translation (Sennrich et al., 2016, ACL) is stronger and costs a translation pass. Both risk altering meaning: negation, quantities and named entities are the usual casualties.
Self-supervised learning. In SimCLR (Chen et al., 2020, ICML) augmentation is not a regulariser but the entire learning signal — the model is trained to identify which two views came from the same image. Their ablation shows composition matters enormously: random crop plus colour distortion is essential, and either alone is much weaker. When augmentation defines the task, its design decides what the representation becomes invariant to, for better or worse.
Evaluation discipline
- Augment the training split only, after splitting.
- Report the un-augmented baseline in every table. It is frequently the winner and frequently omitted.
- Ablate each transform individually. Composed policies hide the one that is hurting.
- Check test-time augmentation separately: averaging predictions over transformed copies of a test input is a different technique with a different cost profile.
- Where invariance is claimed, measure it directly — feed transformed inputs and report prediction agreement, not only accuracy.
Reading
- Shorten and Khoshgoftaar, A survey on Image Data Augmentation for Deep Learning, Journal of Big Data 2019.
- Cubuk et al., RandAugment: Practical automated data augmentation with a reduced search space, 2020 — arxiv.org/abs/1909.13719
- Zhang et al., mixup: Beyond Empirical Risk Minimization, ICLR 2018 — arxiv.org/abs/1710.09412
- Chen, Dobriban and Lee, A Group-Theoretic Framework for Data Augmentation, JMLR 2020 — arxiv.org/abs/1907.10905
- Park et al., SpecAugment, Interspeech 2019 — arxiv.org/abs/1904.08779
What to learn next
- Synthetic data — generating rows from scratch, and what that cannot give you.
- Overfitting and underfitting — the problem augmentation is really treating.
- Image classification — where these transforms do the most work.