Image transforms with torchvision v2
Transforms are the preparation steps between a stored image and a model-ready tensor — resizing, scaling, normalising, plus deliberate random edits that multiply your training data for free.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Transforms are the steps that turn a stored photo into the exact numbers a model expects — and, during training, deliberately vary each photo so the model stops memorising.
Watch a vegetable vendor's child learn tomatoes. They see tomatoes at dawn and noon, half-shadowed, upside down in a pile, wet from washing. After a season, no lighting or angle can fool them.
Now imagine a child shown one studio photograph of one tomato, a thousand times. They have learned that photograph, not tomatoes.
Training transforms give your model the market child's education. Every time a photo is fetched, it is randomly cropped a little differently, sometimes flipped, before the model sees it. Same stored files, endlessly varied views.
Why it exists
Two separate jobs hide in "prepare an image", and transforms do both.
The boring job: models are fussy eaters. A stored photo is whole numbers 0 to 255 at some arbitrary size; the model wants decimals in a narrow range at one fixed size. Every image needs the same standardising steps — resize, convert, normalise (shift the numbers into the range the model was built around).
The clever job: augmentation — random edits that create variety. A flipped mango is still a mango; a slightly re-cropped one too. Each random edit is a free new training example, and the variety fights overfitting directly.
How it works
training: evaluation:
photo photo
→ random crop (different every time) → plain resize (same every time)
→ maybe flip → convert + normalise
→ convert + normalise |
| v
v steady, repeatable input
fresh variation each epochTwo recipes, on purpose. Practice varies; the exam is standard — the same split as train and eval mode, one level down the pipeline.
A real example you have seen
Photo apps that recognise your face from odd angles, half-lit, sunglasses on — that robustness was trained with exactly these random edits. Nobody photographed you ten thousand times; augmentation did.
Remember this
- Transforms standardise images and create random variety.
- Random edits for training; the same plain steps, every time, for evaluation.
- The variety is free extra data — the cheapest overfitting cure there is.
What to learn next
- Datasets that do not fit in memory — feeding these transforms when the images will not all fit in RAM.
- Data augmentation — the augmentation menu beyond images.
- Transfer learning — where the normalisation statistics contract comes from.
Developer — Code and libraries.
Setup
pip install torch torchvisionWritten and tested against torch 2.5 with torchvision 0.20, whose transforms.v2 API is the stable, recommended one. (The older torchvision.transforms still works; v2 is faster, tensor-native, and what new code should use.)
Two pipelines, and proof they behave differently
No downloads — a fake image stands in for a decoded JPEG:
import torch
from torchvision.transforms import v2
torch.manual_seed(0)
# One fake 3-channel image, values 0-255, as uint8 - exactly what a decoded JPEG gives you.
image = torch.randint(0, 256, (3, 96, 96), dtype=torch.uint8)
train_tf = v2.Compose([
v2.RandomResizedCrop(size=(64, 64), antialias=True),
v2.RandomHorizontalFlip(p=0.5),
v2.ToDtype(torch.float32, scale=True), # uint8 0-255 -> float 0.0-1.0
v2.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])
eval_tf = v2.Compose([
v2.Resize(size=(64, 64), antialias=True),
v2.ToDtype(torch.float32, scale=True),
v2.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])
out = train_tf(image)
print("in: ", image.dtype, tuple(image.shape), "range", image.min().item(), "-", image.max().item())
print("out:", out.dtype, tuple(out.shape), f"mean {out.mean().item():.3f}")
a = train_tf(image)
b = train_tf(image)
print("same input, same output?", torch.equal(a, b), "- augmentation is random on purpose")
print("eval twice identical? ", torch.equal(eval_tf(image), eval_tf(image)))in: torch.uint8 (3, 96, 96) range 0 - 255 out: torch.float32 (3, 64, 64) mean 0.238 same input, same output? False - augmentation is random on purpose eval twice identical? True
The exact mean depends on seed and version; the two False/True lines are the point, and they are the point on every machine.
The walkthrough
Compose chains transforms into one callable — the thing you pass as transform= into a dataset, to run inside __getitem__ per sample. That placement matters: per-sample work is what workers parallelise.
RandomResizedCrop picks a random region and scale, then resizes to the target — the standard training crop, varying framing every single call. RandomHorizontalFlip flips a coin per call. Both draw fresh randomness each time, which is the False in the output.
ToDtype(torch.float32, scale=True) converts whole numbers 0-255 into decimals 0.0-1.0. This is the v2 spelling; you will see ToTensor() in older tutorials — deprecated in v2, replace it with exactly this line (plus v2.ToImage() first if starting from a PIL image rather than a tensor).
Normalize(mean=..., std=...) shifts each colour channel to sit around zero. Those six magic numbers are the ImageNet statistics — mandatory when using pretrained backbones, because the pretrained weights assume them. Training from scratch on your own data? Compute your own dataset's statistics instead.
Order is fixed by meaning: geometry on cheap uint8 first, then dtype conversion, then normalisation last. Normalize before ToDtype crashes on integer input; flips after Normalize waste float bandwidth.
Choose augmentations by what cannot change the label
A horizontal flip keeps a mango a mango — but flips "b" into "d". A vertical flip suits satellite tiles, not street photos. Heavy colour shifts are fine for shapes, fatal for ripeness detection where colour is the signal. Every augmentation encodes an assumption: this edit does not change the answer. Choose from your domain's truths, not from a copied recipe. The broader menu — noise, rotation, cutout, and beyond images — lives in data augmentation.
Common mistakes
Augmenting validation. Random edits on the exam set make metrics wobble run to run, and slightly worse on average. Eval gets the plain pipeline, always.
Wrong normalisation with a pretrained model. Off-by-a-lot inputs, quietly mediocre accuracy, no error message. Copy the exact statistics the backbone shipped with — or better, use its bundled recipe: models.ResNet18_Weights.DEFAULT.transforms().
ToTensor and v2 mixing. Old and new APIs interleave badly. Pick v2, use ToImage + ToDtype, and be consistent file-wide.
Forgetting antialias=True. On tensor inputs it matters for quality and for matching PIL-based results; recent torchvision defaults it on for v2, but writing it explicitly keeps behaviour pinned across versions.
Try it yourself
Print out.mean() across 200 calls of train_tf(image) and watch the spread — that spread is your augmentation strength. Then drop RandomResizedCrop for Resize and watch it collapse.
What to learn next
- Datasets that do not fit in memory — feeding these transforms when the images will not all fit in RAM.
- Data augmentation — the augmentation menu beyond images.
- Transfer learning — where the normalisation statistics contract comes from.
Researcher — Mathematics and papers.
Augmentation as invariance encoding
Augmentation is a group-action prior: for transformations $g \in G$ presumed label-preserving, training on $(g(x), y)$ enforces $f(g(x)) \approx f(x)$ — invariance taught through data rather than architecture. Chen et al. (2020a), A Group-Theoretic Framework for Data Augmentation, formalises the variance-reduction view: averaging over orbits of $G$ provably reduces estimator variance when the invariance holds. The failure mode is equally formal: when the label is not $G$-invariant (chirality, digits, colour-coded classes), the orbit crosses decision boundaries and augmentation injects label noise.
Learned policies — AutoAugment (Cubuk et al., 2019), RandAugment (Cubuk et al., 2020, two scalars replacing a search), TrivialAugment (Müller and Hutter, 2021, no tuning at all) — found that policy strength matters more than composition detail. Mixing-based methods leave the invariance frame: mixup (Zhang et al., 2018) trains on convex combinations of image and label; CutMix (Yun et al., 2019) patches regions with proportional label mixing — both are regularisers on the decision surface's geometry rather than encodings of a symmetry, with measurable calibration benefits.
Augmentation is also the engine of self-supervision: SimCLR (Chen et al., 2020b) learns representations by demanding agreement across two augmented views, making the augmentation menu itself the supervisory signal — the strongest evidence that augmentation choice is representational, not cosmetic.
The v2 machinery
v2 transforms dispatch on TVTensor types (Image, BoundingBoxes, Mask, Video — the tv_tensors module): a geometric transform applied to a (image, boxes, mask) structure transforms all three consistently, which v1 could not express — the reason detection and segmentation pipelines required hand-rolled joint transforms for a decade. Plain tensors pass through with image semantics; arbitrary nested structures are traversed. Everything is scriptable/compilable in the tensor path, and uint8-first pipelines (resize on uint8, convert late) are the documented fast path — geometry at one byte per element, arithmetic deferred.
Placement trade-offs: CPU-side per-sample transforms parallelise across workers but compete with decode for cores; batch-level GPU augmentation (apply v2 ops to batches on-device, or DALI/Kornia) trades PCIe bandwidth and GPU cycles for CPU relief — the right choice is workload-dependent and measurable with the tools in the starvation lesson. Randomness follows worker semantics: per-worker seeds, hence exact augmentations depend on num_workers — the reproducibility caveat from the collate lesson applies verbatim.
Reading
- Cubuk et al. (2019, 2020); Müller and Hutter (2021) — the policy line.
- Zhang et al. (2018), mixup; Yun et al. (2019), CutMix.
- Chen et al. (2020b), SimCLR — augmentation as supervision.
- torchvision docs, "Transforms v2: End-to-end object detection/segmentation example" — the TVTensor dispatch model, normative.
What to learn next
- Datasets that do not fit in memory — feeding these transforms when the images will not all fit in RAM.
- Data augmentation — the augmentation menu beyond images.
- Transfer learning — where the normalisation statistics contract comes from.