ResNet
ResNet stacks residual blocks into four stages that halve the picture and double the channels, and it is still the default backbone a decade later.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
ResNet stacks residual blocks into four stages. Each stage halves the picture and doubles the number of channels.
Think about how a building is put up. Four floors, each made of the same repeated room. Each floor is a little narrower and a little more specialised than the one below.
You do not design every room separately. You design one room, and you repeat it.
ResNet is that idea for networks. One block, repeated. Change how many times you repeat it and you get a different sized model, from small to enormous.
Why it took over
Before ResNet, a network with more than about twenty layers trained badly. The skip connection fixed that, and the team behind it went straight to 152 layers.
It won the big 2015 competition and cut the error rate roughly in half compared to the previous year. That alone would have made it famous.
What made it last was something duller. It is easy to explain, easy to resize, and it works. Ten years on, a paper that needs a vision backbone without an argument still uses a ResNet.
The four-stage shape
photo 224 wide
|
[ stem: one big filter, then pooling ] -> 56 wide, 64 channels
|
[ stage 1: repeated blocks ] -> 56 wide, 64 channels
|
[ stage 2: repeated blocks ] -> 28 wide, 128 channels
|
[ stage 3: repeated blocks ] -> 14 wide, 256 channels
|
[ stage 4: repeated blocks ] -> 7 wide, 512 channels
|
[ average everything, then one classifier ] -> 1000 scoresRead the two columns on the right. The picture keeps halving. The channel count keeps doubling.
That is the trade every backbone makes. You give up knowing exactly where something is, in exchange for knowing more about what it is.
The name is the layer count
ResNet-18 has eighteen weighted layers. ResNet-50 has fifty. ResNet-152 has one hundred and fifty-two.
The deeper ones cost less than you would fear. They switch to a thriftier block, which squeezes the channels down before the expensive work and expands them after. This is measured further down: the thrifty block does the same job with about seventeen times fewer numbers.
Where you have already seen this
- Photo apps that group pictures of the same person.
- Shop apps that let you search by uploading a photo of a product.
- Quality inspection cameras on a factory line.
- Almost every research paper that needed "some vision model" between 2016 and today.
The honest part
ResNet is not the most accurate architecture available. Newer designs beat it on every benchmark.
It stays popular because it is predictable, fast, and widely supported, and because the weights people share are trustworthy. In real projects that matters more than a couple of points on a leaderboard.
Remember this
- ResNet repeats one block in four stages, halving the picture and doubling channels each time.
- The number in the name is the layer count.
- Deeper versions use a thrifty block that squeezes channels before the expensive filter.
What to learn next
- The MobileNet family — what changes when the target is a phone.
- Transfer learning in PyTorch — putting a pretrained ResNet to work.
- ConvNeXt — what a ResNet looks like after ten years of lessons.
Developer — Code and libraries.
Setup
pip install "torch==2.5.1" "torchvision==0.20.1"Run against PyTorch 2.5.1 and torchvision 0.20.1 on CPU. No weights are downloaded, so it runs offline.
Take the thing apart
import torch
from torchvision.models import resnet18, resnet34, resnet50, resnet101, resnet152
from torchvision.models.resnet import BasicBlock, Bottleneck
p = lambda m: sum(q.numel() for q in m.parameters())
print("the two block types, both taking 256 channels in and out")
basic = BasicBlock(256, 256) # two 3x3 convolutions
neck = Bottleneck(256, 64) # 1x1 down, 3x3, 1x1 up
print(f" BasicBlock : {p(basic):>9,} parameters")
print(f" Bottleneck : {p(neck):>9,} parameters")
x = torch.zeros(1, 256, 14, 14)
print(f" shapes : {tuple(basic(x).shape)} and {tuple(neck(x).shape)}")
print("\ndepth costs less than you think, because the deep ones use bottlenecks")
for name, fn in [("resnet18", resnet18), ("resnet34", resnet34), ("resnet50", resnet50),
("resnet101", resnet101), ("resnet152", resnet152)]:
m = fn(weights=None)
n_blocks = sum(1 for mod in m.modules() if isinstance(mod, (BasicBlock, Bottleneck)))
kind = "Bottleneck" if isinstance(m.layer1[0], Bottleneck) else "BasicBlock"
print(f" {name:10s} {n_blocks:3d} {kind:10s} blocks {p(m):>12,} parameters")
print("\nwhere the feature map is, stage by stage (input 224x224)")
net = resnet18(weights=None).eval()
seen = {}
for stage in ["layer1", "layer2", "layer3", "layer4"]:
getattr(net, stage).register_forward_hook(
lambda mod, inp, out, s=stage: seen.__setitem__(s, tuple(out.shape)))
with torch.no_grad():
logits = net(torch.zeros(1, 3, 224, 224))
for stage, shape in seen.items():
print(f" {stage}: {shape}")
print(f" logits: {tuple(logits.shape)}")
print("\nthe shortcut needs help when the shape changes")
print(" resnet18.layer1[0].downsample :", net.layer1[0].downsample)
print(" resnet18.layer2[0].downsample :", net.layer2[0].downsample)
print("\nsame architecture, two sets of weights, four points of accuracy apart")
from torchvision.models import ResNet50_Weights
for w in (ResNet50_Weights.IMAGENET1K_V1, ResNet50_Weights.IMAGENET1K_V2):
acc = w.meta["_metrics"]["ImageNet-1K"]
print(f" {str(w).split('.')[-1]:15s} top-1 {acc['acc@1']:.2f} top-5 {acc['acc@5']:.2f}")the two block types, both taking 256 channels in and out BasicBlock : 1,180,672 parameters Bottleneck : 70,400 parameters shapes : (1, 256, 14, 14) and (1, 256, 14, 14) depth costs less than you think, because the deep ones use bottlenecks resnet18 8 BasicBlock blocks 11,689,512 parameters resnet34 16 BasicBlock blocks 21,797,672 parameters resnet50 16 Bottleneck blocks 25,557,032 parameters resnet101 33 Bottleneck blocks 44,549,160 parameters resnet152 50 Bottleneck blocks 60,192,808 parameters where the feature map is, stage by stage (input 224x224) layer1: (1, 64, 56, 56) layer2: (1, 128, 28, 28) layer3: (1, 256, 14, 14) layer4: (1, 512, 7, 7) logits: (1, 1000) the shortcut needs help when the shape changes resnet18.layer1[0].downsample : None resnet18.layer2[0].downsample : Sequential( (0): Conv2d(64, 128, kernel_size=(1, 1), stride=(2, 2), bias=False) (1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) ) same architecture, two sets of weights, four points of accuracy apart IMAGENET1K_V1 top-1 76.13 top-5 92.86 IMAGENET1K_V2 top-1 80.86 top-5 95.43
Reading the output
The bottleneck block uses 70,400 parameters where the basic block uses 1,180,672. Same input, same output, roughly seventeen times cheaper. It gets there by dropping 256 channels to 64 with a 1x1, doing the 3x3 work at 64, and expanding back to 256 with another 1x1. The 3x3 convolution, the expensive part, now sees a quarter of the channels in and a quarter out, so it costs a sixteenth.
ResNet-34 and ResNet-50 have the same number of blocks. Sixteen each. ResNet-50 is deeper only because each bottleneck holds three weighted layers instead of two. The parameter counts are close, 21.8 million against 25.6 million, but ResNet-50 is significantly more accurate. This is the whole argument for the bottleneck.
The stage shapes are the pattern from the beginner block, confirmed. 56, 28, 14, 7 while channels run 64, 128, 256, 512. Every backbone in this section, including the transformers, reproduces these four numbers.
layer1[0].downsample is None and layer2[0].downsample is not. Inside stage 1 nothing changes shape, so the shortcut is a plain identity. Stage 2 halves the map and doubles the channels, so the shortcut needs a 1x1 convolution with stride 2 to match. There are exactly three of these in a ResNet, one per stage transition.
The two weight sets are the most useful line here. Identical architecture, 76.13% against 80.86% top-1. The gap is training recipe alone: longer schedules, better augmentation, label smoothing. Read the architecture score in a 2016 paper and you are reading its recipe as much as its design.
Common mistakes
Fine-tuning without replacing fc. A pretrained ResNet outputs 1000 ImageNet scores. For a 5-class problem you need model.fc = nn.Linear(model.fc.in_features, 5). The rest is covered in transfer learning in PyTorch.
Using your own normalisation values. Pretrained weights expect the exact preprocessing they were trained with. Call weights.transforms() and use what it returns. Guessing costs several points of accuracy and produces no error.
Leaving batch norm in training mode while fine-tuning on a small set. With a batch of 8, batch-norm statistics are noisy and the running averages drift away from the pretrained ones. Call model.eval() for frozen sections, and see train and eval mode.
Assuming resnet50(pretrained=True) still works. That argument was removed. The current form is resnet50(weights=ResNet50_Weights.DEFAULT), and DEFAULT currently resolves to the better V2 weights, so you can silently change accuracy by upgrading torchvision. Pin the exact enum member in production code.
Feeding images much smaller than 224. The stem strides by 2 then pools by 2, so a 32x32 image is 8x8 before stage 1 and 1x1 by stage 4. For small images use a CIFAR-style stem with a 3x3 stride-1 convolution and no max pool.
Try it yourself
Print p(net.layer4) and compare it with p(net.layer1). Stage 4 holds more than half the parameters of the whole network, because parameters scale with the square of the channel count while the feature map is at its smallest. Then work out why freezing stages 1 and 2 during fine-tuning saves so little memory.
What to learn next
- The MobileNet family — what changes when the target is a phone.
- Transfer learning in PyTorch — putting a pretrained ResNet to work.
- ConvNeXt — what a ResNet looks like after ten years of lessons.
Researcher — Mathematics and papers.
The architecture
He, Zhang, Ren and Sun (2015), Deep Residual Learning for Image Recognition, arxiv.org/abs/1512.03385. A 7x7 stride-2 stem, a 3x3 stride-2 max pool, then four stages at strides 4, 8, 16 and 32 relative to the input, then global average pooling and a single linear classifier.
Block counts per stage define the variant: [2,2,2,2] for ResNet-18, [3,4,6,3] for ResNet-34 and ResNet-50, [3,4,23,3] for ResNet-101, [3,8,36,3] for ResNet-152. Depth is concentrated in stage 3, where the feature map is 14x14 and the compute per block is moderate.
The bottleneck block, used from ResNet-50 upward, is $1\times1 \to 3\times3 \to 1\times1$ with an expansion factor of 4: the block's output has four times the channels of its internal width. For internal width $w$ and input $4w$:
$$ \text{params} \approx \underbrace{4w \cdot w}{1\times1} + \underbrace{9 w^2}{3\times3} + \underbrace{w \cdot 4w}_{1\times1} = 17 w^2 $$
against $2 \cdot 9 (4w)^2 = 288 w^2$ for a basic block of the same external width. The measured ratio in the developer block, $1{,}180{,}672 / 70{,}400 \approx 16.8$, matches this once batch-norm parameters are included.
Reported results: 3.57% top-5 error for an ensemble on ImageNet, first place in ILSVRC 2015 classification, and a 28% relative improvement on COCO detection.
Variants worth knowing
| Variant | Change | Reference |
|---|---|---|
| Pre-activation ResNet | BN and ReLU moved before the convolutions, shortcut kept clean | He et al., 2016, 1603.05027 |
| Wide ResNet | Fewer layers, more channels; a 16-layer wide net matches a 1000-layer thin one | Zagoruyko and Komodakis, 2016, 1605.07146 |
| ResNeXt | Grouped 3x3 convolutions add a "cardinality" axis, better than width at fixed FLOPs | Xie et al., 2017, 1611.05431 |
| SE-ResNet | Squeeze-and-excitation channel gating added to each block | Hu et al., 2018, 1709.01507 |
| ResNet-D | Stride moved to the 3x3; average pooling in the projection shortcut; 3x3 stem | He et al., 2019, Bag of Tricks |
The ResNet-D changes are worth singling out. In the original stage transition, the shortcut is a 1x1 convolution with stride 2, which reads one pixel in four and discards the rest before the addition. Replacing it with average pooling followed by a stride-1 1x1 recovers roughly a point of top-1 at negligible cost, and is the default in most timm ResNet variants.
Architecture versus recipe
Wightman, Touvron and Jégou (2021), ResNet strikes back: An improved training procedure in timm, arxiv.org/abs/2110.00476, retrain an unmodified ResNet-50 with LAMB, binary cross-entropy, Mixup, CutMix, RandAugment and a long schedule, reaching 80.4% top-1 at 224x224 with no extra data and no distillation, against the original 76.1%.
This result reframes a decade of comparisons. Any claim that architecture X beats ResNet-50 by four points is indistinguishable from a claim about training recipes unless both were trained the same way. torchvision's own V1 and V2 weights, 76.13% and 80.86%, make the same point inside a single library.
Bello et al. (2021), Revisiting ResNets, reach the same conclusion from the scaling side: with modern regularisation and scaling rules, ResNets match EfficientNets at equal training cost on TPUs.
Transfer behaviour
Kornblith, Shlens and Le (2019), Do Better ImageNet Models Transfer Better?, arxiv.org/abs/1805.08974, find ImageNet top-1 correlates strongly with fixed-feature and fine-tuned transfer accuracy across 12 datasets, but that the correlation weakens sharply for fine-grained tasks and for small target datasets. Regularisation that helps ImageNet, such as label smoothing, can hurt the transferability of penultimate features.
He, Girshick and Dollár (2018), Rethinking ImageNet Pre-training, arxiv.org/abs/1811.08883, show that COCO detection models trained from random initialisation match pre-trained ones given enough schedule, reaching 50.9 AP without external data. Pre-training buys convergence speed rather than a final accuracy ceiling, provided you have enough target data. With a small target set, the speed-up is the whole point.
Why it persists
Three properties, none of them accuracy. Its stage strides of 4, 8, 16 and 32 are what every detection and segmentation neck expects, so FPN, DeepLab and Mask R-CNN plug in without modification. Its 3x3 convolutions are the best-optimised primitive on every inference stack, from cuDNN to CoreML to embedded NPUs. And a decade of published numbers use it as the reference point, which makes it the cheapest baseline to defend in review.
What to learn next
- The MobileNet family — what changes when the target is a phone.
- Transfer learning in PyTorch — putting a pretrained ResNet to work.
- ConvNeXt — what a ResNet looks like after ten years of lessons.