ConvNeXt
ConvNeXt takes a ResNet and swaps in one transformer design choice at a time, ending with a pure convolutional network that matches Swin at the same compute.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
ConvNeXt is an old ResNet rebuilt with modern parts, one part at a time. It ended up matching the newest transformer models.
Think of an old motorcycle that has been kept running for twenty years. New brakes, new tyres, fuel injection, a modern exhaust. The frame is the original one, and it now keeps pace with new bikes.
That is exactly what the ConvNeXt team did. They started with a plain ResNet-50 and changed one thing at a time. Each change borrowed an idea from the transformer models.
At the end they had a network built only from ordinary convolution layers. It matched the transformers point for point.
Why this mattered so much
By 2021 the story was settled, or seemed to be. Transformers had beaten convolutional networks on image benchmarks. The word going round was that convolution was finished.
The ConvNeXt team was suspicious of that story, for a good reason. The transformer results also came with newer training recipes: longer schedules, stronger data augmentation, better optimisers.
So which part won? The architecture, or the training?
They answered it by holding everything else fixed and moving one piece at a time. That answer is the real contribution, more than the network itself.
What they found
Take a plain ResNet-50 and retrain it with a modern recipe. With no design changes at all, it went from about 76 percent correct to about 79 percent.
Three points, from training alone. That is a large share of the gap people had been attributing to architecture.
Then the design changes were added one at a time. Bigger filters. Fewer activation steps. A different normalisation. A different way of shrinking the picture between stages. Each step gave a small gain, and together they reached about 82 percent.
plain ResNet-50, old recipe about 76
same network, modern recipe about 79
plus every design change about 82The changes that mattered most
A wider filter. ResNet uses a three-by-three filter. ConvNeXt uses seven-by-seven, so each layer sees more of the picture at once. That is closer to how attention behaves.
Far fewer activation steps. A ResNet block squashes negatives to zero several times. A ConvNeXt block does it once. Transformer blocks are similarly sparse, and copying that helped.
A wider middle. The block expands to four times as many channels in the middle, then comes back down. A transformer's feed-forward part is built the same way.
Where you have already seen this
- ConvNeXt backbones sit inside many current detection and segmentation systems.
- Several image generation models use a ConvNeXt-style block in their internals.
- If you download a modern "convolutional backbone" today, it is often this one.
The honest part
None of these ideas were new inventions. Large filters, wide middles and layer normalisation all existed already.
The contribution was the discipline of the experiment. Change one thing, measure, keep it or drop it. Then publish the whole ladder, including the step that made things worse. That is unusual and valuable. The paper is worth reading even if you never use the model.
Remember this
- ConvNeXt is a ResNet modernised step by step with transformer design choices.
- A large part of the transformer advantage turned out to be training recipe, not architecture.
- The finished network is pure convolution and matches Swin at the same compute.
What to learn next
- Swin transformer — the model ConvNeXt was measured against.
- Vision transformers — the design ConvNeXt borrowed from.
- Masked image modelling — the pretraining ConvNeXt V2 had to be redesigned for.
Developer — Code and libraries.
Setup
pip install "torch==2.5.1" "torchvision==0.20.1"Run against PyTorch 2.5.1 and torchvision 0.20.1 on CPU. No downloads.
Build the block and check it against the library
import torch
import torch.nn as nn
from torch.utils.flop_counter import FlopCounterMode
p = lambda m: sum(q.numel() for q in m.parameters())
class ConvNeXtBlock(nn.Module):
"""7x7 depthwise, then LayerNorm, then a 4x wide MLP over channels."""
def __init__(self, dim, layer_scale=1e-6):
super().__init__()
self.dw = nn.Conv2d(dim, dim, 7, padding=3, groups=dim)
self.norm = nn.LayerNorm(dim, eps=1e-6)
self.fc1 = nn.Linear(dim, 4 * dim)
self.act = nn.GELU()
self.fc2 = nn.Linear(4 * dim, dim)
self.gamma = nn.Parameter(layer_scale * torch.ones(dim))
def forward(self, x):
shortcut = x
x = self.dw(x)
x = x.permute(0, 2, 3, 1) # NCHW -> NHWC, so LayerNorm sees channels
x = self.fc2(self.act(self.fc1(self.norm(x))))
x = self.gamma * x
return shortcut + x.permute(0, 3, 1, 2) # back to NCHW
blk = ConvNeXtBlock(96)
x = torch.zeros(2, 96, 56, 56)
print("hand-written block :", tuple(blk(x).shape), f"{p(blk):,} parameters")
from torchvision.models.convnext import CNBlock
tv = CNBlock(96, layer_scale=1e-6, stochastic_depth_prob=0.0)
print("torchvision CNBlock:", tuple(tv(x).shape), f"{p(tv):,} parameters")
print("\nwhere the parameters sit in one block of width 96")
for name, mod in [("7x7 depthwise", blk.dw), ("LayerNorm", blk.norm),
("1x1 expand to 4x", blk.fc1), ("1x1 back down", blk.fc2)]:
print(f" {name:18s} {p(mod):>8,}")
print(f" {'layer scale':18s} {blk.gamma.numel():>8,}")
print("\nthe depthwise kernel is cheap, which is why 7x7 is affordable")
for k in (3, 5, 7, 9):
dw = nn.Conv2d(96, 96, k, padding=k // 2, groups=96)
dense = nn.Conv2d(96, 96, k, padding=k // 2)
print(f" {k}x{k}: depthwise {p(dw):>7,} ordinary {p(dense):>9,}")
def gmacs(model, res=224):
model.eval()
c = FlopCounterMode(display=False)
with c, torch.no_grad():
model(torch.zeros(1, 3, res, res))
return c.get_total_flops() / 2 / 1e9
from torchvision.models import convnext_tiny, convnext_small, convnext_base, swin_t, resnet50
print("\nthe comparison the paper was built to win (measured here)")
print(f"{'model':16s} {'params':>12s} {'GMACs':>7s}")
for name, fn in [("resnet50", resnet50), ("swin_t", swin_t), ("convnext_tiny", convnext_tiny),
("convnext_small", convnext_small), ("convnext_base", convnext_base)]:
m = fn(weights=None)
print(f"{name:16s} {p(m):>12,} {gmacs(m):>7.2f}")hand-written block : (2, 96, 56, 56) 79,296 parameters torchvision CNBlock: (2, 96, 56, 56) 79,296 parameters where the parameters sit in one block of width 96 7x7 depthwise 4,800 LayerNorm 192 1x1 expand to 4x 37,248 1x1 back down 36,960 layer scale 96 the depthwise kernel is cheap, which is why 7x7 is affordable 3x3: depthwise 960 ordinary 83,040 5x5: depthwise 2,496 ordinary 230,496 7x7: depthwise 4,800 ordinary 451,680 9x9: depthwise 7,872 ordinary 746,592 the comparison the paper was built to win (measured here) model params GMACs resnet50 25,557,032 4.09 swin_t 28,288,354 4.49 convnext_tiny 28,589,128 4.46 convnext_small 50,223,688 8.68 convnext_base 88,591,464 15.35
Reading the output
The hand-written block matches torchvision to the parameter. 79,296 both ways. If you can rebuild a block from scratch and hit the library's count exactly, you have understood it. If your count differs, the difference is always a bias term, a norm layer or a missing scale parameter.
The 7x7 depthwise convolution is 6% of the block. 4,800 parameters out of 79,296. Almost everything sits in the two 1x1 layers that expand to 384 channels and come back. Kernel size feels like the headline change and is close to free; the channel mixing is where the model lives.
A 7x7 depthwise kernel costs less than a 3x3 dense one. 4,800 against 83,040. Depthwise convolution is what makes large kernels affordable. Without it, ConvNeXt's central design choice would not fit in the budget.
swin_t and convnext_tiny are 4.49 and 4.46 GMACs. That match is not luck. The paper's whole method is to compare at equal compute, so the ConvNeXt widths were chosen to land on Swin's budget. When you read "ConvNeXt beats Swin", this is the sense in which it does.
The two permute calls are the awkward part. LayerNorm normalises the last dimension, and PyTorch images are channel-first. So the tensor is flipped to channels-last and back. torchvision uses a Permute module and a LayerNorm2d helper for the same reason. On a GPU, running the whole model in channels_last memory format removes the cost of these flips.
Common mistakes
Reaching for nn.BatchNorm2d inside the block. ConvNeXt uses LayerNorm, which is one of the measured steps in the paper. Substituting batch norm changes the small-batch behaviour and undoes part of the design.
Applying LayerNorm on a channel-first tensor. nn.LayerNorm(96) normalises the last dimension. On an (N, C, H, W) tensor it will normalise over width, run without error, and train badly. Permute first, or use torchvision.models.convnext.LayerNorm2d.
Dropping layer scale. The gamma parameter initialised at 1e-6 makes each block start as an almost exact identity. It matters for deep and wide variants; removing it destabilises training at the larger sizes.
Expecting a speed win from the FLOP match. Equal GMACs does not mean equal wall clock. On one desktop CPU at four threads, convnext_tiny measured about 80 ms per image and swin_t about 160 ms, at nearly identical arithmetic. Your machine will differ; measure it.
Fine-tuning with the ImageNet learning rate. ConvNeXt was trained with AdamW, a long warm-up, and layer-wise learning-rate decay. Fine-tuning it like a ResNet with SGD and a constant rate usually underperforms.
Try it yourself
Change 4 * dim to 2 * dim in the block and recount. The expansion ratio, not the kernel size, is the parameter dial. Then swap the 7x7 for a 3x3 and note how little the count moves, which is the point the third block of output is making.
What to learn next
- Swin transformer — the model ConvNeXt was measured against.
- Vision transformers — the design ConvNeXt borrowed from.
- Masked image modelling — the pretraining ConvNeXt V2 had to be redesigned for.
Researcher — Mathematics and papers.
The experiment
Liu, Mao, Wu, Feichtenhofer, Darrell and Xie (2022), A ConvNet for the 2020s, CVPR, arxiv.org/abs/2201.03545. The method is an ablation ladder from ResNet-50 to a Swin-T-sized ConvNet, holding FLOPs approximately constant and reporting ImageNet-1K top-1 after every step:
| Step | Top-1 |
|---|---|
| ResNet-50, original recipe | 76.1 |
| Modern training recipe (300 epochs, AdamW, Mixup, CutMix, RandAugment, stochastic depth) | 78.8 |
| Stage compute ratio changed from (3,4,6,3) to (3,3,9,3) | 79.4 |
| Patchify stem: 4x4 stride-4 convolution replaces the 7x7 stem and max pool | 79.5 |
| Depthwise convolution, width raised 64 to 96 | 80.5 |
| Inverted bottleneck | 80.6 |
| Depthwise layer moved up, before the expansion | 79.9 |
| Kernel enlarged to 7x7 | 80.6 |
| ReLU replaced by GELU | 80.6 |
| Fewer activations: one GELU per block | 81.3 |
| Fewer normalisation layers | 81.4 |
| BatchNorm replaced by LayerNorm | 81.5 |
| Separate downsampling layers between stages | 82.0 |
Two entries carry most of the paper's argument.
The first is 76.1 to 78.8. 2.7 points from the training recipe alone, with no architectural change. Comparisons between a 2016 architecture paper and a 2021 one are contaminated by this unless both were retrained. It is the same conclusion reached independently in ResNet strikes back.
The second is the 80.6 to 79.9 drop when the depthwise layer moves up. The paper publishes the regression rather than hiding it. That step is kept because it enables the large kernel that follows, which is the kind of dependency an ablation table can express and prose usually cannot.
Reported headline: 87.8% ImageNet top-1 for the largest variant, and outperforming Swin Transformers on COCO detection and ADE20K segmentation at comparable compute.
What the block actually is
$$ \text{Block}(x) = x + \gamma \odot \bigl( W_2\, \sigma(W_1\, \mathrm{LN}(\mathrm{DWConv}_{7\times7}(x))) \bigr) $$
$\mathrm{DWConv}_{7\times7}$ is a depthwise convolution, $\mathrm{LN}$ is layer normalisation over the channel dimension, $W_1$ expands by 4, $\sigma$ is GELU, $W_2$ projects back, and $\gamma$ is a per-channel LayerScale parameter initialised at $10^{-6}$.
Structurally this is a transformer block with the self-attention replaced by a depthwise convolution. Depthwise convolution mixes information across space with weights that do not depend on content; self-attention does the same with weights that do. Reading it this way, the paper is measuring how much of the transformer's advantage comes from content-dependent mixing rather than from everything else in the block. The answer, at ImageNet scale and this compute budget, is close to none.
Kernel size and the limit of the trend
The paper tested 3, 5, 7, 9 and 11 and found saturation at 7. This sits alongside Ding et al. (2022), RepLKNet, which pushes to 31x31 kernels and reports gains, but requires re-parameterisation with a parallel small kernel to train stably and targets dense prediction where large effective receptive fields help most. The two results are compatible: gains from kernel size depend on the task and on how the network is trained, not on kernel size alone.
ConvNeXt V2
Woo, Debnath, Hu, Chen, Liu, Kweon and Xie (2023), ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, arxiv.org/abs/2301.00808, attach masked autoencoder pretraining to ConvNeXt. Two obstacles had to be solved.
First, a convolution cannot skip masked patches the way a transformer encoder can, so they treat the masked input as a sparse tensor and use sparse convolutions during pretraining, calling the design a fully convolutional masked autoencoder.
Second, naively combining the two "leads to subpar performance". Diagnosing it, they found feature collapse: many channels became inactive. Global Response Normalization was added to increase inter-channel competition, and the combination works. Results: a 3.7M-parameter Atto model at 76.7% top-1, and a 650M Huge model at 88.9% using public data only.
The finding worth carrying away is that pretraining method and architecture are not independent choices. A recipe designed for transformers needed architectural surgery to work on a ConvNet.
Reading this lesson correctly
The correct conclusion from ConvNeXt is not "convolutions win". It is that architecture comparisons are only meaningful when training recipe, data, augmentation, schedule and compute are matched, and that most published comparisons before 2022 were not. Transformers still hold real advantages: they scale better with very large data, they handle multiple modalities in one stack, and their attention maps are directly inspectable. ConvNeXt shows that at ImageNet scale, with a matched recipe, the architectural gap on classification was small.
What to learn next
- Swin transformer — the model ConvNeXt was measured against.
- Vision transformers — the design ConvNeXt borrowed from.
- Masked image modelling — the pretraining ConvNeXt V2 had to be redesigned for.