Padding, stride and output shapes
Four numbers decide how big a feature map comes out of a convolution, and getting them wrong is the most common error a beginner meets in vision code.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Four numbers decide how big the result of a convolution is. They are the picture size, the filter size, the padding and the stride.
Picture a sheet of rolled chapati dough on the counter. You press a round cutter into it, lift it, shift sideways, press again. How many rounds you get depends on three things. The size of the sheet, the width of the cutter, and how far you shift it.
A convolution does the same thing to a picture. The cutter is the filter, a small grid of numbers that slides over the image looking for one pattern.
Two words describe how you move the cutter. Stride is how far you shift it between presses. Padding is extra blank dough you add around the edge so the cutter can reach the corners.
Why this matters more than it sounds
Layers in a network are joined end to end. The size coming out of one layer is the size going into the next.
Get one number wrong and the whole chain breaks. Nearly every beginner in computer vision meets the same error message about shapes that cannot be multiplied.
There is a second reason to care. Without padding, every convolution shaves a thin border off the picture.
Do that twenty times and a small image runs out of pixels entirely. Padding is what lets networks be deep.
How the sliding works
Think of a strip of ten squares and a cutter three squares wide, shifted one square at a time.
strip: [1][2][3][4][5][6][7][8][9][10]
press 1: [=====] covers 1,2,3
press 2: [=====] covers 2,3,4
press 3: [=====] covers 3,4,5
...
press 8: [=====] covers 8,9,10
eight presses fit, so the output strip is eight squares longThe strip got shorter, from ten down to eight. That is the shaving.
Now add one blank square at each end, which is padding of one. The strip becomes twelve squares, ten presses fit, and the output matches the input. This pairing is so useful that it is the default choice almost everywhere.
What stride does
Stride of one means shift by one square, so the presses overlap heavily. Stride of two means skip a square each time, so about half as many presses fit.
That halving is deliberate. Networks use it to shrink the picture on purpose, so later layers have less work to do.
stride one -> output is about the same size as the input
stride two -> output is about half the height and half the widthWhere you have already seen this
- Your phone's portrait mode blurs the background using a network built from these layers.
- A number plate reader at a toll booth shrinks a wide photo down through several stride-two steps.
- Every cartoon filter in a camera app runs a stack of these slides in under a second.
The honest part
The arithmetic has a rounding step in it. When the cutter does not fit a whole number of times, the leftover strip is thrown away.
No warning is printed. Your model trains fine and quietly never looks at the last column of every image. Read that twice, because it catches experienced people too.
Remember this
- Output size depends on the picture size, the filter size, the padding and the stride.
- Padding of one with a three-wide filter keeps the size unchanged, which is why that pairing is everywhere.
- Stride of two halves the height and the width on purpose.
What to learn next
- Receptive fields — how far back into the image one output pixel can see.
- Convolutional neural networks — the operation itself, if any of this felt rushed.
- Tensor shapes and broadcasting — the general rules behind every shape error.
Developer — Code and libraries.
Setup
pip install "torch==2.5.1"Written and run against PyTorch 2.5.1 on CPU. These Conv2d shape rules have been stable for years and are unchanged in the current release.
The formula, checked against PyTorch on every line
import torch
import torch.nn as nn
def out_size(n, k, p, s, d=1):
"""Height (or width) after one conv, from the PyTorch Conv2d formula."""
return (n + 2 * p - d * (k - 1) - 1) // s + 1
cases = [
# (input, kernel, padding, stride, dilation, nickname)
(28, 3, 0, 1, 1, "valid 3x3"),
(28, 3, 1, 1, 1, "same 3x3"),
(28, 5, 2, 1, 1, "same 5x5"),
(28, 3, 1, 2, 1, "stride-2 3x3"),
(28, 7, 3, 2, 1, "ResNet stem"),
(28, 3, 2, 1, 2, "dilated 3x3"),
(7, 2, 0, 2, 1, "2x2 on odd input"),
]
print(f"{'case':18s} {'in':>3s} {'k':>2s} {'p':>2s} {'s':>2s} {'d':>2s} {'formula':>8s} {'pytorch':>8s}")
for n, k, p, s, d, name in cases:
predicted = out_size(n, k, p, s, d)
conv = nn.Conv2d(1, 1, kernel_size=k, padding=p, stride=s, dilation=d)
actual = conv(torch.zeros(1, 1, n, n)).shape[-1]
flag = "ok" if predicted == actual else "MISMATCH"
print(f"{name:18s} {n:3d} {k:2d} {p:2d} {s:2d} {d:2d} {predicted:8d} {actual:8d} {flag}")
print("\n-- padding='same' does the arithmetic for you (stride 1 only) --")
same = nn.Conv2d(1, 1, kernel_size=5, padding="same")
print("5x5 same on 28x28 ->", tuple(same(torch.zeros(1, 1, 28, 28)).shape))
print("\n-- what a stack does to a 224x224 photo --")
stem = nn.Sequential(
nn.Conv2d(3, 64, kernel_size=7, stride=2, padding=3),
nn.MaxPool2d(kernel_size=3, stride=2, padding=1),
nn.Conv2d(64, 128, kernel_size=3, stride=2, padding=1),
nn.Conv2d(128, 256, kernel_size=3, stride=2, padding=1),
)
x = torch.zeros(1, 3, 224, 224)
print(f"{'input':12s} {tuple(x.shape)}")
for layer in stem:
x = layer(x)
print(f"{layer.__class__.__name__:12s} {tuple(x.shape)}")
print("\n-- the silent truncation --")
n, k, p, s = 10, 3, 0, 2
o = out_size(n, k, p, s)
covered = (o - 1) * s + k
print(f"{n}x{n} through a stride-{s} {k}x{k} with no padding -> {o}x{o}")
print(f"the windows reach column {covered} of {n}: the last {n - covered} column is dropped, silently")case in k p s d formula pytorch valid 3x3 28 3 0 1 1 26 26 ok same 3x3 28 3 1 1 1 28 28 ok same 5x5 28 5 2 1 1 28 28 ok stride-2 3x3 28 3 1 2 1 14 14 ok ResNet stem 28 7 3 2 1 14 14 ok dilated 3x3 28 3 2 1 2 28 28 ok 2x2 on odd input 7 2 0 2 1 3 3 ok -- padding='same' does the arithmetic for you (stride 1 only) -- 5x5 same on 28x28 -> (1, 1, 28, 28) -- what a stack does to a 224x224 photo -- input (1, 3, 224, 224) Conv2d (1, 64, 112, 112) MaxPool2d (1, 64, 56, 56) Conv2d (1, 128, 28, 28) Conv2d (1, 256, 14, 14) -- the silent truncation -- 10x10 through a stride-2 3x3 with no padding -> 4x4 the windows reach column 9 of 10: the last 1 column is dropped, silently
Reading the output
Every row says ok. The formula and PyTorch agree on all seven cases, including the dilated one. You can work out any shape on paper before writing the layer.
Padding of one with a 3x3 kernel keeps the size. So does padding of two with a 5x5. The pattern is padding equal to half the kernel size, rounded down, and it works only for odd kernels. That is the whole reason 3x3, 5x5 and 7x7 dominate and 4x4 does not.
Dilation widens the reach without adding weights. The dilated row uses a 3x3 kernel with gaps. It covers a 5x5 area, so it needs padding of two to hold the size. Segmentation models lean on this heavily.
The stem row is the ResNet stem. A 7x7 kernel with stride 2 and padding 3 halves the picture in one step. That is why 224 became 112 in the stack below.
The stack shows the whole trade. Spatial size falls 224, 112, 56, 28, 14 while channels climb 3, 64, 64, 128, 256. Every backbone in this section does this.
The last block is the trap. A 10-wide input with a stride-2 3x3 and no padding produces 4 outputs, and the windows only ever touch the first 9 columns. Column 10 affects nothing. Nothing warns you.
Common mistakes
Forgetting that the batch dimension exists. Conv2d wants (N, C, H, W). A lone image is (C, H, W) and raises Expected 4-dimensional input. Fix it with x.unsqueeze(0).
Hard-coding the Linear input size. After Flatten, the linear layer needs channels * height * width. Change a stride anywhere upstream and it breaks. Use nn.LazyLinear(num_classes), which reads the size on the first forward pass, or print the shapes as this script does.
Using padding='same' with a stride above one. PyTorch raises ValueError: padding='same' is not supported for strided convolutions. The string form works at stride 1 only.
Assuming pooling rounds the same way as convolution. MaxPool2d has a ceil_mode argument that defaults to False, matching convolution. Set it to True and an odd-sized input rounds up instead of down, which changes every shape downstream.
Copying a TensorFlow model layer by layer. TensorFlow's padding="SAME" pads asymmetrically when the numbers do not divide evenly. PyTorch pads symmetrically. Ported models can sit one pixel off and produce slightly wrong outputs with no error at all.
Try it yourself
Take a 224x224 input and design a stack that ends at exactly 7x7, using only stride-2 3x3 convolutions with padding 1. Predict the number of layers on paper first, then check it with the loop above. Then repeat with a 225x225 input and watch where the rounding bites.
What to learn next
- Receptive fields — how far back into the image one output pixel can see.
- Convolutional neural networks — the operation itself, if any of this felt rushed.
- Tensor shapes and broadcasting — the general rules behind every shape error.
Researcher — Mathematics and papers.
The output-size relation
For a single spatial dimension, PyTorch's Conv2d computes:
$$ H_{\text{out}} = \left\lfloor \frac{H_{\text{in}} + 2P - D\,(K - 1) - 1}{S} \right\rfloor + 1 $$
$H_{\text{in}}$ is the input height, $P$ the padding added to each side, $K$ the kernel size, $S$ the stride, and $D$ the dilation. With $D = 1$ this collapses to the familiar $\lfloor (H_{\text{in}} + 2P - K)/S \rfloor + 1$.
The term $D(K-1) + 1$ is the effective kernel size: the width of the input span a dilated kernel touches. A 3x3 kernel at $D = 2$ spans 5 pixels while holding 9 weights.
Size preservation and the parity argument
Size is preserved when $H_{\text{out}} = H_{\text{in}}$ at $S = 1$, which requires
$$ 2P = D\,(K-1). $$
For $D = 1$ this gives $P = (K-1)/2$, an integer only when $K$ is odd. Even kernels need asymmetric padding, and the layer must then choose which side receives the extra row. PyTorch declines to choose; TensorFlow's SAME mode chooses the bottom and the right. This is a genuine source of numerical divergence when porting weights between the two frameworks.
Coverage and the discarded border
The largest input index any window touches is
$$ (H_{\text{out}} - 1)\,S + D\,(K-1) + 1 - P. $$
When this falls below $H_{\text{in}}$, the trailing rows influence no output at all. The floor in the size relation is what produces the loss. MaxPool2d(ceil_mode=True) replaces the floor with a ceiling and pads the deficit implicitly, which is why some ported architectures need it to reproduce their published feature-map sizes.
Transposed convolution, and its ambiguity
The transposed convolution inverts the shape map, not the operation:
$$ H_{\text{out}} = (H_{\text{in}} - 1)\,S - 2P + D\,(K-1) + \text{output padding} + 1 $$
Because the forward floor is many-to-one, several input sizes map to one output size, so the inverse is not unique. The output_padding argument resolves the ambiguity by hand. Dumoulin and Visin (2016), A guide to convolution arithmetic for deep learning, arxiv.org/abs/1603.07285, works through every case with diagrams and remains the reference.
Transposed convolution with $K$ not divisible by $S$ produces the checkerboard artefacts documented by Odena, Dumoulin and Olah (2016), Deconvolution and Checkerboard Artifacts, Distill. Resize-then-convolve avoids them.
Cost
For a layer with $C_{\text{in}}$ input channels, $C_{\text{out}}$ filters and a $K \times K$ kernel:
- Parameters: $C_{\text{out}} C_{\text{in}} K^2 + C_{\text{out}}$, independent of resolution.
- Multiply-accumulates: $H_{\text{out}} W_{\text{out}} \cdot C_{\text{out}} C_{\text{in}} K^2$.
Published FLOP counts for vision models almost always mean multiply-accumulates, one per multiply-and-add pair. PyTorch's torch.utils.flop_counter.FlopCounterMode counts two FLOPs per MAC, so its numbers are twice the figures quoted in papers. Halve them before comparing.
A stride $S > 1$ divides the MAC count by $S^2$ while leaving the parameter count untouched, which is why downsampling early is the cheapest structural saving available.
Grouped and depthwise variants
With $g$ groups, input and output channels are partitioned and each group is convolved independently:
$$ \text{params} = \frac{C_{\text{out}} C_{\text{in}} K^2}{g} + C_{\text{out}}. $$
At $g = C_{\text{in}} = C_{\text{out}}$ this is a depthwise convolution and the cost falls by a factor of $C$. The shape relation is unchanged; only the parameter and MAC counts move. That separation of shape from cost is what the whole MobileNet line exploits.
References
- PyTorch
Conv2ddocumentation, docs.pytorch.org/docs/stable/generated/torch.nn.Conv2d.html — the relation above is quoted from it. - Dumoulin and Visin, A guide to convolution arithmetic for deep learning, 2016.
- Odena, Dumoulin and Olah, Deconvolution and Checkerboard Artifacts, Distill, 2016.
What to learn next
- Receptive fields — how far back into the image one output pixel can see.
- Convolutional neural networks — the operation itself, if any of this felt rushed.
- Tensor shapes and broadcasting — the general rules behind every shape error.