CNN Backbones and Pretraining

Receptive fields

The receptive field is how much of the original picture one number in a feature map can see, and it decides whether your network can recognise a whole object at all.

On this page 6
  1. Why this decides whether your model works
  2. How the keyhole widens
  3. The part almost nobody is told
  4. Where you have already seen this
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A receptive field is how much of the original picture one number in a later layer can see.

Think about looking through a keyhole into a room. You see a small circle of the room, and nothing else. Step back from the door and the circle widens. Eventually you see the whole bed, not one corner of the pillow.

Every number inside a network is looking through a keyhole at the original photo. Layers near the input have tiny keyholes. Layers near the output have wide ones.

That widening is not an accident. It is the reason a network can go from noticing edges to naming a whole dog.

Why this decides whether your model works

Suppose one number in the last layer can only see a patch of thirty by thirty pixels. Ask that network whether the photo shows a cat. It has to judge from a patch the size of a stamp.

It will do badly, and no amount of training will fix it. The information never reached that number.

The opposite failure is real too. A network whose keyhole covers the entire photo may struggle to say where something is. Every output number saw the same thing.

Picking a network is partly picking a keyhole size that matches the objects you care about.

How the keyhole widens

Each layer adds a little. Stack two layers that each look at a three-by-three patch and you get a five-by-five view. The second layer looks at neighbours that themselves looked at neighbours.

  input pixels     ...  a b c d e  ...
                        \ | / \ | /
  layer 1 sees 3         P     Q          P sees a,b,c   Q sees c,d,e
                          \   /
  layer 2 sees P and Q      R             R sees a,b,c,d,e  -> five wide

Depth widens the keyhole slowly, one small step at a time.

Shrinking the picture widens it fast. After a stride-two step, every following layer's step counts double, because one square now stands for two original pixels.

This is why networks shrink the picture. It is the cheapest way to see more of it.

The part almost nobody is told

The keyhole is not evenly bright. The pixels in the middle matter enormously, and the pixels at the rim barely register.

Think of a torch beam on a wall. There is a bright centre and a faint fading edge, not a hard circle.

So the size printed on paper is the widest the network could see. The area it actually relies on is much smaller. This surprises people, and it is measured further down this page.

Where you have already seen this

  • A phone camera detecting a face fails when you stand too close and your face fills the frame.
  • Medical scan tools are built with wide keyholes, because a doctor's finding depends on surrounding tissue.
  • Text detection in a document photo needs a narrow keyhole, because a letter's neighbours are distractions.

Remember this

  • The receptive field is how much of the original image one later number can see.
  • Depth widens it slowly; shrinking the picture widens it quickly.
  • The real influence is concentrated in the middle, so the useful field is smaller than the printed one.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install "torch==2.5.1"

Run against PyTorch 2.5.1 on CPU. It takes a few seconds.

Measure it, do not trust the arithmetic

You can compute the receptive field with a formula. You can also measure it, by pushing a gradient backwards from one output pixel and seeing which input pixels feel it. Doing both is how you find out whether you believe your own formula.

receptive_field.py
import torch
import torch.nn as nn

torch.manual_seed(0)


def stack(n_layers, k=3, stride=1, constant=True):
    """A plain stack of convolutions, no padding, one channel throughout."""
    layers = []
    for _ in range(n_layers):
        conv = nn.Conv2d(1, 1, kernel_size=k, stride=stride, padding=0, bias=False)
        if constant:
            nn.init.constant_(conv.weight, 1.0)   # every input pixel gets a real vote
        layers.append(conv)
    return nn.Sequential(*layers)


def measure_rf(net, size=101):
    """Backprop from ONE output pixel and see which input pixels get gradient."""
    x = torch.zeros(1, 1, size, size, requires_grad=True)
    y = net(x)
    c = y.shape[-1] // 2
    grad = torch.autograd.grad(y[0, 0, c, c], x)[0][0, 0]
    touched = (grad.abs() > 0).any(dim=0)          # columns with any influence
    return int(touched.sum()), grad


def formula_rf(kernels, strides):
    r, jump = 1, 1
    for k, s in zip(kernels, strides):
        r = r + (k - 1) * jump
        jump = jump * s
    return r


print("stacked 3x3 convolutions, stride 1")
print(f"{'layers':>6s} {'formula':>8s} {'measured':>9s}")
for L in (1, 2, 3, 5, 10):
    net = stack(L)
    measured, _ = measure_rf(net)
    print(f"{L:6d} {formula_rf([3] * L, [1] * L):8d} {measured:9d}")

print("\nstride changes everything: 5 layers of 3x3")
for strides in ([1, 1, 1, 1, 1], [2, 1, 1, 1, 1], [2, 2, 1, 1, 1], [2, 2, 2, 2, 2]):
    layers = []
    for s in strides:
        conv = nn.Conv2d(1, 1, 3, stride=s, padding=0, bias=False)
        nn.init.constant_(conv.weight, 1.0)
        layers.append(conv)
    net = nn.Sequential(*layers)
    measured, _ = measure_rf(net, size=201)
    print(f"strides {strides} -> formula {formula_rf([3]*5, strides):3d}, measured {measured:3d}")

print("\none 7x7 versus three stacked 3x3")
print("7x7 reach:", formula_rf([7], [1]), " params:", 7 * 7)
print("3x3 x3 reach:", formula_rf([3, 3, 3], [1, 1, 1]), " params:", 3 * 3 * 3)

print("\neffective receptive field: same net, random weights, averaged over 40 inits")
size = 101
total = torch.zeros(size, size)
for seed in range(40):
    torch.manual_seed(seed)
    net = stack(10, constant=False)
    _, grad = measure_rf(net, size=size)
    total += grad.abs()
centre_row = total[size // 2]
lo, hi = size // 2 - 10, size // 2 + 11
peak = centre_row.max()
for i in range(lo, hi):
    share = (centre_row[i] / peak).item()
    bar = "#" * int(round(share * 40))
    print(f"offset {i - size//2:+3d} |{bar:<40s}| {share:.3f}")
touched = int((centre_row > 0).sum())
print(f"\ntheoretical width: {formula_rf([3]*10, [1]*10)} pixels")
print(f"pixels with any gradient at all: {touched}")
print(f"pixels holding 50% of the total gradient: "
      f"{int((centre_row.sort(descending=True).values.cumsum(0) < 0.5 * centre_row.sum()).sum()) + 1}")
Output
stacked 3x3 convolutions, stride 1
layers  formula  measured
     1        3         3
     2        5         5
     3        7         7
     5       11        11
    10       21        21

stride changes everything: 5 layers of 3x3
strides [1, 1, 1, 1, 1] -> formula  11, measured  11
strides [2, 1, 1, 1, 1] -> formula  19, measured  19
strides [2, 2, 1, 1, 1] -> formula  31, measured  31
strides [2, 2, 2, 2, 2] -> formula  63, measured  63

one 7x7 versus three stacked 3x3
7x7 reach: 7  params: 49
3x3 x3 reach: 7  params: 27

effective receptive field: same net, random weights, averaged over 40 inits
offset -10 |                                        | 0.010
offset  -9 |#                                       | 0.022
offset  -8 |####                                    | 0.109
offset  -7 |#####                                   | 0.129
offset  -6 |#########                               | 0.237
offset  -5 |################                        | 0.408
offset  -4 |##################                      | 0.452
offset  -3 |#########################               | 0.628
offset  -2 |#################################       | 0.826
offset  -1 |################################        | 0.789
offset  +0 |########################################| 1.000
offset  +1 |#####################################   | 0.928
offset  +2 |######################                  | 0.541
offset  +3 |###########################             | 0.668
offset  +4 |################                        | 0.389
offset  +5 |##########                              | 0.257
offset  +6 |#######                                 | 0.182
offset  +7 |#####                                   | 0.133
offset  +8 |##                                      | 0.056
offset  +9 |#                                       | 0.026
offset +10 |                                        | 0.008

theoretical width: 21 pixels
pixels with any gradient at all: 21
pixels holding 50% of the total gradient: 5

The individual bar heights come from PyTorch's random initialisation. Different versions or a different platform will shift them a little. The shape does not shift, and the shape is the point.

Reading the output

Formula and measurement agree everywhere. Ten 3x3 layers reach 21 pixels, exactly as 1 + 10 * 2 predicts. The measurement is not a sanity check you can skip once; it is how you catch a wrong stride in a real architecture.

Strides multiply, they do not add. Five layers reach 11 pixels at stride 1 and 63 pixels when every layer strides by 2. Each stride-2 step doubles the value of every later step. This is the single biggest lever you have over receptive field.

Three 3x3 layers reach as far as one 7x7, using 27 weights instead of 49. This is the VGG argument, and it also inserts two extra non-linearities. It is why 3x3 became the default kernel size and 7x7 survives only in stems.

The measured field is 21 pixels wide, and five of them hold half the influence. Look at the bars: the centre is at 1.000 and the rim is at 0.008, more than a hundred times weaker. The printed receptive field is a bound, not a description.

Common mistakes

Quoting a receptive field larger than the input. ResNet-50's theoretical field is 483 pixels while it is trained on 224x224 crops. The number is not wrong, it is a bound over an infinite plane. Treat "field larger than the image" as "the field is not your bottleneck".

Forgetting that padding inflates the count. Padding adds pixels that carry no image content. A field of 21 on a padded 28x28 map includes zeros near the border.

Assuming a wider field is always better. For dense tasks such as segmentation, a huge field blurs boundaries. Architectures like U-Net deliberately mix a wide-field path with a narrow-field one.

Reaching for dilation without thinking about gaps. Dilated convolutions widen the field cheaply but sample the input on a grid with holes. Stacking several with the same dilation rate produces gridding artefacts, which is why the DeepLab family cycles through different rates.

Try it yourself

Add padding=1 to every layer in stack and re-measure. The measured field should stay at 21 for ten layers, because padding changes the output size and not the reach. Then set stride=2 on the first layer only, and predict the answer before you run it.

What to learn next

Researcher — Mathematics and papers.

Recurrence

For a chain of layers indexed $l = 1 \dots L$ with kernel sizes $k_l$, strides $s_l$ and dilations $d_l$, the receptive field and the jump (input pixels per output step) evolve as:

$$ j_0 = 1, \quad r_0 = 1 $$ $$ j_l = j_{l-1} \cdot s_l, \qquad r_l = r_{l-1} + \bigl(d_l (k_l - 1)\bigr) \cdot j_{l-1} $$

$r_l$ is the receptive field after layer $l$, and $j_{l-1}$ is the cumulative stride of everything before it. The closed form is

$$ r_L = 1 + \sum_{l=1}^{L} d_l (k_l - 1) \prod_{i=1}^{l-1} s_i . $$

The product term is why strides dominate: an early stride-2 multiplies the contribution of every later layer.

Padding does not enter this relation. It shifts the centre of the field, tracked by a second recurrence, and is what determines whether the field of the centre output pixel is aligned with the centre of the input.

Araujo, Norris and Sim (2019), Computing Receptive Fields of Convolutional Neural Networks, Distill, gives the general treatment including branched graphs, where the field of a merge node is the maximum over its inputs. Their table of theoretical fields at the final convolutional layer: AlexNet v2 195, VGG-16 212, ResNet v1-50 483, Inception v2 699, ResNet v1-101 1027, Inception v3 1311, ResNet v1-152 1507. Effective stride is 32 for all of them.

Effective receptive field

Luo, Li, Urtasun and Zemel, Understanding the Effective Receptive Field in Deep Convolutional Neural Networks, NeurIPS 2016, posted to arXiv in January 2017 at arxiv.org/abs/1701.04128, analyse the distribution of $\partial y_c / \partial x_{ij}$ rather than its support.

Their central result: for a stack of $n$ convolutions with independent random weights, the gradient magnitude at the input approaches a two-dimensional Gaussian as $n$ grows, by the central limit theorem applied to the many paths from input to output. Consequently the effective field grows as $O(\sqrt{n})$ while the theoretical field grows as $O(n)$, and the ratio $\text{ERF}/\text{RF}$ shrinks as $O(1/\sqrt{n})$.

Two practical corollaries they establish:

  • Training expands the effective field. The Gaussian result holds at initialisation; learned weights concentrate paths and widen the useful region, though not to the theoretical bound.
  • Subsampling and dilation enlarge the effective field more efficiently than depth, because they raise the jump rather than adding paths.

The measurement in the developer block is a small reproduction of their method: the summed absolute input gradient of a single central output unit, averaged over random initialisations.

Design consequences

Detection. Feature-pyramid designs assign objects to levels whose stride, and therefore field, matches object scale. An anchor whose object is much larger than the level's effective field is unrecoverable by training.

Segmentation. Atrous spatial pyramid pooling, from the DeepLab line (Chen et al., 2017), enlarges the field at fixed stride by running parallel dilated branches at several rates, avoiding the resolution loss that striding would cause.

Transformers. Global self-attention has a full-image receptive field in one layer, which is exactly why it lacks the locality prior of convolution and needs more data or stronger augmentation to compensate. Swin's windowed attention deliberately reintroduces a bounded, growing field.

Measuring it on a real network

The gradient method used above transfers directly. Feed a zero image with requires_grad=True, take a scalar from one spatial position of the target layer, call torch.autograd.grad, and inspect the support and the mass of the input gradient. Zero input is deliberate: it removes data-dependent gating in ReLU units only when biases are zero, so on a trained network use several real images and average, and report the effective field rather than the support.

What to learn next