Feature pyramid networks
An FPN carries deep meaning back down to high-resolution layers, so a detector can find small objects without giving up the understanding it built at low resolution.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A feature pyramid network sends deep understanding back down to the detailed layers, so small objects get both.
The analogy
Think about looking for a friend at a railway station. From the footbridge you see the whole platform and understand the scene. That is a crowd, that is the tea stall, there is the exit.
Walk down to the platform and you can see faces in sharp detail. But you have lost the layout, and cannot tell which end you are at.
The best way to find your friend is both. Keep the map in your head from the bridge, and use the detail you have on the ground. A feature pyramid network gives a detector exactly this combination.
Why it exists
A convolutional network shrinks the picture as it goes deeper. Early layers hold fine detail but understand almost nothing. Deep layers understand a great deal but have thrown the detail away.
Detect from the deep layers only, and small objects disappear. A distant bird occupies less than one grid position, so there is nothing left to detect.
Detect from the shallow layers only, and you have detail with no understanding. The layer knows there is an edge. It has no idea whether the edge belongs to a bird.
Before FPNs, people ran the whole detector several times on resized copies of the picture. That worked and was slow.
How it works
going down (the normal network) coming back up
------------------------------ -----------------
big picture, fine detail <--- detail + meaning <-- detect small things here
| ^
smaller, more meaning <--- |
| ^
smallest, most meaning ----> start here <-- detect big things hereThe network runs down as usual, shrinking the picture. Then a second path climbs back up.
At each step going up, the deep understanding is enlarged and added to the detailed layer sitting at that size. Every level ends up with detail and meaning.
Detection then runs at every level. Small objects are handled where detail is fine. Large objects are handled where the view is wide.
What makes this a good idea
Adding a second path costs very little. No extra image passes, no resizing the input, no running the model five times.
The result is that one network handles objects from a few pixels across to most of the frame. Nearly every detector built since 2017 has one, including the YOLO family.
Where you have seen the result
- A phone camera focusing on a face at the back of a group photo.
- A drone survey counting both trees and individual fruits.
- A road camera reading a number plate and tracking the whole vehicle.
- A medical scan tool flagging a small nodule inside a large organ.
Remember this
- Deep layers understand a lot but hold little detail; shallow layers are the opposite.
- An FPN adds a path that carries deep understanding back down into the detailed layers.
- Detection then runs at several sizes at once, from one pass over the picture.
What to learn next
- RoI pooling and RoI align — pulling a fixed-size patch out of a pyramid level.
- Convolutional neural networks — why deep layers lose detail in the first place.
- Detecting very small objects — where the pyramid runs out of road.
Developer — Code and libraries.
Setup
pip install torch torchvisionWritten against torch 2.5.1 and torchvision 0.20.1. torchvision.ops.FeaturePyramidNetwork takes an OrderedDict of feature maps and returns one with the same keys.
Building a pyramid and proving the top-down path works
import math
from collections import OrderedDict
import torch
from torchvision.ops import FeaturePyramidNetwork
torch.manual_seed(0)
# What a ResNet hands over for a 640x640 image: deep maps are small but rich.
backbone_out = OrderedDict([
("c2", torch.rand(1, 256, 160, 160)), # stride 4
("c3", torch.rand(1, 512, 80, 80)), # stride 8
("c4", torch.rand(1, 1024, 40, 40)), # stride 16
("c5", torch.rand(1, 2048, 20, 20)), # stride 32
])
for k, v in backbone_out.items():
print(f"{k}: {tuple(v.shape)} stride {640 // v.shape[-1]}")
fpn = FeaturePyramidNetwork(in_channels_list=[256, 512, 1024, 2048], out_channels=256)
with torch.no_grad():
pyramid = fpn(backbone_out)
print("\nafter the FPN, every level speaks the same language (256 channels):")
for k, v in pyramid.items():
print(f"{k}: {tuple(v.shape)}")
print(f"\nFPN parameters: {sum(p.numel() for p in fpn.parameters()):,}"
" (a 1x1 lateral conv and a 3x3 output conv per level)")
# Proof that the top-down path really carries deep meaning downwards:
# change ONLY the deepest map and watch every level move, including the shallowest.
with torch.no_grad():
changed = OrderedDict(backbone_out)
changed["c5"] = torch.rand(1, 2048, 20, 20)
pyramid2 = fpn(changed)
print("\nmean change per level after replacing c5 alone:")
for k in pyramid:
print(f" {k}: {(pyramid[k] - pyramid2[k]).abs().mean():.4f}")
# Which level should a given object be detected on? Faster R-CNN's FPN rule.
def level_for(w, h, k0=4, canonical=224):
return int(min(5, max(2, math.floor(k0 + math.log2(math.sqrt(w * h) / canonical + 1e-6)))))
print("\nobject size -> pyramid level (Faster R-CNN FPN assignment):")
for w, h, what in [(20, 20, "a distant bird"), (60, 90, "a helmet"),
(150, 300, "a person"), (400, 380, "a lorry")]:
lvl = level_for(w, h)
print(f" {what:<16}{w:>4} x{h:<4} -> P{lvl}, stride {2**lvl}")
print("\ngrid cells a 24x24 object covers at each stride:")
for stride in [4, 8, 16, 32]:
print(f" stride {stride:>2}: {(24 / stride) ** 2:6.2f} cells")c2: (1, 256, 160, 160) stride 4 c3: (1, 512, 80, 80) stride 8 c4: (1, 1024, 40, 40) stride 16 c5: (1, 2048, 20, 20) stride 32 after the FPN, every level speaks the same language (256 channels): c2: (1, 256, 160, 160) c3: (1, 256, 80, 80) c4: (1, 256, 40, 40) c5: (1, 256, 20, 20) FPN parameters: 3,344,384 (a 1x1 lateral conv and a 3x3 output conv per level) mean change per level after replacing c5 alone: c2: 0.3250 c3: 0.3234 c4: 0.3201 c5: 0.3128 object size -> pyramid level (Faster R-CNN FPN assignment): a distant bird 20 x20 -> P2, stride 4 a helmet 60 x90 -> P2, stride 4 a person 150 x300 -> P3, stride 8 a lorry 400 x380 -> P4, stride 16 grid cells a 24x24 object covers at each stride: stride 4: 36.00 cells stride 8: 9.00 cells stride 16: 2.25 cells stride 32: 0.56 cells
Reading that output
Channel counts collapse to 256 everywhere. The backbone hands over 256, 512, 1024 and 2048 channels. The FPN's 1x1 lateral convolutions project all of them to a common width. That is what lets one shared detection head run on every level. A head with different weights per level would be a different architecture, and a worse one.
Spatial sizes are untouched. 160, 80, 40 and 20 go in and come out. The FPN changes what each cell means, not where cells are.
The top-down test is the important one. Replacing only c5 changed c2 by 0.3250 on average. c2 never touched c5 directly. The signal travelled: c5 was upsampled and added into c4, which was upsampled and added into c3, and so on. That chain is the entire point of the architecture.
These magnitudes come from an untrained FPN with random weights, seeded at zero for reproducibility. Do not read meaning into the exact values. Read meaning into the fact that none of them is zero.
3.3 million parameters. A ResNet-50 backbone is around 25 million. The pyramid costs about an eighth of the backbone and lifts small-object accuracy substantially. That ratio is why it became standard.
The level assignment rule sends a 20-pixel bird and a 60-by-90 helmet to the same level. The formula floors at P2. Everything smaller than roughly 112 pixels on a side lands on the finest level.
This is a known weakness. The finest level carries most of the work and most of the small objects. Detectors for tiny objects add a P1 level, or tile the image instead.
The last table is the reason all of this exists. A 24-pixel object covers 0.56 cells at stride 32. Less than one cell. There is nothing to classify. At stride 4 it covers 36 cells, which is plenty.
Adding an FPN to a real backbone
import torch
from torchvision.models import resnet50
from torchvision.models.detection.backbone_utils import BackboneWithFPN
body = resnet50(weights=None) # weights=None: nothing is downloaded
backbone = BackboneWithFPN(
body,
return_layers={"layer1": "0", "layer2": "1", "layer3": "2", "layer4": "3"},
in_channels_list=[256, 512, 1024, 2048],
out_channels=256,
)
with torch.no_grad():
out = backbone(torch.zeros(1, 3, 640, 640))
for k, v in out.items():
print(f"level {k}: {tuple(v.shape)} stride {640 // v.shape[-1]}")
print("\nbackbone+FPN parameters:", f"{sum(p.numel() for p in backbone.parameters()):,}")level 0: (1, 256, 160, 160) stride 4 level 1: (1, 256, 80, 80) stride 8 level 2: (1, 256, 40, 40) stride 16 level 3: (1, 256, 20, 20) stride 32 level pool: (1, 256, 10, 10) stride 64 backbone+FPN parameters: 26,852,416
BackboneWithFPN adds a pool level by max-pooling the deepest output, giving a stride-64 map for very large objects. Detection models in torchvision expect that key to exist. Worth knowing before you write your own backbone wrapper.
The pyramid variants you will meet in configs
| Name | Change | Trade |
|---|---|---|
| FPN | top-down path with lateral additions | the baseline |
| PANet | adds a second bottom-up path on top | better localisation, more compute |
| BiFPN | learnable weights per input, repeated blocks | used by EfficientDet |
| NAS-FPN | connection pattern found by architecture search | strong, hard to modify |
| PAFPN | PANet-style, used across Ultralytics YOLO models | good speed and accuracy balance |
Every one of them is the same idea: move information between resolutions before detecting. They differ in which direction, how many times, and whether the mixing weights are learned.
Common mistakes
Feeding a plain dict instead of an OrderedDict. Level ordering decides which map is upsampled into which. In Python 3.7 and later a plain dict preserves insertion order, so this usually works. The API asks for OrderedDict, and being explicit costs nothing.
Mismatched in_channels_list. It must line up with the maps you pass, in the same order. A mismatch fails at the first lateral convolution with a shape error that names a layer you did not write.
Detecting from the raw backbone C maps. Those have not been mixed. Use the pyramid outputs, not the backbone outputs.
Assuming more levels is better. Each level costs memory and adds candidate boxes. Levels that produce no positives on your data are pure cost. Print positives per level once and delete the dead ones.
Ignoring where your objects actually live. If everything in your dataset is 20 pixels across, every object lands on one level. The rest of the pyramid is wasted. Consider a finer level or tiling instead.
Try it yourself
Add a stride-2 level to the first snippet by prepending a (1, 128, 320, 320) map called c1. Print the new parameter count and the new total cell count across the pyramid. That cell count is how many candidate positions your head must score every frame. It is the honest price of detecting smaller things.
What to learn next
- RoI pooling and RoI align — pulling a fixed-size patch out of a pyramid level.
- Convolutional neural networks — why deep layers lose detail in the first place.
- Detecting very small objects — where the pyramid runs out of road.
Researcher — Mathematics and papers.
The construction
Let $C_l$ be the backbone output at stride $2^l$. The pyramid is built top-down:
$$ P_L = \mathrm{conv}_{1\times1}(C_L), \qquad P_l = \mathrm{conv}{3\times3}\big(\mathrm{conv}{1\times1}(C_l) + \mathrm{up}{2\times}(P{l+1})\big) $$
Where $\mathrm{up}_{2\times}$ is nearest-neighbour upsampling and $L$ is the deepest level. The $1\times1$ convolution equalises channel count. The $3\times3$ convolution after the addition suppresses the aliasing that nearest-neighbour upsampling introduces.
Lin et al. (2017) report the ablations that matter. Removing the top-down path costs a large amount of AP. Removing the lateral connections costs a large amount as well. Neither path alone is sufficient. The merges must be element-wise additions rather than concatenations, or the parameter budget stops being this small.
Level assignment
The RoI-to-level rule, from the same paper:
$$ k = \left\lfloor k_0 + \log_2!\left(\frac{\sqrt{wh}}{224}\right) \right\rfloor $$
Here $w, h$ are the RoI's width and height and $224$ is the ImageNet pre-training resolution. $k_0 = 4$ is the level a $224 \times 224$ RoI maps to. The result is clamped to the available levels.
The intent is to pool an RoI from a level whose stride leaves it near the backbone's pre-training size. The choice of $224$ is inherited from ImageNet, not derived. It is a reasonable target for a retuning experiment on unusual object sizes.
Effective receptive field, and why the pyramid is not the whole story
A stride-4 level has a large theoretical receptive field after an FPN, because information flowed down from stride 32. The effective receptive field is Gaussian-weighted and much smaller (Luo et al., 2016). Context reaching a shallow cell is weaker than the diagram implies.
This is one reason attention-based necks and encoder layers over multi-scale features became attractive. Deformable DETR's multi-scale deformable attention (Zhu et al., 2021) removes the chain. Every level can query sampling points on every other level directly.
Variants and what they actually changed
- PANet (Liu et al., 2018) adds a bottom-up path after the top-down one. That shortens the route from fine detail to the deepest level. Adopted widely in YOLO necks.
- NAS-FPN (Ghiasi et al., 2019) searched over merge topologies. Strong results, poor interpretability, hard to adapt.
- BiFPN (Tan et al., 2020, EfficientDet) removes single-input nodes and adds a residual edge at the same level. It learns a normalised scalar weight per input: $O = \sum_i \frac{w_i}{\epsilon + \sum_j w_j} \cdot I_i$, with $w_i \geq 0$ enforced by ReLU. Repeating the block is a depth knob.
- Recursive FPN (Qiao et al., 2021, DetectoRS) feeds the pyramid back into the backbone.
Gains across these are single-digit AP and interact heavily with the training recipe. Liu et al. (2022), A ConvNet for the 2020s, is a useful caution here. Much of the reported gap between architectures in this period came from recipes, not structure.
Cost
For $C$ channels and level sizes $H_l \times W_l$, the pyramid adds $\sum_l (C_l^{in} \cdot C + 9 C^2)$ parameters and FLOPs proportional to $\sum_l H_l W_l C^2$. The finest level dominates: at stride 4, $H_1 W_1$ is 16 times $H_3 W_3$ at stride 16.
That is why detectors targeting small objects usually start at stride 8, not stride 4. Adding a P1 level is expensive enough to be a deliberate decision rather than a default.
Papers
- Lin et al., Feature Pyramid Networks for Object Detection, CVPR 2017 — arxiv.org/abs/1612.03144
- Liu et al., Path Aggregation Network for Instance Segmentation, CVPR 2018 — arxiv.org/abs/1803.01534
- Tan et al., EfficientDet: Scalable and Efficient Object Detection, CVPR 2020 — arxiv.org/abs/1911.09070
- Ghiasi et al., NAS-FPN, CVPR 2019 — arxiv.org/abs/1904.07392
- Luo et al., Understanding the Effective Receptive Field in Deep CNNs, NeurIPS 2016 — arxiv.org/abs/1701.04128
What to learn next
- RoI pooling and RoI align — pulling a fixed-size patch out of a pyramid level.
- Convolutional neural networks — why deep layers lose detail in the first place.
- Detecting very small objects — where the pyramid runs out of road.