R-CNN, Fast R-CNN and Faster R-CNN
Three papers in two years turned detection from a slow search over 2,000 crops into one network that proposes and classifies regions in a single pass.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Three versions of the same idea, each one removing a slow step from the version before.
The analogy
Think about finding a name in a thick printed phone directory.
The slowest way is to photocopy every page separately and read each copy on its own. That is R-CNN: cut out two thousand pieces of the picture and run a whole network on each one.
Faster is to read the book once, then look at whichever pages you need. That is Fast R-CNN: run the network on the picture once and pick regions out of the result.
Fastest is to have the book tell you which pages to look at. That is Faster R-CNN: the network proposes the regions itself.
Why each step happened
R-CNN worked, and it was painfully slow. Something outside the network chose two thousand candidate regions. Each region then went through a full network of its own. One picture took most of a minute.
Fast R-CNN noticed the waste. Neighbouring crops overlap heavily, so the same pixels were being processed again and again. Run the network once over the whole picture, then cut the regions out of the result. The paper reports training the deep network nine times faster and testing 213 times faster than R-CNN.
Faster R-CNN removed the last outside step. The region proposals were still coming from a separate hand-written search. Replacing it with a small network sharing the same features made the whole detector one trainable thing.
How it works
R-CNN picture -> outside search -> 2000 crops -> network per crop -> answers
Fast R-CNN picture -> network once -> cut regions out -> small head -> answers
^
outside search still picks the regions
Faster R-CNN picture -> network once -> [ tiny network proposes regions ]
-> cut regions out -> small head -> answersThe stages stayed the same. What changed is how much of the work happens inside one network that can be trained together.
Why anyone still uses this
Two-stage detectors are not the fastest. They are still among the most accurate, and they are steadier on unusual data.
The reason is the second look. Stage one asks a cheap question: is anything here at all? Stage two looks properly at only the survivors, and gets to be careful because there are far fewer of them.
Suppose your job is counting defects on a production line at one picture per second. That carefulness is worth more than speed.
Where you have seen the result
- Quality checks on a factory line flagging cracked parts.
- Satellite tools counting buildings in a district.
- Medical scan software marking regions for a doctor.
- Document tools finding tables and stamps on a scanned page.
Remember this
- R-CNN ran a full network on every candidate crop, which was accurate and very slow.
- Fast R-CNN ran the network once and cut regions out of the result.
- Faster R-CNN moved region proposing inside the network, making the whole thing trainable end to end.
What to learn next
- YOLO versions compared — the one-stage answer to all of this.
- RoI pooling and RoI align — the operation the second stage depends on.
- Object detection — the introduction, if any of this moved too quickly.
Developer — Code and libraries.
Setup
pip install torch torchvisionWritten against torch 2.5.1 and torchvision 0.20.1. Every snippet uses weights=None, so nothing is downloaded and everything runs on CPU.
Opening up a real Faster R-CNN
import torch
from torchvision.models.detection import fasterrcnn_resnet50_fpn
torch.manual_seed(0)
# weights=None downloads nothing. The architecture is what we are inspecting.
model = fasterrcnn_resnet50_fpn(weights=None, weights_backbone=None, num_classes=21)
groups = {}
for name, p in model.named_parameters():
top = name.split(".")[0]
groups[top] = groups.get(top, 0) + p.numel()
print("where the parameters live:")
for k, v in groups.items():
print(f" {k:<12}{v:>12,}")
print(f" {'TOTAL':<12}{sum(groups.values()):>12,}")
print("\nstage 1, the region proposal network:")
print(" proposals kept before NMS:", model.rpn._pre_nms_top_n)
print(" proposals kept after NMS :", model.rpn._post_nms_top_n)
print(" NMS threshold :", model.rpn.nms_thresh)
print(" positive / negative IoU :", model.rpn.proposal_matcher.high_threshold,
"/", model.rpn.proposal_matcher.low_threshold)
print("\nstage 2, the box head:")
print(" RoI pooler :", type(model.roi_heads.box_roi_pool).__name__,
"output", model.roi_heads.box_roi_pool.output_size)
print(" score threshold :", model.roi_heads.score_thresh)
print(" detections per image :", model.roi_heads.detections_per_img)
# Training mode returns four losses, one per learned job.
model.train()
losses = model([torch.rand(3, 320, 320)],
[{"boxes": torch.tensor([[50., 60., 180., 240.]]),
"labels": torch.tensor([3])}])
print("\nthe four losses Faster R-CNN optimises together:")
for k, v in losses.items():
print(f" {k:<22}{v.item():.4f}")
# Eval mode returns detections instead.
model.eval()
with torch.no_grad():
out = model([torch.rand(3, 320, 320)])[0]
print("\neval output keys:", list(out.keys()))
print("shapes:", {k: tuple(v.shape) for k, v in out.items()})where the parameters live:
backbone 26,852,416
rpn 593,935
roi_heads 14,003,305
TOTAL 41,449,656
stage 1, the region proposal network:
proposals kept before NMS: {'training': 2000, 'testing': 1000}
proposals kept after NMS : {'training': 2000, 'testing': 1000}
NMS threshold : 0.7
positive / negative IoU : 0.7 / 0.3
stage 2, the box head:
RoI pooler : MultiScaleRoIAlign output (7, 7)
score threshold : 0.05
detections per image : 100
the four losses Faster R-CNN optimises together:
loss_classifier 3.2444
loss_box_reg 0.0022
loss_objectness 0.7168
loss_rpn_box_reg 0.0111
eval output keys: ['boxes', 'labels', 'scores']
shapes: {'boxes': (100, 4), 'labels': (100,), 'scores': (100,)}Reading that output
The RPN is tiny: 594 thousand parameters against 41 million total. That is under one and a half percent of the model. Faster R-CNN's headline contribution costs almost nothing, because it reuses the backbone features that were already computed.
The four losses are the four learned jobs. loss_objectness is stage one asking "object or not". loss_rpn_box_reg is stage one nudging anchors. loss_classifier is stage two naming the class. loss_box_reg is stage two refining the box.
The loss values come from randomly initialised weights with torch.manual_seed(0). They will move with a different seed or a different torch version. Treat them as a smoke test, not a benchmark. One is worth sanity-checking though: loss_classifier starts near 3.24, and the natural logarithm of 21 classes is about 3.04. An untrained classifier guessing uniformly should sit right about there.
loss_box_reg is almost zero at step one. There is nothing suspicious about that. Regression loss is averaged over positive proposals only, and an untrained RPN produces very few. Watch this number rise before it falls during real training.
Eval mode returned exactly 100 boxes on pure noise. detections_per_img is 100 and score_thresh is 0.05. An untrained model produces near-random scores, plenty of which clear 0.05. A trained model on a real image returns far fewer. If your trained model still returns 100 boxes per image, your score threshold is too low for your use.
_pre_nms_top_n and _post_nms_top_n are dictionaries keyed by mode. Training keeps 2,000 proposals; testing keeps 1,000. More proposals means better recall and slower inference. It is the first knob to turn when small objects are being missed.
The cost R-CNN paid, measured
import time
import torch
import torch.nn as nn
from torchvision.ops import roi_align
torch.manual_seed(0)
torch.set_num_threads(4)
cnn = nn.Sequential(
nn.Conv2d(3, 32, 3, 2, 1), nn.ReLU(),
nn.Conv2d(32, 64, 3, 2, 1), nn.ReLU(),
nn.Conv2d(64, 128, 3, 2, 1), nn.ReLU(),
)
N = 300
crops = torch.rand(N, 3, 64, 64) # the R-CNN way: N separate crops
image = torch.rand(1, 3, 320, 320) # the Fast R-CNN way: one image
with torch.no_grad():
t0 = time.perf_counter()
for i in range(0, N, 25): # batched, which is generous to R-CNN
cnn(crops[i:i + 25])
t1 = time.perf_counter()
features = cnn(image) # one pass over the whole picture
boxes = torch.cat([torch.zeros(N, 1), torch.rand(N, 4) * 200
+ torch.tensor([0., 0., 60., 60.])], dim=1)
roi_align(features, boxes, output_size=7, spatial_scale=0.125, sampling_ratio=2)
t2 = time.perf_counter()
print(f"{N} crops, each through the CNN : {t1 - t0:.3f} s")
print(f"one pass + {N} roi_align crops : {t2 - t1:.3f} s")
print(f"ratio: {(t1 - t0) / (t2 - t1):.1f}x")This one is timing-dependent, so read it with care. The numbers below are one run on four CPU threads of the machine this lesson was written on. Repeated runs on that same machine gave ratios between roughly 2 and 4.5. Your absolute times will be different, and the ratio will move between runs:
300 crops, each through the CNN : 0.053 s one pass + 300 roi_align crops : 0.022 s ratio: 2.5x
The direction is what matters, and even the direction understates the real gap in two ways. Real R-CNN used around 2,000 crops, not 300, and a full VGG-16 rather than three convolutions. The Fast R-CNN paper reports its VGG-16 model testing 213 times faster than R-CNN.
The lineage, with what each paper actually changed
| Year | Model | The step it removed | Still slow because |
|---|---|---|---|
| 2014 | R-CNN | nothing; established the approach | a full CNN per crop, roughly 2,000 per image |
| 2014 | SPPnet | per-crop CNN passes | proposals still external; training not end to end |
| 2015 | Fast R-CNN | per-crop passes, multi-stage training | proposals still external and slow |
| 2015 | Faster R-CNN | external proposals | two stages, so slower than one-stage detectors |
| 2017 | Mask R-CNN | quantisation in RoI pooling; added masks | same two-stage cost |
Two details are worth keeping. R-CNN trained the CNN, an SVM per class and a box regressor as three separate stages. Fast R-CNN made it one multi-task loss. SPPnet had the shared-feature idea before Fast R-CNN. It could not backpropagate through its spatial pyramid pooling into the convolutional layers.
Common mistakes
Assuming num_classes excludes background. In torchvision it includes background as class 0. num_classes=21 means 20 real classes. Getting this wrong shifts every label by one.
Passing targets in eval mode, or omitting them in train mode. The model returns losses in train() and detections in eval(). Calling model(images) while in train() raises an error about missing targets.
Boxes in the wrong format. torchvision detection models want xyxy in absolute pixels, float32, with x2 > x1. See bounding box formats.
Empty images with no objects. Pass boxes as a (0, 4) float tensor and labels as a (0,) int64 tensor. Passing None or an empty list crashes deep inside the assigner.
Fine-tuning without replacing the head. Swap model.roi_heads.box_predictor with a new FastRCNNPredictor sized for your class count. Loading pretrained weights and changing num_classes in the constructor gives you a randomly initialised backbone as well.
Try it yourself
Set model.rpn._post_nms_top_n = {"training": 2000, "testing": 100} and re-run the eval call. Fewer proposals means faster inference and lower recall. Then find the smallest value that still returns every object in an image you care about. That number, not a paper's default, is your setting.
What to learn next
- YOLO versions compared — the one-stage answer to all of this.
- RoI pooling and RoI align — the operation the second stage depends on.
- Object detection — the introduction, if any of this moved too quickly.
Researcher — Mathematics and papers.
R-CNN, 2014
Girshick et al. (2014) combine three separately-trained pieces. Selective search proposes around 2,000 regions. A CNN, fine-tuned on warped 227x227 crops, embeds each one. A per-class linear SVM classifies the result, and a ridge-regression refiner adjusts the box.
The contribution was the demonstration that ImageNet-pretrained CNN features transfer to detection. The cost was structural. Every proposal is an independent forward pass. Features are cached to disk in the hundreds of gigabytes. The three training stages cannot inform each other.
SPPnet, 2014
He et al. (2014) compute the convolutional feature map once per image. Spatial pyramid pooling then runs on each proposal's projected region. A variable-sized region becomes a fixed-length vector.
This is the shared-computation insight. The Fast R-CNN paper identifies its limitation. Backpropagation through SPP into the layers below is inefficient when training samples come from different images. SPPnet therefore froze the convolutional layers.
Fast R-CNN, 2015
Girshick (2015) replaces SPP with single-level RoI pooling and trains everything with one multi-task loss:
$$ L(p, u, t^u, v) = L_{\text{cls}}(p, u) + \lambda \, [u \geq 1] \, L_{\text{loc}}(t^u, v) $$
Here $p$ is the predicted class distribution and $u$ the true class. $t^u$ holds the predicted box offsets for class $u$, and $v$ the target offsets. $[u \geq 1]$ is one for object classes and zero for background, so background proposals contribute no localisation loss. $L_{\text{loc}}$ is smooth $L_1$:
$$ \text{smooth}_{L_1}(x) = \begin{cases} 0.5 x^2 & |x| < 1 \ |x| - 0.5 & \text{otherwise} \end{cases} $$
Quadratic near zero for stable gradients, linear far away so outliers cannot dominate. This is the Huber loss with $\delta = 1$, and it is still the default box loss in torchvision.
The paper's sampling scheme matters as much as the architecture. It draws 128 RoIs from 2 images rather than from 128 images. The shared convolutional computation is then amortised. Correlation between RoIs from the same image was expected to hurt convergence and did not.
Reported: VGG-16 trained 9x faster and tested 213x faster than R-CNN, and 3x/10x faster than SPPnet, with higher mAP.
Faster R-CNN, 2015
Ren et al. (2015) replace selective search with a Region Proposal Network. A 3x3 convolution runs over the shared feature map. Two 1x1 sibling convolutions follow it. At each position they produce $2k$ objectness scores and $4k$ box offsets, for $k$ anchors.
The RPN loss is the same multi-task form. An anchor is positive at IoU above 0.7, or when it is the best anchor for a ground-truth box. It is negative below 0.3.
Training alternates between the RPN and the detector. Approximate joint training is the alternative. It ignores gradients flowing from the RoI pooling coordinates back into the proposal boxes. That approximation is what RoI align later removed; see RoI pooling and RoI align.
Proposal count drops from around 2,000 to 300 at equal or better recall. The proposals are learned rather than generic.
Why two stages remain competitive
The second stage performs cascade refinement: it re-samples features from a box that is already approximately correct. This has two effects a one-stage head does not get.
First, the classification problem is easier because the input is centred on the object. Second, box regression is applied twice, and iterated regression converges towards higher IoU.
Cai and Vasconcelos (2018), Cascade R-CNN, push this further. They use a sequence of heads trained at increasing IoU thresholds: 0.5, 0.6, 0.7. A single head trained at high IoU overfits, because few proposals qualify. A cascade feeds each stage the improved output of the previous one, so proposal quality rises to meet the threshold. This remains one of the reliable ways to raise AP at high IoU.
Sparse R-CNN (Sun et al., 2021) closes the loop with the DETR line. It keeps a small fixed set of learned proposal boxes and proposal features. Iterative heads refine them, with no dense anchors and no NMS.
Papers
- Girshick et al., Rich feature hierarchies for accurate object detection and semantic segmentation, CVPR 2014 — arxiv.org/abs/1311.2524
- He et al., Spatial Pyramid Pooling in Deep Convolutional Networks, ECCV 2014 — arxiv.org/abs/1406.4729
- Girshick, Fast R-CNN, ICCV 2015 — arxiv.org/abs/1504.08083
- Ren et al., Faster R-CNN, NeurIPS 2015 — arxiv.org/abs/1506.01497
- Cai and Vasconcelos, Cascade R-CNN: Delving into High Quality Object Detection, CVPR 2018 — arxiv.org/abs/1712.00726
- Sun et al., Sparse R-CNN, CVPR 2021 — arxiv.org/abs/2011.12450
What to learn next
- YOLO versions compared — the one-stage answer to all of this.
- RoI pooling and RoI align — the operation the second stage depends on.
- Object detection — the introduction, if any of this moved too quickly.