Video Understanding and Tracking
3D convolutions for video
A 3D convolution slides a small cube through a stack of frames, so one filter can respond to movement instead of only to appearance.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A 3D convolution is the CNN filter you already know, made deep enough to cover several frames at once.
Picture a stack of rotis on a plate, one on top of the other. Now push a small cookie cutter straight down through the whole stack.
The cutter does not touch one roti. It touches a column through several of them at once. It comes back with a piece of every layer.
That is a 3D convolution. The stack is your frames. The cutter is the filter. It reads across time and space together.
Why it exists
A normal CNN filter is a flat square. It slides across one picture and finds edges, corners and textures. You met this in convolutional neural networks.
Show it a video and it treats each frame as a separate photo. Whatever it learns about frame five, it learns without knowing anything about frame four.
So it cannot tell walking from standing. It cannot tell a door opening from a door closing. Both look the same in any single frame.
The fix is small and obvious once you see it. Give the filter a third dimension, so it covers a few frames deep as well as wide.
How it works
a stack of frames a 3D filter (a small cube)
┌──────────┐ frame 3
┌──────────┐ frame 2 ┌───┐
┌──────────┐ frame 1 + │▓▓▓│ 3 wide, 3 tall, 3 frames deep
└──────────┘ frame 0 └───┘
│
▼
slide the cube left-right, up-down, AND forwards in time
│
▼
a map of WHERE and WHEN that motion pattern appearedA flat filter can learn "a vertical edge is here". A cube filter can learn something richer. "A vertical edge is here, and one frame ago it was two pixels left."
That second thing is a motion detector. Nobody had to design it. The network works out which motions matter, from labelled examples alone.
The price you pay
Cubes are expensive.
A three-by-three flat filter covers nine numbers. A three-by-three-by-three cube covers twenty-seven. That is three times the settings to learn. Three times the memory too, and three times the arithmetic, at every layer.
On a full network the difference is large enough to matter. It changes what fits on your graphics card and how long training takes.
Where you have already seen this
- Sports apps that label a shot as a cover drive or a pull shot.
- Fitness apps counting your squats from a phone camera.
- Gesture control on a smart TV.
- Content moderation systems flagging violent video.
What is honestly hard here
Video labels are scarce and video files are enormous.
A 3D network has many more settings to learn than a flat one, so it needs more examples. But labelling video is slow and costly, because a human has to watch it.
For years this stalled the whole field. The breakthrough was not a cleverer cube. It was borrowing: take a network already trained on millions of photographs, then stretch its flat filters into cubes. That trick is why modern video models work at all.
Remember this
- A 3D convolution slides a cube through frames, reading space and time together.
- It can detect motion patterns that a flat filter can never see.
- Cubes cost roughly three times as much as squares, and video labels are scarce.
What to learn next
- SlowFast networks — two pathways reading the same video at different speeds.
- Convolutional neural networks — the 2D filter these cubes are built from.
- Transfer learning in PyTorch — the mechanics behind weight inflation.
Developer — Code and libraries.
Setup
pip install torchRun against torch 2.5.1 on CPU, Python 3.10. nn.Conv3d has been stable since early PyTorch and nothing here needs a GPU.
Proving a 2D layer is blind to motion
We build two clips containing exactly the same pixels, arranged differently in time.
import torch
import torch.nn as nn
T, H, W = 8, 12, 12
def moving_bar():
x = torch.zeros(1, 1, T, H, W)
for t in range(T):
x[0, 0, t, :, t + 2] = 1.0 # a bright column that slides right each frame
return x
def still_bar():
x = torch.zeros(1, 1, T, H, W)
x[0, 0, :, :, 5] = 1.0 # the same column, never moving
return x
# A hand-built kernel: "this frame minus the previous frame", spatially identical.
motion = nn.Conv3d(1, 1, kernel_size=(2, 1, 1), bias=False)
with torch.no_grad():
motion.weight.copy_(torch.tensor([-1.0, 1.0]).view(1, 1, 2, 1, 1))
moving, still = moving_bar(), still_bar()
print("clip shape (N, C, T, H, W):", tuple(moving.shape))
print("Conv3d on the moving bar, total response:", motion(moving).abs().sum().item())
print("Conv3d on the still bar, total response:", motion(still).abs().sum().item())
# A 2D convolution applied frame by frame cannot see across time at all.
flat = nn.Conv2d(1, 1, kernel_size=1, bias=False)
with torch.no_grad():
flat.weight.fill_(1.0)
print("\nConv2d per frame, moving bar:", flat(moving[0].transpose(0, 1)).abs().sum().item())
print("Conv2d per frame, still bar:", flat(still[0].transpose(0, 1)).abs().sum().item())
print("identical: a 2D layer literally cannot tell these two clips apart")
c_in = c_out = 64
mid = (3 * 3 * 3 * c_in * c_out) // (3 * 3 * c_in + 3 * c_out) # R(2+1)D sizing rule
layers = {
"Conv2d 3x3": nn.Conv2d(c_in, c_out, 3, padding=1, bias=False),
"Conv3d 3x3x3": nn.Conv3d(c_in, c_out, 3, padding=1, bias=False),
f"R(2+1)D, mid={mid}": nn.Sequential(
nn.Conv3d(c_in, mid, (1, 3, 3), padding=(0, 1, 1), bias=False), # space only
nn.ReLU(),
nn.Conv3d(mid, c_out, (3, 1, 1), padding=(1, 0, 0), bias=False), # time only
),
}
print(f"\n{'layer':22s}{'parameters':>12s}")
for name, m in layers.items():
print(f"{name:22s}{sum(p.numel() for p in m.parameters()):12,d}")
net = nn.Sequential(
nn.Conv3d(3, 16, 3, padding=1), nn.ReLU(), nn.MaxPool3d((1, 2, 2)),
nn.Conv3d(16, 32, 3, padding=1), nn.ReLU(), nn.MaxPool3d(2),
nn.AdaptiveAvgPool3d(1), nn.Flatten(), nn.Linear(32, 5),
)
x = torch.zeros(1, 3, 8, 32, 32)
print(f"\n{'input':18s} -> {tuple(x.shape)}")
for layer in net:
x = layer(x)
print(f"{layer.__class__.__name__:18s} -> {tuple(x.shape)}")clip shape (N, C, T, H, W): (1, 1, 8, 12, 12) Conv3d on the moving bar, total response: 168.0 Conv3d on the still bar, total response: 0.0 Conv2d per frame, moving bar: 96.0 Conv2d per frame, still bar: 96.0 identical: a 2D layer literally cannot tell these two clips apart layer parameters Conv2d 3x3 36,864 Conv3d 3x3x3 110,592 R(2+1)D, mid=144 110,592 input -> (1, 3, 8, 32, 32) Conv3d -> (1, 16, 8, 32, 32) ReLU -> (1, 16, 8, 32, 32) MaxPool3d -> (1, 16, 8, 16, 16) Conv3d -> (1, 32, 8, 16, 16) ReLU -> (1, 32, 8, 16, 16) MaxPool3d -> (1, 32, 4, 8, 8) AdaptiveAvgPool3d -> (1, 32, 1, 1, 1) Flatten -> (1, 32) Linear -> (1, 5)
Reading the output
168 against 0. The 3D kernel is a difference between consecutive frames. On the moving bar it fires everywhere the bar arrived or left. On the still bar every difference is zero. One kernel, two clips, complete separation.
96 against 96. The 2D layer sums the same pixels in both clips and returns the same number. It is not that it performs poorly on motion. It has no mechanism to represent motion at all. This is the argument for Conv3d in one line of output.
110,592 against 36,864 — exactly three times. The ratio is the temporal kernel size. A 3x3x3 cube holds three times the weights of a 3x3 square. Memory and arithmetic scale the same way, which is the entire cost of the upgrade.
R(2+1)D matches the parameter count, and that is deliberate. The factorisation splits one cube into a spatial layer followed by a temporal layer. The middle width mid=144 comes from a formula (Tran et al., 2018) chosen so the two designs have equal capacity. Equal parameters, but an extra ReLU between them, so the factorised version is strictly more expressive at the same size. It also trains more easily.
Watch the shapes move. Spatial size goes 32, 16, 8. Temporal size stays 8, then halves to 4 only when MaxPool3d(2) pools all three axes. Note MaxPool3d((1, 2, 2)) — pooling space while leaving time alone. Early temporal pooling is one of the most common ways to accidentally destroy motion information, so most video architectures delay it.
The trick that made these networks work: inflation
Video labels are scarce. Photograph labels are not. I3D (Carreira and Zisserman, 2017) exploits that.
Take a 2D network trained on ImageNet. For every 3x3 kernel, copy it k times along a new temporal axis and divide by k. The resulting k x 3 x 3 kernel, applied to a clip of identical frames, produces exactly what the 2D kernel produced on one frame.
import torch
w2d = torch.randn(64, 3, 3, 3) # (out, in, H, W) from a 2D model
k = 3
w3d = w2d.unsqueeze(2).repeat(1, 1, k, 1, 1) / k # (out, in, T, H, W)
print(w2d.shape, "->", w3d.shape)
print("sum preserved:", torch.allclose(w3d.sum(dim=2), w2d))torch.Size([64, 3, 3, 3]) -> torch.Size([64, 3, 3, 3, 3]) sum preserved: True
The 3D network starts life as a good image network rather than at random. Every strong 3D convolutional result since 2017 depends on this or on a similar transfer step. Related reading: transfer learning in PyTorch.
Common mistakes
Getting the axis order wrong. PyTorch video tensors are (N, C, T, H, W). Frames read from OpenCV arrive as (T, H, W, C). The permutation is x.permute(3, 0, 1, 2).unsqueeze(0). A wrong permutation trains without error and learns nothing useful.
Pooling time too early. MaxPool3d(2) in the first block halves your temporal resolution before any layer has used it. Use MaxPool3d((1, 2, 2)) in early blocks.
Running out of memory and blaming the model. Activation memory scales with N x C x T x H x W. Halving the clip length saves as much as halving the batch. Try shorter clips before smaller batches, because batch size interacts with batch normalisation.
Using batch normalisation with a batch of one clip. Statistics over a single clip are noise. Use GroupNorm, or gradient accumulation with a larger effective batch — see gradient accumulation.
Training a 3D network from scratch on a small dataset. It will overfit and you will conclude 3D convolutions do not work. Start from inflated or pretrained weights.
Try it yourself
Change the hand-built kernel to [-1, 0, 1] over three frames, and set kernel_size=(3, 1, 1). Predict the response on the moving bar before running it. Then build a clip where the bar moves left and check what the sign of the response does.
What to learn next
- SlowFast networks — two pathways reading the same video at different speeds.
- Convolutional neural networks — the 2D filter these cubes are built from.
- Transfer learning in PyTorch — the mechanics behind weight inflation.
Researcher — Mathematics and papers.
The operation
For input $X$ with $C_{\text{in}}$ channels and kernel $K$ of size $k_t \times k_h \times k_w$:
$$ Y[c_o, t, i, j] = b[c_o] + \sum_{c=0}^{C_{\text{in}}-1} \sum_{p=0}^{k_t-1} \sum_{m=0}^{k_h-1} \sum_{n=0}^{k_w-1} X[c,\, t + p,\, i + m,\, j + n] \cdot K[c_o, c, p, m, n] $$
Where $c_o$ indexes the output channel, $t$ the temporal position, and $(i, j)$ the spatial position. Compared with 2D convolution, the only change is the extra sum over $p$.
Parameters: $C_{\text{out}} C_{\text{in}} k_t k_h k_w + C_{\text{out}}$, a factor $k_t$ above the 2D case.
FLOPs: $\approx 2 \cdot T_{\text{out}} H_{\text{out}} W_{\text{out}} \cdot C_{\text{out}} C_{\text{in}} k_t k_h k_w$, a factor $k_t \cdot T_{\text{out}}$ above 2D convolution on a single frame.
Activation memory grows linearly in $T$, which is why clip length is usually the first hyperparameter sacrificed.
Factorisations
Full 3D convolution is rarely used unfactorised in modern architectures. Three decompositions dominate:
R(2+1)D (Tran et al., CVPR 2018) replaces $k_t \times k_h \times k_w$ with $1 \times k_h \times k_w$ followed by $k_t \times 1 \times 1$, with a hidden width
$$ M_i = \left\lfloor \frac{k_t k_h k_w \, C_{i-1} C_i}{k_h k_w \, C_{i-1} + k_t \, C_i} \right\rfloor $$
chosen so parameter count matches the unfactorised layer. The gain is the additional non-linearity between the two, plus a measurably lower training error — the paper's central empirical claim is that R(2+1)D is easier to optimise, not smaller.
P3D (Qiu et al., ICCV 2017) explores three wiring variants of the same split (serial, parallel, and a residual hybrid) inside a ResNet bottleneck.
S3D (Xie et al., ECCV 2018) applies the separation selectively: 2D convolutions in early layers, separable 3D in later layers. It reports higher accuracy than I3D at lower cost, and gives the useful finding that temporal modelling is more valuable in late layers than early ones. Early layers are learning edges, which do not need a temporal extent.
Architectural lineage
| Year | Model | Contribution |
|---|---|---|
| 2015 | C3D (Tran et al.) | First widely used 3D CNN; fixed 3x3x3 kernels throughout |
| 2017 | I3D (Carreira and Zisserman) | Inflation from ImageNet weights; Kinetics dataset |
| 2017 | P3D (Qiu et al.) | Pseudo-3D residual variants |
| 2018 | R(2+1)D (Tran et al.) | Matched-capacity spatial/temporal factorisation |
| 2018 | S3D (Xie et al.) | Top-heavy separable design; early layers stay 2D |
| 2018 | Non-local (Wang et al.) | Self-attention block inserted into 3D CNNs |
| 2019 | SlowFast (Feichtenhofer et al.) | Two pathways at different frame rates |
| 2020 | X3D (Feichtenhofer) | Progressive expansion along 6 axes; strong accuracy per FLOP |
X3D is the one most worth reading if compute is your constraint. It starts from a tiny 2D image classifier and expands one axis at a time — temporal length, frame rate, spatial resolution, width, bottleneck width, depth — greedily selecting the expansion with the best accuracy-per-FLOP at each step. The result reaches I3D-level accuracy at roughly a fifth of the multiply-adds.
Inflation, formally
Given 2D weights $W^{2D} \in \mathbb{R}^{C_o \times C_i \times k_h \times k_w}$, the inflated kernel is
$$ W^{3D}[c_o, c, p, m, n] = \frac{1}{k_t} W^{2D}[c_o, c, m, n] \quad \forall p $$
The bootstrapping argument: feed a clip made of $k_t$ copies of one image. The inflated 3D layer's response equals the 2D layer's response on that image, so the whole network's behaviour is preserved. The division by $k_t$ is what makes this exact, and omitting it scales activations by $k_t$ at every layer.
I3D reported that inflated Kinetics-pretrained models transferred to UCF-101 and HMDB-51 far better than any from-scratch 3D model, which is what established Kinetics as the standard video pretraining corpus.
Where 3D convolutions stand now
Transformers overtook 3D CNNs on the large benchmarks from 2021 onward — see video transformers. The honest comparison is narrower than headline numbers suggest:
- 3D CNNs carry locality and translation-equivariance priors, so they remain stronger in low-data regimes.
- Their cost is predictable and their memory profile is friendlier to edge deployment.
- X3D-class models still hold competitive accuracy-per-FLOP, which is the metric that matters on device.
Convolutional and attention-based video models have also converged in practice. Video Swin restricts attention to local 3D windows, which is structurally close to a large separable convolution; MViT pools keys and values, which is a stride. The design space is one space, not two camps.
References
- Tran et al., Learning Spatiotemporal Features with 3D Convolutional Networks (C3D), ICCV 2015 — arxiv.org/abs/1412.0767
- Carreira and Zisserman, Quo Vadis, Action Recognition? (I3D), CVPR 2017 — arxiv.org/abs/1705.07750
- Tran et al., A Closer Look at Spatiotemporal Convolutions (R(2+1)D), CVPR 2018 — arxiv.org/abs/1711.11248
- Xie et al., Rethinking Spatiotemporal Feature Learning (S3D), ECCV 2018 — arxiv.org/abs/1712.04851
- Feichtenhofer, X3D: Expanding Architectures for Efficient Video Recognition, CVPR 2020 — arxiv.org/abs/2004.04730
What to learn next
- SlowFast networks — two pathways reading the same video at different speeds.
- Convolutional neural networks — the 2D filter these cubes are built from.
- Transfer learning in PyTorch — the mechanics behind weight inflation.