Video Understanding and Tracking
SlowFast networks
SlowFast runs two pathways over the same video, one seeing few frames in rich detail and one seeing many frames thinly, then wires them together.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
SlowFast watches the same video twice at once. One pass looks carefully at a few frames. The other glances at many frames.
Think about crossing a busy road in an Indian city. Your eyes do two jobs at the same time.
One job is slow and careful: reading the shop signs, noticing a pothole, finding the far pavement. The other job is fast and rough. You catch that a scooter is coming, without registering its colour.
You need both. Detail alone gets you hit. Motion alone leaves you lost.
SlowFast is that arrangement written as a network.
Why it exists
Earlier video networks picked one frame rate and lived with it.
Pick a low rate and you save compute and see rich detail, but fast movement disappears between frames. Pick a high rate and you catch fast movement, but you cannot afford full detail on so many frames.
The insight behind SlowFast is that these two jobs need different things.
What something is changes slowly. A person stays a person from frame to frame. Recognising them needs detail, and detail is expensive, so do it rarely.
How something moves changes quickly. Catching it needs many frames, but each frame can be crude, because you only care about the change.
So use two pathways with different settings, instead of forcing one pathway to compromise.
How it works
video frames: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 ...
│ │ │ │ │
SLOW pathway: X X X X X few frames, MANY channels
(detail, appearance)
FAST pathway: x x x x x x x x x x x x x ... many frames, FEW channels
(motion, timing)
└────────── lateral connections ──────────┘
Fast keeps telling Slow what movedA channel is one of the many parallel feature maps a layer produces. More channels means more capacity to describe appearance.
The Slow pathway takes roughly one frame in sixteen, and keeps all its channels. The Fast pathway takes eight times as many frames, and keeps about one-eighth of the channels.
Because the Fast pathway is so thin, it is cheap despite reading far more frames. It accounts for around a fifth of the total work.
The connections between them
The two pathways do not run in isolation and meet at the end. Throughout the network, the Fast pathway feeds its findings sideways into the Slow one.
These are called lateral connections — links passing information across, rather than forward. The Slow pathway therefore gets motion information at every stage, not only at the finish.
Information flows one way. Fast informs Slow. The paper found adding the reverse direction gave no benefit.
Where you have already seen this
- Sports analysis labelling a serve, a smash or a drop shot.
- Security systems distinguishing walking from running from falling.
- Video platforms auto-tagging clips for search.
- Driver monitoring systems noticing a head nodding off.
What is honestly hard here
Two pathways means two of everything. Two clips to sample, two sets of shapes, two places for a bug to hide.
The lateral connections are the fiddly part. The Fast pathway has eight times as many frames. Its output cannot be attached to Slow directly. Something has to squash time by a factor of eight first. Getting that squash wrong is the standard mistake.
This design is genuinely more complicated than a single network. Read the shapes in the code below twice.
Remember this
- Two pathways: Slow sees few frames with many channels, Fast sees many frames with few channels.
- Fast is deliberately thin, so eight times the frames costs about a fifth of the compute.
- Lateral connections feed motion from Fast into Slow throughout, not only at the end.
What to learn next
- Video transformers — attention taking over from convolution on video.
- Action recognition in practice — running one of these models end to end.
- 3D convolutions for video — the building block both pathways are made of.
Developer — Code and libraries.
Setup
pip install torchRun against torch 2.5.1 on CPU. This builds a miniature SlowFast to expose the sampling, the shapes and the lateral connection. The real model lives in PyTorchVideo and in Detectron2's SlowFast implementation.
The three numbers that define the design
tau is the Slow pathway's frame stride. alpha is how many times more frames Fast takes. beta is Fast's channel fraction. The paper uses 16, 8 and one-eighth.
import torch
import torch.nn as nn
RAW_FRAMES = 64 # what we grabbed from the video
TAU, ALPHA, BETA = 16, 8, 1 / 8 # the three numbers that define SlowFast
slow_idx = torch.arange(0, RAW_FRAMES, TAU) # every 16th frame
fast_idx = torch.arange(0, RAW_FRAMES, TAU // ALPHA) # every 2nd frame
print("Slow pathway frames:", slow_idx.tolist())
print("Fast pathway frames:", fast_idx.tolist())
print(f"Fast sees {len(fast_idx) // len(slow_idx)}x more frames "
f"and carries {BETA:.3f} of the channels\n")
def pathway(c_out, temporal_kernel):
"""Two conv blocks. temporal_kernel=1 means this pathway cannot mix across time."""
pad = temporal_kernel // 2
return nn.Sequential(
nn.Conv3d(3, c_out, (temporal_kernel, 7, 7),
stride=(1, 2, 2), padding=(pad, 3, 3), bias=False),
nn.ReLU(),
nn.Conv3d(c_out, c_out * 2, (temporal_kernel, 3, 3),
stride=(1, 2, 2), padding=(pad, 1, 1), bias=False),
)
C = 64
slow = pathway(C, temporal_kernel=1) # Slow stays purely spatial early on
fast = pathway(int(C * BETA), temporal_kernel=5) # Fast uses a wide temporal kernel
with torch.no_grad():
slow_out = slow(torch.zeros(1, 3, len(slow_idx), 112, 112))
fast_out = fast(torch.zeros(1, 3, len(fast_idx), 112, 112))
print("Slow output (N,C,T,H,W):", tuple(slow_out.shape))
print("Fast output (N,C,T,H,W):", tuple(fast_out.shape))
# The lateral connection: squash Fast's time axis so it can be glued onto Slow.
lateral = nn.Conv3d(fast_out.shape[1], fast_out.shape[1] * 2,
kernel_size=(5, 1, 1), stride=(ALPHA, 1, 1),
padding=(2, 0, 0), bias=False)
with torch.no_grad():
fused = torch.cat([slow_out, lateral(fast_out)], dim=1)
print("after lateral fusion :", tuple(fused.shape))
print(f"\nSlow parameters: {sum(p.numel() for p in slow.parameters()):,}")
print(f"Fast parameters: {sum(p.numel() for p in fast.parameters()):,}")Slow pathway frames: [0, 16, 32, 48] Fast pathway frames: [0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62] Fast sees 8x more frames and carries 0.125 of the channels Slow output (N,C,T,H,W): (1, 128, 4, 28, 28) Fast output (N,C,T,H,W): (1, 16, 32, 28, 28) after lateral fusion : (1, 160, 4, 28, 28) Slow parameters: 83,136 Fast parameters: 11,640
The lateral connection, line by line
This is the part worth slowing down on.
slow_out has 4 frames and 128 channels. fast_out has 32 frames and 16 channels. You cannot concatenate them: the time axes disagree, 4 against 32.
The lateral convolution fixes it with stride=(ALPHA, 1, 1). Striding by 8 along time turns 32 frames into 4, matching Slow exactly. The kernel_size=(5, 1, 1) means each output frame is built from a 5-frame temporal neighbourhood rather than one sampled frame, so nothing is discarded outright.
It also widens the channel count from 16 to 32 (beta * C * 2 in the paper's notation), giving Slow a meaningfully sized motion signal. The result is 128 + 32 = 160 channels at 4 frames.
The paper compared three ways of doing this squash — sampling every eighth frame, concatenating and reshaping, and this strided convolution. The strided convolution worked best, and it is the one every implementation uses.
Where the compute actually goes
The two-pathway design only makes sense if Fast is genuinely cheap. Here is the arithmetic, on a ResNet-50-shaped SlowFast at 224 pixels.
BETA, ALPHA = 1 / 8, 8
def conv(cin, cout, kt, ks, T, S):
return cin * cout * kt * ks * ks * T * S * S # multiply-adds
def stage(cin, width, blocks, kt, T, S):
"""ResNet bottleneck: 1x1 (with temporal kernel kt), 3x3 spatial, 1x1 expand."""
cost, c = 0, cin
for _ in range(blocks):
cost += conv(c, width, kt, 1, T, S)
cost += conv(width, width, 1, 3, T, S)
cost += conv(width, width * 4, 1, 1, T, S)
c = width * 4
return cost, c
def pathway(w, T, stem_kt, res_kt):
"""w scales every channel width; res_kt gives the temporal kernel per stage."""
total = conv(3, int(64 * w), stem_kt, 7, T, 112) # stem
c = int(64 * w)
for width, blocks, S, kt in zip([64, 128, 256, 512], [3, 4, 6, 3],
[56, 28, 14, 7], res_kt):
s, c = stage(c, int(width * w), blocks, kt, T, S)
total += s
return total
slow = pathway(1.0, T=4, stem_kt=1, res_kt=[1, 1, 3, 3])
fast = pathway(BETA, T=4 * ALPHA, stem_kt=5, res_kt=[3, 3, 3, 3])
print(f"Slow pathway : {slow / 1e9:6.1f} GMAC")
print(f"Fast pathway : {fast / 1e9:6.1f} GMAC")
print(f"Fast share : {100 * fast / (slow + fast):4.1f}% of the two-pathway total")
print(f"\nper-layer rule: beta^2 x alpha = {BETA ** 2 * ALPHA:.3f}, "
"so at equal kernels Fast costs an eighth of Slow")Slow pathway : 17.3 GMAC Fast pathway : 4.8 GMAC Fast share : 21.6% of the two-pathway total per-layer rule: beta^2 x alpha = 0.125, so at equal kernels Fast costs an eighth of Slow
This is a cost model, not a benchmark. It counts convolution multiply-adds and ignores downsample shortcuts, normalisation and the lateral connections, so the absolute figures are approximate. The ratio is the point, and 21.6% lands on the paper's stated figure of roughly 20% for the Fast pathway.
The one-line reason sits at the bottom of the output. Cost scales with input channels times output channels, so scaling both by beta costs beta squared. Reading alpha times more frames multiplies by alpha. With beta one-eighth and alpha eight, the product is one-eighth. Eight times the frames for an eighth of the cost, per layer.
Common mistakes
Sampling both pathways independently. They must cover the same time span, from the same clip. Sample the Fast frames, then take every alpha-th of those for Slow. Two independent samplers mean the lateral connections are aligning unrelated moments.
Building the lateral connection with a plain reshape. Reshaping 32 frames into 4 groups of 8 and stacking them along channels is one of the paper's ablated variants, and it scored lower. Use the strided convolution.
Giving the Fast pathway temporal kernel 1. Its entire job is temporal. The Slow pathway is the one that starts with temporal kernel 1 and only gains temporal kernels in later stages.
Pooling time in the Fast pathway. The paper explicitly uses no temporal downsampling in Fast, so its high temporal resolution survives to the lateral connections. Adding a MaxPool3d with temporal stride there erases the reason the pathway exists.
Expecting a from-scratch SlowFast to train on a laptop. The published models were trained on Kinetics with many GPUs. Fine-tune a pretrained checkpoint from PyTorchVideo instead — see action recognition in practice.
Try it yourself
Change BETA to 1/4 and re-run the cost model. Predict the Fast share before running it. Then work out what beta value would make the two pathways cost the same, and check your answer against the code.
What to learn next
- Video transformers — attention taking over from convolution on video.
- Action recognition in practice — running one of these models end to end.
- 3D convolutions for video — the building block both pathways are made of.
Researcher — Mathematics and papers.
The design in symbols
Let $\tau$ be the Slow pathway's temporal stride, $\alpha$ the frame-rate ratio, and $\beta$ the channel ratio. The paper's instantiation uses $\tau = 16$, $\alpha = 8$, $\beta = 1/8$.
From a raw clip of $\tau T$ frames:
- Slow consumes $T$ frames with $C$ channels per stage.
- Fast consumes $\alpha T$ frames with $\beta C$ channels per stage.
Per-layer multiply-adds scale as $C_{\text{in}} C_{\text{out}} \cdot T \cdot H W \cdot k$. The Fast-to-Slow ratio at matched kernels is therefore
$$ \frac{\beta C \cdot \beta C \cdot \alpha T}{C \cdot C \cdot T} = \alpha \beta^2 = 8 \cdot \tfrac{1}{64} = \tfrac{1}{8} $$
which is why the Fast pathway lands near 20% of total compute once its larger temporal kernels are included. The design is not "add a cheap branch" — it is a specific choice of $\alpha\beta^2 \ll 1$ that makes high temporal resolution affordable.
The asymmetry, and why it is the contribution
Both pathways are 3D ResNets. The novelty is that they are configured differently on purpose:
| Slow | Fast | |
|---|---|---|
| Frames | $T$ | $\alpha T$ |
| Channels | $C$ | $\beta C$ |
| Temporal kernels, early stages | 1 (non-temporal) | 3 (temporal) |
| Temporal downsampling | none | none |
| Role | semantics, appearance | motion, timing |
Two ablations from the paper matter more than the headline result.
Slow with temporal kernels in early stages performs worse. Adding temporal convolution to a low-frame-rate pathway degraded accuracy. The interpretation is that when frames are 16 apart, objects have moved far enough that a temporal kernel spanning them is comparing unrelated content.
A Fast pathway with reduced spatial resolution, or with colour removed, still works well. The pathway is genuinely reading motion rather than appearance, which is the evidence for the two-role hypothesis rather than for it being a two-scale ensemble.
Lateral connections
Fusion happens after each of pool1, res2, res3, res4. The Fast feature map of shape ${\alpha T, S^2, \beta C}$ must be mapped to Slow's ${T, S^2, C}$. Three variants were compared:
- Time-to-channel: reshape ${\alpha T, S^2, \beta C}$ into ${T, S^2, \alpha\beta C}$.
- Time-strided sampling: take every $\alpha$-th frame.
- Time-strided convolution: a 3D convolution with kernel $5 \times 1^2$, output channels $2\beta C$, stride $\alpha$.
Variant 3 was best and is standard. Fusion is by concatenation, and the direction is Fast-to-Slow only; adding Slow-to-Fast or making it bidirectional gave no improvement.
Relationship to two-stream networks
SlowFast is often described as two-stream without optical flow, and the paper explicitly rejects that framing. Two-stream networks (Simonyan and Zisserman, 2014) use two different input modalities — RGB and pre-computed flow — with symmetric backbones and late fusion.
SlowFast uses one modality, asymmetric backbones, different temporal rates, and fusion throughout. It has no pre-computed flow, so it is trainable end to end and has no flow-extraction cost at inference. The biological framing in the paper, drawing on the parvocellular and magnocellular pathways in the primate visual system, is offered as motivation rather than evidence.
Results and successors
The paper reports state-of-the-art accuracy on Kinetics, Charades and AVA at publication, and won an ActivityNet challenge task. It became the standard backbone for spatio-temporal action detection on AVA for several years.
Successors and adjacent work:
- X3D (Feichtenhofer, 2020), from the same author, achieves comparable accuracy at far lower cost by expanding a small 2D model along one axis at a time. Where compute is constrained, X3D generally dominates SlowFast.
- MViT and Video Swin (2021–2022) replaced convolutional backbones with attention on the large benchmarks — see video transformers.
- Multi-rate designs persist inside transformers. Several video transformers apply different token sampling rates to different branches, which is the SlowFast idea in a different substrate.
The durable contribution is not the specific architecture. It is the claim, supported by ablation, that spatial semantics and temporal dynamics deserve different capacity allocations. That claim has outlived the convolutional implementation of it.
References
- Feichtenhofer, Fan, Malik and He, SlowFast Networks for Video Recognition, ICCV 2019 — arxiv.org/abs/1812.03982
- Simonyan and Zisserman, Two-Stream Convolutional Networks, NeurIPS 2014 — arxiv.org/abs/1406.2199
- Feichtenhofer, X3D, CVPR 2020 — arxiv.org/abs/2004.04730
- Fan et al., PyTorchVideo: A Deep Learning Library for Video Understanding, ACM MM 2021 — pytorchvideo.org
What to learn next
- Video transformers — attention taking over from convolution on video.
- Action recognition in practice — running one of these models end to end.
- 3D convolutions for video — the building block both pathways are made of.