Video Understanding and Tracking

SlowFast networks

SlowFast runs two pathways over the same video, one seeing few frames in rich detail and one seeing many frames thinly, then wires them together.

On this page 7
  1. Why it exists
  2. How it works
  3. The connections between them
  4. Where you have already seen this
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

SlowFast watches the same video twice at once. One pass looks carefully at a few frames. The other glances at many frames.

Think about crossing a busy road in an Indian city. Your eyes do two jobs at the same time.

One job is slow and careful: reading the shop signs, noticing a pothole, finding the far pavement. The other job is fast and rough. You catch that a scooter is coming, without registering its colour.

You need both. Detail alone gets you hit. Motion alone leaves you lost.

SlowFast is that arrangement written as a network.

Why it exists

Earlier video networks picked one frame rate and lived with it.

Pick a low rate and you save compute and see rich detail, but fast movement disappears between frames. Pick a high rate and you catch fast movement, but you cannot afford full detail on so many frames.

The insight behind SlowFast is that these two jobs need different things.

What something is changes slowly. A person stays a person from frame to frame. Recognising them needs detail, and detail is expensive, so do it rarely.

How something moves changes quickly. Catching it needs many frames, but each frame can be crude, because you only care about the change.

So use two pathways with different settings, instead of forcing one pathway to compromise.

How it works

   video frames:  0  1  2  3  4  5  6  7  8  9  10 11 12 13 14 15 ...
                  │        │        │        │        │
   SLOW pathway:  X        X        X        X        X       few frames, MANY channels
                                                              (detail, appearance)

   FAST pathway:  x  x  x  x  x  x  x  x  x  x  x  x  x ...    many frames, FEW channels
                                                              (motion, timing)

                  └────────── lateral connections ──────────┘
                         Fast keeps telling Slow what moved

A channel is one of the many parallel feature maps a layer produces. More channels means more capacity to describe appearance.

The Slow pathway takes roughly one frame in sixteen, and keeps all its channels. The Fast pathway takes eight times as many frames, and keeps about one-eighth of the channels.

Because the Fast pathway is so thin, it is cheap despite reading far more frames. It accounts for around a fifth of the total work.

The connections between them

The two pathways do not run in isolation and meet at the end. Throughout the network, the Fast pathway feeds its findings sideways into the Slow one.

These are called lateral connections — links passing information across, rather than forward. The Slow pathway therefore gets motion information at every stage, not only at the finish.

Information flows one way. Fast informs Slow. The paper found adding the reverse direction gave no benefit.

Where you have already seen this

  • Sports analysis labelling a serve, a smash or a drop shot.
  • Security systems distinguishing walking from running from falling.
  • Video platforms auto-tagging clips for search.
  • Driver monitoring systems noticing a head nodding off.

What is honestly hard here

Two pathways means two of everything. Two clips to sample, two sets of shapes, two places for a bug to hide.

The lateral connections are the fiddly part. The Fast pathway has eight times as many frames. Its output cannot be attached to Slow directly. Something has to squash time by a factor of eight first. Getting that squash wrong is the standard mistake.

This design is genuinely more complicated than a single network. Read the shapes in the code below twice.

Remember this

  • Two pathways: Slow sees few frames with many channels, Fast sees many frames with few channels.
  • Fast is deliberately thin, so eight times the frames costs about a fifth of the compute.
  • Lateral connections feed motion from Fast into Slow throughout, not only at the end.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Run against torch 2.5.1 on CPU. This builds a miniature SlowFast to expose the sampling, the shapes and the lateral connection. The real model lives in PyTorchVideo and in Detectron2's SlowFast implementation.

The three numbers that define the design

tau is the Slow pathway's frame stride. alpha is how many times more frames Fast takes. beta is Fast's channel fraction. The paper uses 16, 8 and one-eighth.

slowfast_shapes.py
import torch
import torch.nn as nn

RAW_FRAMES = 64                       # what we grabbed from the video
TAU, ALPHA, BETA = 16, 8, 1 / 8       # the three numbers that define SlowFast

slow_idx = torch.arange(0, RAW_FRAMES, TAU)           # every 16th frame
fast_idx = torch.arange(0, RAW_FRAMES, TAU // ALPHA)  # every 2nd frame
print("Slow pathway frames:", slow_idx.tolist())
print("Fast pathway frames:", fast_idx.tolist())
print(f"Fast sees {len(fast_idx) // len(slow_idx)}x more frames "
      f"and carries {BETA:.3f} of the channels\n")


def pathway(c_out, temporal_kernel):
    """Two conv blocks. temporal_kernel=1 means this pathway cannot mix across time."""
    pad = temporal_kernel // 2
    return nn.Sequential(
        nn.Conv3d(3, c_out, (temporal_kernel, 7, 7),
                  stride=(1, 2, 2), padding=(pad, 3, 3), bias=False),
        nn.ReLU(),
        nn.Conv3d(c_out, c_out * 2, (temporal_kernel, 3, 3),
                  stride=(1, 2, 2), padding=(pad, 1, 1), bias=False),
    )


C = 64
slow = pathway(C, temporal_kernel=1)              # Slow stays purely spatial early on
fast = pathway(int(C * BETA), temporal_kernel=5)  # Fast uses a wide temporal kernel

with torch.no_grad():
    slow_out = slow(torch.zeros(1, 3, len(slow_idx), 112, 112))
    fast_out = fast(torch.zeros(1, 3, len(fast_idx), 112, 112))

print("Slow output (N,C,T,H,W):", tuple(slow_out.shape))
print("Fast output (N,C,T,H,W):", tuple(fast_out.shape))

# The lateral connection: squash Fast's time axis so it can be glued onto Slow.
lateral = nn.Conv3d(fast_out.shape[1], fast_out.shape[1] * 2,
                    kernel_size=(5, 1, 1), stride=(ALPHA, 1, 1),
                    padding=(2, 0, 0), bias=False)
with torch.no_grad():
    fused = torch.cat([slow_out, lateral(fast_out)], dim=1)
print("after lateral fusion    :", tuple(fused.shape))

print(f"\nSlow parameters: {sum(p.numel() for p in slow.parameters()):,}")
print(f"Fast parameters: {sum(p.numel() for p in fast.parameters()):,}")
Output
Slow pathway frames: [0, 16, 32, 48]
Fast pathway frames: [0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62]
Fast sees 8x more frames and carries 0.125 of the channels

Slow output (N,C,T,H,W): (1, 128, 4, 28, 28)
Fast output (N,C,T,H,W): (1, 16, 32, 28, 28)
after lateral fusion    : (1, 160, 4, 28, 28)

Slow parameters: 83,136
Fast parameters: 11,640

The lateral connection, line by line

This is the part worth slowing down on.

slow_out has 4 frames and 128 channels. fast_out has 32 frames and 16 channels. You cannot concatenate them: the time axes disagree, 4 against 32.

The lateral convolution fixes it with stride=(ALPHA, 1, 1). Striding by 8 along time turns 32 frames into 4, matching Slow exactly. The kernel_size=(5, 1, 1) means each output frame is built from a 5-frame temporal neighbourhood rather than one sampled frame, so nothing is discarded outright.

It also widens the channel count from 16 to 32 (beta * C * 2 in the paper's notation), giving Slow a meaningfully sized motion signal. The result is 128 + 32 = 160 channels at 4 frames.

The paper compared three ways of doing this squash — sampling every eighth frame, concatenating and reshaping, and this strided convolution. The strided convolution worked best, and it is the one every implementation uses.

Where the compute actually goes

The two-pathway design only makes sense if Fast is genuinely cheap. Here is the arithmetic, on a ResNet-50-shaped SlowFast at 224 pixels.

slowfast_cost.py
BETA, ALPHA = 1 / 8, 8


def conv(cin, cout, kt, ks, T, S):
    return cin * cout * kt * ks * ks * T * S * S       # multiply-adds


def stage(cin, width, blocks, kt, T, S):
    """ResNet bottleneck: 1x1 (with temporal kernel kt), 3x3 spatial, 1x1 expand."""
    cost, c = 0, cin
    for _ in range(blocks):
        cost += conv(c, width, kt, 1, T, S)
        cost += conv(width, width, 1, 3, T, S)
        cost += conv(width, width * 4, 1, 1, T, S)
        c = width * 4
    return cost, c


def pathway(w, T, stem_kt, res_kt):
    """w scales every channel width; res_kt gives the temporal kernel per stage."""
    total = conv(3, int(64 * w), stem_kt, 7, T, 112)   # stem
    c = int(64 * w)
    for width, blocks, S, kt in zip([64, 128, 256, 512], [3, 4, 6, 3],
                                    [56, 28, 14, 7], res_kt):
        s, c = stage(c, int(width * w), blocks, kt, T, S)
        total += s
    return total


slow = pathway(1.0, T=4, stem_kt=1, res_kt=[1, 1, 3, 3])
fast = pathway(BETA, T=4 * ALPHA, stem_kt=5, res_kt=[3, 3, 3, 3])

print(f"Slow pathway : {slow / 1e9:6.1f} GMAC")
print(f"Fast pathway : {fast / 1e9:6.1f} GMAC")
print(f"Fast share   : {100 * fast / (slow + fast):4.1f}% of the two-pathway total")
print(f"\nper-layer rule: beta^2 x alpha = {BETA ** 2 * ALPHA:.3f}, "
      "so at equal kernels Fast costs an eighth of Slow")
Output
Slow pathway :   17.3 GMAC
Fast pathway :    4.8 GMAC
Fast share   : 21.6% of the two-pathway total

per-layer rule: beta^2 x alpha = 0.125, so at equal kernels Fast costs an eighth of Slow

This is a cost model, not a benchmark. It counts convolution multiply-adds and ignores downsample shortcuts, normalisation and the lateral connections, so the absolute figures are approximate. The ratio is the point, and 21.6% lands on the paper's stated figure of roughly 20% for the Fast pathway.

The one-line reason sits at the bottom of the output. Cost scales with input channels times output channels, so scaling both by beta costs beta squared. Reading alpha times more frames multiplies by alpha. With beta one-eighth and alpha eight, the product is one-eighth. Eight times the frames for an eighth of the cost, per layer.

Common mistakes

Sampling both pathways independently. They must cover the same time span, from the same clip. Sample the Fast frames, then take every alpha-th of those for Slow. Two independent samplers mean the lateral connections are aligning unrelated moments.

Building the lateral connection with a plain reshape. Reshaping 32 frames into 4 groups of 8 and stacking them along channels is one of the paper's ablated variants, and it scored lower. Use the strided convolution.

Giving the Fast pathway temporal kernel 1. Its entire job is temporal. The Slow pathway is the one that starts with temporal kernel 1 and only gains temporal kernels in later stages.

Pooling time in the Fast pathway. The paper explicitly uses no temporal downsampling in Fast, so its high temporal resolution survives to the lateral connections. Adding a MaxPool3d with temporal stride there erases the reason the pathway exists.

Expecting a from-scratch SlowFast to train on a laptop. The published models were trained on Kinetics with many GPUs. Fine-tune a pretrained checkpoint from PyTorchVideo instead — see action recognition in practice.

Try it yourself

Change BETA to 1/4 and re-run the cost model. Predict the Fast share before running it. Then work out what beta value would make the two pathways cost the same, and check your answer against the code.

What to learn next

Researcher — Mathematics and papers.

The design in symbols

Let $\tau$ be the Slow pathway's temporal stride, $\alpha$ the frame-rate ratio, and $\beta$ the channel ratio. The paper's instantiation uses $\tau = 16$, $\alpha = 8$, $\beta = 1/8$.

From a raw clip of $\tau T$ frames:

  • Slow consumes $T$ frames with $C$ channels per stage.
  • Fast consumes $\alpha T$ frames with $\beta C$ channels per stage.

Per-layer multiply-adds scale as $C_{\text{in}} C_{\text{out}} \cdot T \cdot H W \cdot k$. The Fast-to-Slow ratio at matched kernels is therefore

$$ \frac{\beta C \cdot \beta C \cdot \alpha T}{C \cdot C \cdot T} = \alpha \beta^2 = 8 \cdot \tfrac{1}{64} = \tfrac{1}{8} $$

which is why the Fast pathway lands near 20% of total compute once its larger temporal kernels are included. The design is not "add a cheap branch" — it is a specific choice of $\alpha\beta^2 \ll 1$ that makes high temporal resolution affordable.

The asymmetry, and why it is the contribution

Both pathways are 3D ResNets. The novelty is that they are configured differently on purpose:

SlowFast
Frames$T$$\alpha T$
Channels$C$$\beta C$
Temporal kernels, early stages1 (non-temporal)3 (temporal)
Temporal downsamplingnonenone
Rolesemantics, appearancemotion, timing

Two ablations from the paper matter more than the headline result.

Slow with temporal kernels in early stages performs worse. Adding temporal convolution to a low-frame-rate pathway degraded accuracy. The interpretation is that when frames are 16 apart, objects have moved far enough that a temporal kernel spanning them is comparing unrelated content.

A Fast pathway with reduced spatial resolution, or with colour removed, still works well. The pathway is genuinely reading motion rather than appearance, which is the evidence for the two-role hypothesis rather than for it being a two-scale ensemble.

Lateral connections

Fusion happens after each of pool1, res2, res3, res4. The Fast feature map of shape ${\alpha T, S^2, \beta C}$ must be mapped to Slow's ${T, S^2, C}$. Three variants were compared:

  1. Time-to-channel: reshape ${\alpha T, S^2, \beta C}$ into ${T, S^2, \alpha\beta C}$.
  2. Time-strided sampling: take every $\alpha$-th frame.
  3. Time-strided convolution: a 3D convolution with kernel $5 \times 1^2$, output channels $2\beta C$, stride $\alpha$.

Variant 3 was best and is standard. Fusion is by concatenation, and the direction is Fast-to-Slow only; adding Slow-to-Fast or making it bidirectional gave no improvement.

Relationship to two-stream networks

SlowFast is often described as two-stream without optical flow, and the paper explicitly rejects that framing. Two-stream networks (Simonyan and Zisserman, 2014) use two different input modalities — RGB and pre-computed flow — with symmetric backbones and late fusion.

SlowFast uses one modality, asymmetric backbones, different temporal rates, and fusion throughout. It has no pre-computed flow, so it is trainable end to end and has no flow-extraction cost at inference. The biological framing in the paper, drawing on the parvocellular and magnocellular pathways in the primate visual system, is offered as motivation rather than evidence.

Results and successors

The paper reports state-of-the-art accuracy on Kinetics, Charades and AVA at publication, and won an ActivityNet challenge task. It became the standard backbone for spatio-temporal action detection on AVA for several years.

Successors and adjacent work:

  • X3D (Feichtenhofer, 2020), from the same author, achieves comparable accuracy at far lower cost by expanding a small 2D model along one axis at a time. Where compute is constrained, X3D generally dominates SlowFast.
  • MViT and Video Swin (2021–2022) replaced convolutional backbones with attention on the large benchmarks — see video transformers.
  • Multi-rate designs persist inside transformers. Several video transformers apply different token sampling rates to different branches, which is the SlowFast idea in a different substrate.

The durable contribution is not the specific architecture. It is the claim, supported by ablation, that spatial semantics and temporal dynamics deserve different capacity allocations. That claim has outlived the convolutional implementation of it.

References

What to learn next