Video Understanding and Tracking
Video transformers
Cutting a video into patches and letting attention connect them works, but attending over every pair of patches is unaffordable, so every video transformer is a way of splitting that attention up.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A video transformer chops a video into small square patches. Every patch then decides which other patches matter to it.
Think about watching a cricket highlight with a friend who missed the match. You do not describe every blade of grass. You link a few things together: the bowler's run-up, the bat's swing, the ball's path, the fielder's dive.
You connected pieces that were far apart in space and in time, and you ignored everything else. Nobody told you which pieces mattered. You worked it out.
That linking-what-matters is called attention, and you met it in attention. A video transformer does it over video patches.
Why it exists
The cube-shaped filters from 3D convolutions have one hard limit. A cube only sees its immediate neighbourhood.
Connecting the bowler at the start of a clip to the fielder at the end is slow. Information must crawl through layer after layer of cubes. Long-range links are possible but slow to build, and easily lost on the way.
Attention has no such limit. Any patch can look at any other patch, in one step, however far away in time.
Transformers had already taken over language for exactly this reason. Vision transformers did the same to still images. Video was next.
How it works
frame 1 frame 2 frame 3
┌─┬─┬─┐ ┌─┬─┬─┐ ┌─┬─┬─┐
├─┼─┼─┤ ├─┼─┼─┤ ├─┼─┼─┤ cut every frame into patches
└─┴─┴─┘ └─┴─┴─┘ └─┴─┴─┘
│ │ │
└──────────────┴──────────────┘
▼
every patch becomes one "token"
▼
each token asks every other token:
"do you matter to me?"
▼
a prediction for the clipA token is one patch turned into a list of numbers. A single frame at a common size becomes 196 patches. Sixteen frames become over three thousand tokens.
The problem that shapes everything
Letting every token look at every other token means counting pairs.
Two hundred tokens gives forty thousand pairs, which is fine. Three thousand tokens gives nearly ten million pairs, which is not. Double the frames and the pair count goes up four times, not two.
Video ran into this wall immediately. Every video transformer you will read about is a different answer to one question: which pairs can we skip?
The most popular answer is to split attention in two. First let each patch look at other patches in the same frame. Then let each patch look at the same position in other frames. Two cheap steps instead of one impossible one.
Where you have already seen this
- Video search that finds a clip from a text description.
- Automatic captions describing what happens in a video.
- Recommendation systems reading video content rather than only titles.
- Editing tools that find every shot containing a particular action.
What is honestly hard here
Transformers do not come with the built-in assumptions that convolutions have.
A convolution knows that nearby pixels belong together, because it is built that way. A transformer starts knowing nothing and has to learn even that from examples.
That makes transformers more flexible and much hungrier for data. On small video datasets a well-tuned convolutional model often still wins. The transformer results you read about usually depend on enormous pretraining, and that part is rarely emphasised.
Remember this
- A video transformer turns patches into tokens and lets attention link any two of them.
- Attention cost grows with the square of the token count, and video has a great many tokens.
- Every design is a scheme for skipping pairs, most often by separating space from time.
What to learn next
- Action recognition in practice — putting one of these models to work.
- Vision transformers — the image model every video transformer extends.
- Attention — the mechanism itself, at its simplest.
Developer — Code and libraries.
Setup
pip install torchRun against torch 2.5.1 on CPU. The Hugging Face section at the end needs pip install transformers and downloads a checkpoint.
The arithmetic that forces every design decision
import torch
import torch.nn as nn
PATCH, IMG = 16, 224
S = (IMG // PATCH) ** 2 # spatial tokens per frame
print(f"{S} patches per frame at {PATCH}x{PATCH} on a {IMG}x{IMG} frame\n")
print(f"{'frames':>7}{'tokens':>9}{'joint pairs':>14}{'divided pairs':>15}{'ratio':>8}")
for T in (1, 8, 16, 32, 64):
joint = (T * S) ** 2 # every token attends to every other token
divided = T * S * S + S * T * T # space inside a frame, then time per position
print(f"{T:7d}{T*S:9d}{joint:14,d}{divided:15,d}{joint/divided:8.1f}x")
class DividedSpaceTime(nn.Module):
"""One TimeSformer block: attend over time first, then over space."""
def __init__(self, dim, heads):
super().__init__()
self.time_attn = nn.MultiheadAttention(dim, heads, batch_first=True)
self.space_attn = nn.MultiheadAttention(dim, heads, batch_first=True)
self.norm1, self.norm2 = nn.LayerNorm(dim), nn.LayerNorm(dim)
def forward(self, x, T, S):
B, _, D = x.shape # x is (batch, T*S, dim)
t = x.reshape(B, T, S, D).permute(0, 2, 1, 3).reshape(B * S, T, D)
t = self.time_attn(t, t, t, need_weights=False)[0] # one sequence per position
x = x + t.reshape(B, S, T, D).permute(0, 2, 1, 3).reshape(B, T * S, D)
x = self.norm1(x)
s = x.reshape(B * T, S, D) # one sequence per frame
s = self.space_attn(s, s, s, need_weights=False)[0]
x = x + s.reshape(B, T * S, D)
return self.norm2(x)
torch.manual_seed(0)
T_, S_, D_ = 8, 196, 192
block = DividedSpaceTime(D_, heads=3)
x = torch.zeros(1, T_ * S_, D_)
print("\ninput tokens :", tuple(x.shape))
with torch.no_grad():
y = block(x, T_, S_)
print("output tokens:", tuple(y.shape))
print("block parameters:", f"{sum(p.numel() for p in block.parameters()):,}")
print("\ntime attention matrix per patch position:", (T_, T_))
print("space attention matrix per frame :", (S_, S_))
# VideoMAE tubelets: a patch spans 2 frames, so the token count halves.
FRAMES, TUBELET = 16, 2
tokens = (FRAMES // TUBELET) * (IMG // PATCH) ** 2
print(f"\nVideoMAE tubelet tokens for {FRAMES} frames: {tokens}")196 patches per frame at 16x16 on a 224x224 frame
frames tokens joint pairs divided pairs ratio
1 196 38,416 38,612 1.0x
8 1568 2,458,624 319,872 7.7x
16 3136 9,834,496 664,832 14.8x
32 6272 39,337,984 1,430,016 27.5x
64 12544 157,351,936 3,261,440 48.2x
input tokens : (1, 1568, 192)
output tokens: (1, 1568, 192)
block parameters: 297,216
time attention matrix per patch position: (8, 8)
space attention matrix per frame : (196, 196)
VideoMAE tubelet tokens for 16 frames: 1568Reading the table
At one frame, splitting attention gains nothing. The ratio is 1.0x, and divided attention is fractionally worse because it adds a degenerate time step over a sequence of length one. Every technique on this page is a video technique, and applying it to images is pure overhead.
At 8 frames the split already saves nearly 8x. At 64 frames it saves 48x. The gap widens because joint attention grows with the square of T while divided attention grows linearly in T for the spatial part.
The two attention matrices have wildly different shapes. Space is 196 by 196. Time is 8 by 8. Concentrating capacity where the tokens actually are is the whole trick, and it is also why long-video transformers still struggle: push T to 1000 and the time matrix becomes the expensive one.
1568 appears twice, for two different reasons. Eight frames at 196 patches gives 1568 tokens. Separately, VideoMAE with 16 frames and a 2-frame tubelet gives 1568 tokens as well. A tubelet is a patch extended through time, so it halves the token count at no accuracy cost. Hugging Face's VideoMAEModel documentation shows a last_hidden_state of [1, 1568, 768] for a 16-frame clip, which is this arithmetic confirmed against the official docs.
Running a real pretrained video transformer
The code above builds the mechanism. To classify an actual video, use a pretrained checkpoint. This one downloads roughly 350 MB and runs on CPU in a few seconds per clip.
import torch
from transformers import VideoMAEVideoProcessor, VideoMAEForVideoClassification
name = "MCG-NJU/videomae-base-finetuned-kinetics"
processor = VideoMAEVideoProcessor.from_pretrained(name)
model = VideoMAEForVideoClassification.from_pretrained(name).eval()
# 16 frames, each 224x224 RGB. Replace with real decoded frames.
video = list(torch.randint(0, 256, (16, 3, 224, 224), dtype=torch.uint8).numpy())
inputs = processor(video, return_tensors="pt")
with torch.no_grad():
logits = model(**inputs).logits
print(model.config.id2label[logits.argmax(-1).item()])No output block here, deliberately. The input is random noise, so the predicted Kinetics class is arbitrary and would change with the seed. Printing an invented label would teach you to expect something that will not happen.
Two API notes. VideoMAEVideoProcessor is the current class in Transformers v5; VideoMAEImageProcessor still exists and older tutorials use it. And the model is fixed at 16 frames with a 2-frame tubelet, set by config.num_frames and config.tubelet_size — feeding a different frame count raises a shape error.
Common mistakes
Reshaping tokens with the wrong axis order. The block above moves between (B, T*S, D) and (B*S, T, D). Get the permute wrong and it still runs, produces the right shapes, and mixes unrelated patches. Test it by feeding a clip of identical frames: the time attention output should then be the same for every frame.
Using .view where .reshape is needed. After a permute, the tensor is not contiguous and .view raises view size is not compatible with input tensor's size and stride. .reshape handles it. See reshape, view and contiguity.
Forgetting temporal position information. Patches carry no inherent notion of which frame they came from. Without a temporal position embedding, the model is order-blind and will score near chance on any dataset requiring temporal order.
Training a video transformer from scratch on a small dataset. It will lose to a pretrained 3D CNN. Start from a checkpoint. This is not a minor optimisation; it is the difference between working and not working.
Assuming attention weights explain the prediction. Attention maps are attractive and are not faithful explanations. Treat them as a debugging aid, not as evidence — explainability covers why.
Try it yourself
Add a joint attention path to the block: flatten to (B, T*S, D) and run one MultiheadAttention over all 1568 tokens. Time both on your machine, then try 32 frames and watch the gap widen exactly as the table predicts.
What to learn next
- Action recognition in practice — putting one of these models to work.
- Vision transformers — the image model every video transformer extends.
- Attention — the mechanism itself, at its simplest.
Researcher — Mathematics and papers.
Factorisation, formally
Let a clip give $T$ frames and $S$ spatial patches per frame, so $N = TS$ tokens. Self-attention over all tokens costs
$$ O(N^2 d) = O(T^2 S^2 d) $$
for embedding dimension $d$. The dependence on $T$ is quadratic, which is what makes long clips unreachable.
Divided space-time attention (Bertasius et al., 2021, TimeSformer) applies temporal attention within each spatial position, then spatial attention within each frame:
$$ O(T S^2 d) + O(S T^2 d) = O(TS \cdot (S + T) \cdot d) $$
For $S = 196$ and $T = 8$ this is a 7.7x reduction, matching the developer table. TimeSformer compared five schemes — space-only, joint space-time, divided, sparse local-global, and axial — and found divided attention best on both Kinetics-400 and Something-Something v2. The paper attributes this to divided attention learning separate parameters for the spatial and temporal roles, giving higher modelling capacity than joint attention at lower cost.
ViViT (Arnab et al., 2021) enumerated four models: joint space-time; factorised encoder (a spatial transformer per frame, then a temporal transformer over per-frame representations); factorised self-attention; and factorised dot-product attention. The factorised encoder gives the best complexity-accuracy trade-off, and is effectively late fusion with a learned temporal aggregator.
ViViT also introduced tubelet embedding: a 3D patch of size $t \times h \times w$ instead of a per-frame 2D patch. This fuses temporal information at the tokenisation step and reduces the token count by a factor of $t$. VideoMAE's tubelet_size=2 is this idea, and it is why a 16-frame clip yields 1568 tokens rather than 3136.
Attention with structure
Two designs reintroduce convolution-like inductive bias rather than removing pairs:
MViT / MViTv2 (Fan et al., ICCV 2021; Li et al., CVPR 2022) build a multiscale feature hierarchy inside the transformer. Pooling attention downsamples keys and values with a strided pooling operator, so resolution drops and channel count rises through the network, exactly as in a CNN. This makes the token count shrink with depth instead of staying constant.
Video Swin (Liu et al., CVPR 2022) restricts attention to non-overlapping local 3D windows, shifting the window partition between blocks so information crosses window boundaries. Attention cost becomes linear in the number of tokens. Structurally this is close to a large separable convolution with content-dependent weights, and the paper argues explicitly that reintroducing locality is why it beats global-attention models at equal compute.
Self-supervised pretraining
Supervised video pretraining is bounded by label availability. The masked-modelling route removed that bound.
VideoMAE (Tong et al., NeurIPS 2022) masks 90–95% of tubelets and reconstructs the missing pixels. Three findings from the paper:
- Extremely high masking ratios work on video, higher than on images, because temporal redundancy makes reconstruction too easy at lower ratios.
- Tube masking — masking the same spatial positions across all frames — is essential. Random per-frame masking lets the model copy from the adjacent frame instead of learning structure.
- It performs well on datasets of only 3,000 to 4,000 videos with no extra data, and the paper reports data quality mattering more than data quantity.
Reported results with a plain ViT backbone and no extra data: 83.9% on Kinetics-400, 75.3% on Something-Something v2, 90.8% on UCF101, 61.1% on HMDB51.
V-JEPA 2 (Meta, June 2025) predicts in representation space rather than pixel space, trained on over a million hours of video. It reports 77.3% top-1 on Something-Something v2 and 39.7 recall-at-5 on Epic-Kitchens-100 action anticipation, and post-trains into an action-conditioned world model used for zero-shot robot planning. The direction of travel is clear: video pretraining objectives are moving from pixel reconstruction to latent prediction.
What the benchmark numbers do and do not show
Three cautions when reading video transformer results:
Kinetics rewards scene recognition. Many Kinetics classes are identifiable from a single frame. Strong Kinetics accuracy is weak evidence of temporal modelling. Something-Something v2, whose classes come in direction-reversed pairs, is the discriminating benchmark.
Evaluation protocol is often buried. A 10 x 3 view protocol adds 2 to 4 points of top-1 over 1 x 1 on Kinetics-400 — comparable to the gap between competing architectures. Comparisons across differing protocols are not comparisons.
Pretraining corpus dominates architecture. ViViT and TimeSformer results depend on ImageNet-21k or JFT initialisation. Liu et al. (2022), A ConvNet for the 2020s, made the parallel argument in the image domain: much of the reported transformer advantage came from training recipes rather than architecture. The same caution applies to video, where compute budgets are larger and ablations rarer.
References
- Bertasius, Wang and Torresani, Is Space-Time Attention All You Need for Video Understanding? (TimeSformer), ICML 2021 — arxiv.org/abs/2102.05095
- Arnab et al., ViViT: A Video Vision Transformer, ICCV 2021 — arxiv.org/abs/2103.15691
- Fan et al., Multiscale Vision Transformers, ICCV 2021 — arxiv.org/abs/2104.11227
- Liu et al., Video Swin Transformer, CVPR 2022 — arxiv.org/abs/2106.13230
- Tong et al., VideoMAE, NeurIPS 2022 — arxiv.org/abs/2203.12602
- Assran et al., V-JEPA 2, 2025 — arxiv.org/abs/2506.09985
What to learn next
- Action recognition in practice — putting one of these models to work.
- Vision transformers — the image model every video transformer extends.
- Attention — the mechanism itself, at its simplest.