Video Understanding and Tracking
Action recognition in practice
Getting a video classifier working is mostly pipeline work, and the two decisions that move your number most are how you sample clips and how you combine their predictions.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Action recognition means naming what someone is doing in a video clip. Image classification names what is in a photo. This names the activity instead.
Think of a friend sending you a six-second video with no message. You watch it and think "she is kneading dough". You did not read a caption. You watched movement and named it.
That naming is the task. The label describes an activity, not an object.
The difference from photo labelling is small to say and large to build. A photo has one moment. A video has hundreds, and the answer lives in how they change.
Why it exists
Enormous amounts of video exist and almost none of it is described.
Nobody is going to sit and write "person opening a fridge" against a million security clips. Nobody will tag every rally in a season of matches. Search, moderation, safety alerts and sports statistics all wait on that labelling being automatic.
How it works
long video
│
▼
pick a few frames <- the sampling decision
│
▼
resize, crop, normalise <- must match what the model was trained on
│
▼
[ video model ]
│
▼
a score for every action label
│
▼
repeat for several clips, average the scores <- the aggregation decision
│
▼
"pouring tea" (34% sure)Only the middle box is the neural network. Everything above and below it is pipeline, and that is where most of your accuracy is won or lost.
The two decisions that matter most
Which frames you show the model. Covered in frame sampling strategies. Show a short window and you may miss the action entirely.
How many clips you try, and how you combine the answers. One clip from one part of the video gives one opinion. Take several clips from across the video and average their opinions, and you get a noticeably better answer.
That second point is worth sitting with. The model does not improve. You asked it more times and pooled the replies, exactly as you would ask several friends rather than one.
A trap in the benchmarks
Many popular action labels can be guessed from a single frame.
"Swimming" needs water and a person. "Playing guitar" needs a guitar. A model can score respectably on such datasets while understanding nothing about movement.
The honest test is a dataset where the same objects appear in both classes and only the order differs. "Pushing something left to right" against "pushing something right to left". Models that fake it collapse there.
Where you have already seen this
- YouTube Shorts recommending clips of a specific activity.
- Fitness apps counting repetitions and correcting your form.
- Retail analytics counting how many people picked up a product.
- Elderly-care sensors detecting a fall.
What is honestly hard here
Video datasets are large enough to hurt.
Kinetics, the standard benchmark, holds hundreds of thousands of clips and takes days to download and terabytes to store. Training from scratch needs many graphics cards for many days.
Almost nobody does that. Almost everybody starts from a released checkpoint and fine-tunes on a few thousand of their own clips. That path works well and is the one to take. Read that twice, because it is the difference between a project that finishes and one that does not.
Remember this
- Action recognition names an activity in a clip, not an object in a frame.
- Most of the accuracy lives in sampling and aggregation, not in the model.
- Start from a pretrained checkpoint. Training from scratch is a research-scale project.
What to learn next
- Video object segmentation — from naming a clip to outlining a thing in it.
- Model evaluation — the metrics behind these accuracy numbers.
- Transfer learning in PyTorch — the fine-tuning recipe in full.
Developer — Code and libraries.
Setup
pip install torch torchvisionRun against torch 2.5.1 and torchvision 0.20.1 on CPU, Python 3.10. The torchvision video model API has been stable since 0.13; the current stable release at the time of writing is the 0.28 line, and these calls are unchanged.
What a video classifier actually is
import torch
import torchvision
from torchvision.models.video import r3d_18, R3D_18_Weights
print("torch", torch.__version__, "| torchvision", torchvision.__version__)
# weights=None keeps this offline: random weights, real architecture, real shapes.
model = r3d_18(weights=None, num_classes=400).eval()
clip = torch.zeros(1, 3, 16, 112, 112) # (N, C, T, H, W)
with torch.no_grad():
logits = model(clip)
print("\ninput :", tuple(clip.shape))
print("logits:", tuple(logits.shape))
print("parameters:", f"{sum(p.numel() for p in model.parameters()):,}")
# The preprocessing the pretrained weights expect, read from the weights metadata
# rather than copied from a blog post. No download happens for this.
meta = R3D_18_Weights.KINETICS400_V1
print("\npretrained preprocessing:", meta.transforms())
print("classes:", len(meta.meta["categories"]),
"| first three:", meta.meta["categories"][:3])torch 2.5.1+cu121 | torchvision 0.20.1+cu121
input : (1, 3, 16, 112, 112)
logits: (1, 400)
parameters: 33,371,472
pretrained preprocessing: VideoClassification(
crop_size=[112, 112]
resize_size=[128, 171]
mean=[0.43216, 0.394666, 0.37645]
std=[0.22803, 0.22145, 0.216989]
interpolation=InterpolationMode.BILINEAR
)
classes: 400 | first three: ['abseiling', 'air drumming', 'answering questions']Read the preprocessing block carefully, because it is the single most common source of silent accuracy loss.
resize_size=[128, 171] is not square. Kinetics videos are widescreen. The weights expect a resize to 128 by 171 followed by a 112 by 112 crop. Resizing to a square first distorts every frame and costs you accuracy with no error message.
Those mean and standard deviation values are Kinetics statistics, not ImageNet ones. ImageNet's familiar [0.485, 0.456, 0.406] is wrong here. Use meta.transforms() rather than typing numbers from memory.
weights=None gives you the architecture with no download. Use it for shape debugging. For real predictions, pass weights=R3D_18_Weights.DEFAULT, which fetches roughly 130 MB once and caches it.
Available models in torchvision.models.video: r3d_18, mc3_18, r2plus1d_18, s3d, mvit_v1_b, mvit_v2_s, swin3d_t, swin3d_s, swin3d_b. The ResNet-style ones are 112-pixel models; the transformer ones are 224-pixel models with different preprocessing. Always read .transforms() for whichever you pick.
Why one clip is not enough
The model is fixed. What changes is how many times you ask it.
import numpy as np
CLASSES = ["pouring tea", "washing hands", "chopping onion", "clapping"]
rng = np.random.default_rng(7)
# Pretend the model scored 10 clips cut from the same video. The truth is class 0.
true = 0
logits = rng.normal(0, 1.6, size=(10, len(CLASSES)))
logits[:, true] += 1.0 # the right class is favoured, but only slightly
def softmax(z):
e = np.exp(z - z.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
probs = softmax(logits)
per_clip = probs.argmax(axis=1)
print("per-clip predictions:", [CLASSES[i] for i in per_clip])
print(f"single-clip accuracy: {(per_clip == true).mean():.2f}")
print("\naggregated over all 10 clips")
print(f" mean of softmax -> {CLASSES[probs.mean(0).argmax()]}")
print(f" mean of logits -> {CLASSES[logits.mean(0).argmax()]}")
print(f" majority vote -> {CLASSES[np.bincount(per_clip, minlength=4).argmax()]}")
np.set_printoptions(precision=3, suppress=True)
print("\nmean softmax per class:", probs.mean(0))per-clip predictions: ['pouring tea', 'clapping', 'chopping onion', 'pouring tea', 'washing hands', 'clapping', 'pouring tea', 'pouring tea', 'chopping onion', 'washing hands'] single-clip accuracy: 0.40 aggregated over all 10 clips mean of softmax -> pouring tea mean of logits -> pouring tea majority vote -> pouring tea mean softmax per class: [0.343 0.219 0.175 0.263]
Four of ten clips were right on their own. All three aggregation rules got the video right.
The mean probability for the correct class is 0.343 against 0.263 for the runner-up. That is a thin margin, and it is realistic. Aggregation buys you a few points on a benchmark; it does not turn a weak model into a strong one.
Which rule to use, in practice:
- Mean of softmax is the standard and the safest default. Each clip contributes a bounded amount.
- Mean of logits lets one very confident clip dominate. Better when the action occurs in a short window and most clips are uninformative.
- Majority vote discards confidence entirely. It behaves poorly with many classes, because votes scatter.
The full pipeline, in order
import torch
from torchvision.models.video import r3d_18, R3D_18_Weights
weights = R3D_18_Weights.DEFAULT # downloads on first use
model = r3d_18(weights=weights).eval()
preprocess = weights.transforms()
# frames: a uint8 tensor (T, C, H, W) decoded from your video, RGB, full resolution.
# Sample several clips of 16 frames each, spread across the video.
def predict(frames, clip_len=16, n_clips=5):
starts = torch.linspace(0, max(len(frames) - clip_len, 0), n_clips).long()
batch = torch.stack([preprocess(frames[s:s + clip_len]) for s in starts])
with torch.no_grad():
probs = model(batch).softmax(-1) # (n_clips, 400)
mean = probs.mean(0)
top = mean.topk(3)
return [(weights.meta["categories"][i], float(p)) for p, i in zip(*top)]No output block, deliberately. The result depends entirely on which video you feed it, and inventing a plausible-looking prediction would be dishonest.
Note preprocess(frames[s:s + clip_len]) operates on a whole clip. The transform applies the same resize and crop to every frame in it, which is what keeps motion intact. Cropping each frame independently destroys the motion the model is looking for.
Fine-tuning on your own classes
The realistic project. Replace the head, freeze most of the backbone, train briefly.
import torch.nn as nn
from torchvision.models.video import r3d_18, R3D_18_Weights
model = r3d_18(weights=R3D_18_Weights.DEFAULT)
for p in model.parameters():
p.requires_grad = False # freeze the pretrained backbone
model.fc = nn.Linear(model.fc.in_features, 5) # 5 classes of your own, trainable
trainable = [n for n, p in model.named_parameters() if p.requires_grad]
print("trainable parameters:", trainable)trainable parameters: ['fc.weight', 'fc.bias']
That runs after the checkpoint downloads. Start here, measure, then unfreeze layer4 if you need more. Unfreezing everything at once on a few thousand clips overfits — the same pattern as in transfer learning in PyTorch.
Fine-tuning a 3D model on a few thousand clips is feasible on one mid-range GPU in hours. Training from scratch on Kinetics is not; the published models used many GPUs for days, and that is stated plainly in the papers.
Common mistakes
Normalising with ImageNet statistics. See the output above. Kinetics models have their own.
Applying random crops per frame during training. Choose the augmentation once per clip. Per-frame randomness is a synthetic camera shake that destroys the motion signal.
Comparing your accuracy against a published number under a different view protocol. A paper's 10 x 3 result is 2 to 4 points above its own 1 x 1 result on Kinetics-400. State your protocol whenever you state a number.
Splitting train and test by clip instead of by source video. Two clips from the same video share background, lighting and person. Splitting by clip leaks, and your reported accuracy is fiction — see train-test split for the general pattern.
Reporting Kinetics-style accuracy as evidence of temporal understanding. Test on Something-Something v2 if that is the claim you want to make.
Downloading Kinetics before checking you need it. It is hundreds of thousands of clips and the download links rot. Fine-tune from a checkpoint on your own data first.
Try it yourself
Modify the aggregation script so only two of the ten clips actually contain the action, with the other eight scoring near-uniformly. Compare mean-of-softmax against mean-of-logits and against taking the maximum over clips. Then work out which real-world situation each rule suits.
What to learn next
- Video object segmentation — from naming a clip to outlining a thing in it.
- Model evaluation — the metrics behind these accuracy numbers.
- Transfer learning in PyTorch — the fine-tuning recipe in full.
Researcher — Mathematics and papers.
The task, and its benchmarks
Trimmed action recognition maps a clip to a single label. The distinction from three neighbouring tasks matters when reading results:
| Task | Output | Benchmarks |
|---|---|---|
| Trimmed recognition | one label per clip | Kinetics, UCF-101, HMDB-51, SSv2 |
| Temporal localisation | label plus start and end times | ActivityNet, THUMOS |
| Spatio-temporal detection | label plus a box per frame | AVA, AVA-Kinetics |
| Multi-label recognition | several concurrent labels | Charades |
Numbers are not comparable across these. A "state of the art on video understanding" claim without naming the task is not a claim.
The Kinetics shortcut problem
Kinetics-400 (Kay et al., 2017) is the default pretraining corpus, and its structure biases the field.
Single-frame baselines reach a substantial fraction of full video model accuracy on Kinetics, because scene and object context are highly predictive of the label. Zhou et al. (2018) and Xie et al. (2018) both documented this. The consequence is that architectural improvements measured only on Kinetics may be improving object recognition rather than temporal reasoning.
Something-Something v2 (Goyal et al., 2017) was constructed to remove that shortcut. Classes are templated with objects abstracted away — "Pushing [something] from left to right" against "Pushing [something] from right to left". Models that discard temporal order are capped near chance on those pairs, and reported SSv2 accuracy is therefore the more informative number for temporal claims.
Evaluation protocol, stated precisely
The convention is K x C: K temporal clips uniformly spaced, C spatial crops each, with predictions averaged over all K·C passes.
1 x 1: centre clip, centre crop. Deployment-realistic, lowest score.10 x 3: the SlowFast and Video Swin convention.4 x 3or5 x 3: common for transformers, where per-pass cost is higher.
Averaging in probability space (mean of softmax) and in logit space give different results, and papers rarely state which. Mean of softmax is the more common and the more conservative.
Two further protocol details that shift numbers by a point or more: whether the spatial crops are three horizontal crops of the full-height frame (SlowFast convention) or corner-plus-centre crops; and whether frames are resized by shorter side or to a fixed size.
Training recipe, and what actually matters
Ablations across the 3D CNN and video transformer literature agree on a rough ordering of what moves accuracy:
- Pretraining corpus. ImageNet-21k or Kinetics-600 initialisation is worth more than any architectural choice at fixed compute.
- Clip length and sampling. Longer temporal receptive fields help most on datasets with long actions, and not at all on Kinetics.
- Regularisation. For transformers: RandAugment, mixup, CutMix, label smoothing, stochastic depth. Video transformers without heavy augmentation overfit badly.
- Resolution. Accuracy per FLOP is usually better spent on temporal extent than on spatial resolution beyond 224 pixels.
For 3D CNNs, batch normalisation statistics are a recurring source of irreproducibility. Small per-GPU batch sizes make batch-norm estimates noisy; synchronised batch norm across GPUs is standard in the SlowFast codebase and its absence changes results measurably.
Efficiency, for deployment
Reported inference cost is often the 1 x 1 figure and the reported accuracy the 10 x 3 figure. Insist on both under the same protocol.
Practical routes to a deployable video classifier:
- X3D-S / X3D-M (Feichtenhofer, 2020): accuracy per FLOP remains competitive and the models are small.
- Frame-level model plus temporal pooling: a strong image backbone with TSN-style segment consensus is a surprisingly hard baseline to beat on scene-driven datasets, and parallelises across frames without effort.
- Adaptive clip selection: SCSampler-style cheap prefiltering, worthwhile on untrimmed long video.
- Quantisation and distillation: see quantization in practice and knowledge distillation. Video models quantise about as well as image models; the memory saving matters more here because activations are larger.
References
- Kay et al., The Kinetics Human Action Video Dataset, 2017 — arxiv.org/abs/1705.06950
- Goyal et al., The "Something Something" Video Database, ICCV 2017 — arxiv.org/abs/1706.04261
- Carreira and Zisserman, Quo Vadis, Action Recognition?, CVPR 2017 — arxiv.org/abs/1705.07750
- Feichtenhofer et al., SlowFast Networks, ICCV 2019 — arxiv.org/abs/1812.03982
- torchvision video models reference — docs.pytorch.org/vision/stable/models.html
What to learn next
- Video object segmentation — from naming a clip to outlining a thing in it.
- Model evaluation — the metrics behind these accuracy numbers.
- Transfer learning in PyTorch — the fine-tuning recipe in full.