Video Understanding and Tracking
Optical flow
Optical flow measures how far every pixel moved between two frames, giving you motion as a picture instead of guessing it from appearance.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Optical flow is an arrow drawn on every pixel, saying how far that pixel moved since the last frame.
Sit by a train window and watch the world go past. Nearby poles rush backwards. Distant hills drift slowly. The sky barely moves at all.
You are reading motion, and you are reading a different amount of motion at every point in your view. That map of "how fast is this bit moving, and which way" is optical flow.
Your brain does this without being asked. A computer has to be told how.
Why it exists
A single picture tells you what is there. Two pictures together tell you what is happening.
Walking and standing still look identical in one frame. So do a door opening and a door closing. So do a ball rising and a ball falling.
Optical flow pulls that missing information out and hands it over as a picture of its own. Feed a model both the colours and the motion, and it can tell those pairs apart.
How it works
The core idea is a guess that is usually true. A patch of the world keeps its brightness as it moves.
frame 1 frame 2
┌───────────┐ ┌───────────┐
│ ▓▓ │ │ ▓▓ │
│ ▓▓ │ ───────► │ ▓▓ │
└───────────┘ └───────────┘
^ ^
patch here same patch, 3 pixels right
flow at that pixel = "3 right, 0 down"So the computer takes a small patch from the first frame. It hunts nearby in the second frame for the patch that looks most like it. The distance it had to travel becomes the arrow.
Do that for every pixel and you have a full flow field. Two numbers per pixel: one for sideways movement, one for up and down.
The problem that never fully goes away
Look at a plain white wall through a drinking straw. Move the wall. Nothing appears to change.
There is no detail in view, so no movement can be measured. This is called the aperture problem — no visible detail means no measurable motion. You cannot tell how something moved without features to track.
Real footage is full of these regions. Blank walls, clear sky, still water, a plain shirt. Flow in those places is a guess, not a measurement. Treating a guess as a measurement is how flow systems fail.
Where you have already seen this
- Video calls staying smooth on a weak connection, because the codec sends motion instead of pixels.
- Phone camera stabilisation smoothing out your shaky hands.
- Slow-motion features that invent frames between the real ones.
- Cricket broadcast graphics tracking the ball's path.
What is honestly hard here
Flow is confidently wrong when the motion is large.
A patch that moved three pixels is easy to find again. A patch that moved forty pixels may be matched to something else entirely. The answer that comes back still looks like a normal answer. There is no error flag.
The standard fix is to shrink the image first, measure the big motion there, then refine. It helps enormously and it does not fully solve the problem. You will see this failure in the code below.
Remember this
- Optical flow gives one arrow per pixel: how far and which way it moved.
- It assumes a patch keeps its brightness while it moves.
- Blank regions and large motions are where it breaks, and it breaks without warning.
What to learn next
- 3D convolutions for video — learning motion instead of measuring it.
- Tracking by detection — where explicit correspondence still wins.
- Video object segmentation — pushing a mask forward using motion.
Developer — Code and libraries.
Setup
pip install opencv-python numpyRun against opencv-python 4.11.0 and numpy 1.26.4. The RAFT section further down needs pip install torch torchvision; it was run against torch 2.5.1 and torchvision 0.20.1.
Dense flow with a known answer
Testing flow on real footage is hard, because you do not know the true motion. So we make our own: take a textured image and shift it by an exact amount.
import cv2
import numpy as np
rng = np.random.default_rng(0)
# Flow needs texture. A blank wall has no measurable motion, as we prove below.
base = rng.integers(0, 256, size=(128, 128), dtype=np.uint8)
frame1 = cv2.GaussianBlur(base, (5, 5), 1.5)
DX, DY = 3, 1 # the true motion, in pixels
frame2 = np.roll(np.roll(frame1, DX, axis=1), DY, axis=0)
def farneback(a, b, levels=3):
return cv2.calcOpticalFlowFarneback(a, b, None, pyr_scale=0.5, levels=levels,
winsize=15, iterations=3, poly_n=5,
poly_sigma=1.2, flags=0)
flow = farneback(frame1, frame2)
print("flow array shape:", flow.shape, "-> (H, W, 2): one (dx, dy) per pixel")
inner = flow[20:-20, 20:-20] # ignore the border, where np.roll wraps
print(f"mean dx in the interior: {inner[..., 0].mean():+.2f} (true {DX:+d})")
print(f"mean dy in the interior: {inner[..., 1].mean():+.2f} (true {DY:+d})")
truth = np.zeros_like(inner)
truth[..., 0], truth[..., 1] = DX, DY
epe = np.linalg.norm(inner - truth, axis=2) # endpoint error, the standard flow metric
print(f"endpoint error: mean {epe.mean():.2f} px, 95th percentile {np.percentile(epe, 95):.2f} px")
flat = np.full((128, 128), 128, np.uint8) # a blank wall, moved 3 pixels
flat_flow = farneback(flat, np.roll(flat, DX, axis=1))[20:-20, 20:-20]
print(f"\nflat grey region moving {DX} px -> measured dx {flat_flow[..., 0].mean():+.2f}")
print("that zero is the aperture problem: no texture, no measurable motion")flow array shape: (128, 128, 2) -> (H, W, 2): one (dx, dy) per pixel mean dx in the interior: +3.00 (true +3) mean dy in the interior: +1.00 (true +1) endpoint error: mean 0.00 px, 95th percentile 0.00 px flat grey region moving 3 px -> measured dx +0.00 that zero is the aperture problem: no texture, no measurable motion
An endpoint error of 0.00 is not a typical result. This is the easiest case that exists: heavy texture everywhere, an exact integer shift, no lighting change, no occlusion, no object boundaries. Real footage produces endpoint errors of one to several pixels. Treat this run as a check that the code is correct, not as evidence that flow is easy.
The flat region returns exactly zero. The wall moved and the algorithm reports no motion. It is not a bug. There is no information in that patch to measure motion from, and the algorithm is right to refuse.
Where it breaks: large motion
The pyramid is the trick that lets flow see beyond a few pixels. Shrink both frames, measure the coarse motion, then refine at full size. levels controls how many times it shrinks.
import cv2
import numpy as np
rng = np.random.default_rng(0)
f1 = cv2.GaussianBlur(rng.integers(0, 256, (256, 256), dtype=np.uint8), (5, 5), 1.5)
def shift(img, dx):
M = np.float32([[1, 0, dx], [0, 1, 0]])
return cv2.warpAffine(img, M, img.shape[::-1], borderMode=cv2.BORDER_WRAP)
def measure(a, b, levels):
flow = cv2.calcOpticalFlowFarneback(a, b, None, 0.5, levels, 15, 3, 5, 1.2, 0)
return flow[40:-40, 40:-40, 0].mean()
print(f"{'true dx':>9}{'1 level':>10}{'3 levels':>10}{'5 levels':>10}")
for dx in (2.5, 8.0, 24.0, 40.0):
f2 = shift(f1, dx)
row = "".join(f"{measure(f1, f2, lv):+10.2f}" for lv in (1, 3, 5))
print(f"{dx:9.1f}{row}") true dx 1 level 3 levels 5 levels
2.5 +2.50 +2.50 +2.50
8.0 +0.58 +8.00 +8.00
24.0 -0.00 +23.99 +23.99
40.0 -0.17 -7.95 -7.95This table is the most useful thing on this page.
Sub-pixel motion works fine at any setting. 2.5 pixels comes back as 2.50. Flow is not restricted to whole pixels.
One level cannot see 8 pixels of motion. It reports 0.58 and then 0.00. Without a pyramid, the search window is the whole search, and motion beyond it is invisible.
Three levels handle 24 pixels. The coarse level sees the motion at a quarter scale, where 24 pixels becomes 3, and the fine levels polish it.
At 40 pixels, every setting fails — and returns -7.95. Read that again. It did not return zero, and it did not raise an error. It returned a confident, precise, wrong number, with the sign flipped.
That is the failure mode you must design around. Any pipeline consuming flow needs its own sanity check, because flow itself will not tell you it failed. Compare forward flow against backward flow: where they do not cancel out, the flow is untrustworthy.
Sparse flow, when you only care about a few points
Dense flow computes a vector for every pixel. If you are tracking a handful of features, Lucas-Kanade is far cheaper.
import cv2
import numpy as np
rng = np.random.default_rng(0)
f1 = cv2.GaussianBlur(rng.integers(0, 256, (200, 200), dtype=np.uint8), (5, 5), 1.5)
M = np.float32([[1, 0, 4.0], [0, 1, -2.0]])
f2 = cv2.warpAffine(f1, M, (200, 200), borderMode=cv2.BORDER_WRAP)
pts = cv2.goodFeaturesToTrack(f1, maxCorners=20, qualityLevel=0.3, minDistance=15)
print("corners found:", len(pts))
nxt, status, err = cv2.calcOpticalFlowPyrLK(f1, f2, pts, None,
winSize=(21, 21), maxLevel=3)
good = status.ravel() == 1 # status 0 means the point was lost
d = (nxt - pts)[good].reshape(-1, 2)
print("tracked ok :", int(good.sum()))
print("median shift: dx=%+.2f dy=%+.2f (true +4.00, -2.00)"
% (np.median(d[:, 0]), np.median(d[:, 1])))corners found: 20 tracked ok : 20 median shift: dx=+4.00 dy=-2.00 (true +4.00, -2.00)
Note status. Lucas-Kanade tells you which points it lost, which dense flow never does. Always check it, and always take a median rather than a mean — one bad point ruins a mean.
The learned option: RAFT
Since 2020 the accurate answer has been a neural network. RAFT builds a table of similarities between all pairs of pixels, then refines a flow estimate over many small steps.
import torch
import torchvision
from torchvision.models.optical_flow import raft_small, Raft_Small_Weights
print("torchvision", torchvision.__version__)
# weights=None keeps this offline: random weights, real architecture, real shapes.
model = raft_small(weights=None).eval()
a = torch.zeros(1, 3, 128, 128) # RAFT wants both frames, normalised to [-1, 1]
b = torch.zeros(1, 3, 128, 128)
with torch.no_grad():
flows = model(a, b) # a LIST, not a tensor
print("refinement steps returned:", len(flows))
print("each step's shape :", tuple(flows[-1].shape), "-> (N, 2, H, W)")
print("parameters :", f"{sum(p.numel() for p in model.parameters()):,}")
print("\npretrained preprocessing:", Raft_Small_Weights.DEFAULT.transforms())torchvision 0.20.1+cu121 refinement steps returned: 12 each step's shape : (1, 2, 128, 128) -> (N, 2, H, W) parameters : 990,162 pretrained preprocessing: OpticalFlow()
Two things to take from this. The model returns a list, one flow field per refinement step, and the last one is the answer. And the flow tensor is (N, 2, H, W), channels-first, where OpenCV gives (H, W, 2), channels-last. Mixing the two conventions is a frequent bug.
This ran with random weights so no download was needed. For real flow you need weights=Raft_Small_Weights.DEFAULT, which downloads a checkpoint. Running RAFT on full-resolution video is slow on CPU and belongs on a GPU.
Common mistakes
Using flow magnitude as a feature without normalising by frame rate. The same motion at 60 frames per second produces half the per-frame displacement of the same motion at 30. Your feature now encodes the camera, not the action.
Ignoring status from Lucas-Kanade. Lost points return stale coordinates. Filter on status == 1 and on err before using anything.
Trusting flow near object boundaries. A pixel on the edge of a moving car belongs to the car in one frame and the road in the next. Its flow is meaningless. This is occlusion, and it is where most of the error lives in every benchmark.
Forgetting to convert to greyscale. calcOpticalFlowFarneback needs single-channel input. Passing a colour frame raises an unhelpful error about channel counts.
Computing flow at full resolution for a model that consumes 224 pixels. Resize first. Flow cost scales with pixel count and you are throwing the detail away regardless.
Try it yourself
Add a backward-flow check to the first script: compute flow from frame2 to frame1, warp it, and add it to the forward flow. Where the sum is far from zero, the flow is unreliable. Run it on the 40-pixel failure case and watch the check catch what the flow field never admitted.
What to learn next
- 3D convolutions for video — learning motion instead of measuring it.
- Tracking by detection — where explicit correspondence still wins.
- Video object segmentation — pushing a mask forward using motion.
Researcher — Mathematics and papers.
The brightness constancy equation
Assume a point's intensity is unchanged as it moves:
$$ I(x, y, t) = I(x + \Delta x,\; y + \Delta y,\; t + \Delta t) $$
A first-order Taylor expansion of the right-hand side gives the optical flow constraint equation:
$$ I_x u + I_y v + I_t = 0 $$
Where $I_x, I_y$ are spatial image gradients, $I_t$ is the temporal gradient, and $(u, v)$ is the flow vector at that pixel.
One equation, two unknowns. The system is underdetermined at every pixel, which is the aperture problem stated algebraically. Only the flow component along the image gradient is observable. Every classical method is a different choice of extra constraint.
The three classical answers
Lucas and Kanade (1981) assume flow is constant over a window $\Omega$ and solve by least squares:
$$ \begin{bmatrix} u \ v \end{bmatrix} = \left( \sum_{\Omega} \begin{bmatrix} I_x^2 & I_x I_y \ I_x I_y & I_y^2 \end{bmatrix} \right)^{-1} \left( -\sum_{\Omega} \begin{bmatrix} I_x I_t \ I_y I_t \end{bmatrix} \right) $$
That matrix is the structure tensor. Its eigenvalues diagnose the pixel: two large eigenvalues means a corner and a well-posed solve; one large means an edge and the aperture problem; two small means flat texture and no solution. This is exactly why goodFeaturesToTrack, which is the Shi-Tomasi corner detector, is the right companion to Lucas-Kanade.
Horn and Schunck (1981) instead add a global smoothness penalty and minimise:
$$ E = \iint \left[ (I_x u + I_y v + I_t)^2 + \alpha^2 \left( |\nabla u|^2 + |\nabla v|^2 \right) \right] dx\, dy $$
Where $\alpha$ trades data fidelity against smoothness. The result is dense but over-smoothed across motion boundaries, which the modern total-variation family (Zach et al., 2007) fixes with an $L^1$ penalty that permits discontinuities.
Farnebäck (2003) approximates each neighbourhood by a quadratic polynomial and derives displacement from how the polynomial coefficients change. This is what calcOpticalFlowFarneback implements, and it is why the method is comparatively robust to a global illumination scaling: a gain change alters the polynomial coefficients in a way partially absorbed by the estimation.
Coarse-to-fine, and its ceiling
All three methods linearise, so they are only valid for displacements smaller than roughly the window size. The standard remedy is a Gaussian pyramid: estimate at scale $2^{-L}$, upsample, warp, and refine.
The range this buys is approximately $w \cdot 2^{L}$ pixels for window size $w$. The table in the developer section is that formula made visible.
The ceiling is structural, not a tuning problem. Small fast-moving objects vanish at coarse pyramid levels before their motion can be estimated. Brox and Malik (2011) address it by injecting descriptor matches at full resolution; RAFT sidesteps it by never linearising in the first place.
RAFT and what changed
Teed and Deng (2020), RAFT: Recurrent All-Pairs Field Transforms, restructured the problem:
- Encode both frames into features at $1/8$ resolution.
- Build an all-pairs correlation volume: the inner product of every feature in frame 1 with every feature in frame 2, giving an $H/8 \times W/8 \times H/8 \times W/8$ tensor, pooled into a 4-level pyramid.
- Initialise flow at zero and run a GRU that repeatedly looks up correlation values at the current flow estimate and emits an update.
Two properties matter. The correlation volume is built once, at full spatial resolution in the feature space, so small fast objects are not destroyed by downsampling. And the iterative updates share weights, so more iterations at test time is a cost-accuracy dial rather than a retraining decision. The torchvision model returning a list of 12 flow fields is this loop made visible.
RAFT won the ECCV 2020 best paper award and reduced Sintel and KITTI errors substantially over the preceding PWC-Net family. Successors worth knowing: GMA (Jiang et al., 2021) adds global motion aggregation for occluded regions; FlowFormer (Huang et al., 2022) replaces the GRU with a transformer over the cost volume; SEA-RAFT (Wang et al., ECCV 2024) simplifies the architecture and adds a mixture-of-Laplace loss, reaching strong Spring and Sintel results at higher throughput.
Metrics and benchmarks
- EPE (endpoint error): mean Euclidean distance between predicted and ground-truth flow vectors. The default on Sintel.
- Fl-all: percentage of pixels with EPE above 3 px and above 5% of the ground-truth magnitude. The default on KITTI, and more informative than EPE because it is scale-aware.
- Sintel splits results into
cleanandfinalpasses;finaladds motion blur and atmospheric effects and is the harder, more realistic number.
Ground truth for real footage is close to impossible to obtain, so the field trains on synthetic data (FlyingChairs, FlyingThings3D, Sintel) and evaluates zero-shot generalisation to KITTI and Spring. Any claim about flow accuracy on real video is a claim about synthetic-to-real transfer.
Flow in video recognition: the honest position
Two-stream networks (Simonyan and Zisserman, 2014) fed RGB and stacked flow to separate CNNs and won large gains. I3D inherited the design, and for several years pre-computed flow was standard.
That has largely reversed. Flow is expensive to compute, cannot be trained end to end when pre-computed, and roughly doubles storage. Modern 3D convolutional and transformer models learn motion representations directly from RGB and close most of the gap. Where flow still earns its keep:
- Tracking and video object segmentation, where explicit correspondence is the task rather than a feature.
- Frame interpolation and video compression, where the flow field is the output.
- Small-data regimes, where flow is a strong hand-designed prior that a network would need far more data to learn.
References
- Lucas and Kanade, An Iterative Image Registration Technique, IJCAI 1981.
- Horn and Schunck, Determining Optical Flow, Artificial Intelligence, 1981.
- Farnebäck, Two-Frame Motion Estimation Based on Polynomial Expansion, SCIA 2003.
- Simonyan and Zisserman, Two-Stream Convolutional Networks, NeurIPS 2014 — arxiv.org/abs/1406.2199
- Teed and Deng, RAFT, ECCV 2020 — arxiv.org/abs/2003.12039
- Jiang et al., Learning to Estimate Hidden Motions with GMA, ICCV 2021 — arxiv.org/abs/2104.02409
- Wang et al., SEA-RAFT, ECCV 2024 — arxiv.org/abs/2405.14793
What to learn next
- 3D convolutions for video — learning motion instead of measuring it.
- Tracking by detection — where explicit correspondence still wins.
- Video object segmentation — pushing a mask forward using motion.