Video Understanding and Tracking
Video object segmentation
You outline an object in the first frame and the system must keep that exact outline on it for the rest of the video, through turns, blur and disappearances.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Video object segmentation means drawing an outline around one object in the first frame. Then you keep that outline on it for every frame after.
Think of colouring inside the lines in a flipbook you drew yourself. On page one you carefully shade the cat. On page two the cat has moved a little, so you shade it there.
By page fifty the cat has turned around, walked behind a chair and come out the other side. You are still shading the same cat, because you have been following it.
Everything about the task is in that word "following". You are not finding a cat on each page independently. You are keeping hold of this cat.
Why it exists
Ordinary image segmentation, covered in image segmentation, labels every pixel in one picture. Run it on each frame separately and two problems appear at once.
It has no memory. Frame ten has no idea what frame nine decided. Outlines flicker and jump.
It has no identity. With two cats in the scene, nothing links the cat in frame ten to the one in frame nine. You wanted one cat and got a fresh guess every frame.
Video object segmentation fixes both by carrying information forward.
How it works
frame 0 frame 1 frame 2 frame 3
┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐
│ (▓▓) │ │ (▓▓) │ │ (▓▓)│ │ ????? │
└────────┘ └────────┘ └────────┘ └────────┘
you draw system system object went
this outline predicts predicts behind a pole
│ │ │ │
└── remembered ───┴── remembered ───┴─── remembered ────┘
what the object looked like, all the way alongThe system keeps a memory of how the object looked in the frames it has already handled. For each new frame it asks: which pixels here look like the thing I have been remembering?
That memory is the whole design. The naive alternative is to push the outline forward using motion. That works for a few frames and then drifts away, as the code below shows.
Two flavours of the task
Semi-supervised means you are given the first-frame outline. The system only has to follow it. This is the common setting and the one described above.
Unsupervised means nothing is given. The system must decide by itself which object is the interesting one, then follow it. Harder, and less used in practice, because usually you know what you want to track.
Where you have already seen this
- Video calls blurring your background while keeping your outline crisp.
- Phone editors that let you change the colour of one person's shirt across a whole clip.
- Sports graphics that highlight one player as they run.
- Film effects that cut an actor out of a shot, once done by hand, frame by frame, for weeks.
What is honestly hard here
Three things break these systems, and none is fully solved.
Disappearing and coming back. An object passes behind a pillar. When it reappears, the system must know it is the same object and not a new one.
Drift. Small errors accumulate. Include a few background pixels this frame. Use them as your memory next frame. The outline slowly swallows the background.
Similar objects. Two people in identical uniforms standing close together. Appearance memory helps far less when the appearances match.
Read that list twice before assuming this task is easy. Every published system fails on all three. The numbers you see quoted are averages over cases where it works.
Remember this
- You give a first-frame outline; the system keeps it on the same object throughout.
- The key mechanism is memory of how the object looked earlier.
- Drift, occlusion and lookalikes are the failure modes, and they are unsolved.
What to learn next
- Tracking by detection — the same identity problem, solved with boxes instead of masks.
- Image segmentation — the single-frame version of this task.
- Optical flow — the motion signal both baselines above depend on.
Developer — Code and libraries.
Setup
pip install opencv-python numpyRun against opencv-python 4.11.0 and numpy 1.26.4. We build the evaluation metrics and two baselines from scratch, because understanding what J&F measures is worth more than calling a library that prints it.
The metrics, implemented
DAVIS, the standard benchmark, scores with two numbers averaged together.
J is region similarity: intersection over union between the predicted and true masks. F is boundary accuracy: an F1 score between the two outlines, allowing a small tolerance so a one-pixel offset is not punished. The headline number is J&F, their mean.
import cv2
import numpy as np
H = W = 96
N = 12
rng = np.random.default_rng(0)
texture = rng.integers(0, 90, (H, W), dtype=np.uint8) # a static, textured background
def scene(t):
"""Frame t and its true mask: a bright disc drifting, growing and shrinking."""
cx, cy = 22 + 2.7 * t, 38 + 1.3 * t
r = 11 + 2.5 * np.sin(t / 2)
yy, xx = np.mgrid[0:H, 0:W]
mask = ((xx - cx) ** 2 + (yy - cy) ** 2) <= r * r
frame = texture.copy()
frame[mask] = 230
return cv2.GaussianBlur(frame, (3, 3), 0.8), mask
def region_j(pred, gt):
"""DAVIS J: intersection over union of the two masks."""
inter = np.logical_and(pred, gt).sum()
union = np.logical_or(pred, gt).sum()
return 1.0 if union == 0 else inter / union
def boundary_f(pred, gt, tol_frac=0.008):
"""DAVIS F: F1 between the two outlines, allowing a small tolerance."""
def outline(m):
m8 = m.astype(np.uint8)
grad = cv2.morphologyEx(m8, cv2.MORPH_GRADIENT, np.ones((3, 3), np.uint8))
return grad.astype(bool)
tol = max(1, int(round(tol_frac * np.hypot(H, W))))
k = np.ones((2 * tol + 1, 2 * tol + 1), np.uint8)
bp, bg = outline(pred), outline(gt)
bp_d = cv2.dilate(bp.astype(np.uint8), k).astype(bool)
bg_d = cv2.dilate(bg.astype(np.uint8), k).astype(bool)
prec = (bp & bg_d).sum() / max(bp.sum(), 1)
rec = (bg & bp_d).sum() / max(bg.sum(), 1)
return 0.0 if prec + rec == 0 else 2 * prec * rec / (prec + rec)
frames, gts = zip(*[scene(t) for t in range(N)])
copy_mask = gts[0].copy() # baseline 1: never update the mask
flow_mask = gts[0].copy() # baseline 2: push the mask along the optical flow
print(f"{'frame':>5}{'copy J':>9}{'copy F':>9}{'flow J':>9}{'flow F':>9}")
for t in range(1, N):
flow = cv2.calcOpticalFlowFarneback(frames[t - 1], frames[t], None,
0.5, 3, 15, 3, 5, 1.2, 0)
dx = np.median(flow[..., 0][flow_mask]) if flow_mask.any() else 0.0
dy = np.median(flow[..., 1][flow_mask]) if flow_mask.any() else 0.0
M = np.float32([[1, 0, dx], [0, 1, dy]])
flow_mask = cv2.warpAffine(flow_mask.astype(np.uint8), M, (W, H)).astype(bool)
print(f"{t:5d}{region_j(copy_mask, gts[t]):9.3f}{boundary_f(copy_mask, gts[t]):9.3f}"
f"{region_j(flow_mask, gts[t]):9.3f}{boundary_f(flow_mask, gts[t]):9.3f}")frame copy J copy F flow J flow F
1 0.702 0.631 0.807 0.952
2 0.501 0.324 0.698 0.676
3 0.359 0.209 0.653 0.516
4 0.241 0.162 0.692 0.613
5 0.137 0.134 0.791 0.810
6 0.054 0.162 0.885 0.961
7 0.000 0.179 0.834 0.959
8 0.000 0.000 0.682 0.720
9 0.000 0.000 0.619 0.605
10 0.000 0.000 0.643 0.643
11 0.000 0.000 0.702 0.688What these two baselines prove
Copying the first mask collapses to zero by frame 7. The object has moved off it entirely. This is the "do nothing" baseline, and any real method must beat it by a wide margin. Publishing a number without this comparison hides how much of the score is from the object barely moving.
Flow propagation tracks position and never fixes shape. The flow J column hovers between 0.6 and 0.9 and does not collapse, because the mask is being carried along by the measured motion. But it never returns to 1.0 either.
The reason is in the scene definition: the disc also grows and shrinks. A translation-only warp cannot represent that. So the mask keeps the wrong size, and J oscillates in step with the radius.
That oscillation is the argument for everything that came after. Propagating a mask by motion is a good start and a permanent ceiling. You need something that re-decides which pixels belong to the object using its appearance, not only where the object went.
J and F disagree, and both matter. At frame 6, flow J is 0.885 while flow F is 0.961. The region overlap is imperfect but the boundary lines up well. At frame 3 the reverse holds. A method can look strong on one and weak on the other, which is why DAVIS reports their mean rather than either alone.
How real systems work
Modern methods replace the flow warp with a memory read. The pattern, established by STM (Oh et al., 2019) and refined by XMem and Cutie:
- Encode each past frame together with its mask into key-value pairs, stored in a memory bank.
- Encode the current frame into queries.
- For each query position, attend over all memory keys and read back a weighted mix of the values.
- Decode the result into a mask for this frame.
- Add this frame and its predicted mask to the memory.
This is cross-attention over time, and step 5 is the drift risk. You are storing your own predictions as if they were ground truth. Every design in this family has a policy limiting what enters memory and for how long, because unbounded memory grows without limit and unfiltered memory poisons itself.
Running one of these needs a GPU and a checkpoint download of several hundred megabytes. Nothing about it is CPU-friendly, and this page will not pretend otherwise. The mechanism is what transfers; the code is a checkout of the authors' repository.
The current practical answer
Since 2024 the general-purpose route is Meta's Segment Anything family. SAM 2 (Meta, August 2024) extended promptable segmentation to video with a memory attention module and a memory bank holding both past frames and object pointers. SAM 3, released 19 November 2025, added promptable concept segmentation: give a noun phrase like "yellow school bus" and it returns masks and consistent IDs for every matching object across the video, where SAM 1 and 2 handled one object per prompt.
For a working system today, that family is the sensible starting point. It is zero-shot, so it needs no training data of your own, and the Ultralytics and Hugging Face wrappers make it a few lines. It also needs a GPU for anything longer than a short clip, and the licence is Meta's own rather than a standard open-source one, so check it before shipping.
Common mistakes
Evaluating with J alone. A blobby mask in roughly the right place scores well on J and poorly on F. Report both.
Feeding your own predicted masks back into memory without any filter. This is how drift becomes collapse. Keep the first frame's ground-truth mask permanently in memory, and be selective about what else you store.
Assuming an object that vanishes is gone. Occlusion is temporary; deletion is permanent. A system that deletes on the first missing frame cannot recover, and re-appearance after occlusion is a large fraction of real cases.
Testing on DAVIS 2016 and claiming a general result. It has one object per video, short sequences, and little occlusion. MOSE and LVOS were built specifically because DAVIS was saturated and unrepresentative.
Using a fixed threshold on the mask probability across a whole dataset. Object scale varies enormously. Tune the threshold per sequence, or use the boundary metric to choose it — model evaluation covers why a single global threshold rarely holds.
Try it yourself
Add a third baseline that also rescales the mask, using the ratio of flow divergence inside the mask to estimate size change. Compare its J against the translation-only version. Then add an occlusion by blanking the disc for two frames, and watch what both baselines do.
What to learn next
- Tracking by detection — the same identity problem, solved with boxes instead of masks.
- Image segmentation — the single-frame version of this task.
- Optical flow — the motion signal both baselines above depend on.
Researcher — Mathematics and papers.
Task definitions
Semi-supervised (one-shot) VOS: given ground-truth masks for frame 0, predict masks for all remaining frames. The dominant setting, and the one DAVIS 2017 and YouTube-VOS evaluate.
Unsupervised VOS: no initialisation; the method must discover and segment the primary object or objects. Evaluated with a bipartite matching between predicted and ground-truth objects before scoring.
Interactive VOS: a human provides scribbles, is shown the result, and refines. Scored as accuracy against interaction budget, which is the metric SAM 2 reports gains on.
Metrics
DAVIS (Perazzi et al., CVPR 2016) defines:
Region similarity $\mathcal{J}$ is the Jaccard index between predicted mask $M$ and ground truth $G$:
$$ \mathcal{J} = \frac{|M \cap G|}{|M \cup G|} $$
Contour accuracy $\mathcal{F}$ is the F-measure between the boundary point sets $c(M)$ and $c(G)$:
$$ \mathcal{F} = \frac{2 P_c R_c}{P_c + R_c} $$
Where $P_c$ is the fraction of predicted boundary points within a tolerance of a ground-truth boundary point, and $R_c$ the converse. The tolerance in the official implementation is a fixed fraction of the image diagonal, computed via a bipartite matching rather than the dilation approximation used in the developer section.
Decay $\mathcal{D}$, less often reported, measures the drop in $\mathcal{J}$ between the first and last quartiles of the sequence. It is the direct measurement of drift and is the more honest number for long-sequence claims.
The headline is $\mathcal{J}\&\mathcal{F}$, the mean of the two.
Architectural lineage
| Year | Method | Idea |
|---|---|---|
| 2017 | OSVOS (Caelles et al.) | Fine-tune a network on frame 0 at test time. Accurate, seconds per sequence |
| 2018 | OSMN, RGMP | Mask propagation without test-time training |
| 2019 | STM (Oh et al.) | Space-time memory: cross-attention over all past frames and masks |
| 2021 | STCN (Cheng et al.) | Affinity computed between frames, not frame-mask pairs; simpler and faster |
| 2022 | XMem (Cheng and Schwing) | Atkinson-Shiffrin-inspired three-tier memory: sensory, working, long-term |
| 2024 | Cutie (Cheng et al.) | Object-level memory reading with query-based object transformer |
| 2024 | SAM 2 (Meta) | Promptable video segmentation, streaming memory, large-scale data engine |
| 2025 | SAM 3 (Meta, November) | Promptable concept segmentation: text prompt returns all instances with IDs |
XMem is the important architectural step for long video. STM's memory grows linearly with sequence length, so a ten-minute video exhausts GPU memory. XMem introduces a bounded working memory plus a compact long-term store with a consolidation policy, making hour-scale sequences feasible. That is a systems contribution as much as a modelling one, and it is why LVOS-style long-video benchmarks became measurable.
SAM 2's contribution is partly the data engine. Its memory attention is stacked transformer blocks conditioning current-frame features on a memory bank of past frames and object pointers, which is architecturally in the STM lineage. The step change came from training on a far larger interactively-collected video mask corpus.
Benchmarks, and why they kept changing
- DAVIS 2016: 50 sequences, single object, short. Saturated by 2019.
- DAVIS 2017: multi-object, 150 sequences. Still the default headline benchmark.
- YouTube-VOS: ~4,000 videos, includes unseen categories at test time, so it measures generalisation rather than memorisation.
- MOSE (2023): built for complex scenes with heavy occlusion and crowded similar objects. Scores drop sharply relative to DAVIS, which is the point.
- LVOS: long videos, minutes rather than seconds. Exposes drift and memory growth.
The pattern is consistent: each benchmark saturates, and the successor isolates a failure mode the previous one hid. Reporting only DAVIS 2017 in 2026 is reporting on the easy case.
Open problems
Drift under self-training. Every memory-based method stores its own predictions. There is no principled uncertainty estimate on a stored mask, so bad memories are weighted like good ones. Confidence-gated memory writing helps and is heuristic.
Identity through long occlusion. Appearance memory decays and motion priors expire. This is the same problem as re-identification in multi-object tracking, and the two literatures have converged slowly — see re-identification embeddings.
Fine structure. Hair, wires, transparent objects, and thin limbs remain poorly segmented, and $\mathcal{F}$ is where it shows. Boundary quality has improved far less than region overlap over the last decade.
Evaluation cost. J&F averages over frames, so a method that is perfect for 90% of a sequence and catastrophic for 10% can outscore one that is consistently good. Per-sequence worst-case reporting is rare and would be more informative.
References
- Perazzi et al., A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation (DAVIS), CVPR 2016.
- Caelles et al., One-Shot Video Object Segmentation (OSVOS), CVPR 2017 — arxiv.org/abs/1611.05198
- Oh et al., Video Object Segmentation using Space-Time Memory Networks (STM), ICCV 2019 — arxiv.org/abs/1904.00607
- Cheng and Schwing, XMem, ECCV 2022 — arxiv.org/abs/2207.07115
- Ravi et al., SAM 2: Segment Anything in Images and Videos, 2024 — arxiv.org/abs/2408.00714
- Meta AI, SAM 3: Segment Anything with Concepts, November 2025 — ai.meta.com/blog/segment-anything-model-3/
What to learn next
- Tracking by detection — the same identity problem, solved with boxes instead of masks.
- Image segmentation — the single-frame version of this task.
- Optical flow — the motion signal both baselines above depend on.