Video Understanding and Tracking

Video object segmentation

You outline an object in the first frame and the system must keep that exact outline on it for the rest of the video, through turns, blur and disappearances.

Read these first

On this page 7
  1. Why it exists
  2. How it works
  3. Two flavours of the task
  4. Where you have already seen this
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Video object segmentation means drawing an outline around one object in the first frame. Then you keep that outline on it for every frame after.

Think of colouring inside the lines in a flipbook you drew yourself. On page one you carefully shade the cat. On page two the cat has moved a little, so you shade it there.

By page fifty the cat has turned around, walked behind a chair and come out the other side. You are still shading the same cat, because you have been following it.

Everything about the task is in that word "following". You are not finding a cat on each page independently. You are keeping hold of this cat.

Why it exists

Ordinary image segmentation, covered in image segmentation, labels every pixel in one picture. Run it on each frame separately and two problems appear at once.

It has no memory. Frame ten has no idea what frame nine decided. Outlines flicker and jump.

It has no identity. With two cats in the scene, nothing links the cat in frame ten to the one in frame nine. You wanted one cat and got a fresh guess every frame.

Video object segmentation fixes both by carrying information forward.

How it works

   frame 0            frame 1            frame 2           frame 3
   ┌────────┐         ┌────────┐         ┌────────┐        ┌────────┐
   │  (▓▓)  │         │   (▓▓) │         │    (▓▓)│        │  ????? │
   └────────┘         └────────┘         └────────┘        └────────┘
   you draw           system            system             object went
   this outline       predicts          predicts           behind a pole
        │                 │                 │                   │
        └── remembered ───┴── remembered ───┴─── remembered ────┘
             what the object looked like, all the way along

The system keeps a memory of how the object looked in the frames it has already handled. For each new frame it asks: which pixels here look like the thing I have been remembering?

That memory is the whole design. The naive alternative is to push the outline forward using motion. That works for a few frames and then drifts away, as the code below shows.

Two flavours of the task

Semi-supervised means you are given the first-frame outline. The system only has to follow it. This is the common setting and the one described above.

Unsupervised means nothing is given. The system must decide by itself which object is the interesting one, then follow it. Harder, and less used in practice, because usually you know what you want to track.

Where you have already seen this

  • Video calls blurring your background while keeping your outline crisp.
  • Phone editors that let you change the colour of one person's shirt across a whole clip.
  • Sports graphics that highlight one player as they run.
  • Film effects that cut an actor out of a shot, once done by hand, frame by frame, for weeks.

What is honestly hard here

Three things break these systems, and none is fully solved.

Disappearing and coming back. An object passes behind a pillar. When it reappears, the system must know it is the same object and not a new one.

Drift. Small errors accumulate. Include a few background pixels this frame. Use them as your memory next frame. The outline slowly swallows the background.

Similar objects. Two people in identical uniforms standing close together. Appearance memory helps far less when the appearances match.

Read that list twice before assuming this task is easy. Every published system fails on all three. The numbers you see quoted are averages over cases where it works.

Remember this

  • You give a first-frame outline; the system keeps it on the same object throughout.
  • The key mechanism is memory of how the object looked earlier.
  • Drift, occlusion and lookalikes are the failure modes, and they are unsolved.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install opencv-python numpy

Run against opencv-python 4.11.0 and numpy 1.26.4. We build the evaluation metrics and two baselines from scratch, because understanding what J&F measures is worth more than calling a library that prints it.

The metrics, implemented

DAVIS, the standard benchmark, scores with two numbers averaged together.

J is region similarity: intersection over union between the predicted and true masks. F is boundary accuracy: an F1 score between the two outlines, allowing a small tolerance so a one-pixel offset is not punished. The headline number is J&F, their mean.

vos_baselines.py
import cv2
import numpy as np

H = W = 96
N = 12
rng = np.random.default_rng(0)
texture = rng.integers(0, 90, (H, W), dtype=np.uint8)     # a static, textured background


def scene(t):
    """Frame t and its true mask: a bright disc drifting, growing and shrinking."""
    cx, cy = 22 + 2.7 * t, 38 + 1.3 * t
    r = 11 + 2.5 * np.sin(t / 2)
    yy, xx = np.mgrid[0:H, 0:W]
    mask = ((xx - cx) ** 2 + (yy - cy) ** 2) <= r * r
    frame = texture.copy()
    frame[mask] = 230
    return cv2.GaussianBlur(frame, (3, 3), 0.8), mask


def region_j(pred, gt):
    """DAVIS J: intersection over union of the two masks."""
    inter = np.logical_and(pred, gt).sum()
    union = np.logical_or(pred, gt).sum()
    return 1.0 if union == 0 else inter / union


def boundary_f(pred, gt, tol_frac=0.008):
    """DAVIS F: F1 between the two outlines, allowing a small tolerance."""
    def outline(m):
        m8 = m.astype(np.uint8)
        grad = cv2.morphologyEx(m8, cv2.MORPH_GRADIENT, np.ones((3, 3), np.uint8))
        return grad.astype(bool)

    tol = max(1, int(round(tol_frac * np.hypot(H, W))))
    k = np.ones((2 * tol + 1, 2 * tol + 1), np.uint8)
    bp, bg = outline(pred), outline(gt)
    bp_d = cv2.dilate(bp.astype(np.uint8), k).astype(bool)
    bg_d = cv2.dilate(bg.astype(np.uint8), k).astype(bool)
    prec = (bp & bg_d).sum() / max(bp.sum(), 1)
    rec = (bg & bp_d).sum() / max(bg.sum(), 1)
    return 0.0 if prec + rec == 0 else 2 * prec * rec / (prec + rec)


frames, gts = zip(*[scene(t) for t in range(N)])
copy_mask = gts[0].copy()          # baseline 1: never update the mask
flow_mask = gts[0].copy()          # baseline 2: push the mask along the optical flow

print(f"{'frame':>5}{'copy J':>9}{'copy F':>9}{'flow J':>9}{'flow F':>9}")
for t in range(1, N):
    flow = cv2.calcOpticalFlowFarneback(frames[t - 1], frames[t], None,
                                        0.5, 3, 15, 3, 5, 1.2, 0)
    dx = np.median(flow[..., 0][flow_mask]) if flow_mask.any() else 0.0
    dy = np.median(flow[..., 1][flow_mask]) if flow_mask.any() else 0.0
    M = np.float32([[1, 0, dx], [0, 1, dy]])
    flow_mask = cv2.warpAffine(flow_mask.astype(np.uint8), M, (W, H)).astype(bool)
    print(f"{t:5d}{region_j(copy_mask, gts[t]):9.3f}{boundary_f(copy_mask, gts[t]):9.3f}"
          f"{region_j(flow_mask, gts[t]):9.3f}{boundary_f(flow_mask, gts[t]):9.3f}")
Output
frame   copy J   copy F   flow J   flow F
    1    0.702    0.631    0.807    0.952
    2    0.501    0.324    0.698    0.676
    3    0.359    0.209    0.653    0.516
    4    0.241    0.162    0.692    0.613
    5    0.137    0.134    0.791    0.810
    6    0.054    0.162    0.885    0.961
    7    0.000    0.179    0.834    0.959
    8    0.000    0.000    0.682    0.720
    9    0.000    0.000    0.619    0.605
   10    0.000    0.000    0.643    0.643
   11    0.000    0.000    0.702    0.688

What these two baselines prove

Copying the first mask collapses to zero by frame 7. The object has moved off it entirely. This is the "do nothing" baseline, and any real method must beat it by a wide margin. Publishing a number without this comparison hides how much of the score is from the object barely moving.

Flow propagation tracks position and never fixes shape. The flow J column hovers between 0.6 and 0.9 and does not collapse, because the mask is being carried along by the measured motion. But it never returns to 1.0 either.

The reason is in the scene definition: the disc also grows and shrinks. A translation-only warp cannot represent that. So the mask keeps the wrong size, and J oscillates in step with the radius.

That oscillation is the argument for everything that came after. Propagating a mask by motion is a good start and a permanent ceiling. You need something that re-decides which pixels belong to the object using its appearance, not only where the object went.

J and F disagree, and both matter. At frame 6, flow J is 0.885 while flow F is 0.961. The region overlap is imperfect but the boundary lines up well. At frame 3 the reverse holds. A method can look strong on one and weak on the other, which is why DAVIS reports their mean rather than either alone.

How real systems work

Modern methods replace the flow warp with a memory read. The pattern, established by STM (Oh et al., 2019) and refined by XMem and Cutie:

  1. Encode each past frame together with its mask into key-value pairs, stored in a memory bank.
  2. Encode the current frame into queries.
  3. For each query position, attend over all memory keys and read back a weighted mix of the values.
  4. Decode the result into a mask for this frame.
  5. Add this frame and its predicted mask to the memory.

This is cross-attention over time, and step 5 is the drift risk. You are storing your own predictions as if they were ground truth. Every design in this family has a policy limiting what enters memory and for how long, because unbounded memory grows without limit and unfiltered memory poisons itself.

Running one of these needs a GPU and a checkpoint download of several hundred megabytes. Nothing about it is CPU-friendly, and this page will not pretend otherwise. The mechanism is what transfers; the code is a checkout of the authors' repository.

The current practical answer

Since 2024 the general-purpose route is Meta's Segment Anything family. SAM 2 (Meta, August 2024) extended promptable segmentation to video with a memory attention module and a memory bank holding both past frames and object pointers. SAM 3, released 19 November 2025, added promptable concept segmentation: give a noun phrase like "yellow school bus" and it returns masks and consistent IDs for every matching object across the video, where SAM 1 and 2 handled one object per prompt.

For a working system today, that family is the sensible starting point. It is zero-shot, so it needs no training data of your own, and the Ultralytics and Hugging Face wrappers make it a few lines. It also needs a GPU for anything longer than a short clip, and the licence is Meta's own rather than a standard open-source one, so check it before shipping.

Common mistakes

Evaluating with J alone. A blobby mask in roughly the right place scores well on J and poorly on F. Report both.

Feeding your own predicted masks back into memory without any filter. This is how drift becomes collapse. Keep the first frame's ground-truth mask permanently in memory, and be selective about what else you store.

Assuming an object that vanishes is gone. Occlusion is temporary; deletion is permanent. A system that deletes on the first missing frame cannot recover, and re-appearance after occlusion is a large fraction of real cases.

Testing on DAVIS 2016 and claiming a general result. It has one object per video, short sequences, and little occlusion. MOSE and LVOS were built specifically because DAVIS was saturated and unrepresentative.

Using a fixed threshold on the mask probability across a whole dataset. Object scale varies enormously. Tune the threshold per sequence, or use the boundary metric to choose it — model evaluation covers why a single global threshold rarely holds.

Try it yourself

Add a third baseline that also rescales the mask, using the ratio of flow divergence inside the mask to estimate size change. Compare its J against the translation-only version. Then add an occlusion by blanking the disc for two frames, and watch what both baselines do.

What to learn next

Researcher — Mathematics and papers.

Task definitions

Semi-supervised (one-shot) VOS: given ground-truth masks for frame 0, predict masks for all remaining frames. The dominant setting, and the one DAVIS 2017 and YouTube-VOS evaluate.

Unsupervised VOS: no initialisation; the method must discover and segment the primary object or objects. Evaluated with a bipartite matching between predicted and ground-truth objects before scoring.

Interactive VOS: a human provides scribbles, is shown the result, and refines. Scored as accuracy against interaction budget, which is the metric SAM 2 reports gains on.

Metrics

DAVIS (Perazzi et al., CVPR 2016) defines:

Region similarity $\mathcal{J}$ is the Jaccard index between predicted mask $M$ and ground truth $G$:

$$ \mathcal{J} = \frac{|M \cap G|}{|M \cup G|} $$

Contour accuracy $\mathcal{F}$ is the F-measure between the boundary point sets $c(M)$ and $c(G)$:

$$ \mathcal{F} = \frac{2 P_c R_c}{P_c + R_c} $$

Where $P_c$ is the fraction of predicted boundary points within a tolerance of a ground-truth boundary point, and $R_c$ the converse. The tolerance in the official implementation is a fixed fraction of the image diagonal, computed via a bipartite matching rather than the dilation approximation used in the developer section.

Decay $\mathcal{D}$, less often reported, measures the drop in $\mathcal{J}$ between the first and last quartiles of the sequence. It is the direct measurement of drift and is the more honest number for long-sequence claims.

The headline is $\mathcal{J}\&\mathcal{F}$, the mean of the two.

Architectural lineage

YearMethodIdea
2017OSVOS (Caelles et al.)Fine-tune a network on frame 0 at test time. Accurate, seconds per sequence
2018OSMN, RGMPMask propagation without test-time training
2019STM (Oh et al.)Space-time memory: cross-attention over all past frames and masks
2021STCN (Cheng et al.)Affinity computed between frames, not frame-mask pairs; simpler and faster
2022XMem (Cheng and Schwing)Atkinson-Shiffrin-inspired three-tier memory: sensory, working, long-term
2024Cutie (Cheng et al.)Object-level memory reading with query-based object transformer
2024SAM 2 (Meta)Promptable video segmentation, streaming memory, large-scale data engine
2025SAM 3 (Meta, November)Promptable concept segmentation: text prompt returns all instances with IDs

XMem is the important architectural step for long video. STM's memory grows linearly with sequence length, so a ten-minute video exhausts GPU memory. XMem introduces a bounded working memory plus a compact long-term store with a consolidation policy, making hour-scale sequences feasible. That is a systems contribution as much as a modelling one, and it is why LVOS-style long-video benchmarks became measurable.

SAM 2's contribution is partly the data engine. Its memory attention is stacked transformer blocks conditioning current-frame features on a memory bank of past frames and object pointers, which is architecturally in the STM lineage. The step change came from training on a far larger interactively-collected video mask corpus.

Benchmarks, and why they kept changing

  • DAVIS 2016: 50 sequences, single object, short. Saturated by 2019.
  • DAVIS 2017: multi-object, 150 sequences. Still the default headline benchmark.
  • YouTube-VOS: ~4,000 videos, includes unseen categories at test time, so it measures generalisation rather than memorisation.
  • MOSE (2023): built for complex scenes with heavy occlusion and crowded similar objects. Scores drop sharply relative to DAVIS, which is the point.
  • LVOS: long videos, minutes rather than seconds. Exposes drift and memory growth.

The pattern is consistent: each benchmark saturates, and the successor isolates a failure mode the previous one hid. Reporting only DAVIS 2017 in 2026 is reporting on the easy case.

Open problems

Drift under self-training. Every memory-based method stores its own predictions. There is no principled uncertainty estimate on a stored mask, so bad memories are weighted like good ones. Confidence-gated memory writing helps and is heuristic.

Identity through long occlusion. Appearance memory decays and motion priors expire. This is the same problem as re-identification in multi-object tracking, and the two literatures have converged slowly — see re-identification embeddings.

Fine structure. Hair, wires, transparent objects, and thin limbs remain poorly segmented, and $\mathcal{F}$ is where it shows. Boundary quality has improved far less than region overlap over the last decade.

Evaluation cost. J&F averages over frames, so a method that is perfect for 90% of a sequence and catastrophic for 10% can outscore one that is consistently good. Per-sequence worst-case reporting is rare and would be more informative.

References

What to learn next