Video Understanding and Tracking

Reading and writing video frames

A video is a stack of pictures plus a clock, and every video model starts by turning that file back into an array of frames you can index.

On this page 7
  1. Why this is the first thing to learn
  2. How it works
  3. The part that surprises people
  4. Where you have already seen this
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A video is a pile of still pictures shown fast, one after another. Your eye reads them as movement.

Think of the flipbook you drew in the corner of a school notebook. Each page held one drawing of a stick figure. Flick the pages with your thumb and the figure walks.

Nothing on any single page moved. The movement lived in the order of the pages and the speed of your thumb.

A video file is that flipbook. Every page is called a frame, meaning one complete still picture.

Why this is the first thing to learn

Models do not read .mp4 files. Models read numbers.

Every video system on earth starts by turning a file back into a stack of pictures. Get this step wrong and everything after it is wrong too.

There is a second reason. A video file does not store every page separately. That would be enormous.

Instead it stores a few full pages, and for the rest it stores only what changed. This trick is called compression — throwing away repeated information to make the file smaller.

How it works

   video file
   ┌──────────────────────────────────────┐
   │ full page │ changes │ changes │ full │   <- what is stored on disk
   └──────────────────────────────────────┘
                     │
                     ▼
              [ decoder ]  rebuilds every page
                     │
                     ▼
   frame 0   frame 1   frame 2   frame 3  ...  <- what your code receives
   ┌────┐    ┌────┐    ┌────┐    ┌────┐
   │    │    │    │    │    │    │    │        each one a full grid of pixels
   └────┘    └────┘    └────┘    └────┘

The decoder is the piece of software that rebuilds full pages from stored changes. Your code never sees the compressed form. It sees finished pictures.

Alongside the pictures, the file carries a clock. Frame rate means how many pictures are shown each second. Thirty is common for phone video, twenty-four for cinema.

The part that surprises people

Jumping to the middle of a video is not free.

To rebuild page four hundred, the decoder may need to start from the last full page before it. Then it replays every stored change in between. Ask for a thousand random pages and you can wait a long time.

Reading straight through, start to finish, is far faster than hopping around. That single fact shapes how every serious video pipeline is built.

Where you have already seen this

  • Scrubbing a YouTube video: the picture goes blocky while the decoder catches up.
  • A CCTV system storing weeks of footage on one small hard disk.
  • WhatsApp shrinking a video before it sends.
  • A dashcam saving hours of driving onto a memory card.

What is honestly hard here

Video is full of small lies that cost people days.

The frame count written in the file header is a note the writer left behind. It is sometimes wrong. Colours may come out of the decoder in an unexpected order.

Two libraries reading the same file can hand you slightly different pixels. None of this is your mistake. It is what video formats are actually like.

Check what you got rather than trusting what you were told.

Remember this

  • A video is frames plus a frame rate, nothing more mysterious.
  • Files store changes between frames, so a decoder must rebuild each picture.
  • Reading in order is fast. Jumping around is slow, and worth designing away.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install opencv-python numpy

Written and run against opencv-python 4.11.0 and numpy 1.26.4 on Python 3.10. These video I/O calls are long-standing across the 4.x line. OpenCV 5.0 restructured modules but kept every function reachable as cv2.<name>().

Write a clip, then read it back

No download, no camera. We create a video, then take it apart.

video_roundtrip.py
import cv2
import numpy as np

W, H, N_FRAMES, FPS = 64, 48, 30, 10

# Build a clip in memory: a coloured square sliding left to right on black.
frames = []
for i in range(N_FRAMES):
    frame = np.zeros((H, W, 3), dtype=np.uint8)   # OpenCV frames are H x W x 3, uint8
    x = 4 + i                                     # the square moves one pixel per frame
    frame[20:28, x:x + 8] = (0, 0, 255)           # BGR order, so this is RED, not blue
    frames.append(frame)

fourcc = cv2.VideoWriter_fourcc(*"mp4v")
writer = cv2.VideoWriter("slider.mp4", fourcc, FPS, (W, H))   # note: (width, height)
print("writer opened:", writer.isOpened())
for f in frames:
    writer.write(f)
writer.release()

cap = cv2.VideoCapture("slider.mp4")
print("capture opened:", cap.isOpened())
print("width      :", cap.get(cv2.CAP_PROP_FRAME_WIDTH))
print("height     :", cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
print("fps        :", cap.get(cv2.CAP_PROP_FPS))
print("frame count:", cap.get(cv2.CAP_PROP_FRAME_COUNT), "(metadata, not always true)")

read = 0
first = None
while True:
    ok, frame = cap.read()          # ok is False at end of stream OR on a decode error
    if not ok:
        break
    if read == 0:
        first = frame
    read += 1
cap.release()

print("frames actually decoded:", read)
print("first frame shape:", first.shape, "dtype:", first.dtype)
b, g, r = first[24, 8]
print(f"pixel inside the square, BGR order: B={b} G={g} R={r}")
rgb = cv2.cvtColor(first, cv2.COLOR_BGR2RGB)
print("same pixel after BGR2RGB       :", tuple(int(v) for v in rgb[24, 8]))
Output
writer opened: True
capture opened: True
width      : 64.0
height     : 48.0
fps        : 10.0
frame count: 30.0 (metadata, not always true)
frames actually decoded: 30
first frame shape: (48, 64, 3) dtype: uint8
pixel inside the square, BGR order: B=0 G=0 R=252
same pixel after BGR2RGB       : (252, 0, 0)

Read that output line by line

(48, 64, 3) is height, width, channels. The writer wanted (64, 48) — width, height. The two are in opposite orders. Swapping them silently produces a video of the wrong shape, or an empty file. This trips up nearly everyone once.

We wrote 255 and got back 252. The pixel is not the pixel we put in. mp4v is a lossy codec, so it discards detail to save space. The exact value depends on which encoder your OpenCV build shipped with, so yours may differ by a few counts. The point is that it is not 255.

That matters more than it sounds. Save preprocessed frames to a compressed video, reload them, and you have added noise to your dataset. Save arrays, or use a lossless codec, when the pixels themselves are the data.

Colours come out as blue, green, red. OpenCV predates the convention everyone else settled on. Matplotlib, PyTorch, Pillow and every pretrained model expect red, green, blue. A model fed BGR loses accuracy without raising an error, which is the worst kind of bug.

Seeking, and why it is not free

seeking.py
import cv2
import numpy as np

cap = cv2.VideoCapture("slider.mp4")
fps = cap.get(cv2.CAP_PROP_FPS)

def square_centre(frame):
    row = frame[24].astype(float).sum(axis=1)      # brightness along one scanline
    xs = np.arange(row.size)
    return (row * xs).sum() / row.sum()            # centre of mass of the bright blob

for seconds in (0.0, 1.0, 2.0):
    target = int(round(seconds * fps))
    cap.set(cv2.CAP_PROP_POS_FRAMES, target)        # ask the decoder to jump
    landed = int(cap.get(cv2.CAP_PROP_POS_FRAMES))  # where it says it is
    ok, frame = cap.read()
    expected = 4 + target + 3.5                     # we know where we drew it
    print(f"t={seconds}s  asked frame {target}  decoder reports {landed}  "
          f"square centre x={square_centre(frame):.1f} (drawn at {expected})")
cap.release()
Output
t=0.0s  asked frame 0  decoder reports 0  square centre x=7.5 (drawn at 7.5)
t=1.0s  asked frame 10  decoder reports 10  square centre x=17.2 (drawn at 17.5)
t=2.0s  asked frame 20  decoder reports 20  square centre x=27.4 (drawn at 27.5)

The seek landed on the right frames here, and the measured position is within half a pixel of where we drew it. That small gap is compression again, blurring the square's edges.

This clip is thirty frames long, with tiny gaps between full pages. On a real hour-long recording, seeking to an arbitrary frame can land you on a nearby one instead. The decoder stopped at the previous full page. Verify positions when accuracy matters.

Common mistakes

Assuming cap.read() returning False means the video ended. It also returns False on a corrupt frame midway through. Count what you decoded and compare against what you expected, as the first script does.

Trusting CAP_PROP_FRAME_COUNT. It is metadata written by whatever produced the file. For streams and some containers it is zero or plainly wrong. If the exact count matters, decode once and count.

Forgetting writer.release(). Without it the file's index is never written and the video is unplayable. The same applies to cap.release() in long-running processes, which otherwise leak file handles.

Passing the size as (height, width) to VideoWriter. It takes (width, height). The symptom is a valid file containing zero frames, with no error message anywhere.

Feeding BGR frames to a pretrained model. Convert with cv2.cvtColor(frame, cv2.COLOR_BGR2RGB) before normalising. Accuracy drops by a few points and nothing warns you.

Decoding a whole video to find one frame. For dataset building, decode once, keep what you need, and never open the file again in the training loop. Random access during training is a common reason a video model starves the GPU — see is the GPU waiting for data.

Try it yourself

Change the codec to cv2.VideoWriter_fourcc(*"FFV1") and write to slider.avi. FFV1 is lossless. Check whether the pixel now reads back as exactly 255, and compare the two file sizes.

What to learn next

Researcher — Mathematics and papers.

Container, codec, stream

Three layers that are routinely confused:

  • Container (.mp4, .mkv, .avi, .webm): the box. It holds streams, timestamps, indexes and metadata.
  • Codec (H.264, H.265/HEVC, VP9, AV1, FFV1): the algorithm compressing one stream.
  • Stream: the encoded bytes of one track — video, audio, subtitles.

A .mp4 extension says nothing about how the video inside is compressed. Failures blamed on "the mp4 format" are almost always codec support failures in the installed FFmpeg build.

Frame types and the group of pictures

Modern codecs emit three frame kinds:

TypeDepends onEffect
I-frame (keyframe)nothingIndependently decodable, large
P-frameearlier framesSmall, forward prediction only
B-frameearlier and later framesSmallest, requires out-of-order decoding

A group of pictures (GOP) is one I-frame plus the dependent frames that follow it. GOP length in typical delivery encodes is 1 to 10 seconds. Random access therefore costs, on average, half a GOP of wasted decoding.

B-frames force a split between two clocks. DTS (decode timestamp) is the order bytes must be fed to the decoder. PTS (presentation timestamp) is the order frames are displayed. They differ whenever B-frames are present, and code assuming decode order equals display order will reorder your video.

Accurate seeking

Container-level seek granularity is the keyframe. Frame-accurate seeking requires seeking to the preceding keyframe, decoding forward, and discarding output until the target PTS. Libraries differ in whether they do this:

  • OpenCV's CAP_PROP_POS_FRAMES delegates to FFmpeg and is not guaranteed frame-accurate across all containers.
  • PyAV exposes container.seek(..., any_frame=False) plus manual decode-forward, which is accurate but explicit.
  • Decord and NVIDIA DALI maintain their own frame indexes for random access, at the cost of an indexing pass.

For research datasets, the reproducible option is to decode once into frames or a fixed-size tensor store, and never seek during training. Every published video benchmark reporting decoding as a bottleneck is describing runtime seeking.

Variable frame rate

Phone cameras, screen recorders and streaming captures produce variable frame rate (VFR) video: inter-frame intervals are not constant. CAP_PROP_FPS then reports an average, and index-based sampling silently distorts time.

Two consequences for video models:

  1. A "16 frames at stride 4" clip covers a different real duration in different videos.
  2. Optical flow magnitudes, which are per-frame displacements, become incomparable across a dataset.

The standard mitigation is transcoding to constant frame rate before dataset construction, and recording that in the dataset card. Kinetics, Something-Something and AVA all specify a fixed rate for this reason.

Colour, and where accuracy leaks

Consumer video is stored as YUV 4:2:0: luma at full resolution, two chroma planes at half resolution in both axes. Conversion to RGB upsamples chroma and applies a matrix depending on the colour primaries (BT.601 for standard definition, BT.709 for high definition, BT.2020 for wide gamut) and on the range convention (limited range 16–235, or full range 0–255).

Getting the matrix or the range wrong produces a systematic shift in every pixel. It is invisible to the eye and measurable in model accuracy. Two decoders disagreeing on BT.601 versus BT.709 is a documented source of irreproducible video benchmarks.

Decoding cost

For a H x W frame at F frames per second, decoding cost scales roughly with H · W · F and with codec complexity. AV1 decode is materially more expensive than H.264 at equal resolution, which matters when a data loader must sustain hundreds of clips per second.

Practical routes when decode becomes the bottleneck:

  • Hardware decode: NVDEC through DALI or PyAV, moving decode off the CPU entirely.
  • Pre-extraction: frames to JPEG, or short clips to .npy / .pt shards. Costs disk, buys determinism.
  • Reduced-resolution decode: many codecs decode to a lower resolution directly, avoiding full-size reconstruction.
  • Compressed-domain models: Wu et al. (2018), Compressed Video Action Recognition, operate on I-frames plus stored motion vectors and residuals, skipping full decode. Accuracy is competitive on Kinetics-class tasks at a large speed advantage, and the approach remains under-used.

Libraries, honestly compared

LibraryStrengthWeakness
OpenCV VideoCaptureEverywhere, one lineSeek accuracy varies; no PTS access
PyAVDirect FFmpeg bindings, full PTS/DTS controlVerbose; you handle frame types yourself
torchvision read_videoTensors, PyTorch-nativeHistorically slow; API has changed across releases
DecordFast random access via its own indexSmaller maintainer base
NVIDIA DALIGPU decode plus augmentation in one graphNVIDIA hardware only, heavier dependency

Pin whichever you choose and record the version in your dataset card. Decoder behaviour is a hyperparameter, and treating it as one prevents a whole class of unreproducible result.

References

  • Wiegand et al., Overview of the H.264/AVC Video Coding Standard, IEEE TCSVT, 2003.
  • Sullivan et al., Overview of the High Efficiency Video Coding (HEVC) Standard, IEEE TCSVT, 2012.
  • Wu et al., Compressed Video Action Recognition, CVPR 2018 — arxiv.org/abs/1712.00636
  • OpenCV video I/O reference — docs.opencv.org

What to learn next

What to learn next

These follow on from what you just read.

  • Video Understanding and Tracking

    Frame sampling strategies

    A model can only look at a handful of frames, so choosing which handful is a design decision that quietly sets the ceiling on your accuracy.

  • Video Understanding and Tracking

    Optical flow

    Optical flow measures how far every pixel moved between two frames, giving you motion as a picture instead of guessing it from appearance.

  • Video Understanding and Tracking

    3D convolutions for video

    A 3D convolution slides a small cube through a stack of frames, so one filter can respond to movement instead of only to appearance.