Video Understanding and Tracking
Reading and writing video frames
A video is a stack of pictures plus a clock, and every video model starts by turning that file back into an array of frames you can index.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A video is a pile of still pictures shown fast, one after another. Your eye reads them as movement.
Think of the flipbook you drew in the corner of a school notebook. Each page held one drawing of a stick figure. Flick the pages with your thumb and the figure walks.
Nothing on any single page moved. The movement lived in the order of the pages and the speed of your thumb.
A video file is that flipbook. Every page is called a frame, meaning one complete still picture.
Why this is the first thing to learn
Models do not read .mp4 files. Models read numbers.
Every video system on earth starts by turning a file back into a stack of pictures. Get this step wrong and everything after it is wrong too.
There is a second reason. A video file does not store every page separately. That would be enormous.
Instead it stores a few full pages, and for the rest it stores only what changed. This trick is called compression — throwing away repeated information to make the file smaller.
How it works
video file
┌──────────────────────────────────────┐
│ full page │ changes │ changes │ full │ <- what is stored on disk
└──────────────────────────────────────┘
│
▼
[ decoder ] rebuilds every page
│
▼
frame 0 frame 1 frame 2 frame 3 ... <- what your code receives
┌────┐ ┌────┐ ┌────┐ ┌────┐
│ │ │ │ │ │ │ │ each one a full grid of pixels
└────┘ └────┘ └────┘ └────┘The decoder is the piece of software that rebuilds full pages from stored changes. Your code never sees the compressed form. It sees finished pictures.
Alongside the pictures, the file carries a clock. Frame rate means how many pictures are shown each second. Thirty is common for phone video, twenty-four for cinema.
The part that surprises people
Jumping to the middle of a video is not free.
To rebuild page four hundred, the decoder may need to start from the last full page before it. Then it replays every stored change in between. Ask for a thousand random pages and you can wait a long time.
Reading straight through, start to finish, is far faster than hopping around. That single fact shapes how every serious video pipeline is built.
Where you have already seen this
- Scrubbing a YouTube video: the picture goes blocky while the decoder catches up.
- A CCTV system storing weeks of footage on one small hard disk.
- WhatsApp shrinking a video before it sends.
- A dashcam saving hours of driving onto a memory card.
What is honestly hard here
Video is full of small lies that cost people days.
The frame count written in the file header is a note the writer left behind. It is sometimes wrong. Colours may come out of the decoder in an unexpected order.
Two libraries reading the same file can hand you slightly different pixels. None of this is your mistake. It is what video formats are actually like.
Check what you got rather than trusting what you were told.
Remember this
- A video is frames plus a frame rate, nothing more mysterious.
- Files store changes between frames, so a decoder must rebuild each picture.
- Reading in order is fast. Jumping around is slow, and worth designing away.
What to learn next
- Frame sampling strategies — choosing which frames a model actually sees.
- How images are stored — the pixel grid every frame turns into.
- OpenCV — the rest of the toolbox around these calls.
Developer — Code and libraries.
Setup
pip install opencv-python numpyWritten and run against opencv-python 4.11.0 and numpy 1.26.4 on Python 3.10. These video I/O calls are long-standing across the 4.x line. OpenCV 5.0 restructured modules but kept every function reachable as cv2.<name>().
Write a clip, then read it back
No download, no camera. We create a video, then take it apart.
import cv2
import numpy as np
W, H, N_FRAMES, FPS = 64, 48, 30, 10
# Build a clip in memory: a coloured square sliding left to right on black.
frames = []
for i in range(N_FRAMES):
frame = np.zeros((H, W, 3), dtype=np.uint8) # OpenCV frames are H x W x 3, uint8
x = 4 + i # the square moves one pixel per frame
frame[20:28, x:x + 8] = (0, 0, 255) # BGR order, so this is RED, not blue
frames.append(frame)
fourcc = cv2.VideoWriter_fourcc(*"mp4v")
writer = cv2.VideoWriter("slider.mp4", fourcc, FPS, (W, H)) # note: (width, height)
print("writer opened:", writer.isOpened())
for f in frames:
writer.write(f)
writer.release()
cap = cv2.VideoCapture("slider.mp4")
print("capture opened:", cap.isOpened())
print("width :", cap.get(cv2.CAP_PROP_FRAME_WIDTH))
print("height :", cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
print("fps :", cap.get(cv2.CAP_PROP_FPS))
print("frame count:", cap.get(cv2.CAP_PROP_FRAME_COUNT), "(metadata, not always true)")
read = 0
first = None
while True:
ok, frame = cap.read() # ok is False at end of stream OR on a decode error
if not ok:
break
if read == 0:
first = frame
read += 1
cap.release()
print("frames actually decoded:", read)
print("first frame shape:", first.shape, "dtype:", first.dtype)
b, g, r = first[24, 8]
print(f"pixel inside the square, BGR order: B={b} G={g} R={r}")
rgb = cv2.cvtColor(first, cv2.COLOR_BGR2RGB)
print("same pixel after BGR2RGB :", tuple(int(v) for v in rgb[24, 8]))writer opened: True capture opened: True width : 64.0 height : 48.0 fps : 10.0 frame count: 30.0 (metadata, not always true) frames actually decoded: 30 first frame shape: (48, 64, 3) dtype: uint8 pixel inside the square, BGR order: B=0 G=0 R=252 same pixel after BGR2RGB : (252, 0, 0)
Read that output line by line
(48, 64, 3) is height, width, channels. The writer wanted (64, 48) — width, height. The two are in opposite orders. Swapping them silently produces a video of the wrong shape, or an empty file. This trips up nearly everyone once.
We wrote 255 and got back 252. The pixel is not the pixel we put in. mp4v is a lossy codec, so it discards detail to save space. The exact value depends on which encoder your OpenCV build shipped with, so yours may differ by a few counts. The point is that it is not 255.
That matters more than it sounds. Save preprocessed frames to a compressed video, reload them, and you have added noise to your dataset. Save arrays, or use a lossless codec, when the pixels themselves are the data.
Colours come out as blue, green, red. OpenCV predates the convention everyone else settled on. Matplotlib, PyTorch, Pillow and every pretrained model expect red, green, blue. A model fed BGR loses accuracy without raising an error, which is the worst kind of bug.
Seeking, and why it is not free
import cv2
import numpy as np
cap = cv2.VideoCapture("slider.mp4")
fps = cap.get(cv2.CAP_PROP_FPS)
def square_centre(frame):
row = frame[24].astype(float).sum(axis=1) # brightness along one scanline
xs = np.arange(row.size)
return (row * xs).sum() / row.sum() # centre of mass of the bright blob
for seconds in (0.0, 1.0, 2.0):
target = int(round(seconds * fps))
cap.set(cv2.CAP_PROP_POS_FRAMES, target) # ask the decoder to jump
landed = int(cap.get(cv2.CAP_PROP_POS_FRAMES)) # where it says it is
ok, frame = cap.read()
expected = 4 + target + 3.5 # we know where we drew it
print(f"t={seconds}s asked frame {target} decoder reports {landed} "
f"square centre x={square_centre(frame):.1f} (drawn at {expected})")
cap.release()t=0.0s asked frame 0 decoder reports 0 square centre x=7.5 (drawn at 7.5) t=1.0s asked frame 10 decoder reports 10 square centre x=17.2 (drawn at 17.5) t=2.0s asked frame 20 decoder reports 20 square centre x=27.4 (drawn at 27.5)
The seek landed on the right frames here, and the measured position is within half a pixel of where we drew it. That small gap is compression again, blurring the square's edges.
This clip is thirty frames long, with tiny gaps between full pages. On a real hour-long recording, seeking to an arbitrary frame can land you on a nearby one instead. The decoder stopped at the previous full page. Verify positions when accuracy matters.
Common mistakes
Assuming cap.read() returning False means the video ended. It also returns False on a corrupt frame midway through. Count what you decoded and compare against what you expected, as the first script does.
Trusting CAP_PROP_FRAME_COUNT. It is metadata written by whatever produced the file. For streams and some containers it is zero or plainly wrong. If the exact count matters, decode once and count.
Forgetting writer.release(). Without it the file's index is never written and the video is unplayable. The same applies to cap.release() in long-running processes, which otherwise leak file handles.
Passing the size as (height, width) to VideoWriter. It takes (width, height). The symptom is a valid file containing zero frames, with no error message anywhere.
Feeding BGR frames to a pretrained model. Convert with cv2.cvtColor(frame, cv2.COLOR_BGR2RGB) before normalising. Accuracy drops by a few points and nothing warns you.
Decoding a whole video to find one frame. For dataset building, decode once, keep what you need, and never open the file again in the training loop. Random access during training is a common reason a video model starves the GPU — see is the GPU waiting for data.
Try it yourself
Change the codec to cv2.VideoWriter_fourcc(*"FFV1") and write to slider.avi. FFV1 is lossless. Check whether the pixel now reads back as exactly 255, and compare the two file sizes.
What to learn next
- Frame sampling strategies — choosing which frames a model actually sees.
- How images are stored — the pixel grid every frame turns into.
- OpenCV — the rest of the toolbox around these calls.
Researcher — Mathematics and papers.
Container, codec, stream
Three layers that are routinely confused:
- Container (
.mp4,.mkv,.avi,.webm): the box. It holds streams, timestamps, indexes and metadata. - Codec (H.264, H.265/HEVC, VP9, AV1, FFV1): the algorithm compressing one stream.
- Stream: the encoded bytes of one track — video, audio, subtitles.
A .mp4 extension says nothing about how the video inside is compressed. Failures blamed on "the mp4 format" are almost always codec support failures in the installed FFmpeg build.
Frame types and the group of pictures
Modern codecs emit three frame kinds:
| Type | Depends on | Effect |
|---|---|---|
| I-frame (keyframe) | nothing | Independently decodable, large |
| P-frame | earlier frames | Small, forward prediction only |
| B-frame | earlier and later frames | Smallest, requires out-of-order decoding |
A group of pictures (GOP) is one I-frame plus the dependent frames that follow it. GOP length in typical delivery encodes is 1 to 10 seconds. Random access therefore costs, on average, half a GOP of wasted decoding.
B-frames force a split between two clocks. DTS (decode timestamp) is the order bytes must be fed to the decoder. PTS (presentation timestamp) is the order frames are displayed. They differ whenever B-frames are present, and code assuming decode order equals display order will reorder your video.
Accurate seeking
Container-level seek granularity is the keyframe. Frame-accurate seeking requires seeking to the preceding keyframe, decoding forward, and discarding output until the target PTS. Libraries differ in whether they do this:
- OpenCV's
CAP_PROP_POS_FRAMESdelegates to FFmpeg and is not guaranteed frame-accurate across all containers. - PyAV exposes
container.seek(..., any_frame=False)plus manual decode-forward, which is accurate but explicit. - Decord and NVIDIA DALI maintain their own frame indexes for random access, at the cost of an indexing pass.
For research datasets, the reproducible option is to decode once into frames or a fixed-size tensor store, and never seek during training. Every published video benchmark reporting decoding as a bottleneck is describing runtime seeking.
Variable frame rate
Phone cameras, screen recorders and streaming captures produce variable frame rate (VFR) video: inter-frame intervals are not constant. CAP_PROP_FPS then reports an average, and index-based sampling silently distorts time.
Two consequences for video models:
- A "16 frames at stride 4" clip covers a different real duration in different videos.
- Optical flow magnitudes, which are per-frame displacements, become incomparable across a dataset.
The standard mitigation is transcoding to constant frame rate before dataset construction, and recording that in the dataset card. Kinetics, Something-Something and AVA all specify a fixed rate for this reason.
Colour, and where accuracy leaks
Consumer video is stored as YUV 4:2:0: luma at full resolution, two chroma planes at half resolution in both axes. Conversion to RGB upsamples chroma and applies a matrix depending on the colour primaries (BT.601 for standard definition, BT.709 for high definition, BT.2020 for wide gamut) and on the range convention (limited range 16–235, or full range 0–255).
Getting the matrix or the range wrong produces a systematic shift in every pixel. It is invisible to the eye and measurable in model accuracy. Two decoders disagreeing on BT.601 versus BT.709 is a documented source of irreproducible video benchmarks.
Decoding cost
For a H x W frame at F frames per second, decoding cost scales roughly with H · W · F and with codec complexity. AV1 decode is materially more expensive than H.264 at equal resolution, which matters when a data loader must sustain hundreds of clips per second.
Practical routes when decode becomes the bottleneck:
- Hardware decode: NVDEC through DALI or PyAV, moving decode off the CPU entirely.
- Pre-extraction: frames to JPEG, or short clips to
.npy/.ptshards. Costs disk, buys determinism. - Reduced-resolution decode: many codecs decode to a lower resolution directly, avoiding full-size reconstruction.
- Compressed-domain models: Wu et al. (2018), Compressed Video Action Recognition, operate on I-frames plus stored motion vectors and residuals, skipping full decode. Accuracy is competitive on Kinetics-class tasks at a large speed advantage, and the approach remains under-used.
Libraries, honestly compared
| Library | Strength | Weakness |
|---|---|---|
OpenCV VideoCapture | Everywhere, one line | Seek accuracy varies; no PTS access |
| PyAV | Direct FFmpeg bindings, full PTS/DTS control | Verbose; you handle frame types yourself |
torchvision read_video | Tensors, PyTorch-native | Historically slow; API has changed across releases |
| Decord | Fast random access via its own index | Smaller maintainer base |
| NVIDIA DALI | GPU decode plus augmentation in one graph | NVIDIA hardware only, heavier dependency |
Pin whichever you choose and record the version in your dataset card. Decoder behaviour is a hyperparameter, and treating it as one prevents a whole class of unreproducible result.
References
- Wiegand et al., Overview of the H.264/AVC Video Coding Standard, IEEE TCSVT, 2003.
- Sullivan et al., Overview of the High Efficiency Video Coding (HEVC) Standard, IEEE TCSVT, 2012.
- Wu et al., Compressed Video Action Recognition, CVPR 2018 — arxiv.org/abs/1712.00636
- OpenCV video I/O reference — docs.opencv.org
What to learn next
- Frame sampling strategies — choosing which frames a model actually sees.
- How images are stored — the pixel grid every frame turns into.
- OpenCV — the rest of the toolbox around these calls.