RTSP and live camera pipelines
A live camera never waits for you, so a pipeline that queues every frame falls further behind for ever; the fix is to drop frames on purpose.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A live camera keeps sending frames whether you are ready or not. Try to process all of them and you fall further behind every second.
The analogy you have lived
You are watching a cricket match on a laggy stream. Your neighbours are watching the same match on television.
You hear them cheer. Four seconds later, your screen shows the six.
Nothing is broken. Your stream is behind, and it stays behind, because the match does not slow down to let you catch up. Every delayed second is a second you never get back.
A camera feeding a model behaves exactly this way, and the gap grows.
Why the gap grows instead of settling
Say a camera sends thirty frames a second, and your model needs sixty milliseconds per frame. You can handle about sixteen. Fourteen more arrive every second than you can process.
If you put them in a queue, the queue grows by fourteen every second. After a minute it holds eight hundred frames. The frame you are analysing is nearly a minute old.
The model is not slow in any meaningful sense; it finishes each frame in sixty milliseconds. The system is unusable anyway, because it is answering questions about the past.
camera: ■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 30 per second, never stops
model: ■ ■ ■ ■ ■ ■ 16 per second
queue: grows by 14 every second, for everThe fix is to throw frames away
This feels wrong the first time. It is the only correct answer.
Keep exactly one frame: the newest. When a new one arrives, discard whatever was waiting. When the model is free, hand it the freshest frame available.
Now you still process about sixteen frames a second. But every frame you process is a few milliseconds old rather than a minute old. Same throughput, completely different product.
For a security camera, a queue means an intruder alert arriving after they have left. Dropping means you might miss a frame and will notice within a fraction of a second.
The second thing that goes wrong
Network video connections break. A cable is knocked, a router restarts, the camera loses power for a moment.
When that happens the reading function does not raise a helpful error. It returns nothing, or blocks for a long time, or keeps returning the last frame it saw for ever.
A pipeline that runs for months has to expect this. Notice that frames have stopped, close the connection, wait, and open it again. This reconnection code is more of the work than the model. It separates a demonstration from a system.
Where you have seen this
- Shop and society cameras feeding a screen at a guard's desk.
- Traffic junction cameras counting vehicles.
- Factory line cameras checking parts as they pass.
- A video doorbell that tells you someone is at the gate.
The doorbell makes the point plainly. A notification thirty seconds late is worse than no notification, because you will walk to the gate for nobody.
The honest part
Deciding what to drop is a product decision, not an engineering one.
Dropping the oldest frames is right for "what is happening now". It is wrong for "count everything that passed", where a missed frame is a missed person.
If you truly need every frame, you cannot process live. You must record and process afterwards, and accept the delay. Pretending otherwise builds a system that silently loses data.
Remember this
- A live camera never waits. An unbounded queue means growing delay, for ever.
- Keep the newest frame and throw the rest away, unless you genuinely need all of them.
- Reconnection handling is most of the real work.
What to learn next
- Vision on Jetson and DeepStream — the hardware pipeline for many streams at once.
- OpenCV — the library doing the capture and decode here.
- Real-time audio pipelines — the same backlog problem, in a different medium.
Developer — Code and libraries.
Setup
pip install numpy opencv-pythonThe simulation below needs no camera and no network, and it reproduces the exact failure that a real RTSP stream produces. Timings are wall-clock and will vary between runs and machines; the pattern is what transfers.
The failure and the fix, side by side
import threading, queue, time
import numpy as np
FPS = 30
DURATION = 3.0 # seconds of "live" video
INFER_MS = 60 # our model needs 60 ms per frame: slower than the camera
def camera(out_q, stop, drop_old):
"""Stands in for an RTSP stream: one frame every 1/FPS seconds, no waiting."""
n = 0
t0 = time.perf_counter()
while not stop.is_set():
n += 1
due = t0 + n / FPS
while time.perf_counter() < due:
time.sleep(0.001)
frame = (n, time.perf_counter()) # frame id and capture time
if drop_old:
try:
out_q.get_nowait() # throw away the stale frame
except queue.Empty:
pass
out_q.put(frame)
def run(drop_old, maxsize):
q = queue.Queue(maxsize=maxsize)
stop = threading.Event()
t = threading.Thread(target=camera, args=(q, stop, drop_old), daemon=True)
t.start()
end = time.perf_counter() + DURATION
lags, seen = [], []
while time.perf_counter() < end:
try:
fid, captured = q.get(timeout=0.5)
except queue.Empty:
break
time.sleep(INFER_MS / 1000) # pretend to run the model
lags.append((time.perf_counter() - captured) * 1000)
seen.append(fid)
stop.set(); t.join(timeout=1)
return lags, seen, q.qsize()
for label, drop, size in (("queue everything (maxsize=0)", False, 0),
("keep only the newest frame", True, 1)):
lags, seen, left = run(drop, size)
produced = int(DURATION * FPS)
print(f"--- {label} ---")
print(f" frames the camera sent : about {produced}")
print(f" frames we processed : {len(seen)}")
print(f" frames skipped : {produced - len(seen)}")
print(f" still waiting in the queue at the end: {left}")
print(f" age of the frame when the model saw it:")
print(f" first {lags[0]:7.0f} ms middle {lags[len(lags)//2]:7.0f} ms"
f" last {lags[-1]:7.0f} ms\n")--- queue everything (maxsize=0) ---
frames the camera sent : about 90
frames we processed : 50
frames skipped : 40
still waiting in the queue at the end: 42
age of the frame when the model saw it:
first 60 ms middle 737 ms last 1391 ms
--- keep only the newest frame ---
frames the camera sent : about 90
frames we processed : 50
frames skipped : 40
still waiting in the queue at the end: 1
age of the frame when the model saw it:
first 61 ms middle 71 ms last 88 msBoth strategies did the same amount of work
Fifty frames processed either way. The model's speed did not change and neither did the number of frames it got through. Any claim that dropping frames "loses throughput" is contradicted by these two lines.
The latency is a different universe. Queueing everything: 60 ms, then 737, then 1,391, still climbing when the clip ended. Dropping: 61, 71, 88 ms — flat.
Extrapolate the first one honestly. Three seconds of video produced 1.4 seconds of lag. Thirty seconds would produce about 14 seconds. An hour would produce close to half an hour. There is no equilibrium, because the arrival rate exceeds the service rate by a constant amount and the queue integrates the difference.
Forty-two frames were still sitting in the queue. They represent work already paid for — decoded, copied, held in memory — that will never be useful. That memory grows without bound too, and long-running camera services that get killed by the out-of-memory killer overnight are almost always this bug.
The same thing with a real RTSP stream
OpenCV's VideoCapture has an internal buffer that behaves like the unbounded queue above. The standard fix is a reader thread that never stops consuming.
import cv2, threading, time
class FreshestFrame:
"""Reads the stream continuously, keeps only the newest decoded frame."""
def __init__(self, url):
self.url, self.frame, self.lock = url, None, threading.Lock()
self.stopped = False
threading.Thread(target=self._loop, daemon=True).start()
def _loop(self):
cap = None
while not self.stopped:
if cap is None or not cap.isOpened():
cap = cv2.VideoCapture(self.url, cv2.CAP_FFMPEG)
cap.set(cv2.CAP_PROP_BUFFERSIZE, 1) # honoured by some backends only
if not cap.isOpened():
time.sleep(2.0) # back off, then retry
continue
ok, f = cap.read()
if not ok: # stream dropped: reconnect
cap.release()
cap = None
continue
with self.lock:
self.frame = f # overwrite, never queue
def read(self):
with self.lock:
return None if self.frame is None else self.frame.copy()
def stop(self):
self.stopped = True
# cam = FreshestFrame("rtsp://user:pass@192.168.1.50:554/stream1")
# while True:
# frame = cam.read()
# if frame is not None:
# run_model(frame)No output block: this needs a live RTSP camera, which this machine does not have, and inventing a frame count would be a claim you cannot check. Everything in it is straightforward to verify against your own camera in a few minutes.
Two details that matter. CAP_PROP_BUFFERSIZE is advisory — the FFmpeg backend often ignores it, which is exactly why the reader thread exists rather than relying on the property. And self.frame.copy() is not optional: without it, the consumer is handed a buffer the reader thread may overwrite mid-inference.
When the model is the slow part, use grab and retrieve
If you want to skip decoding entirely for frames you will not use, grab reads and discards without decoding, and retrieve decodes the last grabbed frame.
for _ in range(skip):
cap.grab() # pull the frame off the wire, do not decode it
ok, frame = cap.retrieve()Decoding is a large share of the cost per frame, so skipping it for discarded frames is real saving. This is the cheap version of what a hardware-accelerated pipeline does properly.
What breaks in production, in order of frequency
The stream stops and nothing notices. cap.read() returns False, or blocks. Add a watchdog: if no new frame has arrived for a few seconds, tear the capture down and rebuild it. The reader loop above does the first half; a real one needs a timestamp check too.
TCP versus UDP. RTSP over UDP is lower latency and drops packets under load, producing smeared or torn frames. Over TCP it retransmits, so frames are clean and latency grows. Force one deliberately:
import os
os.environ["OPENCV_FFMPEG_CAPTURE_OPTIONS"] = "rtsp_transport;tcp"That variable must be set before VideoCapture is constructed.
Waiting for a keyframe. H.264 streams only decode from a keyframe onward. Cameras send one every one to four seconds, so a fresh connection shows nothing for up to several seconds. That is not a bug, and if your reconnection logic treats it as one you will loop for ever.
Credentials in the URL. rtsp://user:pass@host/... puts the password into every log line and every crash report. Read it from the environment and construct the URL at use.
Clock skew. The camera's timestamp and your machine's clock are not the same clock. Use your own arrival time for latency measurements, as the simulation does, unless the stream carries a synchronised timestamp you trust.
Common mistakes
Calling cv2.VideoCapture.read() in the same loop as inference. The buffer fills while the model runs, and you have rebuilt the growing queue inside the library where you cannot see it.
Unbounded queues anywhere in the pipeline. Every stage boundary needs either a bound with a drop policy or explicit backpressure. Memory growth is the symptom.
Reconnecting instantly in a tight loop. A camera that is genuinely down will be hammered. Back off, and cap the delay.
Assuming the frame rate the camera advertises. Cameras drop their own frame rate under poor light because exposure gets longer. Measure the arrival rate; do not read it from a spec sheet.
Using cv2.imshow in a production service. It needs a display, it is slow, and it will block your reader thread.
Try it yourself
Set INFER_MS to 20, below the frame interval of 33 ms, and rerun. Both strategies should now stay flat, because the service rate finally exceeds the arrival rate. Then set it to 35 — barely over — and watch how slowly but relentlessly the queued version drifts. That slow drift is the version that reaches production, because it looks fine in a two-minute demonstration.
What to learn next
- Vision on Jetson and DeepStream — the hardware pipeline for many streams at once.
- OpenCV — the library doing the capture and decode here.
- Real-time audio pipelines — the same backlog problem, in a different medium.
Researcher — Mathematics and papers.
The queueing argument
Frames arrive at rate $\lambda$ and are served at rate $\mu$. Utilisation is $\rho = \lambda/\mu$.
For $\rho < 1$, an M/M/1 queue has expected waiting time
$$ W = \frac{1}{\mu - \lambda} $$
which is finite but grows without bound as $\rho \to 1$. For $\rho \ge 1$ there is no steady state at all: the backlog at time $t$ grows as $(\lambda - \mu)t$, and latency grows linearly for ever. The measured 60 ms, 737 ms, 1,391 ms over three seconds is that linear growth, with $\lambda = 30$ and $\mu \approx 16.7$ giving $\rho \approx 1.8$.
The critical observation is that video streams have no admission control. In a request-response service, $\lambda$ is bounded by clients that wait for responses; backpressure is automatic. A camera does not wait. So $\rho \ge 1$ is not an overload condition to be handled, it is the normal operating state that must be designed for.
The two valid responses are to make $\mu > \lambda$ (a faster model, or fewer streams per machine), or to reduce the effective $\lambda$ by dropping. There is no third option, and a bounded queue with a drop policy is the second one stated precisely.
Drop policies
| Policy | Keeps | Right for |
|---|---|---|
| Drop-oldest (queue size 1) | freshest frame | real-time reaction; the default |
| Drop-newest | oldest frame | never useful for live video |
| Uniform sampling (every $k$th) | even temporal coverage | offline analytics, motion statistics |
| Keyframe-only | I-frames | very low bandwidth; irregular spacing |
| Content-triggered | frames after a cheap detector fires | most efficient; adds a second failure mode |
Content triggering deserves the caution. A cheap motion or background-subtraction gate can cut the frame rate reaching the model by an order of magnitude, and it also becomes a silent single point of failure: a gate that stops firing produces a system that reports nothing while appearing healthy.
Latency accounting
End-to-end latency on a camera pipeline decomposes as
$$ L = t_{\text{exposure}} + t_{\text{encode}} + t_{\text{network}} + t_{\text{jitter buffer}} + t_{\text{decode}} + t_{\text{queue}} + t_{\text{infer}} + t_{\text{post}} $$
Engineering effort concentrates on $t_{\text{infer}}$ and almost never on the rest, which is usually where the time is. Typical contributions on an IP camera: encoder latency of one to three frame intervals, a jitter buffer of 100 to 500 ms (the largest single term in most deployments and frequently configurable), and decode of a few milliseconds with hardware support.
Measure $L$ directly rather than summing estimates. Point the camera at a millisecond-resolution clock display and photograph the screen showing the processed output next to it. The difference is the true end-to-end latency, including every term you forgot.
Protocols
RTSP (RFC 2326) carries only control — DESCRIBE, SETUP, PLAY, TEARDOWN. Media flows over RTP, with RTCP for statistics. RTP over UDP loses packets and never retransmits, so loss appears as macroblock corruption; RTP interleaved over the RTSP TCP connection retransmits and converts loss into delay. Choose explicitly.
WebRTC uses the same RTP layer with congestion control (GCC) and adaptive bitrate, achieving sub-200 ms glass-to-glass. Increasingly the right choice for new systems.
HLS and DASH are segment-based and target several seconds of latency by design. Low-Latency HLS reduces this to two or three seconds. Neither is a real-time protocol; using them for a reaction-time system is a category error.
ONVIF is the discovery and control standard that tells you a camera's RTSP URL, rather than hard-coding a vendor-specific path.
Hardware pipelines
At more than a few streams per machine, decoding in Python becomes the constraint. The production answer is a pipeline where decoded frames never leave device memory:
- GStreamer with
nvv4l2decoderorvaapih264dec, keeping frames in GPU or VA-API surfaces. - NVIDIA DeepStream, which is a GStreamer pipeline with batching, tracking and inference plugins; see vision on Jetson and DeepStream.
- Intel Media SDK / VPL with OpenVINO for the CPU and integrated-GPU path.
The shared principle is zero-copy: hardware decoder output feeds the inference engine directly. A pipeline that decodes to a NumPy array and copies it back to the device has already lost most of the benefit, and as preprocessing on the GPU measures, that copy is not small.
Reliability
A camera service is a long-running networked process and should be engineered as one:
- Watchdog on frame arrival, not on process liveness. A process that is alive and receiving nothing is the common failure.
- Exponential backoff with a cap on reconnection, plus jitter, so that a site-wide outage does not produce a synchronised reconnect storm.
- Per-stream isolation. One camera's failure must not stall the others; separate threads or processes per stream.
- Metrics that reveal the queue. Export frames received, frames processed, frames dropped, queue depth and frame age. Frame age is the one that tells you whether the system is real-time, and it is the one most often missing.
References
- Schulzrinne et al., Real Time Streaming Protocol (RTSP), RFC 2326
- Schulzrinne et al., RTP: A Transport Protocol for Real-Time Applications, RFC 3550
- ONVIF Profile S specification — streaming and discovery
- OpenCV
VideoCapturedocumentation — docs.opencv.org - GStreamer application development manual — pipeline and latency handling
What to learn next
- Vision on Jetson and DeepStream — the hardware pipeline for many streams at once.
- OpenCV — the library doing the capture and decode here.
- Real-time audio pipelines — the same backlog problem, in a different medium.