Shipping Vision Models

Vision on Jetson and DeepStream

A Jetson is one chip where the processor, the graphics unit, the video decoder and the memory all share a single power and memory budget, and DeepStream is the pipeline built to respect that.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have lived
  3. What makes a Jetson different from a desktop
  4. Why a special pipeline exists
  5. What you give up
  6. Where you have seen this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A Jetson is a small computer where everything shares one power budget and one pool of memory. DeepStream is NVIDIA's video pipeline for exactly that.

The analogy you have lived

Think of a small flat with a fixed sanctioned electricity load. One meter, one limit, shared by everything.

The geyser alone is fine. The iron alone is fine. Switch on the geyser, the iron and the air conditioner together and the fuse trips. Nothing was faulty. You exceeded a budget that was always there.

A Jetson works this way. Processor, graphics unit, video decoder and memory all draw on one power budget and one memory pool. On a desktop these feel unlimited. Here, they are a budget you plan around.

What makes a Jetson different from a desktop

One pool of memory, shared. On a desktop the graphics card has its own memory. Moving data between it and main memory costs time. On a Jetson there is one pool. That is a gift, because nothing has to be copied. It is also a trap: a greedy model starves everything else.

Fixed, chosen power modes. You tell the board how much power it may draw. A lower setting runs cooler and quieter and slower. This is a decision you make, and the difference between settings is large.

Dedicated video hardware. Separate circuits for decoding video, so the graphics unit never spends its time unpacking camera streams.

It gets hot and then slows down. Put one in a sealed box on a factory wall, in an Indian summer. It slows itself down to stay safe. Benchmarks from an air-conditioned lab do not survive contact with a real installation.

Why a special pipeline exists

The obvious approach is: read a frame, turn it into an array, hand it to the model. That works, and it wastes most of the board.

Every one of those steps copies the picture. Decoder to main memory. Main memory to your program. Your program to the model. On a device with one shared memory pool and a modest budget, copying is the expense.

DeepStream builds a pipeline where the frame stays in one place. The video hardware decodes it. The resizing hardware shrinks it. The model reads it. The picture never enters your program.

   many cameras
        ↓
   [ decode in dedicated hardware ]
        ↓
   [ resize and gather into a batch ]     ← frame never leaves device memory
        ↓
   [ the model, as a TensorRT engine ]
        ↓
   [ tracking, then your logic ]

Your code plugs in at the end and receives results, not pictures.

What you give up

The pipeline is described in configuration files rather than written in Python. That is a real learning cost. It is smaller than it looks: most projects change a handful of values.

It is NVIDIA hardware only. It moves in step with a particular software release. Mismatched versions are the commonest cause of a board that will not run.

And it is designed for many camera streams. For one camera and a modest frame rate, a plain script is easier and good enough.

Where you have seen this

  • Traffic junctions counting and classifying vehicles across several lanes.
  • Retail stores measuring footfall and queue lengths.
  • Factory lines checking parts on multiple conveyors.
  • Building entrances matching faces against a small list.

The common shape is several cameras, one small box, running for years without anyone visiting it.

The honest part

The figures on the box are best-case. Highest power setting, lowest precision, idealised input.

Your installation sits in a hot enclosure, at a power setting the enclosure allows, running the pipeline you built. Plan for a fraction of the headline number. Measure on the real board, in the real box, before promising anything.

Remember this

  • Everything on a Jetson shares one power budget and one memory pool.
  • The win comes from never copying the frame, which is what DeepStream is built for.
  • Headline performance numbers assume conditions your installation will not have.

What to learn next

Developer — Code and libraries.

Setup

Stated plainly: this lesson has no output blocks. The commands and configuration below need a Jetson, which this machine is not. Everything is written against NVIDIA's current documentation with versions pinned; nothing is presented as an observed result, because inventing a frames-per-second number for hardware I did not run on would be the worst thing this page could do.

Versions, as of August 2026: DeepStream 9.1 requires JetPack 7.2 GA (Jetson Linux 39.2) and supports Jetson Orin and AGX Thor. It bundles CUDA 13.2, TensorRT 10.16.1.7, cuDNN 9.20.0.46, GStreamer 1.24.2 and OpenCV 4.8.0.

Note that TensorRT 10.16 is not the newest standalone TensorRT release, which is 11.2.1. DeepStream pins its own stack. Installing a newer TensorRT alongside it breaks it. This version coupling is the single most common reason a Jetson project stalls on day one.

Know your board first

bash
# which JetPack / Jetson Linux is installed
cat /etc/nv_tegra_release
sudo apt-cache show nvidia-jetpack | head -n 5

# available power modes, and the current one
sudo nvpmodel -q --verbose

# pick the highest mode, then lock clocks to their maximum
sudo nvpmodel -m 0
sudo jetson_clocks

# live view of CPU, GPU, memory, temperature and power draw
sudo tegrastats --interval 1000

tegrastats is the tool you will use most. It prints, per second, per-core CPU load, GR3D_FREQ (GPU utilisation), RAM used out of total, the temperature of each thermal zone, and instantaneous power for each rail.

Three readings to watch:

  • RAM x/y — remember this is shared with the GPU. A model that would fit comfortably in a discrete card's memory can starve the operating system here.
  • GR3D_FREQ — if this sits low while your pipeline is slow, the GPU is waiting for something else, usually decode or memory.
  • Thermal zone temperatures — if they climb toward the limit and your frame rate falls, you are thermally throttled, and the fix is mechanical, not algorithmic.

Build the engine on the board

TensorRT engines are tied to the GPU architecture, the TensorRT version and the driver. An engine built on your desktop will not load on a Jetson. Build it on the target, or on an identical target.

bash
# on the Jetson, with the ONNX you exported earlier
/usr/src/tensorrt/bin/trtexec \
    --onnx=detector.onnx \
    --saveEngine=detector_fp16.engine \
    --fp16 \
    --minShapes=images:1x3x544x960 \
    --optShapes=images:4x3x544x960 \
    --maxShapes=images:8x3x544x960

Set optShapes to the number of camera streams you will actually batch. This is the same argument as in TensorRT for vision models, and it matters more here because you have less headroom to waste.

Builds on a Jetson are slow — the autotuning search is the same work with far less compute. Do it once, in your provisioning step, and cache the result on the device.

The pipeline, as configuration

DeepStream is a GStreamer pipeline described in .txt files. A minimal two-camera setup:

deepstream_app_config.txt
[application]
enable-perf-measurement=1
perf-measurement-interval-sec=5

[source0]
enable=1
type=4                        # 4 = RTSP
uri=rtsp://192.168.1.50:554/stream1
num-sources=1
gpu-id=0
cudadec-memtype=0             # keep decoded frames in device memory

[source1]
enable=1
type=4
uri=rtsp://192.168.1.51:554/stream1
num-sources=1
gpu-id=0
cudadec-memtype=0

[streammux]
batch-size=2                  # must match the number of enabled sources
batched-push-timeout=40000    # microseconds: do not wait longer than one frame
width=960
height=544
live-source=1                 # tells the muxer these streams never pause

[primary-gie]
enable=1
gpu-id=0
batch-size=2
interval=0                    # 0 = infer on every frame; 1 = every other frame
config-file=detector_config.txt

[tracker]
enable=1
tracker-width=960
tracker-height=544
ll-lib-file=/opt/nvidia/deepstream/deepstream/lib/libnvds_nvmultiobjecttracker.so

[sink0]
enable=1
type=1                        # 1 = fakesink; use this when measuring
sync=0                        # never throttle to the display clock
detector_config.txt
[property]
gpu-id=0
net-scale-factor=0.0039215697      # 1/255
model-color-format=0               # 0 = RGB, 1 = BGR
onnx-file=detector.onnx
model-engine-file=detector_fp16.engine
labelfile-path=labels.txt
batch-size=2
network-mode=2                     # 0 = FP32, 1 = INT8, 2 = FP16
num-detected-classes=4
interval=0
gie-unique-id=1
cluster-mode=2                     # 2 = NMS
bash
deepstream-app -c deepstream_app_config.txt

The settings that decide whether it works

batch-size must agree in three places — [streammux], [primary-gie] and the engine's optShapes. Mismatches do not error; they silently serialise your streams and you lose most of the throughput.

sync=0 on the sink. With sync=1 the pipeline paces itself to the display clock, and your measured frame rate is the display's frame rate rather than the pipeline's. This is the most common fake benchmark on Jetson forums.

batched-push-timeout bounds latency. The muxer waits this long to fill a batch. Set it longer than a frame interval and every frame inherits the wait. At 25 frames per second, one interval is 40,000 microseconds, which is the value above.

interval is the cheapest speed knob you have. Setting interval=1 runs the detector on every second frame and lets the tracker carry identities through the gaps. For slow-moving scenes this roughly halves the model cost with little visible difference, and it is worth trying before optimising anything else.

net-scale-factor and model-color-format are your preprocessing contract. They must match training exactly. This is the same trap as train and serve preprocessing mismatch, relocated into an .ini file where nobody reviews it.

cudadec-memtype=0 keeps decoded frames in device memory. Change it and you reintroduce the copy the whole architecture exists to avoid.

Reading results from Python

pyds gives access to the per-frame metadata without ever touching pixels.

python
import pyds

def probe(pad, info, _):
    batch = pyds.gst_buffer_get_nvds_batch_meta(hash(info.get_buffer()))
    frame_node = batch.frame_meta_list
    while frame_node is not None:
        frame = pyds.NvDsFrameMeta.cast(frame_node.data)
        obj_node = frame.obj_meta_list
        while obj_node is not None:
            obj = pyds.NvDsObjectMeta.cast(obj_node.data)
            print(frame.source_id, frame.frame_num, obj.object_id,
                  obj.class_id, round(obj.confidence, 3))
            obj_node = obj_node.next
        frame_node = frame_node.next
    return Gst.PadProbeReturn.OK

obj.object_id is the tracker's persistent identity, which is what you count and measure dwell time with. See tracking metrics for how to tell whether those identities are any good.

The hardware, in numbers

ModuleAI performanceMemory bandwidthMemory
Orin Nano (original)up to 40 TOPS68 GB/s4 or 8 GB
Orin Nano Super67 TOPS102 GB/s8 GB
Orin NXup to 100 TOPS—8 or 16 GB
AGX Orinup to 275 TOPS204 GB/s32 or 64 GB

The Orin Nano Super is worth a note: NVIDIA raised the original Orin Nano's figures from 40 to 67 TOPS purely through a JetPack 6.2 software update and higher clocks, on the same silicon. Existing boards got the increase by reflashing.

Two honest caveats on this table. TOPS figures are INT8 with structured sparsity — the theoretical peak of the tensor cores, not what a real network achieves. And for most vision pipelines the memory bandwidth column predicts real performance better than the TOPS column does, because convolutions at video resolution are memory-bound.

Common mistakes

Mixing versions. Installing a newer TensorRT, or a DeepStream built for a different JetPack. Match the table at the top of this lesson exactly.

Copying an engine from a desktop. It will not load. Build on the target.

Benchmarking with a display sink. Use type=1 (fakesink) with sync=0 when measuring, and only switch to a real sink when you want to look at it.

Testing on the bench and shipping in a box. An open developer kit on a desk and the same module in a sealed IP65 enclosure on a wall are different thermal environments. Measure tegrastats temperatures after an hour in the real enclosure.

Running at the default power mode. Boards frequently ship in a lower mode. Check nvpmodel -q and set it deliberately, then confirm the enclosure can dissipate what that mode draws.

Forgetting the SD card. Developer kits boot from microSD, which is slow and wears out. Production deployments should use NVMe or eMMC.

Reaching for DeepStream for one camera. For a single stream, cv2.VideoCapture plus a TensorRT engine is simpler, easier to debug, and fast enough. DeepStream earns its complexity somewhere around four streams.

Try it yourself

Without a Jetson, the transferable exercise is the version matrix. Take your current model, list every component it depends on — ONNX opset, TensorRT, CUDA, driver — and write down the exact versions. Then check them against the DeepStream 9.1 requirements above. Most Jetson projects fail at this table, before any code runs.

What to learn next

Researcher — Mathematics and papers.

The architecture, and why it changes the maths

A Jetson SoC integrates Arm Cortex-A78AE cores, an Ampere-architecture GPU with tensor cores, dedicated video decode and encode engines (NVDEC/NVENC), an image signal processor, a vision accelerator (PVA) and on the larger modules deep learning accelerators (DLA), all sharing one LPDDR5 memory controller.

Two consequences dominate every design decision.

Memory is unified and coherent. There is no host-to-device copy. cudaHostAlloc with the mapped flag, or the NVMM buffer type GStreamer uses, gives CPU and GPU pointers to the same physical pages. The PCIe transfer term that dominates discrete-GPU pipelines is absent here.

Memory is also the shared bottleneck. Every engine draws from the same bandwidth. An AGX Orin's 204 GB/s is roughly a fifth of a contemporary discrete card's, and it is being divided among decode, preprocessing, inference and the CPU. Arithmetic intensity — FLOPs per byte moved — predicts achieved performance here far better than peak TOPS does.

For a convolution layer with $C_{\text{in}}$, $C_{\text{out}}$, kernel $k$ and output $H \times W$:

$$ \text{arithmetic intensity} \approx \frac{2 H W C_{\text{out}} C_{\text{in}} k^2}{\left(HWC_{\text{in}} + H W C_{\text{out}} + C_{\text{out}}C_{\text{in}}k^2\right)\cdot \text{bytes per element}} $$

Layers below the roofline's ridge point are bandwidth-limited, and on a Jetson the ridge point sits at a much lower intensity than on a discrete card. This is why depthwise-separable architectures, which have low arithmetic intensity, often disappoint on paper-TOPS comparisons: they were designed to reduce FLOPs, and FLOPs were not the constraint.

Zero-copy and the NVMM buffer

GStreamer on Jetson carries video/x-raw(memory:NVMM) capabilities, meaning buffers live in device-accessible memory with a hardware surface handle rather than in ordinary system memory. nvv4l2decoder produces them; nvvideoconvert, nvstreammux and nvinfer consume them.

Inserting any element that requires plain video/x-raw — including an appsink that hands frames to NumPy — forces a copy out of NVMM and back. On a bandwidth-constrained device this single change can dominate the pipeline's cost. It is the specific reason "I rewrote it in Python and it got much slower" is a recurring Jetson forum post.

Power, thermals and honest measurement

nvpmodel selects a power mode: a combination of online CPU cores, and caps on CPU, GPU, memory-controller and DLA clocks. The developer kits ship in a mode below their maximum.

Dynamic voltage and frequency scaling then adjusts within that envelope, and jetson_clocks pins everything to the mode's ceiling — useful for reproducible benchmarking, and generally wrong for deployment, since it removes the thermal headroom that the governor was using.

The measurement protocol that produces numbers you can trust:

  1. Set the power mode the product will ship with, in the enclosure it will ship in.
  2. Run the real pipeline, with real camera streams, for at least thirty minutes.
  3. Log tegrastats throughout and report the frame rate in the final five minutes, not the first.
  4. Report the temperature curve alongside it.

A number from the first minute of a benchtop run is a number about a different machine.

DLA, and its restrictions

AGX Orin includes two deep learning accelerators. They are far more power-efficient than the GPU for supported layers, and they support a restricted operator set — broadly convolutions, pooling, activations and elementwise ops, with constraints on dimensions and precision.

trtexec --useDLACore=0 --allowGPUFallback compiles for DLA with GPU fallback for unsupported layers. The fallback is the catch: each transition between DLA and GPU is a synchronisation point, and a network that falls back repeatedly can be slower than pure GPU execution while using more power. Read the verbose builder log and count the transitions before concluding anything.

Where DeepStream's advantage comes from

Against a naive Python pipeline, DeepStream wins on four axes, roughly in order of magnitude:

  1. No copies out of NVMM. The largest single factor on a bandwidth-limited device.
  2. Cross-stream batching. nvstreammux gathers frames from $N$ cameras into one batch, so the GPU runs one batched inference instead of $N$ small ones. Small-batch GPU inference is dominated by launch overhead.
  3. Hardware decode and hardware resize. NVDEC and the VIC handle work that would otherwise occupy the GPU or the CPU.
  4. A C-level pipeline. No Python interpreter in the per-frame path.

Against a well-built GStreamer pipeline using the same NVIDIA elements, DeepStream's advantage narrows considerably: it is those elements plus batching, tracking, and a metadata model. If you already know GStreamer, building directly on nvv4l2decoder, nvinfer and friends is a legitimate choice with less configuration surface.

Alternatives

  • Holoscan is NVIDIA's newer sensor-processing SDK, aimed at medical and scientific streaming, with a Python-first API and less GStreamer baggage.
  • Triton Inference Server runs on Jetson and gives model management and multi-model serving; it composes with DeepStream via nvinferserver rather than replacing it.
  • Plain TensorRT plus GStreamer for a single stream, as above.
  • ONNX Runtime with the TensorRT execution provider when operator coverage is the problem and you can afford some overhead.

References

  • DeepStream SDK documentation, version 9.1 — docs.nvidia.com/metropolis/deepstream
  • Jetson Linux Developer Guide, power and performance — docs.nvidia.com/jetson
  • NVIDIA, Jetson Orin Nano Developer Kit Gets a Super Boost, 2024 — developer.nvidia.com/blog
  • Williams, Waterman and Patterson, Roofline: An Insightful Visual Performance Model, CACM 2009
  • GStreamer application development manual — pipeline construction and caps negotiation

What to learn next