Shipping Vision Models

TensorRT for vision models

TensorRT recompiles your network for one exact GPU, fusing layers and picking kernels by measurement, and the engine it produces will not run anywhere else.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have lived
  3. What "faster" actually comes from
  4. What you give up
  5. Where this is used
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

TensorRT rebuilds your trained model into a program tuned for one specific NVIDIA graphics card. That makes it much faster and completely unportable.

The analogy you have lived

A shirt off the rack fits everybody a little and nobody exactly. It has to.

A tailor measures your shoulders, your arms, your chest, and stitches a shirt for you. It fits far better. It also fits only you. Hand it to your cousin and it is the wrong shirt.

Your training framework is the off-the-rack shirt. It has to handle any model, any input size, any hardware. TensorRT is the tailor. It measures your exact model on your exact card and stitches something that fits, and only fits, that combination.

What "faster" actually comes from

Three things, and none of them change what the model computes.

Joining steps together. A network is written as many small steps: a convolution, then adding a number, then a shape function. Running them separately means writing results to memory and reading them back three times. Joined into one step, the numbers stay in fast memory and the trip is made once. Memory traffic, not arithmetic, is what most vision models are limited by.

Choosing the fastest way to do each step. There are many ways to compute a convolution, and which is fastest depends on the sizes involved and the card. TensorRT tries several and times them. It does not reason about it; it measures.

Using smaller numbers. Training uses high-precision numbers. Inference rarely needs them. Halving the size of each number halves the memory traffic and lets the card use hardware built for exactly that.

What you give up

The result is locked to the hardware. Build on one card model and it will not load on another. Build it in one datacentre and the identical card in another datacentre is fine; a different card model is not.

It is locked to the software version too. Upgrade the toolkit and your saved engine stops loading. You must rebuild.

Building takes time. Minutes, sometimes far longer, because it is timing many alternatives. That is fine as a step in your release pipeline. It is not something to do when a request arrives.

Shapes get fixed. You declare in advance what input sizes you will send. Send something outside that range and it fails.

So the workflow has a fixed shape. Build the engine as part of your release. Store it as an artefact tied to one card and one toolkit version. Rebuild whenever either changes.

Where this is used

  • Cameras at toll plazas and traffic junctions reading plates in real time.
  • Warehouse robots that must see and react within a fixed budget.
  • Video analysis pipelines processing dozens of streams on one machine.
  • Medical scanners doing reconstruction while the patient is still there.

The common thread is a hard time limit and a fixed, known machine. That is exactly when tailoring pays.

The honest part

The speed-up you get is not a fixed number. It depends on the model, the batch size, the resolution and the card. It ranges from barely worth doing to several times faster.

Anyone quoting a single multiplier is quoting one measurement of one model on one card. Measure yours.

Remember this

  • TensorRT recompiles your model for one card and one toolkit version.
  • Its speed comes from joining steps, timing alternatives, and using smaller numbers.
  • Build in your release pipeline, never at request time, and rebuild after any upgrade.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch onnx onnxruntime           # this part runs anywhere
# TensorRT itself needs an NVIDIA GPU and a matching CUDA toolkit:
#   pip install tensorrt                      # or the .deb / .tar from NVIDIA

Versions referenced below: TensorRT 11.2.1 (the current release as of August 2026, CUDA 13.3 baseline), torch 2.5.1, onnx 1.22.0.

Stated plainly: the export and verification section below was run and its output is real. The TensorRT build and runtime sections were not run on this machine, which has no TensorRT installed. Their code is written against the current NVIDIA documentation and carries no output block, because an invented timing table would be worse than none.

Step one: export a clean ONNX graph

This runs on CPU and is where most TensorRT problems are actually created.

export.py
import numpy as np
import torch, torch.nn as nn
import onnx, onnxruntime as ort

torch.manual_seed(0)


class TinyBackbone(nn.Module):
    """A stand-in for a real vision backbone: convolutions, norm, pooling, head."""
    def __init__(self):
        super().__init__()
        self.body = nn.Sequential(
            nn.Conv2d(3, 16, 3, stride=2, padding=1), nn.BatchNorm2d(16), nn.ReLU(),
            nn.Conv2d(16, 32, 3, stride=2, padding=1), nn.BatchNorm2d(32), nn.ReLU(),
            nn.Conv2d(32, 64, 3, stride=2, padding=1), nn.BatchNorm2d(64), nn.ReLU(),
            nn.AdaptiveAvgPool2d(1), nn.Flatten())
        self.head = nn.Linear(64, 10)

    def forward(self, x):
        return self.head(self.body(x))


model = TinyBackbone().eval()
dummy = torch.randn(1, 3, 224, 224)

torch.onnx.export(
    model, dummy, "backbone.onnx",
    input_names=["images"], output_names=["logits"],
    # Without this, the engine is locked to batch size 1 forever.
    dynamic_axes={"images": {0: "batch"}, "logits": {0: "batch"}},
    opset_version=17,
)

g = onnx.load("backbone.onnx")
onnx.checker.check_model(g)
print(f"opset {g.opset_import[0].version}, ir_version {g.ir_version}")
print(f"nodes in the graph: {len(g.graph.node)}")
print("op types:", sorted({n.op_type for n in g.graph.node}))
print("input :", g.graph.input[0].name,
      [d.dim_param or d.dim_value for d in g.graph.input[0].type.tensor_type.shape.dim])
print("output:", g.graph.output[0].name,
      [d.dim_param or d.dim_value for d in g.graph.output[0].type.tensor_type.shape.dim])

sess = ort.InferenceSession("backbone.onnx", providers=["CPUExecutionProvider"])
for b in (1, 8):
    x = np.random.default_rng(0).standard_normal((b, 3, 224, 224)).astype(np.float32)
    y = sess.run(None, {"images": x})[0]
    with torch.no_grad():
        y_torch = model(torch.from_numpy(x)).numpy()
    print(f"batch {b}: onnx output {y.shape}, "
          f"largest gap from PyTorch {np.abs(y - y_torch).max():.2e}")
Output
opset 17, ir_version 8
nodes in the graph: 9
op types: ['Conv', 'Flatten', 'Gemm', 'GlobalAveragePool', 'Relu']
input : images ['batch', 3, 224, 224]
output: logits ['batch', 10]
batch 1: onnx output (1, 10), largest gap from PyTorch 2.61e-08
batch 8: onnx output (8, 10), largest gap from PyTorch 2.61e-08

What that output already tells you

The three BatchNorm2d layers are gone. The op list contains Conv, Relu, GlobalAveragePool, Flatten and Gemm — no BatchNormalization. Because the model was in eval() mode, batch norm is a fixed affine transform and the exporter folded it into the preceding convolution's weights and bias. Nine nodes for a twelve-module network. TensorRT will fuse further, joining each Conv with its Relu.

Export in train() mode and this does not happen: you get separate BatchNormalization nodes, using batch statistics, and a model that behaves differently from the one you evaluated. Calling .eval() before export is not a style preference.

The input shape is ['batch', 3, 224, 224]. The string in position zero is what makes the batch dimension dynamic. Without dynamic_axes it would be the literal 1, and every downstream tool — including TensorRT — would hard-code a batch of one. Fixing this later means re-exporting.

Both batch sizes agree with PyTorch to 2.61e-08. That is float32 rounding. Verify this before touching TensorRT. If ONNX and PyTorch already disagree, TensorRT is not your problem.

Step two: build an engine

The fastest path is the command-line tool that ships with TensorRT.

bash
# fp32, fixed shape
trtexec --onnx=backbone.onnx --saveEngine=backbone.engine

# fp16, with a dynamic batch range: minimum, the size to optimise for, maximum
trtexec --onnx=backbone.onnx --saveEngine=backbone_fp16.engine --fp16 \
        --minShapes=images:1x3x224x224 \
        --optShapes=images:8x3x224x224 \
        --maxShapes=images:32x3x224x224

# benchmark an engine you already built
trtexec --loadEngine=backbone_fp16.engine --shapes=images:8x3x224x224

No output block: trtexec prints throughput and latency measured on whatever GPU you run it on, and the numbers are meaningless copied from another machine.

The three shape arguments are the part people get wrong. optShapes is what the builder times its kernel choices against. Set it to the batch size you actually serve. An engine with min=1, opt=1, max=32 will accept a batch of 32 and run it with kernels chosen for a batch of 1.

Step three: run it

trt_infer.py
import numpy as np
import tensorrt as trt
from cuda import cudart          # pip install cuda-python

logger = trt.Logger(trt.Logger.WARNING)
runtime = trt.Runtime(logger)
with open("backbone_fp16.engine", "rb") as f:
    engine = runtime.deserialize_cuda_engine(f.read())
context = engine.create_execution_context()

batch = 8
x = np.random.randn(batch, 3, 224, 224).astype(np.float32)
context.set_input_shape("images", x.shape)
out_shape = tuple(context.get_tensor_shape("logits"))
y = np.empty(out_shape, dtype=np.float32)

_, d_in = cudart.cudaMalloc(x.nbytes)
_, d_out = cudart.cudaMalloc(y.nbytes)
_, stream = cudart.cudaStreamCreate()

cudart.cudaMemcpyAsync(d_in, x.ctypes.data, x.nbytes,
                       cudart.cudaMemcpyKind.cudaMemcpyHostToDevice, stream)
context.set_tensor_address("images", int(d_in))
context.set_tensor_address("logits", int(d_out))
context.execute_async_v3(stream)
cudart.cudaMemcpyAsync(y.ctypes.data, d_out, y.nbytes,
                       cudart.cudaMemcpyKind.cudaMemcpyDeviceToHost, stream)
cudart.cudaStreamSynchronize(stream)
print(y.shape, y[0, :4])

No output block, for the same reason: this machine has no TensorRT installed, and printing a fabricated array would be a lie about something you can check in five minutes on real hardware.

The API names are worth memorising because they changed at TensorRT 10 and older tutorials are everywhere. set_input_shape and set_tensor_address with execute_async_v3 are current. The bindings list with execute_async_v2, which you will find in most blog posts, is the old API.

Verifying an engine is correct

Never assume. FP16 has about three decimal digits of precision, and a layer that overflows its range produces inf silently.

python
ref = onnx_session.run(None, {"images": x})[0]     # your fp32 baseline
trt_out = run_engine(x)
print("max abs diff :", np.abs(trt_out - ref).max())
print("argmax agree :", (trt_out.argmax(1) == ref.argmax(1)).mean())

Run that over a few hundred real images, not random noise. Random input does not exercise the activation ranges your model actually sees, which is precisely where FP16 overflow happens.

Common mistakes

Building the engine on a different GPU from the one that serves. An engine built for one architecture will not load on another, and even within an architecture the timing-based kernel choices are made for the card present at build time. Build in a container that matches production, on a matching card.

Committing engines to version control. They are large, opaque, and invalid after any driver or TensorRT upgrade. Commit the ONNX file and the build command; produce the engine in CI.

Ignoring the first-inference cost. The first call after deserialisation allocates workspace and warms caches. Discard several iterations before timing anything, and warm the engine at service startup rather than on the first user request.

Treating a fallback to a slow implementation as normal. If TensorRT cannot map an op it may fall back or refuse. Run the builder with verbose logging and read which layers ran on which precision. An unexpected fallback is usually one exotic op that can be replaced.

Expecting FP16 to be free. It usually is for vision backbones. It is not for models with large dynamic ranges, and the failure is nan in the output rather than a warning.

Try it yourself

Export the same model without dynamic_axes and compare the printed input shape. Then add a nn.Dropout and export in train() mode instead of eval(), and look at how many nodes the graph gains. Both experiments run on CPU and both teach you something you would otherwise discover at engine-build time.

What to learn next

Researcher — Mathematics and papers.

What the builder is actually doing

TensorRT's build phase is a compiler pass over the network graph followed by an empirical autotuning search.

Graph optimisation. Constant folding, dead-layer elimination, and layer fusion. The dominant fusions for vision workloads are vertical — Conv + BatchNorm + activation into a single kernel — and horizontal, merging layers that consume the same input. The arithmetic saving is small; the saving in global-memory traffic is large. A fused convolution-activation reads the input once and writes the output once, where the unfused pair does two of each.

Batch norm folding is exact in inference mode:

$$ y = \gamma\frac{Wx + b - \mu}{\sqrt{\sigma^2 + \epsilon}} + \beta = W'x + b', \quad W' = \frac{\gamma W}{\sqrt{\sigma^2+\epsilon}}, \quad b' = \frac{\gamma(b - \mu)}{\sqrt{\sigma^2+\epsilon}} + \beta $$

which is why the exported graph above already had no BatchNormalization nodes.

Tactic selection. For each layer TensorRT enumerates candidate implementations — direct, implicit GEMM, Winograd, FFT-based, cuDNN, cuBLAS, and its own kernels — and benchmarks each on the actual device with the actual shapes. This is why builds are slow, why they are not deterministic across runs, and why the resulting plan is valid only for the device it was built on. --timingCacheFile persists the measurements across builds of similar networks.

Precision assignment. With --fp16 or --int8, the builder chooses per layer whether the faster precision is used, based on measured speed and, for INT8, the calibration data. Layers can and do stay in FP32.

Memory planning. Activation tensors are assigned overlapping slices of one workspace using liveness analysis, so peak memory is set by the largest simultaneously-live set rather than the sum of all activations.

Precision, and what it costs

FormatBitsRangeNote
FP3232$\pm 3.4\times10^{38}$baseline
TF3219 stored in 32FP32 range, FP16 mantissaautomatic on Ampere and later for matmul
FP1616$\pm 65504$~3 decimal digits; overflow is the practical risk
BF1616FP32 rangefewer mantissa bits, no overflow; Ampere+
INT88per-tensor or per-channel scaleneeds calibration; see the next lesson
FP88E4M3 / E5M2Hopper and later
FP44NVFP4Blackwell; mostly a large-language-model story so far

For convolutional vision backbones FP16 is close to free in accuracy and is the default recommendation. The failure mode to watch is activation overflow in networks with unbounded activations or large residual sums; BF16 removes that risk at some precision cost.

Dynamic shapes

An engine declares one or more optimisation profiles, each giving min, opt and max for every dynamic dimension. Kernel selection is performed at opt. Consequences:

  • Serving a batch far from opt gets kernels chosen for opt, not for the batch you sent.
  • Multiple profiles are supported; each adds build time and memory.
  • Fully dynamic spatial dimensions cost more than dynamic batch, because more kernel families become viable and fewer assumptions hold.

For variable-resolution vision workloads, a small set of fixed buckets with letterboxing usually beats a wide dynamic range.

Quantisation: implicit is gone

This changed and older material is now wrong. Through TensorRT 8, INT8 was configured by supplying an IInt8EntropyCalibrator2 and letting the builder infer per-tensor scales — implicit quantisation. From TensorRT 10.1 that path is deprecated, superseded by explicit quantisation, in which QuantizeLinear and DequantizeLinear nodes are present in the ONNX graph and carry the scales.

The practical consequence is that quantisation is now decided in your export pipeline, not in the TensorRT builder. NVIDIA's TensorRT Model Optimizer (nvidia-modelopt) is the supported tool for producing those Q/DQ graphs. The next lesson, INT8 calibration for vision models, covers what calibration is doing and demonstrates it on CPU where you can watch it.

Alternatives, and when they win

  • Torch-TensorRT compiles inside PyTorch, falling back to eager for unsupported ops. Best when you want most of the gain without leaving the framework.
  • ONNX Runtime with the TensorRT execution provider partitions the graph, giving TensorRT what it supports and running the rest itself. The most forgiving option for models with exotic ops.
  • torch.compile with mode="max-autotune" performs its own fusion and autotuning and requires no export at all. Weaker than TensorRT on convolutional networks, competitive on transformers, and vastly easier to operate.
  • Triton Inference Server is orthogonal: it serves TensorRT engines with dynamic batching, model versioning and concurrent execution, and is usually what you deploy the engine into.

Measuring honestly

Report, always: the exact GPU, driver, CUDA and TensorRT versions; batch size; input resolution; precision; whether preprocessing is included; and whether the number is throughput or latency. A "3x speed-up" without those is not a measurement.

The two numbers to separate are latency at batch 1 (what a single request feels) and throughput at the largest batch that fits your latency budget (what your capacity plan needs). Optimising one frequently harms the other. And as preprocessing on the GPU shows with measured numbers, a heavily optimised engine often ends up waiting on a JPEG decoder.

References

What to learn next