Shipping Vision Models

Preprocessing on the GPU

Once the model is fast, decoding and resizing images on the CPU becomes the whole cost, and moving that work onto the GPU is usually the largest speed-up left.

On this page 10
  1. The short answer
  2. The analogy you have lived
  3. Why this creeps up on people
  4. What is actually slow
  5. The other benefit
  6. The catch you must know about
  7. Where this is used
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Once your model is fast, the slow part is decoding and resizing the pictures. That work can move onto the graphics card too.

The analogy you have lived

Picture a wedding kitchen. The main cook is superb. Give him a tray of chopped vegetables and the dish is ready in half a minute.

But there is one person chopping, in a room down the corridor, with one knife. The cook stands at his stove doing nothing for twenty minutes, then cooks for thirty seconds, then waits again.

Buying a better stove changes nothing. The bottleneck is the chopping, and the fix is to get more knives into the kitchen.

Your graphics card is the cook. Image decoding is the chopping.

Why this creeps up on people

It happens in a specific order, and the order is what makes it surprising.

At the start, the model is slow and everything else looks like rounding error. So you optimise the model. You make it smaller, you compile it, you drop to lower precision. It gets ten times faster.

Now the picture preparation, which never changed, is ten times more important than it was. It was five percent of your time; now it is most of it.

Nothing got worse. You removed the other cost and revealed this one.

What is actually slow

Three steps sit between a file and the model.

Decoding. A JPEG is a compressed file, not pixels. Turning it into pixels is real work, done one image at a time, on one processor core.

Resizing. Two million pixels in, fifty thousand out, for every image.

Scaling the numbers. Cheap on its own, but it touches every pixel again.

The graphics card can do all three, in parallel, across many images at once. Modern NVIDIA cards also have dedicated hardware for JPEG decoding. It does not compete with the model.

The other benefit

There is a second win that is easy to miss. Decode the pictures on the graphics card and they never cross the link as full-size images.

You send the small compressed file instead. A one-megabyte file crossing, rather than six megabytes of pixels. Less traffic on the slowest link in the machine.

   CPU path:  file → [decode on CPU] → big pixel block → across the link → card
   GPU path:  file → across the link → [decode on card] → pixels already there

The catch you must know about

The graphics card resizes slightly differently from the usual image libraries. Not wrongly — differently.

That is exactly the mismatch problem from earlier in this section. Move preparation to the card and you have changed the recipe. Check that the model still sees what it learned from.

Where this is used

  • Video pipelines handling many camera streams on one machine.
  • Training runs where the card sits idle waiting for the data loader.
  • Batch jobs scoring millions of stored photos.
  • Any service where the model was optimised and got no faster.

The honest part

This is worth doing only when preparation is actually your bottleneck. Measure first.

Suppose your model takes half a second and decoding takes five milliseconds. Moving the decoding buys one percent. It costs a week and a new class of bug.

Remember this

  • Optimising the model makes preparation a bigger share of the total, not a smaller one.
  • Decoding and resizing on the card is parallel, and uses dedicated hardware.
  • The card resizes differently, so re-check your numbers against training.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch torchvision pillow numpy

Versions used: torch 2.5.1+cu121, torchvision 0.20.1+cu121, Pillow 11.0.0. This section needs an NVIDIA GPU. The measurements below were taken on an RTX A6000 with an Intel Core i9-13900K. Timings are hardware-dependent and vary between runs on the same machine; across four runs here the ratio was between 6.6 and 7.4 times, while the absolute millisecond figures moved by nearly a factor of two with background load. Treat the ratio as the finding and the milliseconds as an illustration.

Both pipelines, timed properly

gpu_preprocess.py
import io, time, statistics
import numpy as np
import torch
import torch.nn.functional as F
from PIL import Image
from torchvision.io import decode_jpeg

assert torch.cuda.is_available(), "this script needs an NVIDIA GPU"
dev = torch.device("cuda")
print(torch.cuda.get_device_name(0), "| torch", torch.__version__)

# 64 fake 1080p photos, held as encoded JPEG bytes, exactly like a camera sends.
rng = np.random.default_rng(0)
yy, xx = np.mgrid[0:1080, 0:1920]
base = (128 + 100 * np.sin(xx / 37) * np.cos(yy / 41)).astype(np.uint8)
BLOBS = []
for i in range(64):
    a = np.stack([np.roll(base, i * 7, 1), np.roll(base, i * 3, 0), base], -1)
    buf = io.BytesIO()
    Image.fromarray(a).save(buf, "JPEG", quality=85)
    BLOBS.append(buf.getvalue())

MEAN = torch.tensor([0.485, 0.456, 0.406], device=dev).view(1, 3, 1, 1)
STD = torch.tensor([0.229, 0.224, 0.225], device=dev).view(1, 3, 1, 1)


def cpu_pipeline():
    out = []
    for b in BLOBS:
        im = Image.open(io.BytesIO(b)).convert("RGB").resize((224, 224), Image.BILINEAR)
        out.append(torch.from_numpy(np.asarray(im).copy()))
    t = torch.stack(out).permute(0, 3, 1, 2).to(dev, non_blocking=True).float() / 255
    return (t - MEAN) / STD


def gpu_pipeline():
    raw = [torch.frombuffer(bytearray(b), dtype=torch.uint8) for b in BLOBS]
    imgs = decode_jpeg(raw, device=dev)              # nvJPEG, on the GPU
    t = F.interpolate(torch.stack(imgs).float(), size=(224, 224),
                      mode="bilinear", antialias=True, align_corners=False) / 255
    return (t - MEAN) / STD


def bench(fn, n=10):
    fn(); torch.cuda.synchronize()                   # warm up, then time properly
    ts = []
    for _ in range(n):
        torch.cuda.synchronize()
        t0 = time.perf_counter()
        fn()
        torch.cuda.synchronize()
        ts.append((time.perf_counter() - t0) * 1000)
    return statistics.median(ts)


a, b = cpu_pipeline(), gpu_pipeline()
print(f"\nboth pipelines produce {tuple(a.shape)} on {a.device}")
print(f"largest disagreement between them: {(a - b).abs().max().item():.4f}")

c_ms, g_ms = bench(cpu_pipeline), bench(gpu_pipeline)
print(f"\nCPU decode + resize for 64 images : {c_ms:7.1f} ms  "
      f"({64_000/c_ms:6.0f} images/second)")
print(f"GPU decode + resize for 64 images : {g_ms:7.1f} ms  "
      f"({64_000/g_ms:6.0f} images/second)")
print(f"speed-up: {c_ms/g_ms:.1f}x")

model = torch.nn.Sequential(
    torch.nn.Conv2d(3, 32, 3, stride=2, padding=1), torch.nn.ReLU(),
    torch.nn.Conv2d(32, 64, 3, stride=2, padding=1), torch.nn.ReLU(),
    torch.nn.AdaptiveAvgPool2d(1), torch.nn.Flatten(), torch.nn.Linear(64, 10),
).to(dev).eval()

with torch.no_grad():
    m_ms = bench(lambda: model(a))
print(f"\nthe model forward pass itself : {m_ms:7.1f} ms")
print(f"preprocessing share of the total, CPU path: {c_ms/(c_ms+m_ms):.1%}")
print(f"preprocessing share of the total, GPU path: {g_ms/(g_ms+m_ms):.1%}")
Output
NVIDIA RTX A6000 | torch 2.5.1+cu121

both pipelines produce (64, 3, 224, 224) on cuda:0
largest disagreement between them: 0.0637

CPU decode + resize for 64 images :   642.8 ms  (   100 images/second)
GPU decode + resize for 64 images :    95.1 ms  (   673 images/second)
speed-up: 6.8x

the model forward pass itself :     2.2 ms
preprocessing share of the total, CPU path: 99.7%
preprocessing share of the total, GPU path: 97.7%

The number that should stop you

The model takes 2.2 milliseconds. Preprocessing takes 643. Preprocessing is 99.7 percent of the work on the CPU path, and still 97.7 percent after moving it to the GPU.

Read that again in terms of effort. A team could spend a month getting this model from 2.2 ms to 1.1 ms and improve end-to-end throughput by about 0.2 percent. The same team could spend an afternoon on the decode path and get nearly seven times.

This model is small, so the ratio is extreme. On a full ResNet-50 at batch 64 the forward pass is tens of milliseconds rather than two, and preprocessing is still a large share. The direction of the result generalises even where the size of it does not.

Moving to the GPU gave 6.8 times here. Two mechanisms, roughly equally: nvJPEG decodes on dedicated hardware in parallel across the batch, and the resize runs on thousands of cores instead of one. A third benefit does not show in this timing at all — only the compressed bytes cross the PCIe link, not the decoded 1080p frames.

And there is the disagreement: 0.0637. The two pipelines produce measurably different tensors. That is 16 grey levels out of 255, on a normalised tensor. It comes from Pillow's resampling filter differing from F.interpolate with antialias=True, and from decode differences between libjpeg-turbo and nvJPEG. Move preprocessing to the GPU and you have changed the model's input distribution. Re-run your accuracy evaluation. This is exactly the failure described in train and serve preprocessing mismatch, arriving through the back door of an optimisation.

How to time GPU code without lying to yourself

Three rules, all visible in bench above.

Warm up first. The first call allocates memory, loads kernels and initialises the nvJPEG context. Timing it measures startup, not steady state.

Synchronise before and after. CUDA calls are asynchronous. Without torch.cuda.synchronize() you are timing how long it takes to queue the work, which is close to zero and completely meaningless.

Report the median, not the mean. One scheduling hiccup ruins a mean. Better still, report the median and the 95th percentile, because tail latency is what your users experience.

Overlap the copy with the compute

Decoding on the GPU is one of two wins. The other is never letting the GPU idle while data moves.

python
stream = torch.cuda.Stream()
for batch in loader:
    with torch.cuda.stream(stream):
        gpu_batch = batch.to("cuda", non_blocking=True)   # needs pinned memory
    torch.cuda.current_stream().wait_stream(stream)
    output = model(gpu_batch)

non_blocking=True is a no-op unless the source tensor is in pinned (page-locked) memory. In PyTorch that means DataLoader(..., pin_memory=True), or .pin_memory() on the tensor. Set the flag without pinning and you get a synchronous copy plus the illusion that you optimised something.

The heavier tools

For sustained high-throughput pipelines, NVIDIA DALI is the production answer. It builds a preprocessing graph — decode, crop, resize, normalise, augment — and executes it on the GPU with its own threading, plugging into PyTorch, TensorFlow and Triton.

bash
pip install --extra-index-url https://pypi.nvidia.com nvidia-dali-cuda120

No output block: DALI is not installed on this machine, and a fabricated throughput table would be worse than none. What is worth knowing before you reach for it: DALI's resize and JPEG decode do not match Pillow's bit-for-bit either, so the same re-validation applies; and its main advantage over the code above is not raw speed but pipelining, prefetching and handling ragged input sizes without Python overhead.

CV-CUDA is the other option, offering GPU implementations of a wider set of OpenCV-style operations for when your preprocessing is more than decode-resize-normalise.

Common mistakes

Optimising this before measuring it. Profile end to end first. If preprocessing is 5 percent of your time, moving it to the GPU is a week of work for a 5 percent ceiling.

Forgetting to re-validate accuracy. The tensors changed. Your numbers may have too.

Decoding on the GPU while the model is already GPU-bound. nvJPEG uses dedicated hardware on most datacentre cards, but the resize and normalise steps compete with the model for the same cores. If the GPU is already saturated, you have moved the bottleneck rather than removed it.

Using GPU decode for tiny images. The per-call overhead dominates below roughly a few hundred pixels square. Batch them, or leave them on the CPU.

Assuming every image is a JPEG. decode_jpeg with device="cuda" handles JPEG only. PNG, WebP and 16-bit TIFF fall back to the CPU, and a mixed-format input stream will quietly take the slow path for part of every batch.

Timing without synchronisation. The most common way to publish a fake speed-up.

Try it yourself

Change quality=85 to quality=30 when creating the blobs and rerun. Smaller files mean less decode work; predict whether the CPU or the GPU path benefits more. Then replace the tiny model with torchvision.models.resnet50() and watch the preprocessing share fall from 99.7 percent to something you have to think about.

What to learn next

Researcher — Mathematics and papers.

Where the time goes

A JPEG decode is entropy decoding (Huffman, inherently serial per scan), dequantisation, an inverse DCT per 8x8 block, chroma upsampling, and a colour-space conversion. Only the first stage resists parallelisation. libjpeg-turbo attacks the rest with SIMD; nvJPEG maps the block-parallel stages onto the GPU and, on most datacentre and professional cards, onto a dedicated hardware decoder engine that does not consume streaming-multiprocessor time at all.

The resize is memory-bound: $O(HW)$ reads and $O(H'W')$ writes with a small filter kernel. It is close to the ideal GPU workload, and the measured gain is roughly the ratio of achievable memory bandwidth, which on an A6000 against a single CPU core is very large.

Normalisation is elementwise and is worth fusing into the resize output rather than running as a separate pass.

The transfer, and why decode placement matters twice

For a batch of $N$ images at resolution $H \times W$ with 3 channels:

$$ \text{bytes across PCIe} = \begin{cases} N \cdot 3HW & \text{decode on CPU (uint8)} \ N \cdot \text{compressed size} & \text{decode on GPU} \end{cases} $$

At 1920x1080, that is 6.2 MB per image decoded versus roughly 0.3 MB compressed at quality 85 — a factor of about 20. On PCIe 4.0 x16 (about 25 GB/s usable) a batch of 64 decoded frames is roughly 16 ms of pure transfer, which is itself larger than many models' forward pass.

Pinned memory matters because pageable host memory cannot be the source of an asynchronous DMA. The driver stages it through an internal pinned buffer, serialising the copy. Pinning is not free either: it reserves physical pages and over-pinning degrades system performance, so pin the data loader's output buffers, not everything.

Amdahl, stated for this problem

With preprocessing fraction $p$ of total time and a speed-up $k$ on that fraction alone, end-to-end speed-up is

$$ S = \frac{1}{(1-p) + p/k} $$

The measured $p = 0.997$ and $k = 6.8$ give $S \approx 5.6$. Note the asymmetry: at $p = 0.05$, even an infinite $k$ gives $S = 1.05$. This is the arithmetic behind "profile before optimising", and it is why the order in which bottlenecks are removed determines what is worth doing next.

Numerical divergence is not optional

The measured maximum divergence of 0.0637 in normalised units is roughly $16/255$ in pixel terms — twice the standard adversarial perturbation budget of $8/255$. It arises from:

  • Resampling filter differences, the same mechanism analysed by Parmar et al. (2022): whether the kernel support is stretched by the scale factor.
  • IDCT implementation. The JPEG standard specifies the IDCT only to an accuracy tolerance, not bit-exactly. libjpeg-turbo's integer IDCT and nvJPEG's need not agree to the last bit.
  • Chroma upsampling. Fancy (triangular) versus nearest upsampling of the subsampled chroma planes is an encoder-independent decoder choice.

The correct response is a validation gate: re-run the held-out evaluation with the new pipeline and compare, rather than assuming that "the same operations" produce the same numbers.

Design patterns for throughput

Decouple the stages. Decode, resize and inference should be separate stages with bounded queues between them, so a slow stage applies backpressure rather than accumulating latency. The same argument as in RTSP and live camera pipelines.

Batch at the widest point. nvJPEG's batched API amortises launch overhead; batching the resize into one interpolate call over a stacked tensor beats a Python loop over per-image calls by a wide margin, as the code above does.

Prefer fixed output shapes. Ragged inputs force per-image kernel launches. Bucketing by aspect ratio and letterboxing recovers batching.

Consider decoding once and caching. For repeated passes over a fixed dataset — most training runs — decode to a raw or lightly-compressed on-disk format once. This removes the decode cost entirely and usually beats any GPU decode strategy, at the price of disk.

Tooling

ToolScopeNote
torchvision.io.decode_jpeg(device="cuda")JPEG decodein-tree, no extra dependency; JPEG only
NVIDIA DALIfull pipeline graphprefetching, threading, framework plugins
CV-CUDAOpenCV-style ops on GPUbroader operator coverage
nvImageCodecdecode/encode, multiple formatssuccessor to the standalone nvJPEG samples
Triton ensemble modelsserver-side pipelinepreprocessing as a model in the serving graph

References

What to learn next