Shipping Vision Models

Train and serve preprocessing mismatch

The single most common reason a vision model that scored well in a notebook performs worse in production is that the serving code prepares the image differently.

On this page 10
  1. The short answer
  2. The analogy you have lived
  3. What preparation actually does
  4. The mistakes, in order of how much damage they do
  5. Why nobody catches it
  6. The fix that actually works
  7. Where this has bitten people
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

The code that prepares an image for serving must match the code that prepared it during training. Otherwise the model sees a different picture.

The analogy you have lived

Your mother's rice recipe says one cup of rice, two cups of water. At home, with her steel cup, it comes out perfect every time.

You cook it at a friend's house with their cup. Same recipe, same steps, same care. The rice is a mush.

Nothing about the recipe was wrong. The cup changed. And the recipe never wrote down which cup.

A trained model is that recipe. The preprocessing code is the cup.

What preparation actually does

A model does not eat photographs. It eats a fixed-size block of numbers. Getting from one to the other takes four steps, and every one of them can differ between two computers.

   photo file
      |  1. decode it into pixels
      |  2. resize to the size the model expects
      |  3. put the colour channels in the right order
      |  4. shift and scale the numbers
      v
   the block of numbers the model was trained on

Step two has at least six common implementations that give different answers. Step three has two conventions, and half the world uses each. Step four is written down in a config file that people forget to copy.

The mistakes, in order of how much damage they do

Wrong colour order. Some libraries hand you red, green, blue. Others hand you blue, green, red. Swap them and every colour in the image is wrong. This is the worst one, and it is silent. No error, no warning. Only a model that has become mediocre.

Missing or wrong number scaling. Training subtracted a set of fixed values from each colour and divided by another set. Skip that at serving time. The numbers reaching the model are now several times larger than anything it learned from.

Different resizing. Two libraries can both claim "bilinear" resizing and produce visibly different images. One filters out fine detail before shrinking; the other does not. The difference is small per pixel and real.

Wrong size or wrong crop. Training resized the short side and cut a square from the middle. Serving squashed the whole photo into a square. Everything in the image is now a different shape.

Why nobody catches it

None of these throw an error. The tensor is the right shape. The model runs. It returns a confident answer.

The accuracy is a few points lower than the notebook said. The team spends a month blaming the model, the data, or production.

The only reliable way to find it is to compare numbers, not pictures. Take one image, push it through both pipelines, and subtract.

The fix that actually works

Write the preparation once. Ship it with the model, in the same file or the same folder. Never let a training script and a serving script contain two copies of it.

Better still, bake the preparation into the exported model itself, so there is nothing left to get wrong.

Where this has bitten people

  • A phone app whose accuracy dropped because the camera library returned a different colour order.
  • A web service that resized with one library while training used another.
  • A model retrained by a new team member who copied the architecture and not the mean and standard deviation values.
  • A pipeline where images were rotated by the camera's own metadata during training and not during serving.

The honest part

You will not notice this by looking at the images. The two versions look identical to a human eye.

That is exactly why it survives code review, survives testing, and reaches production. The check has to be numeric, and it has to be automatic.

Remember this

  • The model eats numbers, not photos. Four steps turn one into the other, and all four can differ.
  • Colour order and number scaling cause far more damage than resizing does.
  • Write the preparation once, ship it with the model, and test it numerically.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy opencv-python pillow torch

Versions used below: numpy 1.26.4, opencv-python 4.11.0, Pillow 11.0.0, torch 2.5.1. Resizing behaviour has changed between library versions before, which is itself an argument for pinning them.

Measure the disagreement, then measure what it costs

preproc_mismatch.py
import numpy as np, cv2, torch
import torch.nn.functional as F
from PIL import Image

rng = np.random.default_rng(0)


def make_image(i):
    """A 512x512 RGB picture with fine texture, like real photos have."""
    yy, xx = np.mgrid[0:512, 0:512]
    f = 0.20 + 0.02 * (i % 7)
    base = 0.5 + 0.45 * np.sin(f * xx + 0.3 * i) * np.cos(f * yy * 0.9 + 0.2 * i)
    img = np.stack([base, np.roll(base, 3, 0), np.roll(base, 6, 1)], -1)
    return (np.clip(img, 0, 1) * 255).astype(np.uint8)


def pipe_pil(arr):
    """What most training scripts do: PIL, RGB, bilinear, /255."""
    im = Image.fromarray(arr).resize((224, 224), Image.BILINEAR)
    return np.asarray(im, dtype=np.float32) / 255.0


def pipe_cv2(arr):
    """What most C++/OpenCV serving code does: cv2 bilinear on a BGR array."""
    bgr = cv2.cvtColor(arr, cv2.COLOR_RGB2BGR)
    small = cv2.resize(bgr, (224, 224), interpolation=cv2.INTER_LINEAR)
    rgb = cv2.cvtColor(small, cv2.COLOR_BGR2RGB)
    return rgb.astype(np.float32) / 255.0


def pipe_torch_nearest(arr):
    """torch.nn.functional.interpolate, whose antialias default is False."""
    t = torch.from_numpy(arr).permute(2, 0, 1).float()[None] / 255.0
    out = F.interpolate(t, size=(224, 224), mode="bilinear",
                        align_corners=False, antialias=False)
    return out[0].permute(1, 2, 0).numpy()


def pipe_torch_aa(arr):
    t = torch.from_numpy(arr).permute(2, 0, 1).float()[None] / 255.0
    out = F.interpolate(t, size=(224, 224), mode="bilinear",
                        align_corners=False, antialias=True)
    return out[0].permute(1, 2, 0).numpy()


PIPES = {"PIL bilinear": pipe_pil,
         "cv2 INTER_LINEAR": pipe_cv2,
         "torch antialias=False": pipe_torch_nearest,
         "torch antialias=True": pipe_torch_aa}

arr = make_image(0)
ref = pipe_pil(arr)
print("difference from the PIL pipeline, on one 512x512 photo resized to 224:")
print(f"{'pipeline':<24} {'mean |diff|':>11} {'max |diff|':>10} {'>1/255':>8} {'cosine':>8}")
for name, fn in PIPES.items():
    out = fn(arr)
    d = np.abs(out - ref)
    cos = float((out.ravel() @ ref.ravel())
                / (np.linalg.norm(out) * np.linalg.norm(ref)))
    print(f"{name:<24} {d.mean():>11.5f} {d.max():>10.5f} "
          f"{(d > 1/255).mean():>7.1%} {cos:>8.5f}")

# A fixed, seeded linear classifier: same weights on every machine and every run.
W = rng.normal(size=(224 * 224 * 3, 5)).astype(np.float32) / 100
MEAN = np.array([0.485, 0.456, 0.406], np.float32)
STD = np.array([0.229, 0.224, 0.225], np.float32)


def classify(x, normalise=True):
    z = (x - MEAN) / STD if normalise else x
    return int(np.argmax(z.reshape(1, -1) @ W))


imgs = [make_image(i) for i in range(200)]
print("\nout of 200 images, how many change class when the pipeline changes?")
base = [classify(pipe_pil(a)) for a in imgs]
for name, fn in PIPES.items():
    other = [classify(fn(a)) for a in imgs]
    flips = sum(1 for a, b in zip(base, other) if a != b)
    print(f"  trained with PIL, served with {name:<22} {flips:>3}/200")

print("\ntwo mistakes that are not about resizing at all:")
no_norm = sum(1 for a in imgs
              if classify(pipe_pil(a)) != classify(pipe_pil(a), normalise=False))
print(f"  forgot the mean/std normalisation entirely   {no_norm:>3}/200")
swapped = sum(1 for a in imgs
              if classify(pipe_pil(a)) != classify(pipe_pil(a)[..., ::-1].copy()))
print(f"  fed BGR to a model trained on RGB           {swapped:>3}/200")
Output
difference from the PIL pipeline, on one 512x512 photo resized to 224:
pipeline                 mean |diff| max |diff|   >1/255   cosine
PIL bilinear                 0.00000    0.00000    0.0%  1.00000
cv2 INTER_LINEAR             0.00511    0.19608   38.0%  0.99988
torch antialias=False        0.00518    0.19492   53.9%  0.99990
torch antialias=True         0.00108    0.00392    0.0%  1.00000

out of 200 images, how many change class when the pipeline changes?
  trained with PIL, served with PIL bilinear             0/200
  trained with PIL, served with cv2 INTER_LINEAR         3/200
  trained with PIL, served with torch antialias=False    4/200
  trained with PIL, served with torch antialias=True     1/200

two mistakes that are not about resizing at all:
  forgot the mean/std normalisation entirely   105/200
  fed BGR to a model trained on RGB           132/200

Read the first table before the second

cv2.INTER_LINEAR disagrees with PIL on 38 percent of pixels by more than one grey level, with a worst case of 0.196 — that is 50 out of 255. These are not rounding differences. On a picture with fine texture they are structural: PIL applies a low-pass filter scaled to the downsampling factor, and cv2.INTER_LINEAR does not, so it aliases. Use cv2.INTER_AREA when downscaling and the gap narrows dramatically.

torch.nn.functional.interpolate defaults to antialias=False, and it is worse still: 53.9 percent of pixels differ. This is the single most-hit trap in PyTorch serving code. torchvision.transforms.v2.Resize defaults to antialias=True for tensors, so a training script using transforms and a serving script using F.interpolate disagree by default. Setting antialias=True drops the maximum error to 0.00392, which is one grey level.

Cosine similarity is 0.99988 for the worst pipeline. If your test is "are the tensors similar", every one of these passes. Cosine similarity is the wrong test. Use maximum absolute difference and the fraction of pixels above one grey level.

Now read the second table, because it reverses the priorities

Resizing mismatches flipped 1 to 4 predictions out of 200. Real, worth fixing, and small.

Forgetting normalisation flipped 105 out of 200. Feeding BGR to an RGB model flipped 132 out of 200. Two orders of magnitude more damage, from mistakes that take one line to make and one line to fix.

This ordering is the practical lesson. Teams spend days arguing about interpolation filters while the actual production bug is a missing cv2.cvtColor or a config file that never made it into the container. Check the cheap, catastrophic things first.

A caveat stated plainly: the "classifier" here is a fixed random projection, chosen so the numbers reproduce on any machine. A trained network is more robust to small input perturbations than a random projection, so treat the resize flip counts as an upper bound and the relative ordering as the finding.

Ship the preparation inside the model

The most reliable fix is to leave nothing for the serving code to get wrong.

baked_in.py
import torch, torch.nn as nn

class Preprocessed(nn.Module):
    """Takes uint8 NHWC RGB straight from the decoder. No config to forget."""
    def __init__(self, net):
        super().__init__()
        self.net = net
        self.register_buffer("mean", torch.tensor([0.485, 0.456, 0.406]).view(1, 3, 1, 1))
        self.register_buffer("std", torch.tensor([0.229, 0.224, 0.225]).view(1, 3, 1, 1))

    def forward(self, x):                       # x: (N, H, W, 3), uint8, RGB
        x = x.permute(0, 3, 1, 2).float() / 255.0
        x = torch.nn.functional.interpolate(
            x, size=(224, 224), mode="bilinear",
            align_corners=False, antialias=True)   # matches transforms.v2.Resize
        return self.net((x - self.mean) / self.std)

net = nn.Sequential(nn.AdaptiveAvgPool2d(1), nn.Flatten(), nn.Linear(3, 5))
model = Preprocessed(net).eval()

raw = torch.randint(0, 256, (2, 512, 512, 3), dtype=torch.uint8)
with torch.no_grad():
    print("output shape:", tuple(model(raw).shape))
print("buffers travel with the weights:", list(dict(model.named_buffers())))
Output
output shape: (2, 5)
buffers travel with the weights: ['mean', 'std']

register_buffer is the important detail. Buffers are saved in state_dict and exported into ONNX, so the mean and standard deviation cannot be separated from the weights. A plain Python attribute would be lost the moment somebody loads the checkpoint in a different script.

The test that belongs in your CI

python
def test_pipelines_agree():
    img = load_fixture("test_image.png")
    a = training_preprocess(img)
    b = serving_preprocess(img)
    assert np.abs(a - b).max() < 1e-4, f"max diff {np.abs(a - b).max()}"

Check in one real image as a fixture. Compare the tensors, not the predictions. This test costs milliseconds and catches the class of bug that costs weeks.

Common mistakes

Resizing to a square instead of resize-short-side-then-centre-crop. ImageNet models were trained by scaling the short side to 256 and cropping 224 from the middle. Squashing a 16:9 photo into a square changes every aspect ratio in it.

Assuming cv2.imread gives you RGB. It gives you BGR. It has always given you BGR. This is the most reliable bug in computer vision.

Ignoring EXIF orientation. PIL.Image.open does not rotate by the EXIF tag; ImageOps.exif_transpose does. Phone photos are frequently stored sideways with a tag saying which way is up. See image formats and compression artefacts.

Different scaling conventions in different frameworks. Some model families expect [0, 1], some [-1, 1], some raw [0, 255] with only a mean subtracted. Read the model card, do not guess.

Testing with a solid-colour or noise image. Resizing differences appear only on structured, high-frequency content. A grey square passes every comparison.

Try it yourself

Change pipe_cv2 to use cv2.INTER_AREA and rerun the first table. Then remove the cv2.cvtColor calls so the pipeline genuinely returns BGR, and confirm the flip count jumps into the hundreds. Feeling that gap once is what makes you write the CI test.

What to learn next

Researcher — Mathematics and papers.

Why two "bilinear" resizers disagree

Downsampling by a factor $s > 1$ is a decimation. Without prefiltering it violates the Nyquist criterion, and content above the new Nyquist frequency folds back as aliasing.

A correctly implemented resampler convolves with a reconstruction kernel whose support is stretched by the scale factor when minifying:

$$ y[n] = \sum_{k} x[k] \cdot h!\left(\frac{k - s\,n}{s}\right) \Big/ \sum_{k} h!\left(\frac{k - s\,n}{s}\right) $$

For bilinear, $h(t) = \max(0, 1 - |t|)$. With $s = 512/224 \approx 2.29$, the correct kernel spans about 4.6 input pixels. An implementation that keeps the kernel at its unit width samples only 2 input pixels regardless of $s$, discarding most of the input and aliasing the rest.

Pillow stretches the support. cv2.INTER_LINEAR does not (cv2.INTER_AREA does something equivalent for integer-ish factors). torch.nn.functional.interpolate does not unless antialias=True. This single design difference accounts for the entire first table above.

Parmar et al. (2022) showed the same defect measurably shifts FID scores, and Zhang (2019) showed aliasing from strided downsampling inside networks causes prediction instability under one-pixel shifts. It is the same signal-processing failure in three different places.

Secondary differences persist even between correct implementations: the pixel-centre convention (align_corners), the boundary extension rule, and whether the computation is done in uint8 or floating point. These produce differences on the order of one grey level, not fifty.

Framing this as covariate shift

Let $T_{\text{train}}$ and $T_{\text{serve}}$ be the two preprocessing maps and $f$ the model. Training minimises risk under $\mathbb{E}{x \sim \mathcal{D}}\left[\ell(f(T{\text{train}}(x)), y)\right]$; deployment incurs $\mathbb{E}{x \sim \mathcal{D}}\left[\ell(f(T{\text{serve}}(x)), y)\right]$. The gap is a covariate shift induced entirely by engineering, with two properties that make it unusually nasty:

  • It is deterministic and systematic, so it does not average out over a large evaluation set.
  • It is invisible to every input-distribution monitor, because the shift happens after the monitored raw input and before the monitored model output.

The perturbation magnitude is worth putting in context. The measured maximum was 0.196 in $[0,1]$ units, roughly $\ell_\infty = 50/255$. Standard adversarial robustness work operates at $\epsilon = 8/255$. The preprocessing mismatch is a six-times-larger perturbation than the adversarial threat model the field spends most of its effort on, arriving for free, on every request.

Eliminating the class of bug

Bake preprocessing into the exported graph. ONNX supports the required ops; torch.onnx.export on a wrapper module like the one above produces a graph that takes raw uint8. TensorRT and OpenVINO both have preprocessing APIs (IPreprocessing and ov.preprocess.PrePostProcessor) that fold resize, layout conversion and normalisation into the compiled engine, where they also run faster.

Or make the contract explicit and testable. Publish the preprocessing as a versioned artefact alongside the weights, with a golden-tensor fixture: one input image, one expected output tensor, checked byte-for-byte in CI. This is what a model card's "preprocessing" section should contain and rarely does.

Do not rely on numerically equivalent reimplementation. Two teams writing "the same" resize in two languages is the failure mode, not the fix.

What to measure

StatisticWhy
$\max\lvert a - b\rvert$catches structural disagreement; the only reliable one
fraction of elements $> 1/255$scale-free, interpretable as "visibly different pixels"
per-channel mean and std of both tensorscatches normalisation and channel-order errors instantly
cosine similaritynot sufficient; stays above 0.999 through serious errors
KL divergence of output distributions over a held-out setend-to-end check when you cannot inspect tensors

The per-channel statistics deserve emphasis. An RGB/BGR swap leaves the max-difference test looking alarming but ambiguous, while the channel means immediately show which two channels traded places.

Papers and references

What to learn next