Train and serve preprocessing mismatch
The single most common reason a vision model that scored well in a notebook performs worse in production is that the serving code prepares the image differently.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
The code that prepares an image for serving must match the code that prepared it during training. Otherwise the model sees a different picture.
The analogy you have lived
Your mother's rice recipe says one cup of rice, two cups of water. At home, with her steel cup, it comes out perfect every time.
You cook it at a friend's house with their cup. Same recipe, same steps, same care. The rice is a mush.
Nothing about the recipe was wrong. The cup changed. And the recipe never wrote down which cup.
A trained model is that recipe. The preprocessing code is the cup.
What preparation actually does
A model does not eat photographs. It eats a fixed-size block of numbers. Getting from one to the other takes four steps, and every one of them can differ between two computers.
photo file
| 1. decode it into pixels
| 2. resize to the size the model expects
| 3. put the colour channels in the right order
| 4. shift and scale the numbers
v
the block of numbers the model was trained onStep two has at least six common implementations that give different answers. Step three has two conventions, and half the world uses each. Step four is written down in a config file that people forget to copy.
The mistakes, in order of how much damage they do
Wrong colour order. Some libraries hand you red, green, blue. Others hand you blue, green, red. Swap them and every colour in the image is wrong. This is the worst one, and it is silent. No error, no warning. Only a model that has become mediocre.
Missing or wrong number scaling. Training subtracted a set of fixed values from each colour and divided by another set. Skip that at serving time. The numbers reaching the model are now several times larger than anything it learned from.
Different resizing. Two libraries can both claim "bilinear" resizing and produce visibly different images. One filters out fine detail before shrinking; the other does not. The difference is small per pixel and real.
Wrong size or wrong crop. Training resized the short side and cut a square from the middle. Serving squashed the whole photo into a square. Everything in the image is now a different shape.
Why nobody catches it
None of these throw an error. The tensor is the right shape. The model runs. It returns a confident answer.
The accuracy is a few points lower than the notebook said. The team spends a month blaming the model, the data, or production.
The only reliable way to find it is to compare numbers, not pictures. Take one image, push it through both pipelines, and subtract.
The fix that actually works
Write the preparation once. Ship it with the model, in the same file or the same folder. Never let a training script and a serving script contain two copies of it.
Better still, bake the preparation into the exported model itself, so there is nothing left to get wrong.
Where this has bitten people
- A phone app whose accuracy dropped because the camera library returned a different colour order.
- A web service that resized with one library while training used another.
- A model retrained by a new team member who copied the architecture and not the mean and standard deviation values.
- A pipeline where images were rotated by the camera's own metadata during training and not during serving.
The honest part
You will not notice this by looking at the images. The two versions look identical to a human eye.
That is exactly why it survives code review, survives testing, and reaches production. The check has to be numeric, and it has to be automatic.
Remember this
- The model eats numbers, not photos. Four steps turn one into the other, and all four can differ.
- Colour order and number scaling cause far more damage than resizing does.
- Write the preparation once, ship it with the model, and test it numerically.
What to learn next
- Image formats and compression artefacts — the damage that happens before your code even sees the pixels.
- torchvision transforms v2 — the training-side API whose defaults you have to match.
- ONNX — exporting a model with its preprocessing attached.
Developer — Code and libraries.
Setup
pip install numpy opencv-python pillow torchVersions used below: numpy 1.26.4, opencv-python 4.11.0, Pillow 11.0.0, torch 2.5.1. Resizing behaviour has changed between library versions before, which is itself an argument for pinning them.
Measure the disagreement, then measure what it costs
import numpy as np, cv2, torch
import torch.nn.functional as F
from PIL import Image
rng = np.random.default_rng(0)
def make_image(i):
"""A 512x512 RGB picture with fine texture, like real photos have."""
yy, xx = np.mgrid[0:512, 0:512]
f = 0.20 + 0.02 * (i % 7)
base = 0.5 + 0.45 * np.sin(f * xx + 0.3 * i) * np.cos(f * yy * 0.9 + 0.2 * i)
img = np.stack([base, np.roll(base, 3, 0), np.roll(base, 6, 1)], -1)
return (np.clip(img, 0, 1) * 255).astype(np.uint8)
def pipe_pil(arr):
"""What most training scripts do: PIL, RGB, bilinear, /255."""
im = Image.fromarray(arr).resize((224, 224), Image.BILINEAR)
return np.asarray(im, dtype=np.float32) / 255.0
def pipe_cv2(arr):
"""What most C++/OpenCV serving code does: cv2 bilinear on a BGR array."""
bgr = cv2.cvtColor(arr, cv2.COLOR_RGB2BGR)
small = cv2.resize(bgr, (224, 224), interpolation=cv2.INTER_LINEAR)
rgb = cv2.cvtColor(small, cv2.COLOR_BGR2RGB)
return rgb.astype(np.float32) / 255.0
def pipe_torch_nearest(arr):
"""torch.nn.functional.interpolate, whose antialias default is False."""
t = torch.from_numpy(arr).permute(2, 0, 1).float()[None] / 255.0
out = F.interpolate(t, size=(224, 224), mode="bilinear",
align_corners=False, antialias=False)
return out[0].permute(1, 2, 0).numpy()
def pipe_torch_aa(arr):
t = torch.from_numpy(arr).permute(2, 0, 1).float()[None] / 255.0
out = F.interpolate(t, size=(224, 224), mode="bilinear",
align_corners=False, antialias=True)
return out[0].permute(1, 2, 0).numpy()
PIPES = {"PIL bilinear": pipe_pil,
"cv2 INTER_LINEAR": pipe_cv2,
"torch antialias=False": pipe_torch_nearest,
"torch antialias=True": pipe_torch_aa}
arr = make_image(0)
ref = pipe_pil(arr)
print("difference from the PIL pipeline, on one 512x512 photo resized to 224:")
print(f"{'pipeline':<24} {'mean |diff|':>11} {'max |diff|':>10} {'>1/255':>8} {'cosine':>8}")
for name, fn in PIPES.items():
out = fn(arr)
d = np.abs(out - ref)
cos = float((out.ravel() @ ref.ravel())
/ (np.linalg.norm(out) * np.linalg.norm(ref)))
print(f"{name:<24} {d.mean():>11.5f} {d.max():>10.5f} "
f"{(d > 1/255).mean():>7.1%} {cos:>8.5f}")
# A fixed, seeded linear classifier: same weights on every machine and every run.
W = rng.normal(size=(224 * 224 * 3, 5)).astype(np.float32) / 100
MEAN = np.array([0.485, 0.456, 0.406], np.float32)
STD = np.array([0.229, 0.224, 0.225], np.float32)
def classify(x, normalise=True):
z = (x - MEAN) / STD if normalise else x
return int(np.argmax(z.reshape(1, -1) @ W))
imgs = [make_image(i) for i in range(200)]
print("\nout of 200 images, how many change class when the pipeline changes?")
base = [classify(pipe_pil(a)) for a in imgs]
for name, fn in PIPES.items():
other = [classify(fn(a)) for a in imgs]
flips = sum(1 for a, b in zip(base, other) if a != b)
print(f" trained with PIL, served with {name:<22} {flips:>3}/200")
print("\ntwo mistakes that are not about resizing at all:")
no_norm = sum(1 for a in imgs
if classify(pipe_pil(a)) != classify(pipe_pil(a), normalise=False))
print(f" forgot the mean/std normalisation entirely {no_norm:>3}/200")
swapped = sum(1 for a in imgs
if classify(pipe_pil(a)) != classify(pipe_pil(a)[..., ::-1].copy()))
print(f" fed BGR to a model trained on RGB {swapped:>3}/200")difference from the PIL pipeline, on one 512x512 photo resized to 224: pipeline mean |diff| max |diff| >1/255 cosine PIL bilinear 0.00000 0.00000 0.0% 1.00000 cv2 INTER_LINEAR 0.00511 0.19608 38.0% 0.99988 torch antialias=False 0.00518 0.19492 53.9% 0.99990 torch antialias=True 0.00108 0.00392 0.0% 1.00000 out of 200 images, how many change class when the pipeline changes? trained with PIL, served with PIL bilinear 0/200 trained with PIL, served with cv2 INTER_LINEAR 3/200 trained with PIL, served with torch antialias=False 4/200 trained with PIL, served with torch antialias=True 1/200 two mistakes that are not about resizing at all: forgot the mean/std normalisation entirely 105/200 fed BGR to a model trained on RGB 132/200
Read the first table before the second
cv2.INTER_LINEAR disagrees with PIL on 38 percent of pixels by more than one grey level, with a worst case of 0.196 — that is 50 out of 255. These are not rounding differences. On a picture with fine texture they are structural: PIL applies a low-pass filter scaled to the downsampling factor, and cv2.INTER_LINEAR does not, so it aliases. Use cv2.INTER_AREA when downscaling and the gap narrows dramatically.
torch.nn.functional.interpolate defaults to antialias=False, and it is worse still: 53.9 percent of pixels differ. This is the single most-hit trap in PyTorch serving code. torchvision.transforms.v2.Resize defaults to antialias=True for tensors, so a training script using transforms and a serving script using F.interpolate disagree by default. Setting antialias=True drops the maximum error to 0.00392, which is one grey level.
Cosine similarity is 0.99988 for the worst pipeline. If your test is "are the tensors similar", every one of these passes. Cosine similarity is the wrong test. Use maximum absolute difference and the fraction of pixels above one grey level.
Now read the second table, because it reverses the priorities
Resizing mismatches flipped 1 to 4 predictions out of 200. Real, worth fixing, and small.
Forgetting normalisation flipped 105 out of 200. Feeding BGR to an RGB model flipped 132 out of 200. Two orders of magnitude more damage, from mistakes that take one line to make and one line to fix.
This ordering is the practical lesson. Teams spend days arguing about interpolation filters while the actual production bug is a missing cv2.cvtColor or a config file that never made it into the container. Check the cheap, catastrophic things first.
A caveat stated plainly: the "classifier" here is a fixed random projection, chosen so the numbers reproduce on any machine. A trained network is more robust to small input perturbations than a random projection, so treat the resize flip counts as an upper bound and the relative ordering as the finding.
Ship the preparation inside the model
The most reliable fix is to leave nothing for the serving code to get wrong.
import torch, torch.nn as nn
class Preprocessed(nn.Module):
"""Takes uint8 NHWC RGB straight from the decoder. No config to forget."""
def __init__(self, net):
super().__init__()
self.net = net
self.register_buffer("mean", torch.tensor([0.485, 0.456, 0.406]).view(1, 3, 1, 1))
self.register_buffer("std", torch.tensor([0.229, 0.224, 0.225]).view(1, 3, 1, 1))
def forward(self, x): # x: (N, H, W, 3), uint8, RGB
x = x.permute(0, 3, 1, 2).float() / 255.0
x = torch.nn.functional.interpolate(
x, size=(224, 224), mode="bilinear",
align_corners=False, antialias=True) # matches transforms.v2.Resize
return self.net((x - self.mean) / self.std)
net = nn.Sequential(nn.AdaptiveAvgPool2d(1), nn.Flatten(), nn.Linear(3, 5))
model = Preprocessed(net).eval()
raw = torch.randint(0, 256, (2, 512, 512, 3), dtype=torch.uint8)
with torch.no_grad():
print("output shape:", tuple(model(raw).shape))
print("buffers travel with the weights:", list(dict(model.named_buffers())))output shape: (2, 5) buffers travel with the weights: ['mean', 'std']
register_buffer is the important detail. Buffers are saved in state_dict and exported into ONNX, so the mean and standard deviation cannot be separated from the weights. A plain Python attribute would be lost the moment somebody loads the checkpoint in a different script.
The test that belongs in your CI
def test_pipelines_agree():
img = load_fixture("test_image.png")
a = training_preprocess(img)
b = serving_preprocess(img)
assert np.abs(a - b).max() < 1e-4, f"max diff {np.abs(a - b).max()}"Check in one real image as a fixture. Compare the tensors, not the predictions. This test costs milliseconds and catches the class of bug that costs weeks.
Common mistakes
Resizing to a square instead of resize-short-side-then-centre-crop. ImageNet models were trained by scaling the short side to 256 and cropping 224 from the middle. Squashing a 16:9 photo into a square changes every aspect ratio in it.
Assuming cv2.imread gives you RGB. It gives you BGR. It has always given you BGR. This is the most reliable bug in computer vision.
Ignoring EXIF orientation. PIL.Image.open does not rotate by the EXIF tag; ImageOps.exif_transpose does. Phone photos are frequently stored sideways with a tag saying which way is up. See image formats and compression artefacts.
Different scaling conventions in different frameworks. Some model families expect [0, 1], some [-1, 1], some raw [0, 255] with only a mean subtracted. Read the model card, do not guess.
Testing with a solid-colour or noise image. Resizing differences appear only on structured, high-frequency content. A grey square passes every comparison.
Try it yourself
Change pipe_cv2 to use cv2.INTER_AREA and rerun the first table. Then remove the cv2.cvtColor calls so the pipeline genuinely returns BGR, and confirm the flip count jumps into the hundreds. Feeling that gap once is what makes you write the CI test.
What to learn next
- Image formats and compression artefacts — the damage that happens before your code even sees the pixels.
- torchvision transforms v2 — the training-side API whose defaults you have to match.
- ONNX — exporting a model with its preprocessing attached.
Researcher — Mathematics and papers.
Why two "bilinear" resizers disagree
Downsampling by a factor $s > 1$ is a decimation. Without prefiltering it violates the Nyquist criterion, and content above the new Nyquist frequency folds back as aliasing.
A correctly implemented resampler convolves with a reconstruction kernel whose support is stretched by the scale factor when minifying:
$$ y[n] = \sum_{k} x[k] \cdot h!\left(\frac{k - s\,n}{s}\right) \Big/ \sum_{k} h!\left(\frac{k - s\,n}{s}\right) $$
For bilinear, $h(t) = \max(0, 1 - |t|)$. With $s = 512/224 \approx 2.29$, the correct kernel spans about 4.6 input pixels. An implementation that keeps the kernel at its unit width samples only 2 input pixels regardless of $s$, discarding most of the input and aliasing the rest.
Pillow stretches the support. cv2.INTER_LINEAR does not (cv2.INTER_AREA does something equivalent for integer-ish factors). torch.nn.functional.interpolate does not unless antialias=True. This single design difference accounts for the entire first table above.
Parmar et al. (2022) showed the same defect measurably shifts FID scores, and Zhang (2019) showed aliasing from strided downsampling inside networks causes prediction instability under one-pixel shifts. It is the same signal-processing failure in three different places.
Secondary differences persist even between correct implementations: the pixel-centre convention (align_corners), the boundary extension rule, and whether the computation is done in uint8 or floating point. These produce differences on the order of one grey level, not fifty.
Framing this as covariate shift
Let $T_{\text{train}}$ and $T_{\text{serve}}$ be the two preprocessing maps and $f$ the model. Training minimises risk under $\mathbb{E}{x \sim \mathcal{D}}\left[\ell(f(T{\text{train}}(x)), y)\right]$; deployment incurs $\mathbb{E}{x \sim \mathcal{D}}\left[\ell(f(T{\text{serve}}(x)), y)\right]$. The gap is a covariate shift induced entirely by engineering, with two properties that make it unusually nasty:
- It is deterministic and systematic, so it does not average out over a large evaluation set.
- It is invisible to every input-distribution monitor, because the shift happens after the monitored raw input and before the monitored model output.
The perturbation magnitude is worth putting in context. The measured maximum was 0.196 in $[0,1]$ units, roughly $\ell_\infty = 50/255$. Standard adversarial robustness work operates at $\epsilon = 8/255$. The preprocessing mismatch is a six-times-larger perturbation than the adversarial threat model the field spends most of its effort on, arriving for free, on every request.
Eliminating the class of bug
Bake preprocessing into the exported graph. ONNX supports the required ops; torch.onnx.export on a wrapper module like the one above produces a graph that takes raw uint8. TensorRT and OpenVINO both have preprocessing APIs (IPreprocessing and ov.preprocess.PrePostProcessor) that fold resize, layout conversion and normalisation into the compiled engine, where they also run faster.
Or make the contract explicit and testable. Publish the preprocessing as a versioned artefact alongside the weights, with a golden-tensor fixture: one input image, one expected output tensor, checked byte-for-byte in CI. This is what a model card's "preprocessing" section should contain and rarely does.
Do not rely on numerically equivalent reimplementation. Two teams writing "the same" resize in two languages is the failure mode, not the fix.
What to measure
| Statistic | Why |
|---|---|
| $\max\lvert a - b\rvert$ | catches structural disagreement; the only reliable one |
| fraction of elements $> 1/255$ | scale-free, interpretable as "visibly different pixels" |
| per-channel mean and std of both tensors | catches normalisation and channel-order errors instantly |
| cosine similarity | not sufficient; stays above 0.999 through serious errors |
| KL divergence of output distributions over a held-out set | end-to-end check when you cannot inspect tensors |
The per-channel statistics deserve emphasis. An RGB/BGR swap leaves the max-difference test looking alarming but ambiguous, while the channel means immediately show which two channels traded places.
Papers and references
- Parmar et al., On Aliased Resizing and Surprising Subtleties in GAN Evaluation, CVPR 2022 — arxiv.org/abs/2104.11222
- Zhang, Making Convolutional Networks Shift-Invariant Again, ICML 2019 — arxiv.org/abs/1904.11486
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — the training/serving skew framing
- Pillow resampling documentation — pillow.readthedocs.io/en/stable/handbook/concepts.html#filters
- torchvision transforms v2 — docs.pytorch.org/vision/stable/transforms.html
What to learn next
- Image formats and compression artefacts — the damage that happens before your code even sees the pixels.
- torchvision transforms v2 — the training-side API whose defaults you have to match.
- ONNX — exporting a model with its preprocessing attached.