Shipping Vision Models

INT8 calibration for vision models

Running a model in 8-bit needs someone to decide what range each layer's numbers cover, and that decision is made by watching a few hundred real images pass through.

Read these first

On this page 11
  1. The short answer
  2. The analogy you have lived
  3. Why squeeze at all
  4. What the squeezing looks like
  5. Where the images come from
  6. The failure that teaches the most
  7. Two ways to choose the range
  8. Where you have seen this
  9. The honest part
  10. Remember this
  11. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Calibration means learning, from real images, what range of numbers each layer produces. That range decides how those numbers get squeezed into 8 bits.

The analogy you have lived

You have a thermometer with a fixed dial. Two hundred and fifty-six marks, and no more. You get to choose what the lowest and highest marks mean.

Set it from zero to fifty degrees and every weather reading is precise. Each mark is a fifth of a degree.

Set the same dial from zero to a thousand degrees, in case you ever measure a furnace. Now every mark is four degrees, and a warm day and a hot day land on the same mark. The dial is useless for weather.

Choosing the range is the whole problem. Choosing it by looking at real readings first is calibration.

Why squeeze at all

A trained model stores each number in 32 bits. Running in 8 bits makes the model four times smaller. It moves four times less data through memory. And modern chips have hardware built for it.

On phones and small edge devices, this often decides whether the model runs at all. On servers it is a straight cost saving.

What the squeezing looks like

Each layer's numbers get their own range. Everything inside that range is mapped onto the available marks. Anything outside is pushed to the nearest end.

   real values in a layer:   -2.7 ......... 0 ......... 3.1
                              |                          |
   chosen range              low                       high
                              |__________________________|
                                mapped onto 256 marks

Pick the range too narrow and the largest values get flattened. Pick it too wide and every ordinary value crowds into a few marks.

Where the images come from

You cannot know a layer's range by reading the code. It depends on what the model sees.

So you feed it images and watch. A few hundred is usually enough. They must be real images of the kind the model will serve. Not random noise. Not a blank test pattern. Not one class repeated.

This is the step people rush, and it is the step that decides whether the whole exercise works.

The failure that teaches the most

Calibrate on blank images and every layer reports a tiny range, because nothing interesting happened. Then a real photo arrives. Its values fly past the top of that range. Every one lands on the same top mark.

The model does not crash. It does not warn you. It returns confident nonsense.

That is the shape of nearly every quantisation disaster: not a broken model, a badly chosen range.

Two ways to choose the range

Take the smallest and largest value you saw. Simple, and one freak value in one image stretches the range for every image after it.

Ignore the extremes. Cut off the top and bottom fraction of a percent and set the range from what remains. A few values get clipped, and every ordinary value gets more marks. This is usually better, and it is a genuine trade you should measure rather than assume.

Where you have seen this

  • Face unlock and photo search running on a phone with no help from a server.
  • Camera modules in a shop or factory that must be cheap and low-power.
  • Voice assistants that respond without sending anything anywhere.
  • Large models on servers where 8-bit halves the cost per request.

The honest part

Eight-bit is not free, and honest teams measure the cost rather than assuming it away.

For most convolutional vision models it is small. For a few it is not, and the models that suffer tend to be the ones with unusual activation ranges. There is no way to know which yours is except to test it, on your data, against your metric.

Never ship a quantised model on the strength of the file size. Ship it on the strength of a measurement.

Remember this

  • Calibration decides what range of numbers each layer covers, by watching real images.
  • The calibration images must resemble production. A bad set gives a confident, broken model.
  • Always measure accuracy after quantising. The saving is guaranteed; the accuracy is not.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install onnx onnxruntime torch numpy

Versions used: onnxruntime 1.20.1, onnx 1.22.0, torch 2.5.1. This whole example runs on CPU. ONNX Runtime's static quantisation uses the same calibration machinery as every other toolkit, so what you learn here transfers directly to TensorRT and to mobile runtimes.

Quantise the same model four ways

int8_calibration.py
import os
import numpy as np
import torch, torch.nn as nn
import onnxruntime as ort
from onnxruntime.quantization import (quantize_static, CalibrationDataReader,
                                      QuantType, QuantFormat, CalibrationMethod)
from onnxruntime.quantization.shape_inference import quant_pre_process

torch.manual_seed(0)
rng = np.random.default_rng(0)


def shapes(n):
    """n greyscale 32x32 pictures: 0 = disc, 1 = square, 2 = triangle."""
    y = rng.integers(0, 3, n)
    out = np.full((n, 1, 32, 32), 0.15, np.float32)
    gy, gx = np.mgrid[0:32, 0:32]
    for i, k in enumerate(y):
        r = rng.integers(8, 12)
        cy, cx = 16 + rng.integers(-2, 3), 16 + rng.integers(-2, 3)
        if k == 0:
            m = (gy-cy)**2 + (gx-cx)**2 <= r*r
        elif k == 1:
            m = (abs(gy-cy) <= r) & (abs(gx-cx) <= r)
        else:
            m = (gy-cy >= -r) & (gy-cy <= r) & (abs(gx-cx) <= (gy-cy+r)/2)
        out[i, 0][m] = 0.85
    return out + rng.normal(0, .02, out.shape).astype(np.float32), y


net = nn.Sequential(nn.Conv2d(1, 8, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
                    nn.Conv2d(8, 16, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
                    nn.Flatten(), nn.Linear(16*8*8, 3))
x_tr, y_tr = shapes(600)
opt = torch.optim.Adam(net.parameters(), 3e-3)
for _ in range(60):                      # a short, honest training run
    opt.zero_grad()
    loss = nn.functional.cross_entropy(net(torch.tensor(x_tr)), torch.tensor(y_tr))
    loss.backward(); opt.step()
net.eval()

torch.onnx.export(net, torch.zeros(1, 1, 32, 32), "shapes.onnx",
                  input_names=["input"], output_names=["logits"],
                  dynamic_axes={"input": {0: "n"}, "logits": {0: "n"}},
                  opset_version=17)
# Fold constants and fill in every tensor shape; the quantiser needs both.
quant_pre_process("shapes.onnx", "shapes_prep.onnx")


class Reader(CalibrationDataReader):
    """Feeds calibration images to the quantiser, one batch at a time."""
    def __init__(self, data):
        self.it = iter([{"input": d[None]} for d in data])

    def get_next(self):
        return next(self.it, None)


x_cal, _ = shapes(128)                        # representative, from the same source
x_bad = np.full((128, 1, 32, 32), 0.15, np.float32)   # all blank: a useless set
x_te, y_te = shapes(400)


def build(path, cal, method):
    quantize_static("shapes_prep.onnx", path, Reader(cal),
                    quant_format=QuantFormat.QDQ,
                    activation_type=QuantType.QUInt8,
                    weight_type=QuantType.QInt8,
                    per_channel=True,
                    calibrate_method=method,
                    extra_options={"ActivationSymmetric": False})
    return path


def run(path, x):
    s = ort.InferenceSession(path, providers=["CPUExecutionProvider"])
    return s.run(None, {"input": x})[0]


fp32 = run("shapes_prep.onnx", x_te)
acc32 = (fp32.argmax(1) == y_te).mean()
print(f"float32 model: accuracy {acc32:.3f}, "
      f"{os.path.getsize('shapes_prep.onnx')/1024:.1f} KB")

cases = [("good calibration set, MinMax", x_cal, CalibrationMethod.MinMax),
         ("good calibration set, Entropy", x_cal, CalibrationMethod.Entropy),
         ("good calibration set, Percentile", x_cal, CalibrationMethod.Percentile),
         ("blank calibration set, MinMax", x_bad, CalibrationMethod.MinMax)]

print(f"\n{'int8 variant':<34} {'KB':>6} {'acc':>6} {'agree':>7} {'max logit diff':>15}")
for i, (name, cal, method) in enumerate(cases):
    p = build(f"q{i}.onnx", cal, method)
    q = run(p, x_te)
    agree = (q.argmax(1) == fp32.argmax(1)).mean()
    print(f"{name:<34} {os.path.getsize(p)/1024:>6.1f} "
          f"{(q.argmax(1)==y_te).mean():>6.3f} {agree:>6.1%} "
          f"{np.abs(q-fp32).max():>15.4f}")
    os.remove(p)
Output
float32 model: accuracy 1.000, 18.6 KB

int8 variant                           KB    acc   agree  max logit diff
good calibration set, MinMax         10.3  1.000 100.0%          1.5497
Collecting tensor data and making histogram ...
Finding optimal threshold for each tensor using 'entropy' algorithm ...
Number of tensors : 9
Number of histogram bins : 128 (The number may increase depends on the data it collects)
Number of quantized bins : 128
good calibration set, Entropy        10.3  1.000 100.0%          1.5497
Collecting tensor data and making histogram ...
Finding optimal threshold for each tensor using 'percentile' algorithm ...
Number of tensors : 9
Number of histogram bins : 2048
Percentile : (0.0010000000000047748,99.999)
good calibration set, Percentile     10.3  1.000 100.0%          1.5537
blank calibration set, MinMax        10.3  0.335  33.5%         19.9496

What that table says

The blank calibration set produced a 33.5 percent accurate model. With three classes, that is chance. It is the same weights, the same architecture, the same quantisation code. The only thing that changed is which images were shown during calibration, and the model is now worthless.

It agrees with the float model on 33.5 percent of inputs and its logits are off by up to 19.95. Compare that with 1.55 for the good calibration sets — an order of magnitude larger. A blank image produces a narrow range at every layer, so real activations saturate at the top of that range and the layer's output becomes almost constant.

Nothing warned us. No exception, no NaN, no size difference. All four files are 10.3 KB. The only way this failure is ever found is by measuring accuracy after quantising.

All three calibration methods on good data are identical here, to three decimals. MinMax, Entropy and Percentile all give 1.000 accuracy and a maximum logit difference of about 1.55. That is honest and worth stating: on well-behaved activations with no outliers, the choice of method does not matter. It starts to matter when a layer has a heavy tail, which this toy model does not have. Do not carry "Entropy is better" as an article of faith; measure it on your model.

Float32 accuracy is 1.000 because the task is very easy. Three synthetic shapes on a plain background. That is deliberate — it means any accuracy loss you see is caused by quantisation and nothing else. On a real dataset expect the float baseline to be well below 1.000, and judge the INT8 model by the gap, not the absolute number.

The file shrank from 18.6 KB to 10.3 KB, not to a quarter. Weights are 8-bit now, but the QDQ format adds a QuantizeLinear and DequantizeLinear node with its own scale and zero-point around each quantised tensor, and this model is mostly a fully-connected layer. Larger models get closer to the theoretical four times.

The knobs that matter

quant_format=QuantFormat.QDQ inserts explicit quantise and dequantise nodes into the graph. This is what TensorRT 10 and later require, and it is what makes the quantisation decisions visible and portable rather than an internal detail of one runtime. The older QOperator format bakes the operation into fused integer ops and is more runtime-specific.

per_channel=True gives each output channel of a convolution its own weight scale instead of one scale for the whole tensor. For depthwise convolutions this is not optional — per-tensor weight quantisation of a depthwise layer is a well-known accuracy cliff, because channel magnitudes in depthwise layers vary enormously.

activation_type=QuantType.QUInt8 with weight_type=QuantType.QInt8 is the standard asymmetric-activations, symmetric-weights pairing. Activations after a ReLU are non-negative, so an unsigned type with a zero point uses the whole range; weights are roughly centred, so a signed symmetric type avoids a zero-point term in the accumulation.

quant_pre_process folds constants and runs shape inference. Skip it and the quantiser emits warnings and produces a worse model, because it cannot see the shapes it needs.

The workflow that catches problems

python
# 1. Baseline, before anything else.
fp32_metric = evaluate(fp32_model, validation_set)

# 2. Calibrate on data drawn from production, never from training only.
calibration = sample_production_images(n=300)

# 3. Quantise, then measure the SAME metric on the SAME set.
int8_metric = evaluate(int8_model, validation_set)

# 4. Also measure per-layer disagreement, to find the layer that broke.
for name in layer_outputs:
    print(name, np.abs(fp32_acts[name] - int8_acts[name]).max())

Step four is the one that saves days. A single layer accounting for most of the error is common, and the fix is usually to leave that one layer in higher precision rather than to abandon INT8.

Common mistakes

Calibrating on training data when production looks different. Calibration ranges are a property of the input distribution. If production images are darker, or come from a different camera, the ranges are wrong. Sample from production.

Using too few images, or too many. Somewhere between 100 and 1,000 is the usual sweet spot. Below that the ranges are noisy; above it you are paying for a more precise estimate of something that only needs to be roughly right.

Quantising the first and last layers. The input layer sees raw pixel statistics and the output layer produces the logits you threshold. Both are cheap in compute and expensive in accuracy. Most toolkits let you exclude nodes by name; do it.

Forgetting depthwise convolutions. MobileNet-family architectures lose several points with per-tensor weight quantisation and recover almost all of it with per-channel. If your INT8 model collapsed and it has depthwise layers, check this first.

Assuming INT8 is faster. It is faster only where the hardware and the runtime have integer kernels for your ops. On a CPU without VNNI, or a GPU where the runtime falls back, an INT8 model can be slower than FP16. Measure. See latency and throughput.

Shipping without a rollback. Keep the float model deployable behind a flag. Quantisation problems often surface as a slow drift in a business metric rather than a crash.

Try it yourself

Change x_bad from blank images to images of one class only, and predict whether accuracy lands at 33 percent or somewhere in between. Then reduce x_cal from 128 images to 4 and see how much of the good result survives. Both experiments take under a minute and both change how carefully you will pick calibration data.

What to learn next

Researcher — Mathematics and papers.

Affine quantisation

A real tensor $x$ is represented by an integer tensor $q$ with a scale $s > 0$ and a zero point $z \in \mathbb{Z}$:

$$ q = \operatorname{clamp}!\left(\operatorname{round}!\left(\frac{x}{s}\right) + z,\; q_{\min},\; q_{\max}\right), \qquad \hat{x} = s\,(q - z) $$

For unsigned 8-bit, $q_{\min}=0$ and $q_{\max}=255$; for signed, $-128$ and $127$. Given a calibrated range $[\alpha, \beta]$:

$$ s = \frac{\beta - \alpha}{q_{\max} - q_{\min}}, \qquad z = q_{\min} - \operatorname{round}!\left(\frac{\alpha}{s}\right) $$

Symmetric quantisation forces $z = 0$ with $\beta = -\alpha = \max|x|$. This removes the cross-terms in the integer matrix multiply, which is why weights are almost always symmetric and per-channel, while activations after a ReLU are asymmetric and per-tensor.

The error decomposes into two parts: rounding error, bounded by $s/2$ and shrinking as the range narrows, and clipping error, zero inside the range and unbounded outside it. Calibration is the choice of $[\alpha, \beta]$ that balances them, and the two move in opposite directions.

Calibration methods

MethodRuleBehaviour
MinMax$[\min x, \max x]$ over the calibration setzero clipping, maximum rounding error; one outlier ruins it
Percentilee.g. the 99.99th percentile of $\lvert x\rvert$clips a fixed tail; the standard robust default
Entropy (KL)minimise $D_{\mathrm{KL}}(P \Vert Q)$ between the float histogram and the requantised oneNVIDIA's original method; strong on heavy-tailed activations
MSEminimise $\mathbb{E}\lVert x - \hat{x}\rVert^2$ by searching the thresholddirect, slower, frequently the best

The entropy method builds a histogram of $|x|$ (2048 bins in the original TensorRT implementation), then for each candidate threshold quantises the histogram into 128 bins and computes the divergence, taking the threshold that minimises it. The intuition is that the quantised distribution should carry as much information about the original as possible.

The runnable example shows all three agreeing exactly. That is the correct result for activations without heavy tails, and is a useful calibration of expectations: method choice matters when the distribution has outliers, and not otherwise.

Granularity

$$ \text{per-tensor} \;\subset\; \text{per-channel} \;\subset\; \text{per-group} \;\subset\; \text{per-token} $$

For convolutional weights, per-output-channel is nearly free at inference time — the scale folds into the output requantisation — and recovers most of the accuracy lost by per-tensor. Krishnamoorthi (2018) measured the effect directly, and it is largest for depthwise separable convolutions, where per-channel weight ranges span orders of magnitude. This is the technical reason MobileNets were, for years, the standard counterexample to "INT8 is free".

Per-token and per-group schemes belong to the large-language-model literature and do not transfer usefully to convolutional vision.

Post-training quantisation beyond plain calibration

Cross-layer equalisation (Nagel et al., 2019) exploits the positive scaling equivariance of ReLU: for consecutive layers, $f(sx) = s f(x)$ allows rescaling channel $i$ of layer $\ell$ by $s_i$ and channel $i$ of layer $\ell+1$ by $1/s_i$ without changing the function. Choosing $s_i$ to equalise channel ranges makes per-tensor quantisation behave almost like per-channel. It requires no data at all.

AdaRound (Nagel et al., 2020) shows that rounding-to-nearest is not the optimal rounding, and learns a per-weight rounding decision by minimising the layer's output error on a few hundred unlabelled images. Typically recovers a large share of the gap to quantisation-aware training, at a fraction of the cost.

BRECQ (Li et al., 2021) extends this to block-wise reconstruction with second-order information, and is the strongest data-free-ish PTQ family for CNNs.

Quantisation-aware training (Jacob et al., 2018) simulates quantisation in the forward pass and uses the straight-through estimator for the backward pass:

$$ \frac{\partial \hat{x}}{\partial x} \approx \mathbb{1}!\left[\alpha \le x \le \beta\right] $$

It is the accuracy ceiling and it costs a training run. LSQ (Esser et al., 2020) makes the scale itself a learned parameter and is the strongest variant. Reach for QAT only after PTQ plus AdaRound has been measured and found insufficient.

Toolchain state, 2026

The important change for TensorRT users: implicit quantisation via IInt8EntropyCalibrator2 is deprecated from TensorRT 10.1 in favour of explicit quantisation, where QuantizeLinear and DequantizeLinear nodes appear in the ONNX graph and carry the scales. The builder no longer infers ranges; your export pipeline decides them.

That makes the workflow: quantise with a graph-level tool that emits Q/DQ, then let TensorRT compile the resulting graph. NVIDIA's supported tool is TensorRT Model Optimizer (nvidia-modelopt). ONNX Runtime's quantize_static with QuantFormat.QDQ, as used above, produces the same shape of graph and is the version you can run on a laptop.

Elsewhere: PyTorch's torch.ao.quantization has moved to the torch.export-based flow (prepare_pt2e / convert_pt2e), replacing the older eager and FX modes. OpenVINO uses NNCF, whose nncf.quantize() is now the recommended entry point after create_compressed_model() was deprecated. TensorFlow Lite keeps its representative-dataset converter API.

Reporting

State, every time: the calibration set's size and provenance, the method, the granularity, which layers were excluded, the metric before and after on the same held-out set, and the measured latency on the target hardware. A quantisation result without the calibration provenance is not reproducible, and as the blank-calibration row above shows, provenance is the variable that dominates everything else.

Papers

What to learn next