Edge and On-device AI

ONNX

ONNX is a single open file format for trained models, so a model built in PyTorch can run in C++, Java, JavaScript or on a phone without shipping PyTorch with it.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. What is actually inside the file
  5. The runtime is the other half
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

ONNX is one file format that every framework can write and almost every device can read.

The name stands for Open Neural Network Exchange. Exchange is the important word.

The analogy you have already lived

You wrote something in Word and needed to send it to someone who does not have Word. So you exported a PDF.

The PDF opens on their phone, on a shared computer, in a browser. Nobody needed your software. The document stopped depending on the tool that made it.

ONNX is the PDF of trained models.

Why it exists

A model trained in PyTorch is a Python object. To run it, you need Python, and you need PyTorch, which is over a gigabyte installed.

That is fine on a laptop. It is impossible on a phone, in a browser, or inside a C++ camera app.

Before ONNX, every combination of framework and target needed its own converter. Researchers used one tool, product teams used another, and moving between them was a rewriting job.

ONNX cut the problem in half. Every framework writes one format. Every device reads that one format.

   PyTorch  ─┐                        ┌─ phone (Android, iOS)
   TensorFlow┤                        ├─ browser
   scikit-learn ├──►  one .onnx file ──┤─ C++ / Java / C# app
   others   ─┘                        └─ a Raspberry Pi

What is actually inside the file

Two things.

A list of operations, in order. Multiply by this matrix. Add this. Take the larger of this and zero. Each one is a standard operation with a written-down meaning.

The numbers. All the trained weights, stored alongside the operation list.

That is it. No Python, no classes, no code. A different program can read the list, do the operations in order, and get the same answers.

The runtime is the other half

The file alone does nothing. Something has to read it and do the work. That something is a runtime, and the common one is ONNX Runtime.

The important part for edge AI: the runtime is small. It ships as a library of a few megabytes, in C++, Java, C#, JavaScript and Python. Compare that with over a gigabyte for a full training framework.

The honest part

Not every model converts cleanly.

Some models do unusual things in Python. A loop whose length depends on the input. A custom operation. A shape decided while the program runs.

The exporter may fail on those. Worse, it may succeed and record only the one path it happened to take.

That is the failure everyone hits eventually. The export succeeds, and the exported model quietly ignores a branch of your code. Always check the exported model's answers against the original. The developer block shows how, in three lines.

Remember this

  • ONNX is one file format every framework can write.
  • The file holds a list of operations and the trained numbers, and no code.
  • Always compare the exported model's answers with the original before shipping.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch onnx onnxruntime scikit-learn

onnxruntime takes about 35 MB installed and is the only piece a device needs. On a Raspberry Pi or inside an Android app you install that, and not PyTorch — whose CPU-only build is around 500 MB installed, and whose CUDA build is several gigabytes. The Android and iOS builds of ONNX Runtime are smaller still, a few megabytes, because they drop the operators your model does not use.

Export, verify, benchmark

The three steps that matter are all in this script: exporting, checking the answers still match, and finding out whether it got faster.

to_onnx.py
import os, time, numpy as np, torch, torch.nn as nn, torch.nn.functional as F, onnxruntime as ort
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

torch.manual_seed(0); torch.set_num_threads(1)
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X / 16.0, y, test_size=0.3, random_state=0, stratify=y)
Xtr = torch.tensor(Xtr, dtype=torch.float32); ytr = torch.tensor(ytr, dtype=torch.long)
Xte = torch.tensor(Xte, dtype=torch.float32)

model = nn.Sequential(nn.Linear(64, 256), nn.ReLU(), nn.Linear(256, 256), nn.ReLU(), nn.Linear(256, 10))
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for _ in range(30):
    for i in torch.randperm(len(Xtr)).split(64):
        opt.zero_grad(); F.cross_entropy(model(Xtr[i]), ytr[i]).backward(); opt.step()
model.eval()                            # eval mode matters: dropout and batchnorm get baked in as-is

torch.onnx.export(
    model,
    torch.zeros(1, 64),                 # a sample input; only its shape and dtype are used
    "digits.onnx",
    input_names=["pixels"], output_names=["scores"],
    dynamic_axes={"pixels": {0: "batch"}, "scores": {0: "batch"}},   # any batch size, not only 1
    opset_version=13,
)
print("digits.onnx : %.1f KB" % (os.path.getsize("digits.onnx") / 1024))

opts = ort.SessionOptions(); opts.intra_op_num_threads = 1
sess = ort.InferenceSession("digits.onnx", opts, providers=["CPUExecutionProvider"])
for i in sess.get_inputs():  print("input       :", i.name, i.shape, i.type)
for o in sess.get_outputs(): print("output      :", o.name, o.shape, o.type)

x = Xte.numpy()
onnx_scores = sess.run(None, {"pixels": x})[0]
with torch.no_grad():
    torch_scores = model(Xte).numpy()
print("largest disagreement       : %.2e" % np.abs(onnx_scores - torch_scores).max())
print("same digit on all %d images: %s" % (len(x), bool((onnx_scores.argmax(1) == torch_scores.argmax(1)).all())))

def bench(fn, n=400):
    for _ in range(50): fn()
    t = time.perf_counter()
    for _ in range(n): fn()
    return (time.perf_counter() - t) / n * 1000

one_t, one_n = Xte[:1], x[:1]
with torch.no_grad():
    print("PyTorch      : %.3f ms per image" % bench(lambda: model(one_t)))
print("ONNX Runtime : %.3f ms per image" % bench(lambda: sess.run(None, {"pixels": one_n})))
Output
digits.onnx : 332.7 KB
input       : pixels ['batch', 64] tensor(float)
output      : scores ['batch', 10] tensor(float)
largest disagreement       : 0.00e+00
same digit on all 540 images: True
PyTorch      : 0.022 ms per image
ONNX Runtime : 0.012 ms per image

What each line of that output is telling you

332.7 KB, against 334.2 KB for the same weights saved by PyTorch. ONNX is not a compression format. It stores the same weights with slightly different bookkeeping. Anyone who tells you ONNX makes models smaller has confused it with quantisation.

['batch', 64]. The first dimension is a name, not a number, because of dynamic_axes. Leave that argument out and the model accepts exactly one image per call, forever. This is the single most common export mistake.

0.00e+00. Bit-for-bit identical outputs on all 540 images. Do not expect a clean zero on every model — anything under about 1e-5 is normal and fine, because ONNX Runtime is free to fuse and reorder operations. What you are checking is that nothing structural was lost.

0.022 ms against 0.012 ms. Both timings are measurements on one machine, so yours will land somewhere else; the file sizes, the disagreement figure and the accuracy reproduce exactly. What generalises is the ratio: ONNX Runtime is about 1.8 times faster than PyTorch eager mode here, on the same weights and the same one thread. It gets that by fusing operations, pre-planning memory and skipping Python entirely. Note that this is your free speed-up: no quality was traded away.

Making it smaller too

ONNX Runtime ships its own quantiser. It reads an ONNX file and writes a smaller one.

quantise_onnx.py
import os, time, onnxruntime as ort
from onnxruntime.quantization import quantize_dynamic, QuantType
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

quantize_dynamic("digits.onnx", "digits_int8.onnx", weight_type=QuantType.QInt8)

X, y = load_digits(return_X_y=True)
_, Xte, _, yte = train_test_split(X / 16.0, y, test_size=0.3, random_state=0, stratify=y)
Xte = Xte.astype("float32")

print("file                 size     accuracy   ms/image")
for path in ["digits.onnx", "digits_int8.onnx"]:
    opts = ort.SessionOptions(); opts.intra_op_num_threads = 1
    s = ort.InferenceSession(path, opts, providers=["CPUExecutionProvider"])
    acc = (s.run(None, {"pixels": Xte})[0].argmax(1) == yte).mean()
    one = Xte[:1]
    for _ in range(50): s.run(None, {"pixels": one})
    t = time.perf_counter()
    for _ in range(400): s.run(None, {"pixels": one})
    ms = (time.perf_counter() - t) / 400 * 1000
    print("%-18s %6.1f KB    %.4f      %.3f" % (path, os.path.getsize(path) / 1024, acc, ms))
Output
WARNING:root:Please consider to run pre-processing before quantization. Refer to example: https://github.com/microsoft/onnxruntime-inference-examples/blob/main/quantization/image_classification/cpu/ReadMe.md
file                 size     accuracy   ms/image
digits.onnx         332.7 KB    0.9648      0.012
digits_int8.onnx     87.7 KB    0.9648      0.012

The warning is expected and harmless here. It is asking you to run ONNX Runtime's shape-inference and optimisation pass first, which matters for convolutional models and does nothing for three dense layers.

3.8 times smaller, identical accuracy, identical speed. The ms/image column is again a measurement on one machine; the sizes and accuracies are not. Compare that with the PyTorch dynamic quantisation in quantisation in practice, which made the same model five times slower. The runtime matters as much as the technique.

Execution providers

One ONNX file, many backends. ONNX Runtime calls them execution providers, and you pass them in preference order:

python
sess = ort.InferenceSession("digits.onnx", providers=["CPUExecutionProvider"])
print(sess.get_providers())
Output
['CPUExecutionProvider']

On other machines the list is longer: CUDAExecutionProvider on an NVIDIA GPU, CoreMLExecutionProvider on a Mac or iPhone, NnapiExecutionProvider on Android, QNNExecutionProvider on Qualcomm hardware, WebGpuExecutionProvider in a browser. The runtime places each operation on the first provider that supports it and leaves the rest on the CPU.

Two consequences worth knowing. A provider being listed does not mean it ran your model — partial fallback to CPU is silent and common. And a provider that supports every operation can still be slower for a small model, because moving data to an accelerator costs more than the work saved.

Common mistakes

Forgetting model.eval() before export. Dropout and batch-norm are exported in whatever mode the model is in. Exporting in training mode gives you a model that behaves randomly.

Omitting dynamic_axes. You get a model locked to one batch size. It works in testing and fails the first time someone sends two inputs.

Not comparing outputs. Tracing follows the one path your sample input took. if x.sum() > 0: in a forward method becomes whichever branch happened to run. The comparison in the script above is what catches this.

An opset mismatch. opset_version is the version of the ONNX operation set. Old runtimes cannot read new opsets. Pick the lowest version your model needs and check what your target device supports.

Exporting the checkpoint instead of the model. torch.save(model.state_dict()) is a PyTorch file, not an ONNX file. It still needs PyTorch to load.

Try it yourself

Delete the dynamic_axes=... line and rerun. The export succeeds, and sess.get_inputs()[0].shape becomes [1, 64] instead of ['batch', 64]. Then feed the session all 540 images at once:

Output
[ONNXRuntimeError] : 2 : INVALID_ARGUMENT : Got invalid dimensions for input: pixels for the following indices
 index: 0 Got: 540 Expected: 1
 Please fix either the inputs/outputs or the model.

That is what a locked batch dimension looks like in production. It passes every test written with one image.

What to learn next

Researcher — Mathematics and papers.

The specification

ONNX defines a serialisation format and an operator set. A model is a protobuf ModelProto containing a GraphProto: a directed acyclic graph whose nodes are operator invocations and whose edges are named tensors, plus initializer entries holding the trained weights.

Three version numbers travel with every model and all three matter:

  • IR version — the protobuf schema version.
  • Opset version — the version of the operator set. Operators have versioned semantics; Resize-11 and Resize-13 differ in behaviour, not only in signature.
  • Producer version — informational, and the first thing to check when a conversion misbehaves.

Operators live in domains. The default ai.onnx domain holds roughly 190 operators; ai.onnx.ml covers classical machine learning primitives such as tree ensembles and scalers, which is how scikit-learn models convert through skl2onnx.

Tracing versus scripting

torch.onnx.export historically ran torch.jit.trace: execute the model once on a sample input and record the operations that fired. The consequences are exact and worth stating precisely.

  • Data-dependent control flow is resolved, not preserved. A branch on a tensor value records only the taken branch.
  • Python-level loops are unrolled to their traced length.
  • Shapes are specialised unless declared dynamic.
  • Side effects in Python are invisible to the trace.

Scripting (torch.jit.script) preserves control flow by compiling a subset of Python, but accepts a narrower language. PyTorch 2.x adds dynamo=True, which exports through TorchDynamo's graph capture and handles more programs. The failure mode is unchanged in kind: an exported graph is a graph, and anything that was dynamic Python is either captured explicitly or frozen.

The verification step is therefore not a nicety. It is the only mechanism that catches silent specialisation, and it should compare outputs across several inputs that exercise different branches, not one.

Graph optimisation

ONNX Runtime applies transformations at three levels before execution.

  • Basic: constant folding, redundant node elimination, identity removal.
  • Extended: operator fusion — Conv + BatchNorm + ReLU into one kernel, MatMul + Add into Gemm, the several nodes of a GELU or LayerNorm into a single fused op.
  • Layout: transposing weights into the memory order the target kernels prefer, typically NCHW to NHWC for CPU convolution.

Fusion is where most of the measured speed-up over eager frameworks comes from, and it is why exported outputs can differ from the original in the last few bits: fusing changes the order of floating-point operations, and floating-point addition is not associative.

Optimised graphs can be serialised (sess_options.optimized_model_filepath) so the cost is paid once at build time rather than at every application start — a meaningful saving on a device where model load time is a visible part of the user experience.

Quantisation in ONNX Runtime

Two representations, and the difference is operational rather than mathematical.

QDQ format inserts explicit QuantizeLinear and DequantizeLinear node pairs around operations. The graph remains float-typed; the runtime fuses matching Q/DQ pairs into integer kernels where it can. This is the portable representation, and it is what NNAPI, QNN and TensorRT consume.

Operator-oriented format replaces nodes with quantised equivalents such as QLinearConv and QLinearMatMul directly. More compact, less portable.

quantize_dynamic computes activation ranges at runtime and needs no calibration data. quantize_static requires a CalibrationDataReader supplying representative inputs, and produces a model that can run entirely in integer arithmetic — which is what an integer-only NPU requires. The recommended preprocessing step (onnxruntime.quantization.shape_inference.quant_pre_process) runs symbolic shape inference and basic optimisation first; skipping it on convolutional models frequently leaves quantisation unable to fuse, which is what the warning in the developer block is about.

Where ONNX does not reach

  • Training. ONNX Runtime Training exists, but the format's centre of gravity is inference. Model definitions for training remain framework-specific.
  • Very new operators. An architecture published this month may use an operation with no ONNX equivalent, requiring a custom op registered in the runtime, which destroys the portability that was the point.
  • Dynamic-shape-heavy models. Detection heads with data-dependent output counts, and beam search with data-dependent stopping, export awkwardly and often lose their optimisation opportunities.
  • Apple's platform. Core ML remains the route to the Neural Engine. ONNX Runtime's Core ML execution provider bridges this, with partial operator coverage. See Core ML.

Reading

What to learn next