ONNX
ONNX is a single open file format for trained models, so a model built in PyTorch can run in C++, Java, JavaScript or on a phone without shipping PyTorch with it.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
ONNX is one file format that every framework can write and almost every device can read.
The name stands for Open Neural Network Exchange. Exchange is the important word.
The analogy you have already lived
You wrote something in Word and needed to send it to someone who does not have Word. So you exported a PDF.
The PDF opens on their phone, on a shared computer, in a browser. Nobody needed your software. The document stopped depending on the tool that made it.
ONNX is the PDF of trained models.
Why it exists
A model trained in PyTorch is a Python object. To run it, you need Python, and you need PyTorch, which is over a gigabyte installed.
That is fine on a laptop. It is impossible on a phone, in a browser, or inside a C++ camera app.
Before ONNX, every combination of framework and target needed its own converter. Researchers used one tool, product teams used another, and moving between them was a rewriting job.
ONNX cut the problem in half. Every framework writes one format. Every device reads that one format.
PyTorch ─┐ ┌─ phone (Android, iOS)
TensorFlow┤ ├─ browser
scikit-learn ├──► one .onnx file ──┤─ C++ / Java / C# app
others ─┘ └─ a Raspberry PiWhat is actually inside the file
Two things.
A list of operations, in order. Multiply by this matrix. Add this. Take the larger of this and zero. Each one is a standard operation with a written-down meaning.
The numbers. All the trained weights, stored alongside the operation list.
That is it. No Python, no classes, no code. A different program can read the list, do the operations in order, and get the same answers.
The runtime is the other half
The file alone does nothing. Something has to read it and do the work. That something is a runtime, and the common one is ONNX Runtime.
The important part for edge AI: the runtime is small. It ships as a library of a few megabytes, in C++, Java, C#, JavaScript and Python. Compare that with over a gigabyte for a full training framework.
The honest part
Not every model converts cleanly.
Some models do unusual things in Python. A loop whose length depends on the input. A custom operation. A shape decided while the program runs.
The exporter may fail on those. Worse, it may succeed and record only the one path it happened to take.
That is the failure everyone hits eventually. The export succeeds, and the exported model quietly ignores a branch of your code. Always check the exported model's answers against the original. The developer block shows how, in three lines.
Remember this
- ONNX is one file format every framework can write.
- The file holds a list of operations and the trained numbers, and no code.
- Always compare the exported model's answers with the original before shipping.
What to learn next
- TensorFlow Lite — the other dominant on-device format.
- Running models in the browser — the same ONNX file, executed in JavaScript.
- AI on a Raspberry Pi — the same file on a board that costs less than a textbook.
Developer — Code and libraries.
Setup
pip install torch onnx onnxruntime scikit-learnonnxruntime takes about 35 MB installed and is the only piece a device needs. On a Raspberry Pi or inside an Android app you install that, and not PyTorch — whose CPU-only build is around 500 MB installed, and whose CUDA build is several gigabytes. The Android and iOS builds of ONNX Runtime are smaller still, a few megabytes, because they drop the operators your model does not use.
Export, verify, benchmark
The three steps that matter are all in this script: exporting, checking the answers still match, and finding out whether it got faster.
import os, time, numpy as np, torch, torch.nn as nn, torch.nn.functional as F, onnxruntime as ort
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
torch.manual_seed(0); torch.set_num_threads(1)
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X / 16.0, y, test_size=0.3, random_state=0, stratify=y)
Xtr = torch.tensor(Xtr, dtype=torch.float32); ytr = torch.tensor(ytr, dtype=torch.long)
Xte = torch.tensor(Xte, dtype=torch.float32)
model = nn.Sequential(nn.Linear(64, 256), nn.ReLU(), nn.Linear(256, 256), nn.ReLU(), nn.Linear(256, 10))
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for _ in range(30):
for i in torch.randperm(len(Xtr)).split(64):
opt.zero_grad(); F.cross_entropy(model(Xtr[i]), ytr[i]).backward(); opt.step()
model.eval() # eval mode matters: dropout and batchnorm get baked in as-is
torch.onnx.export(
model,
torch.zeros(1, 64), # a sample input; only its shape and dtype are used
"digits.onnx",
input_names=["pixels"], output_names=["scores"],
dynamic_axes={"pixels": {0: "batch"}, "scores": {0: "batch"}}, # any batch size, not only 1
opset_version=13,
)
print("digits.onnx : %.1f KB" % (os.path.getsize("digits.onnx") / 1024))
opts = ort.SessionOptions(); opts.intra_op_num_threads = 1
sess = ort.InferenceSession("digits.onnx", opts, providers=["CPUExecutionProvider"])
for i in sess.get_inputs(): print("input :", i.name, i.shape, i.type)
for o in sess.get_outputs(): print("output :", o.name, o.shape, o.type)
x = Xte.numpy()
onnx_scores = sess.run(None, {"pixels": x})[0]
with torch.no_grad():
torch_scores = model(Xte).numpy()
print("largest disagreement : %.2e" % np.abs(onnx_scores - torch_scores).max())
print("same digit on all %d images: %s" % (len(x), bool((onnx_scores.argmax(1) == torch_scores.argmax(1)).all())))
def bench(fn, n=400):
for _ in range(50): fn()
t = time.perf_counter()
for _ in range(n): fn()
return (time.perf_counter() - t) / n * 1000
one_t, one_n = Xte[:1], x[:1]
with torch.no_grad():
print("PyTorch : %.3f ms per image" % bench(lambda: model(one_t)))
print("ONNX Runtime : %.3f ms per image" % bench(lambda: sess.run(None, {"pixels": one_n})))digits.onnx : 332.7 KB input : pixels ['batch', 64] tensor(float) output : scores ['batch', 10] tensor(float) largest disagreement : 0.00e+00 same digit on all 540 images: True PyTorch : 0.022 ms per image ONNX Runtime : 0.012 ms per image
What each line of that output is telling you
332.7 KB, against 334.2 KB for the same weights saved by PyTorch. ONNX is not a compression format. It stores the same weights with slightly different bookkeeping. Anyone who tells you ONNX makes models smaller has confused it with quantisation.
['batch', 64]. The first dimension is a name, not a number, because of dynamic_axes. Leave that argument out and the model accepts exactly one image per call, forever. This is the single most common export mistake.
0.00e+00. Bit-for-bit identical outputs on all 540 images. Do not expect a clean zero on every model — anything under about 1e-5 is normal and fine, because ONNX Runtime is free to fuse and reorder operations. What you are checking is that nothing structural was lost.
0.022 ms against 0.012 ms. Both timings are measurements on one machine, so yours will land somewhere else; the file sizes, the disagreement figure and the accuracy reproduce exactly. What generalises is the ratio: ONNX Runtime is about 1.8 times faster than PyTorch eager mode here, on the same weights and the same one thread. It gets that by fusing operations, pre-planning memory and skipping Python entirely. Note that this is your free speed-up: no quality was traded away.
Making it smaller too
ONNX Runtime ships its own quantiser. It reads an ONNX file and writes a smaller one.
import os, time, onnxruntime as ort
from onnxruntime.quantization import quantize_dynamic, QuantType
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
quantize_dynamic("digits.onnx", "digits_int8.onnx", weight_type=QuantType.QInt8)
X, y = load_digits(return_X_y=True)
_, Xte, _, yte = train_test_split(X / 16.0, y, test_size=0.3, random_state=0, stratify=y)
Xte = Xte.astype("float32")
print("file size accuracy ms/image")
for path in ["digits.onnx", "digits_int8.onnx"]:
opts = ort.SessionOptions(); opts.intra_op_num_threads = 1
s = ort.InferenceSession(path, opts, providers=["CPUExecutionProvider"])
acc = (s.run(None, {"pixels": Xte})[0].argmax(1) == yte).mean()
one = Xte[:1]
for _ in range(50): s.run(None, {"pixels": one})
t = time.perf_counter()
for _ in range(400): s.run(None, {"pixels": one})
ms = (time.perf_counter() - t) / 400 * 1000
print("%-18s %6.1f KB %.4f %.3f" % (path, os.path.getsize(path) / 1024, acc, ms))WARNING:root:Please consider to run pre-processing before quantization. Refer to example: https://github.com/microsoft/onnxruntime-inference-examples/blob/main/quantization/image_classification/cpu/ReadMe.md file size accuracy ms/image digits.onnx 332.7 KB 0.9648 0.012 digits_int8.onnx 87.7 KB 0.9648 0.012
The warning is expected and harmless here. It is asking you to run ONNX Runtime's shape-inference and optimisation pass first, which matters for convolutional models and does nothing for three dense layers.
3.8 times smaller, identical accuracy, identical speed. The ms/image column is again a measurement on one machine; the sizes and accuracies are not. Compare that with the PyTorch dynamic quantisation in quantisation in practice, which made the same model five times slower. The runtime matters as much as the technique.
Execution providers
One ONNX file, many backends. ONNX Runtime calls them execution providers, and you pass them in preference order:
sess = ort.InferenceSession("digits.onnx", providers=["CPUExecutionProvider"])
print(sess.get_providers())['CPUExecutionProvider']
On other machines the list is longer: CUDAExecutionProvider on an NVIDIA GPU, CoreMLExecutionProvider on a Mac or iPhone, NnapiExecutionProvider on Android, QNNExecutionProvider on Qualcomm hardware, WebGpuExecutionProvider in a browser. The runtime places each operation on the first provider that supports it and leaves the rest on the CPU.
Two consequences worth knowing. A provider being listed does not mean it ran your model — partial fallback to CPU is silent and common. And a provider that supports every operation can still be slower for a small model, because moving data to an accelerator costs more than the work saved.
Common mistakes
Forgetting model.eval() before export. Dropout and batch-norm are exported in whatever mode the model is in. Exporting in training mode gives you a model that behaves randomly.
Omitting dynamic_axes. You get a model locked to one batch size. It works in testing and fails the first time someone sends two inputs.
Not comparing outputs. Tracing follows the one path your sample input took. if x.sum() > 0: in a forward method becomes whichever branch happened to run. The comparison in the script above is what catches this.
An opset mismatch. opset_version is the version of the ONNX operation set. Old runtimes cannot read new opsets. Pick the lowest version your model needs and check what your target device supports.
Exporting the checkpoint instead of the model. torch.save(model.state_dict()) is a PyTorch file, not an ONNX file. It still needs PyTorch to load.
Try it yourself
Delete the dynamic_axes=... line and rerun. The export succeeds, and sess.get_inputs()[0].shape becomes [1, 64] instead of ['batch', 64]. Then feed the session all 540 images at once:
[ONNXRuntimeError] : 2 : INVALID_ARGUMENT : Got invalid dimensions for input: pixels for the following indices index: 0 Got: 540 Expected: 1 Please fix either the inputs/outputs or the model.
That is what a locked batch dimension looks like in production. It passes every test written with one image.
What to learn next
- TensorFlow Lite — the other dominant on-device format.
- Running models in the browser — the same ONNX file, executed in JavaScript.
- AI on a Raspberry Pi — the same file on a board that costs less than a textbook.
Researcher — Mathematics and papers.
The specification
ONNX defines a serialisation format and an operator set. A model is a protobuf ModelProto containing a GraphProto: a directed acyclic graph whose nodes are operator invocations and whose edges are named tensors, plus initializer entries holding the trained weights.
Three version numbers travel with every model and all three matter:
- IR version — the protobuf schema version.
- Opset version — the version of the operator set. Operators have versioned semantics;
Resize-11andResize-13differ in behaviour, not only in signature. - Producer version — informational, and the first thing to check when a conversion misbehaves.
Operators live in domains. The default ai.onnx domain holds roughly 190 operators; ai.onnx.ml covers classical machine learning primitives such as tree ensembles and scalers, which is how scikit-learn models convert through skl2onnx.
Tracing versus scripting
torch.onnx.export historically ran torch.jit.trace: execute the model once on a sample input and record the operations that fired. The consequences are exact and worth stating precisely.
- Data-dependent control flow is resolved, not preserved. A branch on a tensor value records only the taken branch.
- Python-level loops are unrolled to their traced length.
- Shapes are specialised unless declared dynamic.
- Side effects in Python are invisible to the trace.
Scripting (torch.jit.script) preserves control flow by compiling a subset of Python, but accepts a narrower language. PyTorch 2.x adds dynamo=True, which exports through TorchDynamo's graph capture and handles more programs. The failure mode is unchanged in kind: an exported graph is a graph, and anything that was dynamic Python is either captured explicitly or frozen.
The verification step is therefore not a nicety. It is the only mechanism that catches silent specialisation, and it should compare outputs across several inputs that exercise different branches, not one.
Graph optimisation
ONNX Runtime applies transformations at three levels before execution.
- Basic: constant folding, redundant node elimination, identity removal.
- Extended: operator fusion —
Conv + BatchNorm + ReLUinto one kernel,MatMul + AddintoGemm, the several nodes of a GELU or LayerNorm into a single fused op. - Layout: transposing weights into the memory order the target kernels prefer, typically NCHW to NHWC for CPU convolution.
Fusion is where most of the measured speed-up over eager frameworks comes from, and it is why exported outputs can differ from the original in the last few bits: fusing changes the order of floating-point operations, and floating-point addition is not associative.
Optimised graphs can be serialised (sess_options.optimized_model_filepath) so the cost is paid once at build time rather than at every application start — a meaningful saving on a device where model load time is a visible part of the user experience.
Quantisation in ONNX Runtime
Two representations, and the difference is operational rather than mathematical.
QDQ format inserts explicit QuantizeLinear and DequantizeLinear node pairs around operations. The graph remains float-typed; the runtime fuses matching Q/DQ pairs into integer kernels where it can. This is the portable representation, and it is what NNAPI, QNN and TensorRT consume.
Operator-oriented format replaces nodes with quantised equivalents such as QLinearConv and QLinearMatMul directly. More compact, less portable.
quantize_dynamic computes activation ranges at runtime and needs no calibration data. quantize_static requires a CalibrationDataReader supplying representative inputs, and produces a model that can run entirely in integer arithmetic — which is what an integer-only NPU requires. The recommended preprocessing step (onnxruntime.quantization.shape_inference.quant_pre_process) runs symbolic shape inference and basic optimisation first; skipping it on convolutional models frequently leaves quantisation unable to fuse, which is what the warning in the developer block is about.
Where ONNX does not reach
- Training. ONNX Runtime Training exists, but the format's centre of gravity is inference. Model definitions for training remain framework-specific.
- Very new operators. An architecture published this month may use an operation with no ONNX equivalent, requiring a custom op registered in the runtime, which destroys the portability that was the point.
- Dynamic-shape-heavy models. Detection heads with data-dependent output counts, and beam search with data-dependent stopping, export awkwardly and often lose their optimisation opportunities.
- Apple's platform. Core ML remains the route to the Neural Engine. ONNX Runtime's Core ML execution provider bridges this, with partial operator coverage. See Core ML.
Reading
- The ONNX operator specification — github.com/onnx/onnx/blob/main/docs/Operators.md
- ONNX Runtime performance tuning documentation — onnxruntime.ai/docs/performance
- Chen et al., TVM: An Automated End-to-End Optimizing Compiler for Deep Learning, OSDI 2018 — arxiv.org/abs/1802.04799 — the compiler-based alternative to a fixed runtime.
- Reed et al., Torch.fx: Practical Program Capture and Transformation for Deep Learning in Python, MLSys 2022 — on why graph capture from Python is hard.
What to learn next
- TensorFlow Lite — the other dominant on-device format.
- Running models in the browser — the same ONNX file, executed in JavaScript.
- AI on a Raspberry Pi — the same file on a board that costs less than a textbook.