AI glossary

ONNX

In one sentence ONNX is a common file format for trained models, letting a network trained in one framework run anywhere — servers, browsers, phones.

By Updated

ONNX (Open Neural Network Exchange) is a framework-neutral format for trained models: export once, run on anything that speaks it.

It is the PDF of machine learning. A document written in Word once required Word to open; PDF ended that — author in whatever you like, and every reader everywhere renders it identically. Models had the same lock-in: a PyTorch model needed PyTorch, heavy with training machinery, wherever it went. Exporting to ONNX detaches the model from its birthplace: the network becomes a standard file — a graph of operators with weights — that any compliant runtime can execute.

The main executor is ONNX Runtime, which runs the same file on Linux servers, Windows apps, phones, and in the browser via WebAssembly and WebGL — with hardware-specific "execution providers" (CUDA, TensorRT, CoreML, DirectML) slotting in underneath. Two wins beyond portability: the runtime is a lean inference engine, dropping the training-framework baggage, and it applies graph optimisations (fusing operations, constant folding) plus easy quantization, often making the exported model faster than the original. Hugging Face's Optimum packages this for transformers.

python
torch.onnx.export(model, example_input, "model.onnx")
# then, anywhere:
import onnxruntime as ort
sess = ort.InferenceSession("model.onnx")

Honest edges: an export can fail or silently diverge when a model uses exotic or dynamic operations, so always compare outputs after exporting. And for large LLMs specifically, dedicated stacks (vLLM, llama.cpp/GGUF, TensorRT-LLM) dominate; ONNX's home turf is classic vision, speech and tabular networks shipped into apps and edge devices.

Where to go next