OpenVINO and CPU inference
A lot of vision models do not need a GPU at all, and OpenVINO makes an ordinary Intel CPU fast enough to serve them at a fraction of the cost.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
OpenVINO recompiles a model to run fast on an ordinary processor. For many vision models, that is all the hardware you need.
The analogy you have lived
You have a car and a bicycle. For a two-kilometre errand through city traffic, the bicycle wins. Not because bicycles are faster, but because you skip the parking, the fuel queue and the one-way roads.
A graphics card is the car. It is genuinely faster for heavy work. It also comes with parking: a special machine, drivers to match, a much larger bill, and a queue for capacity.
For a small model handling a modest number of requests, the processor you already own gets there first.
When a processor is the right answer
Your model is small. A detector built for phones, a classifier over small crops, an OCR head. These finish in a few milliseconds on a processor.
Your traffic is modest. A few requests a second does not fill a graphics card. You would be paying for a card that idles.
Your model must run on a machine somebody already owns. A shop counter, a factory floor, an office desktop. Nobody is installing a graphics card there.
Your load is spiky. Processors scale up and down easily and cheaply. Cards do not.
What OpenVINO does about it
Left alone, most model files run on a processor in a general, cautious way. OpenVINO rewrites the model to suit the processor it finds.
It joins steps together, so results stay in fast memory instead of going back to main memory between each one.
It uses the processor's wide instructions, which handle many numbers at once instead of one at a time. Modern Intel chips have instructions designed specifically for this kind of arithmetic.
It arranges the data the way the processor prefers. That sounds minor and is not. Memory layout decides how much time the chip spends waiting.
It asks you what you want. Lowest delay for a single request, or the most requests per second overall? Those need different arrangements, and you pick one.
The two goals, and why they conflict
Ask for low delay and OpenVINO puts the whole machine behind one request. It finishes as soon as possible. While it runs, nothing else does.
Ask for high volume and it splits the machine into several independent workers, each handling its own request. Any one request takes longer. Far more get done per minute.
Neither is correct in general. A camera that must react in fifty milliseconds needs the first. A nightly job scoring a million stored photos needs the second.
Where you have seen this
- Cameras at a shop entrance counting visitors on a small box under the counter.
- Document scanners that read forms on the office machine, not in a cloud.
- Quality checks on a factory line, on the industrial computer already installed.
- Any product where sending customer images to a server is not allowed.
That last one matters more each year. Running on the machine means the pictures never leave it.
The honest part
There is a real ceiling. Large models, high resolutions and heavy traffic need a graphics card, and no amount of processor optimisation changes that.
The useful question is not which is better. It is where your model sits. Measure your model on the processor you would actually deploy on, before assuming you need anything more.
Remember this
- Many vision models are small enough to serve on an ordinary processor.
- OpenVINO joins steps, uses wide instructions, and lays memory out for the chip.
- Choose low delay or high volume deliberately; you cannot have both.
What to learn next
- ONNX — the format this whole workflow starts from.
- Latency and throughput — the trade-off the performance hints expose.
- Model serving — putting a compiled model behind an API.
Developer — Code and libraries.
Setup
pip install openvino # tested with 2026.3.1
pip install onnx onnxruntime torch # to produce and cross-check the modelOpenVINO reads ONNX directly, so anything you exported for TensorRT works here unchanged.
Step one: make a model and a reference answer
Run this in an environment with PyTorch and ONNX Runtime. It reuses the export from INT8 calibration, which produced shapes_prep.onnx.
import numpy as np, onnxruntime as ort
rng = np.random.default_rng(0)
x = rng.random((32, 1, 32, 32)).astype(np.float32)
np.save("ov_x.npy", x)
s = ort.InferenceSession("shapes_prep.onnx", providers=["CPUExecutionProvider"])
np.save("ov_ref.npy", s.run(None, {"input": x})[0])
print("saved", x.shape)saved (32, 1, 32, 32)
Step two: load, verify, and measure
import time, statistics
import numpy as np
import openvino as ov
core = ov.Core()
print("OpenVINO", ov.__version__)
print("devices :", core.available_devices)
print("CPU :", core.get_property("CPU", "FULL_DEVICE_NAME"))
# OpenVINO reads ONNX directly. convert_model turns it into OpenVINO IR,
# which loads faster and can be saved next to your service.
model = core.read_model("shapes_prep.onnx")
print("\ninputs :", [(i.any_name, i.partial_shape) for i in model.inputs])
print("outputs:", [(o.any_name, o.partial_shape) for o in model.outputs])
ov.save_model(model, "shapes_ov.xml") # writes shapes_ov.xml + .bin
x = np.load("ov_x.npy")
ref = np.load("ov_ref.npy") # what ONNX Runtime produced
compiled = core.compile_model(model, "CPU")
out = compiled(x)[compiled.output(0)]
print(f"\noutput shape {out.shape}, "
f"largest difference from ONNX Runtime: {np.abs(out - ref).max():.2e}")
def bench(cm, batch, n=50):
req = cm.create_infer_request()
xb = x[:batch]
for _ in range(10):
req.infer({0: xb}) # warm up before timing anything
ts = []
for _ in range(n):
t0 = time.perf_counter()
req.infer({0: xb})
ts.append((time.perf_counter() - t0) * 1000)
return statistics.median(ts)
print("\nlatency by batch size, PERFORMANCE_HINT LATENCY:")
lat = core.compile_model(model, "CPU", {"PERFORMANCE_HINT": "LATENCY"})
for b in (1, 4, 16, 32):
ms = bench(lat, b)
print(f" batch {b:>2}: {ms:7.3f} ms ({b*1000/ms:8.0f} images/second)")
thr = core.compile_model(model, "CPU", {"PERFORMANCE_HINT": "THROUGHPUT"})
print("\nsame model compiled with PERFORMANCE_HINT THROUGHPUT:")
print(" optimal number of infer requests:",
thr.get_property("OPTIMAL_NUMBER_OF_INFER_REQUESTS"))
for b in (1, 32):
ms = bench(thr, b)
print(f" batch {b:>2}: {ms:7.3f} ms ({b*1000/ms:8.0f} images/second)")OpenVINO 2026.3.1-22476-759c5a6ab8c-releases/2026/3
devices : ['CPU', 'GPU.0', 'GPU.1']
CPU : 13th Gen Intel(R) Core(TM) i9-13900K
inputs : [('input', <PartialShape: [?,1,32,32]>)]
outputs: [('logits', <PartialShape: [?,3]>)]
output shape (32, 3), largest difference from ONNX Runtime: 6.68e-06
latency by batch size, PERFORMANCE_HINT LATENCY:
batch 1: 0.098 ms ( 10204 images/second)
batch 4: 0.245 ms ( 16320 images/second)
batch 16: 0.227 ms ( 70361 images/second)
batch 32: 0.333 ms ( 96082 images/second)
same model compiled with PERFORMANCE_HINT THROUGHPUT:
optimal number of infer requests: 24
batch 1: 0.183 ms ( 5461 images/second)
batch 32: 0.690 ms ( 46357 images/second)Timings are hardware-dependent and move between runs. These come from one run on an Intel Core i9-13900K. Across five runs the millisecond figures varied by a factor of two or more with background load, while every comparison drawn below held in all of them. Treat the comparisons as the finding and the milliseconds as an illustration.
Reading that output
The answer matches ONNX Runtime to 6.68e-06. That is float32 accumulation-order noise, not a correctness problem. It is also the check to run first, every time: a fast wrong answer is not an optimisation.
Batching gives an enormous throughput win for a tiny latency cost. Batch 1 takes 0.098 ms; batch 32 takes 0.333 ms. Thirty-two times the work for under four times the time — 10,204 images per second becomes 96,082. A single 32x32 image cannot fill the vector units or hide memory latency, so most of that 0.098 ms is fixed overhead that batching amortises.
OPTIMAL_NUMBER_OF_INFER_REQUESTS is 24. Under the throughput hint, OpenVINO partitioned this 24-core CPU into 24 independent streams. That number is the one you should size your request queue against; running fewer leaves cores idle, running more adds queueing delay without adding work.
The throughput hint is not automatically faster. Driven by one request at a time it is slower at both batch sizes: 0.690 ms against 0.333 ms at batch 32, and 46,357 images per second against 96,082. Each stream now gets a fraction of the machine, so a single batch runs on a fraction of the cores. The throughput hint pays off when you have many concurrent requests to keep those 24 streams busy, and costs you when you have one.
This is the most common OpenVINO mistake: setting THROUGHPUT and benchmarking with one request at a time. You measure the cost and none of the benefit.
Devices list GPU.0 and GPU.1. OpenVINO enumerates Intel integrated and discrete graphics. It does not target NVIDIA cards; for those, see the TensorRT lesson.
The command-line benchmark
benchmark_app ships with the package and is the tool to reach for before writing any timing code of your own.
benchmark_app -m shapes_ov.xml -d CPU -hint latency -shape "[1,1,32,32]" -t 5[Step 10/11] Measuring performance (Start inference synchronously, limits: 5000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). [ INFO ] First inference took 0.35 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] [ INFO ] Count: 49081 iterations [ INFO ] Duration: 5000.86 ms [ INFO ] Latency: [ INFO ] Median: 0.06 ms [ INFO ] Average: 0.08 ms [ INFO ] Min: 0.03 ms [ INFO ] Max: 16.92 ms [ INFO ] Throughput: 9814.51 FPS
Note the gap between the median of 0.06 ms and the maximum of 16.92 ms. Across repeated runs on this machine that ratio sat between roughly 75 and 280 times, and it never went away. The tail is the operating system scheduling other work, and it is what a user with a hard deadline actually experiences. Any latency claim quoting only a mean or a median is hiding it.
The median here is close to the batch-1 figure measured from Python, and generally lands a little below it, because benchmark_app excludes the Python call overhead the earlier script includes. Both are correct answers to different questions; the Python one is closer to what your service will see.
Preprocessing inside the model
OpenVINO can fold resizing, layout conversion and normalisation into the compiled model, which removes an entire class of preprocessing mismatch and runs faster than doing it in NumPy.
from openvino.preprocess import PrePostProcessor, ColorFormat
from openvino import Type, Layout
ppp = PrePostProcessor(model)
ppp.input().tensor() \
.set_element_type(Type.u8) \
.set_layout(Layout("NHWC")) \
.set_color_format(ColorFormat.BGR)
ppp.input().preprocess() \
.convert_element_type(Type.f32) \
.convert_color(ColorFormat.RGB) \
.mean([123.675, 116.28, 103.53]) \
.scale([58.395, 57.12, 57.375])
ppp.input().model().set_layout(Layout("NCHW"))
model = ppp.build()No output block for this fragment: it is a graph transformation whose effect shows up in the compiled model, and the shapes here would need adjusting for the single-channel toy model above. On a real three-channel model it lets you feed raw uint8 BGR straight from cv2.imread with no Python preprocessing at all.
Common mistakes
Benchmarking without warming up. The first inference compiles kernels and allocates buffers. The script above discards ten iterations for this reason.
Setting THROUGHPUT and sending one request at a time. Demonstrated above: it is slower. Match the hint to your traffic shape.
Recompiling per request. compile_model is expensive and create_infer_request is cheap. Compile once at startup, keep a pool of requests.
Assuming int8 helps everywhere. It helps a great deal on CPUs with VNNI or AMX instructions and much less on older ones. Quantise with NNCF (nncf.quantize() — create_compressed_model() is deprecated), then measure on the deployment CPU.
Using the deprecated openvino.runtime namespace. It was marked for removal in 2026.0. Import openvino directly; every example above does.
Comparing CPU and GPU numbers without including preprocessing. As the GPU preprocessing lesson measures, decode and resize frequently dominate. A CPU pipeline that avoids a host-to-device transfer can beat a GPU pipeline end to end even when the raw model is slower.
Try it yourself
Compile the same model with {"PERFORMANCE_HINT": "THROUGHPUT"} and drive it with 24 concurrent InferRequest objects using start_async and wait, instead of one synchronous request. That is the configuration the throughput hint was designed for, and it is the only way to see it win.
What to learn next
- ONNX — the format this whole workflow starts from.
- Latency and throughput — the trade-off the performance hints expose.
- Model serving — putting a compiled model behind an API.
Researcher — Mathematics and papers.
What the CPU plugin actually does
OpenVINO's CPU plugin compiles the intermediate representation into a graph of oneDNN primitives. The transformations that matter:
Operator fusion. Convolution with bias, batch norm and an activation collapses into a single primitive. As in TensorRT, the arithmetic saving is small and the saving in memory traffic is large — a fused chain reads and writes the activation tensor once instead of four times.
Layout selection. The plugin reformats tensors into blocked layouts such as nChw8c or nChw16c, interleaving channel groups so that the innermost loop maps onto a vector register. Reorder nodes are inserted at layout boundaries; a graph that ping-pongs between layouts pays for every crossing, which is why an unsupported op in the middle of a network costs more than its own runtime.
Instruction selection. AVX2, AVX-512, VNNI (int8 dot-product with 32-bit accumulate) and AMX (tile-based matrix units on Sapphire Rapids and later). VNNI is roughly a 4x arithmetic-throughput multiplier for int8 convolution over an AVX-512 float path; AMX is larger again for bf16 and int8. The presence or absence of these instructions is the dominant term in whether quantisation is worth anything on a given CPU.
Threading. Built on oneTBB. Under LATENCY the whole thread pool cooperates on one inference; under THROUGHPUT the plugin partitions cores into streams, each running an independent request. On hybrid CPUs — the i9-13900K above has performance and efficiency cores — the plugin's stream assignment is also making a core-type decision, which is one reason its OPTIMAL_NUMBER_OF_INFER_REQUESTS is worth trusting over a hand-set value.
Latency and throughput are a queueing problem
For $s$ streams each serving requests at rate $\mu$, offered load $\lambda$, and per-request service time $T_s$:
$$ \text{throughput} \approx s \cdot \frac{1}{T_s}, \qquad \text{latency} \approx T_s + \text{queueing delay} $$
$T_s$ under $s$ streams is larger than $T_1$ under one, because each stream has fewer cores. The measurement above quantifies it: 0.333 ms with the whole machine on one batch of 32, and 0.690 ms with the machine split into 24 streams. So splitting is a loss unless $\lambda$ is high enough to keep the streams occupied.
The crossover is where the arrival rate exceeds what a single latency-optimised stream can serve. Below it, use LATENCY. Above it, use THROUGHPUT and size your concurrency to OPTIMAL_NUMBER_OF_INFER_REQUESTS. Publishing a single "images per second" figure without stating which hint and what concurrency produced it is not a measurement.
Tail latency
The benchmark_app output above shows a median of 0.06 ms against a maximum of 16.92 ms. On a general-purpose operating system this is normal and comes from scheduler preemption, page faults, frequency scaling, and on hybrid CPUs from migration between performance and efficiency cores.
For a hard real-time budget the mitigations are outside the inference library: pin threads to cores, isolate those cores from the general scheduler, disable frequency scaling, and use huge pages. Report P50, P95 and P99 always; a P99 that is 75 times the P50 is a capacity-planning fact, not noise.
Choosing between runtimes
| Runtime | Strength | Weakness |
|---|---|---|
| OpenVINO | best on Intel CPUs; preprocessing in-graph; mature int8 | Intel-oriented; no NVIDIA target |
| ONNX Runtime | portable, many execution providers, easy fallback | fewer CPU-specific tricks by default |
| TensorRT | fastest on NVIDIA GPUs | NVIDIA only; engine locked to hardware and version |
torch.compile | no export step; good on transformers | weaker on convolutional CPU inference |
| TFLite / XNNPACK | mobile and ARM | limited operator coverage |
For ARM CPUs — Graviton, Apple silicon, Raspberry Pi — OpenVINO is not the right answer; ONNX Runtime with XNNPACK or ARM Compute Library is. See Raspberry Pi AI.
When CPU serving is the correct engineering decision
Beyond cost, three arguments are frequently decisive and rarely stated:
- Data residency. Inference on the machine that captured the image means the image never transits a network. For medical, financial and several jurisdictions' privacy regimes this is a requirement, not a preference.
- Operational simplicity. No driver-version matrix, no CUDA compatibility, no accelerator-aware scheduling. The failure surface is much smaller.
- Elasticity. CPU capacity is available on demand at fine granularity; GPU capacity is not, and a service with spiky traffic pays for idle accelerators.
The counterargument is equally concrete: above a certain model size and request rate, no amount of CPU optimisation closes the gap, and the honest way to find your crossover is to measure both on your own model.
References
- OpenVINO documentation, 2026 release — docs.openvino.ai
- oneDNN developer guide — layout and primitive selection
- Intel AVX-512 VNNI and AMX programming references
- NNCF (Neural Network Compression Framework) — github.com/openvinotoolkit/nncf
What to learn next
- ONNX — the format this whole workflow starts from.
- Latency and throughput — the trade-off the performance hints expose.
- Model serving — putting a compiled model behind an API.