AI on a Raspberry Pi
A Raspberry Pi runs real models on four ARM cores with no GPU, so the work is choosing a small model, quantising it, and being honest about frames per second.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A Raspberry Pi is a whole computer the size of a playing-card deck. It runs real models slowly, cheaply, and for years without attention.
The analogy you have already lived
Look at the clock on your wall. It does one job. It has done it since somebody hung it there. Nobody restarts it, updates it, or thinks about it.
Your phone could tell the time better. The wall clock wins because it is cheap, it is already there, and it never asks for anything.
A Raspberry Pi is that, for computing. You bolt it to a wall or hide it in a machine, and it does one job forever.
Why people use one for AI
It costs very little. A Pi plus a camera costs less than a mid-range phone. A computer with a graphics card costs far more.
You can leave it somewhere. In a cattle shed, on a factory line, inside a signboard, in a field. A laptop cannot live in those places. A Pi can.
It runs on a small amount of power. A few watts. It can run off a solar panel and a battery.
No cloud bill and no signal needed. Which brings back every reason from why run a model on the device.
What it is not
There is no graphics card in a Raspberry Pi. Nothing like the ones that train models.
A Pi has four small ARM processor cores, the same family as in a mid-range phone. Everything runs on those cores unless you add extra hardware.
So the question is never "can it run this model". It is "how many pictures a second, and is that enough for the job?"
Counting people entering a shop needs perhaps two pictures a second. That is comfortable. Following a cricket ball needs sixty. That is not going to happen.
What the work actually is
train on a laptop or in a free cloud notebook
|
export to ONNX <- one file, no framework attached
|
quantise it <- 8-bit weights, roughly 4x smaller
|
copy the file to the Pi
|
run it with a small runtime, on all four coresNothing is trained on the Pi. Training is a big-computer job. The Pi only answers questions.
The honest part
Everything takes longer than you expect the first time. Installing packages on a Pi is slow. Cameras have their own difficulties. Plan a weekend, not an evening.
Cheap power supplies cause strange faults. A Pi that is not getting enough current slows itself down on purpose to survive. It does not tell you loudly. It gets mysteriously slow, and people blame the model.
The memory card wears out. Writing logs to it constantly will kill it in months. Write logs to memory, or to a USB drive.
A Pi Zero has very little memory. Some models will not load at all. That is a hard stop, not a slow one.
Remember this
- A Pi has four small cores and no graphics card.
- The real question is frames per second, not whether it runs.
- Power supply and memory cause more failures than the model does.
What to learn next
- Latency and throughput — measuring speed so the numbers mean something.
- Power and thermal limits — why the tenth minute is slower than the first.
- Docker for ML — packaging what you built so it survives a reboot.
Developer — Code and libraries.
Setup
On the Pi, running 64-bit Raspberry Pi OS:
sudo apt update && sudo apt install -y python3-pip
pip install onnxruntime numpy psutilonnxruntime has official aarch64 wheels, so this is a download rather than a compile. Use the 64-bit OS — the 32-bit build is slower and has thinner package support.
The benchmark below runs unchanged on your laptop too, which is the point: measure it in both places and compare.
A camera-shaped benchmark
Weights are random. Latency depends on a model's shape, not on what it has learned, so a randomly initialised network measures exactly like a trained one of the same architecture. That keeps this script to a few seconds with no download.
import os, time, statistics as st, numpy as np, psutil, torch, torch.nn as nn, onnxruntime as ort
from onnxruntime.quantization import quantize_static, CalibrationDataReader, QuantType, QuantFormat
torch.manual_seed(0)
cnn = nn.Sequential(
nn.Conv2d(3, 32, 3, stride=2, padding=1), nn.ReLU(),
nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
nn.Conv2d(64, 128, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
nn.Conv2d(128, 128, 3, padding=1), nn.ReLU(),
nn.AdaptiveAvgPool2d(1), nn.Flatten(), nn.Linear(128, 10),
).eval()
torch.onnx.export(cnn, torch.zeros(1, 3, 128, 128), "camera.onnx",
input_names=["image"], output_names=["scores"], opset_version=13)
rng = np.random.default_rng(0)
class Calib(CalibrationDataReader):
"""Real inputs belong here. Noise is fine for a speed test and wrong for an accuracy test."""
def __init__(self, n=20):
self.data = iter([{"image": rng.standard_normal((1, 3, 128, 128), dtype=np.float32)} for _ in range(n)])
def get_next(self):
return next(self.data, None)
quantize_static("camera.onnx", "camera_int8.onnx", Calib(), quant_format=QuantFormat.QDQ,
activation_type=QuantType.QUInt8, weight_type=QuantType.QInt8)
print("camera.onnx : %7.1f KB, %d parameters"
% (os.path.getsize("camera.onnx") / 1024, sum(p.numel() for p in cnn.parameters())))
print("camera_int8.onnx : %7.1f KB" % (os.path.getsize("camera_int8.onnx") / 1024))
print()
frame = rng.standard_normal((1, 3, 128, 128), dtype=np.float32)
proc = psutil.Process()
print("model threads ms/frame frames/second extra RAM")
for path in ["camera.onnx", "camera_int8.onnx"]:
for threads in [1, 2, 4]:
before = proc.memory_info().rss
so = ort.SessionOptions()
so.intra_op_num_threads = threads # a Pi has four cores; try all of them
so.inter_op_num_threads = 1
sess = ort.InferenceSession(path, so, providers=["CPUExecutionProvider"])
for _ in range(20):
sess.run(None, {"image": frame})
rss = (proc.memory_info().rss - before) / 1024 / 1024
times = []
for _ in range(100):
t = time.perf_counter(); sess.run(None, {"image": frame}); times.append((time.perf_counter() - t) * 1000)
ms = st.median(times)
print("%-18s %7d %8.2f %13.0f %6.1f MB" % (path, threads, ms, 1000 / ms, rss))
del sesscamera.onnx : 947.6 KB, 242122 parameters camera_int8.onnx : 246.0 KB model threads ms/frame frames/second extra RAM camera.onnx 1 2.32 432 5.2 MB camera.onnx 2 1.21 824 2.6 MB camera.onnx 4 0.65 1547 2.7 MB camera_int8.onnx 1 0.85 1173 1.6 MB camera_int8.onnx 2 0.48 2091 1.5 MB camera_int8.onnx 4 0.34 2980 1.8 MB
That output is from a desktop, not a Pi. The file sizes and parameter count will reproduce exactly for you; the timings will not, and they are the whole point of running it yourself. Expect a Raspberry Pi to be very much slower — this is a desktop CPU with far higher clocks and wider vector units. Run the script on both and write down the ratio, because that ratio is what you will need for every future estimate.
The extra RAM column is noisy. The first row absorbs one-off allocations that later rows reuse, so read it as an order of magnitude, not a measurement.
Two results worth carrying with you
Threads scale, almost linearly. 2.32 ms on one thread, 0.65 ms on four — a 3.6x speed-up from 4 cores. This is the single biggest free win on a Pi. ONNX Runtime defaults to a sensible thread count, but it does not always guess right in a container or under a process manager. Set intra_op_num_threads explicitly.
Here, at last, int8 was genuinely faster. 2.32 ms down to 0.85 ms on one thread — 2.7 times. Contrast that with quantisation in practice, where int8 made a small dense model five times slower.
The difference is the model and the method. This is a convolutional network quantised statically, so the whole graph runs in integers with no per-call conversion. That is exactly the case ARM CPUs are optimised for, through their integer dot-product instructions. On a Pi the gap is usually larger than on a desktop, because desktop float units are relatively stronger.
Note also that dynamic quantisation would have failed here. quantize_dynamic on a convolutional model produces ConvInteger nodes that ONNX Runtime's CPU backend does not implement:
NOT_IMPLEMENTED : Could not find an implementation for ConvInteger(10) node with name '/0/Conv_quant'
For convolutional models, use quantize_static with a calibration reader. For dense and transformer models, quantize_dynamic is fine.
Which Pi
Published specifications, which decide what is even possible:
| Board | CPU | RAM |
|---|---|---|
| Pi Zero 2 W | 4x Cortex-A53, 1 GHz | 512 MB |
| Pi 3 Model B+ | 4x Cortex-A53, 1.4 GHz | 1 GB |
| Pi 4 Model B | 4x Cortex-A72, 1.5–1.8 GHz | 1–8 GB |
| Pi 5 | 4x Cortex-A76, 2.4 GHz | 4–16 GB |
The 512 MB on a Pi Zero 2 W is the number that stops projects. Runtime, model, camera buffers and the operating system all come out of it. Check extra RAM in your own run and add the model file size and your image buffers before you order anything.
Adding real acceleration
Two options that change the arithmetic completely, both int8-only:
- Raspberry Pi AI HAT+ — a Hailo-8L or Hailo-8 accelerator over PCIe on a Pi 5. Rated at 13 and 26 TOPS respectively. Uses Hailo's own toolchain and its own compiled model format, not ONNX directly.
- Coral USB Accelerator — an Edge TPU on USB. Requires a fully-integer TFLite model compiled with
edgetpu_compiler. Anything the compiler cannot place falls back to the CPU, silently.
Both mean a second conversion step and a second set of operator restrictions. Get the model working on the CPU first, with a measured baseline. Then decide whether you actually need the accelerator, because frequently you do not.
Common mistakes
Under-powering it. A Pi 5 wants a 5.1 V, 5 A supply; a Pi 4 wants 5.1 V, 3 A. Phone chargers cause silent throttling. Check with:
vcgencmd get_throttled0x0 means fine. Anything else is a power or heat problem, and it is the correct first suspect whenever a Pi is inexplicably slow.
Running the 32-bit OS. Fewer wheels, no NEON improvements from the 64-bit ABI, and a hard 4 GB memory ceiling. Use 64-bit.
Writing logs to the SD card in a loop. SD cards have limited write cycles. Log to /dev/shm or a USB SSD, and boot from an SSD if the Pi will run for years.
Benchmarking without a heatsink. A bare Pi 5 under sustained load will throttle within minutes. See power and thermal limits for how to measure that rather than guess at it.
Assuming your laptop's numbers transfer. They do not, and the ratio is not a constant across models either.
Try it yourself
Run the script on your laptop and on a Pi if you can borrow one, and compute frames per second for each. Then pick a real task — counting vehicles, checking a machine, reading a meter — and write down the frame rate it truly needs. Most people discover they needed 2 frames per second and were about to buy an accelerator for 30.
What to learn next
- Latency and throughput — measuring speed so the numbers mean something.
- Power and thermal limits — why the tenth minute is slower than the first.
- Docker for ML — packaging what you built so it survives a reboot.
Researcher — Mathematics and papers.
What the hardware actually offers
The relevant micro-architectural facts, since they determine the achievable numbers rather than only correlating with them:
| Board | Core | ISA | SIMD | Dot product |
|---|---|---|---|---|
| Pi Zero 2 W, Pi 3 | Cortex-A53 | ARMv8-A | NEON, 128-bit | No |
| Pi 4 | Cortex-A72 | ARMv8-A | NEON, 128-bit | No |
| Pi 5 | Cortex-A76 | ARMv8.2-A | NEON, 128-bit | Yes (SDOT/UDOT) |
The last column matters more than the clock speed. ARMv8.2's SDOT instruction computes a four-element int8 dot product with int32 accumulation in one operation. On a Cortex-A76 an int8 convolution can therefore reach roughly four times the elements per cycle of the float32 path, before any memory effects. On a Cortex-A72, int8 must be widened to int16 first, and the advantage largely evaporates.
This is why "quantise for the edge" is good advice on a Pi 5 and much weaker advice on a Pi 3. Hardware support, not the bit width, produces the speed-up.
The VideoCore GPU is not a general compute target. There is no usable OpenCL, and the Vulkan compute path is not a productive route for inference. Treat the Pi as a CPU-only device unless an accelerator is attached.
Threading and memory bandwidth
Near-linear thread scaling — the 3.6x on four threads above — holds only while the working set fits in cache. Once activations exceed the last-level cache, additional threads contend for a single memory controller and scaling flattens.
Pi 4 has LPDDR4 at 3200 MT/s on a 32-bit bus, giving a theoretical peak around 12.8 GB/s and considerably less in practice. Pi 5 uses LPDDR4X-4267, roughly 17 GB/s theoretical. For a convolutional network at batch 1 with $A$ bytes of activations per layer, the memory-bound floor is approximately
$$ t_{\text{layer}} \ge \frac{A_{\text{in}} + A_{\text{out}} + W}{B} $$
with $W$ the layer's weight bytes and $B$ the achieved bandwidth. Increasing input resolution raises $A$ quadratically and therefore hits this floor rapidly. Reducing input resolution is usually a bigger lever on a Pi than reducing model depth, and it is the first thing to try.
Thermal behaviour
The Pi's firmware implements a documented ladder. A Pi 5 begins soft-limiting at 80 °C and hard-limits at 85 °C; a Pi 4 uses 80 and 85 °C similarly, and additionally drops the ARM clock under sustained load without active cooling. vcgencmd get_throttled returns a bitfield where bits 0–3 report current under-voltage, frequency capping, throttling and soft temperature limit, and bits 16–19 report whether each has occurred since boot.
The sticky bits are the useful ones. A benchmark that reports good numbers while bits 16–19 are set is reporting a run that was interrupted by the firmware, and the number is not reproducible.
Under-voltage is the more common cause in the field. USB-C power delivery negotiation failures with non-compliant supplies produce exactly the same symptom as thermal throttling and are not visible without checking the bitfield.
Accelerator economics
| Option | Peak int8 | Model format | Practical constraint |
|---|---|---|---|
| Pi 5 CPU, 4 threads | — | ONNX, TFLite | Bandwidth-bound above modest resolutions |
| Coral USB (Edge TPU) | 4 TOPS | Compiled TFLite, full integer | 8 MB on-chip model memory; overflow streams from host and collapses throughput |
| AI HAT+ (Hailo-8L) | 13 TOPS | Hailo HEF, own compiler | PCIe, Pi 5 only |
| AI HAT+ (Hailo-8) | 26 TOPS | Hailo HEF, own compiler | PCIe, Pi 5 only |
TOPS figures are peak, not achieved. The Edge TPU's 8 MB parameter cache is the constraint that decides real performance: a model that fits runs at near peak, and one that does not streams parameters over USB per inference and can end up slower than the CPU. This is the single most common disappointment with Coral hardware, and it is a capacity threshold rather than a gradient.
Every accelerator here is integer-only. Quantisation is a prerequisite for using any of them, not an optimisation applied afterwards.
Reading
- Raspberry Pi hardware documentation — raspberrypi.com/documentation/computers — for the throttling bitfield and power requirements.
- ARM, Cortex-A76 Software Optimization Guide — instruction throughputs behind the
SDOTclaim. - Warden and Situnayake, TinyML, O'Reilly 2019 — for the class of devices one step below a Pi.
- Reddi et al., MLPerf Inference Benchmark, ISCA 2020 — arxiv.org/abs/1911.02549 — on measuring these systems in a way that is comparable at all.
What to learn next
- Latency and throughput — measuring speed so the numbers mean something.
- Power and thermal limits — why the tenth minute is slower than the first.
- Docker for ML — packaging what you built so it survives a reboot.