Edge and On-device AI

AI on a Raspberry Pi

A Raspberry Pi runs real models on four ARM cores with no GPU, so the work is choosing a small model, quantising it, and being honest about frames per second.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why people use one for AI
  4. What it is not
  5. What the work actually is
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A Raspberry Pi is a whole computer the size of a playing-card deck. It runs real models slowly, cheaply, and for years without attention.

The analogy you have already lived

Look at the clock on your wall. It does one job. It has done it since somebody hung it there. Nobody restarts it, updates it, or thinks about it.

Your phone could tell the time better. The wall clock wins because it is cheap, it is already there, and it never asks for anything.

A Raspberry Pi is that, for computing. You bolt it to a wall or hide it in a machine, and it does one job forever.

Why people use one for AI

It costs very little. A Pi plus a camera costs less than a mid-range phone. A computer with a graphics card costs far more.

You can leave it somewhere. In a cattle shed, on a factory line, inside a signboard, in a field. A laptop cannot live in those places. A Pi can.

It runs on a small amount of power. A few watts. It can run off a solar panel and a battery.

No cloud bill and no signal needed. Which brings back every reason from why run a model on the device.

What it is not

There is no graphics card in a Raspberry Pi. Nothing like the ones that train models.

A Pi has four small ARM processor cores, the same family as in a mid-range phone. Everything runs on those cores unless you add extra hardware.

So the question is never "can it run this model". It is "how many pictures a second, and is that enough for the job?"

Counting people entering a shop needs perhaps two pictures a second. That is comfortable. Following a cricket ball needs sixty. That is not going to happen.

What the work actually is

  train on a laptop or in a free cloud notebook
            |
  export to ONNX                       <- one file, no framework attached
            |
  quantise it                          <- 8-bit weights, roughly 4x smaller
            |
  copy the file to the Pi
            |
  run it with a small runtime, on all four cores

Nothing is trained on the Pi. Training is a big-computer job. The Pi only answers questions.

The honest part

Everything takes longer than you expect the first time. Installing packages on a Pi is slow. Cameras have their own difficulties. Plan a weekend, not an evening.

Cheap power supplies cause strange faults. A Pi that is not getting enough current slows itself down on purpose to survive. It does not tell you loudly. It gets mysteriously slow, and people blame the model.

The memory card wears out. Writing logs to it constantly will kill it in months. Write logs to memory, or to a USB drive.

A Pi Zero has very little memory. Some models will not load at all. That is a hard stop, not a slow one.

Remember this

  • A Pi has four small cores and no graphics card.
  • The real question is frames per second, not whether it runs.
  • Power supply and memory cause more failures than the model does.

What to learn next

Developer — Code and libraries.

Setup

On the Pi, running 64-bit Raspberry Pi OS:

bash
sudo apt update && sudo apt install -y python3-pip
pip install onnxruntime numpy psutil

onnxruntime has official aarch64 wheels, so this is a download rather than a compile. Use the 64-bit OS — the 32-bit build is slower and has thinner package support.

The benchmark below runs unchanged on your laptop too, which is the point: measure it in both places and compare.

A camera-shaped benchmark

Weights are random. Latency depends on a model's shape, not on what it has learned, so a randomly initialised network measures exactly like a trained one of the same architecture. That keeps this script to a few seconds with no download.

pi_bench.py
import os, time, statistics as st, numpy as np, psutil, torch, torch.nn as nn, onnxruntime as ort
from onnxruntime.quantization import quantize_static, CalibrationDataReader, QuantType, QuantFormat

torch.manual_seed(0)
cnn = nn.Sequential(
    nn.Conv2d(3, 32, 3, stride=2, padding=1), nn.ReLU(),
    nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
    nn.Conv2d(64, 128, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
    nn.Conv2d(128, 128, 3, padding=1), nn.ReLU(),
    nn.AdaptiveAvgPool2d(1), nn.Flatten(), nn.Linear(128, 10),
).eval()
torch.onnx.export(cnn, torch.zeros(1, 3, 128, 128), "camera.onnx",
                  input_names=["image"], output_names=["scores"], opset_version=13)

rng = np.random.default_rng(0)

class Calib(CalibrationDataReader):
    """Real inputs belong here. Noise is fine for a speed test and wrong for an accuracy test."""
    def __init__(self, n=20):
        self.data = iter([{"image": rng.standard_normal((1, 3, 128, 128), dtype=np.float32)} for _ in range(n)])
    def get_next(self):
        return next(self.data, None)

quantize_static("camera.onnx", "camera_int8.onnx", Calib(), quant_format=QuantFormat.QDQ,
                activation_type=QuantType.QUInt8, weight_type=QuantType.QInt8)

print("camera.onnx      : %7.1f KB, %d parameters"
      % (os.path.getsize("camera.onnx") / 1024, sum(p.numel() for p in cnn.parameters())))
print("camera_int8.onnx : %7.1f KB" % (os.path.getsize("camera_int8.onnx") / 1024))
print()

frame = rng.standard_normal((1, 3, 128, 128), dtype=np.float32)
proc = psutil.Process()
print("model              threads   ms/frame   frames/second   extra RAM")
for path in ["camera.onnx", "camera_int8.onnx"]:
    for threads in [1, 2, 4]:
        before = proc.memory_info().rss
        so = ort.SessionOptions()
        so.intra_op_num_threads = threads      # a Pi has four cores; try all of them
        so.inter_op_num_threads = 1
        sess = ort.InferenceSession(path, so, providers=["CPUExecutionProvider"])
        for _ in range(20):
            sess.run(None, {"image": frame})
        rss = (proc.memory_info().rss - before) / 1024 / 1024
        times = []
        for _ in range(100):
            t = time.perf_counter(); sess.run(None, {"image": frame}); times.append((time.perf_counter() - t) * 1000)
        ms = st.median(times)
        print("%-18s %7d   %8.2f   %13.0f   %6.1f MB" % (path, threads, ms, 1000 / ms, rss))
        del sess
Output
camera.onnx      :   947.6 KB, 242122 parameters
camera_int8.onnx :   246.0 KB

model              threads   ms/frame   frames/second   extra RAM
camera.onnx              1       2.32             432      5.2 MB
camera.onnx              2       1.21             824      2.6 MB
camera.onnx              4       0.65            1547      2.7 MB
camera_int8.onnx         1       0.85            1173      1.6 MB
camera_int8.onnx         2       0.48            2091      1.5 MB
camera_int8.onnx         4       0.34            2980      1.8 MB

That output is from a desktop, not a Pi. The file sizes and parameter count will reproduce exactly for you; the timings will not, and they are the whole point of running it yourself. Expect a Raspberry Pi to be very much slower — this is a desktop CPU with far higher clocks and wider vector units. Run the script on both and write down the ratio, because that ratio is what you will need for every future estimate.

The extra RAM column is noisy. The first row absorbs one-off allocations that later rows reuse, so read it as an order of magnitude, not a measurement.

Two results worth carrying with you

Threads scale, almost linearly. 2.32 ms on one thread, 0.65 ms on four — a 3.6x speed-up from 4 cores. This is the single biggest free win on a Pi. ONNX Runtime defaults to a sensible thread count, but it does not always guess right in a container or under a process manager. Set intra_op_num_threads explicitly.

Here, at last, int8 was genuinely faster. 2.32 ms down to 0.85 ms on one thread — 2.7 times. Contrast that with quantisation in practice, where int8 made a small dense model five times slower.

The difference is the model and the method. This is a convolutional network quantised statically, so the whole graph runs in integers with no per-call conversion. That is exactly the case ARM CPUs are optimised for, through their integer dot-product instructions. On a Pi the gap is usually larger than on a desktop, because desktop float units are relatively stronger.

Note also that dynamic quantisation would have failed here. quantize_dynamic on a convolutional model produces ConvInteger nodes that ONNX Runtime's CPU backend does not implement:

Output
NOT_IMPLEMENTED : Could not find an implementation for ConvInteger(10) node with name '/0/Conv_quant'

For convolutional models, use quantize_static with a calibration reader. For dense and transformer models, quantize_dynamic is fine.

Which Pi

Published specifications, which decide what is even possible:

BoardCPURAM
Pi Zero 2 W4x Cortex-A53, 1 GHz512 MB
Pi 3 Model B+4x Cortex-A53, 1.4 GHz1 GB
Pi 4 Model B4x Cortex-A72, 1.5–1.8 GHz1–8 GB
Pi 54x Cortex-A76, 2.4 GHz4–16 GB

The 512 MB on a Pi Zero 2 W is the number that stops projects. Runtime, model, camera buffers and the operating system all come out of it. Check extra RAM in your own run and add the model file size and your image buffers before you order anything.

Adding real acceleration

Two options that change the arithmetic completely, both int8-only:

  • Raspberry Pi AI HAT+ — a Hailo-8L or Hailo-8 accelerator over PCIe on a Pi 5. Rated at 13 and 26 TOPS respectively. Uses Hailo's own toolchain and its own compiled model format, not ONNX directly.
  • Coral USB Accelerator — an Edge TPU on USB. Requires a fully-integer TFLite model compiled with edgetpu_compiler. Anything the compiler cannot place falls back to the CPU, silently.

Both mean a second conversion step and a second set of operator restrictions. Get the model working on the CPU first, with a measured baseline. Then decide whether you actually need the accelerator, because frequently you do not.

Common mistakes

Under-powering it. A Pi 5 wants a 5.1 V, 5 A supply; a Pi 4 wants 5.1 V, 3 A. Phone chargers cause silent throttling. Check with:

bash
vcgencmd get_throttled

0x0 means fine. Anything else is a power or heat problem, and it is the correct first suspect whenever a Pi is inexplicably slow.

Running the 32-bit OS. Fewer wheels, no NEON improvements from the 64-bit ABI, and a hard 4 GB memory ceiling. Use 64-bit.

Writing logs to the SD card in a loop. SD cards have limited write cycles. Log to /dev/shm or a USB SSD, and boot from an SSD if the Pi will run for years.

Benchmarking without a heatsink. A bare Pi 5 under sustained load will throttle within minutes. See power and thermal limits for how to measure that rather than guess at it.

Assuming your laptop's numbers transfer. They do not, and the ratio is not a constant across models either.

Try it yourself

Run the script on your laptop and on a Pi if you can borrow one, and compute frames per second for each. Then pick a real task — counting vehicles, checking a machine, reading a meter — and write down the frame rate it truly needs. Most people discover they needed 2 frames per second and were about to buy an accelerator for 30.

What to learn next

Researcher — Mathematics and papers.

What the hardware actually offers

The relevant micro-architectural facts, since they determine the achievable numbers rather than only correlating with them:

BoardCoreISASIMDDot product
Pi Zero 2 W, Pi 3Cortex-A53ARMv8-ANEON, 128-bitNo
Pi 4Cortex-A72ARMv8-ANEON, 128-bitNo
Pi 5Cortex-A76ARMv8.2-ANEON, 128-bitYes (SDOT/UDOT)

The last column matters more than the clock speed. ARMv8.2's SDOT instruction computes a four-element int8 dot product with int32 accumulation in one operation. On a Cortex-A76 an int8 convolution can therefore reach roughly four times the elements per cycle of the float32 path, before any memory effects. On a Cortex-A72, int8 must be widened to int16 first, and the advantage largely evaporates.

This is why "quantise for the edge" is good advice on a Pi 5 and much weaker advice on a Pi 3. Hardware support, not the bit width, produces the speed-up.

The VideoCore GPU is not a general compute target. There is no usable OpenCL, and the Vulkan compute path is not a productive route for inference. Treat the Pi as a CPU-only device unless an accelerator is attached.

Threading and memory bandwidth

Near-linear thread scaling — the 3.6x on four threads above — holds only while the working set fits in cache. Once activations exceed the last-level cache, additional threads contend for a single memory controller and scaling flattens.

Pi 4 has LPDDR4 at 3200 MT/s on a 32-bit bus, giving a theoretical peak around 12.8 GB/s and considerably less in practice. Pi 5 uses LPDDR4X-4267, roughly 17 GB/s theoretical. For a convolutional network at batch 1 with $A$ bytes of activations per layer, the memory-bound floor is approximately

$$ t_{\text{layer}} \ge \frac{A_{\text{in}} + A_{\text{out}} + W}{B} $$

with $W$ the layer's weight bytes and $B$ the achieved bandwidth. Increasing input resolution raises $A$ quadratically and therefore hits this floor rapidly. Reducing input resolution is usually a bigger lever on a Pi than reducing model depth, and it is the first thing to try.

Thermal behaviour

The Pi's firmware implements a documented ladder. A Pi 5 begins soft-limiting at 80 °C and hard-limits at 85 °C; a Pi 4 uses 80 and 85 °C similarly, and additionally drops the ARM clock under sustained load without active cooling. vcgencmd get_throttled returns a bitfield where bits 0–3 report current under-voltage, frequency capping, throttling and soft temperature limit, and bits 16–19 report whether each has occurred since boot.

The sticky bits are the useful ones. A benchmark that reports good numbers while bits 16–19 are set is reporting a run that was interrupted by the firmware, and the number is not reproducible.

Under-voltage is the more common cause in the field. USB-C power delivery negotiation failures with non-compliant supplies produce exactly the same symptom as thermal throttling and are not visible without checking the bitfield.

Accelerator economics

OptionPeak int8Model formatPractical constraint
Pi 5 CPU, 4 threads—ONNX, TFLiteBandwidth-bound above modest resolutions
Coral USB (Edge TPU)4 TOPSCompiled TFLite, full integer8 MB on-chip model memory; overflow streams from host and collapses throughput
AI HAT+ (Hailo-8L)13 TOPSHailo HEF, own compilerPCIe, Pi 5 only
AI HAT+ (Hailo-8)26 TOPSHailo HEF, own compilerPCIe, Pi 5 only

TOPS figures are peak, not achieved. The Edge TPU's 8 MB parameter cache is the constraint that decides real performance: a model that fits runs at near peak, and one that does not streams parameters over USB per inference and can end up slower than the CPU. This is the single most common disappointment with Coral hardware, and it is a capacity threshold rather than a gradient.

Every accelerator here is integer-only. Quantisation is a prerequisite for using any of them, not an optimisation applied afterwards.

Reading

  • Raspberry Pi hardware documentation — raspberrypi.com/documentation/computers — for the throttling bitfield and power requirements.
  • ARM, Cortex-A76 Software Optimization Guide — instruction throughputs behind the SDOT claim.
  • Warden and Situnayake, TinyML, O'Reilly 2019 — for the class of devices one step below a Pi.
  • Reddi et al., MLPerf Inference Benchmark, ISCA 2020 — arxiv.org/abs/1911.02549 — on measuring these systems in a way that is comparable at all.

What to learn next