Edge and On-device AI

What is edge AI?

Edge AI means the model runs on the device in your hand instead of on a server, so the answer never travels over the internet.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. Where you have already used it
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Edge AI means the model runs on your own device, not on somebody's server.

The "edge" is the far end of the network: your phone, your laptop, a camera, or a board taped inside a machine. The opposite is the cloud, a rented computer in a data centre far away.

The analogy you have already lived

Take your phone into a basement car park with no signal. Point it at your own face. It unlocks.

Nothing was sent anywhere. There was no signal to send it over. The whole decision happened inside the phone, faster than you could lift it.

Now think about a shopkeeper who must phone the head office before quoting any price. Fast when the line is free. Useless when the line is down.

Edge AI is the shopkeeper who already knows the prices.

Why it exists

Early machine learning lived in data centres for good reasons. Models were large, and phones were slow.

Three things changed. Phones got dedicated chips for this kind of arithmetic. Techniques appeared for making models much smaller. And people started using AI for things that cannot wait for a network.

A reversing car cannot wait for a server to say "child behind you". A hearing aid cannot upload every sound in the room. A farmer checking a leaf for disease may have no signal at all.

How it works

  CLOUD AI
  photo  ->  mobile network  ->  server  ->  answer  ->  back to your phone
             (needs signal)      (rented)              (200 ms, if lucky)

  EDGE AI
  photo  ->  the chip already in your hand  ->  answer
                                                (a few ms, always)

A model is a file. In cloud AI that file sits on a server. In edge AI that file sits on your device, and the device does the work.

Both use the same trained model. The change is where the file lives and where the arithmetic happens.

Where you have already used it

  • Face unlock on your phone, in a lift with no bars.
  • Your keyboard suggesting the next word before you type it.
  • "Hey Google" and "Alexa" waking up. The wake word is spotted on the device, and only what follows is sent away.
  • Google Translate with a language pack downloaded, reading a signboard through the camera on a plane.
  • Portrait mode, deciding which pixels are you and which are the wall behind you.

What is honestly hard here

The device is the whole budget. A data centre gives you as much memory and power as you can pay for. A phone gives you what fits in a pocket and runs off a battery.

So edge AI is a squeezing problem. A model that works beautifully on a server may be too big, too slow, or too hungry for a phone. Most of this section is about that squeeze, and about what you lose when you squeeze.

Nobody has removed that trade-off. Anyone who tells you a large model runs free on a phone is selling something.

Remember this

  • Edge AI runs the model on the device, not on a server.
  • It wins on speed, privacy, cost and working with no signal.
  • It loses on size and power, which is why compression matters so much here.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn joblib

That is around 40 MB, and no model is downloaded — the dataset ships inside scikit-learn. Nothing in this section needs a GPU.

A complete edge model, end to end

The point of this script is not accuracy. It is the four numbers it prints: file size, accuracy, prediction time, and bytes sent over the network.

edge_digits.py
import os, time, joblib
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True)     # 1797 images of 8x8 pixels, bundled with scikit-learn
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)

clf = LogisticRegression(max_iter=2000).fit(Xtr, ytr)
joblib.dump(clf, "digits.joblib")       # this file is the thing you would ship to a device

print("training images     :", len(Xtr))
print("model file          : %.1f KB" % (os.path.getsize("digits.joblib") / 1024))
print("test accuracy       : %.4f" % clf.score(Xte, yte))

one = Xte[:1]
for _ in range(100):                    # warm-up: the first call pays one-off setup costs
    clf.predict(one)
t0 = time.perf_counter()
for _ in range(1000):
    clf.predict(one)
print("time for one image  : %.3f ms" % ((time.perf_counter() - t0) / 1000 * 1000))
print("bytes sent over the internet:", 0)
Output
training images     : 1257
model file          : 5.9 KB
test accuracy       : 0.9611
time for one image  : 0.031 ms
bytes sent over the internet: 0

Read those numbers

5.9 KB. That is smaller than the icon of most apps. A working handwriting recogniser fits in less space than this page of text. Not every useful model is enormous.

0.9611. It gets 96 of every 100 test digits right. A deep network reaches about 98 on this dataset. You paid two points of accuracy for a model that fits in a text message.

0.031 ms. Your number will differ, possibly by ten times. It depends on your CPU, whether you are on battery, and what else is running. A phone will be slower than a desktop. Treat every timing in this section as "measure it on your own machine".

0 bytes. No account, no API key, no rate limit, no bill. Turn off your wifi and run it again. Nothing changes.

The three quantities that decide everything

Every remaining lesson in this section moves one of these:

QuantityWhat it costs youHow it is measured
Model sizeStorage, download, RAMbytes on disk
LatencyThe user waitingmilliseconds per input
EnergyBattery life, heatjoules per input

A cloud engineer optimises the third for money. An edge engineer optimises it for whether the phone stays cool enough to keep working.

Common mistakes

Assuming the model is the whole download. A 5 MB model shipped with a 40 MB runtime library is a 45 MB app. Count the runtime.

Testing latency on your development machine. A laptop on mains power is the best case, not the average case. Measure on the slowest device you intend to support.

Timing the first call. The first prediction includes memory allocation and thread setup. That is why the loop above warms up before timing.

Forgetting the model still has to reach the device. Edge AI removes the per-request network cost. It adds a one-time download, charged to the user's data pack.

Try it yourself

Replace LogisticRegression with sklearn.ensemble.RandomForestClassifier(n_estimators=200). Run it again. Accuracy rises a little, and the file size rises a lot. Write both numbers down, because that ratio is the central question of this whole section.

What to learn next

Researcher — Mathematics and papers.

What "edge" actually means

The term is borrowed from network topology. Compute is classified by how many hops it sits from the data source.

TierExample hardwareMemoryPower envelope
CloudA100, H100, TPU pod40–80 GB HBM per accelerator300–700 W per accelerator
Edge serverJetson AGX Orin, a small box in a shop32–64 GB15–60 W
DevicePhone SoC, laptop NPU4–16 GB shared1–8 W sustained
MicrocontrollerCortex-M, ESP32256 KB – 8 MB SRAM1–200 mW

The vocabulary is loose. "TinyML" usually means the last row, where the model must fit in SRAM measured in kilobytes and there is no operating system worth the name.

The binding constraint changes as you descend the table. In the cloud it is throughput per rupee. On a phone it is thermal design power and memory bandwidth. On a microcontroller it is the size of the weight array in flash.

Why on-device inference became viable

Three curves crossed in roughly 2017–2020.

Dedicated silicon. Apple's Neural Engine (A11, 2017), Qualcomm's Hexagon tensor accelerator and Google's Edge TPU put integer matrix-multiply units into consumer devices. They are cheap in energy per operation compared with the CPU, and they are usually integer-only. That is why quantisation is not optional on this hardware.

Efficient architectures. MobileNet (Howard et al., 2017) replaced standard convolutions with depthwise-separable ones, cutting multiply-accumulates by roughly a factor of $k^2$ for a $k \times k$ kernel. MobileNetV2 (Sandler et al., 2018) added inverted residuals with linear bottlenecks. EfficientNet (Tan and Le, 2019) formalised compound scaling of depth, width and input resolution.

Compression as standard practice. Post-training quantisation, magnitude pruning and distillation moved from papers into default tooling. The next four lessons are those techniques.

The energy argument, quantified

Horowitz (2014), Computing's Energy Problem, gives the numbers that shape all edge hardware design. For a 45 nm process:

OperationApproximate energy
32-bit integer add0.1 pJ
32-bit float add0.9 pJ
32-bit float multiply3.7 pJ
32-bit SRAM read (8 KB block)5 pJ
32-bit DRAM read640 pJ

The gap that matters is the last row. A DRAM access costs on the order of a thousand times an arithmetic operation. Moving weights, not multiplying them, dominates the energy bill.

That single fact explains most of edge AI. Quantisation helps mainly because it reduces bytes moved. Pruning helps when it lets weights stay resident in on-chip memory. Batching helps in the cloud because it amortises weight movement over many inputs, and helps far less on a device where the batch size is one.

The roofline view

For a layer with $W$ bytes of weights and $F$ floating-point operations, arithmetic intensity is

$$ I = \frac{F}{W} $$

where $I$ is operations per byte of memory traffic. A device with peak compute $P$ (operations per second) and memory bandwidth $B$ (bytes per second) is compute-bound when $I > P/B$, and memory-bound otherwise.

Batch-1 inference through a fully connected layer has $I \approx 0.5$ at float32: two operations per weight, four bytes per weight. Modern SoCs have $P/B$ ratios in the tens to hundreds. Dense layers at batch size one are therefore firmly memory-bound, and extra compute buys nothing. Convolutions reuse each weight across many spatial positions, so their intensity is far higher and they behave differently.

This is why latency and throughput need separate treatment on the edge, and why a model with half the FLOPs is rarely twice as fast in wall-clock time.

Reading

  • Horowitz, Computing's Energy Problem (and what we can do about it), ISSCC 2014 — the source of the energy table above.
  • Howard et al., MobileNets, 2017 — arxiv.org/abs/1704.04861
  • Sandler et al., MobileNetV2, 2018 — arxiv.org/abs/1801.04381
  • Williams, Waterman and Patterson, Roofline: an insightful visual performance model for multicore architectures, CACM 2009.
  • Warden and Situnayake, TinyML, O'Reilly 2019 — the standard practitioner text for microcontroller-class work.

What to learn next

What to learn next

These follow on from what you just read.

  • Edge and On-device AI

    Why run a model on the device?

    On-device inference wins on latency, privacy, cost and working offline, and loses on model size, update speed and available memory.

  • Edge and On-device AI

    Model compression

    Model compression makes a trained model smaller using four levers — a smaller design, fewer bits per weight, removing weights, and training a small model to copy a big one.

  • Edge and On-device AI

    Quantisation in practice

    Quantisation stores each weight in fewer bits, giving a 4x smaller model at 8 bits for almost no quality loss, and a cliff you will fall off below 4 bits.