What is edge AI?
Edge AI means the model runs on the device in your hand instead of on a server, so the answer never travels over the internet.
- 10 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Edge AI means the model runs on your own device, not on somebody's server.
The "edge" is the far end of the network: your phone, your laptop, a camera, or a board taped inside a machine. The opposite is the cloud, a rented computer in a data centre far away.
The analogy you have already lived
Take your phone into a basement car park with no signal. Point it at your own face. It unlocks.
Nothing was sent anywhere. There was no signal to send it over. The whole decision happened inside the phone, faster than you could lift it.
Now think about a shopkeeper who must phone the head office before quoting any price. Fast when the line is free. Useless when the line is down.
Edge AI is the shopkeeper who already knows the prices.
Why it exists
Early machine learning lived in data centres for good reasons. Models were large, and phones were slow.
Three things changed. Phones got dedicated chips for this kind of arithmetic. Techniques appeared for making models much smaller. And people started using AI for things that cannot wait for a network.
A reversing car cannot wait for a server to say "child behind you". A hearing aid cannot upload every sound in the room. A farmer checking a leaf for disease may have no signal at all.
How it works
CLOUD AI
photo -> mobile network -> server -> answer -> back to your phone
(needs signal) (rented) (200 ms, if lucky)
EDGE AI
photo -> the chip already in your hand -> answer
(a few ms, always)A model is a file. In cloud AI that file sits on a server. In edge AI that file sits on your device, and the device does the work.
Both use the same trained model. The change is where the file lives and where the arithmetic happens.
Where you have already used it
- Face unlock on your phone, in a lift with no bars.
- Your keyboard suggesting the next word before you type it.
- "Hey Google" and "Alexa" waking up. The wake word is spotted on the device, and only what follows is sent away.
- Google Translate with a language pack downloaded, reading a signboard through the camera on a plane.
- Portrait mode, deciding which pixels are you and which are the wall behind you.
What is honestly hard here
The device is the whole budget. A data centre gives you as much memory and power as you can pay for. A phone gives you what fits in a pocket and runs off a battery.
So edge AI is a squeezing problem. A model that works beautifully on a server may be too big, too slow, or too hungry for a phone. Most of this section is about that squeeze, and about what you lose when you squeeze.
Nobody has removed that trade-off. Anyone who tells you a large model runs free on a phone is selling something.
Remember this
- Edge AI runs the model on the device, not on a server.
- It wins on speed, privacy, cost and working with no signal.
- It loses on size and power, which is why compression matters so much here.
What to learn next
- Why run a model on the device? — the four reasons, each with its numbers.
- Model compression — the four ways to make a model fit.
- Model deployment — the cloud side of the same problem.
Developer — Code and libraries.
Setup
pip install scikit-learn joblibThat is around 40 MB, and no model is downloaded — the dataset ships inside scikit-learn. Nothing in this section needs a GPU.
A complete edge model, end to end
The point of this script is not accuracy. It is the four numbers it prints: file size, accuracy, prediction time, and bytes sent over the network.
import os, time, joblib
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
X, y = load_digits(return_X_y=True) # 1797 images of 8x8 pixels, bundled with scikit-learn
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)
clf = LogisticRegression(max_iter=2000).fit(Xtr, ytr)
joblib.dump(clf, "digits.joblib") # this file is the thing you would ship to a device
print("training images :", len(Xtr))
print("model file : %.1f KB" % (os.path.getsize("digits.joblib") / 1024))
print("test accuracy : %.4f" % clf.score(Xte, yte))
one = Xte[:1]
for _ in range(100): # warm-up: the first call pays one-off setup costs
clf.predict(one)
t0 = time.perf_counter()
for _ in range(1000):
clf.predict(one)
print("time for one image : %.3f ms" % ((time.perf_counter() - t0) / 1000 * 1000))
print("bytes sent over the internet:", 0)training images : 1257 model file : 5.9 KB test accuracy : 0.9611 time for one image : 0.031 ms bytes sent over the internet: 0
Read those numbers
5.9 KB. That is smaller than the icon of most apps. A working handwriting recogniser fits in less space than this page of text. Not every useful model is enormous.
0.9611. It gets 96 of every 100 test digits right. A deep network reaches about 98 on this dataset. You paid two points of accuracy for a model that fits in a text message.
0.031 ms. Your number will differ, possibly by ten times. It depends on your CPU, whether you are on battery, and what else is running. A phone will be slower than a desktop. Treat every timing in this section as "measure it on your own machine".
0 bytes. No account, no API key, no rate limit, no bill. Turn off your wifi and run it again. Nothing changes.
The three quantities that decide everything
Every remaining lesson in this section moves one of these:
| Quantity | What it costs you | How it is measured |
|---|---|---|
| Model size | Storage, download, RAM | bytes on disk |
| Latency | The user waiting | milliseconds per input |
| Energy | Battery life, heat | joules per input |
A cloud engineer optimises the third for money. An edge engineer optimises it for whether the phone stays cool enough to keep working.
Common mistakes
Assuming the model is the whole download. A 5 MB model shipped with a 40 MB runtime library is a 45 MB app. Count the runtime.
Testing latency on your development machine. A laptop on mains power is the best case, not the average case. Measure on the slowest device you intend to support.
Timing the first call. The first prediction includes memory allocation and thread setup. That is why the loop above warms up before timing.
Forgetting the model still has to reach the device. Edge AI removes the per-request network cost. It adds a one-time download, charged to the user's data pack.
Try it yourself
Replace LogisticRegression with sklearn.ensemble.RandomForestClassifier(n_estimators=200). Run it again. Accuracy rises a little, and the file size rises a lot. Write both numbers down, because that ratio is the central question of this whole section.
What to learn next
- Why run a model on the device? — the four reasons, each with its numbers.
- Model compression — the four ways to make a model fit.
- Model deployment — the cloud side of the same problem.
Researcher — Mathematics and papers.
What "edge" actually means
The term is borrowed from network topology. Compute is classified by how many hops it sits from the data source.
| Tier | Example hardware | Memory | Power envelope |
|---|---|---|---|
| Cloud | A100, H100, TPU pod | 40–80 GB HBM per accelerator | 300–700 W per accelerator |
| Edge server | Jetson AGX Orin, a small box in a shop | 32–64 GB | 15–60 W |
| Device | Phone SoC, laptop NPU | 4–16 GB shared | 1–8 W sustained |
| Microcontroller | Cortex-M, ESP32 | 256 KB – 8 MB SRAM | 1–200 mW |
The vocabulary is loose. "TinyML" usually means the last row, where the model must fit in SRAM measured in kilobytes and there is no operating system worth the name.
The binding constraint changes as you descend the table. In the cloud it is throughput per rupee. On a phone it is thermal design power and memory bandwidth. On a microcontroller it is the size of the weight array in flash.
Why on-device inference became viable
Three curves crossed in roughly 2017–2020.
Dedicated silicon. Apple's Neural Engine (A11, 2017), Qualcomm's Hexagon tensor accelerator and Google's Edge TPU put integer matrix-multiply units into consumer devices. They are cheap in energy per operation compared with the CPU, and they are usually integer-only. That is why quantisation is not optional on this hardware.
Efficient architectures. MobileNet (Howard et al., 2017) replaced standard convolutions with depthwise-separable ones, cutting multiply-accumulates by roughly a factor of $k^2$ for a $k \times k$ kernel. MobileNetV2 (Sandler et al., 2018) added inverted residuals with linear bottlenecks. EfficientNet (Tan and Le, 2019) formalised compound scaling of depth, width and input resolution.
Compression as standard practice. Post-training quantisation, magnitude pruning and distillation moved from papers into default tooling. The next four lessons are those techniques.
The energy argument, quantified
Horowitz (2014), Computing's Energy Problem, gives the numbers that shape all edge hardware design. For a 45 nm process:
| Operation | Approximate energy |
|---|---|
| 32-bit integer add | 0.1 pJ |
| 32-bit float add | 0.9 pJ |
| 32-bit float multiply | 3.7 pJ |
| 32-bit SRAM read (8 KB block) | 5 pJ |
| 32-bit DRAM read | 640 pJ |
The gap that matters is the last row. A DRAM access costs on the order of a thousand times an arithmetic operation. Moving weights, not multiplying them, dominates the energy bill.
That single fact explains most of edge AI. Quantisation helps mainly because it reduces bytes moved. Pruning helps when it lets weights stay resident in on-chip memory. Batching helps in the cloud because it amortises weight movement over many inputs, and helps far less on a device where the batch size is one.
The roofline view
For a layer with $W$ bytes of weights and $F$ floating-point operations, arithmetic intensity is
$$ I = \frac{F}{W} $$
where $I$ is operations per byte of memory traffic. A device with peak compute $P$ (operations per second) and memory bandwidth $B$ (bytes per second) is compute-bound when $I > P/B$, and memory-bound otherwise.
Batch-1 inference through a fully connected layer has $I \approx 0.5$ at float32: two operations per weight, four bytes per weight. Modern SoCs have $P/B$ ratios in the tens to hundreds. Dense layers at batch size one are therefore firmly memory-bound, and extra compute buys nothing. Convolutions reuse each weight across many spatial positions, so their intensity is far higher and they behave differently.
This is why latency and throughput need separate treatment on the edge, and why a model with half the FLOPs is rarely twice as fast in wall-clock time.
Reading
- Horowitz, Computing's Energy Problem (and what we can do about it), ISSCC 2014 — the source of the energy table above.
- Howard et al., MobileNets, 2017 — arxiv.org/abs/1704.04861
- Sandler et al., MobileNetV2, 2018 — arxiv.org/abs/1801.04381
- Williams, Waterman and Patterson, Roofline: an insightful visual performance model for multicore architectures, CACM 2009.
- Warden and Situnayake, TinyML, O'Reilly 2019 — the standard practitioner text for microcontroller-class work.
What to learn next
- Why run a model on the device? — the four reasons, each with its numbers.
- Model compression — the four ways to make a model fit.
- Model deployment — the cloud side of the same problem.