Why run a model on the device?
On-device inference wins on latency, privacy, cost and working offline, and loses on model size, update speed and available memory.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
You run a model on the device when waiting, leaking, paying or losing signal would ruin it.
The analogy you have already lived
You are in a lift. The doors close, the bars vanish, and your music keeps playing because you downloaded it.
The song that was streaming stops. The song already on your phone does not care.
Now imagine your phone had to ask a server before it recognised your face. In that lift, you would be locked out of your own phone.
Downloading the song is what edge AI does with a model.
The four reasons
One: waiting. A round trip to a server takes time, even on a good connection. On a train that becomes half a second. Half a second is fine for a chatbot and unacceptable for a camera drawing a box around a moving face.
Two: privacy. Sending a photo to a server means the photo leaves your phone. A hospital, a bank, or a parent photographing a child has a real reason to care. If the model runs on the device, the picture never travels.
Three: cost. Somebody pays for every server call. If your app is free, that somebody is you. On-device inference costs nothing per use, forever.
Four: signal. In a basement. On a farm, in a tunnel, on a factory floor. In a village with one bar. The cloud is not there. A model on the device works in all of those places.
The honest other side
On-device is not free. You trade four things away.
Size. The model must be small enough to download and to hold in memory. Users do not install a 900 MB app.
Quality. A model small enough for a phone is usually weaker than the big one on the server.
Updates. Fixing a cloud model takes one deployment. Fixing an on-device model means shipping an app update and waiting for people to install it.
Battery. The work still costs energy. It has moved from a server's power supply to the battery in your pocket.
How to decide
Does it need an answer in under about 100 ms? -> device
Is the input private? -> device
Will it run thousands of times a day, for free? -> device
Might the network be missing? -> device
Does it need the largest, strongest model? -> cloud
Does it need data only the server has? -> cloud
Must it change weekly without an app update? -> cloudMany real products do both. The phone handles the common case instantly, and asks the server only when it is unsure. That is called a cascade — a cheap model first, an expensive one only when needed.
Where you have already seen the split
- Your keyboard predicts words on the phone, and syncs your dictionary to the cloud.
- Your camera finds faces on the phone, and Google Photos searches "beach" in the cloud.
- A voice assistant spots the wake word on the phone, then sends the rest of the sentence away.
Remember this
- Go on-device for speed, privacy, zero running cost and no signal.
- Go to the cloud for the biggest model, fresh data and fast updates.
- Most good products do both, with the device answering first.
What to learn next
- Model compression — how a model gets small enough to ship.
- Latency and throughput — measuring the speed claims in this lesson properly.
- Model serving — what you are choosing not to do when you go on-device.
Developer — Code and libraries.
Setup
pip install scikit-learn joblibThe point of this script is to replace hand-waving about "faster" with arithmetic you can check. It measures one thing honestly and labels every assumption.
Putting numbers on the four reasons
import os, time, joblib
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)
clf = LogisticRegression(max_iter=2000).fit(Xtr, ytr)
joblib.dump(clf, "m.joblib")
model_kb = os.path.getsize("m.joblib") / 1024
one = Xte[:1]
for _ in range(100):
clf.predict(one)
t0 = time.perf_counter()
for _ in range(1000):
clf.predict(one)
tiny_ms = (time.perf_counter() - t0) / 1000 * 1000 # the only measured number here
REAL_MS = 25.0 # a real phone vision model. Typical figure, NOT measured by this script.
SERVER_MS = 15 # how long the server itself takes once the request arrives. Assumed.
print("measured here : %.3f ms for this %.1f KB model" % (tiny_ms, model_kb))
print("assumed : %.1f ms for a real phone vision model (typical, not measured here)" % REAL_MS)
print()
# Round-trip times people report on mobile networks. Ranges, not measurements.
NETWORKS = [("home wifi, server nearby", 30), ("4G in a city", 60),
("4G on a moving train", 180), ("patchy 3G", 500), ("no signal at all", None)]
print("network cloud round trip vs tiny model vs real model")
for name, rtt in NETWORKS:
if rtt is None:
print("%-26s never works works" % name)
continue
cloud = rtt + SERVER_MS
print("%-26s %10d ms %8.0fx faster %8.1fx faster" % (name, cloud, cloud / tiny_ms, cloud / REAL_MS))
print()
IMAGE_KB, PER_DAY = 200, 2000 # one modest phone JPEG, checked twice a second all day
mb_day = IMAGE_KB * PER_DAY / 1024
print("2000 photos a day at 200 KB each")
print(" cloud : %.0f MB uploaded per day, %.1f GB per month" % (mb_day, mb_day * 30 / 1024))
print(" device: %.1f KB downloaded once, then 0 MB per month" % model_kb)measured here : 0.030 ms for this 5.9 KB model assumed : 25.0 ms for a real phone vision model (typical, not measured here) network cloud round trip vs tiny model vs real model home wifi, server nearby 45 ms 1483x faster 1.8x faster 4G in a city 75 ms 2472x faster 3.0x faster 4G on a moving train 195 ms 6426x faster 7.8x faster patchy 3G 515 ms 16972x faster 20.6x faster no signal at all never works works 2000 photos a day at 200 KB each cloud : 391 MB uploaded per day, 11.4 GB per month device: 5.9 KB downloaded once, then 0 MB per month
Read the two speed columns against each other
The "vs tiny model" column says 2472x. Quote that number at anybody and you are lying to them. It compares a 5.9 KB logistic regression against a network round trip. Nobody ships that model in a camera app.
The "vs real model" column is the honest one. A real on-device vision model takes roughly 25 ms on a mid-range phone. Against city 4G the win is 3x, not 2472x. Against patchy 3G it is about 20x.
3x is still worth having. It is the difference between a viewfinder that tracks smoothly and one that stutters. But it is a normal engineering win, not a miracle.
Your own tiny_ms will differ from 0.030 and will wobble between runs, and the whole "vs tiny model" column moves with it. That is normal. It is a sub-microsecond measurement on a shared machine, so treat those multipliers as "very large" and nothing more precise.
The data column is where the argument is won
11.4 GB a month, uploaded, on a mobile data pack. That is the whole cloud argument collapsing for anything high-frequency.
Note what the comparison is: 391 MB every day versus 5.9 KB once. The on-device cost is a one-time download charged at install. The cloud cost recurs for as long as the app lives.
The cascade pattern, in code
Most production systems are not one or the other. They run the small model first and escalate only when it is unsure.
import numpy as np
def decide(probs, threshold=0.90):
"""Answer on the device when confident, otherwise ask the server."""
top = float(np.max(probs))
return ("device", int(np.argmax(probs))) if top >= threshold else ("cloud", None)
for probs in ([0.01, 0.02, 0.95, 0.02], [0.30, 0.28, 0.22, 0.20]):
where, answer = decide(np.array(probs))
print(f"top probability {max(probs):.2f} -> handled by {where}, answer {answer}")top probability 0.95 -> handled by device, answer 2 top probability 0.30 -> handled by cloud, answer None
If 90% of inputs are easy, you have removed 90% of your server bill and 90% of your uploads. The threshold is the dial: raise it and quality goes up along with cost.
Common mistakes
Comparing a toy on-device model with a real cloud model. That is the 2472x error above. Compare like with like, or say plainly that you are not.
Ignoring the cold start. Loading a model into memory can take longer than a hundred inferences. Measure app-open-to-first-answer, not steady-state latency.
Treating "on-device" as automatically private. It is private only if you also stop sending logs, analytics and crash reports containing the input. Privacy is a property of the whole system, not of where the matrix multiply happens.
Forgetting model updates. Cloud models improve silently. Device models improve when users update the app, which for some audiences means never. Version your model file and handle old versions.
Try it yourself
Change REAL_MS to 120, which is what a heavier model on an older phone costs. Recompute the table. Find the network speed at which the cloud actually becomes the faster option, and notice that it exists.
What to learn next
- Model compression — how a model gets small enough to ship.
- Latency and throughput — measuring the speed claims in this lesson properly.
- Model serving — what you are choosing not to do when you go on-device.
Researcher — Mathematics and papers.
Framing the decision properly
End-to-end latency for the cloud path is a sum of terms that are usually collapsed into one number and should not be:
$$ L_{\text{cloud}} = t_{\text{capture}} + t_{\text{encode}} + \frac{S}{B_{\uparrow}} + \text{RTT} + t_{\text{queue}} + t_{\text{infer}}^{\text{server}} + \frac{S'}{B_{\downarrow}} $$
- $S$ — request payload size in bytes, $S'$ the response size.
- $B_{\uparrow}, B_{\downarrow}$ — available uplink and downlink bandwidth in bytes per second.
- $\text{RTT}$ — network round-trip time, plus TLS and connection setup on a cold connection.
- $t_{\text{queue}}$ — server-side queueing delay, which grows without bound as utilisation approaches one.
- $t_{\text{encode}}$ — JPEG or Opus encoding, frequently forgotten and frequently 5–15 ms.
The device path is short by comparison:
$$ L_{\text{device}} = t_{\text{capture}} + t_{\text{preprocess}} + t_{\text{infer}}^{\text{device}} $$
Two consequences follow. First, $L_{\text{cloud}}$ has a hard floor set by the speed of light and the mobile radio, independent of how fast your servers are: LTE control-plane wake-up alone costs tens of milliseconds after an idle period. Second, $L_{\text{cloud}}$ has a heavy tail, because $t_{\text{queue}}$ and RTT both have long-tailed distributions. Reporting only the median hides the failure mode users actually notice.
Tail latency is the real argument
For a request touching $n$ independent services, the probability that none exceeds its own p99 is $0.99^n$. At $n = 10$ roughly one request in ten exceeds a p99 somewhere. Dean and Barroso (2013), The Tail at Scale, is the canonical treatment.
On-device inference has no fan-out, no queueing and no radio. Its latency distribution is tight, bounded by thermal state rather than by network variance. For interactive perception — autofocus, gesture tracking, live captions — the p99 is the specification and the median is decoration.
Energy, honestly
On-device is not automatically the lower-energy option. The comparison is between local compute energy and radio energy:
$$ E_{\text{cloud}} \approx P_{\text{radio}} \cdot \left(t_{\text{tx}} + t_{\text{tail}}\right), \qquad E_{\text{device}} \approx P_{\text{compute}} \cdot t_{\text{infer}} $$
The term people omit is $t_{\text{tail}}$: the cellular radio stays in a high-power RRC-connected state for several seconds after a transmission before demoting. A single small request can therefore cost as much energy as several seconds of radio-on time. Balasubramanian et al. (2009), Energy consumption in mobile phones, measured this tail directly and found it dominating the energy cost of small, infrequent transfers.
The practical rule that falls out: many small requests are the worst pattern. If you must use the network, batch and coalesce so the radio wakes once.
Privacy is a formal property, not a location
"The data stayed on the device" is a claim about information flow, and moving the matrix multiply is neither necessary nor sufficient for it.
- Model outputs leak. Membership inference (Shokri et al., 2017) recovers whether a record was in the training set from confidence scores alone.
- Gradients leak. Federated learning keeps raw data local yet transmits updates from which inputs can be reconstructed (Zhu et al., Deep Leakage from Gradients, 2019). Differential privacy is the mitigation with a proof attached; on-device execution alone is not.
- Telemetry leaks. Crash dumps and analytics routinely carry the very input the on-device path was meant to protect.
On-device inference is a strong default and a weak guarantee. If you need a guarantee, state the threat model and use a mechanism with a proof attached.
Cascades and the economics of confidence
The two-stage cascade in the developer block has a clean cost model. With device cost $c_d$, cloud cost $c_c$, and escalation rate $\alpha$ (the fraction of inputs the device defers):
$$ \mathbb{E}[\text{cost}] = c_d + \alpha \, c_c, \qquad \mathbb{E}[\text{error}] = (1 - \alpha)\varepsilon_d^{\text{acc}} + \alpha \varepsilon_c $$
where $\varepsilon_d^{\text{acc}}$ is the device model's error on the inputs it accepts and $\varepsilon_c$ the cloud model's error on those it receives. The whole design problem is making $\varepsilon_d^{\text{acc}} \ll \varepsilon_d$ by deferring exactly the hard inputs.
That requires calibrated confidence, and neural network softmax outputs are famously not calibrated (Guo et al., 2017, On Calibration of Modern Neural Networks). Temperature scaling on a held-out set is the cheapest fix and usually enough. Without it, the escalation threshold is a number with no meaning.
Reading
- Dean and Barroso, The Tail at Scale, CACM 2013.
- Balasubramanian, Balasubramanian and Venkataramani, Energy consumption in mobile phones: a measurement study, IMC 2009.
- Guo et al., On Calibration of Modern Neural Networks, ICML 2017 — arxiv.org/abs/1706.04599
- Shokri et al., Membership Inference Attacks Against Machine Learning Models, IEEE S&P 2017 — arxiv.org/abs/1610.05820
- Kang et al., Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge, ASPLOS 2017 — splitting one network across device and server at the cheapest cut point.
What to learn next
- Model compression — how a model gets small enough to ship.
- Latency and throughput — measuring the speed claims in this lesson properly.
- Model serving — what you are choosing not to do when you go on-device.