Running models in the browser
A browser can run a trained model with no install and no server, using WebAssembly or WebGPU — at the cost of a few megabytes of runtime on the first visit.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A web page can run a model itself, with no app to install and no server to call.
The analogy you have already lived
You sit down at a restaurant, scan the QR code on the table, and the menu opens.
You did not install anything. You did not create an account. The page arrived, worked, and will be gone when you close the tab.
Running a model in the browser is that, for AI. Someone opens a link and a model runs on their phone. Nothing is installed and nothing is uploaded.
Why it exists
Getting an app onto someone's phone is hard. They must find it, trust it, have space for it, and update it.
A web page skips all of that. One link works on an iPhone, an Android phone, a school computer and an old laptop. There is no store approval, and no separate build for each platform.
And because the model runs in the browser, the picture or the text never leaves the device. A page that checks a photo can genuinely promise it did not upload the photo. Anyone can verify that by watching the network tab.
How it works
Browsers cannot run Python. They run JavaScript. They also run WebAssembly — a way to ship compiled code at close to native speed.
your ONNX model file ──┐
├──► the browser ──► the answer, on the device
a WebAssembly runtime ──┘
(downloaded once, cached)The runtime is a program that knows how to execute a model. The browser downloads it once and keeps it in cache. After the first visit, it is already there.
Newer browsers can also use WebGPU, which lets a page use the graphics chip. That is much faster for big models, and it is not available everywhere yet.
The honest part, in numbers
The runtime is not small. For the common one:
- The JavaScript part is about 0.1 MB compressed.
- The WebAssembly part is about 2.4 MB compressed.
So before your model has downloaded a single byte, a first-time visitor has fetched roughly two and a half megabytes.
On home wifi that is a second. On a slow mobile connection it is a real wait, and it comes out of a data pack. That is the price of "no install".
After the first visit it is cached and costs nothing. So this approach suits pages people return to, and suits one-off pages much less.
Where you have already seen it
- Background removal tools that work while you are offline.
- Browser translation of a page you are reading.
- Web-based photo editors that find faces without uploading anything.
- Live subtitle demos that use your microphone and send nothing.
Remember this
- The browser runs the model with no install and no upload.
- The runtime costs a few megabytes on the first visit, then is cached.
- WebGPU is much faster, and is not on every device yet.
What to learn next
- AI on a Raspberry Pi — the same ONNX file on a board that costs less than a phone.
- Latency and throughput — measuring the timings in this lesson properly.
- ONNX — where the model file in this lesson came from.
Developer — Code and libraries.
Setup
Two steps. Export the model in Python, then write one HTML file. No npm, no bundler, no build step.
pip install torch onnx scikit-learnStep 1: export the model
This is the same export as the ONNX lesson. Run it, and keep digits.onnx.
import os, torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
torch.manual_seed(0); torch.set_num_threads(1)
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X / 16.0, y, test_size=0.3, random_state=0, stratify=y)
Xtr = torch.tensor(Xtr, dtype=torch.float32); ytr = torch.tensor(ytr, dtype=torch.long)
model = nn.Sequential(nn.Linear(64, 256), nn.ReLU(), nn.Linear(256, 256), nn.ReLU(), nn.Linear(256, 10))
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for _ in range(30):
for i in torch.randperm(len(Xtr)).split(64):
opt.zero_grad(); F.cross_entropy(model(Xtr[i]), ytr[i]).backward(); opt.step()
model.eval()
torch.onnx.export(model, torch.zeros(1, 64), "digits.onnx",
input_names=["pixels"], output_names=["scores"],
dynamic_axes={"pixels": {0: "batch"}, "scores": {0: "batch"}},
opset_version=13)
print("digits.onnx %.1f KB" % (os.path.getsize("digits.onnx") / 1024))
print("first test image, true label:", int(yte[0]))digits.onnx 332.7 KB first test image, true label: 1
Step 2: one HTML file
Put this next to digits.onnx. The 64 numbers are that first test image, already scaled — laid out in eight rows of eight so you can see the digit in the code.
<!doctype html>
<html>
<head><meta charset="utf-8"><title>Digits in the browser</title></head>
<body>
<pre id="out">loading…</pre>
<script src="https://cdn.jsdelivr.net/npm/onnxruntime-web@1.20.1/dist/ort.min.js"></script>
<script>
const log = t => document.getElementById("out").textContent += t + "\n";
document.getElementById("out").textContent = "";
// One test digit, 64 pixels, already scaled to the range 0 to 1.
const pixels = new Float32Array([
0, 0, 0, 0, 0.6875, 0.875, 0.1875, 0,
0, 0, 0, 0.125, 1, 1, 0.125, 0,
0, 0, 0, 0.6875, 1, 0.875, 0, 0,
0, 0, 0.1875, 1, 1, 0.9375, 0, 0,
0, 0.0625, 0.8125, 1, 1, 0.8125, 0, 0,
0, 0.375, 1, 0.5625, 0.9375, 0.8125, 0, 0,
0, 0, 0, 0, 0.75, 1, 0.0625, 0,
0, 0, 0, 0, 0.5625, 0.875, 0.0625, 0,
]);
async function main() {
ort.env.wasm.numThreads = 1; // multi-threaded wasm needs special headers; see below
const t0 = performance.now();
const session = await ort.InferenceSession.create("digits.onnx");
log("model loaded in " + (performance.now() - t0).toFixed(0) + " ms");
log("inputs : " + session.inputNames.join(", "));
log("outputs: " + session.outputNames.join(", "));
const tensor = new ort.Tensor("float32", pixels, [1, 64]);
const results = await session.run({ pixels: tensor });
for (let i = 0; i < 20; i++) await session.run({ pixels: tensor }); // warm up first
const t1 = performance.now();
for (let i = 0; i < 100; i++) await session.run({ pixels: tensor });
const ms = (performance.now() - t1) / 100;
const scores = Array.from(results.scores.data);
const max = Math.max(...scores); // subtract the max before exp, or it overflows
const exp = scores.map(s => Math.exp(s - max));
const total = exp.reduce((a, b) => a + b, 0);
const probs = exp.map(e => e / total);
const best = probs.indexOf(Math.max(...probs));
log("predicted digit : " + best);
log("confidence : " + (probs[best] * 100).toFixed(1) + "%");
log("time per run : " + ms.toFixed(3) + " ms");
}
main().catch(e => log("ERROR " + e));
</script>
</body>
</html>Step 3: serve it
Opening the file directly with file:// fails — the browser refuses to fetch digits.onnx from a page with no origin. Any static server works:
python -m http.server 8765Then open http://localhost:8765.
model loaded in 56 ms inputs : pixels outputs: scores predicted digit : 1 confidence : 99.2% time per run : 0.018 ms
The digit and the confidence are exact — the same model in Python gives digit 1 at 99.19%. The two timing lines are measurements on one machine and will differ on yours. On the very first visit, before anything is cached, model loaded in was 462 ms rather than 56 ms.
The number worth staring at
0.018 ms in the browser, against 0.012 ms for native ONNX Runtime on the same machine and the same model. WebAssembly is within about one and a half times of native here.
That is the headline fact about browser inference in 2026. It is not a toy. A model that runs at a useful speed on a laptop CPU will run at a usefully similar speed in a browser tab.
What it actually costs to load
Measured against the CDN, for onnxruntime-web@1.20.1:
| File | Raw | Over the wire, compressed |
|---|---|---|
ort.min.js | 0.43 MB | 0.10 MB |
ort-wasm-simd-threaded.wasm | 10.73 MB | 2.38 MB |
digits.onnx | 0.33 MB | 0.33 MB |
Roughly 2.8 MB on a first visit, of which your model is about a tenth. Which WebAssembly file the runtime picks depends on the build and on what the browser supports, so check your own network tab rather than trusting this row.
Two things follow. Serve the runtime from your own domain rather than a CDN if you care about the first-visit time, and set long cache headers. And quantise the model — a 4x smaller model matters much more once you get past toy sizes.
Threads and WebGPU
ort.env.wasm.numThreads = 1 is in the script for a reason. Multi-threaded WebAssembly needs SharedArrayBuffer, which browsers only allow on cross-origin-isolated pages — meaning the server must send Cross-Origin-Opener-Policy: same-origin and Cross-Origin-Embedder-Policy: require-corp. python -m http.server sends neither, so threads would silently fall back to one anyway. Setting it explicitly makes the behaviour visible rather than accidental.
For larger models, WebGPU is the bigger lever:
const session = await ort.InferenceSession.create("digits.onnx", {
executionProviders: ["webgpu", "wasm"], // try the GPU, fall back to WebAssembly
});
console.log(ort.env.webgpu.adapter ? "WebGPU is available" : "using WebAssembly");No output block here, deliberately. Whether WebGPU is used depends on the browser, the operating system, the graphics driver and whether the user has it enabled. Reporting a number from one machine would tell you nothing about yours.
For a model this small, WebGPU would be slower anyway — moving 64 numbers to the GPU costs more than the arithmetic saves. WebGPU starts paying at convolutional and transformer scale.
Common mistakes
Opening the HTML file directly. file:// has no origin, fetch is blocked, and the error message is about CORS rather than about the real cause.
Blocking the main thread. session.run on a large model freezes the page. Move inference to a Web Worker for anything that takes more than a few milliseconds.
Serving .wasm with the wrong MIME type. It must be application/wasm, or streaming compilation fails and the browser falls back to a slower path. Some static hosts get this wrong.
Forgetting the model download is not cached like code. Set explicit Cache-Control headers on the .onnx file, or every visit re-downloads it.
Assuming every device can afford it. A model plus runtime that a laptop swallows can crash a cheap phone with 2 GB of RAM. Test on the worst device you intend to support, not on your own.
Try it yourself
Open the browser's network tab, tick "Disable cache", and reload. Note the total transferred and the time. Then quantise the model with onnxruntime.quantization.quantize_dynamic from the ONNX lesson, point the page at digits_int8.onnx, and reload again. You will have measured the whole trade-off yourself: what the user pays, and what they get.
What to learn next
- AI on a Raspberry Pi — the same ONNX file on a board that costs less than a phone.
- Latency and throughput — measuring the timings in this lesson properly.
- ONNX — where the model file in this lesson came from.
Researcher — Mathematics and papers.
The execution backends
ONNX Runtime Web exposes several execution providers, and their characteristics differ sharply.
| Provider | Mechanism | Precision | Availability |
|---|---|---|---|
wasm | WebAssembly, SIMD, optional threads | fp32 | Everywhere |
webgpu | WebGPU compute shaders | fp32, fp16 where supported | Chrome, Edge, Safari 26; Firefox partial |
webnn | WebNN API to native drivers | Device-dependent | Behind flags in most browsers |
WebAssembly with 128-bit SIMD reaches roughly 50–80% of native throughput for dense linear algebra. The gap comes from a narrower SIMD width than AVX-512, bounds checking, and no access to specialised matrix instructions. Threads require SharedArrayBuffer, gated behind cross-origin isolation since the Spectre mitigations of 2018.
WebGPU replaced the WebGL-shader-abuse approach with real compute shaders, storage buffers and workgroup memory. Reported speed-ups over WebAssembly for transformer inference are large — often one to two orders of magnitude for models above a few hundred million parameters — because the arithmetic is finally on hardware built for it. Two constraints bind in practice: buffer size limits, commonly 128 MB per buffer and around 2 GB total, and fp16 support being an optional feature the adapter may not expose.
WebNN is the W3C attempt to reach the platform's own accelerator — the Neural Engine, an NPU, a DSP — rather than the GPU. It is the only route to the low-power hardware from a web page, and as of 2026 it is not broadly shipping.
The cold-start problem is the real research problem
Steady-state throughput is not what limits browser inference. The first visit is.
Total time to the first answer decomposes as
$$ T_{\text{first}} = \frac{S_{\text{rt}} + S_{\text{model}}}{B} + T_{\text{compile}} + T_{\text{init}} + T_{\text{infer}} $$
- $S_{\text{rt}}$ — runtime bytes, around 2.5 MB compressed for a full WebAssembly build.
- $S_{\text{model}}$ — model bytes.
- $B$ — the visitor's bandwidth, which for a large share of the world is a small number.
- $T_{\text{compile}}$ — WebAssembly compilation, mitigated by streaming compilation, which requires the correct MIME type.
- $T_{\text{init}}$ — session creation: graph optimisation, memory arena allocation, and for WebGPU, shader compilation, which can dominate on first run.
The mitigations are unglamorous and effective: serve from your own origin with long-lived cache headers, store the model in the Origin Private File System or IndexedDB rather than refetching, use a build with only the operators the model needs, and quantise.
Quantisation on the web
Weight-only int8 is the standard choice, and the reasoning is different from native. On WebAssembly there is often no integer speed-up to be had, so the benefit is entirely in bytes transferred and cache footprint — which is exactly the binding constraint. Transformers.js ships int8 and 4-bit ONNX variants for this reason, not for speed.
For WebGPU the calculation changes: fp16 halves both memory traffic and buffer pressure, and is supported when the adapter exposes the shader-f16 feature.
Privacy is verifiable here, which is unusual
Browser inference is the one deployment where a user can check the privacy claim rather than trust it. The network tab shows every request. A Content Security Policy with connect-src 'self' makes exfiltration structurally difficult, and it is auditable by anyone.
This is a stronger position than a native app, where the user has only the store listing and a privacy label to go on. It is worth stating explicitly when it applies, because it rarely does.
Reading
- ONNX Runtime Web documentation — onnxruntime.ai/docs/tutorials/web
- Transformers.js — huggingface.co/docs/transformers.js — the ONNX Runtime Web wrapper that made browser transformers ordinary.
- Haas et al., Bringing the Web up to Speed with WebAssembly, PLDI 2017 — the original design paper.
- Jangda et al., Not So Fast: Analyzing the Performance of WebAssembly vs. Native Code, USENIX ATC 2019 — a careful, sceptical measurement of the native gap.
- The WebGPU and WebNN specifications at w3.org.
What to learn next
- AI on a Raspberry Pi — the same ONNX file on a board that costs less than a phone.
- Latency and throughput — measuring the timings in this lesson properly.
- ONNX — where the model file in this lesson came from.