Machine Learning Inference in the Browser
Local inference trades server cost and round-trip latency for download size and device compute. That trade is excellent for small models — embeddings, classifiers, keyword spotting, background removal, OCR — and terrible for large ones, and the boundary between the two is a number you can compute before writing any code. This topic covers how the WebAssembly execution backends actually work, what weights cost to move and hold, and how to measure whether the result is fast enough for the interaction you have in mind.
Prerequisites
- [ ] A runtime:
onnxruntime-web1.19+,transformers.js3+, or a hand-compiled inference kernel. - [ ] A model in a portable format — ONNX, GGUF or TensorFlow Lite — with known input and output shapes.
- [ ] Cross-origin isolation headers if you intend to use the multi-threaded backend.
- [ ] A device that resembles your users’. A model that runs in 40 ms on a workstation can take 900 ms on a three-year-old phone.
What the Wasm backend is actually doing
An inference runtime compiled to WebAssembly is a graph executor plus a kernel library. The graph
executor reads the model file, resolves shapes, allocates tensors inside linear memory, and walks the
operator list in topological order. Each operator dispatches to a kernel — a matrix multiply, a
convolution, a softmax — written in C++ and compiled with SIMD enabled.
Nothing about that is magic, and the performance profile follows directly from it. Matrix multiplication
dominates almost every model’s runtime, so the backend’s GEMM implementation is nearly the entire story.
A well-tuned SIMD GEMM in WebAssembly reaches roughly 30–60% of what the same code achieves natively,
because v128 gives you four-wide float lanes where a modern CPU offers eight or sixteen, and because
the engine cannot use the newest instruction set extensions.
Getting a session running
The minimal path with onnxruntime-web is four lines, and the defaults are mostly sensible. The
important configuration is the execution provider list and the thread count.
import * as ort from 'onnxruntime-web';
ort.env.wasm.numThreads = crossOriginIsolated ? navigator.hardwareConcurrency : 1;
ort.env.wasm.simd = true;
ort.env.wasm.wasmPaths = '/vendor/ort/'; // self-host so the headers are yours
const session = await ort.InferenceSession.create('/models/classifier-int8.onnx', {
executionProviders: ['wasm'],
graphOptimizationLevel: 'all',
});
const input = new ort.Tensor('float32', pixels, [1, 3, 224, 224]);
const { logits } = await session.run({ pixel_values: input });
Self-hosting the runtime’s own .wasm files matters for the same reason it matters for any other module:
cross-origin isolation requires cooperating headers, and a CDN that does not send
Cross-Origin-Resource-Policy will break the threaded build with an error that mentions neither.
The weight-loading problem
Model weights are the payload, and they behave differently from code. A 40 MB ONNX file has to be
downloaded, decompressed if it was compressed, parsed, and then copied into linear memory as tensors —
which means peak memory during load is often twice the model size. On a phone with a 1.5 GB tab budget,
a 200 MB model fails during load rather than during inference, and the error rarely says so.
Three techniques make this manageable, and they compose: quantize the weights so the file is smaller, stream the file so peak memory stays near the final size rather than double it, and cache the parsed result so the second visit skips the work entirely. The weight loading guide works through each, including the external-data format that keeps weights out of the graph file.
Preprocessing is where accuracy quietly dies
A model is a pure function of its input tensor, and almost every “the model works in Python but not in the browser” report is a preprocessing mismatch rather than a runtime bug. There are four places to get it wrong, and all four produce output that looks statistically reasonable.
Resizing is the first. Training pipelines usually resize with a specific filter — bilinear with
antialiasing, or bicubic — and the browser’s drawImage uses something else. For a classifier the
difference is often a couple of percent of accuracy; for a model that reads text or fine detail it can be
catastrophic. Resize deliberately, with the same filter the training code used, which is one of the
better reasons to own the
resampler.
Normalisation is the second. Most vision models expect channel-wise mean subtraction and division by a standard deviation, in that order, on values already scaled to 0–1. Applying the scale twice, or subtracting the mean after dividing, gives inputs in the wrong range and a model that confidently predicts the wrong class.
Layout is the third: NCHW versus NHWC, and RGB versus BGR. A channel-order swap produces images that
are recognisable to a human and meaningless to a network trained the other way round. The fourth is
dtype — passing uint8 where float32 is expected either throws or, worse, reinterprets the bytes.
Write the preprocessing once, in the module or in a single well-tested JavaScript function, and validate it by feeding a fixture image through both the browser path and the Python path and comparing the tensors element by element. Five minutes of that saves days.
Choosing a backend
WebAssembly is not the only execution provider available in a browser, and it is frequently not the fastest. The honest comparison looks like this:
| Backend | Strength | Weakness |
|---|---|---|
| Wasm (SIMD, threaded) | Works everywhere, predictable, no driver surprises | 3–10× slower than a GPU for large models |
| WebGPU | Large speedups on big matrix work | Availability gaps, driver bugs, shader compile cost |
| WebNN | Uses platform accelerators directly | Still shipping; limited operator coverage |
| WebGL (legacy) | Broad support | Awkward for modern operators, being retired |
The practical pattern is to try WebGPU and fall back to WebAssembly, because the fallback is always available and always correct. For small models the fallback is often faster anyway: a WebGPU session pays shader compilation and buffer upload costs that a 5 ms Wasm inference never recovers. The backend comparison puts numbers against model sizes.
Threads: the biggest single lever
The threaded WebAssembly backend parallelises GEMM across workers and typically delivers 2.5–3.5× on four cores for transformer and convolutional models. That is a larger improvement than almost any other change you can make, and it is gated entirely on two HTTP headers:
Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp
Without them SharedArrayBuffer is unavailable, the runtime silently falls back to one thread, and your
inference is three times slower for reasons no profiler will point at. Check crossOriginIsolated at
startup and log it — this single boolean explains more mysterious inference slowdowns than everything
else combined. Setting the headers correctly, including what they break in embedded third-party content,
is covered in
configuring COOP/COEP headers.
Thread count should not exceed the physical core count. navigator.hardwareConcurrency reports logical
cores, which on a machine with simultaneous multithreading is double the physical count, and
oversubscribing makes GEMM slower rather than faster because the threads fight for the same cache.
Quantization, and what it costs you
Quantizing weights from 32-bit float to 8-bit integer cuts the download by four, cuts memory by four, and usually makes inference faster because the kernels move a quarter of the bytes. For most classification and embedding models the accuracy cost is under one percent — small enough that the only reason not to quantize is that you have not measured it yet.
The cost lands unevenly, though. Models with wide dynamic range in their activations — some audio models, anything with an unbounded output — can degrade badly with naive per-tensor quantization; per-channel quantization usually recovers most of it. Always evaluate the quantized model against a held-out set rather than eyeballing a few examples, because the failure mode is a systematic bias rather than obvious garbage.
Keeping a session alive across visits
Session creation is expensive: fetch, parse, allocate, optimise the graph, warm up the kernels. On a cold visit that is commonly 300 ms to several seconds. Almost all of it can be avoided on the second visit, and the techniques stack.
Cache the model file itself in the Cache Storage API rather than relying on HTTP caching, so you control eviction and can show accurate progress on first download. Cache the compiled WebAssembly module for the runtime as described in caching compiled Wasm modules in IndexedDB. And keep the session object alive for the whole page session — creating one per inference is a mistake that turns a 20 ms operation into a 2 s one.
let sessionPromise = null;
export function getSession() {
sessionPromise ??= ort.InferenceSession.create(MODEL_URL, OPTIONS)
.then(async (s) => { await warmUp(s); return s; }); // one throwaway inference
return sessionPromise;
}
Assigning the promise rather than the resolved session matters: two concurrent callers before the first resolves would otherwise create two sessions, each with its own copy of the weights. On a 50 MB model that is 100 MB of memory and a mysterious doubling of load time.
Latency budgets and what they permit
Before choosing a model, decide what the interaction can tolerate. The numbers below are the ones users actually feel:
- Under 16 ms — can run per animation frame. Keyword spotting, small pose models, simple filters.
- Under 100 ms — feels instantaneous on a click. Embeddings, small classifiers, OCR on a crop.
- Under 1 s — acceptable with a spinner. Segmentation, background removal, a small language model’s first token.
- Over 1 s — needs a progress indicator and a cancel button, and should probably be on a server.
Run the budget backwards: a 224 × 224 image classifier at int8 on the threaded backend takes 15–40 ms on a laptop and 60–150 ms on a mid-range phone, so it fits the second tier on desktop and the third on mobile. Deciding this early prevents the far more expensive discovery that your architecture only works on the machine you developed it on.
Running inference off the main thread
Even a fast inference is too slow for the main thread once it happens repeatedly. A 40 ms run blocks three animation frames; a 300 ms run makes the page unresponsive in a way users describe as broken. Move the session into a worker and keep the main thread for the interface.
The messaging design matters more than it first appears. Send the smallest representation of the input —
an ImageBitmap handle rather than a pixel array, an ArrayBuffer transferred rather than cloned — and
return only what the interface needs: the top five labels rather than a thousand logits, the mask as a
compact ImageBitmap rather than as float data. A worker that sends back a 4 MB float tensor sixty times
a second undoes everything you gained by moving it there.
With the threaded backend the worker will itself spawn workers, which is allowed and works, but it means your thread budget is the pool size multiplied by the threads per session. One inference worker with four backend threads is almost always better than four inference workers with one thread each, because the former shares weights and cache while the latter duplicates them.
Gotchas and failure modes
no available backend found. ERR: [wasm]— the runtime’s.wasmfiles are not whereort.env.wasm.wasmPathssays. Almost always a bundler moving one file and not the other.- Inference is 3× slower in production than locally. Cross-origin isolation is off, so threads
silently disabled. Log
crossOriginIsolated. RangeError: Array buffer allocation failedduring session creation on mobile. The model plus its parsing copy exceeded the tab budget. Quantize, or stream the weights.- Output shape is
[1, 1, 1000]instead of[1, 1000]. Harmless, and a constant source of confusion — squeeze explicitly rather than indexing by guesswork. - The first inference is ten times slower than the rest. Kernel selection and memory arena warm-up. Run one inference on a dummy input after session creation and throw the result away.
- Results differ slightly from the Python reference. Different GEMM ordering changes floating-point accumulation. Compare with a tolerance, not with equality.
Verifying a model actually works
Correctness verification for inference is comparison against a reference run, and it belongs in CI:
# reference logits from the same model, outside the browser
python -c "import onnxruntime as ort, numpy as np; \
s = ort.InferenceSession('classifier-int8.onnx'); \
x = np.load('fixture.npy'); \
np.save('expected.npy', s.run(None, {'pixel_values': x})[0])"
Then assert in a headless browser test that the in-page result matches within a small tolerance — 1e-3 for float32, looser for quantized models. This catches preprocessing mistakes, which are far more common than model bugs: a channel order swap, a missing normalisation, or a resize that used a different filter from the one the model was trained with will all produce plausible-looking nonsense.
Frequently Asked Questions
Can I run a large language model in the browser? Small ones, yes — models in the 0.5 to 3 billion parameter range run with 4-bit quantization, at a few tokens per second on the WebAssembly backend and considerably faster on WebGPU. The blocker is rarely compute; it is the multi-gigabyte download and the memory ceiling of a tab.
Is WebAssembly inference deterministic across devices? Bit-for-bit, no. Threading changes accumulation order in reductions, and different SIMD paths round differently. Results are stable to within a small tolerance, which is fine for classification and wrong for anything that hashes the output.
How much accuracy does int8 quantization cost? Typically 0.2–1% top-1 for vision classifiers with per-channel quantization, and near zero for embedding models measured by retrieval quality. Audio and generative models are more sensitive. Measure on your own evaluation set — the published numbers are for the reference model, not for yours.
Should I ship the model with my bundle or fetch it separately? Separately, always. Models are large, change on a different schedule from your code, and should be cached independently. Bundling a 40 MB model into a JavaScript artifact defeats every caching layer between you and the user.
What about privacy — is local inference actually safer? It is genuinely different: the input never leaves the device, which removes an entire category of compliance and retention questions for photos, audio and documents. It is not a licence to skip consent, and it does not protect the model itself — anything shipped to a browser can be extracted, so treat model weights as public the moment you serve them.
Do I need a different model for mobile?
Usually a different size of the same model. Ship the quantized small variant by default, and only
upgrade to the larger one after measuring on the device — navigator.hardwareConcurrency and
deviceMemory give a crude but useful signal, and a first-run benchmark gives a better one.
Guides in this topic
- Running ONNX models with onnxruntime-web — a working session from file to logits.
- Choosing between the Wasm and WebGPU backends — measured, per model size.
- Quantizing models for Wasm inference — int8 and int4, and how to check the damage.
- Loading large model weights into linear memory — streaming, external data and peak memory.
- Multi-threaded inference with Wasm threads — the headers, the pool and the scaling curve.
- Measuring inference latency in the browser — warm-up, percentiles and honest numbers.
Related
- SharedArrayBuffer, Atomics & Threading — the threading substrate the backend uses.
- Wasm SIMD & Vectorized Computation — why the kernels are fast at all.
- Module Caching & Startup Performance — making the second visit cheap.
← Back to Production Wasm: Workloads & Deployment