Machine Learning Inference in the Browser

Local inference trades server cost and round-trip latency for download size and device compute. That trade is excellent for small models — embeddings, classifiers, keyword spotting, background removal, OCR — and terrible for large ones, and the boundary between the two is a number you can compute before writing any code. This topic covers how the WebAssembly execution backends actually work, what weights cost to move and hold, and how to measure whether the result is fast enough for the interaction you have in mind.

Prerequisites

  • [ ] A runtime: onnxruntime-web 1.19+, transformers.js 3+, or a hand-compiled inference kernel.
  • [ ] A model in a portable format — ONNX, GGUF or TensorFlow Lite — with known input and output shapes.
  • [ ] Cross-origin isolation headers if you intend to use the multi-threaded backend.
  • [ ] A device that resembles your users’. A model that runs in 40 ms on a workstation can take 900 ms on a three-year-old phone.

What the Wasm backend is actually doing

An inference runtime compiled to WebAssembly is a graph executor plus a kernel library. The graph executor reads the model file, resolves shapes, allocates tensors inside linear memory, and walks the operator list in topological order. Each operator dispatches to a kernel — a matrix multiply, a convolution, a softmax — written in C++ and compiled with SIMD enabled.

Nothing about that is magic, and the performance profile follows directly from it. Matrix multiplication dominates almost every model’s runtime, so the backend’s GEMM implementation is nearly the entire story. A well-tuned SIMD GEMM in WebAssembly reaches roughly 30–60% of what the same code achieves natively, because v128 gives you four-wide float lanes where a modern CPU offers eight or sixteen, and because the engine cannot use the newest instruction set extensions.

What runs when you call session.run The model file is parsed into a graph, tensors are allocated inside linear memory, and each operator dispatches to a compiled kernel. Matrix multiplication kernels account for most of the runtime, and their speed depends on SIMD width. model file — graph, weights, metadata graph executor shape inference · tensor arena in linear memory · operator order GEMM kernels 60–90% of total time conv · attention · norm also SIMD, also memory-bound elementwise ops cheap, but many of them Optimising a model for this backend means reducing matrix work — smaller hidden sizes, fewer layers, quantized weights — not tuning the JavaScript around it.

Getting a session running

The minimal path with onnxruntime-web is four lines, and the defaults are mostly sensible. The important configuration is the execution provider list and the thread count.

import * as ort from 'onnxruntime-web';

ort.env.wasm.numThreads = crossOriginIsolated ? navigator.hardwareConcurrency : 1;
ort.env.wasm.simd = true;
ort.env.wasm.wasmPaths = '/vendor/ort/';        // self-host so the headers are yours

const session = await ort.InferenceSession.create('/models/classifier-int8.onnx', {
  executionProviders: ['wasm'],
  graphOptimizationLevel: 'all',
});

const input = new ort.Tensor('float32', pixels, [1, 3, 224, 224]);
const { logits } = await session.run({ pixel_values: input });

Self-hosting the runtime’s own .wasm files matters for the same reason it matters for any other module: cross-origin isolation requires cooperating headers, and a CDN that does not send Cross-Origin-Resource-Policy will break the threaded build with an error that mentions neither.

The weight-loading problem

Model weights are the payload, and they behave differently from code. A 40 MB ONNX file has to be downloaded, decompressed if it was compressed, parsed, and then copied into linear memory as tensors — which means peak memory during load is often twice the model size. On a phone with a 1.5 GB tab budget, a 200 MB model fails during load rather than during inference, and the error rarely says so.

Three techniques make this manageable, and they compose: quantize the weights so the file is smaller, stream the file so peak memory stays near the final size rather than double it, and cache the parsed result so the second visit skips the work entirely. The weight loading guide works through each, including the external-data format that keeps weights out of the graph file.

Preprocessing is where accuracy quietly dies

A model is a pure function of its input tensor, and almost every “the model works in Python but not in the browser” report is a preprocessing mismatch rather than a runtime bug. There are four places to get it wrong, and all four produce output that looks statistically reasonable.

Resizing is the first. Training pipelines usually resize with a specific filter — bilinear with antialiasing, or bicubic — and the browser’s drawImage uses something else. For a classifier the difference is often a couple of percent of accuracy; for a model that reads text or fine detail it can be catastrophic. Resize deliberately, with the same filter the training code used, which is one of the better reasons to own the resampler.

Normalisation is the second. Most vision models expect channel-wise mean subtraction and division by a standard deviation, in that order, on values already scaled to 0–1. Applying the scale twice, or subtracting the mean after dividing, gives inputs in the wrong range and a model that confidently predicts the wrong class.

Layout is the third: NCHW versus NHWC, and RGB versus BGR. A channel-order swap produces images that are recognisable to a human and meaningless to a network trained the other way round. The fourth is dtype — passing uint8 where float32 is expected either throws or, worse, reinterprets the bytes.

Write the preprocessing once, in the module or in a single well-tested JavaScript function, and validate it by feeding a fixture image through both the browser path and the Python path and comparing the tensors element by element. Five minutes of that saves days.

Choosing a backend

WebAssembly is not the only execution provider available in a browser, and it is frequently not the fastest. The honest comparison looks like this:

Backend Strength Weakness
Wasm (SIMD, threaded) Works everywhere, predictable, no driver surprises 3–10× slower than a GPU for large models
WebGPU Large speedups on big matrix work Availability gaps, driver bugs, shader compile cost
WebNN Uses platform accelerators directly Still shipping; limited operator coverage
WebGL (legacy) Broad support Awkward for modern operators, being retired

The practical pattern is to try WebGPU and fall back to WebAssembly, because the fallback is always available and always correct. For small models the fallback is often faster anyway: a WebGPU session pays shader compilation and buffer upload costs that a 5 ms Wasm inference never recovers. The backend comparison puts numbers against model sizes.

Where the backend crossover sits For small models the fixed costs of a GPU session dominate and the WebAssembly backend wins. As matrix sizes grow the GPU's parallelism takes over, and the crossover for typical models sits in the low tens of megabytes. model size / matrix work → latency crossover Wasm — flat fixed cost WebGPU — high fixed cost, better scaling Measure on your own model rather than trusting the shape of this curve: operator coverage and shader compile time move the crossover a long way.

Threads: the biggest single lever

The threaded WebAssembly backend parallelises GEMM across workers and typically delivers 2.5–3.5× on four cores for transformer and convolutional models. That is a larger improvement than almost any other change you can make, and it is gated entirely on two HTTP headers:

Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp

Without them SharedArrayBuffer is unavailable, the runtime silently falls back to one thread, and your inference is three times slower for reasons no profiler will point at. Check crossOriginIsolated at startup and log it — this single boolean explains more mysterious inference slowdowns than everything else combined. Setting the headers correctly, including what they break in embedded third-party content, is covered in configuring COOP/COEP headers.

Thread count should not exceed the physical core count. navigator.hardwareConcurrency reports logical cores, which on a machine with simultaneous multithreading is double the physical count, and oversubscribing makes GEMM slower rather than faster because the threads fight for the same cache.

Quantization, and what it costs you

Quantizing weights from 32-bit float to 8-bit integer cuts the download by four, cuts memory by four, and usually makes inference faster because the kernels move a quarter of the bytes. For most classification and embedding models the accuracy cost is under one percent — small enough that the only reason not to quantize is that you have not measured it yet.

The cost lands unevenly, though. Models with wide dynamic range in their activations — some audio models, anything with an unbounded output — can degrade badly with naive per-tensor quantization; per-channel quantization usually recovers most of it. Always evaluate the quantized model against a held-out set rather than eyeballing a few examples, because the failure mode is a systematic bias rather than obvious garbage.

Keeping a session alive across visits

Session creation is expensive: fetch, parse, allocate, optimise the graph, warm up the kernels. On a cold visit that is commonly 300 ms to several seconds. Almost all of it can be avoided on the second visit, and the techniques stack.

Cache the model file itself in the Cache Storage API rather than relying on HTTP caching, so you control eviction and can show accurate progress on first download. Cache the compiled WebAssembly module for the runtime as described in caching compiled Wasm modules in IndexedDB. And keep the session object alive for the whole page session — creating one per inference is a mistake that turns a 20 ms operation into a 2 s one.

let sessionPromise = null;
export function getSession() {
  sessionPromise ??= ort.InferenceSession.create(MODEL_URL, OPTIONS)
    .then(async (s) => { await warmUp(s); return s; });   // one throwaway inference
  return sessionPromise;
}

Assigning the promise rather than the resolved session matters: two concurrent callers before the first resolves would otherwise create two sessions, each with its own copy of the weights. On a 50 MB model that is 100 MB of memory and a mysterious doubling of load time.

Latency budgets and what they permit

Before choosing a model, decide what the interaction can tolerate. The numbers below are the ones users actually feel:

  • Under 16 ms — can run per animation frame. Keyword spotting, small pose models, simple filters.
  • Under 100 ms — feels instantaneous on a click. Embeddings, small classifiers, OCR on a crop.
  • Under 1 s — acceptable with a spinner. Segmentation, background removal, a small language model’s first token.
  • Over 1 s — needs a progress indicator and a cancel button, and should probably be on a server.

Run the budget backwards: a 224 × 224 image classifier at int8 on the threaded backend takes 15–40 ms on a laptop and 60–150 ms on a mid-range phone, so it fits the second tier on desktop and the third on mobile. Deciding this early prevents the far more expensive discovery that your architecture only works on the machine you developed it on.

Running inference off the main thread

Even a fast inference is too slow for the main thread once it happens repeatedly. A 40 ms run blocks three animation frames; a 300 ms run makes the page unresponsive in a way users describe as broken. Move the session into a worker and keep the main thread for the interface.

The messaging design matters more than it first appears. Send the smallest representation of the input — an ImageBitmap handle rather than a pixel array, an ArrayBuffer transferred rather than cloned — and return only what the interface needs: the top five labels rather than a thousand logits, the mask as a compact ImageBitmap rather than as float data. A worker that sends back a 4 MB float tensor sixty times a second undoes everything you gained by moving it there.

With the threaded backend the worker will itself spawn workers, which is allowed and works, but it means your thread budget is the pool size multiplied by the threads per session. One inference worker with four backend threads is almost always better than four inference workers with one thread each, because the former shares weights and cache while the latter duplicates them.

What a user waits for, in order Latency is not one number. The download, the compile, the first inference and every inference after it are separate costs with separate fixes. weight download the largest number, and the easiest to cache graph preparation one-off: layout, kernel selection, memory planning first inference includes warmup the steady state does not pay steady-state inference the only figure worth quoting as latency Quote the steady state and the first-run cost separately; a single average hides both. Cache weights in the Cache API or OPFS and the largest cost disappears on the second visit.

Gotchas and failure modes

  • no available backend found. ERR: [wasm] — the runtime’s .wasm files are not where ort.env.wasm.wasmPaths says. Almost always a bundler moving one file and not the other.
  • Inference is 3× slower in production than locally. Cross-origin isolation is off, so threads silently disabled. Log crossOriginIsolated.
  • RangeError: Array buffer allocation failed during session creation on mobile. The model plus its parsing copy exceeded the tab budget. Quantize, or stream the weights.
  • Output shape is [1, 1, 1000] instead of [1, 1000]. Harmless, and a constant source of confusion — squeeze explicitly rather than indexing by guesswork.
  • The first inference is ten times slower than the rest. Kernel selection and memory arena warm-up. Run one inference on a dummy input after session creation and throw the result away.
  • Results differ slightly from the Python reference. Different GEMM ordering changes floating-point accumulation. Compare with a tolerance, not with equality.

Verifying a model actually works

Correctness verification for inference is comparison against a reference run, and it belongs in CI:

# reference logits from the same model, outside the browser
python -c "import onnxruntime as ort, numpy as np; \
  s = ort.InferenceSession('classifier-int8.onnx'); \
  x = np.load('fixture.npy'); \
  np.save('expected.npy', s.run(None, {'pixel_values': x})[0])"

Then assert in a headless browser test that the in-page result matches within a small tolerance — 1e-3 for float32, looser for quantized models. This catches preprocessing mistakes, which are far more common than model bugs: a channel order swap, a missing normalisation, or a resize that used a different filter from the one the model was trained with will all produce plausible-looking nonsense.

Frequently Asked Questions

Can I run a large language model in the browser? Small ones, yes — models in the 0.5 to 3 billion parameter range run with 4-bit quantization, at a few tokens per second on the WebAssembly backend and considerably faster on WebGPU. The blocker is rarely compute; it is the multi-gigabyte download and the memory ceiling of a tab.

Is WebAssembly inference deterministic across devices? Bit-for-bit, no. Threading changes accumulation order in reductions, and different SIMD paths round differently. Results are stable to within a small tolerance, which is fine for classification and wrong for anything that hashes the output.

How much accuracy does int8 quantization cost? Typically 0.2–1% top-1 for vision classifiers with per-channel quantization, and near zero for embedding models measured by retrieval quality. Audio and generative models are more sensitive. Measure on your own evaluation set — the published numbers are for the reference model, not for yours.

Should I ship the model with my bundle or fetch it separately? Separately, always. Models are large, change on a different schedule from your code, and should be cached independently. Bundling a 40 MB model into a JavaScript artifact defeats every caching layer between you and the user.

What about privacy — is local inference actually safer? It is genuinely different: the input never leaves the device, which removes an entire category of compliance and retention questions for photos, audio and documents. It is not a licence to skip consent, and it does not protect the model itself — anything shipped to a browser can be extracted, so treat model weights as public the moment you serve them.

Do I need a different model for mobile? Usually a different size of the same model. Ship the quantized small variant by default, and only upgrade to the larger one after measuring on the device — navigator.hardwareConcurrency and deviceMemory give a crude but useful signal, and a first-run benchmark gives a better one.

Guides in this topic

← Back to Production Wasm: Workloads & Deployment