Measuring Inference Latency in the Browser
This guide answers one task: measure how long browser inference actually takes, in a way that survives scrutiny — separating the costs that behave differently, discarding warm-up, reporting percentiles, and collecting the same numbers from real users rather than only from your laptop.
Prerequisites
- [ ] A working inference session — see running ONNX models with onnxruntime-web.
- [ ] A fixed fixture input, so every measurement runs identical work.
- [ ] Somewhere to send field data: an analytics endpoint, or the console during development.
- [ ] A quiet machine for local runs. Benchmarks taken while a build is running are fiction.
Four costs, four numbers
“How fast is inference?” is four questions, and merging them produces a number that answers none of them.
Model download is network-bound, varies by hundreds of times between connections, and is zero on a warm cache. Session creation parses the graph, allocates arenas and applies optimisations; it is paid once per session and is often larger than everything else on a page that runs one inference. Warm-up is the first inference, which selects kernels and touches cold pages, and is routinely five to ten times the steady-state cost. Steady-state inference is the number people mean when they say “inference time”, and it is the only one that reflects the model and backend rather than the surrounding machinery.
Report all four. A team that only tracks the last one will be baffled when users describe a feature as slow; a team that only tracks a total will not know which lever to pull.
A harness that produces defensible numbers
The measurement itself is short, and every line in it is there for a reason.
export async function benchmark(session, feeds, { runs = 50, warmups = 3 } = {}) {
for (let i = 0; i < warmups; i++) await session.run(feeds); // discarded entirely
const samples = [];
for (let i = 0; i < runs; i++) {
const t0 = performance.now();
await session.run(feeds);
samples.push(performance.now() - t0);
}
samples.sort((a, b) => a - b);
const at = (p) => samples[Math.min(samples.length - 1, Math.floor(samples.length * p))];
return {
runs,
p50: +at(0.5).toFixed(2),
p90: +at(0.9).toFixed(2),
p99: +at(0.99).toFixed(2),
min: +samples[0].toFixed(2),
};
}
Three warm-ups rather than one, because kernel selection can happen lazily across the first few runs. Fifty runs rather than ten, because the distribution has a tail and ten samples cannot describe it. And percentiles rather than a mean, because a single 400 ms outlier from a garbage collection drags a mean far more than it drags a median — while being exactly the thing your p99 should show.
Feed the same fixture every time. Varying inputs means varying work — different sequence lengths, different shapes, different early exits — and turns the measurement into a mixture of model behaviour and input distribution.
Timer resolution and what it hides
performance.now() is clamped for security reasons: resolution is commonly 100 microseconds in isolated
contexts and can be coarser elsewhere. For a 30 ms inference that is irrelevant. For a 0.4 ms operator
it is the entire measurement, and you will see quantised values — a suspicious run of identical numbers
is the signature.
When you need to measure something below a millisecond, measure a batch and divide:
const N = 200;
const t = performance.now();
for (let i = 0; i < N; i++) await session.run(feeds);
const perRun = (performance.now() - t) / N; // resolution no longer matters
The tradeoff is that batching hides the distribution, so use it for micro-measurements and the percentile harness for anything you will report. Never mix the two in one table without saying which is which.
Measuring session creation and download separately
Both are single-shot costs, which means percentiles over repeated runs are meaningless — you need repeated page loads instead. Locally, that is a loop over a fresh context; in the field, it is one sample per session.
const t0 = performance.now();
const bytes = await fetchModel(url);
const t1 = performance.now();
const session = await ort.InferenceSession.create(bytes, opts);
const t2 = performance.now();
await session.run(fixture);
const t3 = performance.now();
report({
downloadMs: t1 - t0,
createMs: t2 - t1,
warmupMs: t3 - t2,
cached: performance.getEntriesByName(url)[0]?.transferSize === 0,
});
The transferSize === 0 check distinguishes a cache hit from a fast network, which matters enormously
when interpreting field data: the two produce similar download times and completely different
conclusions about what to optimise.
Field measurement beats lab measurement
Your development machine is the fastest device any of your users will ever have, and it is plugged in. Field data from real sessions is the only way to know what the feature actually costs, and the collection is a handful of fields.
Send: the four timings, crossOriginIsolated, the backend that was selected, navigator.hardwareConcurrency,
navigator.deviceMemory when present, and a coarse device class derived from the user agent. Aggregate
by backend and by device class rather than globally — a global p50 across desktop and mobile describes
nobody.
What you usually learn is uncomfortable and useful: the GPU path is selected less often than expected, isolation fails for a meaningful slice of traffic, and the p95 on mobile is several times the desktop p50. All three change product decisions, and none of them are visible locally.
Comparing against JavaScript honestly
If the point of the exercise is to justify WebAssembly, the comparison has to be fair. Use the same algorithm, the same input, the same warm-up discipline, and measure the JavaScript version after the JIT has warmed — which takes more iterations than most people allow. Include the boundary crossing and the memory copy in the WebAssembly measurement, because users pay for those too.
It is equally important to report where the comparison does not favour the compiled version. Small inputs, string-heavy work and anything called at high frequency with tiny payloads usually lose. A benchmark page that shows both is far more persuasive than one that shows only the win, and it prevents the team from applying the result somewhere it does not hold. The throughput comparison guide works through the methodology in detail.
Gotchas
- Measuring the first run. It includes kernel selection and cold pages. Discard at least three.
awaitinside the timed region doing more than you think. Ifrun()returns a promise resolved from a worker, you are timing postMessage latency as well. That may be what you want — say so.- Comparing across browsers on different machines. Engine differences are real but smaller than hardware differences. Same machine, same session, or the numbers are noise.
- Battery and thermal state. A laptop on battery throttles aggressively. Plug in for lab measurements, and expect field p95 to be worse for exactly this reason.
- Benchmarking with DevTools open. The profiler adds overhead, sometimes 20% or more. Close it before taking numbers.
Performance note
A useful rule from repeated measurement: session creation is roughly 1.5–3 ms per megabyte of model on a laptop and two to four times that on a phone, warm-up is three to ten times a steady-state run, and the p99 is typically 2.5–4× the p50 on a busy page. Those ratios let you sanity-check a new measurement immediately — a p99 fifty times the p50 means something else is happening, usually a garbage collection triggered by allocation inside the inference loop.
Frequently Asked Questions
How many samples do I need? Fifty for a stable median, a few hundred for a meaningful p99. If a run takes 400 ms, take fewer and report only p50 and p90 rather than pretending to a p99 you cannot support.
Should I use console.time?
For quick local checks, fine. For anything recorded, use performance.now() into an array — console.time
writes to the console, which itself costs time and is not available in field collection.
Can I use the Performance panel instead? Yes, and it is the right tool for finding where time goes inside an inference. It is the wrong tool for producing numbers, because the instrumentation overhead inflates them. Profile to diagnose, benchmark to report.
How do I attribute a slowdown to a specific change? Keep a fixture and a committed baseline file of the four numbers, and compare every run against it in CI on the same runner class. Absolute numbers from a shared CI machine are noisy, but the ratio between two runs on the same machine is stable enough to catch a regression of 15% or more, which is where the regressions that matter live.
Is it worth sampling every session in production? No. One in fifty sessions is plenty for stable percentiles at any reasonable traffic level, and it keeps the analytics volume and the client-side cost negligible. Sample deterministically from a session identifier so the same session is either always measured or never measured, rather than contributing a partial picture.
Related
- Building a reproducible Wasm benchmark harness — the same discipline as a CI gate.
- Choosing between the Wasm and WebGPU backends — the decision these numbers inform.
- Measuring Wasm compile time in DevTools — the module-side equivalent.