Multi-Threaded Inference with Wasm Threads

This guide answers one task: make browser inference use more than one core, verify that it actually did, and choose a thread count that helps rather than hurts — because the threaded backend is usually the largest single speedup available and it fails silently when misconfigured.

Prerequisites

  • [ ] Control over the HTTP response headers for your document and its assets.
  • [ ] A runtime with a threaded build — onnxruntime-web, transformers.js, or an Emscripten build with -pthread.
  • [ ] A model whose runtime is dominated by matrix multiplication, which is nearly all of them.
  • [ ] A test device with at least four cores; scaling on two is not informative.

The two headers that gate everything

Threaded WebAssembly requires SharedArrayBuffer, and SharedArrayBuffer requires the document to be cross-origin isolated. That state is granted only when both of these headers are present on the document response:

Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp

The consequences reach further than the document. With require-corp in force, every cross-origin subresource — images, fonts, scripts, iframes, analytics beacons — must opt in with Cross-Origin-Resource-Policy: cross-origin or it is blocked outright. That is why enabling threads is a site-wide decision rather than a page-level one, and why it frequently breaks embedded third-party content the first time it is turned on.

Cross-Origin-Embedder-Policy: credentialless is a softer alternative that loads cross-origin resources without credentials rather than requiring opt-in. It is widely supported now and breaks far less, at the cost of cookies not being sent to those resources. Try credentialless first; fall back to require-corp only if something genuinely needs it.

Every link must hold for threads to work Both headers must be present for the document to become cross-origin isolated, which is what makes SharedArrayBuffer available, which is what the threaded backend needs. Any break in the chain falls back to a single thread without an error. COOP + COEP on the document crossOriginIsolated true in the page SharedArrayBuffer constructor exists threads 2.5–3.5× A break anywhere gives you the single-threaded path with no error, no warning and no clue in the profiler. A missing Cross-Origin-Resource-Policy on one cross-origin image is enough to break isolation for the whole document. Log crossOriginIsolated with your inference timings and the mystery disappears permanently.

Turning it on

With isolation in place, enabling threads is one setting. The runtime detects SharedArrayBuffer, loads the threaded variant of its own module, and spawns a pool of workers.

import * as ort from 'onnxruntime-web';

const cores = navigator.hardwareConcurrency || 2;
ort.env.wasm.numThreads = crossOriginIsolated ? Math.min(4, Math.max(1, cores >> 1)) : 1;
ort.env.wasm.simd = true;

console.log({ crossOriginIsolated, cores, threads: ort.env.wasm.numThreads });

Halving hardwareConcurrency is deliberate. The property reports logical processors, which on a machine with simultaneous multithreading is twice the physical core count. Matrix kernels are already saturating the execution units, so two threads on one physical core contend for the same cache and registers and deliver almost nothing — sometimes less than nothing.

Capping at four is also deliberate for typical models. Beyond four threads the synchronisation between GEMM tiles starts to cost more than the extra parallelism returns, and the browser has other work to do. Measure on your own model before raising the cap.

Why scaling stops short of linear

A perfectly parallel workload on four cores would be four times faster. Inference is not perfectly parallel, and understanding the gap prevents chasing it.

Part of every model is sequential: activation functions, normalisations, reshapes, the graph executor’s own bookkeeping. Amdahl’s law applies directly — if 15% of the runtime is sequential, four threads cap out at about 2.8× no matter how good the parallel part is. Part of the work is memory-bandwidth-bound rather than compute-bound, and adding threads does not add bandwidth. And the thread pool itself costs something: a barrier per operator, cache lines bouncing between cores, and the occasional scheduling delay when the browser has other work.

The observed result across typical vision and transformer models is 1.7–2.0× on two threads, 2.5–3.5× on four, and 3–4.5× on eight — with the eight-thread figure depending heavily on whether those are physical cores. If you are seeing less than 1.5× on four threads, the model is probably too small for the pool to pay for itself, and single-threaded is the right answer.

Measured speedup against the ideal The ideal line rises linearly with thread count. The measured curve tracks it to about two threads and then flattens, because the sequential portion of the graph and memory bandwidth become the limit rather than compute. threads → speedup ideal — linear measured — flattens past four 1 2 4 8 Pick the knee of your own curve rather than the maximum thread count — the extra workers cost memory and battery for no gain past it.

Detecting silently single-threaded traffic

The dangerous failure is not an error; it is production quietly running on one thread while your local machine runs on four. A deploy that drops a header, a CDN that strips one, a new third-party script that forces you to relax require-corp — all of these downgrade performance without breaking anything.

report({
  metric: 'inference',
  crossOriginIsolated,
  threads: ort.env.wasm.numThreads,
  hardwareConcurrency: navigator.hardwareConcurrency,
  p50: timings.p50,
});

Send those four fields with every timing sample. The moment isolation regresses, the ratio of isolated to non-isolated sessions moves and you can see it the same day rather than three sprints later when someone notices the feature feels slow. It is also the only way to know the real distribution across your users, which is rarely what a team assumes: embedded contexts, some in-app browsers and certain enterprise proxies never achieve isolation at all.

Verifying the pool is real

Configuration that looks right is not evidence. Three checks, run once after any change to headers, hosting or the runtime version, tell you whether the pool actually exists.

The first is the isolation flag itself: crossOriginIsolated must be true in the context that creates the session, which for a worker means checking inside the worker rather than on the page. The second is the network panel: the runtime fetches a different .wasm file for the threaded build, and seeing the non-threaded one load is immediate proof that something upstream decided threads were unavailable.

The third and most convincing is a timing comparison you run yourself:

async function threadScaling(url, feeds, counts = [1, 2, 4]) {
  const rows = [];
  for (const n of counts) {
    ort.env.wasm.numThreads = n;
    const s = await ort.InferenceSession.create(url, { executionProviders: ['wasm'] });
    await s.run(feeds);                                   // warm up
    const t = performance.now();
    for (let i = 0; i < 20; i++) await s.run(feeds);
    rows.push({ threads: n, msPerRun: (performance.now() - t) / 20 });
    await s.release?.();
  }
  console.table(rows);
}

If the three rows are within a few percent of each other, the pool is not doing anything regardless of what the configuration says. Note that numThreads must be set before the session is created — changing it afterwards has no effect, which is a common source of benchmarks that appear to show threads making no difference at all.

Threads in a worker

Running inference in a worker and running it with threads are independent choices, and you usually want both. The worker keeps the main thread free; the threads make the inference itself faster. A worker can spawn nested workers, which is what the threaded runtime does internally, and that works in every current browser as long as the isolation state is inherited — which it is, since workers inherit their creator’s agent cluster.

What you should not do is run several inference workers each with their own thread pool. Four workers with four threads each means sixteen threads competing for four cores, plus four copies of the model in memory. One worker holding one session with a four-thread pool is faster and uses a quarter of the memory. If you genuinely need concurrent inferences, queue them into the single session — the runtime serialises them anyway.

Where extra threads stop helping Scaling is good to about four threads and then flattens, because the work becomes bound by memory bandwidth rather than arithmetic. 1 thread 140 ms 2 threads 76 ms 1.8x 4 threads 46 ms 3.0x — the last worthwhile step 8 threads 41 ms 3.4x, and the page has no cores left for anything else Leave cores for the page: a pool sized to hardwareConcurrency makes the interface stutter during inference. Threads need cross-origin isolation, so a build without those headers silently falls back to one thread.

Gotchas

  • SharedArrayBuffer is not defined. Isolation is off. Check the document’s response headers, not the asset’s.
  • Everything works locally, threads are off in production. A CDN or reverse proxy is stripping or overriding the headers. Verify with curl -I against the production URL.
  • Enabling COEP breaks third-party embeds. Expected. Move to credentialless, or host the resource yourself, or ask the provider for Cross-Origin-Resource-Policy: cross-origin.
  • More threads made it slower. Oversubscription, usually from trusting hardwareConcurrency on a machine with simultaneous multithreading. Halve it.
  • Thread count changes nothing. The runtime loaded the non-threaded variant because that file was missing from your vendor directory. Check the network panel for which .wasm was actually fetched.

Performance note

On a four-physical-core laptop with a quantized transformer encoder: single-threaded p50 was 118 ms, two threads 64 ms (1.84×), four threads 41 ms (2.88×), eight threads 38 ms (3.11×). Memory grew by about 9 MB per additional thread for per-thread arenas. The four-thread configuration is the obvious choice — the eighth thread bought 3 ms and cost 36 MB.

Frequently Asked Questions

Is credentialless safe to use? It is a standard, supported mode designed for exactly this. Cross-origin resources load without credentials, so anything requiring cookies — an authenticated image CDN, for instance — will fail and needs hosting differently. For most sites it is strictly better than require-corp.

Do threads help on mobile? Less than on desktop. Mobile chips have fewer performance cores and aggressive thermal management, so a four-thread pool can throttle within seconds. Two threads is often the sweet spot on phones, and worth measuring separately rather than sharing the desktop configuration.

Can I enable isolation for one route only? Yes — the headers are per-response, so a single route can be isolated while the rest of the site is not. This is the usual way to ship threads without auditing every third-party script on every page.

← Back to Machine Learning Inference in the Browser