Running ONNX Models with onnxruntime-web

This guide answers one task: load an ONNX model in a browser with onnxruntime-web, feed it a correctly shaped tensor, and read the result — with the runtime files served from your own origin so threads and SIMD actually engage.

Prerequisites

  • [ ] onnxruntime-web 1.19 or later from npm.
  • [ ] An .onnx model with known input and output names — check them with onnx.checker or Netron.
  • [ ] A build step that can copy files into your static output directory.
  • [ ] Optional but recommended: COOP/COEP headers so the threaded backend is available.

Install and self-host the runtime files

onnxruntime-web ships its execution engine as separate .wasm and .mjs files. The library fetches them at runtime from a path you control, and getting that path wrong is the single most common setup failure. Copy them into your public directory as part of the build rather than relying on a CDN.

npm i onnxruntime-web
# copy the runtime artifacts next to your app's static assets
mkdir -p public/vendor/ort
cp node_modules/onnxruntime-web/dist/*.wasm  public/vendor/ort/
cp node_modules/onnxruntime-web/dist/*.mjs   public/vendor/ort/

Then tell the library where they live, before creating any session:

import * as ort from 'onnxruntime-web';

ort.env.wasm.wasmPaths = '/vendor/ort/';
ort.env.wasm.simd = true;
ort.env.wasm.numThreads = crossOriginIsolated ? Math.min(4, navigator.hardwareConcurrency || 1) : 1;

Self-hosting also removes the cross-origin resource policy problem: with require-corp in force, a cross-origin .wasm without the matching Cross-Origin-Resource-Policy header simply fails to load, and the error points at the wrong thing.

Three things the browser fetches The application bundle, the runtime's own WebAssembly files, and the model weights are three separate downloads with different sizes and cache lifetimes. Serving all of them from your origin keeps cross-origin isolation intact. app bundle tens of kilobytes changes every deploy hashed filename /vendor/ort/*.wasm 2–10 MB, several variants changes only on upgrade cache for a year model.onnx megabytes to hundreds own version, own cache fetch on demand The runtime picks which variant to fetch based on the detected features — SIMD, threads — so several files must be present even though only one loads. Copying only the one that worked in development is the classic cause of a production-only failure on a different browser.

Create a session

Session creation parses the graph, allocates tensors and applies graph optimisations. Do it once, keep the result, and never create a session per inference.

const session = await ort.InferenceSession.create('/models/mobilenet-int8.onnx', {
  executionProviders: ['wasm'],
  graphOptimizationLevel: 'all',
});

console.log(session.inputNames, session.outputNames);
// [ 'pixel_values' ] [ 'logits' ]

Log the input and output names the first time you integrate a model. Guessing them is how you end up with Error: invalid input 'input', and the names are model-specific: exporters variously produce input, input.1, images or pixel_values for what is conceptually the same tensor.

Build the input tensor

The tensor is a flat typed array plus a shape. For a vision model the layout is almost always NCHW — batch, channel, height, width — which means the three colour channels are stored as three contiguous planes rather than interleaved per pixel as they are in ImageData.

function toNCHW(imageData, mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225]) {
  const { data, width: w, height: h } = imageData;
  const out = new Float32Array(3 * w * h);
  const plane = w * h;
  for (let i = 0, p = 0; i < data.length; i += 4, p++) {
    out[p]             = (data[i]     / 255 - mean[0]) / std[0];   // R plane
    out[p + plane]     = (data[i + 1] / 255 - mean[1]) / std[1];   // G plane
    out[p + 2 * plane] = (data[i + 2] / 255 - mean[2]) / std[2];   // B plane
  }
  return new ort.Tensor('float32', out, [1, 3, h, w]);
}

Two things are easy to get wrong here and hard to notice. The mean and standard deviation must match the ones used in training — the values above are the ImageNet defaults, and a model trained differently will produce confident nonsense with them. And the alpha channel is skipped entirely; including it shifts every subsequent plane and produces an image the model sees as scrambled.

Run and read the output

run() takes an object keyed by input name and resolves to an object keyed by output name. The result’s data is a typed array in the output’s declared dtype.

const feeds = { [session.inputNames[0]]: tensor };
const results = await session.run(feeds);
const logits = results[session.outputNames[0]].data;    // Float32Array(1000)

const top = [...logits]
  .map((score, index) => ({ index, score }))
  .sort((a, b) => b.score - a.score)
  .slice(0, 5);

If the model emits logits rather than probabilities — most do — apply softmax yourself before showing numbers to a user. Reporting raw logits as confidence produces values above one and below zero, which looks like a bug to everyone who sees it.

Expected output

A successful first run prints the shapes you expect and a plausible ranking:

inputs  [ 'pixel_values' ]
outputs [ 'logits' ]
shape   [ 1, 1000 ]
top5    [ { index: 281, score: 9.41 },  // tabby cat
          { index: 285, score: 8.77 },
          { index: 282, score: 7.12 }, ... ]

The sanity check that catches preprocessing mistakes is running a fixture image whose expected class you know, and comparing against the same model outside the browser. If the browser’s top class differs from Python’s, the problem is upstream of the runtime — resize filter, normalisation or channel order — essentially every time.

The five steps between a picture and a prediction An image is resized, normalised, reshaped into a planar tensor, passed to the session, and the resulting logits are turned into probabilities. Four of the five steps are yours, and all four are places accuracy is lost. resize 224² match training filter normalise mean and std to NCHW three planes session.run the only part not yours softmax Four of these five boxes are code you wrote, which is why a wrong prediction is far more often a preprocessing bug than a runtime bug. Validate each step against the reference pipeline once, then never touch them again.

Loading the model file yourself

InferenceSession.create() accepts a URL, but it also accepts an ArrayBuffer or a Uint8Array. Taking the fetch into your own hands is worth it as soon as the model is more than a few megabytes, because it gives you three things the URL form cannot: a progress indicator, control over caching, and the ability to load from somewhere other than the network.

async function loadModelBytes(url, onProgress) {
  const cache = await caches.open('models-v1');
  let res = await cache.match(url);
  if (!res) {
    res = await fetch(url);
    await cache.put(url, res.clone());            // keep it for next time
  }
  const total = Number(res.headers.get('content-length')) || 0;
  const reader = res.body.getReader();
  const chunks = [];
  let seen = 0;
  for (;;) {
    const { value, done } = await reader.read();
    if (done) break;
    chunks.push(value);
    seen += value.length;
    if (total) onProgress(seen / total);
  }
  const bytes = new Uint8Array(seen);
  let off = 0;
  for (const c of chunks) { bytes.set(c, off); off += c.length; }
  return bytes;
}

const session = await ort.InferenceSession.create(await loadModelBytes(MODEL_URL, setPct), OPTIONS);

The Cache Storage API is the right place for model files rather than plain HTTP caching, because you decide when an entry is evicted and you can version the cache name when the model changes. Note that the assembled Uint8Array is a full copy of the model in memory, which is briefly doubled while the session parses it — the reason very large models need the streaming approach described in loading large model weights.

Handling models with more than one input

Vision classifiers take one tensor; almost everything else does not. Text models want input_ids and attention_mask; encoder-decoder models add decoder_input_ids; models with past-key-value caching want a tensor per layer. The feeds object simply grows, and the shapes must be exact.

const feeds = {
  input_ids:      new ort.Tensor('int64', BigInt64Array.from(ids.map(BigInt)), [1, ids.length]),
  attention_mask: new ort.Tensor('int64', BigInt64Array.from(ids.map(() => 1n)), [1, ids.length]),
};
const { last_hidden_state } = await session.run(feeds);

Two details trip people up here. Token identifiers are int64 in most exported models, which means BigInt64Array and BigInt values — passing a plain Int32Array throws a type error that names the tensor but not the reason. And the sequence length appears in several shapes at once; if the mask and the identifiers disagree by even one element, the runtime reports a broadcast failure deep inside an operator rather than at the boundary where you can see it.

Warm up before you measure anything

The first run() after session creation allocates arenas, selects kernels and touches cold pages. It is routinely five to ten times slower than the steady state, and benchmarking without discarding it produces numbers that are simply wrong.

const dummy = new ort.Tensor('float32', new Float32Array(3 * 224 * 224), [1, 3, 224, 224]);
await session.run({ [session.inputNames[0]]: dummy });   // discard

Do the warm-up during a moment the user is not waiting — right after session creation, while they are still choosing an input. It also surfaces shape and name errors immediately rather than on the first real interaction.

One session, many inferences Creating the session is the expensive step: it parses the graph, picks kernels and plans memory. It happens once, and every inference after it reuses the result. model bytes fetched and cached InferenceSession graph prepared once input tensor typed array over memory output tensor read and released Creating a session per inference is the commonest mistake and costs hundreds of milliseconds each time. Reuse the input tensor's backing buffer between runs so the allocator stays out of the hot path. Release outputs you no longer need; they hold memory the next inference would otherwise reuse.

Gotchas

  • no available backend found. ERR: [wasm] ... — wasmPaths is wrong, or the variant the runtime wants was not copied. Check the network panel for a 404 on a .wasm or .mjs under your vendor path.
  • Error: invalid input 'x' — the feed key does not match session.inputNames[0]. Never hard-code the name; read it from the session.
  • Threads silently disabled. crossOriginIsolated is false. The runtime falls back to one thread and says nothing, and inference is three times slower than your local measurements.
  • Cannot read properties of undefined (reading 'data') — you indexed the results object with the wrong output name. Log session.outputNames once.
  • Bundler rewrites the worker URL. Some bundlers try to process the runtime’s .mjs worker file. Mark the vendor directory as an external static asset rather than letting the bundler touch it — the same class of problem described in bundling Wasm ESM with Vite.

Performance note

A quantized MobileNet-class model at 224 × 224 runs in roughly 12–25 ms on a laptop with SIMD and four threads, and 45–120 ms on a mid-range phone. Session creation for the same model is 150–400 ms, which is why it must not be repeated. The preprocessing loop above costs about 1.5 ms for a 224 × 224 image in JavaScript — worth moving into the module only if you are running it per video frame, in which case the copy and the conversion fuse into the same pass.

Frequently Asked Questions

Does this work in a Web Worker? Yes, and it is the recommended deployment. Everything in this page runs unchanged in a worker; only the image acquisition has to move, using createImageBitmap and OffscreenCanvas rather than a DOM canvas.

How do I run several inputs at once? Batch them into one tensor with a first dimension greater than one, if the model’s graph allows it. Batching four images typically costs far less than four separate runs, because the GEMM kernels get larger matrices to work with.

Can I use a model that expects dynamic shapes? Yes — pass the actual shape in the tensor and the runtime resolves it per run. Expect a slower first run for each new shape, since kernel selection happens per resolved shape, so a page that feeds arbitrary sizes will warm up repeatedly. Padding to a small set of fixed sizes is usually faster overall.

← Back to Machine Learning Inference in the Browser