Loading Large Model Weights into Linear Memory

This guide answers one task: load a model whose weights are tens or hundreds of megabytes into a browser tab, without the peak memory spike that kills the page on mobile and without leaving the user staring at a blank screen while it happens.

Prerequisites

  • [ ] A model larger than about 30 MB — below that, none of this is worth the complexity.
  • [ ] A runtime that accepts weights as a buffer, such as onnxruntime-web or a hand-compiled kernel.
  • [ ] Cache Storage available, which means a secure context.
  • [ ] A device with limited memory for testing. Everything works on a workstation.

Why peak memory is roughly double the model

The obvious load path allocates the model several times over, and each copy is live simultaneously. The response body arrives as a stream of chunks; those chunks are collected into an array; the array is concatenated into one contiguous ArrayBuffer; that buffer is handed to the runtime, which parses it and copies the weight tensors into linear memory. At the moment of the copy, the original buffer and the runtime’s arena both exist — so a 200 MB model briefly needs 400 MB, plus the chunk array if it has not been released yet.

On a desktop that spike passes unnoticed. On a phone with a tab budget between 512 MB and 1.5 GB it is the difference between working and a crash whose only symptom is the page reloading itself. The error, when you get one at all, is RangeError: Array buffer allocation failed — which names the symptom and nothing about the cause.

The fix is to reduce the number of simultaneous copies rather than the size of any one of them. Three techniques do that, and they combine: preallocate a single destination and stream directly into it, use the runtime’s external-data mechanism so weights never pass through a JavaScript buffer at all, and release every intermediate reference the moment it is no longer needed.

Where the second copy comes from A naive load holds the chunk array, the concatenated buffer and the runtime arena at overlapping times, peaking at roughly twice the model size. Streaming into one preallocated destination keeps the peak close to the model size itself. naive — peak ≈ 2× model chunk array concatenated ArrayBuffer runtime arena in linear memory streamed — peak ≈ 1× model one preallocated destination, written chunk by chunk The overlap in the top row is what fails on a phone — three live allocations whose lifetimes were never designed to be disjoint. Nothing in the bottom row is faster; it simply never holds two copies at once.

Streaming into one preallocated buffer

If the server sends Content-Length, you know the final size before the first byte arrives, which means you can allocate once and fill it. This removes the chunk array and the concatenation in one step.

export async function fetchIntoBuffer(url, onProgress) {
  const res = await fetch(url);
  if (!res.ok) throw new Error(`${url}: ${res.status}`);
  const total = Number(res.headers.get('content-length'));
  if (!total) return new Uint8Array(await res.arrayBuffer());   // no length: fall back

  const out = new Uint8Array(total);                            // the only large allocation
  const reader = res.body.getReader();
  let offset = 0;
  for (;;) {
    const { value, done } = await reader.read();
    if (done) break;
    out.set(value, offset);
    offset += value.length;
    onProgress?.(offset / total);
  }
  if (offset !== total) throw new Error(`short read: ${offset} of ${total}`);
  return out;
}

Two things make this worth the twenty lines. The progress callback is real rather than fake, so the user sees a bar that means something during a forty-second download — which, on a slow connection, is the difference between waiting and leaving. And the short-read check catches truncated responses, which otherwise surface much later as a parse error inside the runtime that gives no hint that the file was incomplete.

Make sure the server is not gzipping an already-compressed model: with Content-Encoding: gzip the Content-Length describes the compressed size, the decoded stream is larger, and the preallocated buffer overflows. Serve model files with no transfer compression and let the format’s own compression do the work.

External data: keeping weights out of the graph file

ONNX supports storing tensors in separate files rather than inside the model protobuf, and runtimes support loading them. This matters at scale because the graph file stays small and parseable while the weights are fetched as plain binary blobs that never need protobuf decoding.

const session = await ort.InferenceSession.create('/models/llm.onnx', {
  executionProviders: ['wasm'],
  externalData: [
    { path: 'llm.onnx_data', data: await fetchIntoBuffer('/models/llm.onnx_data', setPct) },
  ],
});

The graph file might be 400 kB while the data file is 600 MB. Because they are separate resources, they cache separately, they can be fetched in parallel, and a graph change that does not touch weights does not invalidate the large download. For models that ship in several quantized variants this is particularly valuable — one graph, several data files.

Caching so the second visit is instant

A large model should be downloaded once per device, not once per visit. The Cache Storage API is the right tool: it survives reloads, you control eviction, and entries can be checked before you commit to a network request.

const CACHE = 'model-weights-v3';                  // bump when the weights change

async function cachedFetch(url, onProgress) {
  const cache = await caches.open(CACHE);
  const hit = await cache.match(url);
  if (hit) return new Uint8Array(await hit.arrayBuffer());
  const res = await fetch(url);
  await cache.put(url, res.clone());                // store the compressed original
  return streamFrom(res, onProgress);
}

// drop old versions so storage does not grow without bound
for (const key of await caches.keys()) if (key.startsWith('model-weights-') && key !== CACHE) await caches.delete(key);

Version the cache name rather than the entry, so a weight change evicts cleanly. And check the storage quota before caching hundreds of megabytes — navigator.storage.estimate() reports what is available, and writing past the quota throws QuotaExceededError at an unhelpful moment.

First visit and every visit after The first visit pays the full download and stores the bytes in the cache. Later visits read from the cache and go straight to session creation, turning a forty-second wait into a few hundred milliseconds. first visit download 240 MB — show real progress cache.put create session later visits cache.match create session — no network, no progress bar needed Storing the compressed response rather than the decoded bytes keeps the cache small; the decode happens on read, where it is cheap. Always offer a way to clear the cache: a user who downloaded 240 MB should be able to get it back.

Releasing what you no longer need

The last technique is unglamorous and effective: drop references as soon as the runtime has taken ownership. JavaScript will not collect a 240 MB buffer that a module-scope variable still points at, and in long-lived pages that reference is easy to leave behind.

let bytes = await cachedFetch(WEIGHTS_URL, setPct);
const session = await ort.InferenceSession.create(graph, { externalData: [{ path: 'w.bin', data: bytes }] });
bytes = null;                                     // the runtime copied what it needs

Verify rather than assume. Take a heap snapshot after session creation and look for detached buffers of the model’s size — if one is still retained, something is holding it. The usual culprits are a closure capturing the buffer for a progress callback, an entry in a Map used for deduplication, or a console log that pinned the object during development and never got removed.

Getting weights in without doubling memory Growing memory to the final size first, then streaming the weights directly into that region, avoids ever holding the file and the copy at the same time. cached response from Cache API grow once final size up front stream into region chunk by chunk session ready no second copy Buffering the file first doubles peak memory, which is exactly what a phone refuses to allocate. Growing once avoids repeated copies inside the engine as the memory is resized. Weights are immutable, so they can be mapped read-only and shared between sessions on the same page.

Gotchas

  • RangeError: Array buffer allocation failed on mobile. Peak memory, not model size. Stream into a preallocated buffer and drop intermediates.
  • Progress bar jumps from 0% to 100%. The response had no Content-Length, usually because of chunked transfer encoding or a proxy. Report bytes received rather than a percentage in that case.
  • QuotaExceededError from cache.put. The origin’s storage budget is exhausted. Check navigator.storage.estimate() first, and degrade to a network-only path rather than failing.
  • The buffer overflows on out.set(value, offset). The server is compressing the response; Content-Length is the compressed size. Disable transfer compression for model files.
  • Session creation takes minutes on a phone. Parsing hundreds of megabytes of protobuf is expensive. External data avoids most of it, and quantizing first avoids more — see quantizing models for Wasm inference.

Performance note

For a 240 MB quantized model on a mid-range laptop over a 100 Mbit connection: the naive path peaked at about 500 MB of tab memory and took 27 s; the streamed path with external data peaked at about 280 MB and took the same 27 s, since the network dominates. On a second visit with the cache warm, session creation alone was 1.9 s. The lesson is that streaming buys survival, not speed — and caching buys the speed.

Frequently Asked Questions

Should I split the model into several files and fetch them in parallel? Usually yes for very large models: parallel connections use available bandwidth better, and each piece can be cached and retried independently. Keep the pieces to a handful; dozens of requests cost more in overhead than they gain.

Can I load weights directly from the Origin Private File System? Yes, and it is a good fit for very large models the user has downloaded once — see persisting a Wasm database to OPFS for the access patterns. Reading from OPFS avoids the Cache Storage size limits on some platforms.

Does compressing the model file help? Only if it is not already compressed. Quantized weights have high entropy and compress by a few percent at best, so gzip or Brotli on a model file usually costs CPU for nothing. Test before enabling it.

How do I handle a user who navigates away mid-download? Abort the fetch with an AbortController when the view unmounts, and do not cache a partial response — cache.put on an aborted response stores a truncated body that will fail to parse on the next visit with an error pointing at the model rather than at the cache. Delete the entry on any failure path.

Is it worth showing an estimated time remaining? Yes for anything over about ten seconds, and it is cheap: track bytes per second over a rolling window and divide the remainder by it. Users tolerate a long wait they can see the end of far better than a spinner, and the same number tells you when to offer a lower-quality model instead.

← Back to Machine Learning Inference in the Browser