Loading Large Model Weights into Linear Memory
This guide answers one task: load a model whose weights are tens or hundreds of megabytes into a browser tab, without the peak memory spike that kills the page on mobile and without leaving the user staring at a blank screen while it happens.
Prerequisites
- [ ] A model larger than about 30 MB — below that, none of this is worth the complexity.
- [ ] A runtime that accepts weights as a buffer, such as
onnxruntime-webor a hand-compiled kernel. - [ ]
Cache Storageavailable, which means a secure context. - [ ] A device with limited memory for testing. Everything works on a workstation.
Why peak memory is roughly double the model
The obvious load path allocates the model several times over, and each copy is live simultaneously. The
response body arrives as a stream of chunks; those chunks are collected into an array; the array is
concatenated into one contiguous ArrayBuffer; that buffer is handed to the runtime, which parses it and
copies the weight tensors into linear memory. At the moment of the copy, the original buffer and the
runtime’s arena both exist — so a 200 MB model briefly needs 400 MB, plus the chunk array if it has not
been released yet.
On a desktop that spike passes unnoticed. On a phone with a tab budget between 512 MB and 1.5 GB it is
the difference between working and a crash whose only symptom is the page reloading itself. The error,
when you get one at all, is RangeError: Array buffer allocation failed — which names the symptom and
nothing about the cause.
The fix is to reduce the number of simultaneous copies rather than the size of any one of them. Three techniques do that, and they combine: preallocate a single destination and stream directly into it, use the runtime’s external-data mechanism so weights never pass through a JavaScript buffer at all, and release every intermediate reference the moment it is no longer needed.
Streaming into one preallocated buffer
If the server sends Content-Length, you know the final size before the first byte arrives, which means
you can allocate once and fill it. This removes the chunk array and the concatenation in one step.
export async function fetchIntoBuffer(url, onProgress) {
const res = await fetch(url);
if (!res.ok) throw new Error(`${url}: ${res.status}`);
const total = Number(res.headers.get('content-length'));
if (!total) return new Uint8Array(await res.arrayBuffer()); // no length: fall back
const out = new Uint8Array(total); // the only large allocation
const reader = res.body.getReader();
let offset = 0;
for (;;) {
const { value, done } = await reader.read();
if (done) break;
out.set(value, offset);
offset += value.length;
onProgress?.(offset / total);
}
if (offset !== total) throw new Error(`short read: ${offset} of ${total}`);
return out;
}
Two things make this worth the twenty lines. The progress callback is real rather than fake, so the user sees a bar that means something during a forty-second download — which, on a slow connection, is the difference between waiting and leaving. And the short-read check catches truncated responses, which otherwise surface much later as a parse error inside the runtime that gives no hint that the file was incomplete.
Make sure the server is not gzipping an already-compressed model: with Content-Encoding: gzip the
Content-Length describes the compressed size, the decoded stream is larger, and the preallocated buffer
overflows. Serve model files with no transfer compression and let the format’s own compression do the
work.
External data: keeping weights out of the graph file
ONNX supports storing tensors in separate files rather than inside the model protobuf, and runtimes support loading them. This matters at scale because the graph file stays small and parseable while the weights are fetched as plain binary blobs that never need protobuf decoding.
const session = await ort.InferenceSession.create('/models/llm.onnx', {
executionProviders: ['wasm'],
externalData: [
{ path: 'llm.onnx_data', data: await fetchIntoBuffer('/models/llm.onnx_data', setPct) },
],
});
The graph file might be 400 kB while the data file is 600 MB. Because they are separate resources, they cache separately, they can be fetched in parallel, and a graph change that does not touch weights does not invalidate the large download. For models that ship in several quantized variants this is particularly valuable — one graph, several data files.
Caching so the second visit is instant
A large model should be downloaded once per device, not once per visit. The Cache Storage API is the right tool: it survives reloads, you control eviction, and entries can be checked before you commit to a network request.
const CACHE = 'model-weights-v3'; // bump when the weights change
async function cachedFetch(url, onProgress) {
const cache = await caches.open(CACHE);
const hit = await cache.match(url);
if (hit) return new Uint8Array(await hit.arrayBuffer());
const res = await fetch(url);
await cache.put(url, res.clone()); // store the compressed original
return streamFrom(res, onProgress);
}
// drop old versions so storage does not grow without bound
for (const key of await caches.keys()) if (key.startsWith('model-weights-') && key !== CACHE) await caches.delete(key);
Version the cache name rather than the entry, so a weight change evicts cleanly. And check the storage
quota before caching hundreds of megabytes — navigator.storage.estimate() reports what is available,
and writing past the quota throws QuotaExceededError at an unhelpful moment.
Releasing what you no longer need
The last technique is unglamorous and effective: drop references as soon as the runtime has taken ownership. JavaScript will not collect a 240 MB buffer that a module-scope variable still points at, and in long-lived pages that reference is easy to leave behind.
let bytes = await cachedFetch(WEIGHTS_URL, setPct);
const session = await ort.InferenceSession.create(graph, { externalData: [{ path: 'w.bin', data: bytes }] });
bytes = null; // the runtime copied what it needs
Verify rather than assume. Take a heap snapshot after session creation and look for detached buffers of
the model’s size — if one is still retained, something is holding it. The usual culprits are a closure
capturing the buffer for a progress callback, an entry in a Map used for deduplication, or a console
log that pinned the object during development and never got removed.
Gotchas
RangeError: Array buffer allocation failedon mobile. Peak memory, not model size. Stream into a preallocated buffer and drop intermediates.- Progress bar jumps from 0% to 100%. The response had no
Content-Length, usually because of chunked transfer encoding or a proxy. Report bytes received rather than a percentage in that case. QuotaExceededErrorfromcache.put. The origin’s storage budget is exhausted. Checknavigator.storage.estimate()first, and degrade to a network-only path rather than failing.- The buffer overflows on
out.set(value, offset). The server is compressing the response;Content-Lengthis the compressed size. Disable transfer compression for model files. - Session creation takes minutes on a phone. Parsing hundreds of megabytes of protobuf is expensive. External data avoids most of it, and quantizing first avoids more — see quantizing models for Wasm inference.
Performance note
For a 240 MB quantized model on a mid-range laptop over a 100 Mbit connection: the naive path peaked at about 500 MB of tab memory and took 27 s; the streamed path with external data peaked at about 280 MB and took the same 27 s, since the network dominates. On a second visit with the cache warm, session creation alone was 1.9 s. The lesson is that streaming buys survival, not speed — and caching buys the speed.
Frequently Asked Questions
Should I split the model into several files and fetch them in parallel? Usually yes for very large models: parallel connections use available bandwidth better, and each piece can be cached and retried independently. Keep the pieces to a handful; dozens of requests cost more in overhead than they gain.
Can I load weights directly from the Origin Private File System? Yes, and it is a good fit for very large models the user has downloaded once — see persisting a Wasm database to OPFS for the access patterns. Reading from OPFS avoids the Cache Storage size limits on some platforms.
Does compressing the model file help? Only if it is not already compressed. Quantized weights have high entropy and compress by a few percent at best, so gzip or Brotli on a model file usually costs CPU for nothing. Test before enabling it.
How do I handle a user who navigates away mid-download?
Abort the fetch with an AbortController when the view unmounts, and do not cache a partial response —
cache.put on an aborted response stores a truncated body that will fail to parse on the next visit
with an error pointing at the model rather than at the cache. Delete the entry on any failure path.
Is it worth showing an estimated time remaining? Yes for anything over about ten seconds, and it is cheap: track bytes per second over a rolling window and divide the remainder by it. Users tolerate a long wait they can see the end of far better than a spinner, and the same number tells you when to offer a lower-quality model instead.
Related
- Quantizing models for Wasm inference — the cheapest way to make this problem smaller.
- Tracking linear memory growth over time — watching the arena in production.
- Reducing Wasm cold-start latency — the same problem for code rather than weights.