Memory64 and Large Heaps

This guide answers one task: decide whether your module needs a 64-bit address space, and if so, build one — while understanding that raising the ceiling does not raise what a browser will actually give you.

Prerequisites

  • [ ] A workload whose working set genuinely exceeds 4 GB.
  • [ ] A toolchain with the target: LLVM 18+, Rust’s wasm64-unknown-unknown (nightly), Emscripten with -sMEMORY64.
  • [ ] An engine with the proposal: Chrome 133+, Firefox 134+, Node 24+.
  • [ ] A measurement of your actual peak memory, not an estimate.

What the proposal changes

A WebAssembly memory is indexed by a 32-bit integer, which caps it at 4 GiB minus a page. That limit is in the instruction encoding: i32.load takes an i32 address and there is no way to express a larger one.

Memory64 adds a 64-bit variant. A memory declared with the i64 index type is addressed by i64 values, memory.size and memory.grow take and return i64, and the theoretical ceiling becomes 16 exbibytes.

(module
  (memory (;0;) i64 1024 65536)          ;; 64 MiB initial, 4 GiB maximum, 64-bit indexed

  (func (export "read") (param $addr i64) (result i32)
    (i32.load (local.get $addr))))       ;; the address is now an i64

Nothing else about the module changes. Functions, tables, globals and the type system are unaffected; only the memory’s index type and the addresses flowing into load and store instructions differ.

The ceiling, and what actually limits you A 32-bit memory caps at 4 GiB by encoding. A 64-bit memory has an enormous theoretical ceiling, but browsers impose their own much lower limits, so the practical constraint moves rather than disappearing. 32-bit memory 4 GiB — a hard encoding limit 64-bit memory 16 EiB theoretical what a browser gives typically 2–16 GiB, and often less The proposal removes the encoding limit. The tab's budget, the device's memory and the browser's policy remain exactly where they were. Which is why the first question is what your working set actually is, rather than what the ceiling is.

Building for it

Emscripten has supported the target longest and is the easiest route for C and C++.

emcc app.c -sMEMORY64=1 -sINITIAL_MEMORY=1GB -sMAXIMUM_MEMORY=8GB -O3 -o app.js

For Rust the target is wasm64-unknown-unknown and requires a nightly toolchain, since the target is not yet stable:

rustup +nightly target add wasm64-unknown-unknown
cargo +nightly build --release --target wasm64-unknown-unknown

Both produce a module whose pointers are 64-bit, which is the change with the widest consequences: every pointer in the program doubles in size, every struct containing one grows, and every pointer-heavy data structure uses more cache.

What it costs

The costs are real and worth measuring before adopting.

Pointer size doubles. A linked structure, a tree, a hash table of references — all of them grow, and the extra bytes come out of the cache. For pointer-dense workloads the observed slowdown is typically 5–20%, which is the price of the larger address space.

Bounds checks get more expensive. A 32-bit memory can be implemented with a guard region, so bounds checking is often free — the hardware fault handles an out-of-range access. A 64-bit memory cannot rely on that trick everywhere, so engines may insert explicit checks.

Toolchain maturity is lower. Fewer libraries have been tested against a 64-bit WebAssembly target, and the ones making assumptions about pointer width will surface them here.

same workload, same machine
  wasm32, 3.2 GB working set   fails to allocate
  wasm64, 3.2 GB working set   runs, 14% slower than the same code natively on 64-bit

Whether you need it

The honest answer for most modules is no, and the check takes a minute.

Measure peak memory for your realistic worst case. If it is under about 3 GiB, a 32-bit module handles it and the proposal buys nothing. If it is above that, the next question is whether the working set can be bounded instead — streaming, chunking, spilling to storage — because those techniques keep the module portable to every engine and often make it faster as well as smaller.

The workloads that genuinely need the larger space are ones where the whole dataset must be addressable at once: an in-memory analytical engine over tens of gigabytes, a large model’s weights held resident, a scientific simulation whose grid does not decompose. For those, the alternative is not a clever restructuring but a server.

Three answers, depending on the working set Under about three gigabytes, a 32-bit module is fine. Above it, bounding the working set through chunking or spilling is usually better. Memory64 is for workloads whose entire dataset must be addressable at once. under ~3 GiB a 32-bit module is fine runs everywhere no decision to make larger, but decomposable chunk, stream or spill bounded working set usually the better answer must be resident random access over everything no useful decomposition memory64, or a server The middle column covers far more real workloads than teams expect, and the techniques there improve performance as well as footprint.

Measuring your real peak

Everything above turns on one number, and estimating it is how teams end up adopting a proposal they did not need or discovering a limit in production. Measure it.

In a browser, performance.measureUserAgentSpecificMemory() gives an accurate figure where it is available and requires cross-origin isolation. Where it is not, the module’s own memory size is a good proxy for the part you control:

const pages = instance.exports.memory.buffer.byteLength / 65536;
report({ wasmMemoryMB: (pages * 65536) / 1e6 });

Under a standalone runtime the measurement is easier and more precise:

/usr/bin/time -v wasmtime run --dir=. engine.wasm < fixtures/large.bin 2>&1 | grep Maximum
# Maximum resident set size (kbytes): 3412880

Take the measurement on your largest realistic input rather than a typical one, and take it repeatedly — peak memory for a workload that allocates in bursts is not a stable number, and the maximum across a hundred runs is the figure that matters.

Then compare against what the deployment will actually allow. A desktop browser tab commonly tolerates 2–4 GiB, a phone considerably less, and an edge runtime typically 128 MiB. Those three numbers frequently make the decision before the proposal’s ceiling enters the conversation at all.

Expected output

A 64-bit module declares its memory with the i64 index type, which wasm-objdump reports:

wasm-objdump -x dist/app.wasm | grep -A1 'Memory\['
# Memory[1]:
#  - memory[0] i64 pages: initial=16384 max=131072
# and at runtime, the allocation either succeeds or fails cleanly
node --experimental-wasm-memory64 run.mjs
memory: 1.00 GiB initial, 8.00 GiB maximum
allocated working set: 3.24 GiB
RangeError: WebAssembly.Memory(): could not allocate memory     ← on a smaller device

That second outcome is the one to design for. A module that assumes its maximum will be granted fails on exactly the devices that most need a graceful answer.

Handling the allocation that fails

Because the ceiling is now the host’s policy rather than the encoding, a module using memory64 must treat allocation failure as an ordinary path.

memory.grow returns −1 when it cannot grow, and a well-written module checks. Code that assumes success — which most allocators do — turns a recoverable condition into a trap, and the difference is visible to the user as a crash rather than a message.

fn reserve(pages: u64) -> Result<u64, OutOfMemory> {
    let prev = core::arch::wasm64::memory_grow(0, pages as usize);
    if prev == usize::MAX { Err(OutOfMemory) } else { Ok(prev as u64) }
}

On the host side, instantiate with a maximum that reflects what you are willing to give rather than the largest number the proposal allows, and catch the failure at instantiation:

let memory;
try {
  memory = new WebAssembly.Memory({ initial: 16384n, maximum: 131072n, index: 'i64' });
} catch (e) {
  report({ event: 'memory64-unavailable', error: String(e) });
  return loadSmallerBuild();
}

Designing the feature so that a smaller allocation still produces a useful result — a lower-resolution analysis, a subset of the dataset, a slower streaming path — is what turns “this device cannot run it” into “this device runs the reduced version”. That fallback is worth more than the extra address space on most of the devices that will encounter the limit.

What the wider index buys and costs A 64-bit memory removes the four-gigabyte ceiling at the cost of wider pointers, which makes every pointer-heavy structure larger and colder in cache. i32 memory 4 GB ceiling · 4-byte pointers · the universally supported default i64 memory addresses far beyond 4 GB 8-byte pointers, larger structures Move to a 64-bit memory because you need the addresses, not because it sounds like a faster target. Bounds checks stay; the index type widens, and the engine still refuses an access outside the memory.

Gotchas

  • Assuming the ceiling is the limit. The browser’s budget is far lower than the proposal’s ceiling.
  • Pointer-width assumptions in dependencies. Code casting pointers to u32 breaks; the failure is usually a truncated address.
  • Adopting it for headroom. The pointer-size cost is paid always, and the benefit only when you exceed 4 GiB.
  • No handling for a failed allocation. On a phone, the failure is the normal case.
  • Assuming support. Recent and not universal; detect and provide a 32-bit path.
  • Benchmarking only the 64-bit build. Compare against the 32-bit one on the same workload, or the regression is invisible.

Performance note

For a graph workload with a 2.8 GiB working set that fit in both, the 64-bit build ran 11% slower than the 32-bit one and used 19% more memory, entirely from pointer size. At 3.6 GiB the 32-bit build could not run at all. That is the shape of the decision: below the ceiling, 32-bit wins on every axis; above it, 64-bit is the only option that works.

Frequently Asked Questions

Does this help a server-side module? More than a browser one, because a standalone runtime can grant far more memory than a tab. For a WASI service processing large datasets, memory64 removes a real constraint.

Can a module have both a 32-bit and a 64-bit memory? With the multi-memory proposal, yes — which is an appealing combination for a module that wants a large data region and a small, fast one. Support for the combination is newer than either proposal alone.

Will engines raise their limits over time? Probably, gradually, and driven by device memory rather than by the specification. Designing for a bounded working set remains the more durable approach.

← Back to Post-MVP Wasm Proposals in Practice