Async & Event-Loop Integration
WebAssembly is synchronous. A call into a module runs to completion and returns; there is no instruction to suspend, no way to await a promise, and no mechanism for the module to yield control back to the event loop and resume later. The browser, meanwhile, is asynchronous in almost everything that matters — fetching, reading files, decoding media, talking to a worker. This area is about the seam between those two facts: how a module reaches asynchronous capabilities, and how a long computation coexists with an interface that must stay responsive.
Prerequisites
- [ ] A module that needs data from an asynchronous source, or that runs long enough to be noticed.
- [ ] Familiarity with the browser’s task and microtask queues.
- [ ] A worker, for most of the approaches here.
- [ ] A frame budget you are willing to state: 16 ms, or whatever your interface needs.
The constraint, precisely
A WebAssembly function call is an ordinary synchronous call from JavaScript’s point of view. While it runs, nothing else on that thread runs: no rendering, no event handling, no promise callbacks. When it returns, the thread continues.
That has two consequences. A module cannot obtain a value that is only available asynchronously — the
result of a fetch, a file read, a message from another worker — during a call. And a call that takes
longer than a frame blocks the interface for its whole duration.
Everything in this area is one of four responses to those two facts.
Where the time actually goes
Before choosing a strategy it is worth knowing which part of the work is blocking, because the answer often points at a different fix than expected.
A call that takes 400 ms might be 380 ms of computation inside the module, or 380 ms of marshalling on the JavaScript side, or 300 ms of waiting for data that was fetched synchronously. The three have completely different remedies, and a single end-to-end measurement cannot tell them apart.
const t0 = performance.now();
const bytes = await fetchInput(url); // network
const t1 = performance.now();
const ptr = intoWasm(mod, bytes); // marshalling
const t2 = performance.now();
const out = mod.exports.process(ptr, bytes.length); // compute
const t3 = performance.now();
const result = outOfWasm(mod, out); // marshalling back
const t4 = performance.now();
report({ fetchMs: t1 - t0, inMs: t2 - t1, computeMs: t3 - t2, outMs: t4 - t3 });
Four numbers, and the largest one decides the approach. Dominated by compute means chunking or a worker. Dominated by marshalling means a flatter interface or fewer crossings, which neither chunking nor a worker improves. Dominated by the fetch means the module was never the problem.
Teams regularly move a module to a worker and find the interface no better, because the blocking was in the marshalling code that stayed on the main thread. Measuring first avoids that afternoon.
Fetch first, then call
The simplest arrangement, and the one most modules should use: the host performs every asynchronous operation and the module is a pure function over the results.
const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer());
const ptr = mod.exports.alloc(bytes.length);
new Uint8Array(mod.exports.memory.buffer, ptr, bytes.length).set(bytes);
const outLen = mod.exports.process(ptr, bytes.length);
Nothing inside the module is asynchronous, nothing needs a capability, and the same module runs unchanged in a worker, in Node or under a standalone runtime. That portability is a substantial and often overlooked benefit of keeping asynchrony on the host side.
Where the module needs to decide what to fetch, it can return a description of the request for the host to perform, and be called again with the result. That keeps the decision in the module and the capability in the host, which is the same division described in making outbound HTTP requests from WASI.
Microtasks, tasks and where a module call lands
A little precision about the event loop makes the rest of this area easier to reason about.
A JavaScript task runs to completion, then the microtask queue drains, then the browser may render. A module call happens inside whichever task invoked it, and it does not yield — so the microtask queue does not drain during it and rendering does not happen.
That explains several behaviours that otherwise look strange. An await before a module call yields; an
await after it does not help, because the blocking already happened. A promise resolved inside a host
import does not run its continuation until the module call returns and the microtask queue drains. And a
requestAnimationFrame callback that calls a long-running export delays the frame it was supposed to
render.
// this does not do what it looks like it does
async function run() {
await Promise.resolve(); // yields, then continues in a microtask
heavyWasmCall(); // blocks for 400 ms — the frame is gone
await Promise.resolve(); // too late for anything
}
The practical rule is that yielding must happen between pieces of work rather than around them, which is exactly what the chunked design accomplishes. Returning from the module and scheduling the next slice is what gives the browser its chance; anything that stays inside the call does not.
scheduler.yield(), where available, is a cleaner way to express “let the browser do its work and then
continue” than the setTimeout and requestAnimationFrame idioms, and it preserves task priority in a way
those do not.
Chunking, for work that is long but not asynchronous
A computation that takes two seconds does not need asynchrony; it needs to stop occasionally. A module exporting a “do some work and tell me whether you are finished” entry point lets the host drive it across frames.
#[no_mangle]
pub extern "C" fn step(budget_us: u32) -> i32 {
let deadline = now_us() + budget_us as u64;
while now_us() < deadline {
if !advance_one_unit() { return 1; } // finished
}
0 // more to do
}
function pump() {
if (mod.exports.step(8000) === 1) { onDone(); return; } // 8 ms of work
updateProgress(mod.exports.progress());
requestAnimationFrame(pump);
}
The module keeps its own state between calls, which it was going to do anyway, and the interface gets a frame between each slice for rendering and input. For work on the main thread this is the technique that makes a long computation tolerable without a worker.
Workers, the general answer
For anything that runs long enough to matter and does not need the DOM, the worker is the right place. The module blocks the worker’s thread, which nobody is watching, and the main thread stays free.
The design work is in the messaging: coarse messages, transferred rather than cloned payloads, and results in a shape the interface can use directly. A worker that returns a megabyte of raw floats sixty times a second has moved the computation and kept the cost.
Workers also give the only reliable cancellation mechanism a browser offers. There is no way to interrupt a running module, so a deadline means terminating the worker — which is abrupt, complete and the reason a per-execution worker is often worth its few milliseconds.
Asynchrony inside a worker
Moving the module to a worker changes what asynchrony means there, and two details are worth knowing because they differ from the main thread.
Blocking is permitted. Atomics.wait works, a long synchronous call is fine, and a module that runs for a
second harms nothing — the worker has no rendering to delay. That is the whole point of the arrangement,
and it means a worker-hosted module usually does not need chunking at all.
But the worker still has an event loop, and it still cannot receive messages while a call is running. A cancellation message sent while the module is computing sits in the queue until the call returns, which is why terminating the worker is the only cancellation that works mid-call.
// in the worker: this check only runs between calls, never during one
self.onmessage = ({ data }) => {
if (data.type === 'cancel') { cancelled = true; return; } // queued until the current call ends
runLongCall(data.payload);
};
If cancellation needs to reach a running call, the module has to poll a flag in shared memory rather than
rely on messages — which requires SharedArrayBuffer and therefore cross-origin isolation, and is worth
the complexity only when termination is genuinely unacceptable.
// with shared memory, a flag the module can see mid-call
const control = new Int32Array(new SharedArrayBuffer(4));
control[0] = 0; // main thread sets 1 to request a stop
The module reads that flag at a convenient point in its loop and returns early, which gives clean cancellation without destroying the worker — at the cost of the isolation requirement and a module that must be written to check.
Stack switching, for code that blocks by design
Ported code frequently assumes it can block: a read() in a loop, a synchronous network call, a sleep.
Restructuring it is sometimes impractical, and two mechanisms address that.
Emscripten’s Asyncify rewrites the module so it can unwind and rewind its own stack, allowing
emscripten_sleep and friends to yield to the event loop. It works everywhere and costs 30–100% in binary
size plus runtime overhead on every function that can yield.
The JavaScript Promise Integration proposal does the same thing in the engine, without the rewriting — a module can call an import that returns a promise and be suspended until it settles. It is far cheaper and its support is narrower, which is the usual trade for a recent proposal.
Both are covered in calling async JavaScript with JSPI.
Progress, cancellation and the parts users notice
Whichever strategy you choose, two interface affordances turn a long operation from an ordeal into something acceptable, and both need support from the module.
Progress requires the module to expose how far it has got. For a chunked design that is natural — the host is already calling repeatedly and can ask. For a worker it means the module posting periodic updates, which should be throttled: a progress message per item is a message storm, and once every fifty milliseconds is plenty for a human.
#[no_mangle]
pub extern "C" fn progress_permille() -> u32 {
unsafe { (DONE * 1000 / TOTAL.max(1)) as u32 }
}
Cancellation requires either a chunked design — where the host simply stops calling — or a worker that can be terminated. There is no third option in a browser: a module in the middle of a long synchronous call cannot be interrupted, and a flag it checks only helps if it reaches a check.
let cancelled = false;
cancelButton.onclick = () => { cancelled = true; worker.terminate(); recycleWorker(); };
function pump() {
if (cancelled) return;
if (mod.exports.step(8000) === 1) return onDone();
setProgress(mod.exports.progress_permille() / 10);
requestAnimationFrame(pump);
}
Designing the chunk size around the frame budget rather than around the work’s natural units is what makes both affordances smooth. Eight milliseconds of work per frame leaves room for rendering and gives the progress indicator something to show at a rate people read comfortably.
Gotchas and failure modes
- Blocking the main thread “just for a moment”. Sixteen milliseconds is a dropped frame; a hundred is visible.
Atomics.waiton the main thread. Throws; it is only permitted on a worker.- A chunked loop with no deadline. A step that runs until it finishes is not chunked.
- Per-item messages to a worker. The messaging costs more than the work.
- Assuming a promise from a module call. Exports are synchronous unless a binding layer wrapped them.
- Asyncify applied to the whole module. Restrict it with
ASYNCIFY_ONLYor pay for it everywhere.
Verifying responsiveness
The measurement that matters is not how long the work takes but how long the thread was unavailable. Long-task observation reports exactly that.
new PerformanceObserver((list) => {
for (const entry of list.getEntries()) {
report({ metric: 'long-task', ms: entry.duration, attribution: entry.name });
}
}).observe({ entryTypes: ['longtask'] });
A long task is anything over 50 ms, and a module call that produces them is blocking the interface whatever the throughput numbers say. Tracking the count and duration of long tasks alongside the work’s own timing is what distinguishes “fast” from “responsive”, and they are not the same property.
Guides in this topic
- Awaiting JavaScript promises from Rust —
wasm-bindgen-futuresand how it actually works. - Calling async JavaScript with JSPI — suspending a module without rewriting it.
- Keeping the UI responsive during long Wasm tasks — chunking, workers and measuring the result.
- Streaming data into Wasm with ReadableStream — feeding a module without buffering everything.
Frequently Asked Questions
What is the largest call I can make on the main thread? Treat 8 ms as the working limit and 16 ms as the point at which a frame is at risk. Anything that cannot fit inside that budget on the slowest device you support belongs in a worker or behind a chunked entry point, and measuring on that device rather than on yours is the part that gets skipped.
Can a WebAssembly function return a promise?
Not on its own. A binding layer can wrap an export in a JavaScript function that returns one, and
wasm-bindgen’s async fn support does exactly that — but the underlying call is still synchronous.
Should the module know it is being driven asynchronously?
Ideally not. A module exposing init, step and finish can be driven from a frame loop, from a worker
loop, or from a server-side loop with no knowledge of which — and that indifference is what keeps one
module usable in every environment rather than tied to the browser it was first written for.
Does a worker make the computation faster? No. It makes the main thread available. Speed comes from threads inside the module, from SIMD or from a better algorithm; the worker is about responsiveness.
Why is Atomics.wait forbidden on the main thread?
Because blocking the main thread indefinitely would freeze the browser’s own work. On a worker it is
permitted and is how a threaded module’s synchronisation works.
Is Asyncify still worth using? For a port that must ship now and cannot be restructured, yes. For new code, the chunked or worker approaches are cheaper, and JSPI is the direction the platform is going.
How do I decide between chunking and a worker? Chunk when the work belongs to the interface — a computation whose progress the user is watching and whose result they need on the current screen. Use a worker when the work is independent of what the user is doing next, when it may be cancelled, or when it is long enough that a chunked version would take many seconds of degraded interactivity.
Do these techniques compose? They do, and a large application usually uses several: the host fetches, a worker holds the module, the worker’s own loop chunks so it can check for cancellation, and progress flows back on a throttled channel. Each addresses a different problem, so combining them is normal rather than redundant.
Related
- SharedArrayBuffer, Atomics & Threading — parallelism, as distinct from asynchrony.
- Errors & Traps Across the Boundary — failures that cross an asynchronous seam.
- Compiling Wasm in a worker to free the main thread — the same argument for the compile step.
The seam between a synchronous module and an asynchronous host is not difficult once the constraint is explicit — it is only difficult when a design assumes the module can wait, which it never can.
← Back to JS/Wasm Interop & Memory Management