Async & Event-Loop Integration

WebAssembly is synchronous. A call into a module runs to completion and returns; there is no instruction to suspend, no way to await a promise, and no mechanism for the module to yield control back to the event loop and resume later. The browser, meanwhile, is asynchronous in almost everything that matters — fetching, reading files, decoding media, talking to a worker. This area is about the seam between those two facts: how a module reaches asynchronous capabilities, and how a long computation coexists with an interface that must stay responsive.

Prerequisites

  • [ ] A module that needs data from an asynchronous source, or that runs long enough to be noticed.
  • [ ] Familiarity with the browser’s task and microtask queues.
  • [ ] A worker, for most of the approaches here.
  • [ ] A frame budget you are willing to state: 16 ms, or whatever your interface needs.

The constraint, precisely

A WebAssembly function call is an ordinary synchronous call from JavaScript’s point of view. While it runs, nothing else on that thread runs: no rendering, no event handling, no promise callbacks. When it returns, the thread continues.

That has two consequences. A module cannot obtain a value that is only available asynchronously — the result of a fetch, a file read, a message from another worker — during a call. And a call that takes longer than a frame blocks the interface for its whole duration.

Everything in this area is one of four responses to those two facts.

Four ways across the seam The host can fetch first and call with the data, split the work into chunks between frames, move the module to a worker, or use stack switching where available to suspend the module itself. host fetches first await, then call module stays pure simplest, usually right chunk the work a slice per frame module keeps state no worker needed move to a worker blocking is fine there messages in and out the default for long work stack switching the module suspends newest, narrowest support for ported blocking code The first is nearly always the right starting point: a module that takes data and returns data needs none of the others. The rest exist for work that cannot be expressed that way — long computations, streams, and ported code that blocks by design.

Where the time actually goes

Before choosing a strategy it is worth knowing which part of the work is blocking, because the answer often points at a different fix than expected.

A call that takes 400 ms might be 380 ms of computation inside the module, or 380 ms of marshalling on the JavaScript side, or 300 ms of waiting for data that was fetched synchronously. The three have completely different remedies, and a single end-to-end measurement cannot tell them apart.

const t0 = performance.now();
const bytes = await fetchInput(url);                    // network
const t1 = performance.now();
const ptr = intoWasm(mod, bytes);                       // marshalling
const t2 = performance.now();
const out = mod.exports.process(ptr, bytes.length);     // compute
const t3 = performance.now();
const result = outOfWasm(mod, out);                     // marshalling back
const t4 = performance.now();

report({ fetchMs: t1 - t0, inMs: t2 - t1, computeMs: t3 - t2, outMs: t4 - t3 });

Four numbers, and the largest one decides the approach. Dominated by compute means chunking or a worker. Dominated by marshalling means a flatter interface or fewer crossings, which neither chunking nor a worker improves. Dominated by the fetch means the module was never the problem.

Teams regularly move a module to a worker and find the interface no better, because the blocking was in the marshalling code that stayed on the main thread. Measuring first avoids that afternoon.

Fetch first, then call

The simplest arrangement, and the one most modules should use: the host performs every asynchronous operation and the module is a pure function over the results.

const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer());
const ptr = mod.exports.alloc(bytes.length);
new Uint8Array(mod.exports.memory.buffer, ptr, bytes.length).set(bytes);
const outLen = mod.exports.process(ptr, bytes.length);

Nothing inside the module is asynchronous, nothing needs a capability, and the same module runs unchanged in a worker, in Node or under a standalone runtime. That portability is a substantial and often overlooked benefit of keeping asynchrony on the host side.

Where the module needs to decide what to fetch, it can return a description of the request for the host to perform, and be called again with the result. That keeps the decision in the module and the capability in the host, which is the same division described in making outbound HTTP requests from WASI.

Microtasks, tasks and where a module call lands

A little precision about the event loop makes the rest of this area easier to reason about.

A JavaScript task runs to completion, then the microtask queue drains, then the browser may render. A module call happens inside whichever task invoked it, and it does not yield — so the microtask queue does not drain during it and rendering does not happen.

That explains several behaviours that otherwise look strange. An await before a module call yields; an await after it does not help, because the blocking already happened. A promise resolved inside a host import does not run its continuation until the module call returns and the microtask queue drains. And a requestAnimationFrame callback that calls a long-running export delays the frame it was supposed to render.

// this does not do what it looks like it does
async function run() {
  await Promise.resolve();          // yields, then continues in a microtask
  heavyWasmCall();                  // blocks for 400 ms — the frame is gone
  await Promise.resolve();          // too late for anything
}

The practical rule is that yielding must happen between pieces of work rather than around them, which is exactly what the chunked design accomplishes. Returning from the module and scheduling the next slice is what gives the browser its chance; anything that stays inside the call does not.

scheduler.yield(), where available, is a cleaner way to express “let the browser do its work and then continue” than the setTimeout and requestAnimationFrame idioms, and it preserves task priority in a way those do not.

Chunking, for work that is long but not asynchronous

A computation that takes two seconds does not need asynchrony; it needs to stop occasionally. A module exporting a “do some work and tell me whether you are finished” entry point lets the host drive it across frames.

#[no_mangle]
pub extern "C" fn step(budget_us: u32) -> i32 {
    let deadline = now_us() + budget_us as u64;
    while now_us() < deadline {
        if !advance_one_unit() { return 1; }        // finished
    }
    0                                                // more to do
}
function pump() {
  if (mod.exports.step(8000) === 1) { onDone(); return; }   // 8 ms of work
  updateProgress(mod.exports.progress());
  requestAnimationFrame(pump);
}

The module keeps its own state between calls, which it was going to do anyway, and the interface gets a frame between each slice for rendering and input. For work on the main thread this is the technique that makes a long computation tolerable without a worker.

Workers, the general answer

For anything that runs long enough to matter and does not need the DOM, the worker is the right place. The module blocks the worker’s thread, which nobody is watching, and the main thread stays free.

The design work is in the messaging: coarse messages, transferred rather than cloned payloads, and results in a shape the interface can use directly. A worker that returns a megabyte of raw floats sixty times a second has moved the computation and kept the cost.

Workers also give the only reliable cancellation mechanism a browser offers. There is no way to interrupt a running module, so a deadline means terminating the worker — which is abrupt, complete and the reason a per-execution worker is often worth its few milliseconds.

Asynchrony inside a worker

Moving the module to a worker changes what asynchrony means there, and two details are worth knowing because they differ from the main thread.

Blocking is permitted. Atomics.wait works, a long synchronous call is fine, and a module that runs for a second harms nothing — the worker has no rendering to delay. That is the whole point of the arrangement, and it means a worker-hosted module usually does not need chunking at all.

But the worker still has an event loop, and it still cannot receive messages while a call is running. A cancellation message sent while the module is computing sits in the queue until the call returns, which is why terminating the worker is the only cancellation that works mid-call.

// in the worker: this check only runs between calls, never during one
self.onmessage = ({ data }) => {
  if (data.type === 'cancel') { cancelled = true; return; }   // queued until the current call ends
  runLongCall(data.payload);
};

If cancellation needs to reach a running call, the module has to poll a flag in shared memory rather than rely on messages — which requires SharedArrayBuffer and therefore cross-origin isolation, and is worth the complexity only when termination is genuinely unacceptable.

// with shared memory, a flag the module can see mid-call
const control = new Int32Array(new SharedArrayBuffer(4));
control[0] = 0;                     // main thread sets 1 to request a stop

The module reads that flag at a convenient point in its loop and returns early, which gives clean cancellation without destroying the worker — at the cost of the isolation requirement and a module that must be written to check.

Stack switching, for code that blocks by design

Ported code frequently assumes it can block: a read() in a loop, a synchronous network call, a sleep. Restructuring it is sometimes impractical, and two mechanisms address that.

Emscripten’s Asyncify rewrites the module so it can unwind and rewind its own stack, allowing emscripten_sleep and friends to yield to the event loop. It works everywhere and costs 30–100% in binary size plus runtime overhead on every function that can yield.

The JavaScript Promise Integration proposal does the same thing in the engine, without the rewriting — a module can call an import that returns a promise and be suspended until it settles. It is far cheaper and its support is narrower, which is the usual trade for a recent proposal.

Both are covered in calling async JavaScript with JSPI.

Progress, cancellation and the parts users notice

Whichever strategy you choose, two interface affordances turn a long operation from an ordeal into something acceptable, and both need support from the module.

Progress requires the module to expose how far it has got. For a chunked design that is natural — the host is already calling repeatedly and can ask. For a worker it means the module posting periodic updates, which should be throttled: a progress message per item is a message storm, and once every fifty milliseconds is plenty for a human.

#[no_mangle]
pub extern "C" fn progress_permille() -> u32 {
    unsafe { (DONE * 1000 / TOTAL.max(1)) as u32 }
}

Cancellation requires either a chunked design — where the host simply stops calling — or a worker that can be terminated. There is no third option in a browser: a module in the middle of a long synchronous call cannot be interrupted, and a flag it checks only helps if it reaches a check.

let cancelled = false;
cancelButton.onclick = () => { cancelled = true; worker.terminate(); recycleWorker(); };

function pump() {
  if (cancelled) return;
  if (mod.exports.step(8000) === 1) return onDone();
  setProgress(mod.exports.progress_permille() / 10);
  requestAnimationFrame(pump);
}

Designing the chunk size around the frame budget rather than around the work’s natural units is what makes both affordances smooth. Eight milliseconds of work per frame leaves room for rendering and gives the progress indicator something to show at a rate people read comfortably.

Gotchas and failure modes

  • Blocking the main thread “just for a moment”. Sixteen milliseconds is a dropped frame; a hundred is visible.
  • Atomics.wait on the main thread. Throws; it is only permitted on a worker.
  • A chunked loop with no deadline. A step that runs until it finishes is not chunked.
  • Per-item messages to a worker. The messaging costs more than the work.
  • Assuming a promise from a module call. Exports are synchronous unless a binding layer wrapped them.
  • Asyncify applied to the whole module. Restrict it with ASYNCIFY_ONLY or pay for it everywhere.

Verifying responsiveness

The measurement that matters is not how long the work takes but how long the thread was unavailable. Long-task observation reports exactly that.

new PerformanceObserver((list) => {
  for (const entry of list.getEntries()) {
    report({ metric: 'long-task', ms: entry.duration, attribution: entry.name });
  }
}).observe({ entryTypes: ['longtask'] });

A long task is anything over 50 ms, and a module call that produces them is blocking the interface whatever the throughput numbers say. Tracking the count and duration of long tasks alongside the work’s own timing is what distinguishes “fast” from “responsive”, and they are not the same property.

Guides in this topic

Four ways to be asynchronous A module is synchronous by nature. Each approach below borrows asynchrony from somewhere else: the generated glue, a stack-switching proposal, the host loop, or a separate thread. wasm-bindgen-futures the glue drives a Rust future from a promise JSPI the engine suspends the stack at an await point chunked calls the host slices the work and yields between slices worker thread the work leaves the main thread entirely The first three keep the work on the main thread and differ only in who does the slicing. The fourth is the only one that removes the work from the thread the page renders on.

Frequently Asked Questions

What is the largest call I can make on the main thread? Treat 8 ms as the working limit and 16 ms as the point at which a frame is at risk. Anything that cannot fit inside that budget on the slowest device you support belongs in a worker or behind a chunked entry point, and measuring on that device rather than on yours is the part that gets skipped.

Can a WebAssembly function return a promise? Not on its own. A binding layer can wrap an export in a JavaScript function that returns one, and wasm-bindgen’s async fn support does exactly that — but the underlying call is still synchronous.

Should the module know it is being driven asynchronously? Ideally not. A module exposing init, step and finish can be driven from a frame loop, from a worker loop, or from a server-side loop with no knowledge of which — and that indifference is what keeps one module usable in every environment rather than tied to the browser it was first written for.

Does a worker make the computation faster? No. It makes the main thread available. Speed comes from threads inside the module, from SIMD or from a better algorithm; the worker is about responsiveness.

Why is Atomics.wait forbidden on the main thread? Because blocking the main thread indefinitely would freeze the browser’s own work. On a worker it is permitted and is how a threaded module’s synchronisation works.

Is Asyncify still worth using? For a port that must ship now and cannot be restructured, yes. For new code, the chunked or worker approaches are cheaper, and JSPI is the direction the platform is going.

How do I decide between chunking and a worker? Chunk when the work belongs to the interface — a computation whose progress the user is watching and whose result they need on the current screen. Use a worker when the work is independent of what the user is doing next, when it may be cancelled, or when it is long enough that a chunked version would take many seconds of degraded interactivity.

Do these techniques compose? They do, and a large application usually uses several: the host fetches, a worker holds the module, the worker’s own loop chunks so it can check for cancellation, and progress flows back on a throttled channel. Each addresses a different problem, so combining them is normal rather than redundant.

What the user sees during the work A single blocking call holds the main thread for its whole duration. The same work sliced into chunks returns the thread between slices, so rendering and input continue. one blocking call 1,850 ms with no frame rendered and no input handled sliced into 8 ms chunks work work work work frames and input in every gap Total wall-clock time goes up slightly; the number people report as speed goes down a lot. The slice budget is a deadline, not a work count: measure and adapt it, because devices differ tenfold.

The seam between a synchronous module and an asynchronous host is not difficult once the constraint is explicit — it is only difficult when a design assumes the module can wait, which it never can.

← Back to JS/Wasm Interop & Memory Management