Keeping the UI Responsive During Long Wasm Tasks
This guide answers one task: run a WebAssembly computation that takes seconds without the page becoming unresponsive — with progress the user can see and a cancel button that works.
Prerequisites
- [ ] A module call that takes long enough to notice, measured rather than assumed.
- [ ] The ability to change the module’s interface, or a worker you can terminate.
- [ ] A frame budget: 16 ms, or 8 ms if the page animates.
- [ ] Long-task measurement, to confirm the result.
Measure the blocking first
The number that matters is how long the thread was unavailable, not how long the work took — and the browser reports it directly.
new PerformanceObserver((list) => {
for (const e of list.getEntries()) console.log('long task', e.duration.toFixed(0), 'ms');
}).observe({ entryTypes: ['longtask'] });
Anything over 50 ms is a long task by definition, and a single 1,200 ms entry is the signature of a module call that blocked the page. Recording these before and after the change is what turns “it feels better” into evidence.
While measuring, split the call as described in the topic overview: fetch, marshalling, compute, marshalling back. Moving compute to a worker while leaving 300 ms of marshalling on the main thread improves nothing perceptible.
A chunked entry point
The module needs to be able to stop. That means an interface of three functions — start, do some work, finish — with the module keeping its own state between calls.
static mut STATE: Option<Job> = None;
#[no_mangle]
pub extern "C" fn job_begin(ptr: *const u8, len: usize) -> i32 {
unsafe { STATE = Some(Job::new(std::slice::from_raw_parts(ptr, len))) };
0
}
/// Works for at most `budget_us`, returns 1 when finished, 0 when there is more to do.
#[no_mangle]
pub extern "C" fn job_step(budget_us: u32) -> i32 {
let deadline = now_us() + budget_us as u64;
let job = unsafe { STATE.as_mut().unwrap() };
while now_us() < deadline {
if !job.advance() { return 1; }
}
0
}
#[no_mangle]
pub extern "C" fn job_progress_permille() -> u32 {
unsafe { STATE.as_ref().map_or(0, |j| j.progress_permille()) }
}
The budget is passed in rather than hard-coded so the host can tune it — a page that animates wants a smaller slice than one that does not, and the right value differs between devices.
Checking the clock has a cost, so check it every few units of work rather than every one. A loop that calls
now_us() per item can spend more time reading the clock than working.
Choosing the slice size
The budget passed to step is the one tuning knob that matters, and the right value is a property of the
page rather than of the work.
Start from what else the frame has to do. A page that only renders a progress bar can give the module 12 ms of a 16 ms frame; a page running an animation, a video or a canvas of its own should give 4–6 ms. A budget larger than what is left after the page’s own work produces dropped frames however carefully the module behaves.
Then check the granularity. A budget is only honoured at the points where the module checks the clock, so a work unit that takes 30 ms makes an 8 ms budget meaningless — the slice will be 30 ms regardless. Where individual units are large, the unit itself has to be subdivided before chunking helps at all.
// a unit too coarse for the budget to mean anything
fn advance(&mut self) -> bool { self.process_whole_image(); false }
// subdivided, so the budget is honoured
fn advance(&mut self) -> bool {
let rows = 16.min(self.height - self.row);
self.process_rows(self.row, rows);
self.row += rows;
self.row < self.height
}
Finally, measure on the slowest device you support. A budget tuned on a laptop is several times too large on a mid-range phone, where the same eight milliseconds of work takes twenty — which converts a smooth implementation into one that drops every other frame. Passing the budget in from the host rather than compiling it into the module is what lets you set it per device class once you know.
Driving it from the frame loop
requestAnimationFrame gives one call per frame, which is exactly the cadence wanted.
export function runJob(mod, input, { onProgress, onDone }) {
const ptr = intoWasm(mod, input);
mod.exports.job_begin(ptr, input.length);
let cancelled = false;
function pump() {
if (cancelled) return;
if (mod.exports.job_step(8000) === 1) return onDone(outOfWasm(mod));
onProgress(mod.exports.job_progress_permille() / 10);
requestAnimationFrame(pump);
}
requestAnimationFrame(pump);
return () => { cancelled = true; }; // the cancel function
}
Cancellation here is trivial because the host controls the loop: stop calling, and the work stops. That is a real advantage of chunking over a worker, where cancellation means termination.
Where the page is not animating, scheduler.yield() is a better fit than requestAnimationFrame because
it resumes as soon as the browser has done its work rather than waiting for the next frame:
while (mod.exports.job_step(8000) === 0) {
onProgress(mod.exports.job_progress_permille() / 10);
await scheduler.yield(); // where available
}
Or move it to a worker
When the work does not need the DOM and may run for a long time, a worker is simpler than chunking: the module blocks a thread nobody is watching.
// main thread
const worker = new Worker('/job-worker.js', { type: 'module' });
worker.onmessage = ({ data }) => {
if (data.type === 'progress') setProgress(data.value);
if (data.type === 'done') onDone(data.result);
};
worker.postMessage({ input }, [input.buffer]);
cancelButton.onclick = () => { worker.terminate(); recreateWorker(); };
// job-worker.js — blocking is fine here
self.onmessage = async ({ data }) => {
const mod = await ready;
const ptr = intoWasm(mod, data.input);
mod.exports.job_begin(ptr, data.input.length);
while (mod.exports.job_step(50_000) === 0) { // 50 ms slices: nobody is rendering
self.postMessage({ type: 'progress', value: mod.exports.job_progress_permille() / 10 });
}
self.postMessage({ type: 'done', result: outOfWasm(mod) });
};
Note the larger slice inside the worker. The chunking there is not for responsiveness but so that progress can be reported and, in a shared-memory design, so a cancellation flag can be observed.
Expected output
The long-task observer is the proof:
# before
long task 1214 ms
frames rendered during job: 0
result in 1.21 s
# after (chunked, 8 ms budget)
long task: none over 50 ms
frames rendered during job: 78
progress updates: 78
result in 1.34 s
The total is 10% longer — the cost of stopping and restarting — and the page rendered 78 frames instead of none. That trade is almost always worth making, and quoting both numbers is what makes the decision defensible.
Gotchas
- A step function with no deadline. Not chunked; it runs to completion.
- Checking the clock per item. The clock read can dominate; check every few hundred units.
- Marshalling left on the main thread. Moving compute to a worker while leaving a 300 ms copy achieves nothing.
- Progress updates per item. A message storm; throttle to a few per second.
- State in module globals with concurrent jobs. One job at a time, or state per job.
- Cancellation that only sets a flag in a worker. The message is not read until the current call returns; terminate, or use shared memory.
Performance note
For a 1.2 second workload, chunking at 8 ms per slice added about 10% to the total — roughly 150 slices, each paying a boundary crossing and a clock read — and eliminated every long task. A 4 ms budget added 19% and rendered slightly more smoothly; 16 ms added 4% and dropped occasional frames. Eight milliseconds was the best trade on the hardware tested, and finding your own takes one afternoon of measurement.
Frequently Asked Questions
Can I use setTimeout(fn, 0) instead?
It works and is clamped to about 4 ms with nesting, which makes the cadence unpredictable.
requestAnimationFrame matches rendering and scheduler.yield matches the browser’s own scheduling; both
are better.
Should the progress indicator be determinate? Where the module can report progress, yes — a determinate bar makes a long wait feel shorter and tells the user whether anything is happening. Where it genuinely cannot, an indeterminate indicator plus a count of items completed is more honest than a bar that jumps from nothing to everything.
What if the module cannot be changed to expose a step function? Then a worker is the only option, with termination as the cancellation mechanism. That is frequently the situation with a third-party module, and it is why the worker approach is the more general one.
Does chunking hurt cache locality? Slightly — the work is interleaved with rendering, which evicts cache lines. That is part of the 10% overhead measured above, and it is the main reason larger slices are more efficient in a worker where responsiveness is not a concern.
A last point worth stating plainly: responsiveness is a property users notice and benchmarks do not. A build that finishes the whole job 8% faster but freezes the tab for six seconds will be reported as the slower one, every time. Optimise the number people experience.
Related
- Compiling Wasm in a worker to free the main thread — the same argument for the compile step.
- Streaming data into Wasm with ReadableStream — chunking the input as well as the work.
- Sharing memory between Wasm and Web Workers — the shared flag for mid-call cancellation.
← Back to Async & Event-Loop Integration