Graphics, Games & Simulation
Real-time work has a property no other WebAssembly workload shares: a deadline that arrives sixty times
a second and never moves. Everything in this topic follows from that. The compiled module is usually
doing the simulation — physics, particles, pathfinding, animation — while the GPU does the drawing, and
the interesting engineering is in the seam: how state gets from linear memory into draw calls without a
copy, and how the loop stays inside its budget on a machine you do not own.
Prerequisites
- [ ] A toolchain that can produce a web build:
emccfor C and C++ engines,wasm-packfor Rust. - [ ] A canvas with a WebGL 2 or WebGPU context.
- [ ] Familiarity with the browser frame loop —
requestAnimationFrame, and whysetIntervalis wrong. - [ ] A test machine that is not your development machine. Frame budgets are hardware opinions.
The frame budget is the whole design
At 60 Hz you have 16.67 ms per frame for everything: input handling, simulation, the JavaScript that issues draw calls, the browser’s own compositing, and whatever else the page is doing. A comfortable target for your own work is under 10 ms, because the rest is not yours.
That budget decides the architecture. Simulation state lives in linear memory and stays there. Draw
calls read directly from typed-array views over that memory. Nothing is serialised, nothing is converted
to JavaScript objects, and nothing allocates during a frame — an allocation is a garbage collection
waiting to land in the middle of one.
Two architectures, and the one you should prefer
There are two ways to structure a WebAssembly-driven renderer, and the choice shapes everything else.
In the module-drives-everything model, used by Emscripten ports of existing engines, the compiled code
calls OpenGL ES functions that Emscripten translates into WebGL calls. The engine’s original render loop
runs unchanged, driven by emscripten_set_main_loop. This is the fastest route for an existing codebase
and the reason so many native games run in browsers at all.
In the module-simulates, JavaScript-draws model, the compiled code owns state and physics, and a thin JavaScript layer reads that state through typed-array views and issues draw calls itself. It requires writing the renderer, but it composes with the rest of a web application, gives direct access to WebGPU, and keeps the module small.
For a port, take the first. For a new feature inside an existing web application, take the second — it is much easier to integrate, and the performance difference is small because the expensive part was never the draw call submission.
Getting vertex data to the GPU without copying
The critical path is the same in both models: positions, colours and indices live in linear memory, and
the GPU needs them. WebGL’s buffer upload functions accept a typed-array view, and a view over
linear memory is a legitimate argument — so the upload reads straight from the module’s memory with no
intermediate array.
const positions = new Float32Array(memory.buffer, mod.exports.positions_ptr(), count * 3);
gl.bindBuffer(gl.ARRAY_BUFFER, vbo);
gl.bufferSubData(gl.ARRAY_BUFFER, 0, positions); // reads directly from linear memory
Two rules keep this correct. Rebuild the view after anything that might have grown memory, or use a
module that never grows after startup — the second is strongly preferable in a frame loop. And use
bufferSubData into a preallocated buffer rather than bufferData, which reallocates GPU storage and
stalls the pipeline.
Fixed timestep, variable rendering
Simulation and rendering should not share a clock. A physics step that uses whatever time elapsed since the last frame is non-deterministic, behaves differently on a 144 Hz monitor, and explodes when a frame takes 300 ms because the user switched tabs.
The standard structure accumulates real time and consumes it in fixed steps:
const STEP = 1 / 120; // simulate at 120 Hz regardless of display
let acc = 0, last = performance.now();
function frame(now) {
acc += Math.min((now - last) / 1000, 0.25); // clamp: never simulate more than 0.25 s
last = now;
while (acc >= STEP) { mod.exports.step(STEP); acc -= STEP; }
render(acc / STEP); // interpolation factor for smooth motion
requestAnimationFrame(frame);
}
The clamp matters more than it looks. Without it, a tab restored after two minutes in the background tries to simulate twelve thousand steps in one frame, and the page hangs. With it, the simulation simply loses that time, which is the correct behaviour for everything except a deterministic replay.
Memory layout decides throughput
For simulation, the arrangement of data in linear memory matters more than the arithmetic performed on
it. A particle system storing an array of structs — position, velocity, colour and lifetime interleaved
per particle — touches every field on every pass even when a pass only needs positions, and it defeats
vectorisation because the values a SIMD lane wants are not adjacent.
The alternative, a struct of arrays, stores all x coordinates together, then all y, then all velocities.
An integration pass reads two contiguous streams and writes one, which is the ideal shape for both the
cache and v128 instructions:
pub struct Particles {
x: Vec<f32>, y: Vec<f32>,
vx: Vec<f32>, vy: Vec<f32>,
life: Vec<f32>,
}
pub fn integrate(p: &mut Particles, dt: f32) {
for i in 0..p.x.len() {
p.x[i] += p.vx[i] * dt; // two streams in, one out — vectorises cleanly
p.y[i] += p.vy[i] * dt;
}
}
The same layout also makes the GPU upload trivial: a contiguous array of positions is exactly what
bufferSubData wants, with no gather step in between. Converting an existing array-of-structs engine is
invasive, which is an argument for choosing the layout before there is much code rather than after.
Capacity planning follows the same logic. Decide the maximum entity count at startup, allocate for it, and treat exceeding it as a design error rather than a reason to reallocate mid-frame. A pool with a free list gives you the dynamism without the allocation, and it keeps indices stable — which matters, because in a struct-of-arrays layout an index is the identity.
Input, and why it is harder than it looks
Input arrives as DOM events on the main thread, and the simulation wants it as state at a specific simulation step. Bridging that means buffering events and applying them at step boundaries rather than whenever they fire.
The simplest correct approach is a small input state block inside linear memory that JavaScript writes
and the module reads at the start of each step. Key states, pointer position and button flags fit in a
few dozen bytes, and updating them from event handlers costs nothing. Avoid calling into the module from
an event handler directly — the event may arrive mid-frame, and a simulation that mutates during a step
produces inconsistencies that are extremely hard to reproduce.
Pointer lock, gamepad polling and touch all layer on top of the same structure. Gamepads are polled rather than evented, so read them once per frame and write into the same block.
Running the simulation in a worker
Moving the module to a worker is attractive — the main thread then only issues draw calls — but WebGL
contexts are not transferable, so the renderer must either stay on the main thread or use
OffscreenCanvas to move the context into the worker with the module.
The second option is the better one when available: simulation and rendering both live in the worker, and
the main thread handles only input and interface. With shared memory the two can also be split, with the
worker simulating into a SharedArrayBuffer that the main thread reads for drawing — at the cost of
requiring cross-origin isolation and careful synchronisation, since a half-updated state read mid-frame
produces visible tearing.
Porting an existing engine
Engine ports are a well-trodden path, and the surprises are consistent. The main loop must be restructured
— a while (running) loop cannot exist in a browser, because it never yields, so Emscripten’s main-loop
machinery takes over the driving. File access becomes a virtual file system that needs preloading or
fetching. Threads need cross-origin isolation. And the asset bundle, not the code, is usually what makes
the download unacceptable.
The pages under this topic cover each: the loop restructuring in porting a C game loop to Emscripten, and the GPU paths in the WebGL and WebGPU guides.
Assets are usually the real problem
Teams porting a game to the browser tend to spend their effort on the compiled code and then discover that the code was never the issue. A modest native game ships hundreds of megabytes of textures, meshes and audio, and none of that gets smaller by being compiled to WebAssembly.
Three things help, in order of impact. Compress textures in a GPU-native format — Basis Universal transcodes to whatever the device supports, so one asset serves every platform at a fraction of the size of PNG. Stream assets by need rather than bundling them: the first level, the first area, the first few seconds of audio, with the rest fetched while the player is occupied. And separate code from content entirely so that a code update does not invalidate a two-hundred-megabyte asset cache.
The loading experience is part of this. A browser game that shows a progress bar reaching 100% and then sits for eight seconds instantiating is worse than one that reports each phase honestly. Report the fetch, the decode and the instantiation separately, and start the audio context on the first user gesture rather than at load, since browsers will refuse it otherwise.
Measure the total: transfer size, time to first interactive frame, and memory resident once the first scene is running. Those three numbers decide whether anyone stays long enough to see the frame rate you worked on.
Gotchas and failure modes
- Allocating in the frame loop. The single most common cause of irregular stutter. Preallocate.
- Rebuilding typed-array views every frame. Cheap but not free, and unnecessary if memory never grows.
bufferDatainstead ofbufferSubData. Reallocates GPU storage and stalls.- Simulation tied to frame rate. Different behaviour on different displays, and an explosion after a long pause.
- Reading GPU state back per frame.
readPixelsandgetParameterforce a synchronisation point and can cost more than everything else combined. - Assuming 60 Hz. High-refresh displays exist, and so do power-saving modes that drop to 30.
Determinism, replays and multiplayer
Any feature that replays a session, synchronises two clients, or validates a score depends on the simulation producing identical results from identical inputs. WebAssembly helps here more than most platforms: integer arithmetic is exactly specified, and floating-point follows IEEE 754 with defined rounding, so the same module fed the same inputs produces the same bits on every engine.
The caveats are the ones you introduce. Threads change reduction order and therefore floating-point
results, so a deterministic simulation must be single-threaded or must fix its reduction order
explicitly. Anything seeded from Math.random, the wall clock or performance.now is non-deterministic
by construction — seed from a value that is part of the recorded input instead. And iteration over a hash
map whose order depends on pointer values will differ between runs; use a deterministic container in the
simulation path.
Getting this right buys more than multiplayer. A recorded input stream plus a deterministic simulation is the best debugging tool available for a real-time system: a bug report becomes a file that reproduces the failure exactly, on your machine, as many times as you need. Teams that build this early tend to keep it forever, and teams that skip it spend their evenings trying to reproduce something a player saw once.
Verifying frame time honestly
Frame rate is a lagging indicator; frame time distribution is what users perceive. A steady 58 fps with occasional 40 ms frames feels worse than a steady 50.
const times = [];
function frame(now) {
times.push(now - last); last = now;
if (times.length === 600) {
times.sort((a, b) => a - b);
console.log({ p50: times[300], p95: times[570], p99: times[594], worst: times[599] });
times.length = 0;
}
requestAnimationFrame(frame);
}
Report p95 and p99 alongside the median, and instrument the simulation separately from the frame so you know which half is responsible. The browser’s own performance panel is the right tool for finding where the time goes, but these numbers are what you track over time.
Guides in this topic
- Rendering with WebGL from a Wasm module — buffers, views and draw calls.
- Driving WebGPU from Rust Wasm — the modern API from compiled code.
- Porting a C game loop to Emscripten — restructuring a loop that never returns.
- Running a physics engine in Wasm — fixed steps, determinism and budgets.
- Sharing a canvas framebuffer with Wasm — software rendering straight into pixels.
Frequently Asked Questions
Is WebAssembly fast enough for a real game? Yes, for the simulation. Compiled physics, animation and AI run at a large multiple of equivalent JavaScript, and the GPU does the drawing either way. What limits browser games is usually asset size and input latency rather than compute.
Should the renderer be in Wasm or JavaScript? Draw call submission is cheap in both. Keep it wherever integration is easiest — JavaScript for a feature inside a web application, the module for a port where the engine already owns rendering.
Does SIMD help? Substantially for particles, skinning, and any vector maths over arrays — 2–4× is typical. It helps physics broad-phase less than people expect, because that work is branch-heavy rather than arithmetic.
How do I keep the page responsive while the game runs?
Use OffscreenCanvas and put simulation and rendering in a worker. Failing that, keep the frame budget
honest: a game that consistently uses 14 ms of a 16.7 ms frame leaves nothing for the rest of the page.
How much does the boundary crossing cost per frame? Less than people fear. A call into the module is on the order of tens of nanoseconds, so even a few hundred calls per frame are irrelevant. What costs is crossing with data — converting arrays, building objects, or copying buffers — which is why the design keeps state on one side and passes indices and pointers rather than values.
Should audio go through the module too?
Mixing and synthesis, yes, in an AudioWorklet with the same no-allocation discipline as any other
real-time path. Triggering sounds is better done from JavaScript, because the sound system needs to know
about page state — muting when the tab is hidden, respecting the user’s volume — that the simulation
should not care about.
What frame rate should I target? Design for a variable rate and a fixed simulation step. Targeting a specific frame rate bakes in an assumption that high-refresh displays, battery-saver modes and lower-powered devices all violate, and the interpolated fixed-step loop handles every one of them without special cases.
Can I share the simulation between a browser client and a server? Yes, and it is one of the better reasons to put it in a compiled module: the same binary validates moves server-side under a standalone runtime and predicts them client-side in the tab, with identical results because the instruction semantics are identical.
Related
- Wasm SIMD & Vectorized Computation — vectorising the inner loops.
- Media Processing & Codecs in Wasm — the same buffer discipline for pixels.
- Zero-Copy Data Transfer Patterns — why the GPU can read linear memory directly.
← Back to Production Wasm: Workloads & Deployment