Building a Wasm Image Filter Pipeline
This guide answers one task: apply a sequence of image filters — exposure, curves, blur, sharpen, grain — inside a single WebAssembly module, with no copy between stages, no reallocation per adjustment, and a preview that updates while the user drags a slider.
Prerequisites
- [ ] A Rust or C module built for the web, with an arena rather than a general-purpose allocator.
- [ ] A canvas or
ImageBitmapdisplay path. - [ ] Familiarity with typed-array views over
linear memory. - [ ] A test image with a wide tonal range — a blown-out sky and a dark foreground reveal filter bugs that a flat test pattern hides.
One module, two buffers, many stages
The naive design gives each filter its own input and output allocation and copies between them. For a twelve-megapixel image and six filters that is 300 MB of pointless traffic. The standard fix is ping-pong buffering: allocate exactly two full-size surfaces, and alternate which one is the source and which the destination as you walk the stage list.
static mut A: Vec<u8> = Vec::new();
static mut B: Vec<u8> = Vec::new();
#[no_mangle]
pub extern "C" fn run_pipeline(w: usize, h: usize, stages_ptr: *const u8, n: usize) -> *const u8 {
unsafe {
let stages = std::slice::from_raw_parts(stages_ptr, n);
let (mut src, mut dst) = (&mut A as *mut Vec<u8>, &mut B as *mut Vec<u8>);
for &stage in stages {
apply(stage, &(*src), &mut (*dst), w, h);
std::mem::swap(&mut src, &mut dst); // no copy, just a pointer swap
}
(*src).as_ptr() // last write landed in src after the swap
}
}
Two surfaces is the minimum and almost always enough. In-place filters — brightness, saturation, a
colour LUT — do not even need the swap, and a smart apply can skip it for them, halving the traffic
again for a chain that is mostly point operations.
Describing the chain from JavaScript
The stage list has to cross the boundary, and it should cross as bytes rather than as objects. A compact encoding — one byte of opcode plus a fixed number of float parameters — keeps the crossing to a single small copy and avoids any string handling inside the module.
const OPS = { EXPOSURE: 1, CURVES: 2, BLUR: 3, SHARPEN: 4, GRAIN: 5, SATURATION: 6 };
function encodeChain(stages) {
const buf = new DataView(new ArrayBuffer(stages.length * 12));
stages.forEach((s, i) => {
buf.setUint8(i * 12, OPS[s.op]);
buf.setFloat32(i * 12 + 4, s.a ?? 0, true); // little-endian, always
buf.setFloat32(i * 12 + 8, s.b ?? 0, true);
});
return new Uint8Array(buf.buffer);
}
const chain = encodeChain([
{ op: 'EXPOSURE', a: 0.3 },
{ op: 'BLUR', a: 2.5 },
{ op: 'SHARPEN', a: 0.8, b: 1.2 },
]);
Twelve bytes per stage means a twenty-stage chain is 240 bytes — small enough that re-encoding and re-sending it on every slider movement is free. That matters for the preview loop below, where the chain changes sixty times a second while the image does not.
Separable blurs and fused kernels
The two optimisations that matter most in a filter chain are separability and fusion.
A Gaussian blur with radius r is O(r²) per pixel done naively and O(2r) done as a horizontal pass
followed by a vertical one. At radius 12 that is 625 multiply-adds against 50 — a twelve-fold difference
that dwarfs anything else you will do to the code. Every symmetric blur kernel is separable, and box
blurs are better still: three box passes approximate a Gaussian closely enough for photographic work and
run in constant time per pixel regardless of radius, using a running sum.
Fusion is the other half. A chain of three point operations — exposure, saturation, a curve — reads and writes the whole surface three times. Fused into one loop, it reads and writes once and keeps the pixel in registers in between. For a 12-megapixel image that is the difference between 300 MB and 100 MB of memory traffic, and on a memory-bound workload it shows up directly in the wall clock:
// fused point pass: one read, one write, three operations
for px in surface.chunks_exact_mut(4) {
let mut c = [px[0] as f32, px[1] as f32, px[2] as f32];
apply_exposure(&mut c, ev);
apply_saturation(&mut c, sat);
apply_curve(&mut c, &lut);
px[0] = c[0] as u8; px[1] = c[1] as u8; px[2] = c[2] as u8;
}
Detecting fusable runs is a five-line loop over the stage list: collect consecutive point operations, apply them together, and break the run whenever a neighbourhood operation such as a blur appears.
A preview that keeps up with a slider
Running the full chain on a 12-megapixel image takes a few hundred milliseconds, which is far too slow for a drag. The standard answer is a proxy: keep a downscaled copy — typically fitting within 1600 px on the long edge — and run the chain on that during interaction, then run the full-resolution chain once when the user lets go.
let pending = null;
slider.addEventListener('input', () => {
pending = encodeChain(currentStages());
if (!scheduled) { scheduled = true; requestAnimationFrame(renderPreview); }
});
function renderPreview() {
scheduled = false;
const ptr = mod.exports.run_pipeline_proxy(pw, ph, writeChain(pending), pending.length / 12);
blitToCanvas(ptr, pw, ph); // proxy resolution, ~6 ms
}
Coalescing through requestAnimationFrame is what makes this smooth: slider events fire far faster than
frames, and rendering each one means rendering work you are about to throw away. One render per frame,
always using the newest parameters, is both faster and more responsive than rendering every event.
Expected output
A correct pipeline is deterministic: the same stage list over the same source produces byte-identical output every run. That property is worth asserting in a test, because it catches uninitialised scratch memory — the classic bug where the first run is right and the second is subtly different.
const a = sha256(await runChain(src, chain));
const b = sha256(await runChain(src, chain));
console.log(a === b); // true — no state leaked between runs
The other check is stage-order sensitivity. Applying sharpen before blur must produce a different image from blur before sharpen; if it does not, your stage list is not actually being walked in order.
Gotchas
- Radius parameters not scaled for the proxy. A blur radius of 12 on a full-size image is about 2.5 on a 1600 px proxy. Scale every spatial parameter by the same factor or the preview lies.
- Clipping at stage boundaries. Storing 8-bit intermediates after each stage destroys highlight detail that a later stage could have recovered. For chains with more than three or four stages, keep the intermediates in 16-bit or float.
- Grain that changes every frame. A noise stage seeded from a clock makes the preview shimmer. Seed from a fixed value plus the pixel index so the result is stable and reproducible.
- Surface pointers cached across a resize. Loading a different image reallocates the arena and invalidates every pointer and view. Re-query them after any call that can resize.
- Alpha handled inconsistently. Some stages should ignore alpha, some should premultiply first. Decide once, document it next to the opcode table, and test with a semi-transparent image.
Performance note
For a 12-megapixel image, a six-stage chain with one separable blur runs in roughly 220–320 ms single-threaded. Fusing the point operations removes about 25% of that; SIMD on the blur removes another 30–40%. The proxy path — 2.5 megapixels, same chain — lands near 8 ms, which is what makes live adjustment possible at all. Tiling the full-resolution pass across a small worker pool takes the export under 100 ms on a four-core machine, with the caveat that blur stages need overlapping tile borders equal to the kernel radius.
Frequently Asked Questions
Why not do this on the GPU with WebGL or WebGPU? For long chains on large images, the GPU is faster, and it is the right answer for a real-time video effect. WebAssembly wins on precision control, on exact reproducibility across devices, and on not needing a GPU context at all — which matters when the same code must also run server-side.
How do I add a user-supplied filter? Keep the opcode table closed and expose a parameterised LUT stage instead. Accepting arbitrary code from users means accepting a sandbox problem, which is a different project — see loading untrusted plugins safely.
Can the chain run on a stream of video frames? Yes, and the design barely changes: the arena is sized for one frame and reused, so per-frame cost is just the arithmetic. Combine it with WebCodecs frames for a full hardware-decode-to-filtered-output path.
Should the stage list live in JavaScript or in the module? In JavaScript. The user interface owns the edit state, and the module should be a pure function of (pixels, stage list). Keeping filter state inside the module makes undo, presets and serialisation much harder than they need to be, and it makes the module impossible to reuse from a server-side build.
How should presets be stored? As the same encoded chain, base64 or JSON, versioned with the opcode table. Because the encoding is fixed-width and explicit, a preset saved today still applies cleanly after you add new opcodes — as long as you only ever append to the table and never renumber it.
Related
- Resizing images off the main thread — the proxy generation step.
- Writing v128 SIMD intrinsics in Rust — vectorising the inner loops.
- Implementing a bump allocator in Wasm — the arena the surfaces live in.
← Back to Media Processing & Codecs in Wasm