Building a Wasm Image Filter Pipeline

This guide answers one task: apply a sequence of image filters — exposure, curves, blur, sharpen, grain — inside a single WebAssembly module, with no copy between stages, no reallocation per adjustment, and a preview that updates while the user drags a slider.

Prerequisites

  • [ ] A Rust or C module built for the web, with an arena rather than a general-purpose allocator.
  • [ ] A canvas or ImageBitmap display path.
  • [ ] Familiarity with typed-array views over linear memory.
  • [ ] A test image with a wide tonal range — a blown-out sky and a dark foreground reveal filter bugs that a flat test pattern hides.

One module, two buffers, many stages

The naive design gives each filter its own input and output allocation and copies between them. For a twelve-megapixel image and six filters that is 300 MB of pointless traffic. The standard fix is ping-pong buffering: allocate exactly two full-size surfaces, and alternate which one is the source and which the destination as you walk the stage list.

static mut A: Vec<u8> = Vec::new();
static mut B: Vec<u8> = Vec::new();

#[no_mangle]
pub extern "C" fn run_pipeline(w: usize, h: usize, stages_ptr: *const u8, n: usize) -> *const u8 {
    unsafe {
        let stages = std::slice::from_raw_parts(stages_ptr, n);
        let (mut src, mut dst) = (&mut A as *mut Vec<u8>, &mut B as *mut Vec<u8>);
        for &stage in stages {
            apply(stage, &(*src), &mut (*dst), w, h);
            std::mem::swap(&mut src, &mut dst);      // no copy, just a pointer swap
        }
        (*src).as_ptr()                               // last write landed in src after the swap
    }
}

Two surfaces is the minimum and almost always enough. In-place filters — brightness, saturation, a colour LUT — do not even need the swap, and a smart apply can skip it for them, halving the traffic again for a chain that is mostly point operations.

Two surfaces, four stages, zero copies Each stage reads one surface and writes the other, then the roles swap. Nothing is allocated after startup and no intermediate result is ever copied, so the cost of a longer chain is only the arithmetic it adds. surface A source written by blur read by grain surface B written by curves read by sharpen The pointer swap costs nothing; the arithmetic is the only thing that scales with chain length. Point operations can write into their own input and skip the swap entirely, which halves memory traffic for typical chains. Return the pointer to whichever surface holds the final result — do not assume it is always A.

Describing the chain from JavaScript

The stage list has to cross the boundary, and it should cross as bytes rather than as objects. A compact encoding — one byte of opcode plus a fixed number of float parameters — keeps the crossing to a single small copy and avoids any string handling inside the module.

const OPS = { EXPOSURE: 1, CURVES: 2, BLUR: 3, SHARPEN: 4, GRAIN: 5, SATURATION: 6 };

function encodeChain(stages) {
  const buf = new DataView(new ArrayBuffer(stages.length * 12));
  stages.forEach((s, i) => {
    buf.setUint8(i * 12, OPS[s.op]);
    buf.setFloat32(i * 12 + 4, s.a ?? 0, true);     // little-endian, always
    buf.setFloat32(i * 12 + 8, s.b ?? 0, true);
  });
  return new Uint8Array(buf.buffer);
}

const chain = encodeChain([
  { op: 'EXPOSURE', a: 0.3 },
  { op: 'BLUR', a: 2.5 },
  { op: 'SHARPEN', a: 0.8, b: 1.2 },
]);

Twelve bytes per stage means a twenty-stage chain is 240 bytes — small enough that re-encoding and re-sending it on every slider movement is free. That matters for the preview loop below, where the chain changes sixty times a second while the image does not.

Separable blurs and fused kernels

The two optimisations that matter most in a filter chain are separability and fusion.

A Gaussian blur with radius r is O(r²) per pixel done naively and O(2r) done as a horizontal pass followed by a vertical one. At radius 12 that is 625 multiply-adds against 50 — a twelve-fold difference that dwarfs anything else you will do to the code. Every symmetric blur kernel is separable, and box blurs are better still: three box passes approximate a Gaussian closely enough for photographic work and run in constant time per pixel regardless of radius, using a running sum.

Fusion is the other half. A chain of three point operations — exposure, saturation, a curve — reads and writes the whole surface three times. Fused into one loop, it reads and writes once and keeps the pixel in registers in between. For a 12-megapixel image that is the difference between 300 MB and 100 MB of memory traffic, and on a memory-bound workload it shows up directly in the wall clock:

// fused point pass: one read, one write, three operations
for px in surface.chunks_exact_mut(4) {
    let mut c = [px[0] as f32, px[1] as f32, px[2] as f32];
    apply_exposure(&mut c, ev);
    apply_saturation(&mut c, sat);
    apply_curve(&mut c, &lut);
    px[0] = c[0] as u8; px[1] = c[1] as u8; px[2] = c[2] as u8;
}

Detecting fusable runs is a five-line loop over the stage list: collect consecutive point operations, apply them together, and break the run whenever a neighbourhood operation such as a blur appears.

A preview that keeps up with a slider

Running the full chain on a 12-megapixel image takes a few hundred milliseconds, which is far too slow for a drag. The standard answer is a proxy: keep a downscaled copy — typically fitting within 1600 px on the long edge — and run the chain on that during interaction, then run the full-resolution chain once when the user lets go.

let pending = null;
slider.addEventListener('input', () => {
  pending = encodeChain(currentStages());
  if (!scheduled) { scheduled = true; requestAnimationFrame(renderPreview); }
});

function renderPreview() {
  scheduled = false;
  const ptr = mod.exports.run_pipeline_proxy(pw, ph, writeChain(pending), pending.length / 12);
  blitToCanvas(ptr, pw, ph);                       // proxy resolution, ~6 ms
}

Coalescing through requestAnimationFrame is what makes this smooth: slider events fire far faster than frames, and rendering each one means rendering work you are about to throw away. One render per frame, always using the newest parameters, is both faster and more responsive than rendering every event.

Proxy while dragging, full render on release During a drag the chain runs on a downscaled proxy that fits the frame budget. When the interaction ends, the same chain runs once at full resolution, so the user sees immediate feedback and an accurate final image. while dragging — proxy 1600 px ≈ 2.5 megapixels, whole chain one render per animation frame 6–10 ms — comfortably 60 fps on release — full 12 megapixels identical stage list, identical code runs once, in a worker 250–400 ms — with a progress hint Use the same kernels for both so the preview cannot diverge from the result — a proxy rendered by different code is worse than no preview. Scale radius-based parameters by the proxy factor, or a blur will look weaker in the preview than in the export.

Expected output

A correct pipeline is deterministic: the same stage list over the same source produces byte-identical output every run. That property is worth asserting in a test, because it catches uninitialised scratch memory — the classic bug where the first run is right and the second is subtly different.

const a = sha256(await runChain(src, chain));
const b = sha256(await runChain(src, chain));
console.log(a === b);   // true — no state leaked between runs

The other check is stage-order sensitivity. Applying sharpen before blur must produce a different image from blur before sharpen; if it does not, your stage list is not actually being walked in order.

Why the chain belongs inside the module Calling once per filter copies the image in and out four times. Passing a filter list and running the chain inside the module leaves one crossing and one buffer. one call per filter 4 crossings, 8 copies 62 ms one call, filter list 1 crossing, 2 copies 19 ms — the arithmetic is identical The saving is not in the filters; it is in the copies that disappear between them. Keeping intermediates in a reused scratch buffer removes the allocator from the hot path as well.

Gotchas

  • Radius parameters not scaled for the proxy. A blur radius of 12 on a full-size image is about 2.5 on a 1600 px proxy. Scale every spatial parameter by the same factor or the preview lies.
  • Clipping at stage boundaries. Storing 8-bit intermediates after each stage destroys highlight detail that a later stage could have recovered. For chains with more than three or four stages, keep the intermediates in 16-bit or float.
  • Grain that changes every frame. A noise stage seeded from a clock makes the preview shimmer. Seed from a fixed value plus the pixel index so the result is stable and reproducible.
  • Surface pointers cached across a resize. Loading a different image reallocates the arena and invalidates every pointer and view. Re-query them after any call that can resize.
  • Alpha handled inconsistently. Some stages should ignore alpha, some should premultiply first. Decide once, document it next to the opcode table, and test with a semi-transparent image.

Performance note

For a 12-megapixel image, a six-stage chain with one separable blur runs in roughly 220–320 ms single-threaded. Fusing the point operations removes about 25% of that; SIMD on the blur removes another 30–40%. The proxy path — 2.5 megapixels, same chain — lands near 8 ms, which is what makes live adjustment possible at all. Tiling the full-resolution pass across a small worker pool takes the export under 100 ms on a four-core machine, with the caveat that blur stages need overlapping tile borders equal to the kernel radius.

Frequently Asked Questions

Why not do this on the GPU with WebGL or WebGPU? For long chains on large images, the GPU is faster, and it is the right answer for a real-time video effect. WebAssembly wins on precision control, on exact reproducibility across devices, and on not needing a GPU context at all — which matters when the same code must also run server-side.

How do I add a user-supplied filter? Keep the opcode table closed and expose a parameterised LUT stage instead. Accepting arbitrary code from users means accepting a sandbox problem, which is a different project — see loading untrusted plugins safely.

Can the chain run on a stream of video frames? Yes, and the design barely changes: the arena is sized for one frame and reused, so per-frame cost is just the arithmetic. Combine it with WebCodecs frames for a full hardware-decode-to-filtered-output path.

Should the stage list live in JavaScript or in the module? In JavaScript. The user interface owns the edit state, and the module should be a pure function of (pixels, stage list). Keeping filter state inside the module makes undo, presets and serialisation much harder than they need to be, and it makes the module impossible to reuse from a server-side build.

How should presets be stored? As the same encoded chain, base64 or JSON, versioned with the opcode table. Because the encoding is fixed-width and explicit, a preset saved today still applies cleanly after you add new opcodes — as long as you only ever append to the table and never renumber it.

← Back to Media Processing & Codecs in Wasm