Media Processing & Codecs in Wasm
Media is the workload WebAssembly was effectively designed for: large buffers of bytes, arithmetic-heavy inner loops, and mature C implementations that nobody wants to rewrite in JavaScript. A single 1080p frame is roughly 8.3 MB of RGBA; a one-minute stereo track at 48 kHz is 23 MB of float samples. Moving that much data across the boundary carelessly costs more than the processing, so the engineering problem is rarely the codec itself — it is owning the buffers, keeping them in one place, and letting the module work in the memory the data already lives in.
Prerequisites
- [ ] A toolchain that can build a C or Rust media library —
emcc3.1+ orwasm-pack0.13+. - [ ] A dev server that sets
Cross-Origin-Opener-PolicyandCross-Origin-Embedder-Policyif you plan to use threads. - [ ] Chrome 94+ or Firefox 130+ for
WebCodecsinterop examples; everything else works anywhere. - [ ] A test asset set: one short H.264 clip, one large JPEG, one WAV file. Synthetic inputs hide real bottlenecks.
What a media pipeline looks like in memory
The shape that works is always the same. The module owns a region of linear memory sized for the
largest frame you expect, JavaScript writes into that region through a typed-array view, the module
processes in place or into a second region, and JavaScript reads the output through a second view. The
frame never exists twice unless you ask for it twice.
The region pointers are obtained once, at startup, and cached:
const cap = 1920 * 1080 * 4;
const inPtr = mod.exports.alloc_input(cap);
const outPtr = mod.exports.alloc_output(cap);
// reused for every frame for the life of the instance
If your module can guarantee it never grows memory after this point — which a fixed arena can — the
views can be cached too, and the per-frame cost drops to one set() and one function call.
Building a codec library for the browser
Most real pipelines start from an existing C library. The build is unremarkable; the flags that matter are the ones that decide what the JavaScript side can reach and how big the result is.
emcc decoder.c libjpeg/*.c \
-O3 -msimd128 \
-sMODULARIZE=1 -sEXPORT_ES6=1 \
-sEXPORTED_FUNCTIONS='["_alloc_input","_alloc_output","_decode_frame","_free"]' \
-sEXPORTED_RUNTIME_METHODS='["HEAPU8"]' \
-sALLOW_MEMORY_GROWTH=1 -sINITIAL_MEMORY=64MB \
-o decoder.mjs
-msimd128 is usually worth 1.5–3× on pixel loops and costs a few kilobytes; ship it behind a capability
check with a baseline build as covered in
shipping SIMD and baseline builds together.
-sINITIAL_MEMORY is set high deliberately: a decoder that grows from 16 MB to 100 MB during its first
frame pays for several reallocations and invalidates every view in flight. For a Rust pipeline the
equivalent is wasm-pack build --release with RUSTFLAGS="-C target-feature=+simd128" and an arena
crate instead of the default allocator.
Keeping the main thread out of it
Decoding a frame takes milliseconds; decoding thirty takes a frame budget’s worth of blocked UI. Media
pipelines belong in a worker, with the instance and its memory owned entirely by that worker. The page
sends compressed input and receives either a Transferable result or, better, an ImageBitmap the
worker has already produced.
// worker.js — owns the instance, never posts the whole heap back
import init, { decode } from './decoder.mjs';
const mod = await init();
self.onmessage = async ({ data }) => {
const { bytes, id } = data;
const rgba = decode(mod, bytes); // view into linear memory
const bmp = await createImageBitmap(new ImageData(rgba, w, h));
self.postMessage({ id, bmp }, [bmp]); // transferred, not copied
};
Posting the ImageBitmap rather than the pixel array matters: a transfer moves ownership without a copy,
while posting a typed array backed by linear memory forces a structured clone of every byte. The same
reasoning drives
compiling Wasm in a worker to free the main thread
— the compile itself is expensive enough to be worth moving.
Payload budgets for media modules
Media libraries are large. Knowing the rough numbers before you start saves a redesign:
| Module | Typical size (Brotli) | Notes |
|---|---|---|
| JPEG decoder (libjpeg-turbo subset) | 90–160 kB | Comfortable for first load |
| PNG + zlib | 60–110 kB | Often already covered by the browser |
| AVIF / JPEG XL decoder | 400 kB–1.2 MB | Load lazily, cache aggressively |
| Opus encode + decode | 250–400 kB | Worth it for real-time audio |
| ffmpeg.wasm full build | 10–25 MB | Only for explicit user-initiated work |
Nothing on that list should be fetched during first paint except the small ones. The pattern that works is: render the page, let the user choose a file, and only then fetch the codec while showing progress — the fetch and the user’s file picker overlap, so the perceived cost is close to zero. Pair that with caching the compiled module so the second use is instant.
Gotchas and failure modes
RangeError: WebAssembly.Memory(): could not allocate memory— a 4K frame pipeline asking for 512 MB on a mobile device. Cap your working set, process in tiles, and treat allocation failure as a path you handle rather than an exception you log.- Colour space and stride surprises. A decoder that emits YUV planes with row padding will produce
a sheared image if you assume tightly packed RGBA. Read the stride the library reports, never assume
width * 4. - Detached views mid-frame. Any call that can allocate may grow memory. If you cache views, you must either guarantee no growth after startup or rebuild them after every call — see why memory.grow invalidates pointers.
- Audio glitches from garbage collection, not from Wasm. If DSP runs in an
AudioWorklet, the processing must not allocate at all. A single allocation in the audio callback is audible. SharedArrayBuffer is not definedin a threaded build served without cross-origin isolation. Fix the headers as in configuring COOP/COEP headers, or ship the single-threaded build.
Verifying a media pipeline
Correctness for media is not “it looks right” — it is a checksum against a reference implementation. Decode the same asset with the native library and with the compiled module and compare bytes:
# reference, on the host
djpeg -outfile ref.raw test.jpg && sha256sum ref.raw
# under a standalone runtime, the same code path the browser will run
wasmtime run --dir=. decoder.wasm -- test.jpg out.raw && sha256sum out.raw
Identical digests mean the compile did not change behaviour. If they differ, the usual causes are
floating-point differences from -ffast-math, an uninitialised scratch buffer that happened to be zero
natively, or a SIMD path that diverges on edge pixels. For lossy pipelines where exact equality is not
expected, compare PSNR against a stored baseline and fail the build when it regresses — a check that
belongs in CI alongside the
size regression gate.
When the browser already has a decoder
Before compiling anything, check what the platform gives you. Browsers decode JPEG, PNG, WebP, GIF and —
increasingly — AVIF natively, on a background thread, with hardware acceleration on some platforms. A
compiled decoder for those formats will lose to createImageBitmap almost every time, and it costs the
user a download. The cases where your own decoder wins are narrower than they look:
- A format the browser does not support, or supports only in some of the browsers you target. JPEG XL and certain HDR profiles are the current examples, and the gap moves every year.
- Access to intermediate data.
createImageBitmapgives you an opaque handle; if you need raw coefficients, per-plane YUV, or metadata the browser drops, you need the library. - Deterministic output across browsers. Native decoders differ in chroma upsampling and colour management. If two users must see byte-identical pixels — a medical or measurement context — you have to own the decode.
- Encoding, not decoding. The platform’s encoding options are thin: quality and format, with no control over the parameters that matter for a specific corpus. A compiled encoder gives you all of them.
The pragmatic architecture uses both. Try the native path, fall back to the module, and record which path ran so you can see the real distribution in the field rather than guessing at it:
async function decode(blob, mime) {
if (await supportsNatively(mime)) return createImageBitmap(blob); // free, fast
return decodeWithWasm(new Uint8Array(await blob.arrayBuffer())); // our module
}
Streaming instead of whole-file processing
The naive pipeline reads an entire file into memory, hands it to the module, and waits. For a 40 MB video or a 200 MB WAV, that means the page holds the whole file, the module holds a copy, and the peak footprint is comfortably past what a phone will tolerate. Streaming fixes it, and the structure is not much harder.
Read the source through a ReadableStream reader, push each chunk into a fixed input region, and let the
module consume chunks and emit whatever it can. The module keeps its own parser state between calls, so
it needs an explicit “feed” and “drain” API rather than a single process_everything export:
const reader = file.stream().getReader();
const CHUNK = mod.exports.input_capacity();
const inView = new Uint8Array(mod.exports.memory.buffer, mod.exports.input_ptr(), CHUNK);
for (;;) {
const { value, done } = await reader.read();
if (done) break;
for (let off = 0; off < value.length; off += CHUNK) {
const slice = value.subarray(off, off + CHUNK);
inView.set(slice);
mod.exports.feed(slice.length); // parse what we can, buffer the rest
drainOutput(mod); // emit finished frames as they appear
}
}
mod.exports.finish();
The peak memory becomes the chunk size plus whatever internal state the codec keeps, which for most formats is small and bounded. The user-visible effect is larger than the memory saving: the first frame appears while the rest of the file is still downloading, which turns a ten-second wait into an immediate result. This is the same argument that makes streaming instantiation the default for the module itself.
Tiling and threads for large surfaces
Very large images — scanned documents, satellite tiles, print-resolution exports — defeat a single-buffer design. A 12000 × 9000 RGBA surface is 432 MB, which no browser tab should be holding. The answer is tiling: process a band at a time, with an overlap large enough to cover the filter kernel’s radius so seams do not appear at the boundaries.
Tiling also makes threading trivial, because tiles are independent. With a threaded build and a worker pool, each worker takes a tile index, processes into its own slice of the shared output buffer, and signals completion. There is no locking because no two workers touch the same output bytes — the only synchronisation is the barrier at the end, which Atomics handles cleanly.
The speedup is close to linear up to the physical core count and then flattens, and it is worth measuring rather than assuming: on a four-core laptop a tiled blur typically lands at 3.2–3.6× the single-threaded time, with the shortfall going to memory bandwidth rather than to coordination overhead. If your kernel is bandwidth-bound rather than compute-bound — a simple copy or a channel swap — threads will buy almost nothing, and the honest move is to keep the single-threaded build and its smaller payload.
The data layout contract
Every bug that survives the first day of a media integration is a layout misunderstanding. The module and the page have to agree on five things, and none of them is inferable from the byte count alone: pixel format, channel order, stride, origin, and whether values are premultiplied.
Pixel format and channel order are the obvious pair. A C library that says “RGB” may mean packed 24-bit
with no alpha, while ImageData always wants 32-bit RGBA. Swapping red and blue produces an image that
looks plausible in a thumbnail and obviously wrong once someone notices skin tones, which is why this
bug reliably reaches production. Write the conversion explicitly at the boundary, in the module where
it is cheap, rather than in JavaScript where it is a per-pixel loop in the wrong language.
Stride is the one that bites hardest. Many decoders align each row to a 4-, 8- or 16-byte boundary, so a
row of 1021 pixels occupies more bytes than 1021 * 4. Reading such a buffer as if it were tightly
packed shears the image progressively down the frame — the classic diagonal-smear screenshot. Always ask
the library for the stride it used and copy row by row when it differs from the packed width:
// stride-aware copy out of linear memory into a packed RGBA buffer
const packed = new Uint8ClampedArray(w * h * 4);
for (let y = 0; y < h; y++) {
const src = new Uint8ClampedArray(memory.buffer, outPtr + y * stride, w * 4);
packed.set(src, y * w * 4);
}
Origin — whether row zero is the top or the bottom of the image — differs between graphics APIs and image libraries, and a vertically flipped result is usually a one-line fix in the copy loop rather than a reprocessing step. Premultiplied alpha is the subtlest of the five: compositing premultiplied data as if it were straight alpha darkens edges, and the difference only shows on semi-transparent pixels, so a test image with a hard-edged alpha mask will pass while real content looks wrong.
Write these five properties down in a comment next to the export signature, and assert what you can at
runtime in development builds. A module that exports frame_stride() and frame_format() alongside its
pointer costs nothing and removes an entire class of afternoon-long debugging sessions. The broader
version of this discipline — documenting the memory layout as part of the interface — is covered in
returning structs from Wasm to JavaScript.
Frequently Asked Questions
Why is my Wasm decoder slower than the browser’s built-in one? Because the built-in one is native code with SIMD and often hardware acceleration, running off the main thread, with no boundary crossing and no download. A compiled decoder competes on capability, not on raw speed for formats the platform already handles.
How do I avoid a copy when the source is a File?
You cannot avoid the first one — bytes have to get from the file into linear memory — but you can make
it the only one. Read into the module’s input region directly with view.set(), process in place, and
return a pointer rather than a new array. One copy per item is the floor; anything more is a bug.
Can I use ffmpeg.wasm for real-time processing?
Not for anything with a latency budget. The full build is tens of megabytes and its process model
assumes files, not streams. For real-time work compile the specific codec you need with a streaming API,
or use WebCodecs for the decode and Wasm only for the parts the platform does not do.
Does SIMD help audio as much as it helps images? Usually more, because audio kernels are pure float loops over contiguous samples with no branching. FIR filters, mixing and resampling commonly see 2–4×. The constraint in audio is not throughput but jitter: the callback must finish within the quantum every single time.
Guides in this topic
- Transcoding video in the browser with ffmpeg.wasm — the big hammer, and how to load it without hurting first paint.
- Decoding modern image formats with Wasm codecs — AVIF and JPEG XL where the browser has no decoder.
- Running a DSP kernel in an AudioWorklet — real-time audio with a hard no-allocation rule.
- Resizing images off the main thread — a worker, a resampler and an ImageBitmap.
- Feeding WebCodecs frames into Wasm — hardware decode plus compiled post-processing.
- Building a Wasm image filter pipeline — chaining filters without a copy between stages.
Related
- Zero-Copy Data Transfer Patterns — the buffer discipline this topic depends on.
- Wasm SIMD & Vectorized Computation — where the pixel-loop speedups come from.
- Graphics, Games & Simulation — the same buffers, aimed at a canvas.
← Back to Production Wasm: Workloads & Deployment