Feeding WebCodecs Frames into Wasm

This guide answers one task: let the browser’s hardware decoder produce video frames, move each frame into a WebAssembly module for processing, and get the result back onto the screen — at frame rate, with no frame leaks and no accidental copies.

Prerequisites

  • [ ] A browser with WebCodecs: Chrome 94+, Edge 94+, Safari 16.4+, Firefox 130+.
  • [ ] A compiled kernel that processes a pixel or plane buffer in place.
  • [ ] A demuxer for your container — mp4box.js or similar — because VideoDecoder takes encoded chunks, not files.
  • [ ] A worker, ideally with OffscreenCanvas, so none of this runs on the main thread.

The division of labour

WebCodecs gives you the one thing a compiled decoder cannot: access to the platform’s hardware video decoder. It will decode H.264, HEVC, VP9 and AV1 faster and at lower power than any module you can compile, because it is not running on the CPU at all. What it does not give you is processing — no filters, no analysis, no custom colour work.

So the split is clean. VideoDecoder produces VideoFrame objects; your module consumes pixels and produces pixels; a canvas or an encoder consumes the result. The engineering is entirely in the seam between them, because a VideoFrame is an opaque handle that may live in GPU memory, and linear memory is a plain byte array on the CPU.

Hardware decode, compiled processing, canvas output Encoded chunks go to the platform decoder, which emits VideoFrame handles. Each frame is copied once into linear memory where the module processes it, and the result is drawn to a canvas. The copy is the only point where CPU memory is touched. EncodedVideoChunk from the demuxer VideoDecoder hardware, off-CPU copyTo linear memory process planes in place canvas or encoder Every frame you receive must be closed. A VideoFrame holds a hardware buffer, and the decoder stalls once the pool is exhausted — which presents as playback that runs for two seconds and then freezes with no error at all. If your kernel only reads luma — most analysis does — copy only plane 0 and leave chroma on the GPU.

Demuxing: the part nobody mentions

VideoDecoder does not read files. It consumes EncodedVideoChunk objects, each carrying a timestamp, a type — key or delta — and the compressed bytes for one frame. Producing those from an MP4 means parsing the container, and the browser gives you nothing for it. This surprises almost everyone who starts with WebCodecs, because the decoding half looks so complete.

In practice you use a JavaScript demuxer such as mp4box.js, feed it the file in chunks, and turn its sample callbacks into decoder input. The decoder also needs a configuration before the first chunk, including the codec string and the codec-private data — the avcC box for H.264 — which the demuxer extracts for you:

mp4box.onReady = (info) => {
  const track = info.videoTracks[0];
  decoder.configure({
    codec: track.codec,                              // e.g. "avc1.640028"
    codedWidth: track.video.width,
    codedHeight: track.video.height,
    description: avccBoxFrom(mp4box, track),         // required for H.264 in MP4
  });
  mp4box.setExtractionOptions(track.id);
  mp4box.start();
};

mp4box.onSamples = (id, user, samples) => {
  for (const s of samples) {
    decoder.decode(new EncodedVideoChunk({
      type: s.is_sync ? 'key' : 'delta',
      timestamp: (s.cts * 1e6) / s.timescale,        // microseconds
      duration: (s.duration * 1e6) / s.timescale,
      data: s.data,
    }));
  }
};

Two details cause most of the configuration failures. The timestamp must be in microseconds, not milliseconds — getting this wrong produces frames that appear to arrive out of order or all at once. And the description is mandatory for MP4-contained H.264 and HEVC: without it, configure() either throws or the decoder emits nothing, depending on the browser. Annex-B streams from a live source do not need it, which is why code copied from a WebRTC example often fails on a file.

If the container is WebM rather than MP4, the demuxer changes but the shape does not. Either way, budget real time for this half of the work — it is typically more code than the processing kernel itself.

Copying a frame into linear memory

VideoFrame.copyTo(destination) writes the frame’s planes into an ArrayBufferView you supply — which can be a view onto linear memory. That is the whole trick: allocate a region large enough for the frame’s layout once, and copy directly into it every frame.

const layoutFor = (frame) => frame.allocationSize();          // bytes needed for all planes

let region = 0, regionSize = 0;
async function frameIntoWasm(mod, frame) {
  const need = frame.allocationSize();
  if (need > regionSize) {                                     // grows at most once, at startup
    region = mod.exports.alloc_frame(need);
    regionSize = need;
  }
  const view = new Uint8Array(mod.exports.memory.buffer, region, need);
  const layout = await frame.copyTo(view);                     // planes written back to back
  return { region, layout };
}

copyTo resolves with a plane layout array giving the offset and stride of each plane. Do not compute those yourself: for a I420 frame the chroma planes are half resolution, strides are usually padded, and the arithmetic differs per format and per platform. Pass the layout into your module and let it index accordingly.

const { region, layout } = await frameIntoWasm(mod, frame);
mod.exports.process_i420(
  region,
  layout[0].offset, layout[0].stride,   // Y
  layout[1].offset, layout[1].stride,   // U
  layout[2].offset, layout[2].stride,   // V
  frame.codedWidth, frame.codedHeight,
);
frame.close();                           // release the hardware buffer immediately

Closing frames, or watching playback die

The single most common failure in a WebCodecs pipeline is a leaked VideoFrame. Each one holds a buffer from a fixed pool; when the pool empties, the decoder simply stops emitting frames. There is no exception, no warning, and no obvious clue — playback runs for a second or two and then halts.

The rules are simple and absolute. Close every frame you receive from the decoder’s output callback, including frames you decide to skip. Close frames you construct yourself. If you hand a frame to another context, decide explicitly which side owns it, and close it exactly once. Wrapping the whole handler in try/finally is the cheapest insurance available:

const decoder = new VideoDecoder({
  output: async (frame) => {
    try {
      await processFrame(frame);
    } finally {
      frame.close();                     // even if processing threw
    }
  },
  error: (e) => console.error('decode', e),
});

Expected output

A healthy pipeline reports a steady frame count and a stable queue depth. Log both; either one alone can look fine while the other is quietly wrong.

setInterval(() => {
  console.log({ decoded, processed, queue: decoder.decodeQueueSize });
}, 1000);
// { decoded: 30, processed: 30, queue: 2 }     healthy
// { decoded: 41, processed: 12, queue: 58 }    processing is the bottleneck
// { decoded: 12, processed: 12, queue: 0 }     decoder starved — feed it more chunks

A growing decodeQueueSize means your kernel cannot keep up and frames are piling up in the decoder. The fix is either a faster kernel or deliberate frame dropping — process every second frame and close the rest — which is far better than an unbounded queue that ends in an out-of-memory crash.

Reading the queue depth A queue that stays shallow means the kernel keeps pace. A queue that grows means processing is the bottleneck and frames must be dropped deliberately. A queue that is always empty means the demuxer is not supplying chunks fast enough. queue 1–4 healthy — keep going queue climbing drop frames on purpose queue always 0 decoder is waiting on chunks demux further ahead Backpressure is your job: WebCodecs will happily accept more chunks than it can decode, and the queue is the only signal you get. Stop feeding when decodeQueueSize passes a threshold, and resume when it drains — a dozen lines that prevent every memory blow-up in this pipeline.

Working in YUV instead of RGBA

The instinct is to convert every frame to RGBA and work there, because that is what canvases want. Resist it when you can. A 1080p I420 frame is 3.1 MB; the same frame as RGBA is 8.3 MB, so the conversion costs almost three times the memory traffic before your kernel does any work at all.

Many operations do not need the conversion. Brightness, contrast and gamma act on luma only. Motion detection, scene-change detection, optical flow, QR scanning and most computer vision work on the Y plane alone. Blur and sharpen can be applied per plane. For all of those, copy plane 0, process it, and write the result back — the chroma planes never have to leave the GPU.

When you genuinely need RGB — a colour LUT, a chroma key, anything that mixes channels — do the conversion inside the module where it vectorises well, rather than through a canvas round trip. A SIMD I420 to RGBA conversion runs at well over a gigabyte per second, while drawImage plus getImageData costs a GPU readback and a synchronisation point.

Frame in, frame out WebCodecs owns the decoded frame. copyTo writes its planes into the module's memory, the module processes in place, and a new VideoFrame is constructed over the result. VideoFrame owned by WebCodecs copyTo(memory) planes written in module processes in place, no copy new VideoFrame handed onward Close every frame you receive: an unclosed VideoFrame holds a decoder buffer and stalls the pipeline. Allocate the destination region once at the known frame size rather than per frame. Respect the frame's layout — stride is often wider than the visible width, and ignoring it shears the image.

Gotchas

  • copyTo throws NotSupportedError. The frame’s format cannot be read directly — common with certain hardware paths. Fall back to drawing the frame to an OffscreenCanvas and reading it back, which is slower but always works.
  • Playback stalls after ~30 frames. A leaked VideoFrame. Audit every path out of the output callback for a missing close().
  • Colours look washed out. Full-range versus limited-range YUV. Check frame.colorSpace and apply the right coefficients; assuming BT.709 full range for everything is a common and visible mistake.
  • The first frames are garbage. You started decoding at a non-keyframe. Seek to the nearest keyframe and decode forward, discarding output until you reach the target timestamp.
  • await frame.copyTo() slows everything down. It is asynchronous for a reason, but awaiting inside a tight loop serialises the pipeline. Keep two frames in flight so the copy of one overlaps the processing of the other.

Performance note

On a modern laptop, hardware decode of 1080p30 costs under 2 ms per frame and almost no CPU. copyTo into linear memory costs roughly 1–2 ms for a full I420 frame, dominated by memory bandwidth. That leaves around 28 ms of the 33 ms frame budget for your kernel, which is a lot — a SIMD convolution over the luma plane fits comfortably. The pipeline falls over when you add a canvas round trip per frame, which reintroduces a GPU readback and can cost 8–15 ms on its own.

Frequently Asked Questions

Can I use WebCodecs to encode the processed frames again? Yes — construct a VideoFrame from your processed buffer and feed it to a VideoEncoder. That gives a fully hardware-accelerated round trip with your kernel in the middle, which is far faster than any compiled encoder.

Does this work in a worker? It should. VideoDecoder, VideoFrame and OffscreenCanvas are all available in workers, and running there keeps the main thread completely free. Transfer the encoded chunks in and the finished frames out.

How does this compare to ffmpeg.wasm? Different tools. WebCodecs is faster by an order of magnitude and limited to the codecs the platform supports; ffmpeg.wasm handles anything but does everything on the CPU. Use WebCodecs for the common formats and keep the compiled fallback for the rest.

← Back to Media Processing & Codecs in Wasm