Running a DSP Kernel in an AudioWorklet
This guide answers one task: run a compiled DSP kernel — a filter, a reverb, a pitch shifter — inside an
AudioWorkletProcessor, where the callback must complete within a render quantum every single time or
the user hears a click.
Prerequisites
- [ ] A kernel compiled with no runtime dependencies:
wasm-pack build --target weboremccwith-sSTANDALONE_WASMfor the simplest case. - [ ] A secure context —
AudioWorkletis unavailable over plain HTTP except on localhost. - [ ] Understanding that the worklet thread has no
fetch, noWebAssembly.instantiateStreaming, and no DOM. - [ ] Headphones. Audio glitches are invisible on a waveform and obvious in your ears.
The constraint that shapes everything
process() is called on a dedicated real-time audio thread once per render quantum: 128 frames, which at
48 kHz is 2.67 milliseconds. If your callback takes longer than that, the audio graph underruns and the
user hears a click or a dropout. There is no recovery, no backpressure and no queue — the deadline is
hard.
That single fact rules out several things you would do anywhere else. You cannot allocate: a
new Float32Array(n) inside the callback may trigger a garbage collection whose pause is longer than the
budget. You cannot await anything. You cannot post a message and wait for a reply. And you cannot grow
linear memory, because growth reallocates and copies the entire heap. Every buffer the kernel will
ever need must exist before the first callback runs.
Get the module into the worklet scope
The worklet thread cannot fetch. The standard pattern is to fetch the bytes on the main thread, pass them
through port.postMessage, and instantiate synchronously inside the processor — synchronous
instantiation is permitted for buffers under a few megabytes, and a DSP kernel is far below that limit.
// main thread
const ctx = new AudioContext({ latencyHint: 'interactive' });
await ctx.audioWorklet.addModule('/dsp-processor.js');
const bytes = await (await fetch('/dsp.wasm')).arrayBuffer();
const node = new AudioWorkletNode(ctx, 'dsp', { processorOptions: { bytes } });
node.connect(ctx.destination);
// dsp-processor.js — runs on the audio thread
class DspProcessor extends AudioWorkletProcessor {
constructor({ processorOptions }) {
super();
const mod = new WebAssembly.Module(processorOptions.bytes); // synchronous, once
this.inst = new WebAssembly.Instance(mod, {});
const { memory, init, in_ptr, out_ptr } = this.inst.exports;
init(128); // preallocate everything
this.inView = new Float32Array(memory.buffer, in_ptr(), 128);
this.outView = new Float32Array(memory.buffer, out_ptr(), 128);
}
process(inputs, outputs) {
const inCh = inputs[0][0], outCh = outputs[0][0];
if (!inCh) return true;
this.inView.set(inCh);
this.inst.exports.render(128);
outCh.set(this.outView);
return true;
}
}
registerProcessor('dsp', DspProcessor);
Passing the bytes through processorOptions rather than a later postMessage means the instance exists
before the first process() call, which removes an entire class of “the first half second is silence”
bugs. The two views are built once in the constructor and are valid for the life of the node — but only
because init() preallocated and the kernel never grows memory afterwards.
Planar channels, not interleaved
The Web Audio API hands you planar buffers: inputs[0] is an array of channels, each a separate
Float32Array of 128 samples. Most C DSP code expects the same layout, but a lot of older code expects
interleaved stereo. Converting per quantum is cheap, but doing it in JavaScript is unnecessary work on
the audio thread — do it inside the module where it is a tight loop, or better, compile the kernel to
accept planar input directly.
process(inputs, outputs) {
const input = inputs[0];
for (let c = 0; c < input.length; c++) this.chanViews[c].set(input[c]);
this.inst.exports.render_planar(input.length, 128);
const output = outputs[0];
for (let c = 0; c < output.length; c++) output[c].set(this.outViews[c]);
return true;
}
Allocate one region per channel in init() and keep a view per channel. Two channels at 128 frames is
1 kB of linear memory — the cost is irrelevant, and the resulting code has no per-quantum branching on
channel count.
Expected output
A working kernel is silent about its success, so instrument it deliberately. Time the render call and
report the worst case back to the main thread every second or so — never per quantum, because
postMessage from the audio thread is itself work.
const t0 = performance.now();
this.inst.exports.render(128);
const dt = performance.now() - t0;
if (dt > this.worst) this.worst = dt;
if (++this.n % 375 === 0) { this.port.postMessage({ worstMs: this.worst }); this.worst = 0; }
// console, once per second
{ worstMs: 0.41 } // 15% of the 2.67 ms budget — comfortable
{ worstMs: 0.44 }
{ worstMs: 2.91 } // over budget: something allocated, or the machine was loaded
A worst case above about half the quantum is a warning even if you hear nothing: it means a busy machine will glitch. Investigate before shipping rather than after a support ticket.
Getting audio in and out of the graph
A kernel is only useful once it is wired into a real signal path, and the surrounding graph decides how
much headroom your callback actually has. Three sources cover almost every case: a MediaElementSource
for playback of a file, a MediaStreamSource for microphone input, and an OscillatorNode or buffer
source for synthesis. Each connects to your node the same way, and each brings a different latency
profile with it.
Microphone input is the demanding one. The browser adds its own capture buffer ahead of your processor,
and requesting a low latencyHint shortens that buffer at the cost of a smaller safety margin — exactly
the margin your kernel is competing for. If you are building a live effect, start with
latencyHint: 'interactive' and only move to a numeric hint after you have measured the worst-case
render time on real hardware. Playback of a file is far more forgiving, because nothing downstream is
waiting on a human.
const stream = await navigator.mediaDevices.getUserMedia({
audio: { echoCancellation: false, noiseSuppression: false, autoGainControl: false },
});
const src = ctx.createMediaStreamSource(stream);
src.connect(node); // your worklet node
node.connect(ctx.destination);
Disabling the built-in processing matters when your kernel is the effect: echo cancellation and automatic
gain will fight anything you do, and they are enabled by default. For an analysis-only kernel you
usually leave them on and connect the node to a GainNode with zero gain so the graph still pulls
quanta through your processor without routing them to the speakers.
One more practical detail: an AudioContext starts suspended until a user gesture resumes it, so your
first render never happens until someone clicks. Wire the resume() call to the same button that starts
the feature, and check ctx.state before assuming silence is a bug in your kernel.
Parameters and control changes
Audio parameters change while the graph runs, and the temptation is to postMessage a new value from the
UI. That works but arrives whenever the thread gets around to it, which means a knob turn can land mid
quantum and produce a zipper noise. Two better options exist.
For sample-accurate automation, declare parameterDescriptors and read the parameters argument in
process(). The browser hands you either a single value or a 128-element array when the value is being
automated, and you pass it into the kernel so the change is applied per sample. For less critical
controls — a preset switch, a bypass toggle — a SharedArrayBuffer of control values written by the main
thread and read by the processor gives lock-free, allocation-free updates, using the same
Atomics discipline
you would use anywhere else.
Whichever you pick, smooth the parameter inside the kernel. Jumping a filter cutoff from 200 Hz to 2 kHz between quanta is audible as a click regardless of how the value arrived; a one-pole smoother over a few milliseconds costs two multiplies per sample and removes the artefact entirely.
Gotchas
AudioWorkletNodeconstructed beforeaddModuleresolves. Await the module registration first, or you get “Failed to construct ‘AudioWorkletNode’: unregistered name”.- The first second is silent. The instance was created from a message that arrived after the first
callbacks. Pass the bytes in
processorOptionsinstead. - Views come back zero-length. Something inside the kernel called the allocator and grew memory.
Preallocate in
init(), and make the kernel’s allocator a fixed arena so growth is impossible. - Works on desktop, glitches on mobile. The quantum is the same but the CPU is not. Measure the worst case on the slowest device you support, not the fastest.
process()returns false and the node dies. Returningfalsetells the browser this processor will never produce output again. Returntrueunless you mean it.
Performance note
A biquad cascade over 128 frames costs a few microseconds — the interesting cost is everything around it.
The two set() calls copy 512 bytes each and are negligible. What actually shows up in the worst case is
the browser’s own graph overhead plus scheduling jitter, which on a loaded machine can be a millisecond
on its own. That is why the target is under half the quantum: your kernel’s time is only part of what has
to fit inside it.
Frequently Asked Questions
Can I use threads inside an AudioWorklet?
No. The worklet scope has no Worker constructor, and spawning work elsewhere then waiting for it would
violate the deadline anyway. Parallelism in audio comes from splitting the graph across nodes, not from
threading a single kernel.
Is ScriptProcessorNode easier?
It is deprecated, it runs on the main thread, and it introduces latency measured in tens of milliseconds.
Anything new should use AudioWorklet; the extra setup is an hour once.
How big can the Wasm module be?
Synchronous new WebAssembly.Module() is allowed on this thread for reasonably sized buffers, and DSP
kernels are typically 10–80 kB. If your kernel is megabytes, compile it down or split the analysis work
into a regular worker and keep only the real-time path in the worklet.
Related
- Sharing memory between Wasm and Web Workers — the control-channel pattern in detail.
- Implementing a bump allocator in Wasm — building the arena this page depends on.
- Media Processing & Codecs in Wasm — the rest of the media toolkit.
← Back to Media Processing & Codecs in Wasm