Resizing Images Off the Main Thread
This guide answers one task: resize user-supplied images in the browser with a compiled resampler running in a worker, producing better quality than the platform’s default downscale without blocking the page or allocating a fresh heap for every file.
Prerequisites
- [ ] A resampler compiled to Wasm —
resizefrom the Rustimagecrate,pica’s kernel, or a small Lanczos implementation of your own. - [ ] A worker-capable build (
wasm-pack build --target webworks; so doesemcc -sENVIRONMENT=worker). - [ ]
OffscreenCanvassupport if you want the worker to produce aBlobdirectly — Chrome 69+, Firefox 105+, Safari 16.4+. - [ ] A handful of test images including one very wide panorama and one with an alpha channel.
Why not just use the canvas
ctx.drawImage(bitmap, 0, 0, w, h) resizes in one line, on the GPU, for free. It is the right answer
more often than people building Wasm resizers like to admit. The reasons to reach for a compiled
resampler are specific:
- Quality on large downscales. Browsers use a bilinear or trilinear filter, which aliases badly when the scale factor is below about 0.5. A Lanczos or Mitchell kernel keeps edges crisp where bilinear turns them into mush.
- Deterministic output. Two browsers produce visibly different pixels from the same
drawImage. If the resized image is going to be hashed, compared, or used as a fingerprint, that is fatal. - Colour and precision control. Resampling in linear light rather than sRGB avoids the darkening that plagues naive downscales of high-contrast content. The canvas will not do that for you.
- No canvas at all. In a worker without
OffscreenCanvas, or in a Node-side build sharing the same code, a pure-memory resampler is the only option.
If none of those apply, use the canvas and skip the download. If one does, the rest of this page is the implementation.
Separable passes and the arena
A two-dimensional resample is done as two one-dimensional passes — horizontal, then vertical — because a
separable kernel turns an O(k²) neighbourhood into two O(k) sweeps. That means three buffers in
linear memory: source, intermediate and destination. Size them once for the largest image you accept
and reuse them forever.
// lib.rs — arena sized at init, never reallocated afterwards
static mut SRC: Vec<u8> = Vec::new();
static mut TMP: Vec<f32> = Vec::new();
static mut DST: Vec<u8> = Vec::new();
#[no_mangle]
pub extern "C" fn init(max_w: usize, max_h: usize) {
unsafe {
SRC = vec![0; max_w * max_h * 4];
TMP = vec![0.0; max_w * max_h * 4];
DST = vec![0; max_w * max_h * 4];
}
}
#[no_mangle]
pub extern "C" fn resize(sw: usize, sh: usize, dw: usize, dh: usize) -> *const u8 {
unsafe {
horizontal_pass(&SRC[..sw * sh * 4], &mut TMP, sw, sh, dw); // sw → dw
vertical_pass(&TMP, &mut DST, dw, sh, dh); // sh → dh
DST.as_ptr()
}
}
The arena is what makes the tenth image as fast as the first. A version that allocates per call spends
its time in the allocator and triggers memory.grow, which detaches every view the worker holds —
covered in why memory.grow invalidates pointers.
Choose the maximum dimensions deliberately: 4096 × 4096 RGBA is 67 MB per buffer, and three of those is
already more than a phone will tolerate. Tiling, or a hard cap on accepted input, is the honest answer
for very large sources.
Wiring it into a worker
The worker owns the instance, receives a Blob or ImageBitmap, and returns a result the main thread can
use without touching pixels itself.
// resize-worker.js
import init, { resize, src_ptr, dst_ptr, initArena } from './resizer.js';
const mod = await init();
initArena(4096, 4096);
self.onmessage = async ({ data: { bitmap, dw, dh, id } }) => {
const { width: sw, height: sh } = bitmap;
const off = new OffscreenCanvas(sw, sh);
const ctx = off.getContext('2d', { willReadFrequently: true });
ctx.drawImage(bitmap, 0, 0);
const src = ctx.getImageData(0, 0, sw, sh).data;
new Uint8Array(mod.memory.buffer, src_ptr(), src.length).set(src);
resize(sw, sh, dw, dh);
const out = new Uint8ClampedArray(mod.memory.buffer, dst_ptr(), dw * dh * 4);
const outCanvas = new OffscreenCanvas(dw, dh);
outCanvas.getContext('2d').putImageData(new ImageData(out.slice(), dw, dh), 0, 0);
const blob = await outCanvas.convertToBlob({ type: 'image/jpeg', quality: 0.85 });
self.postMessage({ id, blob });
};
The single slice() in there is deliberate: putImageData needs a buffer it owns, and the view points
into linear memory which the next job will overwrite. It is the only copy in the pipeline beyond the
unavoidable one on the way in.
Expected output
The result should be dimensionally exact and visibly sharper than a canvas downscale of the same image. Check both:
const blob = await resizeInWorker(file, 800, 600);
console.log(blob.type, blob.size);
// image/jpeg 84213
const bmp = await createImageBitmap(blob);
console.log(bmp.width, bmp.height);
// 800 600
For a quality check that does not depend on your eyes, resize a synthetic test chart — concentric circles or a zone plate — and look for moiré. Bilinear produces obvious ring artefacts at small scales; a windowed sinc produces almost none. This is the kind of comparison worth capturing once and keeping as a reference image in the repository.
Resampling in linear light
The detail that separates a good resizer from a correct one is gamma. Image data is stored in sRGB, which is perceptually spaced rather than linear in intensity. Averaging two sRGB values is not the same as averaging the light they represent, and on high-contrast content — white text on black, a picket fence, a starfield — a naive average visibly darkens the result.
Doing it properly means converting each source sample to linear, resampling, and converting back:
#[inline] fn srgb_to_linear(c: f32) -> f32 {
if c <= 0.04045 { c / 12.92 } else { ((c + 0.055) / 1.055).powf(2.4) }
}
#[inline] fn linear_to_srgb(c: f32) -> f32 {
if c <= 0.0031308 { c * 12.92 } else { 1.055 * c.powf(1.0 / 2.4) - 0.055 }
}
Two powf calls per channel per sample is expensive enough to matter, so real implementations use a
256-entry lookup table for the forward direction and a slightly larger one for the reverse. With tables,
the correction costs under 10% and removes an artefact that is otherwise impossible to explain to a
designer. Alpha, if present, must be premultiplied before resampling and un-premultiplied afterwards, or
transparent pixels will bleed their colour into their neighbours.
Gotchas
getImageDatais slow without the hint. Pass{ willReadFrequently: true }when creating the context, or the browser keeps the surface on the GPU and every read stalls on a readback.- Off-by-one dimensions. Rounding the target size independently on each axis breaks the aspect ratio. Compute one scale factor and derive both dimensions from it.
- Colour shift on images with an ICC profile. The canvas applies the profile during decode; your resampler does not know about it. For colour-critical work, decode with the profile preserved and handle conversion explicitly.
- Memory climbing across many files. Usually a cached view rebuilt per job, or
ImageBitmapobjects that were neverclose()d. Both are easy to miss and easy to see in a heap snapshot. - The worker is idle but jobs queue up. One worker processes one image at a time; a drop of fifty
files needs a small pool, sized to
navigator.hardwareConcurrencyand no larger.
Performance note
Resampling a 12-megapixel photo down to 1600 px wide with a Lanczos-3 kernel takes roughly 180–300 ms
single-threaded on a modern laptop, against 8–15 ms for drawImage. Almost all of the difference is the
kernel width, not the fact that it is WebAssembly — the same algorithm in JavaScript takes 900 ms to
1.4 s. SIMD on the horizontal pass typically recovers another 2–2.5×, since four output pixels can share
one set of loads. The decode and encode around it are frequently larger than the resample itself, which
is why doing all three in the worker matters more than optimising any one of them.
Frequently Asked Questions
Should I resize before or after encoding? Always resize the decoded pixels, then encode once at the target size. Encoding at full size and letting the browser scale the result wastes both time and quality.
How do I generate several sizes from one source? Decode once, then run the resampler once per target from the same source buffer. The decode dominates, so producing four sizes costs barely more than producing one.
Is createImageBitmap with resizeQuality: 'high' good enough?
Often, yes — it is the cheapest quality upgrade available and needs no module at all. Compare it against
your resampler on your own content before committing to the extra payload.
What about animated images? Resize each frame through the same arena and re-encode the animation afterwards. The per-frame cost is small, but the encode is not — an animated WebP of a hundred frames is a genuinely long job and belongs behind explicit progress reporting rather than a spinner.
Can the worker read the file directly instead of receiving a bitmap?
Yes, and it is usually better: transfer the File object itself, call createImageBitmap(file) inside
the worker, and the main thread never touches image data at all. The only reason to decode on the main
thread is if you need to display the original immediately as well.
Related
- Building a Wasm image filter pipeline — chaining stages over the same arena.
- Loading Wasm in a Web Worker with ESM — the module-loading half of this setup.
- Autovectorizing loops for Wasm SIMD — making the passes faster.
← Back to Media Processing & Codecs in Wasm