Resizing Images Off the Main Thread

This guide answers one task: resize user-supplied images in the browser with a compiled resampler running in a worker, producing better quality than the platform’s default downscale without blocking the page or allocating a fresh heap for every file.

Prerequisites

  • [ ] A resampler compiled to Wasm — resize from the Rust image crate, pica’s kernel, or a small Lanczos implementation of your own.
  • [ ] A worker-capable build (wasm-pack build --target web works; so does emcc -sENVIRONMENT=worker).
  • [ ] OffscreenCanvas support if you want the worker to produce a Blob directly — Chrome 69+, Firefox 105+, Safari 16.4+.
  • [ ] A handful of test images including one very wide panorama and one with an alpha channel.

Why not just use the canvas

ctx.drawImage(bitmap, 0, 0, w, h) resizes in one line, on the GPU, for free. It is the right answer more often than people building Wasm resizers like to admit. The reasons to reach for a compiled resampler are specific:

  • Quality on large downscales. Browsers use a bilinear or trilinear filter, which aliases badly when the scale factor is below about 0.5. A Lanczos or Mitchell kernel keeps edges crisp where bilinear turns them into mush.
  • Deterministic output. Two browsers produce visibly different pixels from the same drawImage. If the resized image is going to be hashed, compared, or used as a fingerprint, that is fatal.
  • Colour and precision control. Resampling in linear light rather than sRGB avoids the darkening that plagues naive downscales of high-contrast content. The canvas will not do that for you.
  • No canvas at all. In a worker without OffscreenCanvas, or in a Node-side build sharing the same code, a pure-memory resampler is the only option.

If none of those apply, use the canvas and skip the download. If one does, the rest of this page is the implementation.

Why the kernel matters below half scale A bilinear filter samples too few source pixels when the scale factor is small, so fine detail aliases into moiré. A windowed sinc kernel weights a wider neighbourhood and keeps edges intact at the cost of more arithmetic per output pixel. bilinear at 0.33× samples 2 × 2 source pixels ignores 5 of every 9 rows moiré on text and fabric ≈ 4 operations per output pixel Lanczos-3 at 0.33× samples an 18 × 18 neighbourhood every source pixel contributes clean edges, no aliasing ≈ 40 operations per output pixel Ten times the arithmetic on a tenth of the pixels: a 0.33× downscale of a 12-megapixel image is only 1.3 megapixels of output, so the total cost stays small. The expensive case is upscaling, where the output is larger than the input and the kernel width multiplies a bigger number.

Separable passes and the arena

A two-dimensional resample is done as two one-dimensional passes — horizontal, then vertical — because a separable kernel turns an O(k²) neighbourhood into two O(k) sweeps. That means three buffers in linear memory: source, intermediate and destination. Size them once for the largest image you accept and reuse them forever.

// lib.rs — arena sized at init, never reallocated afterwards
static mut SRC: Vec<u8> = Vec::new();
static mut TMP: Vec<f32> = Vec::new();
static mut DST: Vec<u8> = Vec::new();

#[no_mangle]
pub extern "C" fn init(max_w: usize, max_h: usize) {
    unsafe {
        SRC = vec![0; max_w * max_h * 4];
        TMP = vec![0.0; max_w * max_h * 4];
        DST = vec![0; max_w * max_h * 4];
    }
}

#[no_mangle]
pub extern "C" fn resize(sw: usize, sh: usize, dw: usize, dh: usize) -> *const u8 {
    unsafe {
        horizontal_pass(&SRC[..sw * sh * 4], &mut TMP, sw, sh, dw);   // sw → dw
        vertical_pass(&TMP, &mut DST, dw, sh, dh);                     // sh → dh
        DST.as_ptr()
    }
}

The arena is what makes the tenth image as fast as the first. A version that allocates per call spends its time in the allocator and triggers memory.grow, which detaches every view the worker holds — covered in why memory.grow invalidates pointers. Choose the maximum dimensions deliberately: 4096 × 4096 RGBA is 67 MB per buffer, and three of those is already more than a phone will tolerate. Tiling, or a hard cap on accepted input, is the honest answer for very large sources.

Wiring it into a worker

The worker owns the instance, receives a Blob or ImageBitmap, and returns a result the main thread can use without touching pixels itself.

// resize-worker.js
import init, { resize, src_ptr, dst_ptr, initArena } from './resizer.js';
const mod = await init();
initArena(4096, 4096);

self.onmessage = async ({ data: { bitmap, dw, dh, id } }) => {
  const { width: sw, height: sh } = bitmap;
  const off = new OffscreenCanvas(sw, sh);
  const ctx = off.getContext('2d', { willReadFrequently: true });
  ctx.drawImage(bitmap, 0, 0);
  const src = ctx.getImageData(0, 0, sw, sh).data;

  new Uint8Array(mod.memory.buffer, src_ptr(), src.length).set(src);
  resize(sw, sh, dw, dh);
  const out = new Uint8ClampedArray(mod.memory.buffer, dst_ptr(), dw * dh * 4);

  const outCanvas = new OffscreenCanvas(dw, dh);
  outCanvas.getContext('2d').putImageData(new ImageData(out.slice(), dw, dh), 0, 0);
  const blob = await outCanvas.convertToBlob({ type: 'image/jpeg', quality: 0.85 });
  self.postMessage({ id, blob });
};

The single slice() in there is deliberate: putImageData needs a buffer it owns, and the view points into linear memory which the next job will overwrite. It is the only copy in the pipeline beyond the unavoidable one on the way in.

Expected output

The result should be dimensionally exact and visibly sharper than a canvas downscale of the same image. Check both:

const blob = await resizeInWorker(file, 800, 600);
console.log(blob.type, blob.size);
// image/jpeg 84213

const bmp = await createImageBitmap(blob);
console.log(bmp.width, bmp.height);
// 800 600

For a quality check that does not depend on your eyes, resize a synthetic test chart — concentric circles or a zone plate — and look for moiré. Bilinear produces obvious ring artefacts at small scales; a windowed sinc produces almost none. This is the kind of comparison worth capturing once and keeping as a reference image in the repository.

Where each step of the resize runs Decoding happens on the browser's own threads via createImageBitmap, the resample happens inside the worker's linear memory, and encoding happens in the worker through OffscreenCanvas. The main thread only passes handles. main thread — handles only File → ImageBitmap transfer worker getImageData → arena two separable passes OffscreenCanvas.convertToBlob — encode here too Blob Encoding in the worker matters as much as resampling there: JPEG encoding a 12-megapixel image on the main thread is a visible freeze on its own. The main thread never sees a pixel array — only an ImageBitmap going out and a Blob coming back.

Resampling in linear light

The detail that separates a good resizer from a correct one is gamma. Image data is stored in sRGB, which is perceptually spaced rather than linear in intensity. Averaging two sRGB values is not the same as averaging the light they represent, and on high-contrast content — white text on black, a picket fence, a starfield — a naive average visibly darkens the result.

Doing it properly means converting each source sample to linear, resampling, and converting back:

#[inline] fn srgb_to_linear(c: f32) -> f32 {
    if c <= 0.04045 { c / 12.92 } else { ((c + 0.055) / 1.055).powf(2.4) }
}
#[inline] fn linear_to_srgb(c: f32) -> f32 {
    if c <= 0.0031308 { c * 12.92 } else { 1.055 * c.powf(1.0 / 2.4) - 0.055 }
}

Two powf calls per channel per sample is expensive enough to matter, so real implementations use a 256-entry lookup table for the forward direction and a slightly larger one for the reverse. With tables, the correction costs under 10% and removes an artefact that is otherwise impossible to explain to a designer. Alpha, if present, must be premultiplied before resampling and un-premultiplied afterwards, or transparent pixels will bleed their colour into their neighbours.

Where the resize should happen The arithmetic takes the same time in either place. What differs is whether the main thread is holding still while it happens. main thread 410 ms of frozen UI for a batch of twelve images worker + transfer 34 ms transfer cost only; the page keeps rendering throughout Transfer the buffer rather than copying it — an ArrayBuffer moves in microseconds regardless of size. One worker is usually enough; a pool helps only when the batch is large and the device has cores to spare.

Gotchas

  • getImageData is slow without the hint. Pass { willReadFrequently: true } when creating the context, or the browser keeps the surface on the GPU and every read stalls on a readback.
  • Off-by-one dimensions. Rounding the target size independently on each axis breaks the aspect ratio. Compute one scale factor and derive both dimensions from it.
  • Colour shift on images with an ICC profile. The canvas applies the profile during decode; your resampler does not know about it. For colour-critical work, decode with the profile preserved and handle conversion explicitly.
  • Memory climbing across many files. Usually a cached view rebuilt per job, or ImageBitmap objects that were never close()d. Both are easy to miss and easy to see in a heap snapshot.
  • The worker is idle but jobs queue up. One worker processes one image at a time; a drop of fifty files needs a small pool, sized to navigator.hardwareConcurrency and no larger.

Performance note

Resampling a 12-megapixel photo down to 1600 px wide with a Lanczos-3 kernel takes roughly 180–300 ms single-threaded on a modern laptop, against 8–15 ms for drawImage. Almost all of the difference is the kernel width, not the fact that it is WebAssembly — the same algorithm in JavaScript takes 900 ms to 1.4 s. SIMD on the horizontal pass typically recovers another 2–2.5×, since four output pixels can share one set of loads. The decode and encode around it are frequently larger than the resample itself, which is why doing all three in the worker matters more than optimising any one of them.

Frequently Asked Questions

Should I resize before or after encoding? Always resize the decoded pixels, then encode once at the target size. Encoding at full size and letting the browser scale the result wastes both time and quality.

How do I generate several sizes from one source? Decode once, then run the resampler once per target from the same source buffer. The decode dominates, so producing four sizes costs barely more than producing one.

Is createImageBitmap with resizeQuality: 'high' good enough? Often, yes — it is the cheapest quality upgrade available and needs no module at all. Compare it against your resampler on your own content before committing to the extra payload.

What about animated images? Resize each frame through the same arena and re-encode the animation afterwards. The per-frame cost is small, but the encode is not — an animated WebP of a hundred frames is a genuinely long job and belongs behind explicit progress reporting rather than a spinner.

Can the worker read the file directly instead of receiving a bitmap? Yes, and it is usually better: transfer the File object itself, call createImageBitmap(file) inside the worker, and the main thread never touches image data at all. The only reason to decode on the main thread is if you need to display the original immediately as well.

← Back to Media Processing & Codecs in Wasm