Compressing Wasm with Brotli for Delivery
This guide answers one task: make a .wasm file transfer in a third of its bytes, by compressing it at
build time and serving the compressed variant correctly — without breaking streaming compilation or adding
CPU cost to every request.
Prerequisites
- [ ] A built
.wasmfile and a server or CDN whose configuration you control. - [ ] The
brotlicommand-line tool. - [ ] A build step where the compression can run.
- [ ]
curl, to verify what is actually served.
Why WebAssembly compresses so well
A .wasm file is a structured binary with a great deal of repetition: LEB128-encoded indices, repeated
opcode sequences, a name section full of similar identifiers, and data segments that are often text. That
is exactly what a dictionary-based compressor exploits.
Typical ratios are 30–40% for Brotli at maximum quality and 35–45% for gzip, meaning a 400 kB module transfers as roughly 140 kB. For a large module that is hundreds of kilobytes of network time removed for no runtime cost at all.
W=dist/engine.wasm
printf "raw %8d\n" "$(stat -c%s $W)"
printf "gzip %8d\n" "$(gzip -9 -c $W | wc -c)"
printf "brotli %8d\n" "$(brotli -q 11 -c $W | wc -c)"
raw 412688
gzip 161204
brotli 139872
Making the module more compressible
Compression works on what the compiler produced, and a few choices upstream change how much it can find.
Stripping the name section removes a block of text that compresses well but is still bytes nobody needs
in production — typically 15–25% of a module built with debug information. Running wasm-opt before
compressing helps twice: the output is smaller to begin with, and its more regular instruction sequences
compress slightly better than unoptimised ones.
Data segments are where the surprises are. A module embedding a large lookup table of random-looking values compresses poorly, because there is no redundancy to exploit; the same table generated at runtime from a small seed costs a few instructions and nothing in transfer. Conversely, a table of text or of mostly-zero values compresses to almost nothing, so moving it out of the binary gains little.
# see whether the data section is the problem
wasm-objdump -h dist/engine.wasm | grep -E 'Data|Code|Custom'
# check how each compresses by stripping and comparing
wasm-opt --strip-debug -o /tmp/stripped.wasm dist/engine.wasm
brotli -q 11 -c /tmp/stripped.wasm | wc -c
The general principle is that compression rewards the same things good engineering does — less duplication, less dead weight, fewer embedded artefacts nobody reads — so size work upstream and compression downstream compound rather than substituting for one another.
Precompress at build time
Brotli at quality 11 is slow — seconds for a large file — which makes it unsuitable for compressing on every request. Compressing once at build time gives the same result at no request-time cost.
# in the build
brotli -q 11 -f -o dist/engine.wasm.br dist/engine.wasm
gzip -9 -f -k dist/engine.wasm # produces engine.wasm.gz, for older clients
ls -l dist/engine.wasm*
Keep the uncompressed file too. A client that sends no Accept-Encoding, an intermediary that strips it,
and your own tooling all need it, and it costs only disk.
# nginx serves the precompressed variant when the client accepts it
brotli_static on;
gzip_static on;
location ~ \.wasm$ {
add_header Content-Type application/wasm;
add_header Cache-Control "public, max-age=31536000, immutable";
}
The _static directives look for engine.wasm.br and engine.wasm.gz next to the file and serve
whichever the client accepts, setting Content-Encoding appropriately. No compression happens at request
time.
Why on-the-fly compression is the wrong default
Most servers can compress dynamically, and for HTML that is the right choice because the content changes.
A .wasm file does not change between requests, so compressing it per request repeats identical work.
At quality 11 that work is substantial: compressing a 400 kB module takes 1–3 seconds of CPU. A server configured to do that dynamically either falls back to a much lower quality — losing most of the benefit — or becomes a bottleneck under load.
brotli -q 5 dist/engine.wasm → 152 kB in 0.09 s
brotli -q 11 dist/engine.wasm → 140 kB in 2.41 s
Quality 5 is a reasonable dynamic setting and still leaves 12 kB on the table per request, forever. Precompressing at 11 captures all of it once.
Compression and streaming compilation
A common concern is whether compression defeats instantiateStreaming. It does not.
The browser decompresses the response as it arrives and feeds the decompressed bytes to the WebAssembly compiler, so the two optimisations compose: the transfer is smaller and compilation overlaps it. There is no configuration needed and no reason to choose between them.
// works identically with or without Content-Encoding
const { instance } = await WebAssembly.instantiateStreaming(
fetch('/assets/engine.a91c3f.wasm'),
imports,
);
What does matter is the Content-Type, which must still be application/wasm on the compressed response.
Some server configurations set the type from the file extension and see .br, producing
application/x-brotli and a streaming failure. Set the type explicitly for the route rather than relying
on extension mapping.
Fitting it into the build
Compression belongs in the same step that produces the artifact, so the compressed variants can never be stale relative to the binary.
#!/usr/bin/env bash
# scripts/build.sh
set -euo pipefail
cargo build --profile web --target wasm32-unknown-unknown
wasm-opt -Oz --strip-debug -o dist/engine.wasm target/wasm32-unknown-unknown/web/engine.wasm
HASH=$(sha256sum dist/engine.wasm | cut -c1-8)
mv dist/engine.wasm "dist/engine.$HASH.wasm"
brotli -q 11 -f "dist/engine.$HASH.wasm"
gzip -9 -f -k "dist/engine.$HASH.wasm"
printf "built engine.%s.wasm raw %d br %d\n" "$HASH" "$(stat -c%s dist/engine.$HASH.wasm)" "$(stat -c%s dist/engine.$HASH.wasm.br)"
Hashing the filename in the same step is what makes the immutable cache header safe, and producing all
three variants together means a deploy that copies dist/ cannot ship a mismatched set.
If your deployment target compresses for you — several CDNs and object stores do, given the right configuration — verify what it produces rather than assuming. A CDN compressing at quality 5 when your build produced a quality 11 variant it then ignored is a common and invisible waste.
Verify what is actually served
Configuration you have not checked is configuration you do not have, and the checks are one command each.
curl -sI -H 'Accept-Encoding: br' https://example.com/assets/engine.a91c3f.wasm
# HTTP/2 200
# content-type: application/wasm
# content-encoding: br
# content-length: 139872
# cache-control: public, max-age=31536000, immutable
# and confirm the uncompressed variant still works
curl -sI https://example.com/assets/engine.a91c3f.wasm | grep -i 'content-encoding' || echo "no encoding — correct"
Add both to a post-deploy smoke test. A CDN change, a proxy addition or a framework upgrade can quietly stop serving the precompressed variant, and the symptom is a slower load that nobody attributes to anything.
Gotchas
- Compressing dynamically at quality 11. Seconds of CPU per request; precompress instead.
Content-Typefrom the.brextension. Breaks streaming compilation with an unhelpful error.- Precompressed files not deployed. The build produces them and the deploy step copies only
.wasm. - Double compression. An already-compressed response compressed again by an intermediary wastes CPU and can corrupt the encoding.
- Compressing an already-compressed payload. A model file or a media asset gains nothing; check before adding CPU.
- No uncompressed fallback. A client that does not advertise support gets nothing.
Performance note
For the 413 kB module above on a 50 Mbit connection: uncompressed transfer took 1.9 s, gzip 780 ms and Brotli 670 ms, with decompression adding under 15 ms in every case. Precompressing at build time added 2.4 s to the build and nothing to any request. On a slower mobile connection the difference between uncompressed and Brotli was over four seconds, which is the scenario that makes this worth doing rather than a marginal gain.
Frequently Asked Questions
Is Zstandard worth using instead?
Browser support for Content-Encoding: zstd is now broad, and it compresses comparably to Brotli much
faster. For precompressed static files the compression speed does not matter, so Brotli’s slightly better
ratio wins; for dynamic compression Zstandard is the better trade.
Does this apply to the glue JavaScript too?
Yes, and to everything else you serve. The .wasm is called out here because it is often the largest
single asset and because its content type interacts with streaming compilation.
Does compression help a cached module? Not on a repeat visit, where nothing is transferred at all. It helps every first visit, which for a public site is most visits.
Should I compress inside the application instead? Almost never. Transport compression is handled by the browser transparently; compressing in application code means decompressing in application code, which adds a copy and loses streaming.
Related
- Serving Wasm files with the right headers — the other delivery settings.
- Reducing Wasm bundle size with wasm-opt — making the uncompressed file smaller first.
- Catching size regressions in CI — tracking the compressed number.
Compression is the cheapest size win available and the one most often left switched off by accident.
← Back to Wasm Optimization Flags & Size Reduction