Catching Size Regressions in CI
This guide answers one task: make a pipeline fail when a WebAssembly module grows more than it should, so payload increases are a decision someone made rather than something discovered six months later.
Prerequisites
- [ ] A reproducible release build that produces the same bytes from the same commit.
- [ ]
brotli, because compressed size is what users download. - [ ] A place to store a baseline — a committed file is the simplest.
- [ ] A tolerance you have agreed on, rather than one invented during an incident.
Measure what the user downloads
Three numbers exist and only one of them matters for a gate. The raw .wasm size is what the build
produces. The compressed size is what crosses the network. The in-memory size is neither, and is not what
this check is about.
Gate on the compressed size, computed the same way every time, and record the raw size alongside it for diagnosis — a change that leaves compressed size flat while raw size grows usually means more of something highly compressible, such as repeated error strings.
RAW=$(stat -c%s dist/engine.wasm)
GZ=$(brotli -q 11 -c dist/engine.wasm | wc -c)
echo "raw=$RAW compressed=$GZ"
Include the glue JavaScript if your users download it. A module that shrinks while its generated bindings grow has not improved anything, and a gate that measures only the binary will report success.
A baseline and a tolerance
The baseline is the size at the last accepted state; the tolerance is how much growth is allowed without a conversation. A committed file is the simplest store and has the useful property that updating it appears in the diff, so growth is reviewed rather than absorbed.
# .size-baseline
wasm_compressed=53248
glue_compressed=11776
#!/usr/bin/env bash
# scripts/check-size.sh
set -euo pipefail
source .size-baseline
TOL=${SIZE_TOLERANCE:-0.03} # 3% by default
wasm_now=$(brotli -q 11 -c dist/engine_bg.wasm | wc -c)
glue_now=$(brotli -q 11 -c dist/engine.js | wc -c)
check() {
local name=$1 now=$2 base=$3
awk -v n="$now" -v b="$base" -v t="$TOL" -v name="$name" 'BEGIN {
delta = (n - b) / b
printf "%-16s %8d → %8d (%+.2f%%)\n", name, b, n, delta * 100
if (delta > t) { printf " regression: %s exceeds %.0f%% tolerance\n", name, t * 100; exit 1 }
}'
}
check wasm "$wasm_now" "$wasm_compressed"
check glue "$glue_now" "$glue_compressed"
A three percent tolerance is a reasonable default: large enough that a compiler update does not fail the build, small enough that a real addition is visible. Make it configurable so a deliberate large change can pass with an explicit override and a note in the commit message.
Reporting rather than only failing
A gate that fails is useful; a gate that reports on every change is better, because it makes size a thing people see rather than a thing they trip over.
# emit a markdown table for a pull request comment
{
echo "| artifact | baseline | this build | change |"
echo "| --- | ---: | ---: | ---: |"
printf "| engine.wasm | %d | %d | %+.2f%% |\n" \
"$wasm_compressed" "$wasm_now" \
"$(awk -v n=$wasm_now -v b=$wasm_compressed 'BEGIN{print (n-b)/b*100}')"
} > size-report.md
| artifact | baseline | this build | change |
| ----------- | -------: | ---------: | ------: |
| engine.wasm | 53248 | 54912 | +3.13% |
| engine.js | 11776 | 11776 | +0.00% |
Posting that on every pull request changes behaviour. A developer who sees “+3.13%” next to their change asks whether the dependency they added was necessary, which is a conversation that does not happen when the number is invisible.
Finding what grew
When the gate fires, the question is which change caused it, and two tools answer it quickly.
twiggy attributes bytes to functions and data, and its diff mode compares two builds directly:
twiggy top -n 20 dist/engine.wasm
twiggy diff baseline.wasm dist/engine.wasm -n 20
Delta Bytes │ Item
─────────────┼───────────────────────────────────────────
+8912 │ core::fmt::Formatter::pad
+3104 │ ::fmt
+1288 │ data[3]
-412 │ engine::parse_records
That output is typical and immediately diagnostic: a formatting or error-display path was pulled in, which
usually means a format! or a Display implementation reached code that had previously avoided it.
Bisecting works when the attribution is unclear:
git bisect start HEAD HEAD~30
git bisect run sh -c 'cargo build --release --target wasm32-unknown-unknown &&
test $(brotli -q 11 -c target/wasm32-unknown-unknown/release/engine.wasm | wc -c) -lt 54000'
The changes that grow a module
Knowing the usual culprits makes a regression report much faster to act on, because most growth comes from a small set of causes.
Formatting and error display are the most common in Rust. A single format! in a code path that previously
had none can pull in tens of kilobytes of core::fmt machinery, and a Display implementation on an
error type does the same. Returning error codes rather than formatted messages from the module, and
formatting on the JavaScript side, keeps that machinery out entirely.
A new dependency is the second. Crates vary enormously in how much they bring: a small numeric crate may add a kilobyte, while one with a builder API and rich errors can add fifty. Check the delta before and after adding one rather than after a month of accumulated changes.
Panic unwinding is the third. Without panic = "abort" the binary carries landing pads and unwinding
tables throughout, which for a mid-sized module is commonly 15–25% of it.
Debug information is the fourth and the easiest to fix: a name or .debug_* custom section left in a
release build inflates it substantially and does nothing for users. wasm-opt --strip-debug or the
toolchain’s own strip flag removes it.
# see what a candidate dependency costs before committing to it
brotli -q 11 -c dist/engine.wasm | wc -c # before
cargo add some-crate && cargo build --release --target wasm32-unknown-unknown
brotli -q 11 -c dist/engine.wasm | wc -c # after
Running that two-line comparison before adding a dependency takes a minute and has, on more than one project, been the reason a convenient crate was replaced with thirty lines of hand-written code.
Updating the baseline honestly
The baseline should move when growth is accepted, and moving it should be visible.
Update it in the same commit as the change that grew the module, with the reason in the message. Do not regenerate it automatically on the main branch — an auto-updating baseline ratchets upward one accepted change at a time and reports no regression ever, which is the failure mode the dashed line in the figure above describes.
# accepted a deliberate increase
./scripts/update-baseline.sh
git add .size-baseline
git commit -m "Add AVIF decode path (+18 kB compressed, accepted)"
A baseline with a history is also a useful artifact in itself: git log -p .size-baseline reads as a
record of every deliberate payload decision the project has made.
Expected output
A passing run is quiet and a failing one is specific:
./scripts/check-size.sh
wasm 53248 → 53696 (+0.84%)
glue 11776 → 11776 (+0.00%)
artifact ok
./scripts/check-size.sh
wasm 53248 → 61440 (+15.38%)
regression: wasm exceeds 3% tolerance
Gotchas
- Gating on raw size. Compression ratios vary; the user pays for compressed bytes.
- Ignoring the glue. A module that shrinks while its bindings grow has not improved.
- Auto-updating the baseline. Guarantees the gate never fires.
- A tolerance set during an incident. Ends up at 25% and stops meaning anything.
- Different
brotliversions between local and CI. Produces different numbers; pin it. - Non-reproducible builds. Makes every comparison noisy and the gate untrustworthy.
Performance note
The whole check — build, compress, compare — adds about 3 s to a pipeline for a 150 kB module, most of it
the maximum-quality compression. Running brotli -q 5 instead is ten times faster and gives a number
within a few percent, which is fine for a relative comparison provided the baseline was produced the same
way. For the number you quote publicly, use quality 11.
Frequently Asked Questions
What tolerance should I use? Three percent for an established module, wider while a project is young and changing shape. The important part is that it is agreed in advance rather than adjusted to make a build pass.
Should the gate block a merge? Yes, with an easy override. A warning that does not block gets ignored within a month; a block with a documented way to accept the growth produces exactly the conversation the gate exists to start.
Where should the report be published? Wherever the team already looks — a pull request comment is usually enough, and a long-run chart is a bonus rather than a requirement.
Does this work for a multi-artifact build? Yes — extend the baseline file with one entry per artifact and loop. A build producing baseline and SIMD variants should track both, since they grow for different reasons.
Related
- Analyzing Wasm size with twiggy — attributing the growth.
- Validating binaries with wasm-validate — the other artifact check.
- Shrinking Rust Wasm with cargo profiles — getting the number down once it is up.
The gate is five minutes of setup and it is the difference between knowing your payload and guessing it.
← Back to Testing & Verifying Wasm Builds