Running Wasm Tests in Headless Browsers
This guide answers one task: run a WebAssembly test suite in a continuous integration pipeline, across more than one browser engine, fast enough and reliably enough that people trust the result.
Prerequisites
- [ ] A browser test suite that passes locally — see unit testing with wasm-bindgen-test.
- [ ] A CI provider where you can install packages and cache directories.
- [ ] Ten minutes of tolerance for the first pipeline run, which will fail.
- [ ] A decision about which engines you actually support.
Which browsers, and why
Testing everywhere is expensive and mostly redundant. WebAssembly semantics are specified and the engines agree; what differs is the platform around them — the APIs your glue uses, the timing, and occasional quirks in newer features.
A defensible default is two engines in the pipeline on every change and a third on a schedule. Chromium covers the largest share of users and has the best tooling. Firefox catches the genuine engine differences, particularly around newer proposals. WebKit catches Safari-specific platform behaviour, which is where the real surprises live, and is the most awkward to run in CI — which is why it often belongs on a nightly job rather than on every push.
A pipeline that works
The essential parts are: install the Rust target, install the browsers, restore a cache so neither is repeated, and run the suite per engine in parallel.
# .github/workflows/test.yml
name: test
on: [push, pull_request]
jobs:
native:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
- run: cargo test --all-features
browser:
runs-on: ubuntu-latest
needs: native # do not pay for browsers if the logic is broken
strategy:
fail-fast: false
matrix:
browser: [chrome, firefox]
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
with: { targets: wasm32-unknown-unknown }
- uses: Swatinem/rust-cache@v2
- uses: jetli/wasm-pack-action@v0.4
- run: wasm-pack test --headless --${{ matrix.browser }}
Two decisions in that file matter more than the rest. needs: native means the expensive job never runs
for a change that fails a three-second test, which over a month saves a great deal of pipeline time. And
fail-fast: false means a Firefox failure does not cancel the Chromium run, so you see both results and
can tell whether the problem is engine-specific.
Caching the right things
A cold run installs a Rust toolchain, downloads crates, compiles dependencies and provisions a browser. Caching turns a six-minute job into a ninety-second one.
Cache the cargo registry and the build directory keyed on the lockfile — Swatinem/rust-cache does this
correctly and handles the target directory’s quirks. Cache the wasm-pack and wasm-bindgen-cli
binaries, which are otherwise downloaded every run. Browsers are usually preinstalled on hosted runners;
where they are not, pin the version and cache the download.
- uses: actions/cache@v4
with:
path: |
~/.cargo/bin/wasm-pack
~/.cargo/bin/wasm-bindgen
key: wasm-tools-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
Pin the browser version on any job whose results you compare over time. A silent browser upgrade that changes behaviour is indistinguishable from a regression in your code, and finding that out takes a day.
Flakiness, and how to keep it out
Browser tests fail intermittently for reasons that have nothing to do with the code, and a suite with a 2% flake rate is one people stop believing.
The usual causes are all avoidable. A test that waits a fixed number of milliseconds for something
asynchronous will eventually wait too little on a loaded runner — wait for a condition instead. A test that
depends on another test’s leftover state fails when the order changes. A test that fetches from the network
fails when the network does. And a test that races a requestAnimationFrame or a timer against an
assertion fails under load.
// bad: a fixed delay
#[wasm_bindgen_test]
async fn renders_eventually_bad() {
start_render();
sleep_ms(50).await; // hope
assert!(is_rendered());
}
// good: wait for the condition, with a bounded timeout
#[wasm_bindgen_test]
async fn renders_eventually() {
start_render();
wait_until(|| is_rendered(), 2000).await.expect("render did not complete");
assert!(is_rendered());
}
When a test does flake, quarantine it rather than retrying the whole job. A retry hides the signal and doubles the cost; moving the test to a separate job that does not block a merge keeps the main suite trustworthy while you fix it.
Getting useful output from a failure
A CI failure you cannot diagnose from the log is a failure you will reproduce locally anyway, which wastes the run. Three things make the log sufficient most of the time.
Install the panic hook so a Rust panic reports its message and location rather than an address. Capture the
browser console — wasm-pack forwards it, but a custom harness may not — because that is where
instantiation errors and unhandled rejections appear. And on failure, upload the built artifacts and any
screenshots as job artifacts, so the exact binary that failed is available rather than one you rebuild
later from the same commit and slightly different tooling.
- name: upload artifacts on failure
if: failure()
uses: actions/upload-artifact@v4
with:
name: wasm-${{ matrix.browser }}
path: |
pkg/
target/wasm32-unknown-unknown/debug/*.wasm
Running the same suite locally
A pipeline nobody can reproduce locally becomes a remote debugging exercise. Two small investments keep the local and CI paths close enough that a CI failure is reproducible in one command.
Put the commands in a script rather than in the pipeline configuration, and have the pipeline call the script. Then the same entry point runs on a laptop and on a runner, and a change to how tests run happens in one place.
#!/usr/bin/env bash
# scripts/test.sh
set -euo pipefail
cargo test --all-features
for browser in "$@"; do
wasm-pack test --headless "--$browser"
done
./scripts/check-artifact.sh
./scripts/test.sh chrome # locally, one engine
./scripts/test.sh chrome firefox # what CI runs
Pin the tool versions the script depends on in a file the script reads, so a developer with a different
wasm-bindgen does not get a different result. A version mismatch between a laptop and a runner produces
the most frustrating category of CI failure: one that does not reproduce, for a reason nobody thinks to
check.
Where a failure genuinely happens only in CI, the fastest route is usually to reproduce the runner rather than the test — run the same commands inside the same container image locally, which most providers publish. That turns a guessing game into an ordinary debugging session.
Expected output
A healthy pipeline run looks like this, and the timings are the thing to watch over time:
native ✓ 3.1s
browser (chrome) ✓ 41.8s
browser (firefox) ✓ 52.4s
artifact checks ✓ 2.2s
total (parallel) ✓ 58.9s
# and a failure that is diagnosable from the log alone
browser (firefox) ✗ 39.2s
test browser::detects_simd ... FAILED
panicked at 'assertion failed: simd_supported()', tests/browser.rs:48:5
Stack:
__rust_start_panic
my_crate::tests::detects_simd
Gotchas
- Browser not installed on the runner. Hosted runners usually have Chrome, rarely Firefox, almost never WebKit. Install explicitly rather than assuming.
wasm-bindgen-cliversion drift. Must match thewasm-bindgencrate exactly; letwasm-packmanage it or pin both.- No cache. A six-minute job that should take ninety seconds, on every push.
fail-fastcancelling the other engine. You lose the information that distinguishes an engine-specific failure from a real one.- Retrying flaky tests automatically. Hides the problem and doubles the cost.
- Running the browser suite before the native one. Pays for browsers to tell you something a three-second job already knew.
Performance note
For a mid-sized crate: native tests 3.1 s, browser suite 42 s on Chromium and 52 s on Firefox, artifact checks 2.2 s, with the browser jobs running in parallel for a wall-clock total under a minute. Without caching the same pipeline took 5 minutes 40 seconds, almost all of it recompiling dependencies — which makes the cache configuration the highest-value ten lines in the file.
Frequently Asked Questions
Should I use Playwright instead of the built-in runner? For end-to-end tests of a page, yes — Playwright drives a real application and handles multiple browsers well. For unit tests of a crate, the built-in runner is simpler and faster. Most projects end up with both, covering different layers.
How do I test on real Safari rather than WebKit? Only on macOS runners, which are slower and more expensive. Running WebKit on Linux catches most of the engine differences; a scheduled macOS job before a release covers the rest.
Can I run the suite against a browser version I pin? Yes, and you should for any comparison over time. Install a specific version rather than using whatever the runner image ships, or a browser upgrade will read as a code regression.
Related
- Unit testing Rust Wasm with wasm-bindgen-test — writing the tests this runs.
- Setting up CI/CD for Rust Wasm projects — the wider pipeline.
- Catching size regressions in CI — the artifact job alongside these.
← Back to Testing & Verifying Wasm Builds