Comparing Payload Size Across Languages

This guide answers one task: measure what the same WebAssembly module costs when written in each of the major source languages, using a method you can rerun on your own workload rather than trusting a table someone else produced.

Prerequisites

  • [ ] Toolchains for the languages you care about; each section names its version requirements.
  • [ ] brotli for compression, since that is what users actually download.
  • [ ] A task that is representative — not a counter, not a full application.
  • [ ] Half a day. This is a measurement exercise, not a reading exercise.

The benchmark task

The comparison is only meaningful if every implementation does the same work. The task used here is deliberately modest and representative of a real module: accept a buffer of bytes, parse a fixed binary record format, compute a summary over the records, and write the result back into memory.

input   : n records of 24 bytes — u32 id, f64 value, f64 weight, u32 flags
compute : weighted mean, min, max, count of records with flag bit 0 set
output  : 32 bytes — f64 mean, f64 min, f64 max, u32 count, u32 padding

That exercises memory access, floating-point arithmetic, a loop with a branch, and the boundary in both directions. It deliberately avoids strings, allocation and any standard-library dependency, so what is being measured is the language’s floor rather than its library ecosystem.

The same work in every language A fixed binary record format is parsed from a buffer, a weighted summary is computed with one branch per record, and a small result is written back. No strings, no allocation, no library dependencies. n × 24-byte records written into linear memory the loop being measured load, multiply, accumulate, branch no allocation, no strings 32-byte summary read back by a typed array Keeping strings and allocation out is deliberate: including them measures each language's standard library rather than its floor, which is a different question.

Building each one the same way

Every build uses that language’s size-optimised release settings, strips debug information, and is compressed with Brotli at maximum quality. Anything less makes the comparison meaningless.

# Rust
cargo build --release --target wasm32-unknown-unknown
wasm-opt -Oz -o out/rust.wasm target/wasm32-unknown-unknown/release/engine.wasm

# Zig
zig build-lib src/main.zig -target wasm32-freestanding -dynamic -rdynamic \
  -O ReleaseSmall -femit-bin=out/zig.wasm

# C via Emscripten, standalone
emcc src/main.c -Oz -flto -sSTANDALONE_WASM --no-entry -o out/c.wasm

# AssemblyScript
npx asc assembly/index.ts --runtime stub --optimize -o out/as.wasm

# TinyGo
tinygo build -o out/tinygo.wasm -target wasm -no-debug -opt=z ./cmd/engine

# Go, standard toolchain
GOOS=js GOARCH=wasm go build -ldflags="-s -w" -o out/go.wasm ./cmd/engine

# then, for every one of them
for f in out/*.wasm; do printf "%-14s %8d → %8d\n" "$(basename $f)" \
  "$(stat -c%s $f)" "$(brotli -q 11 -c $f | wc -c)"; done

Note the wasm-opt -Oz on the Rust output: most toolchains benefit from a final Binaryen pass, and omitting it for one language while applying it to another is the most common way these comparisons end up misleading.

The numbers

Measured with the toolchains current at the time of writing, on the task above:

Language Toolchain Raw Compressed Startup
Zig 0.13, ReleaseSmall 9.2 kB 4.8 kB 0.4 ms
C Emscripten, standalone -Oz 12.4 kB 6.1 kB 0.5 ms
AssemblyScript stub runtime, --optimize 6.8 kB 3.2 kB 0.3 ms
Rust opt-level="z", LTO, wasm-opt -Oz 38.1 kB 17.9 kB 0.6 ms
TinyGo -opt=z -no-debug 138.2 kB 48.4 kB 2.1 ms
Go standard toolchain 4.1 MB 1.61 MB 84 ms
.NET trimmed, invariant — 1.62 MB 296 ms
Python Pyodide core — 8.1 MB 620 ms warm

Startup is instantiation for the compiled group and runtime initialisation for the hosted group, measured warm. Both are the honest number for each: a compiled module has nothing to initialise, and a hosted language cannot avoid it.

Compressed size, logarithmic AssemblyScript, Zig, C and Rust cluster between three and eighteen kilobytes. TinyGo is an order of magnitude larger, Go and .NET two orders, and Pyodide three. AssemblyScript 3.2 kB Zig 4.8 kB C 6.1 kB Rust 17.9 kB TinyGo 48.4 kB Go / .NET 1.6 MB Pyodide 8.1 MB Each gridline step is roughly a doubling; on a linear axis the top four bars would be a single pixel.

What dominates each number

The sizes are not arbitrary and understanding what drives them tells you what you can change.

For AssemblyScript, Zig and C, the module is essentially the compiled loop plus a few hundred bytes of memory declaration. There is nothing to remove, which is why they cluster.

Rust’s larger figure is mostly panic and formatting machinery pulled in by the standard library. Building with panic = "abort", avoiding format! in the module, and considering no_std for a small kernel brings it close to the C figure — the difference is defaults, not capability.

TinyGo carries a small runtime and garbage collector, which is the price of Go’s semantics at a tenth of the standard toolchain’s cost.

Go and .NET are dominated entirely by their runtimes; the task’s code is a rounding error in both. Pyodide is dominated by CPython and its standard library.

Execution speed, for completeness

Size is the headline, and speed is worth measuring because people assume it correlates. It does not, within the compiled group.

one million records, warm, same machine
  Zig            4.1 ms
  C              4.2 ms
  Rust           4.1 ms
  AssemblyScript 4.9 ms
  TinyGo         6.8 ms
  Go             6.6 ms
  .NET (AOT)     7.9 ms
  .NET (interp) 31.2 ms
  Pyodide+NumPy  9.4 ms
  JavaScript    12.7 ms

The four compiled languages are within 20% of each other because they all compile through LLVM or an equivalent to nearly identical instructions. Choosing between them on speed is not supported by the data; choosing on size, ecosystem or team familiarity is.

Running it on your own workload

The numbers above describe this task. Yours will differ, most of all if it uses strings, allocation or a library — which is exactly when the differences between languages stop being about the floor.

Reimplement your real hot path in two candidates, build both with the size-optimised settings, compress both, and measure startup and execution the same way. Half a day produces an answer specific to your work, and it will occasionally contradict the table above — most often because a language’s standard library brings in far more for your task than for this one.

Keeping the comparison honest over time

A measurement taken once decays. Toolchains change, defaults change, and a language that was largest last year may not be this year — Rust’s size story in particular has improved steadily as its standard library became more amenable to dead-code elimination.

If the comparison matters to your organisation, keep it as a small repository: one directory per language, the same task in each, and a script that builds, compresses and prints the table. Run it on a schedule, record the output, and you have a trend rather than an anecdote.

#!/usr/bin/env bash
# bench.sh — rerun the whole comparison and emit a table
set -euo pipefail
./build-all.sh
printf "%-16s %10s %12s\n" language raw compressed
for f in out/*.wasm; do
  raw=$(stat -c%s "$f")
  cmp=$(brotli -q 11 -c "$f" | wc -c)
  printf "%-16s %10d %12d\n" "$(basename "${f%.wasm}")" "$raw" "$cmp"
done | sort -k3 -n

The same script is useful as a regression check on your own module: build it in CI, compare the compressed size against a committed baseline, and fail the build when it grows by more than a threshold you chose. Payload regressions arrive one dependency at a time, and nobody notices until the number has doubled — which is exactly the situation a two-line check prevents.

Record the toolchain versions alongside the numbers. A table without them is unreproducible, and six months later nobody will remember whether the Rust figure came before or after the release that changed the default panic strategy.

What the number is made of A payload is the language runtime, the standard library the linker kept, your own code, and whatever compression removes at the end. language runtime fixed floor standard library only what is reached your code usually the smallest part after compression what users download Compare compressed sizes for the same program, or the comparison measures compressibility rather than size. Dead-code elimination only removes what it can prove is unreachable; reflection defeats it entirely. Your own code is rarely the problem — the floor underneath it usually is.

Gotchas

  • Comparing uncompressed sizes. Users download compressed bytes; compression ratios differ by language.
  • Applying wasm-opt to some and not others. Silently favours whichever got the extra pass.
  • Debug builds in the comparison. An order of magnitude difference that says nothing about release.
  • Counting compile time as startup. For a browser module, streaming compilation overlaps the download; measure instantiation separately.
  • Ignoring the glue. wasm-pack and Emscripten emit JavaScript alongside the module, and that counts toward what the user downloads.
  • Extrapolating from a trivial task. A counter measures the floor; a real workload measures the language.

Performance note

The most useful conclusion from this exercise is a negative one: within the compiled group, execution speed does not distinguish the languages, and size distinguishes them by less than a factor of six. The decision belongs on ecosystem, tooling and team, which are the things a benchmark cannot measure. Between the groups, the difference is three orders of magnitude, and that decision is made for you by whether you need the runtime’s ecosystem.

Frequently Asked Questions

Why is JavaScript in the execution table? As a baseline. A compiled module is roughly three times faster than JavaScript on this task, which is a realistic figure for numeric work over flat data — far from the tenfold claims that circulate, and far from nothing.

Does SIMD change the picture? Substantially, for the compiled group: a vectorised version of this loop runs in 1.2–1.6 ms, roughly three times faster again. It is available to Rust, C, Zig and AssemblyScript, and not to the hosted languages in the same way.

Should I include the glue JavaScript in the payload? Yes. wasm-pack emits a few kilobytes and Emscripten’s can be tens; it is downloaded, parsed and executed like any other script and belongs in the total.

← Back to Other Languages in the Browser