Quantizing Models for Wasm Inference
This guide answers one task: convert a float32 model into an int8 model that downloads four times faster, uses a quarter of the memory, and runs faster on the WebAssembly backend — while proving how much accuracy you gave up rather than hoping it was fine.
Prerequisites
- [ ]
onnxruntimeandonnxruntime-toolsinstalled in a Python environment. - [ ] The float32 model, plus an evaluation set with labels — a few hundred examples is enough.
- [ ] A calibration set of 100–500 representative inputs if you use static quantization.
- [ ] A way to run the quantized model in the browser and compare against the reference.
What quantization actually changes
A float32 weight takes four bytes and covers an enormous dynamic range. An int8 weight takes one byte and covers 256 values, so the conversion stores a scale factor — and sometimes a zero point — alongside each group of weights and reconstructs an approximation at runtime.
The saving is not only size. Integer kernels move a quarter of the bytes through the cache, and on the
WebAssembly backend the i8x16 SIMD lanes process sixteen values per instruction where f32x4 processes
four. That is why an int8 model on this backend is commonly 1.5–3× faster as well as 4× smaller —
a rare case where the cheap option is also the fast one.
Dynamic quantization: start here
Dynamic quantization converts weights offline and computes activation scales at runtime. It needs no calibration data, it is one command, and for transformer-style models it is usually within a fraction of a percent of the float32 model.
python -m onnxruntime.quantization.preprocess --input model.onnx --output model-prep.onnx
python - <<'PY'
from onnxruntime.quantization import quantize_dynamic, QuantType
quantize_dynamic(
"model-prep.onnx", "model-int8.onnx",
weight_type=QuantType.QInt8,
per_channel=True, # almost always worth it
reduce_range=False, # set True only for old x86 targets
)
PY
The preprocessing step matters more than its obscurity suggests: it runs shape inference and folds constants, and skipping it leaves nodes that the quantizer cannot handle, producing a model that is “quantized” but still mostly float. Check the result rather than trusting the exit code — the file size tells you immediately whether it worked.
Static quantization when activations matter
Static quantization also quantizes activations, using ranges measured from calibration data. It is faster at runtime because nothing is computed per inference, and it is more fragile because the calibration set has to resemble real input.
from onnxruntime.quantization import quantize_static, CalibrationDataReader, QuantFormat
class Reader(CalibrationDataReader):
def __init__(self, samples): self.it = iter([{"pixel_values": s} for s in samples])
def get_next(self): return next(self.it, None)
quantize_static(
"model-prep.onnx", "model-int8-static.onnx",
Reader(calibration_samples),
quant_format=QuantFormat.QDQ,
per_channel=True,
)
Use between 100 and 500 calibration samples drawn from the same distribution as production input. Too few and the measured ranges are noisy; too many and you spend an hour for no further benefit. Include the awkward cases — very dark images, very short texts, silence — because a range measured only on typical input clips the tails and produces exactly the failures users report.
Measuring what you lost
Quantization without an evaluation is a guess. Run both models over the same held-out set and compare the metric you actually care about, not a proxy.
import numpy as np, onnxruntime as ort
ref = ort.InferenceSession("model.onnx")
q = ort.InferenceSession("model-int8.onnx")
agree = 0
for x, y in eval_set:
a = ref.run(None, {"pixel_values": x})[0].argmax()
b = q.run(None, {"pixel_values": x})[0].argmax()
agree += (a == b)
print(f"top-1 agreement: {agree / len(eval_set):.3%}")
Agreement with the float model is a more useful number than absolute accuracy, because it isolates the damage quantization did from whatever the model’s baseline errors already were. Below about 99% agreement on a classifier, investigate rather than ship: it usually means one layer has an unusually wide weight distribution and should be excluded from quantization.
Excluding the layers that hurt
Not every node benefits. The first convolution and the final classifier layer are frequently sensitive, and excluding them costs a few percent of the size saving while recovering most of the accuracy.
quantize_dynamic(
"model-prep.onnx", "model-int8.onnx",
weight_type=QuantType.QInt8,
per_channel=True,
nodes_to_exclude=["/classifier/Gemm", "/features/0/Conv"],
)
Finding the culprits is a bisection: quantize half the nodes, evaluate, and narrow down. In practice two
or three passes find them, and the node names come straight from the graph — Netron or
onnx.load().graph.node will list them.
Per-tensor versus per-channel scales
The choice of how finely you store scales is the single largest quality lever in the whole process, and it costs almost nothing to get right. A per-tensor scale stores one number for an entire weight tensor, so every output channel shares the same quantisation step. A per-channel scale stores one number per output channel, which is a few hundred extra floats in a file that just lost tens of megabytes.
Why it matters so much: convolution and linear layers routinely have channels whose weights span very different magnitudes. One channel might range over ±0.03 while its neighbour ranges over ±1.4. A single shared scale has to cover the widest channel, which means the narrow channel’s weights all collapse into a handful of the 256 available integer values — sometimes into two or three. That channel’s contribution becomes nearly random, and because it is one channel among hundreds, the damage shows up as a small accuracy drop rather than an obvious failure.
Per-channel quantisation gives each channel its own step size, so the narrow one gets fine resolution and the wide one gets coarse resolution, exactly as each needs. The measured difference on typical vision models is between 0.5 and 3 percentage points of top-1 accuracy — the difference between shipping and not shipping. There is no meaningful runtime cost, because the kernel folds the scale into the accumulation it was already doing.
Set per_channel=True unless you have measured that it is unsupported by your target backend. A handful
of older mobile execution paths only handle per-tensor, and the symptom is a model that loads and then
produces uniformly wrong results.
Keeping the pipeline reproducible
Quantization is a build step, and treating it as one prevents the situation where nobody can reproduce the model that is currently in production. Check the float model, the calibration set and the quantizer invocation into version control, and produce the quantized artifact from CI rather than from someone’s laptop.
# quantize.sh — the only supported way to produce a shipping model
set -euo pipefail
python -m onnxruntime.quantization.preprocess --input src/model.onnx --output build/model-prep.onnx
python tools/quantize.py --in build/model-prep.onnx --out build/model-int8.onnx --per-channel
python tools/evaluate.py --ref src/model.onnx --quant build/model-int8.onnx --min-agreement 0.99
sha256sum build/model-int8.onnx | tee build/model-int8.sha256
The evaluation step with a threshold is what turns this from a script into a gate: if a model change drops agreement below the bar, the build fails rather than shipping a quietly worse model. The hash goes into the filename you deploy, so the browser’s cache key changes exactly when the weights do — the same discipline as versioning any other artifact.
When quantization makes things slower
Three situations reliably produce a quantized model that is bigger, slower or both, and all three surprise people who expected a free win.
The first is a graph full of quantize and dequantize pairs. If the quantizer could not fuse them, every operator converts back and forth, and the conversion costs more than the integer arithmetic saves. Look at the node count before and after — a large increase is the symptom.
The second is int4 on a backend without native int4 kernels. The weights unpack to int8 or float before every multiply, so you pay unpacking on every inference to save download size once. That trade is correct for a model where the download dominates and wrong for one that runs continuously.
The third is a model dominated by operators that have no integer implementation at all — some attention variants, unusual normalisations. The quantized weights are stored small and converted to float on load, giving the download saving with none of the speed.
Gotchas
- The file barely shrank. The preprocessing step was skipped, so most nodes were not quantizable.
Run
quantize_dynamicon the preprocessed model, not the raw export. Quantization parameters are not specifiedat runtime. A QDQ model was produced without calibration for some tensor. Either supply calibration data or use dynamic quantization.- Accuracy collapses on a specific input class. Calibration data did not cover it. Widen the calibration set rather than tuning the quantizer.
- The browser refuses to load the model. Some quantized formats are not supported by every execution provider — notably on the GPU path. Test the quantized model on every backend in your fallback chain.
- Results differ between the Python check and the browser. Expected, within tolerance: the browser’s kernels accumulate differently. Compare argmax agreement rather than raw values.
Performance note
For a 92 MB vision model, per-channel dynamic int8 produced a 23 MB file, cut browser session creation from 620 ms to 190 ms, and reduced p50 inference on the threaded WebAssembly backend from 71 ms to 34 ms, with 99.6% top-1 agreement against the float model. The download saving is what users feel first; the inference saving is what makes the feature usable on a phone at all.
Frequently Asked Questions
Should I quantize before or after other graph optimisations? After. Run shape inference and constant folding first — that is what the preprocess step does — so the quantizer sees the simplest possible graph and can fuse more.
Is float16 a middle ground? On the WebAssembly backend, not usefully: there are no native float16 kernels, so values are widened to float32 on the fly. It halves the download, which is real, but you get none of the speed benefit that int8 brings.
Can I quantize a model I did not train? Yes, that is the normal case. You need an evaluation set representative of your use, not the original training data, and the agreement metric above tells you whether the result is acceptable.
Related
- Loading large model weights into linear memory — what to do when the file is still too big.
- Running ONNX models with onnxruntime-web — running the quantized result.
- Analyzing Wasm size with twiggy — the same size discipline applied to code.