Cold Start Characteristics of Server-Side Wasm

This guide answers one task: find out what “cold start” actually costs for your module, by measuring compilation, load and instantiation separately rather than as one opaque number — and then reduce the ones that matter.

Prerequisites

  • [ ] Wasmtime 24+ installed, or another runtime you can embed.
  • [ ] A module representative of your workload, not a hello-world.
  • [ ] A machine that is not doing anything else while you measure.
  • [ ] A clear question: which of the three costs is on your request path?

Three costs, often confused

Compilation turns the binary into machine code. It is proportional to module size and is the most expensive step by far — tens to hundreds of milliseconds for a real module. On managed platforms it happens at deploy time and never appears in a request.

Module load maps an already-compiled artifact into a process. With a pre-compiled artifact this is close to a file map and takes microseconds to low milliseconds, depending on size and whether the page cache is warm.

Instantiation creates an instance from a loaded module: allocating memory, initialising data segments, running the start function. For a small module it is tens of microseconds; for one with large data segments or a big initial memory, it can be milliseconds.

Three costs, three different orders of magnitude Compilation dominates and happens at deploy time on managed platforms. Loading a precompiled artifact is fast. Instantiation is the only cost on a per-request path for most deployments. compilation — 40–400 ms, at deploy time load — 0.2–6 ms inst. instantiation — 20–800 µs, per request Quoting the top bar as "cold start" describes a cost most deployments never pay on a request, and hides the one they pay on every request.

Measure them separately

Wasmtime’s API separates the steps cleanly, so timing each is a few lines. Run each many times and report percentiles, not a single sample.

use std::time::Instant;
use wasmtime::*;

fn main() -> Result<()> {
    let engine = Engine::default();
    let bytes = std::fs::read("service.wasm")?;

    let t0 = Instant::now();
    let module = Module::new(&engine, &bytes)?;              // compilation
    let compile = t0.elapsed();

    let mut linker = Linker::new(&engine);
    wasmtime_wasi::add_to_linker_sync(&mut linker, |s| s)?;

    let mut inst_times = Vec::new();
    for _ in 0..1000 {
        let t = Instant::now();
        let mut store = Store::new(&engine, wasi_ctx());
        let _instance = linker.instantiate(&mut store, &module)?;   // instantiation only
        inst_times.push(t.elapsed());
    }
    inst_times.sort();
    println!("compile {:?}  inst p50 {:?}  p99 {:?}",
             compile, inst_times[500], inst_times[990]);
    Ok(())
}

Creating a fresh Store each iteration matters: reusing one measures something else entirely, because the memory allocation is the bulk of what instantiation does.

Pre-compile, and stop paying for compilation

Wasmtime can serialise a compiled module to a file and load it later without recompiling. That turns the largest cost into a build-time step, which is exactly what managed platforms do internally.

# at build time
wasmtime compile service.wasm -o service.cwasm

# at run time, loading is a map rather than a compile
wasmtime run --allow-precompiled service.cwasm
// embedding: load the precompiled artifact
let module = unsafe { Module::deserialize_file(&engine, "service.cwasm")? };

The artifact is specific to the runtime version and the target architecture — deserialising one produced by a different Wasmtime build is rejected, which is correct and occasionally surprising in a deployment pipeline that upgrades the runtime without rebuilding artifacts. Make the compiled artifact part of your build output, keyed by both versions.

deserialize_file is unsafe because the runtime trusts the artifact’s contents; it must come from your own build, not from a user.

Where instantiation time actually goes

For a module with a small initial memory and few data segments, instantiation is dominated by allocating and zeroing the memory. Three things make it slower.

A large INITIAL_MEMORY means more to allocate and zero on every instance. A module reserving 256 MB because it might need it pays for that reservation per request. Large data segments — embedded lookup tables, static strings, a bundled configuration file — are copied into memory at instantiation, and a few megabytes of them is milliseconds. And a start function that does real work is, by definition, work on the request path.

# see what the module asks for before it runs
wasm-objdump -x service.wasm | grep -A2 'Memory\['
# memory[0] pages: initial=4096 max=8192        ← 256 MB per instance, per request

wasm-objdump -x service.wasm | grep -c '^ - segment'
# 214                                            ← data segments copied at instantiation

Reducing the initial memory to what a typical request needs, and letting it grow for the unusual ones, is often the single largest instantiation win available.

What instantiation spends its time on Allocating and zeroing linear memory dominates for most modules, followed by copying data segments. A start function that does work adds directly to per-request latency. lean module — 40 µs memory data heavy module — 3.1 ms 256 MB initial memory allocated and zeroed 214 data segments copied start function Seventy-five times the cost, all of it paid on every request, and none of it visible in a benchmark that measures only the handler.

Pooling and copy-on-write

Two runtime features attack instantiation directly.

A pooling allocator pre-reserves memory slots and reuses them, avoiding a fresh mapping per instance. The gain is largest exactly where it matters — high instance churn — and the cost is a fixed reservation sized for your concurrency.

Copy-on-write initialisation maps the module’s initial memory image rather than copying it, so data segments cost nothing until written. For a module with megabytes of static tables this can reduce instantiation by an order of magnitude.

let mut config = Config::new();
config.allocation_strategy(InstanceAllocationStrategy::pooling());
config.memory_init_cow(true);                    // map the image, copy lazily
let engine = Engine::new(&config)?;

Both are host-side settings rather than module changes, which makes them unusually cheap to try. Measure with your own module: the benefit depends entirely on how much initial memory and static data it has.

Cold in a different sense: the first request to a machine

Everything above concerns a process that already holds your module. There is a second kind of cold start that dominates the numbers people actually observe: the first request to reach a machine that has not served your code recently.

That path includes fetching the compiled artifact from wherever the platform stores it, mapping it, and creating the first instance — plus, on some platforms, starting a process or an isolate to hold it. It is the number responsible for the occasional slow request in an otherwise fast service, and it scales with artifact size in a way that steady-state instantiation does not.

Three things reduce it. A smaller artifact transfers and maps faster, which is one of the few situations where shrinking a server-side binary has a latency payoff rather than only a deployment one. Keeping the module warm in the locations that matter — through traffic, or through the platform’s own mechanisms where they exist — avoids the path entirely for most requests. And splitting rarely used functionality into a separate service means the commonly hit path carries a smaller artifact.

Measuring it requires cooperation from the platform, since you cannot force a specific machine to be cold. The practical approach is to deploy, wait, send a small number of requests to a location you have not been using, and look at the tail: the slowest few requests after a deploy are the cold ones, and their distribution is what you are trying to improve.

first requests after a deploy, one location
  p50    2.6 ms
  p95   14.1 ms
  max   84.7 ms    ← artifact fetch + first instance

A max several times the p95 immediately after a deploy is normal and is exactly the cost this section describes. A max like that during steady traffic means machines are going cold between requests, which is a capacity or routing question rather than a module one.

Expected output

A measurement run across configurations makes the decision obvious:

module: service.wasm (2.1 MB, 256 MB initial memory, 214 data segments)

  compile from source          412 ms
  deserialize precompiled      3.9 ms
  instantiate, default        3.14 ms p50   4.82 ms p99
  instantiate, cow             0.42 ms p50   0.71 ms p99
  instantiate, cow + pooling   0.09 ms p50   0.18 ms p99
  instantiate, cow + pooling, initial memory 16 MB
                               0.04 ms p50   0.07 ms p99

The last two lines are the interesting ones: host configuration got a 35× improvement, and a one-line change to the module’s initial memory doubled it again.

What precompiling removes On a cold request the binary must be compiled before anything runs. Precompiling to a serialised module removes almost all of that, leaving instantiation and the request itself. cold, compile at load compile 31 ms init 6 work precompiled module load init 6 ms work — the compile has already happened at build time A serialised module is engine-specific and version-specific; it must be produced by the same runtime build. After precompiling, instantiation and the module's own init dominate, so that is where to look next.

Gotchas

  • Measuring with a reused store. You will measure almost nothing and conclude instantiation is free.
  • Benchmarking a hello-world. A trivial module instantiates in microseconds regardless; use yours.
  • Precompiled artifact from a different runtime version. Rejected at load. Key artifacts by runtime version and architecture.
  • Pooling reservation too small. Instantiation fails under load rather than slowing down. Size it for peak concurrency.
  • Large initial memory “just in case”. Paid on every instance, forever.
  • Counting compilation in a platform comparison. If the platform compiles at deploy, it is not part of request latency and including it makes the comparison meaningless.

Performance note

For the module above, moving from compile-on-load to a precompiled artifact with copy-on-write and pooling took per-request startup from 415 ms to 90 µs — a factor of more than four thousand, achieved entirely with host configuration and one module setting. That is why edge platforms can afford an instance per request: they are all doing exactly this underneath.

Frequently Asked Questions

How does this compare with a container cold start? A container cold start is typically 100 ms to several seconds, dominated by image pull and process startup. A pooled WebAssembly instance from a precompiled module is four to six orders of magnitude faster, which is the entire argument for the model.

Does module size affect instantiation? Indirectly. Size drives compilation, which is off the request path when precompiled. What affects instantiation is initial memory and data segments, which correlate with size but are not the same thing — a large module with a small memory instantiates quickly.

Is lazy compilation worth using? Sometimes, for a module where most functions are never called in a given request. It reduces upfront compilation at the cost of a stall on the first call to each function, which shows up as latency variance. For a precompiled deployment it is usually unnecessary.

← Back to Serverless & Edge Deployment