Web Crypto vs Wasm for Hashing

This guide answers one task: decide whether to hash with the browser’s built-in SubtleCrypto or with a compiled module, based on measurements rather than instinct — and know which of the two is a mistake in each direction.

Prerequisites

  • [ ] A secure context; crypto.subtle is unavailable otherwise.
  • [ ] A compiled hash module for comparison — BLAKE3 or SHA-256 from a Rust crate works well.
  • [ ] A benchmark harness that discards warm-up and reports percentiles.
  • [ ] Inputs at several sizes: 64 B, 4 kB, 1 MB, 64 MB. The answer changes across them.

What each one is good at

crypto.subtle.digest is native code, often with hardware acceleration for SHA-2, running outside the JavaScript heap. Per byte, nothing you compile will beat it for the algorithms it implements. Its weaknesses are structural rather than computational: every call is asynchronous, it is one-shot rather than incremental, and it supports a fixed list of algorithms.

A compiled module is synchronous, incremental if you write it that way, and supports whatever you compile. Per byte it is typically 1.5–4× slower than the native SHA-256 implementation, and considerably faster if you choose an algorithm the platform does not have — BLAKE3 with SIMD outruns native SHA-256 on large inputs on most machines.

Overhead versus throughput The platform digest has higher fixed cost per call because it is asynchronous, and higher throughput per byte. A compiled module has almost no per-call overhead and slightly lower throughput, so each wins at a different input size. crypto.subtle.digest per call: ~0.05–0.2 ms (async) per byte: fastest for SHA-2 one-shot only, no streaming wins on big buffers compiled module per call: microseconds (sync) per byte: 1.5–4× slower for SHA-2 incremental, any algorithm wins on many small inputs Hashing ten thousand short strings is dominated by per-call cost; hashing one large file is dominated by throughput. Measure the shape you actually have.

Measure both, properly

The comparison is only meaningful with warm-up discarded and the same input for both. Note that the platform call is asynchronous, so the measurement includes promise scheduling — which is real cost your application pays.

async function benchSubtle(bytes, n = 200) {
  await crypto.subtle.digest('SHA-256', bytes);              // warm up
  const t0 = performance.now();
  for (let i = 0; i < n; i++) await crypto.subtle.digest('SHA-256', bytes);
  return (performance.now() - t0) / n;
}

function benchWasm(mod, ptr, len, n = 200) {
  mod.exports.sha256(ptr, len);                              // warm up
  const t0 = performance.now();
  for (let i = 0; i < n; i++) mod.exports.sha256(ptr, len);
  return (performance.now() - t0) / n;
}

For the module, make sure the input is already in linear memory before timing, or you are measuring the copy rather than the hash. If your real workload does include that copy — because the data arrives as a JavaScript ArrayBuffer — then include it deliberately and say so.

Expected output

Typical numbers on a modern laptop, SHA-256, same input for both:

size      subtle (ms)   wasm (ms)   winner
64 B          0.061       0.0021    wasm  (29×)
4 kB          0.068       0.0094    wasm  (7×)
256 kB        0.28        0.42      subtle
1 MB          0.94        1.61      subtle
64 MB        58.2        99.4       subtle

The crossover for SHA-256 sits somewhere around 100 kB on most machines. Below it, the asynchronous call overhead dominates and the module is dramatically faster; above it, native throughput wins and the gap widens with size.

Many small hashes: the module’s real win

The case where a compiled hash is unambiguously better is a loop over many small inputs — deduplicating records, building a Merkle tree, computing content addresses for a list of chunks. Each await in that loop costs a microtask round trip, and ten thousand of them is tens of milliseconds of pure overhead before any hashing happens.

// 10,000 short strings
// subtle:  await in a loop            → 610 ms
// subtle:  Promise.all batched        → 240 ms
// wasm:    synchronous loop           →  38 ms

Batching with Promise.all recovers some of it but not most, because the per-call cost is not only scheduling. The synchronous module simply does not have the problem, and the difference is large enough to change which features are feasible.

Streaming input the platform cannot handle

crypto.subtle.digest takes a complete buffer. To hash a 2 GB file you would have to hold all of it in memory, which is not acceptable in a tab. A compiled hash with an incremental interface handles it in bounded memory:

const h = mod.exports.hash_init();
const reader = file.stream().getReader();
const view = new Uint8Array(mod.exports.memory.buffer, mod.exports.chunk_ptr(), CHUNK);
for (;;) {
  const { value, done } = await reader.read();
  if (done) break;
  for (let off = 0; off < value.length; off += CHUNK) {
    const slice = value.subarray(off, off + CHUNK);
    view.set(slice);
    mod.exports.hash_update(h, slice.length);
  }
}
const digest = readDigest(mod, mod.exports.hash_final(h));

Peak memory is one chunk. This is the single most common legitimate reason to compile a hash: not speed, but the ability to process data you cannot hold.

One-shot versus incremental The platform digest needs the entire input in memory at once. An incremental compiled hash consumes fixed-size chunks and keeps only its internal state, so file size stops being a memory constraint. one-shot entire 2 GB input resident in memory → allocation failure incremental chunk chunk chunk state: 200 bytes → peak memory is one chunk The same argument applies to hashing a stream that has no end — a live upload, a generated sequence — where a one-shot API cannot be used at all.

Algorithms the platform does not have

The last category is straightforward: if you need BLAKE3, BLAKE2b, SHA-3, xxHash or a domain-specific construction, SubtleCrypto cannot help and the decision is made for you.

BLAKE3 is worth calling out because it inverts the throughput conclusion. With SIMD and its tree structure, a compiled BLAKE3 commonly reaches 2–4 GB/s single-threaded, ahead of native SHA-256 on the same machine, and it parallelises across workers for large inputs. If you control both ends — content addressing inside your own system rather than interoperating with something that specifies SHA-256 — it is usually the better choice on every axis except familiarity.

Content addressing: a worked decision

A concrete case makes the tradeoffs easier to see. Suppose an application splits uploads into 1 MB chunks and computes a digest per chunk to deduplicate against what the server already holds. A 500 MB upload is 500 hashes of 1 MB each.

With crypto.subtle, that is 500 asynchronous calls at roughly 0.94 ms of hashing plus call overhead — around 520 ms in total, all of it off the main thread’s critical path but each result arriving as a microtask that interleaves with rendering. With a compiled SHA-256 the same work takes about 810 ms of synchronous compute, which must be in a worker or the page freezes. With BLAKE3 it takes around 160 ms, and the tree structure means several workers can hash disjoint ranges and combine the results.

The decision follows from a question that has nothing to do with speed: does the digest have to interoperate? If the server, or a specification, or another client expects SHA-256, the algorithm is fixed and the only choice is where to run it — and for 1 MB chunks the platform wins. If the digest is internal to your system, BLAKE3 in a worker is three times faster, streams naturally, and parallelises.

That is the shape of almost every real decision here. Interoperability constrains the algorithm; the algorithm mostly determines the implementation; and the measurement only settles the remaining question of where the work runs.

Combining both in one application

Nothing requires a single choice. A realistic application uses SubtleCrypto for the standard operations it covers — verifying a signature, deriving a key, hashing a moderate buffer — and a module for the specific workload the platform handles badly.

Keep the boundary explicit. A small hash.js module that exposes digestBuffer, digestStream and digestMany, routing each to whichever implementation is right, means call sites do not encode the decision and you can revisit it with a measurement rather than a refactor. It also gives you one place to put the lazy loading: the compiled module should not be fetched at all by a session that only ever hashes one small buffer.

Where each implementation wins WebCrypto calls are asynchronous with a fixed per-call cost, which dominates for small inputs. For large buffers the native implementation pulls far ahead of any module. 64-byte inputs Wasm 2.1M hashes/s WebCrypto 190k 1 MB buffer Wasm 780 MB/s WebCrypto 2,400 MB/s streaming 500 MB Wasm 760 MB/s WebCrypto 2,380 MB/s — hardware accelerated Use WebCrypto for bulk data and a module for many small hashes or an algorithm WebCrypto does not offer. WebCrypto also needs a secure context; a module works anywhere, which occasionally decides it.

Gotchas

  • crypto.subtle is undefined. Insecure context. It is not a browser support problem; it is HTTP.
  • Comparing a cold module against a warm platform call. Instantiate, warm up, then measure.
  • Measuring with the input outside linear memory. The copy is then in your measurement. Decide whether it belongs there.
  • Assuming SHA-1 is available for a checksum. Some browsers restrict it in SubtleCrypto; a module has no such restriction, which is occasionally the only reason to use one.
  • Hashing a File by reading it entirely. Stream it. The one-shot API cannot, and memory will not forgive you.
  • Using a fast non-cryptographic hash where a cryptographic one is needed. xxHash is excellent for hash tables and useless against an adversary.

Performance note

Measured on an M2 laptop: SHA-256 over 1 MB took 0.94 ms with crypto.subtle and 1.61 ms with a compiled Rust implementation, while BLAKE3 over the same input took 0.31 ms — three times faster than the platform’s own SHA-256. Over 64 B inputs in a loop, the module was roughly 29× faster than the platform because the comparison is almost entirely call overhead. Both conclusions are correct simultaneously, which is why a single “which is faster” answer is always wrong.

Frequently Asked Questions

Should I replace SubtleCrypto everywhere with a module? No. For occasional hashes of substantial buffers it is faster, native, free to ship and better reviewed. Reach for a module when the workload is many small hashes, a stream, or an algorithm it lacks.

Does hashing in a worker change the comparison? It removes the main-thread blocking concern from the synchronous module, which is otherwise the one real argument against it. Both APIs are available in workers.

What about HMAC and key derivation? Use the platform. SubtleCrypto implements HMAC and PBKDF2 natively with non-extractable key handling, which is a security advantage a module cannot match — the key never enters your JavaScript heap at all.

← Back to Cryptography & Untrusted Code