Driving WebGPU from Rust Wasm

This guide answers one task: run Rust code compiled to WebAssembly that talks to WebGPU — creating a device, uploading buffers, running a render or compute pipeline — and understand which parts of a native wgpu program have to change for the browser.

Prerequisites

  • [ ] Rust with the wasm32-unknown-unknown target and wasm-pack 0.13+.
  • [ ] wgpu 22+ with the webgl fallback feature if you want one.
  • [ ] A browser with WebGPU: Chrome 113+, Edge 113+, Safari 18+, or Firefox with it enabled.
  • [ ] A secure context. WebGPU is unavailable over plain HTTP outside localhost.

Why wgpu rather than raw bindings

You can call WebGPU directly from Rust through web-sys, and it works, but the API is asynchronous in several places and JavaScript-shaped throughout. wgpu implements the same API surface with Rust ergonomics and, more importantly, compiles to native backends as well — so one codebase runs on Vulkan, Metal and Direct3D when built natively and on WebGPU when built for the web.

That portability is the main argument for this stack. A renderer written once can be developed and profiled as a native binary, where the tooling is far better, and shipped to a browser without a separate implementation. The shaders are the same too, since WGSL compiles everywhere wgpu runs.

One renderer, several backends The same Rust source targets native graphics APIs when compiled for the host and WebGPU when compiled to WebAssembly. Shaders in WGSL are shared, so only the entry point and the surface creation differ. your renderer + WGSL native: Vulkan / Metal good tooling, fast iteration browser: WebGPU same code, async setup browser: WebGL 2 fallback, reduced features

Set up the crate

The web target needs a handful of extra dependencies and a different entry point. Keep them behind a cfg so the native build stays clean.

[lib]
crate-type = ["cdylib", "rlib"]

[dependencies]
wgpu = "22"
winit = "0.30"
pollster = "0.3"

[target.'cfg(target_arch = "wasm32")'.dependencies]
wasm-bindgen = "0.2"
wasm-bindgen-futures = "0.4"
web-sys = { version = "0.3", features = ["Document", "Window", "Element", "HtmlCanvasElement"] }
console_error_panic_hook = "0.1"

console_error_panic_hook is not optional in practice. Without it a Rust panic surfaces as RuntimeError: unreachable executed with no message, and you will spend an afternoon finding out which unwrap it was.

Device creation is asynchronous, and that shapes everything

Natively you can block on adapter and device creation with pollster::block_on. In a browser you cannot block at all — the event loop is the only thing that can make progress, so blocking on a promise deadlocks the page.

The entry point therefore has to be async, driven by wasm_bindgen_futures:

#[cfg_attr(target_arch = "wasm32", wasm_bindgen(start))]
pub async fn start() {
    #[cfg(target_arch = "wasm32")]
    console_error_panic_hook::set_once();

    let instance = wgpu::Instance::default();
    let canvas = web_sys::window().unwrap().document().unwrap()
        .get_element_by_id("gpu-canvas").unwrap()
        .dyn_into::<web_sys::HtmlCanvasElement>().unwrap();
    let surface = instance.create_surface(wgpu::SurfaceTarget::Canvas(canvas)).unwrap();

    let adapter = instance.request_adapter(&wgpu::RequestAdapterOptions {
        compatible_surface: Some(&surface), ..Default::default()
    }).await.expect("no adapter — WebGPU unavailable or blocked");

    let (device, queue) = adapter.request_device(&wgpu::DeviceDescriptor {
        required_limits: wgpu::Limits::downlevel_webgl2_defaults().using_resolution(adapter.limits()),
        ..Default::default()
    }).await.unwrap();
}

Requesting downlevel_webgl2_defaults limits is a deliberate conservatism: it keeps the code within what the WebGL fallback can also do, so a device without WebGPU still runs. Drop it only when you know you need the higher limits.

Uploading data the module owns

Buffer writes go through the queue, and they take a byte slice — which for data already in linear memory is simply a Rust slice. No copy into JavaScript happens at any point.

let vertex_buffer = device.create_buffer(&wgpu::BufferDescriptor {
    label: Some("vertices"),
    size: (MAX_VERTS * std::mem::size_of::<Vertex>()) as u64,
    usage: wgpu::BufferUsages::VERTEX | wgpu::BufferUsages::COPY_DST,
    mapped_at_creation: false,
});

// per frame: write only the live range
queue.write_buffer(&vertex_buffer, 0, bytemuck::cast_slice(&vertices[..live]));

bytemuck::cast_slice reinterprets a slice of your vertex struct as bytes with no copy, provided the struct is #[repr(C)] and plain old data. Getting that annotation wrong is the usual cause of geometry that renders as noise — Rust is free to reorder fields otherwise.

Shaders in WGSL

WGSL is the only shader language WebGPU accepts, and wgpu compiles it for native backends too, so one shader source serves every target.

let shader = device.create_shader_module(wgpu::ShaderModuleDescriptor {
    label: Some("main"),
    source: wgpu::ShaderSource::Wgsl(include_str!("shader.wgsl").into()),
});
struct VsOut { @builtin(position) pos: vec4f, @location(0) color: vec4f };

@vertex
fn vs_main(@location(0) position: vec2f, @location(1) color: vec4f) -> VsOut {
  var out: VsOut;
  out.pos = vec4f(position, 0.0, 1.0);
  out.color = color;
  return out;
}

@fragment
fn fs_main(in: VsOut) -> @location(0) vec4f { return in.color; }

Shader compilation errors surface at pipeline creation with a line and column, which is a genuine improvement over WebGL’s opaque log strings. Read them; they are usually precise.

Setup happens once, submission happens every frame Adapter, device, shader and pipeline creation are asynchronous one-time costs paid at startup. Per frame, only buffer writes, a render pass and a queue submission occur, all of them synchronous. startup — asynchronous, once request adapter request device compile shaders build pipeline per frame — synchronous write_buffer encode render pass queue.submit Pipeline creation is the expensive step and it is easy to do accidentally per frame — build every pipeline you need at startup and keep them.

Reading results back from a compute pass

WebGPU’s compute pipelines make the GPU available for general work, and reading the result back is the part that differs most from native code. Mapping a buffer is asynchronous, and the callback runs when the GPU has finished — you cannot wait for it inline.

let slice = staging.slice(..);
let (tx, rx) = futures_channel::oneshot::channel();
slice.map_async(wgpu::MapMode::Read, move |r| { let _ = tx.send(r); });
device.poll(wgpu::Maintain::Wait);          // no-op in the browser; the event loop drives it
rx.await.unwrap().unwrap();
let data = slice.get_mapped_range();
process(&data);
drop(data);
staging.unmap();

Native code often calls device.poll(Maintain::Wait) and proceeds. In the browser that call does nothing, because progress happens when control returns to the event loop — so the await is doing the real work and the structure must be genuinely asynchronous rather than a blocking call wearing an async coat.

Handling device loss and resize

Two events will happen to a long-running WebGPU application, and neither is optional to handle.

A device can be lost — the browser reclaims it when a tab is backgrounded for long enough on some platforms, a driver resets, or the user switches graphics adapters. Every resource created from that device becomes invalid at once, and continuing to use them produces validation errors on every call. The recovery is to request a new adapter and device and rebuild everything, which is far easier if resource creation lives in one function you can call twice.

let lost = device.lost.clone();
wasm_bindgen_futures::spawn_local(async move {
    let info = lost.await;
    web_sys::console::warn_1(&format!("device lost: {:?}", info.reason).into());
    rebuild_everything().await;
});

Resize is the other. The canvas’s backing store does not follow its CSS size, and a surface configured for the wrong dimensions renders blurry or clipped with no error to explain it. Observe the element, multiply by devicePixelRatio, clamp to the device’s maximum texture dimension, and reconfigure:

let w = (css_width * dpr).round().min(limits.max_texture_dimension_2d as f64) as u32;
surface.configure(&device, &wgpu::SurfaceConfiguration { width: w, height: h, ..config });

Reconfiguring on every resize event during a drag is wasteful; debounce it to one call when the size settles, and skip it entirely when the computed dimensions have not changed, which happens more often than you would expect because observers fire for reasons other than a real resize.

Expected output

A successful start logs the adapter and its limits, which is also the fastest way to see whether you got a hardware device or a fallback:

adapter: Apple M2 (Metal) — type: IntegratedGpu
limits : max_buffer_size 2147483648, max_compute_workgroup_size_x 1024
surface: Bgra8UnormSrgb, present mode Fifo
frame  : p50 3.1 ms, p95 4.4 ms

An adapter reporting Cpu as its type means a software fallback, which will be very slow. Detect it and either warn or switch to the WebGL path deliberately rather than shipping a renderer that runs at four frames per second.

Who owns what in a WebGPU frame The device, the queue and every buffer belong to the browser. Rust holds handles to them and encodes commands; the submission is a single crossing per frame. device handle owned by the browser encode pass in Rust, no crossing queue.submit one crossing per frame GPU executes nothing in the module Encoding in Rust and submitting once keeps the per-frame boundary cost flat regardless of scene size. Buffer writes are the exception: each one crosses, so batch updates into a single staging write. Device loss is a normal event, not an error case — handle it or the canvas stays blank after a sleep.

Gotchas

  • request_adapter returns None. WebGPU is unavailable, blocked by driver policy, or the context is insecure. Always have a fallback path.
  • Blocking on a future. pollster::block_on deadlocks the page. Everything must be awaited.
  • Panics with no message. Install console_error_panic_hook before anything else.
  • Vertex struct not #[repr(C)]. Field reordering produces garbage geometry that still draws.
  • Pipeline created per frame. Expensive; build at startup and cache.
  • Canvas size and surface configuration out of sync. Reconfigure the surface on resize, or you get a stretched or clipped image with no error.

Performance note

For a scene of ten thousand instanced quads, the WebGPU path on an integrated GPU rendered in roughly 1.1 ms per frame against 1.8 ms for the equivalent WebGL 2 path, with the gap widening substantially for compute-heavy work where WebGL has no equivalent at all. Startup is the cost: adapter and device acquisition plus pipeline compilation took 180–420 ms, which is why it belongs behind a loading state rather than in the middle of an interaction.

Frequently Asked Questions

Should I ship WebGPU only, or both paths? Both, for now. wgpu can target WebGL 2 from the same code with a feature flag, so the fallback costs a build configuration rather than a second renderer.

Is compute worth it over a threaded Wasm build? For large, regular, data-parallel work, yes, by a wide margin. For irregular or branch-heavy simulation, threads in WebAssembly often win because the GPU’s advantage depends on uniform control flow.

How do I debug a WGSL shader in the browser? Compilation errors are reported precisely at pipeline creation. For runtime behaviour, write intermediate values to a storage buffer and read them back — there is no shader debugger, so instrumenting the shader is the practical approach.

← Back to Graphics, Games & Simulation