Running Python in the Browser with Pyodide

This guide answers one task: run Python code in a browser using Pyodide, with the multi-megabyte runtime loaded at a moment the user is willing to wait, data crossing the boundary efficiently, and none of it blocking the interface.

Prerequisites

  • [ ] Pyodide 0.26 or later, self-hosted or from a CDN.
  • [ ] A worker, because runtime initialisation takes seconds.
  • [ ] A clear user action to trigger loading — this is not a background convenience.
  • [ ] Realistic expectations about payload: the core is 6–12 MB compressed.

What Pyodide is

Pyodide is CPython compiled to WebAssembly, together with a package system and a set of precompiled scientific libraries. It is a genuine Python interpreter: the same semantics, the same standard library for the parts that make sense in a browser, and the ability to import numpy and have it work.

That capability has no equivalent elsewhere. If a feature needs pandas or scikit-learn in a browser, Pyodide is not one option among several — it is the option, and the payload is what it costs.

The corollary is that Pyodide is the wrong tool for anything small. Running a hundred lines of arithmetic does not justify six megabytes and two seconds of startup; a compiled module does that job in twenty kilobytes and a millisecond.

What arrives when you load Pyodide The CPython interpreter compiled to WebAssembly forms the base, with the standard library above it, optional packages fetched on demand, and the user's code on top. Only the last layer is small. your Python code — kilobytes packages loaded on demand — numpy 4 MB, pandas 9 MB Python standard library CPython compiled to WebAssembly — the floor, 6–12 MB Packages are fetched only when imported, so the cost depends entirely on which ones your code actually uses.

Load it in a worker, on demand

Initialisation is measured in seconds, and on the main thread that is a frozen page. Put it in a worker and start it when the user does something that implies they want it.

// py-worker.js
importScripts('https://cdn.jsdelivr.net/pyodide/v0.26.2/full/pyodide.js');

let pyodide = null;
async function ready() {
  pyodide ??= await loadPyodide({ indexURL: 'https://cdn.jsdelivr.net/pyodide/v0.26.2/full/' });
  return pyodide;
}

self.onmessage = async ({ data: { id, code, globals } }) => {
  try {
    const py = await ready();
    for (const [k, v] of Object.entries(globals ?? {})) py.globals.set(k, v);
    const result = await py.runPythonAsync(code);
    self.postMessage({ id, result: result?.toJs ? result.toJs() : result });
  } catch (error) {
    self.postMessage({ id, error: String(error) });
  }
};

Self-hosting the runtime is worth considering for a production deployment: it removes a third-party dependency from your critical path, lets you set your own cache headers, and avoids cross-origin issues if you later enable isolation for threads. The cost is hosting roughly 200 MB of package artifacts, of which each user downloads only what they import.

Installing packages

Two mechanisms exist, and the difference matters. loadPackage fetches a precompiled package from Pyodide’s own distribution — these are the scientific libraries with C extensions, compiled to WebAssembly ahead of time. micropip installs pure-Python packages from PyPI at runtime.

# in the worker, before running user code
import micropip
await micropip.install(["pydantic", "rich"])       # pure Python, from PyPI
await pyodide.loadPackage(['numpy', 'pandas']);     // precompiled, from the distribution

A package with C extensions that is not in the distribution cannot be installed, because there is nothing to compile it with in the browser. That is the most common disappointment with Pyodide and it is worth checking before designing a feature around a specific library.

Moving data across the boundary

Small values convert automatically: numbers, strings, lists and dictionaries become their JavaScript equivalents and back. Large arrays are where the work is, and Pyodide provides a zero-copy path for typed data.

// JavaScript → Python, no copy for typed arrays
const samples = new Float64Array(buffer);
pyodide.globals.set('samples', samples);
import numpy as np
arr = np.asarray(samples.to_py())     # to_py copies; see below for the alternative
mean = float(arr.mean())

The conversion functions are explicit for a reason: an automatic deep conversion of a large structure is a silent copy, and Pyodide makes you ask for it. For genuinely large numeric data, use pyodide.ffi.to_js with the buffer-sharing options, or write the data into the Python heap once and operate on it there rather than passing it back and forth.

Proxies need releasing. A JavaScript reference to a Python object holds it alive across the boundary, and Pyodide cannot know when you are finished with it:

const df = await py.runPythonAsync('load_dataframe()');
const rows = df.toJs({ dict_converter: Object.fromEntries });
df.destroy();                                        // release the proxy, or it leaks
Three ways a value crosses Small values convert automatically. Large typed arrays can share memory without copying. Complex Python objects arrive as proxies that must be explicitly destroyed to avoid leaking. small values numbers, strings, lists converted automatically nothing to manage typed arrays shared buffer, no copy right for bulk numeric data ask explicitly object proxies a handle into the Python heap every access calls back destroy() or it leaks The third column is where memory problems come from: a proxy held in a JavaScript closure keeps a Python object alive indefinitely.

Calling JavaScript from Python

The interop goes both ways, and the Python side is unusually comfortable: browser APIs appear as ordinary Python objects through the js module.

from js import document, fetch, console
from pyodide.ffi import create_proxy

console.log("hello from Python")

async def load_and_render(url):
    resp = await fetch(url)
    data = await resp.json()
    el = document.querySelector("#out")
    el.textContent = f"{len(data.to_py())} records"

def on_click(event):
    console.log("clicked", event.target.id)

# a callback passed to JavaScript must be a proxy, and must be kept alive
proxy = create_proxy(on_click)
document.querySelector("#go").addEventListener("click", proxy)

The create_proxy call is the part people miss. Passing a Python function directly to an event listener gives JavaScript a temporary that Python may collect, after which the listener fires into freed memory. Creating an explicit proxy and keeping a reference to it — in a module-level variable, a list, or wherever the listener’s lifetime is managed — is required, and calling destroy() on it when the listener is removed prevents the mirror-image leak.

Note that in a worker there is no document, which is the usual arrangement for anything long-running. Python in a worker computes and returns results; the main thread renders them. That separation is worth keeping even when it would be convenient to touch the DOM directly, because it is what keeps a slow script from freezing the page.

Expected output

A working setup reports the phases separately, which is what you want to show the user as well:

pyodide: runtime loaded in 1840 ms (cached: 620 ms)
pyodide: loadPackage(['numpy']) in 410 ms
python: >>> import numpy as np; np.arange(10).mean()
        4.5
python: script completed in 12 ms

The gap between cold and cached startup is the number that decides whether this is acceptable: 1.8 s on a first visit and 0.6 s afterwards is fine behind a button, and unacceptable on page load.

Filesystem, network and the sandbox

Python code expects a filesystem, and Pyodide emulates one in memory. Files written there vanish on reload unless you mount something persistent.

pyodide.FS.mkdir('/data');
pyodide.FS.mount(pyodide.FS.filesystems.IDBFS, {}, '/data');
await new Promise((res) => pyodide.FS.syncfs(true, res));      // load from IndexedDB
// … Python writes to /data …
await new Promise((res) => pyodide.FS.syncfs(false, res));     // persist back

Network access is the browser’s, not Python’s. requests and urllib do not work as they would on a server; use pyodide.http.pyfetch, which wraps the browser’s fetch, and expect CORS to apply exactly as it does to any other request from the page. Code copied from a server-side script will fail here, and the error will be about sockets rather than about the browser.

Subprocesses, threads and signals are similarly absent or limited. Anything that shells out is not going to work.

What loads before your script The interpreter is fetched and compiled, the standard library is mounted, any packages are installed, and only then does the first line of Python execute. interpreter binary several megabytes standard library mounted as files packages fetched on demand your code runs first line at last Load it in a worker and show something useful meanwhile; the sequence takes seconds on a good connection. Install only the packages actually used — each one is a separate download on top of the interpreter. The result is worth it when the alternative is reimplementing a scientific library nobody wants to port.

Gotchas

  • Loading on the main thread. Seconds of frozen page. Use a worker.
  • A package with C extensions not in the distribution. Cannot be installed; check before designing around it.
  • Proxies never destroyed. The Python heap grows until the tab does.
  • requests instead of pyfetch. Server-side networking code does not transfer.
  • Expecting files to persist. Mount IDBFS and call syncfs, or accept that they will not.
  • Measuring cold startup only once. Cached startup is a completely different number and is the one most users experience.

Performance note

On a laptop: 1.84 s to load the runtime cold over a fast connection, 620 ms from cache, plus 410 ms to load NumPy. Once running, numeric work in NumPy was roughly 1.5–3× slower than the same operation natively — impressively close, because NumPy’s inner loops are compiled C rather than interpreted Python. Pure Python loops were 5–20× slower than native, which is the interpreter overhead you would expect.

Frequently Asked Questions

Is this suitable for a production feature? For the right feature, yes — data exploration tools, teaching environments, notebook interfaces and anything where the user has explicitly asked to run Python. It is not suitable as an implementation detail users never asked for.

Can I run untrusted Python with it? The WebAssembly sandbox contains it, so it cannot reach the host, but it can consume unbounded memory and loop forever inside the worker. Terminate the worker on a deadline, exactly as with any untrusted module.

How do I show progress while it loads? Pass a message callback to loadPyodide and report each phase, and report package loading separately — users tolerate a long wait far better when the interface names what it is doing rather than showing an undifferentiated spinner for four seconds.

Does it support threads? Partially and with caveats, requiring cross-origin isolation. Most deployments run it single-threaded and rely on the compiled libraries’ own vectorisation for speed.

← Back to Other Languages in the Browser