Running Python in the Browser with Pyodide
This guide answers one task: run Python code in a browser using Pyodide, with the multi-megabyte runtime loaded at a moment the user is willing to wait, data crossing the boundary efficiently, and none of it blocking the interface.
Prerequisites
- [ ] Pyodide 0.26 or later, self-hosted or from a CDN.
- [ ] A worker, because runtime initialisation takes seconds.
- [ ] A clear user action to trigger loading — this is not a background convenience.
- [ ] Realistic expectations about payload: the core is 6–12 MB compressed.
What Pyodide is
Pyodide is CPython compiled to WebAssembly, together with a package system and a set of precompiled
scientific libraries. It is a genuine Python interpreter: the same semantics, the same standard library
for the parts that make sense in a browser, and the ability to import numpy and have it work.
That capability has no equivalent elsewhere. If a feature needs pandas or scikit-learn in a browser, Pyodide is not one option among several — it is the option, and the payload is what it costs.
The corollary is that Pyodide is the wrong tool for anything small. Running a hundred lines of arithmetic does not justify six megabytes and two seconds of startup; a compiled module does that job in twenty kilobytes and a millisecond.
Load it in a worker, on demand
Initialisation is measured in seconds, and on the main thread that is a frozen page. Put it in a worker and start it when the user does something that implies they want it.
// py-worker.js
importScripts('https://cdn.jsdelivr.net/pyodide/v0.26.2/full/pyodide.js');
let pyodide = null;
async function ready() {
pyodide ??= await loadPyodide({ indexURL: 'https://cdn.jsdelivr.net/pyodide/v0.26.2/full/' });
return pyodide;
}
self.onmessage = async ({ data: { id, code, globals } }) => {
try {
const py = await ready();
for (const [k, v] of Object.entries(globals ?? {})) py.globals.set(k, v);
const result = await py.runPythonAsync(code);
self.postMessage({ id, result: result?.toJs ? result.toJs() : result });
} catch (error) {
self.postMessage({ id, error: String(error) });
}
};
Self-hosting the runtime is worth considering for a production deployment: it removes a third-party dependency from your critical path, lets you set your own cache headers, and avoids cross-origin issues if you later enable isolation for threads. The cost is hosting roughly 200 MB of package artifacts, of which each user downloads only what they import.
Installing packages
Two mechanisms exist, and the difference matters. loadPackage fetches a precompiled package from
Pyodide’s own distribution — these are the scientific libraries with C extensions, compiled to
WebAssembly ahead of time. micropip installs pure-Python packages from PyPI at runtime.
# in the worker, before running user code
import micropip
await micropip.install(["pydantic", "rich"]) # pure Python, from PyPI
await pyodide.loadPackage(['numpy', 'pandas']); // precompiled, from the distribution
A package with C extensions that is not in the distribution cannot be installed, because there is nothing to compile it with in the browser. That is the most common disappointment with Pyodide and it is worth checking before designing a feature around a specific library.
Moving data across the boundary
Small values convert automatically: numbers, strings, lists and dictionaries become their JavaScript equivalents and back. Large arrays are where the work is, and Pyodide provides a zero-copy path for typed data.
// JavaScript → Python, no copy for typed arrays
const samples = new Float64Array(buffer);
pyodide.globals.set('samples', samples);
import numpy as np
arr = np.asarray(samples.to_py()) # to_py copies; see below for the alternative
mean = float(arr.mean())
The conversion functions are explicit for a reason: an automatic deep conversion of a large structure is
a silent copy, and Pyodide makes you ask for it. For genuinely large numeric data, use
pyodide.ffi.to_js with the buffer-sharing options, or write the data into the Python heap once and
operate on it there rather than passing it back and forth.
Proxies need releasing. A JavaScript reference to a Python object holds it alive across the boundary, and Pyodide cannot know when you are finished with it:
const df = await py.runPythonAsync('load_dataframe()');
const rows = df.toJs({ dict_converter: Object.fromEntries });
df.destroy(); // release the proxy, or it leaks
Calling JavaScript from Python
The interop goes both ways, and the Python side is unusually comfortable: browser APIs appear as ordinary
Python objects through the js module.
from js import document, fetch, console
from pyodide.ffi import create_proxy
console.log("hello from Python")
async def load_and_render(url):
resp = await fetch(url)
data = await resp.json()
el = document.querySelector("#out")
el.textContent = f"{len(data.to_py())} records"
def on_click(event):
console.log("clicked", event.target.id)
# a callback passed to JavaScript must be a proxy, and must be kept alive
proxy = create_proxy(on_click)
document.querySelector("#go").addEventListener("click", proxy)
The create_proxy call is the part people miss. Passing a Python function directly to an event listener
gives JavaScript a temporary that Python may collect, after which the listener fires into freed memory.
Creating an explicit proxy and keeping a reference to it — in a module-level variable, a list, or
wherever the listener’s lifetime is managed — is required, and calling destroy() on it when the listener
is removed prevents the mirror-image leak.
Note that in a worker there is no document, which is the usual arrangement for anything long-running.
Python in a worker computes and returns results; the main thread renders them. That separation is worth
keeping even when it would be convenient to touch the DOM directly, because it is what keeps a slow script
from freezing the page.
Expected output
A working setup reports the phases separately, which is what you want to show the user as well:
pyodide: runtime loaded in 1840 ms (cached: 620 ms)
pyodide: loadPackage(['numpy']) in 410 ms
python: >>> import numpy as np; np.arange(10).mean()
4.5
python: script completed in 12 ms
The gap between cold and cached startup is the number that decides whether this is acceptable: 1.8 s on a first visit and 0.6 s afterwards is fine behind a button, and unacceptable on page load.
Filesystem, network and the sandbox
Python code expects a filesystem, and Pyodide emulates one in memory. Files written there vanish on reload unless you mount something persistent.
pyodide.FS.mkdir('/data');
pyodide.FS.mount(pyodide.FS.filesystems.IDBFS, {}, '/data');
await new Promise((res) => pyodide.FS.syncfs(true, res)); // load from IndexedDB
// … Python writes to /data …
await new Promise((res) => pyodide.FS.syncfs(false, res)); // persist back
Network access is the browser’s, not Python’s. requests and urllib do not work as they would on a
server; use pyodide.http.pyfetch, which wraps the browser’s fetch, and expect CORS to apply exactly as
it does to any other request from the page. Code copied from a server-side script will fail here, and the
error will be about sockets rather than about the browser.
Subprocesses, threads and signals are similarly absent or limited. Anything that shells out is not going to work.
Gotchas
- Loading on the main thread. Seconds of frozen page. Use a worker.
- A package with C extensions not in the distribution. Cannot be installed; check before designing around it.
- Proxies never destroyed. The Python heap grows until the tab does.
requestsinstead ofpyfetch. Server-side networking code does not transfer.- Expecting files to persist. Mount IDBFS and call
syncfs, or accept that they will not. - Measuring cold startup only once. Cached startup is a completely different number and is the one most users experience.
Performance note
On a laptop: 1.84 s to load the runtime cold over a fast connection, 620 ms from cache, plus 410 ms to load NumPy. Once running, numeric work in NumPy was roughly 1.5–3× slower than the same operation natively — impressively close, because NumPy’s inner loops are compiled C rather than interpreted Python. Pure Python loops were 5–20× slower than native, which is the interpreter overhead you would expect.
Frequently Asked Questions
Is this suitable for a production feature? For the right feature, yes — data exploration tools, teaching environments, notebook interfaces and anything where the user has explicitly asked to run Python. It is not suitable as an implementation detail users never asked for.
Can I run untrusted Python with it? The WebAssembly sandbox contains it, so it cannot reach the host, but it can consume unbounded memory and loop forever inside the worker. Terminate the worker on a deadline, exactly as with any untrusted module.
How do I show progress while it loads?
Pass a message callback to loadPyodide and report each phase, and report package loading separately —
users tolerate a long wait far better when the interface names what it is doing rather than showing an
undifferentiated spinner for four seconds.
Does it support threads? Partially and with caveats, requiring cross-origin isolation. Most deployments run it single-threaded and rely on the compiled libraries’ own vectorisation for speed.
Related
- Comparing payload size across languages — where Pyodide sits.
- Working with datasets larger than memory — for data too big for the tab.
- Loading Wasm in a Web Worker with ESM — the worker setup.
← Back to Other Languages in the Browser