Using the Emscripten File System API

This guide answers one task: make fopen, fread and the rest work inside a browser for code compiled with Emscripten — choosing the right file system for the job, getting data in and out, and persisting what should survive a reload.

Prerequisites

  • [ ] Emscripten 3.1.60 or later, activated in your shell.
  • [ ] C or C++ code that uses standard file APIs.
  • [ ] A clear idea of which files are inputs, which are outputs and which must persist.
  • [ ] A local server; the virtual file system’s data file will not load from file://.

Four file systems, four jobs

Emscripten’s virtual file system is a layer under the standard library, and you mount different backends at different paths depending on what each is for.

MEMFS is the default and lives entirely in linear memory. It is fast, it is the right choice for scratch files and outputs, and it disappears on reload.

IDBFS persists to IndexedDB, and is how a user’s work survives a refresh. It is not automatic: writes land in memory and reach IndexedDB only when you call syncfs.

NODEFS maps a real directory and exists only under Node, which makes it useful for testing the same module outside a browser.

WORKERFS mounts File and Blob objects a user selected, without copying their contents into linear memory — the option that makes a two-gigabyte input possible.

One virtual tree, several backends The virtual file system mounts different backends at different paths: memory for scratch, IndexedDB for persistence, a host directory under Node, and user-selected files without copying them. the virtual file system MEMFS /tmp in linear memory fast, volatile IDBFS /data IndexedDB behind it needs syncfs NODEFS /host a real directory Node only WORKERFS /in user files, not copied worker only

Preloading files at build time

The simplest way to give code the files it expects is to bundle them. Emscripten emits a .data file alongside the module and populates the file system before main runs.

emcc app.c --preload-file assets@/assets -o app.js
FILE *f = fopen("/assets/config.txt", "r");      // just works

The @ separates the host path from the virtual path, which is worth using explicitly — without it the files land at a path derived from your build directory, and the code then works on your machine and nowhere else.

The whole .data file downloads before main runs, so this suits tens of megabytes at most. Beyond that, fetch and write files yourself so the application starts while data arrives.

Writing files from JavaScript

For inputs the user provides, write them into the file system before calling into the module.

const bytes = new Uint8Array(await file.arrayBuffer());
Module.FS.mkdir('/work');
Module.FS.writeFile('/work/input.bin', bytes);

Module.ccall('process_file', 'number', ['string'], ['/work/input.bin']);

const out = Module.FS.readFile('/work/output.bin');       // Uint8Array
Module.FS.unlink('/work/input.bin');                      // free the memory
Module.FS.unlink('/work/output.bin');

Note the unlink calls. MEMFS holds file contents in linear memory, so a file that is no longer needed is memory you cannot reclaim any other way. A long-running session that writes files without removing them grows until the tab fails.

Export what you need at build time, or these APIs will not be present:

emcc app.c -sEXPORTED_RUNTIME_METHODS='["FS","ccall","cwrap"]' -o app.js

Persisting with IDBFS

IDBFS makes a directory survive a reload, with one significant caveat: nothing is written to IndexedDB until you ask.

Module.FS.mkdir('/data');
Module.FS.mount(Module.IDBFS, {}, '/data');

// load what was persisted, before using the directory
await new Promise((res, rej) => Module.FS.syncfs(true, (e) => (e ? rej(e) : res())));

// … the module reads and writes /data …

// persist back — this is the call people forget
await new Promise((res, rej) => Module.FS.syncfs(false, (e) => (e ? rej(e) : res())));

The boolean argument is the direction: true populates the file system from IndexedDB, false writes it back. Getting it backwards on startup wipes the user’s data, which is a sufficiently unpleasant bug to be worth a comment in the code.

Sync at meaningful points — after a save action, on a timer, and on visibilitychange rather than on beforeunload, which browsers increasingly do not guarantee will run.

Large inputs without copying

WORKERFS mounts File or Blob objects so the module reads them lazily rather than holding them in linear memory. For a media or archive tool, this is the difference between handling a two-gigabyte input and failing to allocate.

// in a worker
FS.mkdir('/in');
FS.mount(WORKERFS, { files: [file] }, '/in');       // file is a File from an input element

Module.ccall('scan_archive', 'number', ['string'], [`/in/${file.name}`]);

Reads are synchronous from the module’s point of view and backed by the browser’s file object underneath, so peak memory is whatever the module buffers rather than the file’s size. It is read-only and worker-only, which is the trade.

Copy it in, or mount it Writing a file into the memory file system holds its entire contents in linear memory. Mounting it through WORKERFS lets the module read ranges on demand, so peak memory is bounded by what the code buffers. FS.writeFile — copied 2 GB resident in linear memory — fails on any device WORKERFS — mounted read buffer the rest stays in the browser's file object, read on demand Code using fseek and fread works unchanged against either; only the memory profile differs, which is why this is a mount choice rather than a code change.

Lazy loading a file over the network

Between preloading everything and mounting a local file there is a third option: a virtual file backed by an HTTP range request, so the module reads parts of a remote file without downloading all of it.

Module.FS.createLazyFile('/data', 'atlas.bin', '/assets/atlas.bin', true, false);
// the module can now fopen("/data/atlas.bin") and seek anywhere in it

The file appears to the module as an ordinary read-only file. Behind it, each read issues a range request for the bytes the module asked for, which means a hundred-megabyte asset costs only the regions actually touched. A texture atlas, a font collection, a database index and a sprite sheet are all natural fits.

Two conditions have to hold. The server must honour Range requests and expose Accept-Ranges, and — in the browser’s main thread — the underlying request is synchronous, which is deprecated and blocks. In a worker the synchronous request is permitted and does not block anything the user sees, which is another reason this pattern belongs off the main thread.

Measure before adopting it. If the module’s access pattern touches most of the file anyway, the range requests are slower than one sequential download and the lazy file has made things worse. It wins when access is genuinely sparse, which is exactly when a naive preload would have been wasteful.

Expected output

A working setup shows the module reading and writing paths as if they were real:

[fs] mounted IDBFS at /data
[fs] syncfs(populate) restored 3 files, 1.2 MB
[app] opened /assets/config.txt
[app] wrote /work/output.bin (412 kB)
[fs] syncfs(persist) wrote 4 files
console.log(Module.FS.readdir('/data'));
// [ '.', '..', 'project.db', 'settings.json', 'cache.bin' ]
console.log(Module.FS.stat('/data/project.db').size);
// 1048576
Which file system to mount where MEMFS lives in linear memory and vanishes on reload. IDBFS persists through IndexedDB but needs an explicit sync. NODEFS maps a real directory and only exists outside the browser. MEMFS (default) in linear memory; fast, counts against the heap, gone on reload IDBFS persists via IndexedDB; requires an explicit FS.syncfs call NODEFS maps a real directory; Node only, never available in a browser --preload-file baked into a .data file fetched beside the binary Preloaded data downloads before main() runs, so a large bundle delays startup rather than the first read. FS.syncfs is asynchronous both ways; a write not followed by a sync is lost on the next reload.

Gotchas

  • FS is not defined. Not exported. Add FS to EXPORTED_RUNTIME_METHODS.
  • syncfs direction reversed on startup. Overwrites persisted data with an empty file system.
  • Files never unlinked. MEMFS contents occupy linear memory for the life of the instance.
  • Preloading hundreds of megabytes. All of it downloads before main runs.
  • WORKERFS on the main thread. Not available; it is worker-only by design.
  • Relying on beforeunload to sync. Not reliably fired; sync on visibilitychange and periodically.

Performance note

Writing a 100 MB file with FS.writeFile took about 90 ms and added 100 MB to the heap permanently until unlinked. Mounting the same file through WORKERFS took under a millisecond and added nothing, with reads costing roughly the same per byte as reading from memory once the browser’s cache was warm. For anything above a few tens of megabytes the mount is not an optimisation but the only workable approach.

Frequently Asked Questions

Can I use OPFS instead? Emscripten has an OPFS backend, and it is attractive for large persistent data because it avoids IndexedDB’s overhead. It carries the same worker requirement as OPFS generally — see persisting a Wasm database to OPFS for how that access model works.

Is there a quota on the persisted data? Yes, the origin’s storage quota, shared with everything else the origin stores. Request persistent storage and check navigator.storage.estimate() before writing large amounts, and handle a failed write rather than assuming space is available.

How do I test file-handling code outside a browser? Build for Node and mount NODEFS over a fixture directory. The same C code then runs against real files in an ordinary test, which is far faster to iterate on than a browser test.

Can I avoid the file system entirely? Often, and it is worth considering. If the C code’s only use of files is reading one input and writing one output, exposing functions that take a pointer and a length is smaller, faster and removes the whole virtual file system from the build — -sFILESYSTEM=0 then strips several kilobytes of runtime. The file system earns its place when the code genuinely walks directories, seeks, or uses a library that insists on paths.

Does the file system work with pthreads? Yes, with the usual caveats about which thread performs the operation. Keep file access on one thread rather than sharing handles across them.

The rule of thumb worth remembering: MEMFS for scratch, IDBFS for what must survive, WORKERFS for what is too big to copy, and no file system at all when the code can take a pointer instead.

← Back to C/C++ to Wasm with Emscripten