Plugin Systems & Extensibility
A plugin system is a promise: third parties write code, your application runs it, and neither side can break the other too badly. WebAssembly is unusually well suited to that promise because a module arrives with no ambient authority, a declared import list you can inspect before running anything, and memory that cannot reach yours. What it does not give you is an interface — and the interface is where plugin systems succeed or fail.
Prerequisites
- [ ] A part of your application that genuinely benefits from third-party extension.
- [ ] A host environment with resource limits: a worker in a browser, or a runtime with fuel on a server.
- [ ] A versioning plan, before the first external plugin exists.
- [ ] Willingness to support the interface you publish for a long time.
Deciding what should be a plugin at all
Not every extension point deserves a compiled module. The machinery has a cost — a sandbox, a versioned interface, documentation, a distribution path, a support burden — and it is worth paying only where the extension genuinely needs to run code.
Three categories reliably justify it. Transformations over data your application holds: a custom export format, a filter, a scoring function, a linter rule. Integrations with systems you do not know about, where the plugin translates between your model and something external. And domain logic that varies per customer in ways configuration cannot express — pricing rules, validation, routing decisions.
Three categories reliably do not. Anything expressible as configuration should be configuration: a declarative rule is easier to write, safer to run, and possible to validate ahead of time. Anything needing rich access to your application’s internals will end up with an interface so wide that the sandbox stops meaning anything. And anything performance-critical on your own hot path is better as code you control, because a plugin’s cost is unpredictable by construction.
A useful test before committing: write three plugins yourself, for three genuinely different purposes, against the interface you are proposing. If any of them needs a capability you did not plan for, the design is not finished. If all three are trivial, the extension point may not need code at all.
The interface is the product
Everything a plugin author experiences is the interface: what they implement, what they can call, and how data crosses. Once published, changing it breaks every plugin in existence, so the design deserves more care than the runtime plumbing around it.
Three shapes are common. A function interface exports a handful of well-known symbols that the host calls — the simplest and the most rigid. An event interface lets the plugin register for hooks and receive callbacks, which suits editors and pipelines. A component interface described in an interface definition language generates bindings for both sides, which is where the ecosystem is heading via the component model and WIT.
;; the minimal function interface: the host calls run, the plugin calls back through host imports
(module
(import "host" "log" (func $log (param i32 i32)))
(import "host" "read_input" (func $read (param i32 i32) (result i32)))
(memory (export "memory") 2)
(func (export "abi_version") (result i32) (i32.const 1))
(func (export "run") (param i32) (result i32) ...))
Host functions are capabilities, not conveniences
Every function you import into a plugin is permission granted permanently. Once an interface offers
http_get, every plugin can make network requests and you cannot take it back without breaking them.
Design the import list as if it were a permissions manifest, because it is one. Start with the smallest
set that makes a useful plugin possible, and prefer narrow, purpose-built functions over general ones: a
read_current_selection that returns the text the user has selected is far safer than a
read_document(path) that takes an arbitrary path, and it is also easier for plugin authors to use
correctly.
Where a capability is genuinely needed but dangerous, mediate it. A plugin that needs network access can
be given fetch_from_allowlist(index) where the host holds the list of permitted endpoints. The plugin
gets its functionality, you keep the decision.
Data crossing: keep it boring
Plugin interfaces attract clever serialisation schemes and should resist them. The plugin and the host
share nothing but bytes in linear memory, and the two ends may be written in different languages, so
the encoding must be trivially implementable everywhere.
The pattern that works is length-prefixed byte buffers with an explicit allocator on the plugin side. The host asks the plugin to allocate, writes bytes into the returned pointer, calls the entry point, and reads a pointer and a length back. JSON inside those buffers is usually fine — it costs a parse but every language has one — and a binary format is worth it only when profiling says so.
Versioning, and the promise you are making
The first external plugin is the moment your interface becomes a public API. From then on, three rules keep the ecosystem from breaking.
Export a version the host checks before the first call, and refuse to run a plugin whose version you do not understand — a clear error at load time is infinitely better than undefined behaviour later. Add capabilities by adding new imports that plugins may optionally use, never by changing an existing signature. And when a breaking change is unavoidable, run both interface versions side by side for a period long enough for authors to migrate, with the host selecting by the declared version.
The versioning guide works through the mechanics, including how to detect optional imports without breaking older modules.
Resource limits are not optional
A plugin that loops forever, allocates without bound, or calls your host functions a million times a second is not necessarily malicious — it is more often a bug — and the outcome is the same either way.
In a browser the only reliable control is a worker you terminate on a deadline. On a server, a runtime with fuel metering or epoch interruption pre-empts the module cleanly and tells you why. Memory is bounded by instantiating with an explicit maximum, and host-side accumulation — logs, output buffers, event queues — must be capped by the host because the plugin has no reason to.
Distribution and trust
A plugin has to reach your application from somewhere, and that path is part of the design. A registry you operate lets you scan modules, check their import lists automatically, and revoke a bad one. A sideloading path — the user picks a file — is simpler and pushes the trust decision to the user, which is appropriate for a developer tool and not for a consumer product.
Whatever the path, validate before running. WebAssembly.Module.imports tells you exactly what a module
wants before a single instruction executes, and rejecting a module that asks for something outside your
allowed set is a five-line check that catches both attacks and mistakes.
Signing is worth considering once a registry exists: the host verifies a signature over the module bytes before compiling, so a compromised distribution channel cannot substitute a different module. It is more machinery than most systems need on day one and much easier to add before there are ten thousand plugins than after.
Discovery, configuration and the user’s mental model
The runtime story is only half a plugin system. Users have to find plugins, understand what they do, configure them, and know which one produced a result they did not expect.
Metadata should travel with the module rather than beside it. A custom section, or a small manifest the host reads before instantiating, carries the name, version, author, a description and the list of capabilities the plugin claims to need. Comparing the claimed capabilities against the module’s actual import list is a cheap and useful consistency check — a plugin that declares no network access and imports a fetch function is either lying or built wrong, and either way you do not want to run it.
Configuration is best expressed as data the host passes in rather than as state the plugin stores. A
plugin that reads its settings from a buffer at the start of each invocation is stateless, testable and
trivially safe to run per tenant. A plugin that keeps configuration in linear memory between calls
needs its instance kept alive, which reintroduces every isolation question you had just answered.
Attribution matters more than it first appears. When a plugin transforms a document, throws an error, or takes four seconds, the user needs to know which one did it. Record the plugin identity alongside every result and every log line, surface it in errors, and expose per-plugin timing in whatever diagnostics you offer. Systems that skip this end up with support conversations that begin with nobody knowing which of eleven installed plugins is responsible.
Performance: instantiate, do not compile
Compilation is expensive and instantiation is cheap. Compile a plugin’s module once — ideally caching the compiled result — and instantiate per invocation or per tenant. Instantiating a compiled 200 kB module takes well under a millisecond, which makes a fresh instance per call entirely affordable and removes every question about state leaking between invocations.
Where invocations are frequent and small, the host call overhead starts to matter. Batch where the interface allows it: a plugin that processes an array of items in one call is dramatically faster than one called per item, and designing the interface that way from the start is free.
Errors, and telling three kinds apart
A plugin invocation can fail in three ways, and conflating them produces a system nobody can operate.
A plugin error is the plugin reporting that it could not do the job — an unsupported input, a validation failure, a missing configuration value. This is a normal outcome and should travel back as a structured result the host can present. Reserve a return code or a result variant for it rather than using a trap, because a trap discards any explanation the plugin wanted to give.
A plugin fault is the module misbehaving: a trap from an out-of-bounds access, a division by zero, an
unreachable from a panic, an exhausted memory. The host catches these as exceptions, records which
plugin faulted, and disables it after a few repetitions rather than letting a broken plugin fail every
request forever. A fault is the plugin author’s bug and the message should say so clearly enough that
they can act on it.
A host error is your side failing: the module could not be fetched, the interface version is unsupported, a limit was misconfigured. These need different handling entirely, because retrying will not help and the plugin author cannot fix them.
try {
const res = await runPlugin(id, input, { timeoutMs: 2000 });
if (!res.ok) return { kind: 'plugin-error', plugin: id, message: res.message };
return { kind: 'ok', value: validate(res.value) };
} catch (e) {
if (e.name === 'RuntimeError') { recordFault(id, e); return { kind: 'plugin-fault', plugin: id }; }
return { kind: 'host-error', error: e };
}
Keeping the three separate in logs and metrics is what makes a plugin ecosystem operable: you can see at a glance whether a spike is one broken plugin, a bad input distribution, or something wrong with your own deployment.
Gotchas and failure modes
- A capability added casually. Every import is permanent. Think of the list as a manifest.
- No version export. The first breaking change becomes a support crisis.
- State reused between invocations. Residual
linear memoryleaks data between tenants and makes plugins behave differently depending on what ran before. - Unbounded host-side buffers. A plugin that logs in a loop exhausts the host, not itself.
- Trusting plugin output. Validate it exactly as you would input from the network.
- Blocking the main thread. In a browser, plugins belong in a worker regardless of how fast you expect them to be.
Verifying a plugin system
The tests that matter are adversarial. Write plugins that misbehave deliberately and assert that the host survives each one:
test('infinite loop is terminated', async () => {
await expect(runPlugin(loopForeverWasm, {}, { timeoutMs: 500 })).rejects.toThrow('timeout');
});
test('memory bomb fails cleanly', async () => {
await expect(runPlugin(allocateForeverWasm, {})).rejects.toThrow(/memory/);
});
test('unexpected import is rejected before execution', async () => {
await expect(loadPlugin(wantsNetworkWasm)).rejects.toThrow(/unexpected import/);
});
test('output is validated', async () => {
await expect(runPlugin(returnsGarbageWasm, {})).rejects.toThrow(/invalid result/);
});
Keep those fixtures in the repository. They are the regression suite for the property that actually matters, and they will catch the day someone removes a limit to make a test pass.
Guides in this topic
- Designing a Wasm plugin interface — exports, imports and data crossing.
- Loading untrusted plugins safely — validation, isolation and the load path.
- Versioning a Wasm plugin API — evolving without breaking.
- Building a plugin host with Extism — using an existing framework instead.
- Limiting plugin CPU and memory use — fuel, epochs, deadlines and caps.
Frequently Asked Questions
Why WebAssembly rather than sandboxed JavaScript? Because containment is structural rather than contractual. A JavaScript sandbox has to enumerate and remove capabilities from a large platform surface; a WebAssembly module starts with none and receives only what you import. It also lets plugin authors use whichever language they prefer.
Can plugins be written in any language? Any language that compiles to WebAssembly, which today includes Rust, C, C++, Go, Zig, AssemblyScript and, with more overhead, Python and JavaScript via embedded interpreters. Publish a template for each language you intend to support — nothing grows a plugin ecosystem faster.
Should I use the component model? If you are starting now and can accept the tooling’s maturity, it removes most of the hand-written marshalling and gives you a typed interface definition. For a system that must work in every browser today, the plain module interface remains the pragmatic choice.
How do I debug a plugin I did not write? Give authors a local harness that runs their module against the same host implementation with verbose logging. Most plugin bugs are interface misunderstandings, and a harness that reproduces the host’s behaviour exactly turns support requests into self-service.
How many plugins can run at once?
As many as you have instances for, but the useful limit is lower. Each instance holds its own
linear memory, and each concurrent execution needs a worker or a runtime thread. A pool sized to the
core count, with a queue in front of it, gives predictable behaviour under load and prevents a burst of
plugin work from starving the application itself.
Should plugins be able to call each other? Prefer not, at least initially. Chaining plugins through the host keeps the composition explicit, inspectable and cancellable, and it avoids a plugin taking a dependency on another plugin’s presence and version. If composition is genuinely needed, express it as a pipeline the host runs rather than as direct calls between modules.
What does a good plugin failure look like to a user? A named plugin, a short explanation, and the application continuing to work without it. Anything that stops the whole feature because one optional extension misbehaved will be remembered as your bug rather than the plugin author’s.
Related
- Cryptography & Untrusted Code — the containment story in depth.
- Wasm Component Model & WIT Bindings — the typed interface future.
- Serverless & Edge Deployment — running the same modules server-side.
← Back to Production Wasm: Workloads & Deployment