Plugins

Plugin design and testing

Design a production plugin around one domain, restartable lifecycle hooks, explicit control operations, and honest compatibility evidence.

A plugin is trusted TypeScript that implements an executable local service. It owns the provider behavior, durable model, public API, control operations, and any callbacks it sends. The runtime owns instance isolation, storage locations, lifecycle orchestration, operation adapters, tracked transport, virtual time, and control authentication.

This guide is the design contract for a plugin intended to grow without splitting into parallel implementations. For a guided first implementation, build the status plugin.

Start with a behavior claim

Do not begin by copying an entire provider API. Write down the application behavior the first version must support:

  • which real client, SDK, or HTTP calls the application makes;
  • which state transitions those calls cause;
  • which callbacks, signatures, delays, or retries the application observes;
  • which privileged actions a test needs to arrange or inspect that world;
  • what is deliberately unsupported.

That claim determines the first coherent slice. An endpoint returning a plausible object is not a slice if the next request cannot observe the resulting state. A broad route inventory backed by fixtures usually creates more compatibility claims than it proves.

Keep the public surface and control surface distinct. Public routes imitate the integration used by the application. Operations arrange, trigger, and inspect the local world. Both must reach the same domain behavior. Operations and emulated APIs develops this boundary, while Compatibility is a tested claim defines the evidence required before saying a provider behavior is supported.

Keep the dependency direction simple

provider-shaped routes ─┐
                        ├── domain behavior ── plugin persistence
control operations ─────┘           │
                                    └── callback or delivery adapter

Routes parse provider-shaped requests and map provider-shaped responses. Operations validate developer-facing inputs and map stable control results. Neither adapter calls the other. Domain services own rules and transactions; persistence owns durable representation; delivery code owns request encoding, signing, and retry policy.

This separation matters most when behavior grows. If a route and an operation each implement "create customer," they will eventually disagree about validation, IDs, events, or transactions. If both call one domain method, their response shapes can differ without creating two services.

Split files when responsibilities have different reasons to change, not merely to create more folders. A small plugin can stay in one file. A larger one will commonly have an API adapter, operations, domain services, persistence, lifecycle, and delivery code. The dependency direction above should remain visible regardless of file count.

Importing the plugin, or a config that mounts it, must be inert. It may construct schemas, route tables, operation definitions, and frozen descriptors. It must not change the working directory or environment, write files, print output, open a database, start a server, or leave a timer or socket running. Open process resources in start, where the runtime can own their lifetime. If tests need fault injection, accept narrow dependencies in a plugin factory instead of adding test-only routes, operations, or response fields.

Only import authoring APIs from the localhost2137 package root. A missing public capability is a contract question, not a reason to couple a plugin to runtime internals.

Minimal authoring shape

This checked file-backed service is a small, working public authoring shape, not the entire authoring API. It puts durable state below both a Hono route and typed operations:

examples/status-plugin/src/status-plugin.ts
import { readFile, writeFile } from "node:fs/promises";
import { Hono } from "hono";
import { defineOperation, definePlugin, type PluginEnv } from "localhost2137";
import { z } from "zod";

const statusSchema = z.object({
	message: z.string().nullable(),
	state: z.enum(["operational", "degraded", "outage"]),
});
const setStatusInput = z.object({
	message: z.string().optional(),
	state: statusSchema.shape.state,
});

type Config = Readonly<Record<string, never>>;
type State = Readonly<{ statusPath: string }>;
type Status = z.output<typeof statusSchema>;

const initialStatus: Status = { message: null, state: "operational" };
const operation = defineOperation<"status", State, Config>();

const readStatus = operation({
	description: "Read the current status",
	input: z.object({}),
	output: statusSchema,
	run: (context) => loadStatus(context.state.statusPath),
});

const setStatus = operation({
	description: "Set the status exposed to the application",
	input: setStatusInput,
	output: statusSchema,
	run: async (context, input) => {
		const status: Status = {
			message: input.message ?? null,
			state: input.state,
		};
		await saveStatus(context.state.statusPath, status);
		return status;
	},
});

const api = new Hono<PluginEnv<State, Config>>();
api.get("/v1/status", async (context) => {
	const { state } = context.get("lh");
	return context.json(await loadStatus(state.statusPath));
});

export const statusPlugin = definePlugin({
	api,
	configSchema: z.object({}),
	connection: ({ baseUrl, instanceId, serviceKey }) => {
		const apiUrl = `${baseUrl}/${instanceId}/${serviceKey}`;
		return {
			env: { STATUS_API_URL: apiUrl },
			values: { apiUrl },
		};
	},
	description: "Local status service",
	id: "status",
	lifecycle: {
		create: (context) => saveStatus(context.storage.path("status.json"), initialStatus),
		start: (context): State => ({
			statusPath: context.storage.path("status.json"),
		}),
	},
	operations: { readStatus, setStatus },
	stateVersion: 1,
});

async function loadStatus(path: string): Promise<Status> {
	return statusSchema.parse(JSON.parse(await readFile(path, "utf8")));
}

async function saveStatus(path: string, status: Status): Promise<void> {
	await writeFile(path, `${JSON.stringify(status)}\n`, "utf8");
}

Bind defineOperation once for one literal plugin ID, state type, and config type. The runtime rejects unbound operations or operations produced by mixed binders. The Hono app is a shared route table, so request handlers read the selected instance from context.get("lh"); they never close over mutable instance state.

Identity and configuration are contracts

The plugin ID and configured service key answer different questions:

  • id identifies the implementation that owns stored data. It is stable across mounts and releases.
  • The service key is chosen by the user. It selects the route, storage namespace, CLI target, and typed instance property for one mount.

Changing a plugin ID under an existing key is rejected rather than opening another plugin's data. Renaming a service key creates a different mount; it is not a data migration. Both IDs are lowercase, URL-safe names, but they do not need to be equal.

configSchema defines deployment choices rather than mutable world state. Parsed config and seed values must end as JSON-compatible plain data and are frozen before callbacks receive them. Keep runtime resources, test doubles, clocks, and callback functions out of config; pass private factory dependencies when the implementation genuinely needs them.

Treat operation keys and descriptions as public authoring metadata. Keys are camel-case JavaScript identifiers and must remain distinct after conversion to kebab-case CLI names. Descriptions and Zod field descriptions should say what an action means, because humans, generated CLI help, and coding agents all read them.

Configuration covers mount identity, config validation, and what changes do to existing worlds.

Lifecycle and durable state

Lifecycle hooks are ownership boundaries, not convenient startup callbacks:

HookContract
createInitialize absent durable state. It may run again after an interrupted attempt, so repeating it must be safe. Do not open long-lived process resources.
updateMigrate stopped durable state from from to to. A failed pair can be retried unchanged.
startOpen resources for every active process generation and return live State. Clean up anything already opened if this hook itself throws.
onStartedReconcile durable running work after every configured service has started and before the instance becomes ready.
seedApply schema-validated baseline state only when seeding is explicitly requested.
onTimeAdvancedReconcile one already committed time window idempotently by advanceId, from, and to.
stopClose resources belonging to one successful start, including during cleanup after a later startup hook fails.

A production lifecycle can keep each ownership rule visible in one place. This checked example opens and closes a database, migrates stopped storage, seeds transactionally, and reconciles durable work at startup and after committed clock advancement:

plugins/stripe/src/lifecycle.ts
import type { Lifecycle, RunningPluginContext } from "localhost2137";
import type { StripeConfig, StripeSeed } from "./config.js";
import { createStripeServices, seedStripeServices } from "./domain/stripe-services.js";
import { StripeDatabase } from "./persistence/database.js";
import { assertCurrentDatabaseVersion, migrateDatabase } from "./persistence/migrations.js";
import type { StripePluginDependencies } from "./plugin-dependencies.js";
import type { StripeState } from "./state.js";
import { StripeWebhookDispatcher } from "./webhooks/webhook-dispatcher.js";

type StripeLifecycle = Lifecycle<StripeState, StripeConfig> & {
	readonly seed: (
		context: RunningPluginContext<StripeState, StripeConfig>,
		seed: StripeSeed,
	) => Promise<void> | void;
};

export function createStripeLifecycle(dependencies: StripePluginDependencies): StripeLifecycle {
	return {
		create(context) {
			dependencies.recordLifecycle?.("create");
			dependencies.beforeCreate?.(context);
			withDatabase(context.storage.path("stripe.sqlite"), (database) => {
				migrateDatabase(database.raw());
			});
		},
		async onTimeAdvanced(context, advance) {
			const eventIds = context.state.services.billing.reconcileTimeAdvance(advance);
			await dependencies.afterTimeReconciled?.(context, advance);
			await context.state.webhooks.reconcile(context, eventIds);
		},
		async onStarted(context) {
			await context.state.webhooks.reconcile(
				context,
				context.state.services.billing.pendingWebhookEventIds(),
			);
		},
		seed(context, seed) {
			dependencies.recordLifecycle?.("seed");
			context.state.database.transaction(() => {
				seedStripeServices(context.state.services, seed, context.clock.now());
			});
		},
		start(context) {
			dependencies.recordLifecycle?.("start");
			const database = new StripeDatabase(context.storage.path("stripe.sqlite"));
			try {
				assertCurrentDatabaseVersion(database.raw());
				return Object.freeze({
					database,
					services: createStripeServices(database, context.config),
					webhooks: new StripeWebhookDispatcher(database, context.config, {
						...(dependencies.webhookDeliveryTimeoutMs === undefined
							? {}
							: { timeoutMs: dependencies.webhookDeliveryTimeoutMs }),
					}),
				});
			} catch (cause) {
				database.close();
				throw cause;
			}
		},
		stop(context) {
			dependencies.recordLifecycle?.("stop");
			dependencies.beforeStop?.(context);
			context.state.database.close();
		},
		update(context, version) {
			dependencies.recordLifecycle?.(`update:${version.from}:${version.to}`);
			withDatabase(context.storage.path("stripe.sqlite"), (database) => {
				migrateDatabase(database.raw());
			});
		},
	};
}

function withDatabase(path: string, work: (database: StripeDatabase) => void): void {
	const database = new StripeDatabase(path);
	try {
		work(database);
	} finally {
		database.close();
	}
}

Only create and start are required. A failed start does not earn a later stop, because the service never reached running state. Use try/catch inside start to close a database, listener, or other resource opened before the failure. Once start succeeds, the runtime calls stop during normal shutdown and startup cleanup; stop should close exactly what the returned live state owns.

onStarted is for recovery that requires a running service, such as draining a durable outbox. It runs after all services have started, so every configured service has reached running state. If it fails, readiness fails and the runtime stops the services that started. Make recovery idempotent: a process can end after an external effect but before the durable acknowledgement.

Pair seedSchema with lifecycle.seed, or omit both. Seeding is not an automatic startup migration, and a hook is not automatically rolled back if it partially changes storage. Perform related durable changes in a plugin-owned transaction where possible. Recovery depends on whether seed was requested during create, reset, or in place; follow the outer mutation.

Version the durable format, not the package

stateVersion is a positive integer describing the plugin's stored representation. It is not the npm version and should not change for route additions, operation additions, or internal refactors that leave old data readable.

When stored state is older, the runtime calls update before start. update may receive any older stored version that the release claims to support, so apply the required migrations in order rather than assuming only the immediately previous version. The runtime records the new version only after the hook succeeds; if the hook fails, the same from and to can arrive again. Migrations therefore need restartable steps or a storage engine that applies them transactionally.

A newer stored version is rejected rather than silently downgraded. Keep real historical fixtures for every supported upgrade path and verify that user-visible data survives. A state-version-1 plugin has no honest predecessor: do not publish version 2 or invent an old schema merely to satisfy a generic upgrade test.

Separate live state from durable state

The State returned by start is live process state, not the durable format. It can contain open database handles, repositories, domain services, and delivery dispatchers. It is discarded after stop; anything that must survive restart belongs under context.storage.

Call context.storage.path() with a portable relative path. Absolute paths, empty or dot segments, and traversal outside the service root are rejected. Do not derive storage from the current working directory or manually include the instance ID or service key—the runtime already selects the isolated root.

Own time and asynchronous work

All lifecycle, route, and operation code receives the instance clock through context.clock.now(). Use it for domain timestamps, due dates, and retry schedules. Using new Date() for those values creates a second time source and breaks pinned-clock tests.

Delivery attempt timeouts and retry policy belong to the plugin. A wall-clock timeout can keep one network attempt from hanging; it is an operational limit and does not advance with the instance clock. The runtime and control operations can impose separate wall-clock safety limits around their own work, but those limits do not define the plugin's provider-facing callback behavior.

The base context used by create, update, and start deliberately has no live state, fetch, or task tracker. The running context used by routes, operations, seed, onStarted, onTimeAdvanced, and stop adds them.

Use context.fetch for outbound HTTP. It is tracked automatically, carries the instance abort signal, and produces delivery diagnostics. Register every other promise that may outlive the current stack with context.tasks.track(label, promise). Await or return work needed for the current result; when scheduling background delivery, handle its rejection deliberately. An untracked fire-and-forget promise makes instance.idle() lie and can continue after reset or shutdown.

Durable asynchronous behavior needs both halves: persist enough intent to retry before releasing the current action, then reconcile pending work in onStarted or onTimeAdvanced. Task tracking makes the current process observable; it does not make an in-memory promise survive a crash.

Honor context.signal in long-running work and use context.log.info() for plugin diagnostics instead of printing to stdout. Keep credentials and provider payloads out of log attributes. Virtual time and asynchronous work covers scheduling, idempotency, and idle semantics in depth.

Derive connection metadata per mount

connection() runs for a concrete baseUrl, instanceId, serviceKey, and parsed config. Derive URLs from those values; do not hard-code the default port, the dev instance, or the assumption that the mount key equals the plugin ID.

Return two deliberate projections:

  • values is JSON-compatible typed metadata for instance[serviceKey].connection;
  • env contains only string values suitable for an application process.

Keep the two projections consistent when they expose the same endpoint. Environment names are global to the child process, so use stable plugin-specific names and test two mounts for collisions. The user can set exportEnv: false for one mount without removing typed connection values. Connection values and environment export defines the complete merge behavior.

Keep routes and operations honest

Operation inputs are Zod objects. Outputs may use any Zod schema, but the parsed result must be JSON-compatible because it can cross typed, CLI, and HTTP adapters. Output validation happens after run; a schema that describes the intended result but a function that returns something else is a plugin defect, not a client error.

Translate expected domain failures at the adapter boundary:

  • a provider route returns that provider's status, headers, and error body;
  • an operation throws LocalhostError with a stable upper-snake-case code, safe message, HTTP status, and only useful JSON details.
plugins/slack/src/operations.ts (source excerpt)
function runSlackOperation<Value>(
	dependencies: SlackPluginDependencies,
	operation: string,
	context: RunningPluginContext<SlackState, SlackConfig>,
	run: () => Value,
): Value {
	dependencies.beforeOperation?.(operation, context);
	try {
		return dependencies.transformOperationResult
			? dependencies.transformOperationResult(operation, run())
			: run();
	} catch (cause) {
		if (!(cause instanceof SlackError)) throw cause;
		throw new LocalhostError(`SLACK_${cause.code.toUpperCase()}`, cause.message, {
			cause,
			details: { slackError: cause.code },
			status: slackOperationStatus(cause),
		});
	}
}

function slackOperationStatus(error: SlackError): number {
	if (error.code === "channel_not_found" || error.code === "user_not_found") return 404;
	if (error.code === "name_taken" || error.code === "not_in_channel") return 409;
	return 400;
}

Keep the original error as cause for internal diagnosis, but do not expose secrets in the safe message or details. Unexpected operation errors become the generic PLUGIN_EXECUTION_FAILED boundary. Do not make public routes return control-plane error shapes merely because operations use them.

Test application-facing routes using real encodings: query strings, form bodies, headers, pagination, signatures, and target SDK behavior where relevant. Hono makes route construction convenient; it does not prove provider compatibility.

Contract and semantic tests

@localhost2137/plugin-testkit supplies shared runtime-integration cases. Register them with the test runner so each behavior remains visible:

plugins/slack/test/slack-contract.test.ts
import { describe, it } from "vitest";
import { createPluginContractCases } from "@localhost2137/plugin-testkit";
import { slackContractFixture } from "./contract/slack-contract-harness.js";

describe("Slack plugin contract", () => {
	for (const contractCase of createPluginContractCases(slackContractFixture)) {
		it(contractCase.name, contractCase.run, 30_000);
	}
});

The fixture asks one production plugin factory to build base and deliberately faulted variants. The testkit then owns the runtime, instances, daemon processes, control calls, cleanup, and assertions. It checks import purity, schema paths, operation introspection and validation, Hono context, multi-instance isolation, connection values, environment collisions, tracked fetch, seed and reset, storage containment, lifecycle recovery, restart persistence, upgrades, and future-version rejection. Optional durability cases exercise pending delivery and committed time-advance recovery in real child processes.

Fault variants may inject a dependency or historical state version. They must not replace the production factory or add a test-only public surface. Otherwise the contract proves a harness that users never run.

The current durability fixture requires positive versions ordered old < current < future. For a new version-1 plugin, use focused public-surface tests for creation, restart persistence, reset, future-version rejection, and the other applicable behaviors. Add the shared upgrade case when a real prior format exists.

The generic contract proves runtime integration, not provider fidelity. Add semantic suites for the exact supported endpoints and operations, domain transactions, ID and timestamp rules, failure shapes, callback signatures, retry schedules, SDK behavior, restart recovery, and every meaningful unsupported boundary. A useful release claim should point to one of those executable tests rather than to the amount of code in the plugin.