Diagnose a failing scenario

Reproduce one fixed world and stop at the first config, control, provider, delivery, or application boundary that differs.

Keep the config path, instance, service, application action, and expected effect fixed while you diagnose. A passing rewrite that bypasses the original SDK, HTTP, or callback path is not a fix.

Run one concrete triage loop

The commands below target the crash-course localhost.config.ts, its slack service, persistent dev instance, and src/read-workspace.ts application.

Before starting or changing anything, inspect the project:

pnpm exec localhost doctor --json

doctor is read-only. It loads config, inspects runtime discovery and stored manifests, and reports:

JSON pathDecision it supports
config.loaded, config.path, config.storageRootIs this the intended project and storage root?
runtime.state, runtime.urlIs the daemon absent, healthy, or unhealthy?
storage.instances[].id, .status, .seedStatus, .servicesDoes dev exist with slack and a usable lifecycle state?
issues[].code, .instanceId, .serviceKey, .messageWhich project or stored-world boundary needs action first?

An absent daemon is normal before dev starts. If config loaded and no earlier project issue blocks startup, keep it in the foreground:

# Terminal 1
pnpm exec localhost dev

From another terminal, walk the selected world from control metadata to the real application path:

# 1. Prove that this daemon exposes slack on dev
pnpm exec localhost describe slack --instance dev --json

# 2. Inspect the app-facing projection; do not paste it into a report
pnpm exec localhost env --instance dev --format json

# 3. Reproduce the provider-shaped application request
pnpm exec localhost run --instance dev -- pnpm dev

# 4. Read only this service's latest bounded evidence
pnpm exec localhost logs slack --instance dev --tail 50 --json

Stop at the first command whose result differs from the scenario:

First failureBroken boundaryNext evidence
doctor cannot load the intended configConfig discovery or importSelected config path and issue code
doctor reports absent/unhealthy runtime after dev should be readyRuntime discoveryRuntime issue code and foreground daemon error
describe slack failsControl access, dev selection, or service mountExisting instance/service names in the error or doctor report
env points elsewhereConnection projectionInstance and service segments in the exported URL
The app request failsApplication adapter or provider-shaped routeCompare the app error with the presence or absence of a matching request entry. Absence keeps the failure at the application boundary; a present entry can be traced inside the provider-shaped boundary.
The app succeeds but a callback is missingDelivery or receiverdelivery entry, receiver evidence, then tracked completion
Local behavior differs from a claimed provider routeCompatibility surfaceInstalled plugin's documented support entry

The service key is the property under services, not necessarily a package name. describe and generated exec help are the installed control contract. Do not guess operation names from another plugin version.

Trace each boundary through logs

The log command already filters by instanceId: dev, serviceKey: slack, and at most 50 newest entries. Its JSON contains a top-level droppedEntries count and bounded entries with these useful fields:

kind
status
message
correlationId
instanceId
serviceKey
wallTime
virtualTime
attributes

The runtime emits request, operation, delivery, and plugin entries. status, message, and safe attributes distinguish work within the same second. wallTime is when the entry was recorded; virtualTime is the selected world's clock.

Correlation IDs are scoped to one boundary, not propagated through a whole scenario. A control request and the operation it invokes share one ID. A provider request, each delivery, a plugin log, and a later idle request can have different IDs. Move between them using the fixed instance, service, time, and safe resource/event IDs. Do not search one correlation ID end to end.

Logs are a bounded ring, not an audit trail. If droppedEntries is non-zero, reproduce the smallest scenario and read the new tail. Body-like attributes are omitted and known sensitive values are redacted, but review every retained entry before sharing it.

Treat asynchronous failures as evidence

When the action schedules callbacks or plugin-owned background work, the test should call await instance.idle(). A rejection identifies a control failure, timeout, cancellation, or tracked work failure. Inspect nearby delivery and plugin entries for the same world and time; the runtime does not emit a separate task log entry.

A successful idle wait proves only that plugin-tracked work is idle. The application may still own a queue, transaction, or job, so wait for its explicit completion boundary before asserting. Work due at a later virtual time also needs a positive clock advance, not a wall-clock sleep. Virtual time and asynchronous work covers committed windows and reconciliation.

For LIFECYCLE_CONFLICT, read the message and current instance/clock state before choosing another transition. INSTANCE_MUTATION_COMMITTED means the mutation took effect but finalization failed; do not retry it blindly. Seed failures have action-specific recovery rules in Seed test worlds.

Handle project failures

A config fingerprint mismatch means the daemon resolved different config; save the intended file and restart that daemon. CONFIG_DEFAULT_EXPORT_MISSING means the selected module loaded without a default config. CONFIG_IMPORT_FAILED means that module or one import threw. Its serialized diagnostic identifies the selected config path but omits the retained private cause, so inspect that module and its imports directly.

Recover a storage lock safely

One live daemon owns a storage root. If startup reports a live owner, use it, stop it normally, or choose another storage root. Never remove that lock. If the error reports a stopped owner or damaged lock, first confirm no runtime or test process uses the root. Only then remove the exact lock path printed by the error. Do not turn it into a wildcard or delete .localhost2137 to silence a problem.

Reset or destroy only an explicitly selected world when that lifecycle action is the intended recovery. Preserve unexpected manifests and pending cleanup evidence until it is recorded.

Copy a useful redacted report

diagnostic-report.md
# localhost2137 diagnostic report

## Scope

- Config path (shortened before sharing):
- Instance ID: `dev`
- Service key: `slack`
- Application action: `pnpm exec localhost run --instance dev -- pnpm dev`
- Expected application result:
- Observed application result:

## Versions

- Node.js:
- localhost2137:
- Emulator plugin:
- Provider SDK or HTTP client:

## First broken boundary

- Boundary:
- Error code and message:
- Boundary-scoped correlation ID:

## Project diagnostic

- `doctor` status:
- `runtime.state`:
- `issues[].code` values:
- `dev` status, seed status, and service keys:

## Relevant bounded logs

- `droppedEntries`:
- Retained `request`, `operation`, `delivery`, or `plugin` entries:

## Compatibility claim

- Installed plugin documentation entry, when relevant:

## Removed before sharing

- Runtime control token
- App-facing connection credentials and provider secrets
- Personal data and unrelated request bodies
- Unrelated environment variables and log entries

Keep only evidence needed to reproduce the first broken boundary. The report is not a dump of the environment or storage directory. CLI workflow and reference defines command behavior, Runtime boundaries identifies claims that need another test, and Local security model defines the local trust boundary.