Skip to content

refactor(daemon): move singletons into per-daemon layers - #57

Merged
FreshlyBrewedCode merged 26 commits into
36-agent-runtime-servicefrom
38-move-singletons-into-layers
Oct 2, 2026
Merged

FreshlyBrewedCode merged 26 commits into
36-agent-runtime-servicefrom
38-move-singletons-into-layers

Conversation

@FreshlyBrewedCode

@FreshlyBrewedCode FreshlyBrewedCode commented Sep 20, 2026 •

Copy link
Copy Markdown
Owner

Part of #32 · Closes #38

Four maps lived at module scope: server/runs.ts's active-run registry, server/pubsub.ts's subscribers, lib/dedupe.ts's dedupeRegistry, and lib/workspace.ts's refreshGates. So two daemons could not run in one process, and every feature since #13 had to add another optional test-seam field (dedupeRegistry?, registry?, now?, beforeStart?). This PR turns them into per-daemon Effect services composed into the daemon's ManagedRuntime (ADR 0009 §5) and removes those seams. The daemon also cancels its active runs on shutdown.

What changed

Per-daemon state as Effect layers

  • RunRegistry, RunPubSub, DedupeRegistry and RefreshGates are now Context.Services, each with a Layer that builds fresh state.
  • src/server/daemon-runtime.ts merges them with AgentRuntimeLayer (feat(runtime): add Effect composition root and agent runtime service #56) into DaemonLayer. startDaemon builds one ManagedRuntime per daemon and builds the layers eagerly at startup (await runtime.context()).
  • startTrackedRun, dispatchChildRun, makeScheduleFire and the HTTP handlers read services from the runtime with serviceOf(runtime, Service). The HTTP handlers stay plain functions. The scheduler loop runs on the daemon runtime and reads DedupeRegistry and Clock.Clock from context.
  • dedupeRegistry?, registry?, now? and beforeStart? are gone. Tests build a fresh runtime with createTestDaemon(adapter).
  • src/server/two-daemons.test.ts starts two full daemons in one process. It asserts separate run lists, registries, pubsub streams and dedupe keys, and that the two runtimes share no service instance.

Cancel active runs on shutdown

  • DaemonHandle.stop() (idempotent) runs these steps in order:
    1. interrupt the scheduler
    2. RunRegistry.shutdown(timeout): close admission (DaemonShuttingDownError, which HTTP returns as 503), cancel running handles and reserved slots, and wait a bounded time for them to settle
    3. server.stop (live SSE clients still see the cancellations)
    4. runtime.dispose()
    5. wait briefly for cancelled runs to record their end
  • An acquireRelease finalizer on the registry layer also cancels runs if the runtime is disposed without stop().
  • factory serve's SIGINT/SIGTERM handler calls stop(). Runs are recorded as RunCancelled instead of reading back as interrupted after a restart, and exec and opencode child processes are killed.
  • Agent-step cancel used to wait up to ~20s, because Stream.fromAsyncIterable's finalizer awaited the generator's return(), which queues behind a pending next(). The adapter iterable is now wrapped: next() races the abort signal, and return() resolves immediately while the adapter tears down in the background. That teardown is capped at 1s (AGENT_TEARDOWN_GRACE_MS), because @tanstack/ai can hold an aborted stream open forever; the sandbox kills opencode's process group on abort. POST /cancel during an agent step now takes ~25ms instead of ~22s.

Also

  • startTrackedRun checks admission before claiming the dedupe key, so a start refused at the concurrency limit no longer holds its key. Before this, an overlap: "skip" schedule stuck on skipped-overlap until restart. This bug is also on main.
  • Every start now reserves a registry slot, so shutdown can see runs that are still allocating a workspace. A second start with a run id that is still reserving is refused instead of starting twice.
  • Restored the refactor(runtime): move headless permission setup into the agent adapter #55 prepareWorkspace wiring that an earlier rebase had dropped, along with the rationale comments and Represent domain failures as Schema.TaggedError #34 tests deleted in that rebase.

Notes for reviewers

  • Atomicity (D24/D29, Dedupe keys on run start #15): all layers are synchronous, and services are resolved before the critical section. The closed check, admission, dedupe claim and slot reservation run as one synchronous stretch with no await or yield between them. M1 and L1 hold the reservation window open with a gate in refreshGates instead of depending on clone timing. Two tests cover a synchronous double start: one with a limit (ConcurrencyLimitError) and one with the same run id (refused).
  • startRun takes an AgentRunner (runFork/runPromise) instead of ManagedRuntime<AgentRuntime>. ManagedRuntime isn't assignable from a runtime that provides more services, and this avoids a cast.

Verification

  • bun run check: 347 pass, 0 fail. bun run test:e2e: 15/15.
  • New tests in src/server/shutdown.test.ts:
    • an exec run cancelled by stop
    • a reserved slot cancelled by stop
    • stop is idempotent
    • dispose without stop still cancels
    • an adapter that never yields again is cancelled in <500ms
    • stop finishes within its bound when teardown hangs
  • Smoke test against a real daemon with opencode/big-pickle:
    • A normal run completes with structured output and usage intact.
    • With an agent step waiting on the model plus an exec sleep 300 run, SIGTERM and SIGINT exit 143/130 in about 1s. A start sent during shutdown gets 503. After restart both runs read RunCancelled, and no opencode or sleep processes are left.

Stack

  1. refactor(cli): replace hand-rolled parsers with effect/unstable/cli #52: Parse CLI with Effect CLI (issue Parse CLI arguments with effect/unstable/cli #33)
  2. refactor(errors): represent domain failures as Schema.TaggedError #53: Domain errors as tagged errors (issue Represent domain failures as Schema.TaggedError #34)
  3. refactor(runtime): move chunk interpretation into the agent adapter #54: Move chunk interpretation into adapter (issue Move chunk interpretation into the agent adapter #35)
  4. refactor(runtime): move headless permission setup into the agent adapter #55: Move headless permission setup into the agent adapter (issue Move headless permission setup out of the workspace allocator #37)
  5. feat(runtime): add Effect composition root and agent runtime service #56: Agent runtime service (issue Add an Effect composition root and make the agent runtime a service #36)
  6. refactor(daemon): move singletons into per-daemon layers #57: Move singletons into per-daemon layers (issue Move the daemon's remaining singletons into layers #38) ← this PR

🤖 Generated with Claude Code

FreshlyBrewedCode and others added 18 commits October 2, 2026 09:22
The dedupe registry is now always created via createDedupeRegistry()
and passed to consumers. This is the first step in moving the four
module-level singletons (dedupe, pubsub, active registry, refresh
gates) into per-daemon instances so two daemons can coexist in one
process.

Part of #38
…bSub()

Subscribers are now per-instance, so two daemons in one process do not
share event fan-out state. Consumers receive a PubSub instance rather
than calling module-level publish/subscribe.

Part of #38
Refresh gates are now per-instance via createRefreshGates() and passed
to allocateWorkspace. This allows two daemons to maintain independent
mirror-refresh queues.

Part of #38
…t seams

The active-run registry is now per-daemon via createRunRegistry().
startTrackedRun and dispatchChildRun accept DaemonServices (registry,
pubsub, dedupeRegistry, refreshGates) as required parameters.

Removed optional test-seam fields:
- dedupeRegistry? from StartTrackedRunOptions and DispatchEnv
- beforeStart? from StartTrackedRunOptions

The slot reservation and dedupe-key claim remain synchronous
check-then-set with no await between, preserving concurrency safety.

Part of #38
SchedulerDeps now requires dedupeRegistry and clock as mandatory fields.
The clock comes from Effect's Clock service (currentTimeMillisUnsafe),
replacing the bespoke now?: () => number test seam. Tests provide a
Clock instance instead of injecting a now function.

Part of #38
startDaemon creates DaemonServices (registry, pubsub, dedupe, refresh
gates) and passes them to serve() and the scheduler. http.ts receives
services as a required ServerOptions field and uses them throughout.

This completes the core refactor - two daemons can now coexist in one
process with independent state. Tests still need updating to provide
services instead of using test seams.

Part of #38
Tests now create their own DaemonServices instances instead of using
module-level singletons or test seams. The beforeStart test seam is
gone - tests use slow workspace allocation or direct registry
manipulation to test the reserved-but-not-started window.

Scheduler tests use Effect's Clock interface instead of the bespoke
now?: () => number test seam.

All 250 tests pass.

Part of #38
Proves that two daemons can run in one process without sharing run
state. Each daemon has its own registry, pubsub, dedupe registry, and
refresh gates. Tests verify:
- Independent run registries (each daemon sees only its own runs)
- Independent pubsub (events don't cross between daemons)
- Independent dedupe state (same key can be claimed in both daemons)

Part of #38
A 0-byte .tanstack-projected-* file under a stray data/src/factory/...
path leaked into 2f97451 from a sandboxed run writing through an
absolute path that mirrored the checkout. .gitignore already covers
.tanstack-projected-* and data/ (added in PR #55), which is why
nothing flagged it since; git rm drops it and its now-empty parent
dirs. No other stray data/ or .factory/workspaces/ paths are tracked
on this branch.

Part of #32
Replace the hand-rolled fake Clock.Clock with effect/testing's
TestClock, per issue #38's "tests use TestClock rather than an
injected now". tickOnce only ever reads the clock synchronously
(currentTimeMillisUnsafe()) and never suspends on Effect.sleep, so
none of TestClock's fiber-coordination machinery is exercised here —
but building it via Effect.runPromise(Effect.scoped(TestClock.make()))
and bridging setTime back into the plain-async fixture (ADR 0009 §5)
composes cleanly, so there was no need to keep the fake.

Part of #32
The #34 exhaustiveness test landed while `fixture` was still synchronous;
#38 made it async for `TestClock`. Rebasing the stack put the two together,
which typecheck caught.
…ollision

The domainErrorMessage helper was removed on the #34 layer; DedupeKeyError
now carries its message as a constructed field.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…vices

Rebasing onto the reworked agent-runtime layer: startRun is synchronous
again, the daemon handle disposes its managed runtime via `stop()`, the
opencode default lives only in AgentRuntimeLayer, and the doc comments
restored below are kept on the per-daemon run registry and its options.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…dropped

- restore `prepareWorkspace` on StartTrackedRunOptions and the startRun call
  (allocated clone workspaces and the legacy `{dir, clone}` POST are prepared
  again, #37)
- restore the deleted rationale comments in http.ts, scheduler.ts, daemon.ts
  and runs.ts (SSE keepalive vs Bun's idleTimeout, the `development: false`
  router crash, D29 admission, DaemonOptions.config, runOnStart consumption)
- drop the unrelated `error` field added to DispatchCollision
- use Effect's default Clock instead of a hand-rolled live clock

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… weakened

The reserved-but-not-started window is now held deterministically by
pre-seeding the daemon's refresh gate for the workspace mirror, replacing the
removed `beforeStart` seam rather than relying on git clone timing. Restores
the strong M1/L1 assertions, the #34 TaggedError block, the DedupeKeyError
`_tag` check and the deleted test comments; adds a synchronous second-start
refusal check; drops the stale `adapter` DispatchEnv cast; and makes the
two-daemon test check dedupe holding and pubsub isolation while runs are live.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A start refused by the concurrency limit threw after claiming its dedupe key
and outside the try that releases it, so the key stayed held by a run that
never started. A scheduled fire with `overlap: "skip"` at the limit then
reported `skipped-overlap` forever. Admission is now checked first; claim and
slot reservation still happen with no await in between.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@FreshlyBrewedCode
FreshlyBrewedCode force-pushed the 38-move-singletons-into-layers branch from 0bbf3f1 to 82f4f5d Compare October 2, 2026 09:41
FreshlyBrewedCode and others added 7 commits October 2, 2026 10:42
The run registry, run-event pubsub, dedupe registry and mirror refresh
gates become Context.Service classes with layers that build fresh state
per layer build. DaemonLayer composes them with AgentRuntimeLayer into
the daemon's ManagedRuntime (server/daemon-runtime.ts), replacing the
hand-threaded DaemonServices bag.

The plain-async boundaries (HTTP handlers, startTrackedRun, ctx.dispatch,
the scheduler's fire) take the daemon runtime and resolve services with
serviceOf, a synchronous runSync lookup over all-synchronous layers, so
startTrackedRun still reserves its slot and claims its dedupe key with no
await between check and set. The scheduler loop is forked on the daemon
runtime and resolves its dedupe registry and Clock from context; the
unused DaemonOptions.clock seam is gone. startRun accepts any runtime that
can run AgentRuntime effects (AgentRunner), since ManagedRuntime is
invariant in its services.

Tests build a fresh daemon runtime per case (createTestDaemon); the
two-daemon test also asserts no service instance is shared.

Refs #38

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
DaemonHandle.stop() (and so factory serve's SIGINT/SIGTERM handler) now
shuts the run registry down before the runtime goes away: it refuses new
starts (DaemonShuttingDownError, 503 over HTTP), cancels every running
handle, marks every reserved slot for cancel-on-start, and waits up to
shutdownTimeoutMs (default 5s) for the runs to settle. Runs therefore
persist RunCancelled and their ctx.exec children are killed, instead of
being orphaned and read back as interrupted after a restart.

Order: interrupt the scheduler, shut the registry down while the daemon
runtime is still live, stop the HTTP server (so SSE tails see the
cancellations), dispose the runtime. stop() is idempotent. The registry
layer's finalizer runs the same shutdown as a safety net when a runtime is
disposed without stop().

Every start now reserves its registry slot, not only admission-limited
ones, so a run still allocating is visible to shutdown either way.

Refs #38

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Stream.fromAsyncIterable closes the adapter iterable with a scope
finalizer that awaits iter.return(), and an async generator's return()
queues behind its pending next(). Interrupting a step that was waiting on
the model therefore blocked until opencode's next chunk, and the
onInterrupt abort that would have hurried it along only ran after that
finalizer: a POST /cancel took ~20s to record RunCancelled, and shutdown
blew its budget.

abortableIterable wraps the adapter stream: next() races the abort
signal, and return() aborts, starts the adapter's own return() without
awaiting it, and resolves immediately. The adapter teardown (killing the
opencode process) still runs; AgentStepHandle.teardown exposes it,
RunHandle.settled waits for every abandoned step's teardown after the
result, and registry shutdown now waits for settled rather than only for
the registry to empty, so stop() does not exit under a live agent.

Refs #38

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…utdown

When the shutdown wait times out, stop() used to resolve as soon as the
runtime was disposed, and factory serve's process.exit could land between
AgentStepFinished{cancelled} and RunCancelled, so the run read back as
interrupted. After disposal (which interrupts whatever still held the
runs), stop() now waits a further bounded 2s for the cancelled runs to
settle via RunRegistry.awaitCancelled.

Refs #38

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Since every start reserves its slot, a reserved entry for the same run id
can only be a concurrent start with the same explicit runId, and that path
skipped admission and started a duplicate run. Any existing entry, running
or reserved, now throws "already active".

Refs #38

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
serviceOf resolves with runSync, which only works over an already-built
or synchronously buildable context. startDaemon now awaits
runtime.context() before serving, so a failing (or future asynchronous)
layer surfaces at startup rather than on the first request, and every
later lookup reads the built context. createTestDaemon already builds
eagerly by resolving its services at construction; documented.

Refs #38

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Measured live against opencode: after an abort, @tanstack/ai's chat
engine can sit on the aborted stream indefinitely, so the adapter's
generator teardown never settles, even though the abort has already
killed the opencode process (the local-process sandbox kills the process
group on the spawn's abort signal). Waiting on that teardown made every
shutdown with an agent run use its whole budget. RunHandle.settled now
waits at most AGENT_TEARDOWN_GRACE_MS (1s) per abandoned step.

Refs #38

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@FreshlyBrewedCode
FreshlyBrewedCode merged commit 1660be4 into main Oct 2, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Move the daemon's remaining singletons into layers

1 participant