Parent
#62
What happened
Found in #67's live leg (docs/findings/14-acp-live-leg.md, "agent crash"). When an ACP agent
process dies by SIGKILL in the middle of a tool call, the step fails correctly
(AgentStepChunkError: claude exited (code 137); stderr: …). The agent's tool processes keep
running, reparented to init, until the tool finishes by itself.
| agent |
process tree mid-turn |
killed |
left running |
| Claude |
daemon → claude-agent-acp (bun) → claude → zsh -c … → python3 |
claude-agent-acp (kill -9) |
claude, zsh and python3. They stayed alive for the whole 120 s sleep and exited only when it ended |
| opencode |
daemon → opencode acp → python3 |
opencode acp (kill -9) |
python3 |
A tool with no natural end, such as a dev server or tail -f, would never exit.
Cancel, a crash of the claude binary itself (SIGTERM), and daemon shutdown all leave nothing
behind. The agents clean up their own children whenever they get to run their own teardown.
Why the adapter can't clean up today
acpAdapter's finally kills only child.pid, which is already dead. Process groups:
- Claude:
claude shares the daemon's process group. Its shell calls setsid() and gets its own
group. A SIGTERM to claude cleans up the shell and the tool (verified).
- opencode: each shell tool gets its own process group (
pgid == pid), directly under
opencode acp.
Options
- Spawn the agent
detached (its own process group) and signal the whole group (-pid) when it
tears down. This fixes Claude, since claude gets the SIGTERM and cleans up, but not opencode.
- Record the agent's descendants (
/proc/*/stat ppid walk) while the step runs, and kill the
recorded set when the agent dies. This covers both agents. It needs Linux and is racy for tools
that start between two walks.
- Leave it to a future container sandbox (ADR 0013 §6), where the container's end kills
everything.
Related: shutdown grace vs. the adapter's kill timer
On shutdown the daemon waits AGENT_TEARDOWN_GRACE_MS (1 s) per step, but the adapter's
post-cancel kill timer is 2 s. An agent that ignored session/cancel would outlive the daemon
unless it exits when its stdin closes. In the live leg, both agents settled the cancel in
milliseconds: the daemon exited 108 ms after SIGTERM, and nothing was alive 3 s later. So
this stays theoretical, but the fix above should also cover it, for example by sending
SIGKILL to the agent's group on process exit.
Parent
#62
What happened
Found in #67's live leg (
docs/findings/14-acp-live-leg.md, "agent crash"). When an ACP agentprocess dies by
SIGKILLin the middle of a tool call, the step fails correctly(
AgentStepChunkError: claude exited (code 137); stderr: …). The agent's tool processes keeprunning, reparented to init, until the tool finishes by itself.
claude-agent-acp(bun) →claude→zsh -c …→python3claude-agent-acp(kill -9)claude,zshandpython3. They stayed alive for the whole 120 ssleepand exited only when it endedopencode acp→python3opencode acp(kill -9)python3A tool with no natural end, such as a dev server or
tail -f, would never exit.Cancel, a crash of the
claudebinary itself (SIGTERM), and daemon shutdown all leave nothingbehind. The agents clean up their own children whenever they get to run their own teardown.
Why the adapter can't clean up today
acpAdapter'sfinallykills onlychild.pid, which is already dead. Process groups:claudeshares the daemon's process group. Its shell callssetsid()and gets its owngroup. A
SIGTERMtoclaudecleans up the shell and the tool (verified).pgid == pid), directly underopencode acp.Options
detached(its own process group) and signal the whole group (-pid) when ittears down. This fixes Claude, since
claudegets theSIGTERMand cleans up, but not opencode./proc/*/statppid walk) while the step runs, and kill therecorded set when the agent dies. This covers both agents. It needs Linux and is racy for tools
that start between two walks.
everything.
Related: shutdown grace vs. the adapter's kill timer
On shutdown the daemon waits
AGENT_TEARDOWN_GRACE_MS(1 s) per step, but the adapter'spost-cancel kill timer is 2 s. An agent that ignored
session/cancelwould outlive the daemonunless it exits when its stdin closes. In the live leg, both agents settled the cancel in
milliseconds: the daemon exited 108 ms after
SIGTERM, and nothing was alive 3 s later. Sothis stays theoretical, but the fix above should also cover it, for example by sending
SIGKILLto the agent's group on process exit.