Skip to content

fix(runtime): a SIGKILLed agent leaves its in-flight tool processes running #74

Description

@FreshlyBrewedCode

Parent

#62

What happened

Found in #67's live leg (docs/findings/14-acp-live-leg.md, "agent crash"). When an ACP agent
process dies by SIGKILL in the middle of a tool call, the step fails correctly
(AgentStepChunkError: claude exited (code 137); stderr: …). The agent's tool processes keep
running, reparented to init, until the tool finishes by itself.

agent process tree mid-turn killed left running
Claude daemon → claude-agent-acp (bun) → claude → zsh -c … → python3 claude-agent-acp (kill -9) claude, zsh and python3. They stayed alive for the whole 120 s sleep and exited only when it ended
opencode daemon → opencode acp → python3 opencode acp (kill -9) python3

A tool with no natural end, such as a dev server or tail -f, would never exit.

Cancel, a crash of the claude binary itself (SIGTERM), and daemon shutdown all leave nothing
behind. The agents clean up their own children whenever they get to run their own teardown.

Why the adapter can't clean up today

acpAdapter's finally kills only child.pid, which is already dead. Process groups:

  • Claude: claude shares the daemon's process group. Its shell calls setsid() and gets its own
    group. A SIGTERM to claude cleans up the shell and the tool (verified).
  • opencode: each shell tool gets its own process group (pgid == pid), directly under
    opencode acp.

Options

  • Spawn the agent detached (its own process group) and signal the whole group (-pid) when it
    tears down. This fixes Claude, since claude gets the SIGTERM and cleans up, but not opencode.
  • Record the agent's descendants (/proc/*/stat ppid walk) while the step runs, and kill the
    recorded set when the agent dies. This covers both agents. It needs Linux and is racy for tools
    that start between two walks.
  • Leave it to a future container sandbox (ADR 0013 §6), where the container's end kills
    everything.

Related: shutdown grace vs. the adapter's kill timer

On shutdown the daemon waits AGENT_TEARDOWN_GRACE_MS (1 s) per step, but the adapter's
post-cancel kill timer is 2 s. An agent that ignored session/cancel would outlive the daemon
unless it exits when its stdin closes. In the live leg, both agents settled the cancel in
milliseconds: the daemon exited 108 ms after SIGTERM, and nothing was alive 3 s later. So
this stays theoretical, but the fix above should also cover it, for example by sending
SIGKILL to the agent's group on process exit.

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions