Fix Kimi subagent tasks being killed after ~5 min with only a SIGINT message - #686
Conversation
…g stops it The Kimi adapter only counted stdout as a sign of life. In print mode the CLI never forwards subagent events and sends tool progress to stderr, so a main agent waiting on an Agent tool call went quiet for five minutes and was killed mid task. On Windows that kill reads as signal SIGINT, and the unresponsive reason was dropped because stderr already held tool noise. Now stderr counts as life, and so does any write under the run's own session journal, meaning the session it created or resumes, located with the same workdir key the CLI uses. A watchdog stop is recorded in its own field and reported as a timeout, and a timed out resume is no longer retried from scratch.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
…t-proof budget The watchdog tests killed healthy runs on loaded runners: the fake CLI is noisy every 30ms once up, so the only silent stretch is node's own boot, and a 5x100ms budget from spawn is shorter than a cold start on Windows CI (or a busy Linux box — reproduced locally). Budget is now 20 ticks (2s); the alive scenarios stay busy for 3s so a watchdog ignoring the journal is still caught. Verified 3 concurrent full-file runs under 4 CPU hogs, 33/33 each. Also clearInterval before the awaited _stopProcess: a tick landing inside the multi-second escalation re-logged, re-killed and kept inflating timedOutMs past the budget (the 800 seen against a 500 budget in CI).
zomux
left a comment
There was a problem hiding this comment.
Approve. The diagnosis and fix are right: Kimi's print mode is silent on stdout while a subagent works (thinking never printed, text/tool calls held to step end, subagent events never forwarded, tool progress on stderr), so a stdout-only watchdog killed healthy runs; the reason was lost because the fallback only fired on empty stderr; and a killed resume was re-run from scratch. Counting stdout, stderr and this run's session journal as life — with the bucket located via a faithful port of the CLI's encodeWorkDirKey and a fail-safe fallback to the old behavior when it can't be found — is the correct shape, and the live Windows verification (20-step foreground subagent, ~8 min, stdout silent, result delivered) matches the code.
Two things fixed on-branch before merge:
- The new watchdog tests were flaky under load — that was the red Windows CI, and it reproduced on a busy Linux box. The fake CLI is noisy every 30ms once up, so the only silent stretch is node's own boot, and a 5x100ms budget from spawn is shorter than a cold start on a loaded runner. Budget is now 20 ticks (2s), alive scenarios stay busy 3s so a journal-blind watchdog is still caught. Verified 3 concurrent full-file runs under 4 CPU hogs: 33/33 each.
- clearInterval before the awaited _stopProcess (pre-existing shape, not a regression): a tick landing inside the multi-second kill escalation re-logged, re-killed and inflated timedOutMs past the budget — the '800 vs 500' in the CI failure.
Separately: the out-of-scope finding that api-gateway.openagents.org drops every tool call after the first for kimi-k2.6 deserves its own issue with priority — it hits every Kimi user on the default config.
The kill count could never accumulate on Windows: after each silent tick lastDataMs was set to now, parking the next tick's elapsed exactly on the interval boundary, where an early-firing timer read as 'data just arrived' and reset silences to 0. A hung run was then never stopped — CI showed the 'stops a quiet run' test letting the 10s fake run to completion. The data handlers already reset silences directly, so lastDataMs can simply mean 'last real sign of life'. Pre-existing on develop; exposed by the new tests.
Publishes everything merged since the last core: #686 (Kimi watchdog no longer kills working subagent runs; watchdog actually fires on Windows), #676 (Stop sticks on one press across all adapters), #694 (runtime detection no longer wakes WSL), #703 (Cline runs on Windows without an npm .cmd shim), #704 (Hermes preflight + Windows install), #693 (Codex attachments, relay stall deadlines, resume order), #695 (live model listing for the picker).


Where this came from
On 2026-09-14 a tester using Workspace on Windows reported the problem. They gave a Kimi agent (Kimi Code CLI) a longer task: "find multi-agent collaboration projects from the past month". About 5 minutes in, the run stopped. The thread showed only
Kimi Code CLI terminated by signal SIGINT., with no hint of why.The logs on the tester's machine show OpenAgents' own watchdog killed a Kimi run that was still working:
daemon.log:Watchdog: kimi silent 300s on <channel> — killingAgenttool. The subagent made 41 tool calls in those 5 minutes and was still returning results one second before the kill.Root cause
The watchdog only treated stdout as a sign of life. Kimi Code CLI's
-p --output-format stream-jsonmode is often silent on stdout even while it works:So while the main agent waits on a subagent, stdout can stay silent for minutes, and after 5 minutes the run was treated as hung.
The reason for the kill was lost. The watchdog was meant to add "Kimi became unresponsive" when it killed a run, but only did so when stderr was empty. In the reported run, stderr already held tool noise (
which: no gh), so the reason was skipped. The error classifier found noerror:line and fell back to a generic message. On Windows the kill iskill('SIGINT'), so the user saw only the signal name.A resumed session that was killed got re-run from scratch. The adapter read "resume failed with no output" as a stale session, cleared it and ran the whole task again.
How to reproduce
Agenttool, and have the subagent run 20 commands of 20 seconds each, one at a time. The prompt is below.Prompt used to reproduce and verify
The Windows verification used the Chinese version of this prompt.
Before
Kimi Code CLI terminated by signal SIGINT.. Nothing said it was a timeout or what to do next.After
<KIMI_CODE_HOME>/sessions/wd_<dir name>_<first 12 hex of the path's sha256>/<session id>/.daemon.logthen recordsWatchdog: kimi silent 300s … killing. When journal writes keep a run alive, it logs… its session is still being written — keeping it aliveonce.Kimi showed no progress for about 5 min and was stopped. Send the request again, or break a long task into smaller steps.Changes
packages/agent-connector/src/adapters/kimi.jstimedOutMsfield on the run result.packages/agent-connector/src/adapters/kimi-stream.jsencodeKimiWorkDirKey, matching Kimi Code CLI 0.42.0's directory naming.classifyKimiErrorgains atimeoutkind.packages/agent-connector/test/kimi.test.js: new test cases, listed below.Verification
node --test test/kimi.test.js. New coverage:npm testfails,wsl.test.js › an agent that only exists inside the distro. It depends on the test machine having~/.openagents/runtimes/claudeinstalled and touches none of the files in this PR.api.moonshot.cnand the prompt above:daemon.logshowskeeping it aliveand nokilling.tool_usesteps followed byend_turn.Out of scope (found while testing, to be handled separately)
api-gateway.openagents.org, every tool call after the first in a turn is dropped. The subagent stops after one step, or the main agent fails withonly thinking content … finishReason=tool_calls. The same requests sent directly to Moonshot work.LLM_*settings silently override theKIMI_*values a user entered. The agent's own env and the sharedkimi.envstack on top of each other, so a user who configured a direct connection can end up on the gateway without knowing.sdk/src/openagents/adapters/kimi.py).