Skip to content

Fix Kimi subagent tasks being killed after ~5 min with only a SIGINT message - #686

Merged
zomux merged 3 commits into
developfrom
bugfix/kimi-watchdog
Sep 15, 2026
Merged

zomux merged 3 commits into
developfrom
bugfix/kimi-watchdog

Conversation

@QuanCheng-QC

Copy link
Copy Markdown
Collaborator

Where this came from

On 2026-09-14 a tester using Workspace on Windows reported the problem. They gave a Kimi agent (Kimi Code CLI) a longer task: "find multi-agent collaboration projects from the past month". About 5 minutes in, the run stopped. The thread showed only Kimi Code CLI terminated by signal SIGINT., with no hint of why.

The logs on the tester's machine show OpenAgents' own watchdog killed a Kimi run that was still working:

  • daemon.log: Watchdog: kimi silent 300s on <channel> — killing
  • Kimi's session journal: the main agent had handed the search to a subagent through the Agent tool. The subagent made 41 tool calls in those 5 minutes and was still returning results one second before the kill.

Root cause

  1. The watchdog only treated stdout as a sign of life. Kimi Code CLI's -p --output-format stream-json mode is often silent on stdout even while it works:

    • Thinking is never printed.
    • Assistant text and tool calls are held back until the step ends.
    • Subagent events are not printed at all.
    • Tool progress goes to stderr.

    So while the main agent waits on a subagent, stdout can stay silent for minutes, and after 5 minutes the run was treated as hung.

  2. The reason for the kill was lost. The watchdog was meant to add "Kimi became unresponsive" when it killed a run, but only did so when stderr was empty. In the reported run, stderr already held tool noise (which: no gh), so the reason was skipped. The error classifier found no error: line and fell back to a generic message. On Windows the kill is kill('SIGINT'), so the user saw only the signal name.

  3. A resumed session that was killed got re-run from scratch. The adapter read "resume failed with no output" as a stale session, cleared it and ran the whole task again.

How to reproduce

  • A Kimi agent running in CLI mode, i.e. Kimi Code CLI installed on the machine.
  • A turn in which Kimi keeps working but prints nothing to stdout for more than 5 minutes. The most reliable way is to have the main agent start a subagent in the foreground through the Agent tool, and have the subagent run 20 commands of 20 seconds each, one at a time. The prompt is below.
  • The logic is platform-independent; on Windows the failure shows up as SIGINT.
Prompt used to reproduce and verify

The Windows verification used the Chinese version of this prompt.

This is a stability test. Use the Agent tool to start a subagent (run_in_background must be false; wait for it in the foreground) and hand it the task below exactly as written. Do not call any other tool yourself until it finishes.

Subagent task: run the 20 commands below in order, one Bash tool call per command, starting each only after the previous one has finished. You must run all 20 before returning; skipping any of them fails the test. Do not combine commands, and do not write a loop or a script.
1. node -e "setTimeout(()=>console.log('step 1'),20000)"
2. node -e "setTimeout(()=>console.log('step 2'),20000)"
3. node -e "setTimeout(()=>console.log('step 3'),20000)"
4. node -e "setTimeout(()=>console.log('step 4'),20000)"
5. node -e "setTimeout(()=>console.log('step 5'),20000)"
6. node -e "setTimeout(()=>console.log('step 6'),20000)"
7. node -e "setTimeout(()=>console.log('step 7'),20000)"
8. node -e "setTimeout(()=>console.log('step 8'),20000)"
9. node -e "setTimeout(()=>console.log('step 9'),20000)"
10. node -e "setTimeout(()=>console.log('step 10'),20000)"
11. node -e "setTimeout(()=>console.log('step 11'),20000)"
12. node -e "setTimeout(()=>console.log('step 12'),20000)"
13. node -e "setTimeout(()=>console.log('step 13'),20000)"
14. node -e "setTimeout(()=>console.log('step 14'),20000)"
15. node -e "setTimeout(()=>console.log('step 15'),20000)"
16. node -e "setTimeout(()=>console.log('step 16'),20000)"
17. node -e "setTimeout(()=>console.log('step 17'),20000)"
18. node -e "setTimeout(()=>console.log('step 18'),20000)"
19. node -e "setTimeout(()=>console.log('step 19'),20000)"
20. node -e "setTimeout(()=>console.log('step 20'),20000)"
When all of them are done, return the list of step numbers you saw.

Once the subagent returns, reply to me with that list exactly as it came back.

Before

  • The Kimi process was killed at about 5 minutes, and whatever the subagent had done was lost.
  • The thread showed only Kimi Code CLI terminated by signal SIGINT.. Nothing said it was a timeout or what to do next.
  • If the run was resuming a session, the adapter then started the whole task over again.

After

  • The watchdog now looks for three signs of life instead of one: stdout output, stderr output, or a new write to this run's Kimi session journal. Any one of them means Kimi is still working.
    • The journal is located with Kimi CLI's own naming scheme: <KIMI_CODE_HOME>/sessions/wd_<dir name>_<first 12 hex of the path's sha256>/<session id>/.
    • Only the session this run created, or the one it resumes, is watched. Writes to other sessions under the same working directory do not count.
  • A run is stopped only after all three have been quiet for 5 minutes. daemon.log then records Watchdog: kimi silent 300s … killing. When journal writes keep a run alive, it logs … its session is still being written — keeping it alive once.
  • A watchdog stop now tells the user why: Kimi showed no progress for about 5 min and was stopped. Send the request again, or break a long task into smaller steps.
  • A timeout no longer triggers the "stale session, retry from scratch" path.
  • When the session directory cannot be found, behavior falls back to what it was. For example, a future Kimi CLI may change its layout, or the CLI may run inside WSL. The watchdog then watches stdout and stderr only, which is never worse than before.

Changes

  • packages/agent-connector/src/adapters/kimi.js
    • The watchdog checks stdout, stderr and the session journal.
    • Existing sessions are recorded before spawning, so only this run's session directory is watched.
    • New timedOutMs field on the run result.
    • A timed-out resume is no longer re-run.
    • Watchdog timings are instance properties so tests can shorten them.
  • packages/agent-connector/src/adapters/kimi-stream.js
    • New encodeKimiWorkDirKey, matching Kimi Code CLI 0.42.0's directory naming.
    • classifyKimiError gains a timeout kind.
  • packages/agent-connector/test/kimi.test.js: new test cases, listed below.

Verification

  • Unit tests: all 33 pass with node --test test/kimi.test.js. New coverage:
    • The directory naming rule, checked against a real Windows session directory name.
    • The timeout message.
    • A fake Kimi CLI in three scenarios: a subagent writing to a newly created session keeps the run alive; writing to the resumed session keeps it alive; writes only to some other session do not, and the run is stopped with the timeout message.
    • With session detection deliberately disabled, the first two scenarios fail, confirming the tests catch this bug.
  • Full suite: 1 of 1627 cases in npm test fails, wsl.test.js › an agent that only exists inside the distro. It depends on the test machine having ~/.openagents/runtimes/claude installed and touches none of the files in this PR.
  • On a real Windows machine, with Kimi Code CLI connecting directly to api.moonshot.cn and the prompt above:
    • The subagent ran 20 steps in the foreground over about 8 minutes, with stdout silent throughout.
    • daemon.log shows keeping it alive and no killing.
    • The subagent's journal records 20 tool_use steps followed by end_turn.
    • The thread received the full list, step 1 through step 20.

Out of scope (found while testing, to be handled separately)

  • Tool calls get lost when kimi-k2.6 is called through the OpenAgents gateway. Through api-gateway.openagents.org, every tool call after the first in a turn is dropped. The subagent stops after one step, or the main agent fails with only thinking content … finishReason=tool_calls. The same requests sent directly to Moonshot work.
  • Generic LLM_* settings silently override the KIMI_* values a user entered. The agent's own env and the shared kimi.env stack on top of each other, so a user who configured a direct connection can end up on the gateway without knowing.
  • Nothing shows in the thread while a subagent is working. Journal writes reset the silence counter, so the "Still working..." status never appears.
  • The Python Kimi adapter has not been updated yet (sdk/src/openagents/adapters/kimi.py).

…g stops it

The Kimi adapter only counted stdout as a sign of life. In print mode the CLI
never forwards subagent events and sends tool progress to stderr, so a main
agent waiting on an Agent tool call went quiet for five minutes and was killed
mid task. On Windows that kill reads as signal SIGINT, and the unresponsive
reason was dropped because stderr already held tool noise.

Now stderr counts as life, and so does any write under the run's own session
journal, meaning the session it created or resumes, located with the same
workdir key the CLI uses. A watchdog stop is recorded in its own field and
reported as a timeout, and a timed out resume is no longer retried from scratch.
@vercel

vercel Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
openagents-workspace Ready Ready Preview Sep 15, 2026 4:38am UTC

Request Review

@QuanCheng-QC

QuanCheng-QC commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator Author

复现:
image

修复后:
image

…t-proof budget

The watchdog tests killed healthy runs on loaded runners: the fake CLI is
noisy every 30ms once up, so the only silent stretch is node's own boot,
and a 5x100ms budget from spawn is shorter than a cold start on Windows CI
(or a busy Linux box — reproduced locally). Budget is now 20 ticks (2s);
the alive scenarios stay busy for 3s so a watchdog ignoring the journal is
still caught. Verified 3 concurrent full-file runs under 4 CPU hogs, 33/33
each.

Also clearInterval before the awaited _stopProcess: a tick landing inside
the multi-second escalation re-logged, re-killed and kept inflating
timedOutMs past the budget (the 800 seen against a 500 budget in CI).

@zomux zomux left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approve. The diagnosis and fix are right: Kimi's print mode is silent on stdout while a subagent works (thinking never printed, text/tool calls held to step end, subagent events never forwarded, tool progress on stderr), so a stdout-only watchdog killed healthy runs; the reason was lost because the fallback only fired on empty stderr; and a killed resume was re-run from scratch. Counting stdout, stderr and this run's session journal as life — with the bucket located via a faithful port of the CLI's encodeWorkDirKey and a fail-safe fallback to the old behavior when it can't be found — is the correct shape, and the live Windows verification (20-step foreground subagent, ~8 min, stdout silent, result delivered) matches the code.

Two things fixed on-branch before merge:

  1. The new watchdog tests were flaky under load — that was the red Windows CI, and it reproduced on a busy Linux box. The fake CLI is noisy every 30ms once up, so the only silent stretch is node's own boot, and a 5x100ms budget from spawn is shorter than a cold start on a loaded runner. Budget is now 20 ticks (2s), alive scenarios stay busy 3s so a journal-blind watchdog is still caught. Verified 3 concurrent full-file runs under 4 CPU hogs: 33/33 each.
  2. clearInterval before the awaited _stopProcess (pre-existing shape, not a regression): a tick landing inside the multi-second kill escalation re-logged, re-killed and inflated timedOutMs past the budget — the '800 vs 500' in the CI failure.

Separately: the out-of-scope finding that api-gateway.openagents.org drops every tool call after the first for kimi-k2.6 deserves its own issue with priority — it hits every Kimi user on the default config.

The kill count could never accumulate on Windows: after each silent tick
lastDataMs was set to now, parking the next tick's elapsed exactly on the
interval boundary, where an early-firing timer read as 'data just arrived'
and reset silences to 0. A hung run was then never stopped — CI showed the
'stops a quiet run' test letting the 10s fake run to completion. The data
handlers already reset silences directly, so lastDataMs can simply mean
'last real sign of life'. Pre-existing on develop; exposed by the new tests.
@zomux
zomux merged commit 8e48624 into develop Sep 15, 2026
14 checks passed
zomux pushed a commit that referenced this pull request Sep 21, 2026
Publishes everything merged since the last core: #686 (Kimi watchdog no
longer kills working subagent runs; watchdog actually fires on Windows),
#676 (Stop sticks on one press across all adapters), #694 (runtime
detection no longer wakes WSL), #703 (Cline runs on Windows without an
npm .cmd shim), #704 (Hermes preflight + Windows install), #693 (Codex
attachments, relay stall deadlines, resume order), #695 (live model
listing for the picker).

This branch was successfully deployed

1 active deployment
Preview — 18731faf Deployed Sep 15, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants