Skip to content

fix: land the timeout-audit review fixes for the merged #63 - #75

Open
asto18089 wants to merge 32 commits into
pinvou3-cleanfrom
round6-followup
Open

asto18089 wants to merge 32 commits into
pinvou3-cleanfrom
round6-followup

Conversation

@asto18089

@asto18089 asto18089 commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

#63 (merged as d349f25) closed mid-review, so its final review wave never reached pinvou3-clean. This PR lands that wave plus the fixes from the review rounds that followed, on top of the current base (pinvou3-clean @ 61cb769). After review stabilized the branch was squashed into one commit; the PR diff is the single source of truth. Full findings and evidence for the original audit are in the round-6 review comment on #63.

What's in the diff

  • exec — headless runs resolve request_user_input via cancel_user_input instead of parking forever once the engine wait became unbounded; the engine half of the cancellation contract is pinned red-green (cancel_user_input_resolves_the_headless_wait_deterministically), while the exec-side arm itself is untested — run_exec_agent has no test harness (neither do its sibling arms).
  • snapshot — every pre-git window on a wedged mount is bounded (workspace canonicalize, the safety classifier's re-canonicalization, the first-init size walk, honoring max_workspace_gb = 0); a killed git init heals instead of wedging (HEAD-less, half-written, leftover config.lock/HEAD.lock all recover); a killed git update-ref/pack-refs leaves a ref lock the open sweep now clears like index.lock; wedge failures surface as errors with stderr evidence and stage-keyed remedy hints; drain captures cap at 16 MiB per stream.
  • runtime — every approval exit publishes approval.decided and is race-free on every reachable path (delivery races, monitor-death settlement, channel-closed exit; the one panic-only window is disclosed under Known limitations); the approval.required-emit-failure rollback removes only a same-thread match, so a reused id cannot strand another waiter; compat streams forward the approval posture.
  • tasks — liveness heartbeats short-circuit before the manager-wide state lock. (The human-wait clock pause was taken back out and now lives in feat(tasks): pause the wall and idle clocks for human waits, answerably #78: as first landed it changed task-lifecycle behavior and regressed the TUI.)
  • vision — the envelope timeout classifies as ToolError::Timeout; the inference permit is acquired inside the envelope.
  • client — envelope-exceeded requests mark connection health and run the recovery probe on every dialect arm (including the Anthropic transport arm); the non-streaming budget got an injectable test seam.
  • app-server — a dropped stream resumes from the last consumed seq (bounded reconnects) before falling back to a best-effort interrupt, on every turn surface; the /health boot probe bounds each attempt.
  • tools — drain aborts are one helper covering every exit including cancel/drop; timeouts SIGKILL the child's process group so forked commands no longer outlive the call.
  • rlm — sub-query budget floored at 1s; the session default references CHILD_TIMEOUT_SECS.
  • mcp — the send-timeout error states the budget instead of tokio's Elapsed display.
  • docs — the 1800s timeout family anchors one complete roster; MCP/SUBAGENTS disclose the per-leg budget and wall-clock interactions; RUNTIME_API.md restore contract corrected.
  • tests — the load-bearing fixes above have regression pins (mutation-verified red-green; pins added in the round-7 follow-ups are listed there), timing margins widened, env-lock isolation on the Anthropic endpoint tests. Named exceptions with no dedicated pin: the two heartbeat short-circuits (their regression surface is lock contention, not observable state) and the vision permit's move inside the envelope (the envelope timeout classification itself is pinned).

Round-9 review follow-ups (eight commits on top of round-8)

From the fresh review wave of the full diff:

  • tasks — a timed-out gate no longer fabricates exit_code = 128+SIGKILL: the integer reached hook exit_code conditions through tool metadata and violated the documented reported_tool_exit_code contract (an exit_code condition must never match a fabricated value; the code doubles as the OOM-kill convention, so OOM hooks fired on every gate timeout). The timeout signal stays timed_out: true; the absence is pinned by a real timed-out gate. Also fixes the bool-assert comparison the stricter local clippy flags in the neighbouring test.
  • snapshot — recovery now covers the commit path's ref locks: update-ref HEAD (commit and prune) and pack-refs take refs/heads/**.lock / packed-refs.lock, a killed git leaves them behind on exactly the wedged mounts this PR targets, nothing ages them out, and the resulting ErrorKind::Other failure fell through the once-per-workspace notice gate — silent undo-history loss. The open sweep clears them on the index.lock staleness rule; fresh locks stay untouched. The pin covers stale removal (flat and nested) and fresh-kept.
  • client — the Anthropic transport's new failure marking is gated off the isolated Auto-classifier request: a read-only inspection must not write the shared connection health or fire a real /models probe (the isolated-dispatch contract the generic send paths already enforce). Pinned with a two-stall isolation test (expect(0) on the probe).
  • runtime — see the runtime bullet above for the rollback guard; the monitor-death settlement's dropped-publish log now names the approval id (the only one of the six error logs missing it).
  • tools — the group-kill's ESRCH arm falls through to the child-only kill: a direct child can setpgid itself out of its own group and stay alive, and the early return left the un-timed wait running past the budget.
  • tools — the plugin truncation refusal bounds the echoed stdout (codewhale_hooks::bounded_text, 512 chars, control bytes stripped) instead of embedding up to 16 MiB of captured NULs in the error text.
  • docs (snapshot) — the paths module doc no longer claims "cheap to call repeatedly" nor a turn-pipeline pre-canonicalization that does not exist; resolution pays one bounded probe per call, so hot paths should cache.
  • docs (mcp, en + zh) — supersedes the round-8 send-leg disclosure: the TUI pool's send leg IS deadline-bounded (the request budget wraps the send), so the Server Fields paragraph drops the false sentence; the unbounded blocking write is now disclosed where it lives, in a new engine-side stdio proxy paragraph (fixed 30s handshake / 120s request / 1800s tool-call budgets).

Review also disclosed two drive-bys the earlier sections had not named: the app-server drop(config) of the resolved endpoint config before the upstream round trip (releases the config read guard so writes are not blocked for up to UPSTREAM_TOTAL_TIMEOUT), and the core/turn.rs hunks (notify-gate widening for ErrorKind::Other init failures, the stage-keyed remedy-hint router, and their marker-based tests), which belong to the snapshot bullet above.

Round-10 review follow-ups (six commits on top of a clean rebase onto pinvou3-clean @ 44921cf / #80)

From the next fresh review wave of the full diff:

  • snapshot — the open-path ref-lock sweep now also clears a stale .git/HEAD.lock: git update-ref HEAD takes the HEAD lock and the resolved branch's lock one after another, so a git killed mid-update can leave a lone HEAD.lock, and a leftover one never makes the side repo unready — the init-path sweep (previously the only one covering it) never runs again on an initialized repo, so every later snapshot would fail at the ref update with the loss surfacing nowhere. The pin covers stale removal (branch, nested, HEAD) and fresh-kept (packed-refs and HEAD), mutation-checked red against the reverted sweep.
  • client — the round-9 isolation gate covered only the Anthropic transport-failure arm; check_anthropic_response still marked shared connection health (and could fire the /models recovery probe) for HTTP 4xx/5xx envelopes on the isolated Auto-classifier request, while the generic dialect's isolated dispatch path writes nothing. The status arm's failure and success marks are gated the same way, pinned by a two-500 isolation test (expect(0) on the probe), mutation-checked red.
  • tasks (turn.rs) — the default arm of the stage-keyed remedy router no longer promises a stale index.lock for every non-init git timeout: an interrupted update-ref/pack-refs leaves a ref lock instead, and the sweep covers all of them.
  • docs/comments — the debounce-flush starvation note no longer claims the heartbeat stream cannot starve the flush (heartbeats never arm dirty, but their ~200ms cadence still restarts the 250ms debounce, so dirty state armed just before a silent-tool phase stays unpersisted for that whole phase); the group-kill ESRCH comment no longer describes a setpgid-escape scenario (the child leads its own group and a leader cannot leave it — the fall-through stays as a defensive last resort on any kill error); the snapshot open-path comment no longer claims the whole path degrades "within one probe budget" (several bounded legs each degrade within their own); RUNTIME_API scopes posture to execution-policy denies (the full-access auto-approve also forces outcomes but carries auto: true, not a posture); MCP (en + zh) says the engine-side stdio proxy ignores the timeout settings — per-server command/args/env still apply.

Testing

cargo fmt --check; clippy -D warnings on codewhale-tui / codewhale-app-server / codewhale-core / codewhale-mcp; on the round-9 tree cargo test -p codewhale-tui --lib is 11975 tests / 13 ignored and -p codewhale-app-server --lib 103/0, green in isolation and green modulo the documented remote_control load flakes plus two sandbox::process_hardening no_new_privs tests that fail identically on the unmodified wave head (environmental on this box); the load-bearing pins were each mutation-checked to fail red (round-9 adds four more pins — gate exit-code absence, ref-lock sweep, anthropic isolation, rollback guard — each verified red against its reverted fix; round-10 adds two — the HEAD.lock sweep above and the anthropic status-arm isolation, likewise red-verified).

Round-7 review follow-ups (nine commits on top of the squashed wave)

  • snapshot — the 16 MiB git-capture cap is now truncation evidence: the restore-diff path listing and the snapshot history parsers refuse a capture that ends with the cap note instead of deciding on a truncated listing (an over-cap ls-tree would have made restore delete files that are present in the snapshot; the pre-restore safety snapshot kept the loss recoverable, but the resulting tree was still wrong). The cap note itself is pinned by a real-child drain test.
  • tools — drain_truncated now matches both truncation notes on both streams, so a plugin whose stdout hit the capture cap can no longer pass a cut-off output as a success; the escalation message names both causes.
  • app-server — a turn that ends in stream failure (resume budget exhausted or writer gone) now writes the last consumed seq back to the thread cursor, so the thread's next message no longer replays the failed turn's events from the store.
  • tasks — the GIT_INIT_FAILED remedy hint now leads with the stale-lock sweep (~1 h self-heal) and names directory removal only as the persistent-failure fallback; the old text pointed straight at removal, which discarded the workspace's whole undo history for a self-healing state (a concurrent second init lands in the same arm transiently).
  • runtime — a dropped approval.decided publish on the interrupted exit is logged at error severity instead of .ok().
  • tasks — the trailing drain loop skips liveness heartbeats like the main loop's short-circuit, instead of taking the manager-wide state lock for a no-op.
  • client — non_streaming_request_envelope stays private; the client::anthropic child module can see the parent's private items.
  • mcp — the engine-side stdio proxy's unbounded send leg is now disclosed at Connection::send, not only in this description.
  • docs — the per-leg MCP budget claim is scoped to stdio (an HTTP server is bounded by one total timeout; en + zh); RUNTIME_API scopes "every resolution is published as approval.decided" to approvals whose approval.required event was published, and states the raw-vs-compat posture shapes; zh_hans/SUBAGENTS syncs the stale default_max_steps example (120 → 0/unbounded) and the missing omitted/zero sentences. Also declared here: the zh_hans sync and restore's target_short get(..12) panic-safety hardening were drive-by fixes in the squashed wave that this description had not called out.

Round-8 review follow-ups (eight commits on top of round-7)

  • tools — drain_truncated anchors the match to the capture tail: both writers append their note as the buffer's last bytes, so output that merely quotes a note mid-body is no longer truncation evidence (the same tail contract the snapshot capture's capture_hit_the_cap pins); mid-buffer negatives are pinned too, and a real-child end-to-end test now covers the plugin surface refusing a 17 MiB unparseable stdout instead of passing the cut capture as a success.
  • runtime — the four remaining .ok()-swallowed approval.decided publishes (auto-approve, auto-review deny, user allow, user deny) log at error severity with the approval id, matching the round-7 forced-arm fix; no control-flow change. The settle path's per-entry comment now matches live-receiver settlement instead of contradicting its own block comment.
  • tasks — the GIT_INIT_FAILED hint names what the removal fallback costs (the workspace's snapshot history), and its test pins sweep-before-remove ordering so the destructive-first regression cannot come back green.
  • tasks — the debounce flush's starvation shape is disclosed at the select arm: biased polls the event arm first and the debounce sleep restarts every iteration, so a stream denser than the debounce interval defers persistence to a stream pause or the trailing flush (heartbeats never arm dirty; dense unpersisted deltas can). Debt disclosed, not yet fixed.
  • docs — SUBAGENTS (en + zh): only an explicit zero max_steps ignores an operator default; an omitted budget falls back to it (the previous sentence contradicted the pinned resolve_max_steps(Scout, None, Some(90)) == 90 and the priority list above it — round-7 had translated the wrong sentence into Chinese). RUNTIME_API: the approval.required emit-failure parenthetical now says "the tool call never executes" instead of pinning one deny path (the auto-approve branch stops the call via the monitor cancel, not an explicit deny). MCP (en + zh): the budget contract discloses the unbounded stdio send leg, matching the in-code disclosure.

Known limitations (deliberate, follow-up scale; pre-existing on base unless noted)

  • Windows job-object group-kill parity (child-only kill there, module-doc disclosed).
  • The contested inference-permit shape in web_search / voice / prompt_suggestion.
  • The pandoc.rs old-shape timeout; the two bounded runners' parallel drain implementations.
  • The engine-side MCP stdio proxy's unbounded send leg; debounce starvation while heartbeats stream (both now disclosed in code and docs, fixes are follow-up scale).
  • A pre-existing test-isolation race (CODEWHALE_MAX_OUTPUT_TOKENS is set process-wide by env-locked tests while unlocked readers can interleave) — separate hygiene PR.
  • A panic-only window in the approval contract: if a waiter unwinds between delivery/removal and the approval.decided publish, settlement cannot rescue the entry (it only rescues map-resident ones), leaving a journal approval.required without a decided — disclosed, follow-up scale; closing it means restructuring the delivery linearization.
  • The permit-outside-the-bound shape is wider than the three named sites: create_message, create_message_stream, translate, synthesize_speech, fim_completion, the no-cache classifier request, and provider-native search all acquire the inference permit before the bounded request — pre-existing on base, nothing new introduced here.
  • open_existing reports a half-initialized side repo as "no restore points" instead of the underlying error until the next open heals it (deliberate healing-first design; a review-observed truthfulness nit).
  • A pre-existing app-server asymmetry (round-10 review): a turn that ends server-side in Failed/Interrupted also skips the seq-cursor writeback (stream_result? propagates before the insert), so the next message replays that turn's event tail once (turn-id filtering keeps it correct); the round-7 writeback covers stream failure only. Self-heals after one turn.
  • A pre-existing app-server window: a writer death between turn creation and the first stream byte orphans the turn with no interrupt and no cursor writeback (byte-identical on base; turn-id filtering keeps the next message correct, but the busy thread rejects messages until the orphan settles).

No-Issue: review-fix replay for the merged #63; no separate tracking issue exists.

@JensenChen28 JensenChen28 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复核当前 head 后发现 1 个阻塞问题:快照初始化路径宣称已完整限制挂死文件系统探测,但仍会立即执行一次无界的 workspace canonicalize,因此原来的永久卡住风险尚未闭合。其余变更与当前 required checks 未发现新的阻塞项。

Comment thread crates/tui/src/snapshot/repo.rs
@asto18089 asto18089 changed the title fix: land the round-6 review fixes for the merged timeout audit (#63) fix: land the timeout-audit review fixes for the merged #63 Sep 27, 2026
@asto18089

Copy link
Copy Markdown
Collaborator Author

评审 thread 清理(2026-09-28):45 个提交压缩为 1 个(61cb769be → 1f69acd,最终树与压缩前完全一致,CI 结论可直接沿用),round-6b/7/8 等逐轮流水账评论已删除,其结论并入 PR 描述的 per-area 清单与 Known limitations;原 round-6 完整审阅记录仍在 #63。清理前引用的旧 commit SHA 已全部失效。附带 PR #78(human-wait-pragma)已 rebase 到新 head,内容不变。JensenChen28 的 CHANGES_REQUESTED(快照安全检查无界二次 canonicalize)已由该压缩提交内的 snapshot 部分修复,inline 回帖已改为自包含内容,请复审。

@JensenChen28 JensenChen28 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复核当前 head 1f69acd:上轮指出的快照初始化无界二次 canonicalize 已修复,安全分类整体通过同一有界文件系统入口执行;新增的流断线续传、超时进程组回收和审批结算路径未发现新的阻塞问题。验证:git diff --check 通过;快照安全分类、app-server 中途断流续传、Unix 超时进程组回收 3 项针对性测试均实际执行且各 1 passed / 0 failed;当前 required checks(Signed-off-by、Gitleaks、check、gate)均通过。当前环境缺少 rustfmt 组件,格式结论采用 CI check 回执。

@asto18089

Copy link
Copy Markdown
Collaborator Author

Fresh full re-review of the current head (1ee9b80, 18 commits, directly on 61cb769 — no rebase needed, mergeable). 11 area reviewers over the whole diff plus a full diff-vs-body accounting; every load-bearing finding below re-verified by hand against the tree.

Bottom line: the PR does real, root-caused work and is honest about its debts (all seven "Known limitations" verified present and pre-existing on base; no smuggled change requires splitting; commits and DCO clean; CI green). Four MAJOR findings need a decision before merge — two are objective fixes, two are completeness/claims judgments.

Verification run

  • cargo fmt --check clean; -p codewhale-tui --lib 11970 passed / 0 failed / 13 ignored; -p codewhale-app-server --lib 103/0; all 10 CI checks green on this head.
  • clippy --all-targets --all-features -D warnings on this box (stable 1.98.1) fails with 8 errors, but 7 are in files this PR does not touch (pre-existing under the newer local toolchain; CI is green, so CI's shape differs). One is on a line this PR adds: crates/tui/src/tools/tasks.rs:1426 assert_eq!(meta["spawn_error"].is_string(), false, …) trips clippy::bool_assert_comparison — assert!(!…) fixes it. Worth folding in so the next stable bump doesn't eat it.

MAJOR

M1 — gate timeout now fabricates exit_code = 137, violating the documented hook contract. crates/tui/src/tools/tasks.rs:647-651: the new run_bounded_child_observed path maps a timed-out gate to status.code().or(Some(128 + 9)). That integer flows into tool metadata (tasks.rs:722), and the consumer contract at crates/tui/src/tui/tool_routing.rs:607-617 (reported_tool_exit_code, pinned by the Hmbown#455 test) exists precisely so "an exit_code condition never matches on a fabricated value". Base recorded None on timeout (timed_out: true was already in the same metadata), and TaskGateRecord.exit_code is Option<i32>, so the "schema always carries an integer" justification doesn't hold. Side effect: a user hook on exit_code == 137 (the conventional OOM code) now fires on every gate timeout. Fix: keep None on the timeout arm (or exempt timed_out at the consumer), and pin it.

M2 — the MCP doc's unbounded-send-leg disclosure is false for the client the paragraph documents. docs/MCP.md:447-449 (and zh_hans 350-351): "The send leg itself is not deadline-bounded (disclosed debt)…" sits in the paragraph that documents the TUI pool's per-server connect_timeout/execute_timeout/read_timeout — but that pool's send leg IS bounded: tokio::time::timeout(request_budget, self.send(…)) at crates/tui/src/mcp.rs:2188-2196 (pre-existing), and the same-PR comment at mcp.rs:534-541 says the opposite of the doc. The genuinely unbounded leg is in the other client — codewhale-mcp's Connection::send (crates/mcp/src/stdio_client.rs:757-771) — which does not consume these knobs at all (fixed 30/120/1800s). As written the sentence is not true of either implementation and contradicts the "per leg… total toward twice execute_timeout" sentence right before it. Fix: move/scope the disclosure to the engine-side stdio proxy (en+zh).

M3 — "the cancellation contract is pinned red-green" overstates coverage of the exec fix. The approval.rs diff is 79 added test lines, 0 deletions: cancel_user_input_resolves_the_headless_wait_deterministically pins the pre-existing engine behavior (its own comment says so), and passes with the fix reverted. Nothing anywhere drives Event::UserInputRequired through a headless run_exec_agent — reverting the actual fix (exec_agent.rs:704-77) leaves the suite green, which is exactly the park-forever bug. Either add a headless-exec pin (inject an engine that emits UserInputRequired, assert the run settles) or soften the body claim to "the engine half of the contract is pinned".

M4 — a killed git update-ref leaves a ref lock nothing sweeps: the same silent-loss class this PR's snapshot work targets. The commit path ends in update-ref HEAD (snapshot/repo.rs:539, prune at :938), which takes refs/heads/<branch>.lock (or packed-refs.lock under gc). The PR's own rationale applies verbatim: on a wedged mount the 300s SIGKILL "cannot run git's lock cleanup" and "git never ages lockfiles out" — but the sweeps cover only config.lock/HEAD.lock at init (repo.rs:348) and index.lock at open (repo.rs:1124). A leftover ref lock makes every future snapshot fail at commit, and because that failure is ErrorKind::Other with no marker, the notify gate (core/turn.rs:482) early-returns — silent undo-history loss, the exact §2.7 failure mode. Judgment call: fix here (add ref-lock patterns to the sweep + a marker) or file as the declared follow-up scale; but the current state is an unfixed sibling of the class the PR claims to have healed.

MINOR (should fix, non-blocking)

  • client/anthropic.rs:231-246: the new transport-failure mark_request_failure + maybe_probe_recovery fires unconditionally, including on the isolated Auto-classifier path — the isolated clone shares connection_health via Arc, contradicting the documented "no shared connection-health writes" isolation contract at client.rs:3070-3072 (and maybe_probe_recovery issues a real /models GET). One-line fix: gate on isolated_request_state. Blast radius small: ConnectionState only gates probing/logging.
  • snapshot/repo.rs:216-227: open_existing now maps a half-initialized side repo to Ok(None) → export UI says "No workspace restore points…" where base surfaced a real error. Next open_or_init heals it, so small — but it's a truthfulness regression for exactly the state this PR introduces.
  • runtime_threads.rs:3887-3904: "race-free by construction" has one residual window — if the monitor future unwinds between entry-removal and the approval.decided publish, settlement (which only rescues map-resident entries) can't help; journal keeps required without decided forever. Narrow (needs a panic inside the waiter arm) and fail-closed for the caller.
  • runtime_threads.rs:9399: the approval.required-emit-failure rollback uses unconditional cancel_pending_approval, which can remove a live entry registered by another waiter under a reused id — the thing the PR's own new doc on cancel_closed_pending_approval forbids. Requires durable-append failure + cross-thread id reuse; outcome is fail-closed.
  • Known-limitations undercount: the "contested inference-permit shape" list (web_search / voice / prompt_suggestion) misses 7 more pre-existing permit-outside-bound sites, including the main chat path (client.rs:3295, 3312, 2383, 2741, 4075, 3221, provider_native_search.rs:117). Nothing new introduced — but the disclosure reads as exhaustive when it isn't.

Body corrections / disclosures wanted

  • app-server/chat_completions.rs drop(config) before the upstream round trip is the one behavioral unclaimed hunk in the diff (releases the config read guard so writes aren't blocked up to 1800s). Benign and audit-adjacent — but say so in the body.
  • core/turn.rs (+109: notify-gate widening for ErrorKind::Other init failures, the stage-keyed remedy-hint router, 5 marker tests) is part of the claimed snapshot work but never named in the body; the round-7/8 hint rewrites land there, not in tools/tasks.rs.
  • Declared drive-bys all present as stated; the undeclared ones are trivial comment/test-only (cli/update.rs math fix, sandbox/backend.rs, web/fetch.rs, web/contract.rs schema pin, mcp.rs http_total comment, repl/runtime.rs, snapshot reader-counter test-isolation rework). Nothing unrelated enough to split.

Verified good (evidence-backed)

  • Snapshot truncation evidence is airtight: both decision-making consumers (tree_paths, list) refuse cap-tailed captures before parsing; the note is appended atomically under the buffer lock and can never be evicted; refusal failure modes are fail-closed everywhere (restore aborts before checkout, prune warn-and-continue).
  • Process-group kill correct on Unix (process_group(0), kill(-pgid), kill-before-drain, no speculative group signal, pid-recycle excluded), Windows child-only parity disclosed; drain_truncated tail contract pinned by unit negatives that fail under contains.
  • App-server resume: exclusive-since_seq arithmetic verified at the seam, budget 1+2 connections, cursor writeback on every stream-failure exit before ?, turn_id-filtered replay, all three pins revert-sensitive. /health probe genuinely bounded per attempt now (base could hang forever on an accept-and-stall child).
  • Approval contract: all exits publish or error-log with ids, zero .ok() remain, settle-vs-deliver linearizes on the map mutex, compat posture forwarding complete; RUNTIME_API wording matches code line-for-line.
  • The 1800s roster is complete (13 family members enumerated and cross-checked); zh_hans carries no drift; SUBAGENTS omitted-vs-zero now matches the pinned resolve_max_steps(Scout, None, Some(90)) == 90.
  • Commits: 18 = 1+9+8 as described, all DCO-signed, messages match diffs, rework chains (drain tail anchor, hint text, docs) end coherent; repo squash-merges, so intermediate bisect risk is moot.

Recommendation: land M1 + M2 (objective), resolve M3 (pin or reword), decide M4 (fix or declared follow-up), fold in the two cheap MINORs (isolated-path health gate, tasks.rs:1426 clippy) and the body corrections. Everything else can merge as-is after that.

@asto18089

Copy link
Copy Markdown
Collaborator Author

All four MAJORs and the actionable MINORs/NITs are now fixed in eight commits pushed to the branch (26 total; CI re-running). Every new pin was mutation-checked red against its reverted fix, then green; on the round-9 tree cargo test -p codewhale-tui --lib is 11974/0 (13 ignored), -p codewhale-app-server --lib 103/0, cargo fmt --check clean, and workspace clippy --all-features -D warnings clean (the --all-targets strict form now fails only on the seven pre-existing untouched-file lints from the newer local toolchain — the PR-added tasks.rs:1426 violation is fixed).

M1 — fabricated exit code: fix(tasks) 241c073 — a timed-out gate records no exit_code again (timed_out: true stays the signal); absence pinned by a real timed-out gate.

M2 — false MCP doc claim: docs(mcp) 797fbd6 — the false "send leg is not deadline-bounded" sentence is dropped from the Server Fields paragraph (that pool's send leg IS bounded); the unbounded blocking write is now disclosed where it lives, in a new engine-side stdio proxy paragraph (en+zh, fixed 30/120/1800s budgets).

M3 — overstated exec pin: resolved by rewording the body (no harness exists for run_exec_agent, and building one is follow-up scale) — the body now says the engine half is pinned and the exec-side arm is not.

M4 — unswept ref locks: fix(snapshot) b8aa807 — the open sweep now clears stale refs/heads/**.lock and packed-refs.lock on the index.lock staleness rule (depth-capped walk); pinned for stale removal (flat + nested) and fresh-kept.

MINORs/NITs fixed:

  • 9a78845 fix(client) — Anthropic transport failure marking gated off the isolated classifier path; pinned (two stalls must leave shared health untouched and fire no probe).
  • e5c3264 fix(runtime) — the emit-failure rollback removes only a same-thread match (cancel_thread_pending_approval), so a reused id can't strand a foreign waiter (pinned both sides); the monitor-death settlement log now names the approval id.
  • 509b97c fix(tools) — the group-kill ESRCH arm falls through to the child-only kill (an escaped setpgid child no longer outlives the budget).
  • 39b531a fix(tools) — the plugin truncation refusal echoes at most 512 bounded characters instead of up to 16 MiB of capture.
  • 12dd44c docs(snapshot) — the paths module doc no longer claims "cheap to call repeatedly" or a nonexistent turn-pipeline pre-canonicalization.

Not fixed, disclosed in the body's Known limitations instead (judgment calls, follow-up scale): the panic-only lost-publish window between delivery and publish (closing it means restructuring the delivery linearization), the open_existing half-init → "no restore points" truthfulness nit (healing-first design), the wider permit-outside-bound family (7 more pre-existing sites), and the pre-existing app-server orphan window between turn creation and first stream byte. The drop(config) and core/turn.rs drive-bys are now named in the body as well.

#63 (d349f25) merged mid-review, so its final review wave never
reached `pinvou3-clean`. This PR carries that wave plus the fixes
from the subsequent review rounds onto the current base, as one
commit; per-area detail lives in the PR description.

- exec: headless runs resolve `request_user_input` via
  `cancel_user_input` instead of parking forever once the engine
  wait became unbounded; the cancellation contract is pinned.
- snapshot: every pre-git window on a wedged mount is bounded
  (workspace canonicalize, the safety classifier's
  re-canonicalization, the first-init size walk, honoring
  `max_workspace_gb = 0`); a killed `git init` heals instead of
  wedging (HEAD-less, half-written, leftover config.lock/HEAD.lock);
  wedge failures surface as errors with stderr evidence and
  stage-keyed remedy hints; drain captures cap at 16 MiB per stream.
- runtime: every approval exit publishes `approval.decided` and is
  race-free by construction (delivery races, monitor-death
  settlement, channel-closed exit); compat streams forward the
  approval posture.
- tasks: liveness heartbeats short-circuit before the manager-wide
  state lock. The human-wait clock pause lives in #78 instead: as
  first landed it changed task-lifecycle behavior and regressed the
  TUI, so it was taken back out.
- vision: the envelope timeout classifies as `ToolError::Timeout`,
  and the inference permit is acquired inside the envelope.
- client: envelope-exceeded requests mark connection health and run
  the recovery probe on every dialect arm, including the Anthropic
  transport arm; the non-streaming budget got an injectable test
  seam.
- app-server: a dropped stream resumes from the last consumed seq
  (bounded reconnects) before falling back to a best-effort
  interrupt, on every turn surface; the `/health` boot probe bounds
  each attempt.
- tools: drain aborts are one helper covering every exit including
  cancel/drop; timeouts SIGKILL the child's process group so forked
  commands no longer outlive the call.
- rlm: the sub-query budget is floored at 1s and the session default
  references `CHILD_TIMEOUT_SECS`.
- mcp: the send-timeout error states the budget instead of tokio's
  `Elapsed` display.
- docs: the 1800s timeout family anchors one complete roster;
  MCP/SUBAGENTS disclose the per-leg budget and wall-clock
  interactions; the RUNTIME_API.md restore contract is corrected.
- tests: every fix above has a regression pin (mutation-verified
  red-green), timing margins are widened, and the Anthropic
  endpoint tests hold the env lock.

Deliberately not fixed here (pre-existing on base unless noted,
follow-up scale): Windows job-object group-kill parity; the
contested inference-permit shape in web_search/voice/
prompt_suggestion; the pandoc.rs old-shape timeout; the two bounded
runners' parallel drain implementations; the engine-side MCP stdio
proxy's unbounded send leg; debounce starvation while heartbeats
stream; the `CODEWHALE_MAX_OUTPUT_TOKENS` test-isolation race.

Verified: `cargo fmt --check`; clippy `-D warnings` on
codewhale-tui, codewhale-app-server, codewhale-core and
codewhale-mcp; `cargo test -p codewhale-tui --lib` 11969/0 and
`-p codewhale-app-server --lib` 103/0 on the final tree; the
load-bearing pins were each mutation-checked to fail red.

No-Issue: review-fix replay for the merged #63; no separate
tracking issue exists. The full findings ledger is the round-6
review comment on #63.

Signed-off-by: asto <asto18089@126.com>
…ry parsers

The 16 MiB drain cap appended a truncation note that no parser
consumed: an over-cap `ls-tree -r` listing would silently shrink the
restore diff's target set, making restore delete files that exist in
the snapshot (the pre-restore safety snapshot makes the loss
recoverable, but the resulting tree is still wrong), and a truncated
`git log` could end in a half-parsed pseudo entry. Both parsers now
refuse a capture that ends with the cap note, and the note itself is
pinned by a real-child drain test.

Signed-off-by: asto <asto18089@126.com>
…ugin surface

`drain_truncated` only matched the drain-grace note and only on
stderr, so a plugin whose stdout hit the 16 MiB capture cap came back
as `Ok(success)` carrying cut-off JSON plus a tail note no detector
read. The detector now matches both truncation notes on both streams,
the plugin escalation message names both causes, and a combination
test pins the detector.

Signed-off-by: asto <asto18089@126.com>
…am failure

The cursor was only persisted on the `Ok` path, so a thread whose turn
died past the resume budget replayed the failed turn's events from the
store on every retry (turn_id filtering kept it correct; the cost grew
with the history). The error path now persists the last consumed seq,
pinned by extending the exhausted-resume test.

Signed-off-by: asto <asto18089@126.com>
The GIT_INIT_FAILED hint told the user to remove the side-repo
directory first, but that state does clear itself: the open-path sweep
clears a lock once it has been stale for an hour (this PR's own
mechanism, pinned by its own tests), and a second host initializing
the same workspace concurrently lands here transiently. Following the
old hint discarded the workspace's whole undo history for a state that
heals within the hour; the hint now leads with waiting and names
removal only as the persistent-failure fallback. The hint test's
negative assertion pinned the wrong premise and now pins the new one.

Signed-off-by: asto <asto18089@126.com>
…wing it

The interrupted/channel-closed exit published `approval.decided` with
`.ok()`, so a failed publish silently stranded the client's pending
UI; it now logs at error severity like the settle path does.

Signed-off-by: asto <asto18089@126.com>
The heartbeat short-circuit covered the main event loop, but the
post-execution drain called `apply_execution_event` directly, taking
the manager-wide state lock for a no-op against the invariant the
short-circuit states.

Signed-off-by: asto <asto18089@126.com>
`non_streaming_request_envelope` was widened to `pub(crate)` for a
caller in the `client::anthropic` child module, which can already see
the parent's private items.

Signed-off-by: asto <asto18089@126.com>
…in code

The debt was only stated in the PR description; `Connection::send` now
carries the disclosure where a reader of the code will meet it.

Signed-off-by: asto <asto18089@126.com>
- MCP: the per-leg budget shape holds for stdio servers only; HTTP
  servers are bounded by a single total timeout. Scope the sentence
  and state the HTTP bound (en + zh).
- RUNTIME_API: scope "every resolution is published as
  `approval.decided`" to approvals whose `approval.required` event was
  published and disclose the event-failure exit; the compat stream
  maps an absent `posture` to `null` while the raw stream leaves it
  absent.
- zh_hans/SUBAGENTS: the example `default_max_steps = 120` and the
  missing omitted/zero sentences were stale against the en doc and the
  code (defaults are zero/unbounded); sync them.

Signed-off-by: asto <asto18089@126.com>
The drain detector matched its notes anywhere in the buffer, while both
writers append their note as the buffer's last bytes and stop retaining
past it. Anchor the match to the tail — the same contract the snapshot
capture's capture_hit_the_cap pins — so a stream that merely quotes a
note mid-body is not truncation evidence, and pin the stricter
semantics with mid-buffer negative cases.

Also close the integration gap the round-7 review flagged: a real
child flooding 17 MiB of unparseable stdout through the plugin surface
must be refused as truncated, not reported as a silent success.

Signed-off-by: asto18089 <asto18089@126.com>
…them

Four approval.decided publishes still ended in .ok(), so a storage
failure silently dropped the event that tells clients to clear their
pending approval UI — the same failure mode the forced-resolution arm
already logs loudly. Log all four (auto-approve, auto-review deny, user
allow, user deny) with the approval id; no control-flow change.

Signed-off-by: asto18089 <asto18089@126.com>
…ement

The per-entry comment still claimed the send cannot deliver because the
receiver died with the monitor, contradicting the block comment above
it: entries whose receiver is still live are settled too, and the deny
is then pushed through the decision channel (the first settle fixture
pins exactly that).

Signed-off-by: asto18089 <asto18089@126.com>
The init-failure remedy already led with the self-healing sweep, but the
user-facing hint never said what the removal fallback costs — that
warning lived only in a code comment. State it inline, and pin the
ordering with a test assertion so a regression that puts removal before
the sweep cannot come back green.

Signed-off-by: asto18089 <asto18089@126.com>
The biased select polls the event arm before the flush arm and the
debounce sleep restarts every iteration, so a stream denser than the
debounce interval defers persistence until the stream pauses or the
loop's trailing flush lands it. Heartbeats never arm dirty, so the
silent-tool tick stream cannot starve it; a dense run of unpersisted
deltas can. Disclose it where the delay lives instead of leaving the
debt only in the commit message.

Signed-off-by: asto18089 <asto18089@126.com>
"Omitted or zero max_steps remains unbounded even when an operator
default is configured" is half false: only an explicit zero ignores the
operator default (resolve_max_steps keeps 0 as unbounded), while an
omitted budget falls back to it — pinned by
resolve_max_steps(Scout, None, Some(90)) == 90 and contradicted by the
priority list two paragraphs up. State the real rule in English and
sync the Chinese doc that round-7 translated from the wrong sentence.

Signed-off-by: asto18089 <asto18089@126.com>
…g one path

"The tool call is denied" only literally describes the registration
branch; when the auto-approve branch's required emit fails, the tool
call is stopped by the monitor's cancel instead. The externally
observable contract is the same either way — the tool never executes —
so state that.

Signed-off-by: asto18089 <asto18089@126.com>
The budget paragraph read as if both stdio legs were bounded, but the
send leg has no deadline — a server that stops draining stdin can park
the write past every budget (disclosed at the write site itself in
round-7). Mirror that caveat in the English and Chinese docs so the
contract matches the code.

Signed-off-by: asto18089 <asto18089@126.com>
The gate moved onto the bounded runner and mapped a timed-out child to
the conventional 128+SIGKILL so the record schema always carried an
integer. That integer reaches hook `exit_code` conditions through
tool metadata, and `reported_tool_exit_code` exists precisely so an
`exit_code` condition never matches a fabricated value (128+SIGKILL
doubles as the OOM-kill code, so OOM hooks fired on every gate
timeout). Base recorded no exit code for a timed-out gate; `timed_out:
true` in the same metadata stays the timeout signal. Pin the absence
with a real timed-out gate, and fix the bool-assert comparison in the
neighbouring spawn-evidence test while there.

Signed-off-by: asto18089 <asto18089@126.com>
The recovery work covered config.lock/HEAD.lock at init and index.lock
at open, but the commit path's `update-ref HEAD` (and prune's, and
pack-refs) takes refs/heads/<branch>.lock or packed-refs.lock, and a
git killed mid-update leaves those behind on exactly the wedged mounts
this recovery targets. Nothing ages them out, every later snapshot
fails at the ref update, and the failure is ErrorKind::Other with no
marker, so the once-per-workspace notice gate swallows it: silent
undo-history loss. Sweep both ref-lock shapes on the same staleness
rule as index.lock when the side repo opens; fresh locks stay untouched.

Signed-off-by: asto18089 <asto18089@126.com>
The new transport-failure marking in the Anthropic arm fires on every
request, including the isolated Auto-classifier one. That clone shares
connection_health through the Arc, so a read-only inspection wrote the
shared health and could fire a real /models recovery probe — the same
contract the generic send paths enforce by dispatching isolated
requests to send_with_isolated_retry. Gate the transport arm on
isolated_request_state and pin the isolation: two stalls from an
isolated client must leave the shared health untouched and fire no
probe.

Signed-off-by: asto18089 <asto18089@126.com>
Two approval-settlement follow-ups:

The approval.required-emit-failure rollback removed its registration
with a blanket map remove. Providers reuse tool-call ids across
threads, so under a reused id the entry may already belong to another
thread's live waiter and the rollback would strand it — the exact case
the cancel_closed_pending_approval doc forbids. Remove only a
same-thread match (cancel_thread_pending_approval), which retires the
now-unused blanket helper, and pin both sides of the guard.

The monitor-death settlement logged a dropped approval.decided publish
without naming the approval — the one path where several approvals can
be stranded at once, and the only one of the six error logs missing
the id. Add it.

Signed-off-by: asto18089 <asto18089@126.com>
kill_the_run returned without any kill when kill(-pgid) failed with
ESRCH, on the reasoning that an empty group means the child is dead.
That holds only for a child that stayed in its group: a direct child
can setpgid itself into another same-session group and stay alive, so
the empty group left the un-timed child.wait() running past the budget
with nothing enforced. Fall through to the child-only kill on ESRCH —
a no-op on a genuinely dead child, the last chance to stop an escaped
one.

Signed-off-by: asto18089 <asto18089@126.com>
The truncation refusal embedded the whole captured stdout in the
ToolError message, so a 16 MiB capped capture rode the error text
toward the transcript. Echo the first bounded characters instead
(codewhale_hooks::bounded_text also strips control bytes, so a capture
of raw NULs echoes nothing); the diagnostic job is the note plus a
preview.

Signed-off-by: asto18089 <asto18089@126.com>
snapshot_dir_for now runs the workspace canonicalization under the
bounded probe on every call, so the module doc's "cheap to call
repeatedly" promise and the comment's premise that the turn pipeline
already canonicalized the workspace no longer described this code.
State the actual contract: every caller pays one bounded probe, hot
paths should cache, and the raw-path fallback hashes identically for
canonical input.

Signed-off-by: asto18089 <asto18089@126.com>
The send-leg disclosure landed in the Server Fields paragraph, but
that paragraph documents the TUI pool's per-server settings — and that
pool's send leg is deadline-bounded (the request budget wraps the send
at mcp.rs call_method), so the sentence was not true of the client the
settings govern and contradicted the per-leg sentence above it. The
unbounded blocking write belongs to the engine-side stdio proxy, which
does not consume these knobs at all; drop the false sentence and give
the proxy its own short paragraph with its fixed budgets, in English
and Chinese.

Signed-off-by: asto18089 <asto18089@126.com>
update-ref HEAD takes the HEAD lock and the resolved branch's lock one
after another, so a git killed mid-update can leave a lone HEAD.lock.
A leftover one does not make the side repo unready, so the init-path
sweep never sees it again on an initialized repo, and the open-path
ref-lock sweep only covered packed-refs.lock and refs/heads/**.lock —
every later snapshot then fails at the ref update with the loss
surfacing nowhere, the exact wedge this sweep exists to age out. Add
the HEAD lock to the open sweep and pin both the stale-removed and
fresh-kept rule for it (mutation-checked red against the reverted
sweep).

Signed-off-by: asto <asto18089@126.com>
The round-9 isolation gate covered only the transport-failure arm;
check_anthropic_response still marked shared connection health (and
could fire the recovery probe) for HTTP 4xx/5xx envelopes on the
isolated Auto-classifier request, while the generic dialect's isolated
dispatch path writes nothing. Gate the status arm's failure and
success marks the same way, and pin it with a two-500 isolation test
whose expect(0) probe would go red against the reverted gate.

Signed-off-by: asto <asto18089@126.com>
The default arm of the stage-keyed remedy router promised a stale
index.lock for every non-init git timeout, but an interrupted
update-ref/pack-refs leaves a ref lock (HEAD.lock, refs/heads/**.lock,
packed-refs.lock) instead — the sweep covers all of them, so the hint
should not send a debugging user to one file name.

Signed-off-by: asto <asto18089@126.com>
Heartbeats never arm dirty, but their ~200ms cadence still restarts
the 250ms debounce, so dirty state armed just before a silent-tool
phase stays unpersisted for that whole phase — the old note claimed
the tick stream cannot starve the flush at all.

Signed-off-by: asto <asto18089@126.com>
The child leads its own process group and a group leader cannot leave
it, so the setpgid-escape scenario the ESRCH comment described is not
reachable; the fall-through stays correct as a defensive last resort
on any kill error.

Signed-off-by: asto <asto18089@126.com>
RUNTIME_API: posture rides execution-policy denies only (the
full-access auto-approve also forces outcomes but carries auto: true,
not a posture). MCP: the engine-side stdio proxy ignores the timeout
settings, not the whole server config — command/args/env still apply
(en + zh_hans).

Signed-off-by: asto <asto18089@126.com>
@asto18089

Copy link
Copy Markdown
Collaborator Author

Fresh full-diff review wave (round-10), run from scratch after rebasing the branch onto the advanced base (pinvou3-clean @ 44921cf / #80 — rebase clean, the tree delta is exactly #80, CI green on the rebased head before this push). The wave: five area-partitioned deep reviewers (snapshot+turn, approval contract, runtime-threads+app-server, tools process/plugin/tasks, client/vision/rlm/mcp/exec) plus four cross-cutting auditors (smuggled changes, body truthfulness, docs/zh parity, performance), with every load-bearing claim re-verified against the code and the findings re-verified by hand before anything was changed.

Verdict: the PR is sound at its core; the wave found one incomplete heal and one incomplete isolation gate, both fixed and mutation-pinned in this push, plus five small corrections. No smuggled changes, no perf regressions, no material body falsehoods.

What the wave confirmed

  • Root-cause quality: every claimed fix traces to real code and a real failure mode (the bounded probes, the readiness-predicate re-init, the approval-exit enumeration across every reachable path, the seq-resume exactly-once accounting, the process-group kill's pre-exec setpgid placement). The audit of the base-vs-PR known limitations confirmed each "pre-existing" claim is byte-identical on base.
  • Body truthfulness: every named pin exists on the branch, every numeric claim (16 MiB, 512 chars, depth 16, 1s floor, 30s/120s/1800s, 3600s staleness) checks out against the code, and every "disclosed in code" item is actually disclosed at the named site.
  • No smuggles: all 26 commits map to the declared scope or an explicitly disclosed drive-by; zero changes to CI, workflows, Cargo.toml/lock, or submodules; no debris, no debug leftovers; the only unsafe added is the declared group kill with its SAFETY comment. (The compat posture forwarding is a public-surface addition — it is covered by the body's runtime bullet and RUNTIME_API.md; flagging it here for explicit sign-off.)
  • Docs: every technical statement introduced by the docs hunks is true against the code; the 1800s roster is complete (every real budget site in crates/ maps to a roster entry); en/zh hunks are semantically aligned, and the zh side additionally repairs pre-existing drift.
  • Performance: the heartbeat short-circuit really avoids the manager-wide lock (~5 acquisitions/s per silent tool removed); drop(config) really released a read guard previously held across up to 1800s of upstream round trip on a write-preferring RwLock; the recovery probe is cooldown-deduplicated (no stampede); the 16 MiB caps are lazily grown and enforced on both streams and both writer sites; drain_truncated got faster (tail-anchored ends_with vs the old O(n·m) window scan).

Findings fixed in this push (six commits)

  1. [MAJOR, snapshot] The open-path ref-lock sweep missed .git/HEAD.lock. git update-ref HEAD — which every snapshot commit runs — takes the HEAD lock and the resolved branch's lock one after another (verified empirically: a lone leftover of either individually fails the next update), and a leftover HEAD.lock never makes the side repo unready, so the init-path sweep that covers it never runs again on an initialized repo. A kill mid-update could therefore still wedge snapshots permanently and silently — exactly the class the round-9 sweep was added to heal. Now swept on the same staleness rule, with the pin covering stale branch/nested/HEAD removal and fresh-kept for both packed-refs and HEAD (mutation-checked red against the reverted sweep). → 891004b
  2. [MINOR, client] The round-9 isolation gate covered only the transport-failure arm; check_anthropic_response still wrote shared connection health (and could fire the /models probe) for HTTP 4xx/5xx on the isolated Auto-classifier request, while the generic dialect's isolated path writes nothing. The status arm's failure and success marks are gated the same way now, pinned by a two-500 isolation test whose expect(0) probe goes red against the reverted gate (mutation-checked). → 3c649d5
  3. [MINOR, tasks] The default arm of the remedy-hint router promised a stale index.lock for every non-init git timeout, but an interrupted update-ref/pack-refs leaves a ref lock; the hint now names the lockfile family the sweep actually clears. → e620466
    4–6. Comment/docs corrections: the debounce-flush starvation note claimed the heartbeat stream cannot starve the flush, but heartbeats' ~200ms cadence restarts the 250ms debounce, so dirty state armed before a silent-tool phase stays unpersisted for that whole phase (→ 379b950); the ESRCH comment described a setpgid-escape that POSIX makes unreachable for a group leader — reworded as the defensive fallback it is (→ b8de056); the open-path comment claimed "one probe budget" where several bounded legs each have their own (folded into 891004b); RUNTIME_API's posture sentence now scopes to execution-policy denies (the full-access auto-approve forces outcomes too but carries auto: true), and MCP (en+zh) now says the engine-side proxy ignores the timeout settings — command/args/env still apply (→ 0362c2c).

Disclosed, not fixed (follow-up scale)

  • A turn that ends server-side in Failed/Interrupted still skips the seq-cursor writeback (stream_result? propagates before the insert) — pre-existing on base, self-heals after one replayed turn; the PR's writeback covers stream failure only, as its bullet says.
  • The notify-gate widening and the /health per-attempt bound have no dedicated regression pins (the latter needs a spawnable stub server).
  • bounded_text still walks the full capture to emit its 512-char echo (one transient full-size pass on the error path — still strictly better than base's retained full capture), and an all-control-byte capture yields an empty echo with no cut marker.
  • The new envelope-exceed test writes process-global retry state without retry_status::test_guard() — the same exposure two pre-existing tests already have; a hygiene PR could take the guard alongside lock_test_env().
  • send_with_isolated_retry applies no envelope of its own (pre-existing, untouched by this PR; noting for the permit/bound follow-up list).

Verification on this head (0362c2c)

cargo fmt --check clean; clippy -D warnings green on the four documented lib targets (the six --all-targets errors are the known pre-existing lints in untouched files, local-toolchain only); cargo test -p codewhale-tui --lib 11975 passed / 0 failed / 13 ignored and -p codewhale-app-server --lib 103 / 0. Environment note: at libtest's default stack this box aborts in acp_server::tests::new_session_starts_with_empty_messages with a stack overflow — reproduced byte-identically on base pinvou3-clean, so it is a local debug-frame/stack-size quirk, not a PR regression; the numbers above are with RUST_MIN_STACK=33554432. Both new pins were each reverted to red and restored to green. CI on the rebased head was green; the run for this push is in flight.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants