Skip to content

fix(explore): exclude unready todos from dispatch suggestions - #6022

Merged
huangruiteng merged 10 commits into
loopx-project:mainfrom
Jim-jimu:fix/explore-todo-readiness
Oct 9, 2026
Merged

huangruiteng merged 10 commits into
loopx-project:mainfrom
Jim-jimu:fix/explore-todo-readiness

Conversation

@Jim-jimu

@Jim-jimu Jim-jimu commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Goal And Delivered Outcome

  • Reproduced defect: Explore can select B while its completion dependency A is unfinished, suggest acquiring B's execution lease, and omit B's waiting condition from the Agent's turn context.
  • Before → after: both planners now exclude unready Todos from current dispatch, worker bundles and resource allocation. Waiting conditions remain visible in bounded diagnostics. Completing A makes B selectable on the next plan; handing A to another Agent does not.
  • Intended base: main. Standalone reproduced defect; no issue is closed. Complements the execution-admission checks in fix(todos): enforce canonical completion dependencies during execution #5884.

Implemented Against

  • Specification: docs/reference/canonical-lease-renew.md (handover and completion-dependency admission) and docs/reference/protocols/task-graph-projection-v0.md (typed Todo topology), revision f03a9acf28b726949a855ab54e2f4f1d06799a3f.
Criterion Disposition Implementation / validation
Unfinished completion dependencies prevent execution Implemented in planning suggestions Both planners reuse todo_item_is_actionable_open; readiness regressions and public CLI journeys
Handover preserves dependencies Preserved Transfer A, observe B still waiting, complete A, then acquire and complete B
Successor lineage is not a completion dependency Preserved Lineage compatibility case; no new graph interpretation
  • Self-check: reviewed planner, context and execution owners; reproduced the defect on the base; verified isolated File/SQLite CLI behavior. Shared readiness projection and exclusions between the existing planners without adding a dependency evaluator.

Scope And Continuation

  • Complete within this scope: Todo/worker branch selection, waiting diagnostics, bounded Agent context and empty-selection guidance. No successor required.
  • Additive packet fields preserve actionability, status and resume conditions. The Todo plan records the total rejection count; compact context includes up to three rejected candidates and an omission count.
  • Existing Explore opt-in, analysis-only mode, quota, claims and execution leases retain their authority. General graph scheduling and priority inheritance are outside scope.

Validation

  • Tested revision: 48b1898f3f30e80ea8c5d376e591e56ed14bc10e, including upstream f03a9acf. Older broad runs are identified below.
  • Run state: finished
  • Input classes: synthetic, public_fixture
Check kind Result Public-safe evidence / limitation
regression_parity passed Same 25 new regression cases: 23 target-defect failures and 2 compatibility passes on unmodified production at f03a9acf; all 25 pass with the fix, included in the 304 below
unit / integration passed Final revision: 304 Python tests across Explore, canonical dependency execution, cold-source imports and backup retention; command below
real_entrypoint / real_backend passed tests/control_plane/test_canonical_dependency_execution.py: public CLI with isolated canonical File and SQLite 3.51.3 stores; rejected B handover, successful A handover, unchanged wait, A completion, replanning and B completion
unit / integration passed Final revision: 36 TypeScript tests in cold_source_import.test.ts, todo_execution_dependency.test.ts and content_digest_single_owner.test.ts; Node 22.22.3, concurrency 1
static passed Final revision: Ruff, mypy, TypeScript typecheck, diff check, semantic advisory and CLI output-budget regression
integration passed Final revision: risk-selected premerge with --from-git-diff --git-diff-base origin/main: 5 direct and 19 selected checks passed, no manual holds; includes public-boundary checks
unit / integration passed Historical full TypeScript run on the feature contents committed as 81d0382e, before upstream integration: 4,252 passed, 32 PostgreSQL-dependent skips, 0 failed
unit / integration failed Historical full Python attempt before integration stopped at approximately 52%: 10,406 passed, 263 skipped, 7 failed. All seven failures reproduce on unmodified f03a9acf and the final revision in the two files listed below

Final focused Python command:

LOOPX_USAGE_PING=0 .venv/bin/python -m pytest -q \
  tests/test_explore*.py tests/capabilities/test_explore*.py \
  tests/control_plane/test_canonical_dependency_execution.py \
  tests/control_plane/test_cold_source*.py \
  tests/control_plane/test_chat_cold_source_import.py \
  tests/control_plane/test_backup_registered_source_retention.py
  • Existing failures: tests/cli_commands/test_periodic_report_intent_capability_chain.py and tests/cli_commands/test_todo_complete_settlement_capabilities.py assert missing report intents / settlement completion evidence. They are outside this diff and remain unresolved.
  • Coverage includes false/missing readiness, completion and Monitor-change conditions, resource capacity, shared-scope bundles, lineage, foreign ownership, analysis-only/off behavior and bounded diagnostic omissions.
  • Gaps: the full Python attempt was interrupted for upstream integration and was not completed; the full TypeScript suite was not rerun after integration. No live-model, packaged App/Lark or PostgreSQL journey was measured. No provider storage or execution-admission rule changes.

Frontend / Visual Evidence

  • UI impact: none. Agent-facing CLI/context projections change; no visual surface is modified.
  • Before / after, states / viewports and attention review: N/A.
  • Source data: none.

Type of Change

  • Bug fix
  • Documentation update
  • Test update

LoopX Area

  • Capability or extension (Explore)
  • Control plane (Agent context projection)

Technical Direction

  • Core control-plane hardening: align Explore dispatch suggestions with existing Todo readiness.

Shared-authority RFC fixture impact

N/A: no provider migration, runtime-routing or shared-authority RFC completion claim. Existing File/SQLite fixtures exercise the affected public readers and unchanged execution lifecycle.

Boundary Checklist

  • The diff and this description contain no private state, credentials, raw traces, internal links or local machine paths.
  • No maintainer-owned benchmark work was duplicated.
  • The change stays within the reproduced planning defect.
  • UI impact is marked none.
  • Every contribution commit includes a DCO Signed-off-by trailer.

mika and others added 3 commits October 8, 2026 22:50
Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>
…iness

Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>
…iness

Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>
Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>
Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>
loopx-agent
loopx-agent previously approved these changes Oct 9, 2026

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

使用 Explore 安排下一步工作的 Agent 与操作者。

前置 Todo A 尚未完成时,旧版可能建议领取等待 A 的 B,造成无效调度;本版把 B 留在诊断中,A 转交仍须等待,A 真正完成后 B 才重新可选。

已验证两个 planner 和 turn-context 的等待诊断、File/SQLite 真实 claim/lease/handoff/completion,以及关闭状态的完整输出对照。

不授予 spawn、claim、lease、quota 或外部发布权限,不声称模型实际采用、全套 CI 或长期运行已验收。

改动思路

共用既有 readiness 和 typed projection,可在一个规则位置维护两 planner,避免新的并行调度状态。

当前边界是正确的计划与等待诊断、现有测试维护和 CI 日志可观察性,不启动执行或扩张权限。

沿用统一的 Todo readiness,先排除不能行动的候选,再评分、分配资源和组 worker bundle。等待状态进入有界诊断;实际执行继续由原 quota/claim/lease 入口裁决。

规范依据:docs/reference/canonical-lease-renew.md,不可变版本 ec21f7da8b9a861fa73477f0ad69e8a4f30e5389。已在阅读实现之前读此版本;对应当前有界改动的接受项如下:

  • todo_done: Actual prerequisite done releases the canonical completion wait; transfer and successor lineage do not. implemented.
  • todo_dependency_pending: Unsatisfied wait fences execution, and release remains available. implemented.
  • resume_ready=true: Read the same current task after actual prerequisite completion before normal acquisition. implemented.

并对照同版本 docs/reference/protocols/task-graph-projection-v0.md:successor lineage 不是完成依赖。

具体改动

审查了全部 13 文件(+336/-80),包括最终加入的 workflow 和 periodic-report 测试维护。两 planner 共享 readiness/exclusion helper;TS turn-context 保留等待条件和总省略数。原 periodic 规则已明确普通完成不产生 milestone;基线七个旧期望失败在相同本地命令复现,本版对齐既有规则,实际 native successor 正例仍通过。CI job ceiling 从 45 改为 60 分钟且输出改为 verbose,未删 collection、test deadline、coverage 或汇总失败门槛;真实两 worker 的分片命令在其余测试仍等待时已输出失败名称。

独立验证:627 focused +71 adjacent Python tests; 158 semantic checks; TS typecheck; premerge five direct checks plus catalog/risk/boundary checks. Baseline native readiness oracle fails both File/SQLite; H passes. Base original periodic tests: seven stale expectation failures; current corresponding tests pass.

额外的完整 public CLI packet 对照覆盖 off、analysis-only、等待、完成和大量无关 Todo;off 仅归一化临时路径及独立验证的路径 hash 后完全一致。File/SQLite 原生 A 转交仍等待,A 完成后 B 领取/完成。另一个拥有更多 fallback 的基线 fixture 本来就没选 B,记录为合法分支,未伪称其失败。

未来改动可维护性检查:本次已把 readiness/exclusion 收敛为两 planner 共同使用的既有规则适配,不需要再建一套调度状态。

对主干的风险

No current material blocker found. 不授予 spawn、claim、lease、quota 或外部发布权限,不声称模型实际采用、全套 CI 或长期运行已验收。 Hosted four-shard full timing not qualified; job ceiling increase preserves tests/coverage. Author description still cites an older tested head, so this review judges the complete current head independently.

未查询、等待或继承 GitHub CI。全仓套件、真实模型对 guidance 的采纳和长期运行未做;现有 UI/消息展示入口没有新增操作,不声称 packaged frontend/Lark 验收。

我的整体评价

该精确 head 的当前有界结果成立,未发现阻塞缺陷,批准代码评审。合并仍是独立权限与仓库策略动作;本次不合并、不升级运行中的安装。

Reviewer: model_agent; runtime_reported; model=gpt-6.1-sol; provider=OpenAI; reasoning_effort=xhigh; execution_observation_id=de1d923ded17a66b1ad1ad14fa37aee1f64208c77ede42beac3b565c69abe4a5

English verdict: APPROVE - exact head 22d5827. Independently validated the whole current diff, real owning entrypoints/backends, recovery and baseline defect sensitivity; no CI was consulted. Parent profile/product qualification and merge authority remain separate.

…iness

Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>
Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>
Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; reasoning_effort=xhigh

APPROVE:未发现当前变更造成的阻塞问题;现有 CLI 输出预算仍为单独的 premerge hold。精确 head d8f2fd85b412b6d81200a2feb2179fbb427e6bfa,全差异基线 5e5f0576b15587a9233ce7a2d9571940bf22ad94。本次重新判断完整 16 文件,没有继承旧 head 22d5827553568202f4b502512005f6985ce1e8ad 的批准。

动机

使用 Explore 安排下一步工作的 Agent 与操作者。前置 Todo A 尚未完成时,旧版可能建议领取等待 A 的 B,造成无效调度;本版把 B 留在诊断中,A 转交仍须等待,A 真正完成后 B 才重新可选。已验证两个 planner 和 turn-context 的等待诊断、File/SQLite 真实 claim/lease/handoff/completion,以及关闭状态的完整输出对照。 这是完成现有计划入口的可用修复,不授予 spawn、claim、lease、quota 或外部发布权限,不声称模型实际采用、全套 CI 或长期运行已验收。

改动思路

从真实 Todo 读模型取得当前等待状态,在排序、worker bundling 和资源分配之前应用既有 todo_item_is_actionable_open;等待者只进入拒绝诊断。共用既有 readiness 和 typed projection,可在一个规则位置维护两 planner,避免新的并行调度状态。 TS 仍负责有界 turn-context 投影,Python 仍是既有 planner provider,没有第二套依赖解析或手工同步状态。当前边界是正确的计划与等待诊断、现有测试维护和 CI 日志可观察性,不启动执行或扩张权限。 转交 A 只是换负责人,只有完成 A 才能解除 B 的等待;新的计划也不能替代随后的 quota、claim 和 lease 检查。

具体改动

关键代码讲解

_candidate_readiness 直接复用当前规范谓词,保留 status、resume_when 和 resume_ready。_execution_exclusion 共用于两个 planner;Monitor 继续使用自己的非探索 lane,其他不可执行 Todo 在容量计算之前排除。_build_worker_branch_candidates 不把等待任务并入共同 scope 的工作包。projectExploreTurnContext 只取前三个拒绝项,并从完整 rejection count 派生遗漏数;拒绝项没有领取/租约建议。

基于 docs/reference/canonical-lease-renew.md 在 5e5f0576b15587a9233ce7a2d9571940bf22ad94 的既有规格,todo_done、todo_dependency_pending、resume_ready=true 三项均 implemented:等待条件限制执行,转交保留依赖,前置任务真正完成后重新读回再领取。该文档及 task-graph-projection 契约与上轮基线字节相同。本次覆盖 README、两 planner、TS 投影和相关回归;还检查完整差异中的 CI shard 45→60 分钟作业上限及 verbose 失败名输出、双语验证文档、过期 periodic milestone 测试对齐、真实 mutation locator 修复和归因测试纯格式清理。测试级超时、coverage、worker 数和执行权限均未放宽。三条 Explore 生产模块与旧 head 字节相同,新增验证维护也重新运行。

对主干的风险

独立验证 872 项当前 focused/改动测试和 667 项语义/census 测试通过,TS typecheck、Ruff、diff hygiene 通过。使用隔离的真实 File/SQLite 原生 CLI 复核依赖等待→A 转交后仍等待→A 完成→B 可执行,未改活动 Goal。相同三任务 oracle 在精确基线两个 provider 都失败;带额外 fallback 的合法场景两版都可正确选择其他工作,不能拿它伪造缺陷复现。默认关闭时三个实际入口完整输出在仅核对并归一化 fixture 路径/路径摘要后相同;analysis-only 仍不给命令。缺 readiness、拒绝 acceptance、共同 scope 和诊断遗漏均已覆盖。

语义与 CI 对齐

预检查 advisory 未检测到受支持的新词汇 carrier,独立完整语义检查通过;该观察不能替代契约解释。这里复用既有状态词汇与 typed projection,没有新 authority 或机器义务。原生 premerge 五条直接检查通过,19 项选中检查中 18 项通过,CLI 输出预算失败。相同命令在 5e5f0576b15587a9233ce7a2d9571940bf22ad94 和 d8f2fd85b412b6d81200a2feb2179fbb427e6bfa 的四个 Todo-list JSON 行以相同增长数值失败,故归为 pre_existing_unrelated;这些是未改的 Todo-list reader 路径,Explore 的实际变更不变量有独立通过证据。预算未被放宽,premerge 仍保留红状态;没有查询或等待 CI,也不宣称全套/hosted shard 时长已验收。

我的整体评价

long_horizon 与 user_experience 均 improved:等待任务不再占执行容量,前置任务完成后正常恢复,独立任务仍能继续。完整设计保持 proportionate,未来维护已通过共用 readiness/exclusion helper 改善;没有理由再增加新调度器。额外 mutation locator 测试保护当前真实 mutation runner,格式整理不改变归因规则,整体没有新增状态机制。APPROVE 与单独的 premerge/合并资格 hold 同时成立;全套测试、实际模型采纳、packaged frontend/Lark 和长期运行仍未由本次验证,不能把计划正确当成这些目标完成。

English verdict: APPROVE - d8f2fd8; full current diff reviewed, canonical readiness reused and real File/SQLite lifecycle plus feature-off parity validated. 872 focused and 667 semantic tests pass. Existing CLI budget failures reproduce identically at base/head and remain a separate premerge hold; CI not consulted.

@huangruiteng
huangruiteng merged commit 073be70 into loopx-project:main Oct 9, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants