Skip to content

test(ark): read canonical todo inventory - #6041

Merged
huangruiteng merged 1 commit into
mainfrom
codex/fix-ark-continuity-todo-readback
Oct 9, 2026
Merged

huangruiteng merged 1 commit into
mainfrom
codex/fix-ark-continuity-todo-readback

Conversation

@Duang777

@Duang777 Duang777 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

Goal And Delivered Outcome

Author Declaration

  • Written by: model_agent, OpenAI Codex.

Implemented against

  • Specification and revision: no separate written specification; the JSON inventory contract introduced by a40161adf and its existing contract test are the basis.
  • Criteria:
Criterion (spec clause) Disposition Symbol / path Test or command
Canonical records remain complete in todos; duplicate lane projections may use response-local references. implemented test_fresh_host_reconstructs_frontier_from_durable_loopx_state uv run --extra test pytest -q tests/test_ark_managed_agent_host.py
  • Self-check before submission: read the Todo list projection and reference contract, reproduced the base failure before this change, then ran the affected test under Python 3.11 and the complete focused files under Python 3.13. Production code and response shape are unchanged.

Scope And Continuation

  • Completed scope and remaining work: updates one stale Ark assertion. Complete within this scope.
  • Slice boundary / successor: N/A; the existing Todo reference tests own the response contract.

Validation

  • Tested revision: f12258a621378b08642a632889de50a49ceb449d
  • Run state: finished
  • Input classes: synthetic
Check kind Result Public-safe evidence / limitation
regression_parity passed The base, #6030 head, and current main fail at the stale lane-record read; the changed test passes. The public #6030 Optional Ark run records the original failure.
real_entrypoint passed Python 3.11 exact continuity test, 1 passed; the test invokes fresh CLI processes for write and readback.
unit passed uv run --extra test pytest -q tests/test_ark_managed_agent_host.py, 13 passed.
unit passed uv run --extra test pytest -q tests/control_plane/test_todo_list_record_references.py, 4 passed.
static passed Ruff on the changed file and git diff --check.
  • Coverage and gaps: the changed assertion, its real CLI path, and the dedicated JSON reference contract pass. No production path changed.

Frontend / Visual Evidence

  • UI impact: none
  • Before: N/A
  • After: N/A
  • States and viewports shown: N/A
  • Source data: none
  • Attention review: N/A

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Refactoring (no functional changes)
  • Documentation update
  • Test update

LoopX Area

  • Control plane (goals, todos, quota, scheduler, registry, runtime)
  • Benchmark boundary (adapters, runners, verifiers, scoring, evidence)
  • Capability or extension (providers, adapters, skills)
  • Public docs or presentation surface (README, protocols, dashboard)
  • Build, packaging, installer, or CI
  • Host or runtime integration

Technical Direction

  • Direction / acceptance reference, when applicable: N/A

Shared-authority RFC fixture impact

  • Production-scale fixture schema: N/A
  • Semantic dimensions changed, or reviewed no-impact rationale: no production or fixture schema change.
  • Provider conformance arms run: N/A
  • Read-only legacy/file/PostgreSQL three-arm rehearsal (required for promotion, runtime-routing, or compatibility-projection changes): N/A

Boundary Checklist

  • Neither the diff nor this PR body/comments/attachments disclose private state, credentials, raw traces or verifier output, internal links, or local machine paths (including .loopx/, .codex/goals/, and live ACTIVE_GOAL_STATE.md).
  • I did not duplicate maintainer-owned benchmark work unless a maintainer split out a public issue for it.
  • I kept the change scoped to the linked issue/task.
  • I completed the visual evidence section for UI changes, or marked UI impact none.
  • Every commit includes a DCO Signed-off-by trailer.

Signed-off-by: Duang777 <duangjl007@gmail.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: actor_kind=model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.

精确 head f12258a621378b08642a632889de50a49ceb449d,基线 4bdcb2eeed64582b99cfc46b213ed755839a69b8;本结论覆盖全量 PR 差异。

动机

维护 Ark 托管宿主连续性测试、验证替换进程能否恢复已保存任务进度的贡献者。 原测试把任务的分组视图当完整记录读取,而该视图已使用同一响应内的引用,因而在读取任务标识时报错;新测试从顶层完整任务清单选取原任务,继续验证保存的进度说明。 同一连续性用例在不可变 base 报 KeyError: todo_id,在精确 head 通过;包含真实新进程 CLI 与三种状态源引用负例的17项相关测试全部通过。 本次只修复一个现有测试的读取入口,不改变 Todo 格式、连续性准入、Ark 运行代码或断言,也没有启动真实云端 Agent 或执行线上状态迁移。

改动思路

不改会使已有连续性验收停在记录形状错误;新增兼容解引用或恢复旧重复记录都会扩张边界。直接读取既有 canonical todos,并保留原任务过滤和持久 note 断言,是完整的单行测试维修。 生产 TS 投影是现有唯一记录形状 owner;测试使用其已保留的完整清单,不新增解引用适配器或恢复重复记录。真正的验收仍是替换进程选择同一任务并读取已保存 note,而非仅不报错。

具体改动

关键内容讲解

完整变更只有 tests/test_ark_managed_agent_host.py 第505行,+1/-1:迭代入口从 todo_readback["agent_todos"]["items"] 改成 todo_readback["todos"]。没有独立书面规范需要映射;判断依据是修改前已有的“新宿主从持久状态恢复 frontier”测试及当前 canonical list 契约,先读这些原有消费者和引用验证,再看修复。

生产 owner loopx/control_plane/todos/context_projection.ts::projectTodoListPayload 保留顶层 todos,把角色视图中相同完整记录压成响应内 $ref。旧测试直接取引用的 todo_id 会失败;新入口仍通过原任务 id 过滤完整记录。前后两次 should_run、选中的任务身份以及保存 note 的最终断言没有删改;真实 CLI 子进程和隔离 fixture 也未替换。另有 test_todo_list_record_references.py 对实际 legacy/File/SQLite 清单、scoped/limit、exact/thin 与源失败路径作独立覆盖。

对主干的风险

独立相同用例在 base 4bdcb2eeed64582b99cfc46b213ed755839a69b8 失败:KeyError: todo_id 出现在旧角色视图的 item 访问;它不是宿主持久状态丢失的证据,也不是本 PR 的新回归。在精确 head 运行 uv run --extra test python -m pytest -q tests/test_ark_managed_agent_host.py tests/control_plane/test_todo_list_record_references.py,17项全部通过,其中包括原用例。测试通过真正的新 Python 进程写入和读取隔离状态,未使用生产 Goal 进行测试。

最强反例是通过删掉 note/identity 断言、吞掉异常或制造空记录让测试变绿;完整单行差异排除了这些捷径,独立引用 suite 的 source/projection 失败路径也通过。所选解释器与 loopx.__file__ 都来自精确 head checkout,diff check 通过。没有新增运行行为、状态词汇、decoder 或部署依赖;未查询或等待 CI,也未启动真实 Ark 云端执行。 当前失败、未测与实际通过按各自范围保留;不根据远端 CI、作者声明或本地元数据推导证据。公开数据、私有材料、权限语义与 guidance/obligation 均按实际变更边界检查。

我的整体评价

APPROVE,无阻塞发现。long_horizon 保持生产行为,既有测试再次验证替换进程后的持久进度;user_experience 改善贡献者诊断与验收路径,避免把展示引用当宿主故障。未来维护检查选择复用当前 canonical inventory,无须额外抽象或生产兼容分支。相同作者当前三个 open PR 的范围为连续性测试、备份 provenance 与 cadence reservation,本项没有新增重复 smoke,原有覆盖和断言保留。批准覆盖该测试维修,不把17项成功等同真实云端宿主、模型采用、整个产品验收或合并授权。

English verdict: APPROVE - head f12258a. The one-line existing-test repair reads canonical todos instead of a compressed role-view reference, preserving all original task identity, admission and durable-note assertions. The same test fails at the immutable base with KeyError: todo_id; the exact-head Ark-host and real legacy/File/SQLite reference suite passes17 tests. Production serialization and host behavior are unchanged. No duplicate smoke, new decoder, cloud Agent execution or CI query is involved. CI was not consulted. This approval records the verified test-only repair, not merge or cloud-provider authority.

@huangruiteng
huangruiteng merged commit f552398 into main Oct 9, 2026
19 of 23 checks passed
@huangruiteng
huangruiteng deleted the codex/fix-ark-continuity-todo-readback branch October 9, 2026 20:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants