Skip to content

fix(control-plane): resume unstarted cadence reservations - #5966

Open
Duang777 wants to merge 16 commits into
loopx-project:mainfrom
Duang777:codex/fix-cadence-reservation-resume
Open

Duang777 wants to merge 16 commits into
loopx-project:mainfrom
Duang777:codex/fix-cadence-reservation-resume

Conversation

@Duang777

@Duang777 Duang777 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Goal And Delivered Outcome

  • Outcome basis / optional anchor: Reproduced managed cadence crash-recovery stall.
  • Goal/source and gap: A matching durable reserved record still passed through the minimum-interval check, although that reservation had advanced its own eligibility time.
  • Observable before → after, with the validation row that proves it: Before this change, a process crash before confirmAutomationStart blocked the same request until the interval elapsed. The same unstarted reservation now resumes after a 1 ms restart, while a different request still receives minimum_interval_wait and a confirmed start remains closed.
  • Issue/task and intended base: Self-contained bug fix against loopx-project/loopx:main.

Author Declaration

  • Written by: OpenAI model agent, directed by a human operator.

Implemented against

  • Specification and revision: No written specification; the request in this PR is the basis.
  • Criteria:
Criterion (spec clause) Disposition Symbol / path Test or command
The same unstarted reservation resumes immediately implemented admitAutomationStart automation_cadence.test.ts
A different request still obeys the interval implemented admitAutomationStart automation_cadence.test.ts
Crashes before the first journal or attempt record recover implemented managed turn executor Two focused test_loopx_turn_executor.py cases
Resume keeps the original interval anchor implemented held reservation branch Existing cadence state assertions
  • Self-check before submission: Reviewed admission, durable reservation, journal, host mutation, and confirmation ordering. Ran focused TypeScript and Python tests, TypeScript typecheck, and diff-driven premerge. This PR does not change cadence persistence schemas or confirmed-start behavior.

Scope And Continuation

  • Completed scope and remaining work: Complete within unstarted cadence reservation recovery.
  • Slice boundary / successor: N/A for this defect; the change reuses the existing request identity and durable state machine.

Validation

  • Tested revision: c1dea898b62f5d90233d4b87b68ac5e3cb4b37fa
  • Run state: finished
  • Input classes: synthetic
Check kind Result Public-safe evidence / limitation
unit passed node --no-warnings --experimental-sqlite --experimental-strip-types --test tests/control_plane_ts/automation_cadence.test.ts: 7 passed.
integration passed Two managed executor crash-recovery tests passed with a 1 ms restart.
static passed npm run -s typecheck:control-plane.
real_entrypoint passed loopx canary premerge --from-git-diff --git-diff-base upstream/main: 16 selected checks passed with no manual hold.
real_backend not_applicable Cadence persistence uses the existing file-backed contract; no external provider path changed.
  • Coverage and gaps: The tests cover unconfirmed resume before both journal boundaries, a competing request during cooldown, stale triggers, and confirmed starts. No persistence migration is required.

Frontend / Visual Evidence

  • UI impact: none
  • Before: N/A
  • After: N/A
  • States and viewports shown: N/A
  • Source data: none
  • Attention review: N/A

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Refactoring (no functional changes)
  • Documentation update
  • Test update

LoopX Area

  • Control plane (goals, todos, quota, scheduler, registry, runtime)
  • Benchmark boundary (adapters, runners, verifiers, scoring, evidence)
  • Capability or extension (providers, adapters, skills)
  • Public docs or presentation surface (README, protocols, dashboard)
  • Build, packaging, installer, or CI
  • Host or runtime integration

Technical Direction

  • Direction / acceptance reference, when applicable: Core control-plane hardening.

Shared-authority RFC fixture impact

N/A. This PR changes the existing cadence state transition, not shared-authority routing or schemas.

Boundary Checklist

  • Neither the diff nor this PR body/comments/attachments disclose private state, credentials, raw traces or verifier output, internal links, or local machine paths (including .loopx/, .codex/goals/, and live ACTIVE_GOAL_STATE.md).
  • I did not duplicate maintainer-owned benchmark work unless a maintainer split out a public issue for it.
  • I kept the change scoped to the linked issue/task.
  • I completed the visual evidence section for UI changes, or marked UI impact none.
  • Every commit includes a DCO Signed-off-by trailer (git commit -s).

Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>
Signed-off-by: Duang777 <duangjl007@gmail.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh

动机

运行受调度间隔约束的 LoopX Turn 的操作者,需要让尚未启动宿主的崩溃任务继续执行。 进程在写入调度预留后、宿主尝试落盘前退出时,旧版重启会被自己的间隔阻挡;新版以原请求身份立即恢复,同时保持原间隔锚点。 同一套真实文件存储与执行器反例在基线两次等待,在候选版两次于 1ms 后恢复;宿主、写回、扣额度与调度各执行一次,重放没有重复效果。 本 PR 不降低新请求的间隔,不扩大手动、租约、额度或宿主权限,也不宣称 App heartbeat 和非 Turn 启动路径已完成验收。

改动思路

该修复复用现有 TypeScript 间隔 owner,并由既有宿主尝试日志与 lane 锁承担执行隔离;不新增恢复开关或平行判断。 当前 PR 只修复未启动的原请求恢复,App hook、非 Turn 启动器和完整设置界面的后续验收仍归既有 RFC owner。 一次预留先占住原间隔,只有宿主尝试日志落盘后才确认为已启动;恢复原预留不会领取第二个间隔。既有 kernel lane 锁在 admission 前包住完整执行,journal 与确认在 host 前落盘。另一请求、已确认请求和更新触发事件仍分别经过原拒绝规则。

具体改动

发布前发现 head 变化,旧草稿未发布。完整当前 PR diff 和受影响 owner 已重新核验,下述实际入口、基线对照、负例、focused suite 与 premerge 均在这里列出的新 B/H 上重跑。

精确 head 9f0028d43622ddf38516169aebaeb16c1ebc7e83,真实 merge base 04accc4b40990ed4db29dd0c97c8c3a6519bda3b。全 diff 三文件6 additions/8 deletions:automation_cadence.ts 的 held-reserved 分支直接返回 resumed;TS测试检查1ms恢复仍保持1000的锚点;两处 Python crash测试把时钟从120秒后收紧到1ms后。

规范是 accepted automatic admission RFC,spec_revision=04accc4b40990ed4db29dd0c97c8c3a6519bda3b。§§2/5/7; M2 的 durable floor、尝试确认、并发隔离和 replay边界得到验证。作者“无书面规范”的声明不能替代这份契约。Appendix: two-phase implementation ledger 的历史 ledger 有一条需同步,见下文;M2/M3 host promotion 的M3设置、App hook及非Turn路径没有在本PR被激活或宣称完成。

关键符号:admitAutomationStart(:121)先验证输入和 stale/started,再允许同一reserved identity;managed_cadence_start(:22)仍仅适配 turn_key:attempt 到同一TS owner;run_loopx_turn_once(:1316)先取得lane/journal锁,写attempt再confirm,随后才调用host。新分支没有重写started_at或降低effective floor。

对主干的风险

最强反例是立即恢复导致重叠宿主、第二次扣额度,或偷偷前移间隔。独立同一脚本在不可变B/H分别中断“预留后、首journal前”与“attempt落盘前”:B两次 interval_wait、四类效果各0;H两次committed、四类效果各1,原锚点不变,再次replay不增加效果。真实kernel lane被占用时,两版都在host/额度之前返回 turn_lane_in_flight。宿主结果/validator及写回效果用合成回调,TS文件存储、生产executor与journal是真实入口;没有调用模型或改动活跃Goal。

同一固定RPC矩阵在两版的完整输出相等:无policy时不建文件、不同请求等待、confirmed即使manual也拒绝、future trigger错误及不同agent不互锁。新B83个executor测试、新H83个executor与10个cadence/lane测试、7个TS测试通过;H的、TS typecheck、premerge16 selected加direct checks通过。没有查询或等待CI。首个premerge误用不存在的参数,未执行检查;修正为现有 --git-diff-base 后完成验证,未把入口错误算成PR缺陷。

[P2] 修正现有 RFC Appendix 的恢复时机。 ledger仍写“once the floor is reached”,而本次实际契约是立即恢复同一未启动预留。PR body已披露新行为,独立证据也证明floor锚点及confirmed保守边界保留;但当前canonical契约仍给出相反恢复条件,可能让调用者等待或求manual bypass。请在原句说明 immediate same-identity recovery、anchor不变和confirmed fail-closed。无需新增配置或另一份规范。

未新增Enum、共享状态或协议字段;typed reserved/started owner被复用,admission是机器约束。advisory未发现新vocabulary carrier,不能因此代替上面的语义对照。默认未配置路径、存储兼容、错误precedence和effect replay保持。整体App调度准出仍未测。

我的整体评价

REQUEST_CHANGES。代码的实际未启动恢复缺口已修复;当前还需让同一owner的既有契约准确说明新恢复时机,这是一处有界文档修正,不要求补齐App或整个RFC。future-facing pass考虑了interval与journal归属:它们承担不同权威,已有共享TS owner足够,无需再拆builder或Python规则。请修正现有ledger后以新准确head读回;无需增加配置或额外机制。control-plane合并仍由maintainer处理。

English verdict: REQUEST_CHANGES — 9f0028d; same unstarted reservations resume at1ms without shifting the floor or duplicating host/settlement effects. Immutable-base oracle fails as expected; head93 executor/cadence/lane tests,7 TS cases,typecheck and risk-based premerge pass. The canonical RFC still says recovery waits for the floor, contradicting this new default; correct that bounded contract sentence. CI was not consulted.

Signed-off-by: Duang777 <duangjl007@gmail.com>
@Duang777

Duang777 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the RFC mismatch in 949b298: the contract now states that the same unstarted reservation resumes immediately, keeps the original interval anchor, and remains fail-closed after confirmation. Revalidated 83 Python executor tests, 7 cadence TypeScript tests, control-plane typecheck, and diff checks.

@Duang777
Duang777 requested a review from loopx-agent October 9, 2026 14:37
…eservation-resume

Signed-off-by: Duang777 <duangjl007@gmail.com>
@Duang777

Duang777 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Synced the reviewed fix to main@ec21f7da8 in 641855d9c473af2b99b3dbfdd0d299873084e6b0. The RFC correction remains unchanged. Revalidation passed: 83 Python executor tests, 7 cadence TypeScript tests, control-plane typecheck, 16 risk-selected premerge checks, the public-boundary check, and all direct checks.

@Duang777

Duang777 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

CI triage update for Python Tests run 37954626954: the failures inspected so far are baseline drift, not changes from this PR. The three failures in test-shard (2)/(4) are stale periodic-report/Todo completion assertions already corrected on main by 28a18f6; the two current focused files pass 8 tests on main. The Stage 2C mutant failure is fixed by merged #6036, and the remaining installed-source provenance failure is tracked in #6040. test-shard (3) is still running. I am holding branch sync and CI reruns until that run completes and #6040 lands; this PR branch remains unchanged.

@Duang777

Duang777 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Final update for Python Tests run 37954626954: test-shard (3) reached 93% without logging a test failure, then hit the 45-minute job timeout configured at this head. Its only error is ##[error]The operation was canceled. This adds no PR-specific failure. The actionable failures remain the baseline issues reported above; current main now contains the #5986 and #6036 fixes, while #6040 is still open. I am not rerunning CI or syncing this branch until that remaining baseline fix lands.

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh

动机

受最短调度间隔约束的 managed Turn 操作者,需要恢复尚未真正启动宿主的中断任务。 进程在预留调度间隔后、宿主尝试记录落盘前退出,旧版重启被自己的间隔阻挡;新版以同一身份立即恢复,保留原间隔锚点。 实际两处中断对照中,基线均等待,新版在 1 毫秒后继续;宿主、写回、扣额度与调度各一次,重放不增加效果。 本 PR 不缩短新请求的间隔,不扩大权限,也不宣称 App 定时器、外部模型和非 Turn 启动器已完成资格验证。

本批只修复同一未启动预留的恢复;当前还需把既有中文语义镜像同步为立即恢复、原锚点不变、已确认启动拒绝。

改动思路

同一 TypeScript cadence owner、既有文件锁及 Turn journal 已足够;无需新增恢复开关或 Python 规则。 本批只修复同一未启动预留的恢复;当前还需把既有中文语义镜像同步为立即恢复、原锚点不变、已确认启动拒绝。 准入先持久预留,生产执行器在现有 lane 锁内落盘宿主尝试,再确认预留,最后调用宿主。恢复复用原请求,不领取第二个间隔;真正的新请求仍受原下限约束。

具体改动

完整 exact head 641855d9c473af2b99b3dbfdd0d299873084e6b0,不可变 merge base ec21f7da8b9a861fa73477f0ad69e8a4f30e5389;四文件 +12/-13。admitAutomationStart 的 held/reserved 分支直接返回 resumed;TS 用例从等满一小时改为1ms恢复,并断言另一个请求仍等待;Python 的两个真实持久化中断用例收紧到1ms;英文 RFC Appendix 改为立即恢复并保留原锚点。没有新状态、配置、持久化格式或平行决策 owner。

规范:accepted RFC,spec_revision=ec21f7da8b9a861fa73477f0ad69e8a4f30e5389。逐项映射:§2 重启/重试不丢约束;§5 预留、确认和起始锚点;§7 默认关闭与权限隔离;§9 真实临时文件/重启拒绝,当前路径均有证据。Language mirror 要求仍未满足:英文首段明确把中文版列为语义镜像,中文版附录仍保留相反的恢复时间条件。M3设置、App hook和非Turn宿主推广继续归既有RFC owner,本批未宣称完成。

对主干的风险

最强反例是立即恢复造成宿主重叠、重复扣额度或把间隔起点前移。我独立使用相同临时文件/生产 executor 与 TypeScript store,在“预留后、首journal前”和“journal存在、attempt前”两处中断:B均interval_wait、四类效果均0;H均committed、四类效果各1,原started_at不变;再次replay仍各1。宿主返回、validator和结算回调用合成结果,持久化/准入/journal及执行器是真实代码;没有调用模型或使用活动Goal。

固定输入的另10行B/H完整输出相同:无policy不建文件,新请求等待,同Goal不同automation不能绕过Goal下限,未来/陈旧trigger拒绝,confirmed即使带manual理由也拒绝,不同Agent独立,缺失phase按started拒绝,精确floor新请求允许。93项Python、7项TS、typecheck,以及17/17原生风险检查与4项direct检查通过。未查询CI。最初私有探针误写confirm方法名,纠正为现有confirm_start后完成两臂验证;该探针错误保留,不归因PR。

[P2] 同步中文恢复契约。 当前head的中文附录第175–176行仍写“同一 Turn 身份在满足时间下限后仍可恢复”,而英文和实际实现已改为立即恢复。请只在同一段同步“立即恢复、原间隔锚点不变”,同时保留已确认和缺phase旧记录fail-closed。这样调用者无需在两种语言中选择不同操作条件;不需要新增机制或额外审批。

我的整体评价

REQUEST_CHANGES,剩余阻断是上述有界中文伴随修正。运行时恢复本身已通过独立反例:减少启动前中断后的等待,未降低新请求约束,也未增加重复效果。可预期长程恢复与效率方向正向;模型成本、App定时器到hook覆盖和长期真实宿主收益仍未测。

future-facing pass 已检查 cadence/journal/lane 的相邻边界:分别承担时间、尝试事实和并发隔离,现有共享owner足够,没有值得在本批再拆出的框架。经验提示仅用于核对“恢复后继续有用工作”的证据,不继承旧结论,也不计为记忆净效用。运行时合并由维护者处理。

English verdict: REQUEST_CHANGES — 641855d; the immediate same-reservation recovery is independently verified at two journal boundaries with an unchanged anchor and exactly-once effects. 93 Python,7 TS,typecheck,10 identical base/head negative rows and17 native checks pass. The accepted Chinese semantic mirror still says to wait for the floor; synchronize that bounded paragraph. CI was not consulted.

@BigDataDZ BigDataDZ left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent verification on native Windows 10 (zh-CN, cp936), Python 3.12, Node 22.23.3, head 641855d9c applied onto 647e216 (94 commits before the merge base; no overlap in the touched files).

The fix works through the real stack on Windows. With the patched automation_cadence.ts, test_reserved_managed_start_recovers_after_death_before_the_attempt_record passes: the death-before-attempt-record reservation resumes at +1 ms of simulated time with the original anchor, and the different-request rejection still holds. Failing-before reproduced: with the base cadence owner both tightened cases fail (the reserved start waits out its own interval).

Windows-specific finding in the first case — test brittleness, not a logic defect. test_reserved_managed_start_recovers_after_death_before_the_first_journal_write fails at its "no journal written" assertion (line 1321) because glob("*.json") matches the cross-runtime lock holder sidecar <journal>.json.lock.holder.json left behind by the simulated death. After a real crash that sidecar legitimately persists, so the glob will bite anywhere the holder survives; suggesting the assertion filter out .lock.holder.json names (or that holder sidecars not end in .json). With that filtered, the assertion's actual intent - no Turn journal written - holds.

The remaining zh-CN mirror gap, precisely located: docs/architecture/rfcs/automatic-execution-admission-v0.zh-CN.md, appendix "实现记录" paragraph (lines 174-178) still reads 「两步之间进程退出时,同一 Turn 身份在满足时间下限后仍可恢复」 - the opposite recovery-time condition. The English appendix at this head says a crash between the two phases leaves the reservation immediately resumable by the same identity without moving the original interval anchor, with confirmed starts fail-closed and phase-less records read as attempted. Syncing those two sentences (plus the 「缺少阶段字段」 fail-closed reading, which the zh-CN text already carries) closes the last review requirement.

Everything else in the second review round I can corroborate: no new state, config, or parallel decision owner; the Python-side behavior change flows entirely through the existing TS cadence owner served by the Effect runtime.

@Duang777

Copy link
Copy Markdown
Collaborator Author

Confirmed both review findings. I prepared the bounded follow-up locally: the Chinese mirror now states that the same Turn identity resumes immediately without moving the original interval anchor, while confirmed and phase-less records remain fail-closed. The no-journal assertion now excludes legitimate .lock.holder.json sidecars left by a crash.

Revalidation passed for both Python crash-boundary cases, all 7 TypeScript cadence cases, and diff checks. I am holding the push until #6068 fixes the canonical Todo baseline and #6040 lands; then I will merge current main once and push, instead of starting another CI run with known baseline failures.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants