Skip to content

docs(skills): route effectiveness and trajectory efficiency to benchmark analysis - #6033

Merged
huangruiteng merged 3 commits into
mainfrom
codex/benchmark-analysis-routing
Oct 9, 2026
Merged

huangruiteng merged 3 commits into
mainfrom
codex/benchmark-analysis-routing

Conversation

@loopx-agent

@loopx-agent loopx-agent commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Benchmark score growth and agent trajectory efficiency can be mistaken for process-performance regressions, leading to profiling before the task-level cause is understood. Route those questions to loopx-benchmark, which now distinguishes useful progress, model throughput, control interactions, recovery and evaluator/feedback delay. loopx-performance-diagnosis remains the focused workflow for demonstrated costs in an owned process.

The two packaged skills and their invocation metadata now share this boundary. The introductory scope explicitly includes analysis of active runs, matching the detailed operating lanes. Active analysis remains provisional; hidden evaluator evidence stays unavailable until solver and scoring are terminal. Official feedback is limited to the declared run protocol, avoiding an unconditional ban on legitimately released feedback. No runtime, scoring, scheduling or access mechanism changes.

Validation at 1425603fee4d8aba5d2383975ab07fdbd5aa6ad8: 71 existing packaged-metadata, skill-install and delivery-parity tests passed; both skill-creator validations passed with the checkout interpreter; diff/public-boundary checks passed. Routing scenarios were inspected manually; independent model-routing behavior and task outcome gains are not claimed. The related refactor is confined to the existing skill boundary, with no new capability or framework.

The scoped host skill install was read back against the candidate source; the refined benchmark body and invocation metadata match, and unrelated installed files remain unchanged. This is candidate instruction delivery, not merged-main adoption or measured model behavior.

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Exact head: 56ea6575a2ac22c7f4c8b2f6ad35a702ca969dbf. 无阻塞发现;一个 P2 入口措辞建议见下文。

动机

LoopX 实验操作者分析“得分涨得慢”或“Agent 重复读取 Todo”时,需要先判断哪些工作产生有效进展,再判断成本来自哪里。 以前 benchmark skill 主要列运行管理与结束后分析,performance skill 直接介绍 profiler;现在两边入口都明确将任务效果、得分增长和轨迹效率交给 benchmark 分析,仅将已有成本证据的进程问题交给 profiler。

得分增长曲线变平,可能是任务策略、反馈滞后或模型吞吐变化;多读了几次正文也可能是在履行 freshness 或恢复义务。仅凭 slope、字节或调用数选 profiler、压短上下文,都不足以解释任务为什么没有进展。已核验两份 skill 的实际安装正文和调用元数据,并完成场景指令演练;这证明指引可被交付且语义一致,不证明独立模型已正确选择或实际任务收益提高。 本 PR 不修改运行时、评分、调度或访问机制,不启动实验,不授权读取活动任务的隐藏评测材料。

改动思路

在现有 benchmark playbook 增加有效进展分析路径,在 performance playbook 与双方 invocation metadata 写出同一分工,没有第三份能力或平行分析框架。先对齐原生成功标准、版本、模型、任务、预算、反馈、采样、evaluator 与并发;再从代表性轨迹把实际任务工作、控制读取、恢复、模型吞吐和评测/反馈等待分开。找到有证据的进程成本时,才委托已有 profiling 工作流,并把结果带回任务效果分析。

活动期只读已授权 solver/runtime 观察及已释放的分数 projection,结论保持 provisional;求解结束且评分完成后,才在现有 analyst 权限内分析完整私有证据。原生协议释放的 official response 与隐藏测试/评测实现不是同一权限。正文保留 solver 排除、普通 microbenchmark 排除、安装不授予权限、launch source fence、write preview 和 countable matched comparison;新入口不接管这些 owner。

具体改动

四个文件 +54/-8,全部是 agent-consumed instruction/metadata;不是普通无行为的文档,也没有 runtime 代码改动。benchmark description、调用 prompt 与新增分析段共同负责任务效果/轨迹效率,performance description、开头与调用 prompt 负责已有证据的进程成本。现有七步 profiling 过程仍相同,未增加默认 profiler、预算豁免、后台服务或数据上传。

关键内容讲解

  • loopx-benchmark/SKILL.md 的新分析段要求 common elapsed window、source/model/task/budget/feedback 等可比身份,缺失/无效分数与零分分开,并沿任务 native ranking 选 best;以 early/middle/late、stalls/counterexamples 和分母约束选择性叙事。
  • 新分解段将 strategy、token throughput、control overhead、tool failures、evaluator/feedback delay 分开,把 capture/grading/delivery 与 solver 活动对齐。没有 telemetry 保留 unknown;单独 score slope 或 token rate 不作因果结论。
  • 新 intervention 段要求在既有权限内做有区分力的小实验,改一个因素、披露 confound,优先任务收益和技术深度;重复正文删减先保留 consumer obligations,并要求规划/恢复最终回到有用工作。
  • 两份 agents/openai.yaml 与两份正文一致交付。benchmark invocation prompt 比基线短,但 board-before-selection、preview writes、countability/matched arms 等义务保留在同一次安装的完整正文中;不能只看 metadata 通过就认为模型已理解。
  • Source, integrity, and artifact boundaries 将“求解期一律禁止官方反馈”改成“只允许 declared run protocol 实际释放的 official feedback”。隐藏 tests/verifier/gold 仍禁止,完整私有 analyst 证据仍需 solver terminal AND scoring complete,安装本 skill 仍不授予权限。

接受依据 docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md,不可变 revision ab64c5d1c36a88c587a54ac26cb1cda195d5f965,§6.2–6.5;以及同版 docs/development/testing-and-quality.md 的 Roadmap-Aligned Optimization。原文没有细项编号,以原文标题/短语标识:

  • Native outcomes:implemented,保留任务原生 score/validation 标准,不拿控制调用数替代成功。
  • Efficiency:implemented,同时解释 raw score 与时间/token/工作成本,少调用不自动等于更好任务结果。
  • Long-horizon control quality:implemented,代表轨迹、控制税、恢复/重复和 evidence 使用仍需区分,不能凭文字相似度或 keyword 得到语义结论。
  • Allowed feedback carries only the benchmark-declared response.:implemented,修正无条件禁反馈与 accepted protocol 的冲突,不开放其 backing hidden materials,也不升级 eligibility。
  • Acceptance requires all four feedback/timing combinations:out_of_scope,这个 instruction PR 不实现或宣称四组合 runner/toolkit qualification,既有研究验收继续负责。

独立重复现有 3 文件测试:71 passed,双方 skill-creator 校验通过。额外在 base/head 用实际 workflow-skills CLI、独立临时 root 验证 preview 无写、install、四文件逐字节源一致、idempotent readback、删除 metadata 后 unavailable/stale 检测与安装恢复、最后 uninstall;生命周期一致。未改用户实际 skill 目录。下面语义演练是 reviewer 对已安装指令的人工核验,不是另一模型的路由实验:

场景 应有路径与实际指令结论
活动实验得分斜率下降 benchmark 对齐窗口/预算/反馈/工作内容,候选原因保持可证伪;不直接 profile,不读隐藏材料。
Todo 正文重复读取 benchmark 先判是否有信息/恢复价值;删减前保留 consumer obligations。一般非实验任务不会被强制套 board。
已测量单条 owned CLI 变慢 performance 保留 uninstrumented baseline、真实进程 profiling、intervention 和原语义/延迟复验。
runner 下单题 solver 继续原 task/contract,不从 benchmark 词语获得 operator 角色或 board discovery 义务。
协议已释放官方反馈,但索要隐藏 tests/verifier 可以用已释放 response;仍拒绝隐藏 backing files 与其它 trial。
solver 已结束、scoring 未完成 仍不读完整隐藏 evaluator 证据,不写终态 case insight;授权已释放 projection 可作 provisional 观察。
普通 library microbenchmark、无 Goal/board task normal tools;有实际进程成本才用 performance,不引入 campaign bookkeeping。

对主干的风险

最强风险是 description 扩大后,把单题 solver、活动期 score 或 profiling hotspot 错当 operator 权限/因果证据。完整正文与上述反例保留了界限;“已授权或已释放”不是任意读取新材料的授予。正式释放反馈也不证明四组合资格,缺失 runtime/integrity evidence 不能由新文案补成真。相关 source/fence、countability、隐私、preview 与 quiet/continuation 条款逐项比较,而非只看新段落。

P2 非阻塞措辞建议: benchmark 开头第15行仍用 “run management or post-run analysis” 概括选择条件;正文新增的 active effectiveness/trajectory lane 已明确允许授权活动期分析。建议把开头也写出这一路,避免只读简介时误以为必须等结束。后文边界清楚,所以这不是已观察到的模型误路由,也不阻塞本次完整文本结论。

语义与 CI 对齐

这是本地 playbook 语义变更,没有新 enum、receipt、quota/scheduler 或 typed state owner。frontmatter/YAML 可发现性、真实 install/readback 和四文件 public/private 扫描通过,diff whitespace 通过。diff advisory 无代码源文件,不能证明指导语义正确;已经人工比较 actor、trigger、authority、active/terminal 时间条件、write ordering、结果资格和恢复/continuation。最初错误调用无 __main__ 的 package、system Python 缺 yaml 均保留为验证环境错误,改用真实 entrypoint 和兼容 checkout interpreter 后通过,源码未改。native review policy 的 wait_for_ci=false 下未读取或等待 GitHub CI。

我的整体评价

APPROVE,当前 task 的 instruction delivery 已完成(goal_achieved)。 长程与体验判断均为 improved:从任务真实成功标准出发解释有效进展,避免没有进程成本证据就先安装 profiler,也保留技术 profiler 的正常可达路径。复用现有两份技能与安装 owner 的范围相称,相关 future-facing pass 是明确定义分工与双向回传,没有添加第三套状态/框架;开头 wording 的小建议可后续完善。

这份结论支持“已交付可读的统一路由与权限指引”。实际模型会否选对、采用了多少、是否提升任务结果或节省总成本仍未测,不因 71 项测试或安装 hash 推断。当前用户宿主未升级,未触发真实 benchmark、profiling 或隐藏材料读取;没有把所有 parent study 验收放到这个文案 PR 上。

English verdict: APPROVE - 56ea657; coherent benchmark-effectiveness versus evidenced process-cost routing, with solver/active/terminal and declared-feedback boundaries preserved. 71 tests and paired real CLI install/readback/recovery/uninstall passed. P2: align the opening scope summary with the explicit active-analysis lane; independent model adoption and outcome gains remain untested.

…sis-routing

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Exact head: 1425603fee4d8aba5d2383975ab07fdbd5aa6ad8;base ec21f7da8b9a861fa73477f0ad69e8a4f30e5389。无阻塞发现。前一结论仅适用于 56ea6575a2ac22c7f4c8b2f6ad35a702ca969dbf;本轮独立复核,原 P2 开头范围建议已落实。

动机

LoopX 实验操作者分析“得分涨得慢”或“Agent 重复读取 Todo”时,需要先判断哪些工作产生有效进展,再判断成本来自哪里。 以前 benchmark skill 主要列运行管理与结束后分析,performance skill 直接介绍 profiler;现在两边入口都明确将任务效果、得分增长和轨迹效率交给 benchmark 分析,仅将已有成本证据的进程问题交给 profiler。

已在当前 source 的真实 workflow-skills CLI 独立安装、读回正文及 metadata,并覆盖缺 metadata 后恢复;开头明确包含活动期分析,整份指引语义一致。此为指令交付与人工场景核验,独立模型选择及任务收益未测。 本 PR 不修改运行时、评分、调度或访问机制,不启动实验,不授权读取活动任务的隐藏评测材料。

得分增长慢、重复读取正文或 token 变多,不自动代表进程变慢。先判断有效工作与反馈延迟,再决定是否需要 profiler,才能避免把低调用数当好结果。新版开头明确写出 active-run analysis,读者无需猜测是否一定要等实验终止。

改动思路

任务效果归已有 benchmark playbook,有证据的 owned-process 成本归已有 performance workflow;明确双向分工比新增第三技能或默认 profiler 更贴合目标。 本次边界是现有两份 skill 与 invocation metadata 的一致交付、活动期入口及权限说明;独立模型的实际选择、效果和总成本继续保持未测。

复用 benchmark 的 native outcome/board/权限与 performance 的既有七步测量流程。活动期只用已授权 solver/runtime 观察和协议已释放 score projection,结论 provisional;solver terminal AND scoring complete 后才能在既有 analyst 权限内读完整私有材料。安装或“benchmark”字样不授予 operator、runner、网络、私有数据或付费执行权限。

具体改动

四个现有 instruction/metadata 文件 +58/-11;生产 runtime、schema、quota/scheduler 未改。相对上次 head,仅 benchmark 开头段增加 “analysis of an active run”;全部四路径在两版 base 中无变化,相关 installer/packaging owner 保持。新结论从全 PR 的明确分工重新判断,旧验证保留原 revision,当前 suite 和真实交付路径重新执行。

  • loopx-benchmark/SKILL.md description、开头与新五步分析:对齐 source/model/task/budget/feedback/sampling/evaluator/concurrency及同 elapsed window,missing/invalid 不作零;代表性 early/middle/late/stalls/counterexamples带分母,分开 task work、control reads、恢复和评测/送达等待。
  • 原因判断区分 task strategy、model throughput、control overhead、tool failure、feedback delay;没有 telemetry 保留 unknown。干预在既有 authority 内改一个因素、披露 confound,删重复上下文前保留 consumer obligations,并验证规划/恢复仍回到有用工作。
  • benchmark invocation metadata 与 loopx-performance-diagnosis 正文/metadata 写出同一分工。performance 既有七步 profiler 流程原样保留;缩短调用 prompt 后,完整安装正文仍保留 board-before-selection、preview writes、countable matched comparison 和 source fence。
  • 原求解期一律禁止 official feedback 改为仅允许 declared run protocol 实际释放的 response;hidden tests/verifier/gold/backing files及 unrelated trials仍禁止。新 intro 不放宽这一边界。

接受依据 docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md,不可变 revision ec21f7da8b9a861fa73477f0ad69e8a4f30e5389,§6.1–6.5;同版 docs/development/testing-and-quality.md Roadmap-Aligned Optimization。它们是修改前的 accepted 文本。

  • Native outcomes:implemented;完整安装文本与场景核验支持此指引条款。
  • Efficiency:implemented;完整安装文本与场景核验支持此指引条款。
  • Long-horizon control quality:implemented;完整安装文本与场景核验支持此指引条款。
  • Allowed feedback carries only the benchmark-declared response.:implemented;完整安装文本与场景核验支持此指引条款。
  • Acceptance requires all four feedback/timing combinations:out_of_scope;This is packaged instruction routing; RFC11.3 bounded slice and current task frame do not implement runner qualification, scoring or new feedback permissions. Existing toolkit integrity gap and study acceptance remain unchanged.

独立当前 source 三份测试 71 passed / 5.70s,两份 creator validation 通过。对 current base/new head 各用实际 workflow-skills CLI、独立临时 target 验 preview 无写、install、四文件逐字节源一致、idempotent readback;删除 metadata 后 ready=false,重新 install 恢复原源,再 uninstall。全程没有改用户真实 skill home。

人工重走七个指令场景:活动平分曲线走 provisional benchmark analysis;重复 Todo 读先核信息/恢复义务;已有测量的 owned CLI 慢走 performance;单题 solver 不获取 board/operator 角色;协议 response 可用而隐藏 backing files拒绝;solver结束但评分未完仍不读全隐藏证据;普通无 Goal/board microbenchmark 用原 task 工具。它们是已安装文本的 reviewer 语义判断,不是独立模型路由实验。

对主干的风险

扩大的入口可能被仅看关键词的模型误当角色或权限。完整正文明确保留 solver/ordinary task 排除、authority、active/terminal 时间条件、write preview、source qualification、隐私与 continuation;这只是可交付指引,不能冒充机器 gate 或模型 comprehension。最初“必须 post-run”的开头歧义现已在同一段消除。

全 diff/public-boundary、source/install 读回通过;没有新共享闭合集合或第二 decision owner。语义检查是完整条款比较:actor、trigger、order、授权、scope、evidence、continuation逐项保留,并披露官方反馈条款的有意修正。按当前 wait_for_ci=false 未读取、轮询或等待 GitHub CI。

我的整体评价

APPROVE,goal_achieved 仅指本次 instruction delivery。 长程与体验为 improved 的指引判断:先解释有效工作、再定位实际进程成本,避免无依据 profiling;活动入口与后文一致,保持明确恢复和权限边界。任务效果归已有 benchmark playbook,有证据的 owned-process 成本归已有 performance workflow;明确双向分工比新增第三技能或默认 profiler 更贴合目标。 本次边界是现有两份 skill 与 invocation metadata 的一致交付、活动期入口及权限说明;独立模型的实际选择、效果和总成本继续保持未测。

future-facing pass 已在原开头段完成,没有增加模块、状态或自动操作。实际独立模型选对率、采用率、得分与总成本仍未测;作者声明的 scoped 本机安装不替代我的隔离交付证据,也不等于 merged-main/全研究资格。没有合并、全局升级或启动实验。

English verdict: APPROVE - 1425603; coherent task-effectiveness versus evidenced process-cost routing with the prior active-analysis opening ambiguity resolved. 71 tests and paired actual CLI install/readback/failure-recovery/uninstall passed; solver, declared-feedback and terminal/scoring boundaries retained. Independent model adoption and outcome/cost gains remain untested.

@huangruiteng
huangruiteng merged commit 0d1c9eb into main Oct 9, 2026
3 of 5 checks passed
@huangruiteng
huangruiteng deleted the codex/benchmark-analysis-routing branch October 9, 2026 16:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants