Repository navigation
feat(benchmark): return complete official feedback on best-only improvements - #6031
Conversation
…ements Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
… Codex Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Exact head: aaa724aebc49d534c89203a3054ed090f56e70a6. 无阻塞发现。下面的 P2 是基线已有问题,建议继续修复;不撤销本次通过结论。
动机
运行已声明 best-only 的求解 Agent 与实验维护者,在自动评测产生严格改善时需要知道改善的具体依据。
best-only 指只在原生选择规则判定提交严格改善时通知求解者。以前只能知道某个提交快照胜出,无法从反馈中看到官方分数和诊断;现在同一胜出提交的完整官方结果随来源快照进入下一次模型请求。 缺少正文会让求解者收到“变好了”的消息,却仍不知道哪项诊断值得采用。已实测官方正文进入真实 Codex 的下一次请求及 resume,且保留原工具输出;这证明信息可见,不证明模型已正确采用。
本 PR 不启动或停止实验,不改变评分,不证明分数或效率提升。 新 epoch 的安装、完整 formatter、任务镜像与各 worker 实际采用,以及 matched study 的效果和成本仍须单独验收。
改动思路
保留原生 select_best 作为唯一排序 owner,在既有 SForge provider 和共同部署的 reader 上扩展官方正文传输,比开放求解者评测权限或重建评分体系更贴合目标。
控制器从原生 API 取得正文,核对原任务、提交、分数、通过率、评测进程 epoch 与 evaluator 后,再验证提交归档及普通 worker 可读的副本。成功发布之后才更新 incumbent;失败保留原状态重试。首次有效 incumbent、平分、无效和未准入快照仍不通知,不增加评测机会。本次可回滚边界是胜出结果的来源校验、发布、实际宿主上下文输送和协议元数据;现有组的保存、停止、排空与新组准入属于后续独立操作。
这是一条 Python 原生 provider 传输链路,复用了既有采样器、publisher、SForge hook 与 report exporter;没有另建通用控制面或评分决策源。当前 producer/reader 一起部署,使用同一版本常量;持久化旧报告则继续保留原字段缺省,避免兼容代码与历史真相混为一谈。
具体改动
完整 diff 为 15 个文件、+418/-61:8 个生产文件 +134/-18,6 个测试文件 +237/-36,README +47/-7。除了正文传输,run/profile 与导出记录实际新尝试的 feedback_payload;缺省旧记录不补值,冲突/未知值拒绝。README 披露新 best-only 的全文行为和新 epoch 边界。未新增依赖、CLI 参数、全局能力开关或闲置模块。
关键代码讲解
feedback.py:198 BestFeedbackPublisher._official_result:读取所选 submission 的完整官方 JSON,核对 task/submission/score/pass-rate,再次读取 admission 核对原 epoch/evaluator。update继续负责 source archive/copy 哈希和发布顺序;错误不推进 incumbent。online_sampling.py:51 OnlineSampler.qualify:实际服务返回的 task/image key/selection/direction 与 task-file/module 摘要进入原生准入校验,未改变 capture 或评测容量规则。image key 不等于镜像内容不可变证明。feedback_hook.py:21 deliver:校验 v2、来源归档、evaluator 与正文摘要,完整 JSON 进入 additionalContext。未知嵌套字段、Unicode、null/bool 保留;原工具输出与 Stop 不被覆盖。把正文标为 data 是指导,不能证明模型不受其中指令影响。export_report.py:167 export:只有实际源数据包含新 payload 时才投影,旧 native/blind/best-only 报告保持字节一致。没有据新代码猜测旧运行采用新协议。sforge.py:160 SForgeWorker._install_worker:仍只向 best-only 普通 worker 安装当前 hook,沿原宿主入口校验可读性;真实 Codex 下一次请求和 resume 已复验。
接受依据是 docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md,不可变 revision de3de98d14d3c6f735c156e49f01c491f0d355f8,§6.5 及 §11.3。原文没有条款编号,下面以原文短语识别:
- Allowed feedback carries only the benchmark-declared response.:本次传输实际官方 API JSON,绑定原 snapshot;不授予隐藏测试、参考答案、评测实现、凭据或其它 trial 的访问。native API 本身省略/限制的内容不由本 PR 重建。
- Blind background evaluation:本次未扩大 blind 入口权限;实际隔离与恢复用例和 base/head prompt/report 对照覆盖相关路径。
- Feedback availability is the only changed factor:任务、原生选择、预算、采样节奏与原 source pins 保留。改变的是新 best-only 明确声明的反馈正文;不把不同反馈协议混入同一历史组。
- Diagnostic observations retain their original qualification status:旧记录未重标,没有将缺失 integrity 证据升级成 eligible。
- Acceptance requires all four feedback/timing combinations:deferred。本 PR 交付可用 provider seam,未交付 toolkit 四组合资格判定;既有 benchmark-toolkit typed integrity policy / E3 matched-study 验收继续负责,不以本次 green tests 代替。
独立实测在确切 head 与固定 native source 运行 7 个文件:318 passed,零 skip。含真实 native judge、Docker、普通 worker、Codex hook、下一次 synthetic Responses 请求及 resume;评测服务网络和宿主秘密文件对 worker 的不可见路径已测。任务/grader/Responses 是合成夹具,故这不是付费模型的理解或真实题目效果证明。额外 base/head probe 使用实际 exporter 与文件锁:blind prompt 和 native、blind、历史 best-only 全部 5 个导出文件的字节摘要相同,head 的官方正文未知字段完整保留。
对主干的风险
最强新风险是把错误提交或新 epoch 的报告当作原提交反馈;HTTP、task/submission/score/pass-rate/epoch/evaluator 与 archive 的负例已覆盖,失败先拒绝,后续对原状态重试成功。全文也增加上下文成本,每次胜出增加两次控制器 GET;它不增加评测提交/slots,超时有界,但正文 token 数和模型净收益未量化。hash 证明相等,不单独证明外部真实性、镜像内容、模型采用或实验资格。
P2 非阻塞建议,基线已有:混合代次 reader 可倒退 cursor。 feedback_hook.py:63 之前已在锁外读取 latest。独立强制交错:旧 reader 先读 auto-2 并在锁前暂停,新 reader 读 auto-3、交付并记 cursor,随后旧 reader 获锁把 cursor 改回 auto-2;下一次 auto-3 又被交付。immutable base 和当前 head 都复现同一顺序及错误,新增 checksum 没有改变这个因果路径。现有同包并发用例通过不能排除它,全文会放大重复上下文成本。建议把 latest 的读取/校验纳入同一 session 锁,或锁内重新读取并丢弃旧候选;加入上述混合代次真实 file/flock 交错回归,验证 cursor 不倒退、latest 不重复。
语义与 CI 对齐
改动是 provider 本地 wire v2 与明确声明的反馈正文,不是新增共享控制面词汇。逐项比对了任务/actor、异步 source、禁止覆盖 live workspace、继续本地工作与非完成提示;native/blind 路径保持,没把 guidance 宣称为机器保证。semantic full-tree 检查通过,覆盖 26/26 vocabulary 和 51/51 owner;第一次因缺 TypeScript 开发依赖退出,准备依赖后同一源码检查通过,失败记录保留。diff advisory 排除了 benchmark,扫描为零,不能拿它证明语义无变化;局部协议及文本已人工对照。compile、diff whitespace、public/private 边界检查通过。按本次原生 review policy 的 wait_for_ci=false,未读取、轮询或等待 GitHub CI,结论依据上述本机验证。
我的整体评价
APPROVE;delivery judgment 为 justified_increment。 用户体验为 improved:普通 worker 通过现有请求看到真正有用的官方诊断,无须新参数、轮询工具或额外确认。长程效果为 accepted_tradeoff:严格胜出过滤、原生排序、来源校验、失败重试、历史协议和恢复链路保留;代价是更多正文 token 与胜出时控制器读请求。公开记录了已有 cursor race,不宣称所有并发恰好一次。
架构保留既有 native/provider owner,生产增量与真实调用相称;未为共同部署的 transient wire 保留多套 decoder,也没有迁移历史报告。相关 future-facing pass 复用了同一版本常量并避免第二排序源;锁内 latest 读取建议属于小范围后续修复。完整 formatter、task-image 与每个实际 worker 的安装/readback,以及四组合资格和 matched-study 分数/成本,由既有操作及研究验收继续负责。本次传输改善有证据;长程净效果与效率增益尚未证明。未合并、升级或启动/停止真实实验。
English verdict: APPROVE - aaa724a; full official-result delivery and source/epoch binding verified through the real native/Docker/Codex path (318 tests), with exact disabled/historical parity. Nonblocking P2: a mixed-generation cursor race is unchanged from base. Live adoption, model utility and study qualification remain outside this provider slice.
…ial-feedback Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Exact head: 78051b77160083d1ee1f6a09db24f1c2880ecbc9;base ec21f7da8b9a861fa73477f0ad69e8a4f30e5389。无阻塞发现。先前评审仅适用于 aaa724aebc49d534c89203a3054ed090f56e70a6;本次独立复核新 head,原 P2 并发问题已修复。
动机
运行已声明 best-only 的求解 Agent 与实验维护者,在自动评测产生严格改善时需要知道改善的具体依据。 以前只能知道某个提交快照胜出,无法从反馈中看到官方分数和诊断;现在同一胜出提交的完整官方结果随来源快照进入下一次模型请求。
已实测完整官方正文进入真实 Codex 的下一次请求及 resume;旧读者倒退 cursor 的并发反例已由同一锁内读取修复。信息可见和避免重复送达有证据,实际模型采用与得分收益未证明。 本 PR 不启动或停止实验,不改变评分,不证明分数或效率提升。 完整 formatter/task-image 与实际新组采用、四组合资格及 matched-study 的净分数/成本仍由原研究验收负责。
完整诊断可能帮助求解者理解哪些改动有效;代价是更长上下文与胜出时的控制器 HTTP 读取。不能据传输成功推断模型用对、得分上升或总成本下降。
改动思路
保留原生 select_best 作为唯一排序 owner,在既有 SForge provider 和共同部署的 reader 上扩展官方正文传输,比开放求解者评测权限或重建评分体系更贴合目标。 本次可回滚边界是胜出结果的来源校验、发布、实际宿主上下文输送和协议元数据;现有组的保存、停止、排空与新组准入属于后续独立操作。
只对已声明 best-only 的原生严格胜出提交返回其原始官方 API JSON;初始 incumbent、平局、退步、无效或未准入提交保持静默。控制器捕获并绑定 source、task、submission、epoch 与 evaluator,再经现有 SForge hook 放入请求上下文。普通 worker 不获得评测服务、凭据或其它 trial 的读取权。
这轮 future-facing pass 已在原 per-session lock 内完成:先取锁,再读 latest、校验、输出及更新 cursor。没有第二排序源、时间戳规则或新 generation journal;共同部署的 transient wire 只保留当前 v2,持久旧报告的缺省字段仍原样保留。
具体改动
全 PR 15 文件 +473/-75:8 个生产文件 +149/-32,6 份测试 +275/-36,README +49/-7。与上次 head 相比,只有 hook 锁范围、三种事件回归和 README 说明改变;两版 main 在全部15路径及相关 benchmark owner 上无中间改变。旧证据保留原 revision,新结论不继承旧批准。
feedback.update/_official_result复用原生select_best,只在严格改善时读取完整官方结果,核对 task/submission/score/pass-rate、epoch、evaluator 与已提交 source;失败不提交 incumbent,原任务可重试。online_judge.create_app、online_sampling.qualify记录本机 native source/task-spec、image key 和原准入事实;image key 不等于镜像内容不可变证明。feedback_hook.deliver(第21行)在三个 Codex hook 上输送完整 JSON、result digest 与 source checkpoint;锁前旧读取已删除。install保留原 Stop hook,原工具输出与本地工作继续。prompts明确异步反馈对应已提交快照,不盲目覆盖更新的工作树、不宣告完成;结果正文是 data,不能授予评测权限。未知嵌套 Unicode/null/bool 字段仍保留;“complete”以 native API 实际返回为边界,其自身省略/截断不被补造。run、SForgeWorker._install_worker、export_report为新 best-only 尝试记录 official-result payload,拒绝冲突元数据,旧导出不补标签。README 披露这一 treatment/wire 变化和锁范围。
接受依据:docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md,不可变 ec21f7da8b9a861fa73477f0ad69e8a4f30e5389,§6.5 Declared feedback and evaluation timing、§11.3 Required delivery slice;源码前的接受文本未被此 PR 修改。
- Allowed feedback carries only the benchmark-declared response.:implemented;当前源码与上述真实入口验证支持此有界条款。
- Blind background evaluation:implemented;当前源码与上述真实入口验证支持此有界条款。
- Feedback availability is the only changed factor:implemented;当前源码与上述真实入口验证支持此有界条款。
- Diagnostic observations retain their original qualification status:implemented;当前源码与上述真实入口验证支持此有界条款。
- Acceptance requires all four feedback/timing combinations:deferred;This bounded provider transport PR does not implement benchmark-toolkit protocol qualification or promote any campaign evidence.
独立当前版本验证:321 passed / 118.59s / zero skips。实际 native judge/API、普通 Docker 用户、Codex0.160.0 下一请求及 resume 均覆盖;task/grader 和 Responses 是合成边界。source/evaluator/epoch/result mismatch、API失败→原状态重试、网络/host-secret隔离、初始/平局静默继续通过。
另用同一个实际文件/flock脚本对当前 main、前一 candidate、新 head 做对照:前两者均 auto3→迟到auto2→再次auto3;新 head 只送 auto3,迟到读者与下一 hook 都静默、cursor 保持 auto3。把新回归直接运行在旧 candidate,PostToolUse/SessionStart/UserPromptSubmit 三例都在旧输出非空断言失败;新版本三例通过。历史 native/blind/best-only 全部导出文件、blind prompt 在三版逐字节相同,没把语义差异正规化掉。
对主干的风险
主要代价仍是完整正文的 token 与胜出时额外 HTTP 请求。当前修复消除了已证明的锁竞争重复正文,但不提供模型理解保证,也不宣称进程在输出后、cursor持久化前崩溃时具有模型确认/所有故障恰好一次保证。边界按现有回执与研究验收保留。新锁内空包可能创建本地 lock 文件;只影响显式 best-only hook,不影响 native/blind 的默认路径或评分规则。
Full-tree semantic、compile、diff/public-boundary 检查通过;共享闭合集合没有新扩展,v2 属 provider 局部 wire。advisory 排除 benchmark、结果零不能证明语义无变化,因此实际对照正文的 actor、scope、异步 source、继续工作、不覆盖和非完成义务。按当前 wait_for_ci=false 未读取、轮询或等待 GitHub CI。
我的整体评价
APPROVE;justified_increment。 体验 improved:完整有用诊断随原请求交付,无额外参数、轮询或确认;并发等待不会再倒退 cursor。长程 accepted_tradeoff:来源校验、原生排序、恢复、旧证据保留,有界 token/HTTP 代价由已接受 declared feedback 设计承载;实际任务净效果和效率仍未知。
保留原生 select_best 作为唯一排序 owner,在既有 SForge provider 和共同部署的 reader 上扩展官方正文传输,比开放求解者评测权限或重建评分体系更贴合目标。 本次可回滚边界是胜出结果的来源校验、发布、实际宿主上下文输送和协议元数据;现有组的保存、停止、排空与新组准入属于后续独立操作。 当前真实传输及并发修复可独立审查和回滚;完整新组/formatter/task-image采用、四组合资格、真实模型 uptake 与 matched score/cost 属原有后续验收。本次没有合并、全局升级或启动真实实验。
English verdict: APPROVE - 78051b7; complete source-bound official feedback verified through the real native/Docker/Codex path (321 tests), with unchanged native/blind/historical output. The prior mixed-generation cursor race is independently reproduced on base and old head and eliminated in all three hook events. Live uptake, score/cost effects and study qualification remain untested.
Goal And Delivered Outcome
Best-only currently tells the solver that a snapshot improved, but withholds the official explanation of that result. A selected strict improvement now delivers the complete native result API response alongside its verified submitted-source checkpoint, directly into the next Codex request.
main.Author Declaration
docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md, “Declared feedback and evaluation timing” atde3de98d14d3c6f735c156e49f01c491f0d355f8.feedback.py; native ranking, full response and non-improvement testsSelf-check: reviewed the changed provider paths, negative cases, public boundary and default change. Python remains the SForge-specific transport owner; no second generic control-plane decision source is introduced. The dependency-free wire reader owns the local schema/payload identifiers used by publication and receipts; this bounded consolidation avoids a new protocol module.
Scope And Continuation
Complete within the provider-delivery slice. New attempts use packet v2 and record
feedback_payload=official-resultin both runtime/profile receipts and portable settings. Historical exports retain absent payload metadata instead of being relabeled. This intentionally changes the best-only treatment; pinned old attempts retain their code and protocol.Concurrent hooks now read the current packet under their existing session lock. A delayed reader can no longer deliver an older generation after a newer one and roll back the cursor, replaying the newest full response. This reproduces on the current
main; the deterministic regression fails on the prior candidate and passes after the repair. It does not estimate live race incidence or task-score gains.Native scoring, selection, sampling cadence and admission opportunity stay unchanged. Actual solver uptake, cost and score comparisons remain separate operational/research work. No benchmark campaign was launched or canceled for this PR; no merge is requested by the validation result alone.
Validation
78051b77160083d1ee1f6a09db24f1c2880ecbc9(tested working-tree source bytes committed unchanged).git diff --checkThe admission provenance covers native Python source, task specification, image key, selection and direction. It does not attest immutable image contents. “Complete” means the unmodified native result API response; the native API's own omissions/truncation remain. Hidden files, raw judge output, evaluator code and unrelated trials are not exposed. Model visibility is proven; real benchmark adoption/effectiveness is untested.
Frontend / Visual Evidence
UI impact: none. Only provider CLI receipts, harness context and protocol documentation change; no frontend/Lark/product entrypoint or first viewport changes.
Type of Change
LoopX Area / Technical Direction
Shared-authority RFC fixture impact
N/A: no shared authority schema, provider routing or core TypeScript rule changes.
Boundary Checklist