Skip to content

feat(benchmark): return complete official feedback on best-only improvements - #6031

Merged
huangruiteng merged 5 commits into
mainfrom
codex/edgebench-official-feedback
Oct 9, 2026
Merged

huangruiteng merged 5 commits into
mainfrom
codex/edgebench-official-feedback

Conversation

@loopx-agent

@loopx-agent loopx-agent commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Goal And Delivered Outcome

Best-only currently tells the solver that a snapshot improved, but withholds the official explanation of that result. A selected strict improvement now delivers the complete native result API response alongside its verified submitted-source checkpoint, directly into the next Codex request.

  • Basis: maintainer-requested best-only feedback revision; target base main.
  • Before → after: identity-only notification → source-bound official result, retaining native selection, silent baseline/ties/regressions, fixed sampling and evaluator isolation.

Author Declaration

  • Written by: model_agent (OpenAI Codex).
  • Specification: this accepted request plus docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md, “Declared feedback and evaluation timing” at de3de98d14d3c6f735c156e49f01c491f0d355f8.
Criterion Disposition Owner / validation
Disclose only the selected native improvement's declared response implemented feedback.py; native ranking, full response and non-improvement tests
Bind result to submitted source and task implemented sampler admission; task/submission/score/pass-rate checks; epoch and evaluator provenance
Deliver through the harness without a polling tool call implemented installed system hook; real judge → publisher → Docker → staged Codex → next request and resume
Preserve blind/native behavior and authority boundaries implemented explicit-mode/runtime tests, denied judge access, host-only credentials, no new solver submission route
Scientific treatment uptake and matched-study effectiveness out_of_scope requires separately qualified runs; synthetic transport is not score evidence

Self-check: reviewed the changed provider paths, negative cases, public boundary and default change. Python remains the SForge-specific transport owner; no second generic control-plane decision source is introduced. The dependency-free wire reader owns the local schema/payload identifiers used by publication and receipts; this bounded consolidation avoids a new protocol module.

Scope And Continuation

Complete within the provider-delivery slice. New attempts use packet v2 and record feedback_payload=official-result in both runtime/profile receipts and portable settings. Historical exports retain absent payload metadata instead of being relabeled. This intentionally changes the best-only treatment; pinned old attempts retain their code and protocol.

Concurrent hooks now read the current packet under their existing session lock. A delayed reader can no longer deliver an older generation after a newer one and roll back the cursor, replaying the newest full response. This reproduces on the current main; the deterministic regression fails on the prior candidate and passes after the repair. It does not estimate live race incidence or task-score gains.

Native scoring, selection, sampling cadence and admission opportunity stay unchanged. Actual solver uptake, cost and score comparisons remain separate operational/research work. No benchmark campaign was launched or canceled for this PR; no merge is requested by the validation result alone.

Validation

  • Tested revision: 78051b77160083d1ee1f6a09db24f1c2880ecbc9 (tested working-tree source bytes committed unchanged).
  • Run state: finished. Input classes: synthetic, public_fixture.
Check kind Result Evidence / limitation
unit / integration passed 321 tests across feedback, hook, online sampling, prompts, SForge runtime, portable reports and the opt-in Docker suite
real_backend / real_entrypoint passed Real native judge and Docker; staged Codex 0.160.0 with synthetic Responses. Full official diagnostics entered the next request and native resume, preserving tool output
regression_parity passed Native rank and silent baseline/tie/regression cases; blind/native registration; wrong bindings/epoch/evaluator and unavailable API retain the incumbent for retry; source permission failures and concurrent hook deduplication, including a deterministic mixed-generation reader race in all three hook events
static passed Changed-source semantic advisory, full-tree semantic vocabulary smoke, Python compilation, public boundary scan and git diff --check

The admission provenance covers native Python source, task specification, image key, selection and direction. It does not attest immutable image contents. “Complete” means the unmodified native result API response; the native API's own omissions/truncation remain. Hidden files, raw judge output, evaluator code and unrelated trials are not exposed. Model visibility is proven; real benchmark adoption/effectiveness is untested.

Frontend / Visual Evidence

UI impact: none. Only provider CLI receipts, harness context and protocol documentation change; no frontend/Lark/product entrypoint or first viewport changes.

Type of Change

  • New feature
  • Breaking change (new-attempt best-only treatment / wire payload)
  • Documentation update
  • Test update

LoopX Area / Technical Direction

  • Benchmark boundary and host/runtime integration.
  • Long-horizon benchmark research: declared-feedback boundary and E1/E2 provider conformance. This does not close E3 matched-study readiness.

Shared-authority RFC fixture impact

N/A: no shared authority schema, provider routing or core TypeScript rule changes.

Boundary Checklist

  • No private state, credentials, raw traces, verifier output, internal links or local machine paths in the diff/body.
  • Maintainer-directed scope; related native fixes remain with their existing owners.
  • Scope is limited to source-bound best-only official feedback.
  • UI impact is none.
  • Every commit carries DCO sign-off.

…ements

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
… Codex

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Exact head: aaa724aebc49d534c89203a3054ed090f56e70a6. 无阻塞发现。下面的 P2 是基线已有问题,建议继续修复;不撤销本次通过结论。

动机

运行已声明 best-only 的求解 Agent 与实验维护者,在自动评测产生严格改善时需要知道改善的具体依据。

best-only 指只在原生选择规则判定提交严格改善时通知求解者。以前只能知道某个提交快照胜出,无法从反馈中看到官方分数和诊断;现在同一胜出提交的完整官方结果随来源快照进入下一次模型请求。 缺少正文会让求解者收到“变好了”的消息,却仍不知道哪项诊断值得采用。已实测官方正文进入真实 Codex 的下一次请求及 resume,且保留原工具输出;这证明信息可见,不证明模型已正确采用。

本 PR 不启动或停止实验,不改变评分,不证明分数或效率提升。 新 epoch 的安装、完整 formatter、任务镜像与各 worker 实际采用,以及 matched study 的效果和成本仍须单独验收。

改动思路

保留原生 select_best 作为唯一排序 owner,在既有 SForge provider 和共同部署的 reader 上扩展官方正文传输,比开放求解者评测权限或重建评分体系更贴合目标。

控制器从原生 API 取得正文,核对原任务、提交、分数、通过率、评测进程 epoch 与 evaluator 后,再验证提交归档及普通 worker 可读的副本。成功发布之后才更新 incumbent;失败保留原状态重试。首次有效 incumbent、平分、无效和未准入快照仍不通知,不增加评测机会。本次可回滚边界是胜出结果的来源校验、发布、实际宿主上下文输送和协议元数据;现有组的保存、停止、排空与新组准入属于后续独立操作。

这是一条 Python 原生 provider 传输链路,复用了既有采样器、publisher、SForge hook 与 report exporter;没有另建通用控制面或评分决策源。当前 producer/reader 一起部署,使用同一版本常量;持久化旧报告则继续保留原字段缺省,避免兼容代码与历史真相混为一谈。

具体改动

完整 diff 为 15 个文件、+418/-61:8 个生产文件 +134/-18,6 个测试文件 +237/-36,README +47/-7。除了正文传输,run/profile 与导出记录实际新尝试的 feedback_payload;缺省旧记录不补值,冲突/未知值拒绝。README 披露新 best-only 的全文行为和新 epoch 边界。未新增依赖、CLI 参数、全局能力开关或闲置模块。

关键代码讲解

  • feedback.py:198 BestFeedbackPublisher._official_result:读取所选 submission 的完整官方 JSON,核对 task/submission/score/pass-rate,再次读取 admission 核对原 epoch/evaluator。update 继续负责 source archive/copy 哈希和发布顺序;错误不推进 incumbent。
  • online_sampling.py:51 OnlineSampler.qualify:实际服务返回的 task/image key/selection/direction 与 task-file/module 摘要进入原生准入校验,未改变 capture 或评测容量规则。image key 不等于镜像内容不可变证明。
  • feedback_hook.py:21 deliver:校验 v2、来源归档、evaluator 与正文摘要,完整 JSON 进入 additionalContext。未知嵌套字段、Unicode、null/bool 保留;原工具输出与 Stop 不被覆盖。把正文标为 data 是指导,不能证明模型不受其中指令影响。
  • export_report.py:167 export:只有实际源数据包含新 payload 时才投影,旧 native/blind/best-only 报告保持字节一致。没有据新代码猜测旧运行采用新协议。
  • sforge.py:160 SForgeWorker._install_worker:仍只向 best-only 普通 worker 安装当前 hook,沿原宿主入口校验可读性;真实 Codex 下一次请求和 resume 已复验。

接受依据是 docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md,不可变 revision de3de98d14d3c6f735c156e49f01c491f0d355f8,§6.5 及 §11.3。原文没有条款编号,下面以原文短语识别:

  • Allowed feedback carries only the benchmark-declared response.:本次传输实际官方 API JSON,绑定原 snapshot;不授予隐藏测试、参考答案、评测实现、凭据或其它 trial 的访问。native API 本身省略/限制的内容不由本 PR 重建。
  • Blind background evaluation:本次未扩大 blind 入口权限;实际隔离与恢复用例和 base/head prompt/report 对照覆盖相关路径。
  • Feedback availability is the only changed factor:任务、原生选择、预算、采样节奏与原 source pins 保留。改变的是新 best-only 明确声明的反馈正文;不把不同反馈协议混入同一历史组。
  • Diagnostic observations retain their original qualification status:旧记录未重标,没有将缺失 integrity 证据升级成 eligible。
  • Acceptance requires all four feedback/timing combinations:deferred。本 PR 交付可用 provider seam,未交付 toolkit 四组合资格判定;既有 benchmark-toolkit typed integrity policy / E3 matched-study 验收继续负责,不以本次 green tests 代替。

独立实测在确切 head 与固定 native source 运行 7 个文件:318 passed,零 skip。含真实 native judge、Docker、普通 worker、Codex hook、下一次 synthetic Responses 请求及 resume;评测服务网络和宿主秘密文件对 worker 的不可见路径已测。任务/grader/Responses 是合成夹具,故这不是付费模型的理解或真实题目效果证明。额外 base/head probe 使用实际 exporter 与文件锁:blind prompt 和 native、blind、历史 best-only 全部 5 个导出文件的字节摘要相同,head 的官方正文未知字段完整保留。

对主干的风险

最强新风险是把错误提交或新 epoch 的报告当作原提交反馈;HTTP、task/submission/score/pass-rate/epoch/evaluator 与 archive 的负例已覆盖,失败先拒绝,后续对原状态重试成功。全文也增加上下文成本,每次胜出增加两次控制器 GET;它不增加评测提交/slots,超时有界,但正文 token 数和模型净收益未量化。hash 证明相等,不单独证明外部真实性、镜像内容、模型采用或实验资格。

P2 非阻塞建议,基线已有:混合代次 reader 可倒退 cursor。 feedback_hook.py:63 之前已在锁外读取 latest。独立强制交错:旧 reader 先读 auto-2 并在锁前暂停,新 reader 读 auto-3、交付并记 cursor,随后旧 reader 获锁把 cursor 改回 auto-2;下一次 auto-3 又被交付。immutable base 和当前 head 都复现同一顺序及错误,新增 checksum 没有改变这个因果路径。现有同包并发用例通过不能排除它,全文会放大重复上下文成本。建议把 latest 的读取/校验纳入同一 session 锁,或锁内重新读取并丢弃旧候选;加入上述混合代次真实 file/flock 交错回归,验证 cursor 不倒退、latest 不重复。

语义与 CI 对齐

改动是 provider 本地 wire v2 与明确声明的反馈正文,不是新增共享控制面词汇。逐项比对了任务/actor、异步 source、禁止覆盖 live workspace、继续本地工作与非完成提示;native/blind 路径保持,没把 guidance 宣称为机器保证。semantic full-tree 检查通过,覆盖 26/26 vocabulary 和 51/51 owner;第一次因缺 TypeScript 开发依赖退出,准备依赖后同一源码检查通过,失败记录保留。diff advisory 排除了 benchmark,扫描为零,不能拿它证明语义无变化;局部协议及文本已人工对照。compile、diff whitespace、public/private 边界检查通过。按本次原生 review policy 的 wait_for_ci=false,未读取、轮询或等待 GitHub CI,结论依据上述本机验证。

我的整体评价

APPROVE;delivery judgment 为 justified_increment。 用户体验为 improved:普通 worker 通过现有请求看到真正有用的官方诊断,无须新参数、轮询工具或额外确认。长程效果为 accepted_tradeoff:严格胜出过滤、原生排序、来源校验、失败重试、历史协议和恢复链路保留;代价是更多正文 token 与胜出时控制器读请求。公开记录了已有 cursor race,不宣称所有并发恰好一次。

架构保留既有 native/provider owner,生产增量与真实调用相称;未为共同部署的 transient wire 保留多套 decoder,也没有迁移历史报告。相关 future-facing pass 复用了同一版本常量并避免第二排序源;锁内 latest 读取建议属于小范围后续修复。完整 formatter、task-image 与每个实际 worker 的安装/readback,以及四组合资格和 matched-study 分数/成本,由既有操作及研究验收继续负责。本次传输改善有证据;长程净效果与效率增益尚未证明。未合并、升级或启动/停止真实实验。

English verdict: APPROVE - aaa724a; full official-result delivery and source/epoch binding verified through the real native/Docker/Codex path (318 tests), with exact disabled/historical parity. Nonblocking P2: a mixed-generation cursor race is unchanged from base. Live adoption, model utility and study qualification remain outside this provider slice.

…ial-feedback

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Exact head: 78051b77160083d1ee1f6a09db24f1c2880ecbc9;base ec21f7da8b9a861fa73477f0ad69e8a4f30e5389。无阻塞发现。先前评审仅适用于 aaa724aebc49d534c89203a3054ed090f56e70a6;本次独立复核新 head,原 P2 并发问题已修复。

动机

运行已声明 best-only 的求解 Agent 与实验维护者,在自动评测产生严格改善时需要知道改善的具体依据。 以前只能知道某个提交快照胜出,无法从反馈中看到官方分数和诊断;现在同一胜出提交的完整官方结果随来源快照进入下一次模型请求。

已实测完整官方正文进入真实 Codex 的下一次请求及 resume;旧读者倒退 cursor 的并发反例已由同一锁内读取修复。信息可见和避免重复送达有证据,实际模型采用与得分收益未证明。 本 PR 不启动或停止实验,不改变评分,不证明分数或效率提升。 完整 formatter/task-image 与实际新组采用、四组合资格及 matched-study 的净分数/成本仍由原研究验收负责。

完整诊断可能帮助求解者理解哪些改动有效;代价是更长上下文与胜出时的控制器 HTTP 读取。不能据传输成功推断模型用对、得分上升或总成本下降。

改动思路

保留原生 select_best 作为唯一排序 owner,在既有 SForge provider 和共同部署的 reader 上扩展官方正文传输,比开放求解者评测权限或重建评分体系更贴合目标。 本次可回滚边界是胜出结果的来源校验、发布、实际宿主上下文输送和协议元数据;现有组的保存、停止、排空与新组准入属于后续独立操作。

只对已声明 best-only 的原生严格胜出提交返回其原始官方 API JSON;初始 incumbent、平局、退步、无效或未准入提交保持静默。控制器捕获并绑定 source、task、submission、epoch 与 evaluator,再经现有 SForge hook 放入请求上下文。普通 worker 不获得评测服务、凭据或其它 trial 的读取权。

这轮 future-facing pass 已在原 per-session lock 内完成:先取锁,再读 latest、校验、输出及更新 cursor。没有第二排序源、时间戳规则或新 generation journal;共同部署的 transient wire 只保留当前 v2,持久旧报告的缺省字段仍原样保留。

具体改动

全 PR 15 文件 +473/-75:8 个生产文件 +149/-32,6 份测试 +275/-36,README +49/-7。与上次 head 相比,只有 hook 锁范围、三种事件回归和 README 说明改变;两版 main 在全部15路径及相关 benchmark owner 上无中间改变。旧证据保留原 revision,新结论不继承旧批准。

  • feedback.update/_official_result 复用原生 select_best,只在严格改善时读取完整官方结果,核对 task/submission/score/pass-rate、epoch、evaluator 与已提交 source;失败不提交 incumbent,原任务可重试。
  • online_judge.create_app、online_sampling.qualify 记录本机 native source/task-spec、image key 和原准入事实;image key 不等于镜像内容不可变证明。
  • feedback_hook.deliver(第21行)在三个 Codex hook 上输送完整 JSON、result digest 与 source checkpoint;锁前旧读取已删除。install 保留原 Stop hook,原工具输出与本地工作继续。
  • prompts 明确异步反馈对应已提交快照,不盲目覆盖更新的工作树、不宣告完成;结果正文是 data,不能授予评测权限。未知嵌套 Unicode/null/bool 字段仍保留;“complete”以 native API 实际返回为边界,其自身省略/截断不被补造。
  • run、SForgeWorker._install_worker、export_report 为新 best-only 尝试记录 official-result payload,拒绝冲突元数据,旧导出不补标签。README 披露这一 treatment/wire 变化和锁范围。

接受依据:docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md,不可变 ec21f7da8b9a861fa73477f0ad69e8a4f30e5389,§6.5 Declared feedback and evaluation timing、§11.3 Required delivery slice;源码前的接受文本未被此 PR 修改。

  • Allowed feedback carries only the benchmark-declared response.:implemented;当前源码与上述真实入口验证支持此有界条款。
  • Blind background evaluation:implemented;当前源码与上述真实入口验证支持此有界条款。
  • Feedback availability is the only changed factor:implemented;当前源码与上述真实入口验证支持此有界条款。
  • Diagnostic observations retain their original qualification status:implemented;当前源码与上述真实入口验证支持此有界条款。
  • Acceptance requires all four feedback/timing combinations:deferred;This bounded provider transport PR does not implement benchmark-toolkit protocol qualification or promote any campaign evidence.

独立当前版本验证:321 passed / 118.59s / zero skips。实际 native judge/API、普通 Docker 用户、Codex0.160.0 下一请求及 resume 均覆盖;task/grader 和 Responses 是合成边界。source/evaluator/epoch/result mismatch、API失败→原状态重试、网络/host-secret隔离、初始/平局静默继续通过。

另用同一个实际文件/flock脚本对当前 main、前一 candidate、新 head 做对照:前两者均 auto3→迟到auto2→再次auto3;新 head 只送 auto3,迟到读者与下一 hook 都静默、cursor 保持 auto3。把新回归直接运行在旧 candidate,PostToolUse/SessionStart/UserPromptSubmit 三例都在旧输出非空断言失败;新版本三例通过。历史 native/blind/best-only 全部导出文件、blind prompt 在三版逐字节相同,没把语义差异正规化掉。

对主干的风险

主要代价仍是完整正文的 token 与胜出时额外 HTTP 请求。当前修复消除了已证明的锁竞争重复正文,但不提供模型理解保证,也不宣称进程在输出后、cursor持久化前崩溃时具有模型确认/所有故障恰好一次保证。边界按现有回执与研究验收保留。新锁内空包可能创建本地 lock 文件;只影响显式 best-only hook,不影响 native/blind 的默认路径或评分规则。

Full-tree semantic、compile、diff/public-boundary 检查通过;共享闭合集合没有新扩展,v2 属 provider 局部 wire。advisory 排除 benchmark、结果零不能证明语义无变化,因此实际对照正文的 actor、scope、异步 source、继续工作、不覆盖和非完成义务。按当前 wait_for_ci=false 未读取、轮询或等待 GitHub CI。

我的整体评价

APPROVE;justified_increment。 体验 improved:完整有用诊断随原请求交付,无额外参数、轮询或确认;并发等待不会再倒退 cursor。长程 accepted_tradeoff:来源校验、原生排序、恢复、旧证据保留,有界 token/HTTP 代价由已接受 declared feedback 设计承载;实际任务净效果和效率仍未知。

保留原生 select_best 作为唯一排序 owner,在既有 SForge provider 和共同部署的 reader 上扩展官方正文传输,比开放求解者评测权限或重建评分体系更贴合目标。 本次可回滚边界是胜出结果的来源校验、发布、实际宿主上下文输送和协议元数据;现有组的保存、停止、排空与新组准入属于后续独立操作。 当前真实传输及并发修复可独立审查和回滚;完整新组/formatter/task-image采用、四组合资格、真实模型 uptake 与 matched score/cost 属原有后续验收。本次没有合并、全局升级或启动真实实验。

English verdict: APPROVE - 78051b7; complete source-bound official feedback verified through the real native/Docker/Codex path (321 tests), with unchanged native/blind/historical output. The prior mixed-generation cursor race is independently reproduced on base and old head and eliminated in all three hook events. Live uptake, score/cost effects and study qualification remain untested.

@huangruiteng
huangruiteng merged commit 41f0089 into main Oct 9, 2026
3 checks passed
@huangruiteng
huangruiteng deleted the codex/edgebench-official-feedback branch October 9, 2026 16:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants