Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 49 additions & 7 deletions benchmark/edgebench/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -171,14 +171,17 @@ but short probes are not a prerequisite for running the intended protocol.
`--feedback best-only` is now the default for **new attempts across all five
workers**. This is an explicit protocol change from native, not a demonstrated
score improvement. Existing trials and pinned study manifests keep their modes.
Use `--feedback native` to retain the previous default or `--feedback blind` as
Selected improvements now include the complete official result API response,
rather than the earlier notification-only payload. This changes the new-run
best-only treatment; old pinned attempts must keep their original protocol.
Use `--feedback native` to retain agent-requested evaluation or `--feedback blind` as
an evaluator-feedback-free control. Harbor is unchanged.

| Mode | Agent-visible evaluator feedback | Evaluation access |
| --- | --- | --- |
| native | Exact score, pass rate, counts, summary, metrics and failed names; `--details` exposes per-check messages; `--list` shows active-submission history | Agent may submit within the native cooldown/budget; automatic samples are hidden |
| blind | None; public task files, local tests and compiler feedback remain available | Host evaluates fixed automatic samples; agent has no judge route or credentials |
| best-only | Latest strict improvement notification and the corresponding submitted-source checkpoint; no score, delta, diagnostics or negative-result status | Fixed capture cadence; one evaluator and latest pending capture per run; agent cannot request extra evaluations |
| best-only | Latest strict improvement's complete official result API response and corresponding submitted-source checkpoint; baseline, ties, regressions and unsuccessful evaluations stay silent | Fixed capture cadence; one evaluator and latest pending capture per run; agent cannot request extra evaluations |

Best-only supports non-game, offline tasks with `score_first`,
`valid_then_score` or `pass_rate_first` selection, including maximizing and minimizing scores. It
Expand All @@ -199,7 +202,17 @@ a lost response cannot create duplicate evaluations. An unexpected judge restart
holds that attempt for reconciliation instead of silently replaying it.

The publisher accepts only that sampler's admitted submission/round identities
and verifies the original source digest. Offline/history-only results cannot
and verifies the original source digest. Before publication it fetches the
selected submission's official result, verifies task/submission/score/pass-rate
against native history, and checks the admission epoch and evaluator provenance.
Provenance includes native Python source and task-specification digests, the judge
image key, selection policy and score direction; the image key is not an immutable
image-content attestation. Result-body digests bind hook delivery to that response.
An unavailable or mismatched response preserves the incumbent for retry.
The response is returned in full as provided by the native API, including any
summary, metrics and diagnostics the task supplies. This does not expand access
to raw judge output, hidden files, evaluator code or unrelated trials.
Offline/history-only results cannot
establish the baseline or change the online incumbent. The solver command's exit
pauses capture/delivery; official outer resume reuses the same publisher and lane.
After native worker cleanup, the host sends an epoch-bound release for the exact
Expand Down Expand Up @@ -249,17 +262,19 @@ archives retain their original permissions; evaluator and secret directories
are not made public. The hook script and system hook configuration are also
checked for ordinary-worker readability before handoff.

All five workers receive the allowlisted notification directly through managed
All five workers receive the source-bound official result directly through managed
Codex `PostToolUse`, `SessionStart` and `UserPromptSubmit` hooks. A short synchronous
reader adds `additionalContext` before the next model request, preserving the
original tool result. It does not force a wake, interrupt active reasoning or
change Stop/continuation behavior. The official worker keeps its native Stop
hook. Delivery is serialized and deduplicated per Codex session; a fresh session
receives the latest checkpoint, while resume receives only a new one. The
reader neither queries the evaluator nor interprets arbitrary packet prose.
Official report text is evaluation data, not executable instructions or authority.

This transport is qualified with staged Codex 0.160.0, the actual system hook in
a disposable container and a synthetic Responses endpoint: a new result arriving
Qualification uses a real native judge and publisher, staged Codex 0.160.0, the
actual system hook in a disposable container and a synthetic Responses endpoint:
a complete official result arriving
during a tool call enters the next model request; a subsequent result enters a
resumed request. Earlier Codex versions must be qualified before admission.
Inspect session input and subsequent checkpoint adoption separately: injection
Expand All @@ -270,12 +285,37 @@ A notification looks like this (digest abbreviated for illustration):

```json
{
"schema_version": "edgebench_best_feedback_v1",
"schema_version": "edgebench_best_feedback_v2",
"latest": {
"kind": "new_best",
"snapshot_id": "auto-7",
"source_sha256": "<SHA-256>",
"source_archive": "/opt/edgebench-feedback/auto-7-<SHA-256>.tar.gz",
"run_id": "example-run",
"task_id": "example-task",
"online_epoch": "<judge-process-epoch>",
"evaluator": {
"native_source_sha256": "<SHA-256>",
"task_spec_sha256": "<SHA-256>",
"judge_image_key": "example.judge.task:version",
"selection": "score_first",
"score_direction": "maximize"
},
"result_sha256": "<SHA-256 of the canonical official_result JSON>",
"official_result": {
"submission_id": "example-submission",
"status": "completed",
"error": null,
"report": {
"task_id": "example-task",
"submission_id": "example-submission",
"valid": true,
"score": 3,
"pass_rate": 0.75,
"summary": "Task-provided official summary",
"metrics": {"coverage": 0.75}
}
},
"message": "This evaluated snapshot strictly improved the task's native ranking among valid scored snapshots. It may differ from your current files; keep using local validation."
}
}
Expand All @@ -287,6 +327,8 @@ failures retry before committing an improvement. The host-only `best-only-host`
artifacts record file publication and health; never mount or copy them into the
worker. Worker-local session hook receipts under `/logs/agent/best-feedback-delivery`
record emission, not model acknowledgement.
Each hook reads the current packet while holding its session delivery lock, so a
waiting reader cannot overwrite a newer delivery cursor and replay old feedback.
Treat missing archive/delivery evidence as an unqualified treatment, not as a
successful best-only trial. Public notifications do not include these errors.

Expand Down
16 changes: 16 additions & 0 deletions benchmark/edgebench/export_report.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,15 @@
from loopx.capabilities.benchmark_toolkit.experiment_board import (
normalize_benchmark_experiment_board_row,
)
from .feedback_hook import FEEDBACK_PAYLOAD


def _validate_feedback_payload(runner, profile):
payload = runner.get("feedback_payload")
if payload != profile.get("feedback_payload"):
raise ValueError("Settings profile disagrees with feedback payload")
if payload is not None and (payload != FEEDBACK_PAYLOAD or runner["feedback"] != "best-only"):
raise ValueError("Unsupported feedback payload for this mode")


def digest(path: Path) -> str:
Expand Down Expand Up @@ -89,6 +98,7 @@ def _validate_run_data(row, config, summary, points):
if config[key] != row[row_key]:
raise ValueError(f"Settings identity mismatch: {key}")
runner, profile = config["runner"], config["worker_profile"]
_validate_feedback_payload(runner, profile)
if (
runner["runner_commit"] != row["runner_revision"]
or runner["model"] != row["model_id"]
Expand Down Expand Up @@ -204,6 +214,7 @@ def export(runs_root: Path, selection: list, output: Path, *, observed_at: str):
raise ValueError(f"Profile disagrees with runtime: {key}")
if row.get("runner_revision") != receipt["runner_commit"]:
raise ValueError("Board and runner revision disagree")
_validate_feedback_payload(receipt, profile)
if number(final["best_score"]) != number(receipt["best_score"]):
raise ValueError("Final and runtime scores disagree")
score_entries = [
Expand Down Expand Up @@ -257,6 +268,11 @@ def export(runs_root: Path, selection: list, output: Path, *, observed_at: str):
"feedback",
)
}
# Old notification-only attempts retain absent/unknown payload metadata;
# never relabel them using today's default or operator-supplied settings.
if "feedback_payload" in receipt:
config["runner"]["feedback_payload"] = receipt["feedback_payload"]
config["worker_profile"]["feedback_payload"] = profile["feedback_payload"]
row["status"], row["observed_at"] = "completed", observed_at
row["metrics"]["best_score"] = {
"value": final["best_score"],
Expand Down
47 changes: 42 additions & 5 deletions benchmark/edgebench/feedback.py
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
"""Host-owned, positive-only projection of SForge's automatic evaluations.
"""Host-owned, best-only delivery of SForge's official evaluation response.

Provider policy only: SForge still captures, evaluates and selects final scores.
The worker receives an allowlisted notification and its own submitted source,
never judge reports, credentials, scores or negative-result metadata.
The worker receives the selected improvement's complete native result and its
own submitted source, never credentials, other runs or non-improving reports.
"""
from __future__ import annotations

Expand All @@ -16,6 +16,7 @@

import requests
from sforge.harness.selection import select_best
from .feedback_hook import SCHEMA_VERSION

FEEDBACK_MODES = ("native", "blind", "best-only")
FEEDBACK_ROOT = PurePosixPath("/opt/edgebench-feedback")
Expand Down Expand Up @@ -88,7 +89,7 @@ def start(self, backend, handle):
return
if self.token is None or self.stop_event.is_set():
raise RuntimeError("Feedback registration is missing or already closed")
self._publish_json(backend, handle, {"schema_version": "edgebench_best_feedback_v1", "latest": None})
self._publish_json(backend, handle, {"schema_version": SCHEMA_VERSION, "latest": None})
self._record()
self.sampler.start(backend, handle, self.token)
self.thread = threading.Thread(target=self._loop, args=(backend, handle), daemon=True)
Expand Down Expand Up @@ -161,6 +162,9 @@ def update(self, history, backend, handle):
admission = self.sampler.admitted()[candidate["submission_id"]]
if admission["source_sha256"] != digest:
raise ValueError("Native archive differs from admitted online capture")
result = self._official_result(candidate)
result_sha256 = hashlib.sha256(json.dumps(result, sort_keys=True,
ensure_ascii=False, allow_nan=False).encode()).hexdigest()
remote = FEEDBACK_ROOT / f"{snapshot}-{digest}.tar.gz"
prepare_feedback_root(backend, handle)
backend.copy_to_container(handle, archive, remote)
Expand All @@ -174,9 +178,13 @@ def update(self, history, backend, handle):
if readable.exit_code or readable.output.split()[:1] != [digest]:
raise RuntimeError("Could not verify best-only source checkpoint for worker")
packet = {
"schema_version": "edgebench_best_feedback_v1",
"schema_version": SCHEMA_VERSION,
"latest": {"kind": "new_best", "snapshot_id": snapshot,
"source_sha256": digest, "source_archive": str(remote),
"run_id": self.run_id, "task_id": self.task_id,
"online_epoch": self.sampler.epoch,
"evaluator": self.sampler.evaluators[self.task_id],
"official_result": result, "result_sha256": result_sha256,
"message": _MESSAGE},
}
if self.stop_event.is_set():
Expand All @@ -187,6 +195,35 @@ def update(self, history, backend, handle):
self.notifications += 1
self._record(packet)

def _official_result(self, candidate):
# Fetch only the sampler-admitted, selected native submission. Do not
# read raw judge logs, verifier code or another run's result directory.
response = self.session.get(f"{self.judge_url}/api/v1/result/{candidate['submission_id']}",
timeout=(3, 5))
response.raise_for_status()
result = response.json()
report = result.get("report")
if (result.get("submission_id") != candidate["submission_id"]
or result.get("status") != "completed" or result.get("error") is not None
or not isinstance(report, dict) or report.get("valid", True) is not True
or report.get("task_id") != self.task_id
or report.get("submission_id") != candidate["submission_id"]
or type(report.get("score")) not in (int, float)
or not math.isfinite(report["score"]) or report["score"] != candidate["score"]
or report.get("pass_rate") != candidate.get("pass_rate")):
raise ValueError("Official result does not match the selected native history entry")
# A restarted service can reuse in-memory submission IDs. Its epoch and
# evaluator must still match the pre-solver admission, even on retry.
response = self.session.get(f"{self.judge_url}/api/v1/best-only/admission",
params={"admin_secret": self.admin_secret}, timeout=(3, 5))
response.raise_for_status()
admission = response.json()
evaluator = admission.get("evaluators", {}).get(self.task_id)
if (admission.get("epoch") != self.sampler.epoch or not evaluator
or evaluator != self.sampler.evaluators.get(self.task_id)):
raise ValueError("Official evaluator changed; reconcile the original trial")
return result

def _record(self, packet=None):
# Host-only health/evidence. Do not mount this directory into the worker.
value = {"best_score": self.score, "notifications": self.notifications, "errors": self.errors}
Expand Down
Loading
Loading