Skip to content

template: confirmation pass with more paired rounds on unsure benchmarks - #54

Merged
ArjunS07 merged 2 commits into
survey/phase1-template-20261008from
feat/lsv-confirm-pass
Oct 8, 2026
Merged

ArjunS07 merged 2 commits into
survey/phase1-template-20261008from
feat/lsv-confirm-pass

Conversation

@ArjunS07

@ArjunS07 ArjunS07 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Purpose

One timing session of 3 paired rounds leaves enough round-to-round noise that, on bottleneck#304, 15 to 45 of 805 benchmarks that the patch does not touch look slower than their noise threshold $\theta_b$ with no change. The noise between sessions is not the problem. This change adds a confirmation pass: more paired rounds in the same container, only on the benchmarks the first pass left unsure.

How it works

flowchart TD
    A["first pass: lsv_measure.py --rounds 3<br/>(gate step: measure)"] --> B{"/tests/confirm.json?"}
    B -- no --> E[snapshots, pytest, parser as before]
    B -- yes --> C["write_confirm_set:<br/>|median ln s_b| > theta_b, plus oracle_up"]
    C -- empty --> E
    C -- ids --> D["lsv_measure.py --only confirm_set.json --rounds N<br/>--out lsv_measure_confirm.json<br/>(gate step: confirm)"]
    D --> E
Loading
  • The confirm set is every benchmark whose first-pass median $\ln(t_{base}/t_{patched})$ has $|x| &gt; \theta_b$ (either direction), plus oracle_up. $\theta_b$ comes from confirm.json theta, else theta_default. When the first pass measured nothing, there is no second pass.
  • The second pass uses the same base copy at /workspace/.fc_base and the same A-B-B-A order. It does not rebuild: the first pass already built both trees (FC_REBUILD_CMD="" for this one call).
  • It holds the measure gate in the same way as the first pass (mg_acquire confirm / mg_release). The acquire is logged as confirm 1|0 in measure_gate.txt, so the pipeline's ungated check covers it.
  • A failed confirmation pass logs a warning. The first pass's results and the reward stay as they are.

confirm.json, written by the survey pipeline into a copy of the task:

{"rounds": 6, "theta": {"bench.A.time_x-2": 0.041}, "theta_default": 0.0296, "oracle_up": ["bench.A.time_x-2"], "theta_source": "record"}

lsv/lsv_measure_confirm.json has the same layout as lsv_measure_results.json (paired mode), plus three top-level keys:

{
  "benchmarks": {
    "bench.A.time_x-2": {
      "name": "bench.A.time_x-2", "baseline": 0.0123, "current": 0.0101, "delta_pct": -17.9, "params": {"p0": 100},
      "paired": {"base_times": [0.0124, 0.0122, 0.0123, 0.0125, 0.0121, 0.0123],
                 "patched_times": [0.0101, 0.0100, 0.0102, 0.0101, 0.0099, 0.0102],
                 "log_ratios": [0.205, 0.199, 0.187, 0.213, 0.200, 0.187],
                 "median_log_ratio": 0.200, "mad_log_ratio": 0.009, "speedup": 1.22}
    }
  },
  "selected_count": 1, "total_count": 805, "skipped_count": 804, "dropped_count": 0, "dropped": [],
  "timing": {"total_s": 412.0, "phases": null}, "error": null,
  "paired": {"rounds": 6, "min_pairs": 3, "order": ["base", "patched", "patched", "base", "..."]},
  "only_count": 2, "confirm_count": 2, "theta_source": "record"
}

Per-round values were already stored in both files: paired.base_times and paired.patched_times hold one value per round in round order (null where a side gave no time), and paired.log_ratios holds $\ln(t_{base}/t_{patched})$ for each round where both sides have a time. To combine the passes, a grader takes, per benchmark, the median over the first pass's log_ratios and the confirmation pass's log_ratios together (3 + N values).

Changes

  1. lsv_measure.py --only FILE
    • Effect: times only the listed benchmark ids, intersected with what LSV selects. The file is a JSON list of ids, or {"ids": [...], ...} whose other keys go to the top level of the output. Ids use the form in lsv_measure_results.json (<name> or <name>-<param index>). asv times all parameter combinations of a selected benchmark; only the listed ids are kept. Paired mode only: without the base copy the output has an error and no benchmarks.
    • Before: no way to time a subset.
  2. lsv_measure.py --out NAME
    • Effect: writes the result to OUTPUT_DIR/NAME. With a name other than the default, lsv_results.json is not touched, and every exit path (no change, shadowed project, error) writes the file.
    • Before: always lsv_measure_results.json.
  3. confirm_set and write_confirm_set in lsv_measure.py
    • Effect: test.sh writes lsv/confirm_set.json (ids, confirm_count, theta_source) and gets the number of rounds (rounds, default 6; 0 when the set is empty).
    • Before: none.
  4. test.sh confirmation pass
    • Effect: runs only when /tests/confirm.json exists and the set is not empty, inside the measure gate. test_timings.json gets lsv_confirm_s.
    • Before: one pass.
  5. tests/docker/test_lsv_confirm_pass.py
    • Effect: tests the set rule, the id mapping, the --only filter and output (with asv stubbed), and the test.sh block (gate around the call, no rebuild, no pass without confirm.json or with an empty set). It also checks that setup.sh, lsv_init.py and environment/ never name confirm.json, and that the only copy out of /tests in setup.sh is rebuild.sh.

Usage

Nothing changes without tests/confirm.json. The survey pipeline writes it when CONFIRM_ROUNDS > 0 (formula-code/formulacode-verified-rl, branch feat/confirm-json). By hand:

echo '{"rounds": 6, "theta_default": 0.0296, "theta_source": "floor"}' > <task>/tests/confirm.json

Verification

  • tests/docker: 431 passed, 10 skipped (includes the 10 new tests).
  • lsv_measure.py loads and runs confirm_set / only_names on Python 3.8 (uv run --no-project --python 3.8).
  • Not run in a real task container: a trial on this machine would compete for CPU with the live survey's timing.

Notes

  • Leak safety: confirm.json names the oracle's improved benchmarks. Harbor removes /tests before a non-oracle agent runs and uploads tests/ again for the verifier, so only test.sh reads it, at verification. setup.sh copies only rebuild.sh out of /tests; the new test checks this.
  • parser.py and the reward still read only the first pass. Combining the passes is up to the grader.
  • The survey pipeline's capture copy replaces each python /tests/lsv_measure.py call with true and required exactly 2 replacements. This template has 3; the pipeline branch above accepts 2 or more.

With tests/confirm.json, test.sh times again, in the same container and inside the measure gate, the benchmarks whose
first-pass median log speedup is beyond theta plus the oracle's improved ones. lsv_measure.py gets --only and --out.
@ArjunS07
ArjunS07 merged commit fd286c6 into survey/phase1-template-20261008 Oct 8, 2026
0 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant