Repository navigation
template: confirmation pass with more paired rounds on unsure benchmarks - #54
Merged
Merged
Conversation
With tests/confirm.json, test.sh times again, in the same container and inside the measure gate, the benchmarks whose first-pass median log speedup is beyond theta plus the oracle's improved ones. lsv_measure.py gets --only and --out.
ArjunS07
merged commit Oct 8, 2026
fd286c6
into
survey/phase1-template-20261008
0 of 3 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
One timing session of 3 paired rounds leaves enough round-to-round noise that, on bottleneck#304, 15 to 45 of 805 benchmarks that the patch does not touch look slower than their noise threshold$\theta_b$ with no change. The noise between sessions is not the problem. This change adds a confirmation pass: more paired rounds in the same container, only on the benchmarks the first pass left unsure.
How it works
flowchart TD A["first pass: lsv_measure.py --rounds 3<br/>(gate step: measure)"] --> B{"/tests/confirm.json?"} B -- no --> E[snapshots, pytest, parser as before] B -- yes --> C["write_confirm_set:<br/>|median ln s_b| > theta_b, plus oracle_up"] C -- empty --> E C -- ids --> D["lsv_measure.py --only confirm_set.json --rounds N<br/>--out lsv_measure_confirm.json<br/>(gate step: confirm)"] D --> Eoracle_up.confirm.jsontheta, elsetheta_default. When the first pass measured nothing, there is no second pass./workspace/.fc_baseand the same A-B-B-A order. It does not rebuild: the first pass already built both trees (FC_REBUILD_CMD=""for this one call).mg_acquire confirm/mg_release). The acquire is logged asconfirm 1|0inmeasure_gate.txt, so the pipeline's ungated check covers it.confirm.json, written by the survey pipeline into a copy of the task:{"rounds": 6, "theta": {"bench.A.time_x-2": 0.041}, "theta_default": 0.0296, "oracle_up": ["bench.A.time_x-2"], "theta_source": "record"}lsv/lsv_measure_confirm.jsonhas the same layout aslsv_measure_results.json(paired mode), plus three top-level keys:{ "benchmarks": { "bench.A.time_x-2": { "name": "bench.A.time_x-2", "baseline": 0.0123, "current": 0.0101, "delta_pct": -17.9, "params": {"p0": 100}, "paired": {"base_times": [0.0124, 0.0122, 0.0123, 0.0125, 0.0121, 0.0123], "patched_times": [0.0101, 0.0100, 0.0102, 0.0101, 0.0099, 0.0102], "log_ratios": [0.205, 0.199, 0.187, 0.213, 0.200, 0.187], "median_log_ratio": 0.200, "mad_log_ratio": 0.009, "speedup": 1.22} } }, "selected_count": 1, "total_count": 805, "skipped_count": 804, "dropped_count": 0, "dropped": [], "timing": {"total_s": 412.0, "phases": null}, "error": null, "paired": {"rounds": 6, "min_pairs": 3, "order": ["base", "patched", "patched", "base", "..."]}, "only_count": 2, "confirm_count": 2, "theta_source": "record" }Per-round values were already stored in both files:$\ln(t_{base}/t_{patched})$ for each round where both sides have a time. To combine the passes, a grader takes, per benchmark, the median over the first pass's
paired.base_timesandpaired.patched_timeshold one value per round in round order (nullwhere a side gave no time), andpaired.log_ratiosholdslog_ratiosand the confirmation pass'slog_ratiostogether (3 + N values).Changes
lsv_measure.py --only FILE{"ids": [...], ...}whose other keys go to the top level of the output. Ids use the form inlsv_measure_results.json(<name>or<name>-<param index>). asv times all parameter combinations of a selected benchmark; only the listed ids are kept. Paired mode only: without the base copy the output has an error and no benchmarks.lsv_measure.py --out NAMEOUTPUT_DIR/NAME. With a name other than the default,lsv_results.jsonis not touched, and every exit path (no change, shadowed project, error) writes the file.lsv_measure_results.json.confirm_setandwrite_confirm_setinlsv_measure.pylsv/confirm_set.json(ids,confirm_count,theta_source) and gets the number of rounds (rounds, default 6; 0 when the set is empty).test.shconfirmation pass/tests/confirm.jsonexists and the set is not empty, inside the measure gate.test_timings.jsongetslsv_confirm_s.tests/docker/test_lsv_confirm_pass.py--onlyfilter and output (with asv stubbed), and the test.sh block (gate around the call, no rebuild, no pass withoutconfirm.jsonor with an empty set). It also checks thatsetup.sh,lsv_init.pyandenvironment/never nameconfirm.json, and that the only copy out of/testsinsetup.shisrebuild.sh.Usage
Nothing changes without
tests/confirm.json. The survey pipeline writes it whenCONFIRM_ROUNDS > 0(formula-code/formulacode-verified-rl, branchfeat/confirm-json). By hand:Verification
tests/docker: 431 passed, 10 skipped (includes the 10 new tests).lsv_measure.pyloads and runsconfirm_set/only_nameson Python 3.8 (uv run --no-project --python 3.8).Notes
confirm.jsonnames the oracle's improved benchmarks. Harbor removes/testsbefore a non-oracle agent runs and uploadstests/again for the verifier, so only test.sh reads it, at verification.setup.shcopies onlyrebuild.shout of/tests; the new test checks this.parser.pyand the reward still read only the first pass. Combining the passes is up to the grader.python /tests/lsv_measure.pycall withtrueand required exactly 2 replacements. This template has 3; the pipeline branch above accepts 2 or more.