From 731f5521184a5c225036e9e7bf1bc4d2513d2871 Mon Sep 17 00:00:00 2001 From: Nathan Walker Date: Sat, 26 Sep 2026 10:48:52 -0400 Subject: [PATCH] Plan: #170 merged, refresh slice 1 shipped (I066) Co-Authored-By: Claude Opus 5.5 --- benchmarks/CAMPAIGN_PLAN.md | 10 +++++++--- benchmarks/REFRESH_PLAN.md | 1 + 2 files changed, 8 insertions(+), 3 deletions(-) diff --git a/benchmarks/CAMPAIGN_PLAN.md b/benchmarks/CAMPAIGN_PLAN.md index d67c77d..15fc4ae 100644 --- a/benchmarks/CAMPAIGN_PLAN.md +++ b/benchmarks/CAMPAIGN_PLAN.md @@ -29,7 +29,7 @@ A shipped preset is a frontier win but not a default win. - Delivery is PRs only (the maintainer, 2026-09-18): no merges to `main`, no PyPI releases from these sessions until he says otherwise. Ships land as pull requests; merging and releasing stay his call. - AMENDED 2026-09-21 (the maintainer, in chat): "if tests come back bit identical, we can very safely merge any PRs that are just incrementing markdown files, I don't need to give input on that if it's just like hey, here's something we learned. If we are changing actual code I need a PR though." Read strictly: the loop merges its own PR (merge commit, then deletes the branch) ONLY when every changed path is a `.md` file and `chimeraboost/` + `tests/` are identical to main, so the last green suite and identity snapshot still hold. Any other file in the diff — library, tests, config, or a script under `benchmarks/` — waits for him. Releases stay his call, always. - WIDENED the same day (the maintainer, in chat, asked whether probe scripts count as code): "benchmarks can be self-mergable as well." The rule now: the loop merges its own PR when every changed path is a `.md` file or lives under `benchmarks/`, and `chimeraboost/` + `tests/` are identical to main. Still his: `chimeraboost/`, `tests/`, packaging and CI config, releases. `benchmarks/tabarena/` stays hands-off (sealed). A self-merged PR that changes how a gate SCORES (`run_benchmarks.py`, `compare_runs.py`, `synth_report.py`, `synthgen/`) or a policy file (a skill, `AGENTS.md`) says so in the step report. -- **FOCUS 2026-09-24 (the maintainer, in chat): "next i want us to focus on issues rather than pareto efficiency stuff".** After the #163 rung (I063) closes, the loop's rungs come from the open GitHub issues, ahead of every beam family; F7's Q4/Q1 and the other ACTIVE families are parked until he says otherwise. Order (bugs first, per his standing preference): (1) #84, `warmup(background=True)` sets its notice flag inside the thread, so a fit in that window still prints the cold-compile notice; set it in the caller before `start()` (library, his merge). (2) #81, `benchmarks/research/`: the cascade self-test's anchors are no longer off by default (`linear_leaves` is auditioned, `early_stopping_rounds` moves the curve) and several `ideas.py` entries set flags removed on 2026-06-15 (benchmarks-only, self-merge). (3) #131, `refresh(X, y)` with an opt-in stored training set, structure pinned, leaves replayed on old + new rows (library). (4) #113, random effects slice 2: slopes, a second grouping column, the entity-ID auto-route; gate `grouped_suite.py`, the auto-route on hc `--decide` (library). (5) #45, loose ends: a fresh sweep, and his call on the fully merged remote branches. #109 (random effects) is the umbrella slice 1 shipped from; #113 carries its remaining scope, so closing it is his call. The harness open items H(13) and H(14) ride along as small benchmarks-only fixes when a rung has room. UPDATE 2026-09-25: #84 and #81 are merged; #131 slice 1 is PR #170, PARKED OPEN by the maintainer ("not sure I want it implemented"), which does NOT hold the one-campaign-PR-at-a-time rule, since he asked the loop to keep going; #113: the maintainer's go 2026-09-25 ("Yes 113 is 'another issue' so you may work on it"); its four design calls go with the stated recommendations (one slope column with independent variances; the real-set covariate by the correlation rule; the entity-ID auto-route demoted to a kill-or-keep probe; fix the employee alias leak and report slice 1's old and new numbers). Next: #113 (random slopes), then #45. UPDATE 2026-09-25 (later): the queue is empty (#84, #81, #45 closed; #113, #109 closed after I067/I068; #131 = PR #170, which the maintainer is testing by hand). The maintainer: "yeah we can continue on the quantile thread" (option B): F7 is ACTIVE again, first target the head's cap-bound low-noise losses (I070). +- **FOCUS 2026-09-24 (the maintainer, in chat): "next i want us to focus on issues rather than pareto efficiency stuff".** After the #163 rung (I063) closes, the loop's rungs come from the open GitHub issues, ahead of every beam family; F7's Q4/Q1 and the other ACTIVE families are parked until he says otherwise. Order (bugs first, per his standing preference): (1) #84, `warmup(background=True)` sets its notice flag inside the thread, so a fit in that window still prints the cold-compile notice; set it in the caller before `start()` (library, his merge). (2) #81, `benchmarks/research/`: the cascade self-test's anchors are no longer off by default (`linear_leaves` is auditioned, `early_stopping_rounds` moves the curve) and several `ideas.py` entries set flags removed on 2026-06-15 (benchmarks-only, self-merge). (3) #131, `refresh(X, y)` with an opt-in stored training set, structure pinned, leaves replayed on old + new rows (library). (4) #113, random effects slice 2: slopes, a second grouping column, the entity-ID auto-route; gate `grouped_suite.py`, the auto-route on hc `--decide` (library). (5) #45, loose ends: a fresh sweep, and his call on the fully merged remote branches. #109 (random effects) is the umbrella slice 1 shipped from; #113 carries its remaining scope, so closing it is his call. The harness open items H(13) and H(14) ride along as small benchmarks-only fixes when a rung has room. UPDATE 2026-09-25: #84 and #81 are merged; #131 slice 1 is PR #170, PARKED OPEN by the maintainer ("not sure I want it implemented"), which does NOT hold the one-campaign-PR-at-a-time rule, since he asked the loop to keep going; #113: the maintainer's go 2026-09-25 ("Yes 113 is 'another issue' so you may work on it"); its four design calls go with the stated recommendations (one slope column with independent variances; the real-set covariate by the correlation rule; the entity-ID auto-route demoted to a kill-or-keep probe; fix the employee alias leak and report slice 1's old and new numbers). Next: #113 (random slopes), then #45. UPDATE 2026-09-25 (later): the queue is empty (#84, #81, #45 closed; #113, #109 closed after I067/I068; #131 = PR #170, which the maintainer is testing by hand). The maintainer: "yeah we can continue on the quantile thread" (option B): F7 is ACTIVE again, first target the head's cap-bound low-noise losses (I070). UPDATE 2026-09-26: PR #170 merged (fa577cc) at the maintainer's go-ahead; #131 stays open for slices 2-7. ## Screening ladder @@ -1281,7 +1281,10 @@ salvaging as its own PR, since it helps every default fit. the merge conflicts then i'll merge the PR". Un-parked: main merged into the branch; the conflicts were CHANGELOG (entries on both sides, all kept), this file and `REFRESH_PLAN.md` (log lines on both sides, all -kept); no code overlapped. Awaiting his merge. +kept); no code overlapped. MERGED 2026-09-26 as PR #170 (fa577cc); +branch deleted; identity snapshot 186/186 on main, no rebaseline (the +store and `refresh` are opt-in, the kernel bit-identical). Issue #131 +stays open for slices 2-7 (`REFRESH_PLAN.md`). #### I065 2026-09-24 issue #81 (research cascade: dead self-test anchor, stale `ideas.py` flags; BENCH tooling + test, pre-registered) why now: the focus rule's second issue. PR #168 merged (3e02ab1), #84 @@ -6128,7 +6131,8 @@ next: F1 S0 entry, then step-0 compute items in sequence (tests → attribution `campaign/issue113-random-slopes` (I067)?; (e) close #113 (explored: I067, I068)?; (f) two genuinely open plan items: GPU_PLAN.md Phase 0 (the "hardware and honesty check") and HIGHCARD_PLAN.md's deferred KDD98 set. - PR #170 (refresh) stays parked open by his choice. + PR #170 (refresh) stays parked open by his choice. (Merged 2026-09-26, + fa577cc.) - CLOSED 2026-09-21: remote `e2/forced-cross-features`, `method/e2-prereg`, `loop-scaffolding` and the five merged `campaign/*` rung branches deleted (each verified fully merged into main first; the four stacked rung branches carried only content-free merge commits on top of commits main already has). The local copies of the five went with them. - Stale remotes, the maintainer's call (status verified 2026-09-21). Fully merged into main, safe to delete: `bench/portable-no-openml-api`, `refactor/readable-comments`, `worktree-tabarena-030-readiness`, `issue106-predict-thresh` (its local copy is 1 commit ahead — the calibration study PR #108 carries), `record/f2-loop-20260918` (checked out in `.record-worktree`). NOT merged: `docs/attribution-humility` (5 ahead), `docs/user-focused` (4 ahead), `bbstats-patch-1` (1 ahead), `whitepaper` (PR #108, open). Local worktree branch `f2/subgate-race` still waits on the I017 kill being confirmed. diff --git a/benchmarks/REFRESH_PLAN.md b/benchmarks/REFRESH_PLAN.md index fdf227a..06c2554 100644 --- a/benchmarks/REFRESH_PLAN.md +++ b/benchmarks/REFRESH_PLAN.md @@ -136,3 +136,4 @@ Still open, for their slices: into the branch (conflicts only in CHANGELOG and two plan files; no code overlapped), so the salvage note above no longer applies. Slices 2-7 are not started. +- 2026-09-26: **slice 1 SHIPPED**, merged as PR #170 (fa577cc).