From a5945b0a0b424467c4d2b978ffc3ff1d87b274c6 Mon Sep 17 00:00:00 2001 From: Nathan Walker Date: Fri, 25 Sep 2026 19:20:17 -0400 Subject: [PATCH 1/5] Plan: I075 pre-registers Q12, the head retrains its winner on all rows (the maintainer's pick) Co-Authored-By: Claude Opus 5.5 --- benchmarks/CAMPAIGN_PLAN.md | 44 ++++++++++++++++++++++++++++++++++++- benchmarks/QUANTILE_PLAN.md | 6 +++-- 2 files changed, 47 insertions(+), 3 deletions(-) diff --git a/benchmarks/CAMPAIGN_PLAN.md b/benchmarks/CAMPAIGN_PLAN.md index 9183530..8a74f20 100644 --- a/benchmarks/CAMPAIGN_PLAN.md +++ b/benchmarks/CAMPAIGN_PLAN.md @@ -186,7 +186,7 @@ fact: 2026-09-25 | standing quantile BASE = `results/quantile-20260925-172250.js | F3 | Classifier forced-cross | KILLED 2026-09-21 (S3, I021) | gr binary engaged 10W-13L, median −0.04%: the race earns its fee on the classifier. Knob stays opt-in (PR #117), no rung-1 pin | | F5 | hc-Brier gap vs CatBoost | **SHIPPED 2026-09-22 (I046, PR #145, d14bf38): `cat_count_features` on by default** (public 5W-1L, decide hc 6W-1L, bit-identical elsewhere, chart refreshed) — earlier: I038 library form, I039 S3 PASS | the published chart (`public_pareto.png`) refresh: DONE 2026-09-24 on 0.33.0 (`results/20260924-090541.json`, `--public --seeds 3`, 22 datasets, ~4 h: the default avg rank 1.94 [1.64-2.24] against CatBoost 1.84 [1.43-2.24] and LightGBM 2.23, at a median 5.6x against 47.2x and 1.0x; `docs/benchmarks.md` prose updated; `pareto.png` re-rendered from `results/20260924-130232.json` and `docs/PROJECT_STATUS.md` synced); S7 (multiclass CTR width) is the family's next idea if picked. History: **A per-categorical count column closes 47% of CatBoost's hc edge** (I035): 4W-1L on the gap sets, +0.43% Brier median, gains ordered by cardinality, sf-police and Traffic unanimous across seeds. The gap is the encoder (CatBoost on our TS keeps none of its edge); not the prior target, Counter, permutations or quantization (TS quantization kills on big sets, +1.7–3.6% on the two small controls — a small-data pointer, parked). `cat_count_features` (opt-in, card ≥ 256, invisible to the cross and linear-leaf races; I038) on the decision tier (I039): gr 0-0-59 exact ties, the 7 hc sets without a qualifying column exact ties, the engaged 7 **6W-1L** at +0.20% median (sf-police +0.73%, Traffic +0.86% Brier; employee_salaries +2.45%, wine-reviews +0.56% RMSE), hc@time 4-0, fit ×1.09 on hc (engaged median 1.165). PR up with the flag OFF. The random-effects alternative (per-column ANOVA λ for the TS, I040) KILLED: uncapped it collapses the small controls (−3.8 / −9.6%), capped at 10 it is a flat wash and still costs kick and eucalyptus; the count column keeps evidence the shrinkage deletes. Next: the maintainer's go on /experiment S4 for the default flip; meanwhile R4 S0 | | F6 | Ordered-TS train/test moment mismatch (shortlist R1) | KILLED 2026-09-21 (S1b, I025) — closed as barrier B18 | The defect is real (rare categories over-trusted, reliability 0.63–0.70) and two transform-side fixes both went 7W-5L against a bar of 8: the gain is sf-police (9 of 9 fits, +0.29% to +0.53%) and nothing else. Nothing ships; the open door is the Counter feature, which belongs to R3 | -| F7 | Multi-quantile head (`ChimeraBoostQuantileRegressor`) | ACTIVE again 2026-09-25 (the maintainer: "we can continue on the quantile thread"; PARKED 2026-09-24 for GitHub issues; ACTIVE from 2026-09-23). Phase 1, the bench, MERGED (I056, PR #156); Q0 DONE (I057, PR #157 merged): uncapped and validation-rescaled pay, depth 6 misses on its guard alone, recentred fails. Q3 CLOSED (I058, barrier B23: a larger rate loses); depth 6 + early-stopping-row calibration PAID (31W-5L, +0.33%, 90% error 3.43 → 0.47); PR #158 merged. **Q5 PASSED (I059): the head's default is now depth 6 + `conformalize="auto"`** (gr 31W-5L, +0.33%, 90% coverage error 3.43 → 0.47); PR #159 merged, snapshot rebaselined. **Q2 CLOSED (I060)**: no probe clears the tightened gate (sign p < 0.05); the five low-noise sets need resolution (bins) on two and a squared-error location on two; PR #160 merged. **Q6 T0 PASSED (I061)**: the validation-chosen head wins 23W-7L-6T (p 0.005), recovers 99% of the oracle, and tops the 59-key frontier above CatBoost MQ (0.6006 @ 7.2× against 0.5982 @ 129×); PR #161 merged. **Q7 PASSED (I062): the audition is the head's default** (identity 177/177 with the bench arm; gr 23W-7L-6T against the Q5 default, p 0.005; CatBoost MQ now 9W-27L against us, off the frontier); PR #162 merged, **released in 0.33.0**. **Q8 T0 PASSED (I070)**: a fixed-width (S) and a scaled-residual (N) audition candidate, gr 18W-1L-17T against the Q7 default (p 7.6e-5) at 1.11x the fit; S alone and uncapped rounds fail. **Q9 PASSED (I071): S and N in the library default** (identity 177/177; CatBoost MQ 29W-7L, RigidShift 35W-1L; the 59-key chart 0.6046 @ 8.3x); merged as PR #177 (a6a90d2); flag: hc `@time` 90% coverage error 2.12 -> 4.69 (Moneyball, a pointer). **Q10 T0 (I072) DIAGNOSED**: RoNGBa's lead on visualizing_soil and SGEMM is the S/N centre model's accuracy (our grid on its mean beats it); pol is a width gap. **Q11 T0-a (I073)**: centre settings on five sets: 254 bins matches RoNGBa's lead (visualizing_soil +24%, SGEMM +7%) but costs cpu_act 2.5%; 8000 rounds +5% on visualizing_soil only; bagging +30% (ceiling, not a default). **T0-b (I073b) FAILS**: as extra candidates the 254-bin centre still costs cpu_act 2.5% (validation cannot see the overfit); 8000 rounds closed unrun (the centre is cap-bound on 1 of 59 keys). The RoNGBa line is CLOSED. **Q4 CLOSED (I074)**: the N candidate already is the spread-aware categorical encoding; `catscale` at 10k now 16.0 against CatBoost MQ's 30.9 (excess CRPS x1000). **I063 DONE (PR #167)**: RoNGBa is the NGBoost opponent (issue #163) | next: the maintainer's pick between Q1 (P16; coverage already solved by calibration, a small CRPS claim) and the no-`eval_set` refit design (`QUANTILE_PLAN.md`, Still open); until then F7 holds. Program: `QUANTILE_PLAN.md` "Campaign 2026-09-23" | +| F7 | Multi-quantile head (`ChimeraBoostQuantileRegressor`) | ACTIVE again 2026-09-25 (the maintainer: "we can continue on the quantile thread"; PARKED 2026-09-24 for GitHub issues; ACTIVE from 2026-09-23). Phase 1, the bench, MERGED (I056, PR #156); Q0 DONE (I057, PR #157 merged): uncapped and validation-rescaled pay, depth 6 misses on its guard alone, recentred fails. Q3 CLOSED (I058, barrier B23: a larger rate loses); depth 6 + early-stopping-row calibration PAID (31W-5L, +0.33%, 90% error 3.43 → 0.47); PR #158 merged. **Q5 PASSED (I059): the head's default is now depth 6 + `conformalize="auto"`** (gr 31W-5L, +0.33%, 90% coverage error 3.43 → 0.47); PR #159 merged, snapshot rebaselined. **Q2 CLOSED (I060)**: no probe clears the tightened gate (sign p < 0.05); the five low-noise sets need resolution (bins) on two and a squared-error location on two; PR #160 merged. **Q6 T0 PASSED (I061)**: the validation-chosen head wins 23W-7L-6T (p 0.005), recovers 99% of the oracle, and tops the 59-key frontier above CatBoost MQ (0.6006 @ 7.2× against 0.5982 @ 129×); PR #161 merged. **Q7 PASSED (I062): the audition is the head's default** (identity 177/177 with the bench arm; gr 23W-7L-6T against the Q5 default, p 0.005; CatBoost MQ now 9W-27L against us, off the frontier); PR #162 merged, **released in 0.33.0**. **Q8 T0 PASSED (I070)**: a fixed-width (S) and a scaled-residual (N) audition candidate, gr 18W-1L-17T against the Q7 default (p 7.6e-5) at 1.11x the fit; S alone and uncapped rounds fail. **Q9 PASSED (I071): S and N in the library default** (identity 177/177; CatBoost MQ 29W-7L, RigidShift 35W-1L; the 59-key chart 0.6046 @ 8.3x); merged as PR #177 (a6a90d2); flag: hc `@time` 90% coverage error 2.12 -> 4.69 (Moneyball, a pointer). **Q10 T0 (I072) DIAGNOSED**: RoNGBa's lead on visualizing_soil and SGEMM is the S/N centre model's accuracy (our grid on its mean beats it); pol is a width gap. **Q11 T0-a (I073)**: centre settings on five sets: 254 bins matches RoNGBa's lead (visualizing_soil +24%, SGEMM +7%) but costs cpu_act 2.5%; 8000 rounds +5% on visualizing_soil only; bagging +30% (ceiling, not a default). **T0-b (I073b) FAILS**: as extra candidates the 254-bin centre still costs cpu_act 2.5% (validation cannot see the overfit); 8000 rounds closed unrun (the centre is cap-bound on 1 of 59 keys). The RoNGBa line is CLOSED. **Q4 CLOSED (I074)**: the N candidate already is the spread-aware categorical encoding; `catscale` at 10k now 16.0 against CatBoost MQ's 30.9 (excess CRPS x1000). **I063 DONE (PR #167)**: RoNGBa is the NGBoost opponent (issue #163) | next: **Q12 (I075), the maintainer's pick 2026-09-25 ("1 is fine")**: retrain the audition's winner on all rows, as `refit_full` does for the point model; bench probe arms, then the decide run. Q1 (P16) stays queued behind it. Program: `QUANTILE_PLAN.md` "Campaign 2026-09-23" | ### F1 — Cross-feature cost trim v2 status: KILLED 2026-08-16 at S2 (I007 at k=6, I008 at k=12) — closed as barrier B16 @@ -453,6 +453,48 @@ Recommended pick: **R1, R2, R3, R4 + H(1)(4)(5)**. R1 and R2 have free probes an ## Iteration log (append-only) +#### I075 2026-09-25 F7 Q12 T0 (the head retrains its winner on all rows after the audition, as the point regressor's `refit_full` does; BENCH-only probe arms, pre-registered) +why now: the maintainer, 2026-09-25, picking option 1 of three: "1 is +fine". Called without an `eval_set`, the head holds back +`validation_fraction` of the rows and never trains on them; the point +regressor retrains on 100% after early stopping (`refit_full`). Checked +first (`scratchpad/carve_check.py`, visualizing_soil and pol, seed 0): +without an `eval_set` both the regressor and the head carve exactly the +suite's shared split (predictions bit-identical to the `eval_set` fits), +so today's field arm IS the library called without an `eval_set`, and the +regressor's own retrain cuts its test RMSE 4.8% and 5.1% there. +barriers: B2 (the rung-3 refit amplifies a bad audition) is about the +point model's audition budget; here the choice is untouched and only the +winner is retrained. B13 and the `refit_full` docstring's reason for +skipping `loss="Quantile"` (keep the conformal holdout honest): every +calibration quantity (factors, offsets, `q_s`, `q_n`, the spread model) +stays from the 80% fit, computed on rows it never trained on; only the +delivered centre (and in arm (b) the head) is retrained, so intervals may +run slightly wide, which the coverage guard reads. +arms (`quantile_suite.py`, muse task `20260925-quantile-q12-refit-probe.md`): +(a) `ChimeraBoostQuantileRefitCentre`: `_fit_audition(with_s, with_n)` +unchanged on the shared split; when R, S or N wins, its centre is replaced +by the same `ChimeraBoostRegressor` fitted on all training rows without an +`eval_set` (it carves the same split and retrains with its default replay +refit). Offsets, spread model, `q_s`, `q_n` and factors stay. H and B +winners are unchanged. (b) `ChimeraBoostQuantileRefitAll`: (a), plus H, B +and R's head retrained from scratch on all rows with early stopping off, +`n_estimators = min(ceil(t_star / 0.8), 2000)` (`_refit_on_full`'s rule, +t_star the trees the early-stopped head kept), rate pinned, the 80% fit's +factors applied. fit_s counts only what the library would add (the +retrain, not a repeat of the 80% fit). +forecast: (1) (a) engages where R, S or N wins (about half the gr fits), +wins >= 70% of the engaged sets (median +1% to +3%), gr sign p < 0.05, +coverage guard ok, fit +5% to +10%. (2) (b) engages everywhere, gr >= 26 +of 36 (median +1% to +2%), p < 0.05, fit +25% to +45%. (3) visualizing_soil +>= +5% and pol >= +3% under (a). (4) hc no worse than gr (a pointer). +decision: an arm passing the default gate (gr two-sided sign p < 0.05, +median > 0, coverage guard ok) becomes the library rung Q13 (the Q7/Q9 +pattern, the arm its exact oracle). Both pass: (b) only if it also beats +(a) head to head on gr at p < 0.05; otherwise (a). Neither: the question +closes. +verdict: PENDING(muse bench task, then the decide run) + #### I074 2026-09-25 F7 Q4 T0-a (does the five-candidate default still lose `catscale` to CatBoost? a synthetic recheck before any spread-TS work; bench-only, pre-registered) why now: Q4's premise (2026-09-23, `results/quantile-synth-20260923-090818.json`) is that the ordered TS is a per-category MEAN, blind to a category that diff --git a/benchmarks/QUANTILE_PLAN.md b/benchmarks/QUANTILE_PLAN.md index e0499d6..a3846e2 100644 --- a/benchmarks/QUANTILE_PLAN.md +++ b/benchmarks/QUANTILE_PLAN.md @@ -541,8 +541,10 @@ offset on real data (a CRPS tie), which made it Q0's subject. refit of the winner after the audition, keeping the calibration taken before it, is the candidate. `quantile_suite.py` cannot measure it: it passes the shared split as an `eval_set`, which the head must not train - on. Needs a no-`eval_set` protocol first; the maintainer's call whether - to open it. + on. Needs a no-`eval_set` protocol first. **OPENED 2026-09-25 as Q12** + (the maintainer picked it over Q1): the head's own carve equals the + suite's shared split bit for bit, so the probe is a retrain added to + today's arm (`CAMPAIGN_PLAN.md` I075). - RESOLVED 2026-09-23 (Q5, I059): `docs/quantiles.md` "How it compares" is re-measured against the new default, with the fixed-width baseline and NGBoost added. From 7c2c9e93bccfe13dcbc6d8bf979a94283bc52121 Mon Sep 17 00:00:00 2001 From: Nathan Walker Date: Fri, 25 Sep 2026 19:58:27 -0400 Subject: [PATCH 2/5] Quantile head: opt-in refit_full, and probe arms that measure it (Q12, I075) ChimeraBoostQuantileRegressor gains refit_full=False. When True and the fit used the automatic early-stopping split (no user eval_set, early stopping on, conformalize "auto" or False), the winner is retrained on all rows after the audition finishes: an R/S/N winner's centre becomes the same ChimeraBoostRegressor fitted on all rows without an eval_set (its own carve, identical to the audition's, and its own replay refit); the head booster of an H/B/R winner is retrained from scratch at min(ceil(t_star / 0.8), n_estimators) rounds with the rate pinned. Calibration factors, offsets, residual quantiles and the spread model stay from the held-out fit. refit_ records what was retrained. Off by default: the identity snapshot is 186/186. quantile_suite.py passes each arm the original training rows (split.full) and adds the probe arms ChimeraBoostQuantileRefitCentre and ChimeraBoostQuantileRefitAll; quantile_synth.py passes the rows too. 11 new tests; 1235 passed. Co-Authored-By: Claude Opus 5.5 --- benchmarks/quantile_suite.py | 72 ++++++++++++++- benchmarks/quantile_synth.py | 3 +- chimeraboost/quantile_api.py | 133 +++++++++++++++++++++++++++- tests/test_quantile_head.py | 167 +++++++++++++++++++++++++++++++++++ tests/test_quantile_shap.py | 29 ++++++ tests/test_quantile_suite.py | 40 ++++++++- 6 files changed, 438 insertions(+), 6 deletions(-) diff --git a/benchmarks/quantile_suite.py b/benchmarks/quantile_suite.py index 6dba34d..6b1c995 100644 --- a/benchmarks/quantile_suite.py +++ b/benchmarks/quantile_suite.py @@ -96,6 +96,24 @@ RONGBA_LEAVES = 31 +class _SplitWithFull(tuple): + """The shared ``(Xf, Xv, yf, yv)`` split, carrying the training rows. + + Unpacks and indexes as the plain 4-tuple every arm already takes; + ``.full`` is the ``(Xtr, ytr)`` the shared split was carved from, in + their original order, so a probe arm can fit on all training rows + while the library's own carve reproduces the shared split. + """ + + def __new__(cls, split, full): + obj = super().__new__(cls, split) + obj.full = full + return obj + + def __init__(self, split, full): + pass + + def _fit_head_model(split, cat, threads, taus, n_estimators, **params): """Construct and fit the suite's head. @@ -811,6 +829,55 @@ def _fit_chimera_default_uncapped(split, Xte, cat, threads, taus): return Q, fit_s, time.time() - t, m.best_iteration_ +_AUDITION_CHOICE = {"head": 0, "bins": 1, "recentred": 2, + "fixed": 3, "scaled": 4} + + +def _refit_extras(m): + """The refit probes' extras: the audition choice as the audition arms + encode it (0-4 for H, B, R, S, N), whether the centre and the head + were retrained, and the head's retrained round count.""" + aud = m.audition_ or {} + refit = m.refit_ or {} + return {"audition_choice": _AUDITION_CHOICE.get(aud.get("selected")), + "refit_centre": bool(refit.get("centre", False)), + "refit_head": bool(refit.get("head", False)), + "refit_rounds": refit.get("rounds")} + + +def _fit_chimera_refit(split, Xte, cat, threads, taus, scope): + """The Q12 probe: the head fitted on ALL training rows with + ``refit_full=True`` -- the audition runs on the library's own carve, + which reproduces the shared split, and the winner is then retrained + on every row. ``scope="centre"`` retrains only the R/S/N centre.""" + Xtr, ytr = split.full + m = ChimeraBoostQuantileRegressor( + quantiles=taus, n_estimators=rb.MAX_ITERS, + early_stopping_rounds=rb.PATIENCE, thread_count=threads, + random_state=0, refit_full=True) + if scope == "centre": + m._refit_scope = "centre" + t = time.time() + m.fit(Xtr, ytr, cat_features=cat or None) + fit_s = time.time() - t + t = time.time() + Q = m.predict(Xte) + pred_s = time.time() - t + return Q, fit_s, pred_s, m.best_iteration_, _refit_extras(m) + + +def _fit_chimera_refit_all(split, Xte, cat, threads, taus): + """The head retrained on all rows: centre and head alike (Q12 arm b).""" + return _fit_chimera_refit(split, Xte, cat, threads, taus, scope="all") + + +def _fit_chimera_refit_centre(split, Xte, cat, threads, taus): + """The head retrained on all rows: the R/S/N centre only, every head + booster untouched (Q12 arm a).""" + return _fit_chimera_refit(split, Xte, cat, threads, taus, + scope="centre") + + PROBES = { "ChimeraBoostQuantileUncapped": _fit_chimera_uncapped, "ChimeraBoostQuantileDepth6": _fit_chimera_depth6, @@ -829,6 +896,8 @@ def _fit_chimera_default_uncapped(split, Xte, cat, threads, taus): "ChimeraBoostQuantileAuditionS": _fit_chimera_audition_s, "ChimeraBoostQuantileAuditionSN": _fit_chimera_audition_sn, "ChimeraBoostQuantileDefaultUncapped": _fit_chimera_default_uncapped, + "ChimeraBoostQuantileRefitAll": _fit_chimera_refit_all, + "ChimeraBoostQuantileRefitCentre": _fit_chimera_refit_centre, } @@ -944,7 +1013,8 @@ def run_one(ds_name, seed, taus, threads, models): "regression") # One early-stopping split, shared by every arm, so no model is judged on # more data than another. Same carve the harness uses. - split = rb._val_split(Xtr, ytr, "regression", 0) + split = _SplitWithFull(rb._val_split(Xtr, ytr, "regression", 0), + (Xtr, ytr)) meta = {"task": "quantile", "n_train": int(Xtr.shape[0]), "n_total": int(X.shape[0]), "n_features": int(X.shape[1]), diff --git a/benchmarks/quantile_synth.py b/benchmarks/quantile_synth.py index 77fb8bb..72bacb5 100644 --- a/benchmarks/quantile_synth.py +++ b/benchmarks/quantile_synth.py @@ -219,7 +219,8 @@ def run_one(regime, n, seed, taus, threads, models): random_state=seed) Xtr, Xte, ytr, yte = X[idx_tr], X[idx_te], y[idx_tr], y[idx_te] Qo_te = Q_or[idx_te] - split = rb._val_split(Xtr, ytr, "regression", 0) + split = qs._SplitWithFull(rb._val_split(Xtr, ytr, "regression", 0), + (Xtr, ytr)) cat = [X.shape[1] - 1] if regime == "catscale" else None oracle_crps = float(qm.crps(yte, Qo_te, taus)) diff --git a/chimeraboost/quantile_api.py b/chimeraboost/quantile_api.py index 3b6646f..f0643d7 100644 --- a/chimeraboost/quantile_api.py +++ b/chimeraboost/quantile_api.py @@ -2,8 +2,8 @@ Kept out of ``sklearn_api`` because it shares that module's input validation but almost none of its fit machinery: no loss family, no linear leaves, no -cross features, no bagging, no full-data refit. It imports the validation -helpers and keeps the same flat, module-function style. +cross features, no bagging, and only an opt-in full-data refit. It imports +the validation helpers and keeps the same flat, module-function style. """ import warnings @@ -265,6 +265,15 @@ def _pick_audition_winner(ordered): return best_name +def _keep_es_state(new_booster, old_booster): + """Keep the early-stopped fit's visible state on a refit replacement, + as ``sklearn_api._refit_on_full`` does: the training and validation + curves, and the budget early stopping chose.""" + new_booster.train_history_ = old_booster.train_history_ + new_booster.valid_history_ = old_booster.valid_history_ + new_booster.best_iteration_ = old_booster.best_iteration_ + + def _phi_centre(phi, mi, mw): """Centre of attributions along the level axis, mirroring ``_centre``.""" if mw == 0.0: @@ -345,6 +354,25 @@ class ChimeraBoostQuantileRegressor(BaseEstimator): ``model_`` stays the fitted head booster whatever wins (the 254-bin booster for a bins win); for S and N it delivers nothing -- ``predict`` serves the centre/spread models -- but stays inspectable. + refit_full : bool, default False + Retrain the winner on all rows after early stopping chose the + budget, as ``ChimeraBoostRegressor`` does. Acts only when the fit + used the automatic early-stopping split (no user ``eval_set``, + early stopping on, the carve succeeded) and ``conformalize`` is + ``"auto"`` or False; with ``conformalize=True``, a user + ``eval_set``, or no carve the fit is exactly what it is without + it. An R, S or N winner's centre is then replaced by a default + ``ChimeraBoostRegressor`` fitted on all rows without an + ``eval_set`` -- offsets, spread model, residual quantiles, floor + and calibration factors stay from the audition fit -- and the + head booster of an H, B or R winner is retrained from scratch on + all rows with early stopping off, at ``min(ceil(t_star / (1 - + validation_fraction)), n_estimators)`` rounds with the + early-stopped learning rate pinned, where t_star is the rounds + the early-stopped head kept. ``audition_``, + ``conformal_scale_``, ``best_iteration_`` and + ``validation_history_`` keep the early-stopped fit's values; + ``refit_`` records what was retrained. Attributes ---------- @@ -364,6 +392,12 @@ class ChimeraBoostQuantileRegressor(BaseEstimator): "crps": {name: score}}`` for the candidates that ran, or ``None`` when no audition ran (no evaluation rows, ``conformalize`` True or False, or ``audition=False``). + refit_ : dict or None + ``{"centre": bool, "head": bool, "rounds": int or None}`` -- + whether the centre and the head booster were retrained on all + rows, and the head's retrained round count (None when the head + was not retrained). None when nothing was retrained + (``refit_full=False``, or a fit the refit does not cover). Notes ----- @@ -383,7 +417,8 @@ def __init__(self, quantiles=None, n_estimators=2000, learning_rate=None, quantize_gradients=True, early_stopping=True, validation_fraction=0.2, split_projection="rotate", exact_splits=False, conformalize="auto", - calibration_fraction=0.2, audition=True): + calibration_fraction=0.2, audition=True, + refit_full=False): self.quantiles = quantiles self.n_estimators = n_estimators self.learning_rate = learning_rate @@ -409,6 +444,11 @@ def __init__(self, quantiles=None, n_estimators=2000, learning_rate=None, self.conformalize = conformalize self.calibration_fraction = calibration_fraction self.audition = audition + self.refit_full = refit_full + # Probe-only scope, set after construction (never a constructor + # parameter): "centre" retrains only the R/S/N centre and leaves + # every head booster untouched. + self._refit_scope = "all" def __sklearn_is_fitted__(self): return hasattr(self, "model_") @@ -467,6 +507,82 @@ def _carve_es_split(self, X, y, sample_weight, eval_set, groups): es_rounds = 50 return X, y, sample_weight, eval_set, es_rounds + def _refit_full_active(self, auto_split): + """True when the full-data refit fires: asked for, on the automatic + split, with no honest holdout at stake (``conformalize=True`` keeps + its pristine fold instead).""" + return (bool(self.refit_full) and bool(auto_split) + and self.conformalize is not True) + + def _maybe_refit_full(self, auto_split, taus, es_rounds, cat_features, + X_full, y_full, sw_full, groups_full): + """Retrain the winner on all rows, after the audition (or the single + head's calibration) finished exactly as without the refit. The + centre of an R, S or N winner is replaced by a default + ``ChimeraBoostRegressor`` fitted on the full rows without an + ``eval_set``; the head booster of an H, B or R winner -- or of the + single head when no audition ran -- is retrained from scratch with + early stopping off. Records what happened in ``refit_``.""" + if not self._refit_full_active(auto_split): + return + selected = self._audition_selected() + centre_done = False + if selected in ("recentred", "fixed", "scaled"): + self._refit_audition_centre(es_rounds, cat_features, X_full, + y_full, sw_full, groups_full) + centre_done = True + head_done, rounds = False, None + if self._refit_scope != "centre" and selected in ( + None, "head", "bins", "recentred"): + rounds = self._refit_head_booster(taus, selected, cat_features, + X_full, y_full, sw_full) + head_done = True + if not centre_done and not head_done: + return + self.refit_ = {"centre": centre_done, "head": head_done, + "rounds": rounds} + + def _refit_audition_centre(self, es_rounds, cat_features, X_full, + y_full, sw_full, groups_full): + """Replace the R/S/N centre with the same regressor fitted on the + full rows without an ``eval_set``, so it carves the same split and + retrains with its own default refit. Offset, spread model, + residual quantiles, floor and factors stay from the audition fit; + the early-stopped curves stay on the replacement, as the + regressor's own refit keeps them.""" + from .sklearn_api import ChimeraBoostRegressor + old = self._centre_model_ + centre = ChimeraBoostRegressor( + n_estimators=self.n_estimators, + early_stopping_rounds=es_rounds, + thread_count=self.thread_count, random_state=self.random_state) + centre.fit(X_full, y_full, cat_features=cat_features, + sample_weight=sw_full, groups=groups_full) + _keep_es_state(centre.model_, old.model_) + self._centre_model_ = centre + + def _refit_head_booster(self, taus, selected, cat_features, X_full, + y_full, sw_full): + """Retrain the winner's head booster from scratch on the full rows: + the same construction (254 bins for a bins win), early stopping + off, ``min(ceil(t_star / (1 - validation_fraction)), + n_estimators)`` rounds with the early-stopped learning rate pinned. + Returns the retrained round count.""" + winner = self.model_ + t_star = len(winner.trees_) + frac = max(1.0 - float(self.validation_fraction), 1e-9) + rounds = min(int(np.ceil(t_star / frac)), int(self.n_estimators)) + booster = self._make_mq_booster( + taus, None, cat_features, len(y_full), + max_bins=254 if selected == "bins" else None) + booster.n_estimators = rounds + booster.learning_rate = float(winner.lr_) + booster.fit(X_full, y_full, cat_features=cat_features, + sample_weight=sw_full) + _keep_es_state(booster, winner) + self.model_ = booster + return rounds + def _make_mq_booster(self, taus, es_rounds, cat_features, n_rows, max_bins=None): """Resolve the auto defaults and construct the booster.""" @@ -530,6 +646,7 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None, self._median_idx_ = _median_index(taus) self.conformal_scale_ = np.ones(taus.shape[0]) self.audition_ = None + self.refit_ = None self._centre_model_ = None self._centre_off_ = None self._fixed_q_ = None @@ -544,8 +661,15 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None, cal, X, y, sample_weight, groups = self._carve_calibration_fold( X, y, sample_weight, groups) + # Kept for the optional full-data refit below: the auto split + # reassigns X/y, but the refit retrains on every row. + X_full, y_full, sw_full, groups_full = X, y, sample_weight, groups + had_user_eval = eval_set is not None X, y, sample_weight, eval_set, es_rounds = self._carve_es_split( X, y, sample_weight, eval_set, groups) + # The carve is the only path that sets eval_set when the user did + # not, so this is True exactly when the automatic split was used. + auto_split = not had_user_eval and eval_set is not None self.model_ = self._make_mq_booster(taus, es_rounds, cat_features, len(y)) @@ -573,6 +697,9 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None, except ValueError: pass # uncertifiable: keep the raw grid, never raise + self._maybe_refit_full(auto_split, taus, es_rounds, cat_features, + X_full, y_full, sw_full, groups_full) + return self def _wants_audition(self, eval_set): diff --git a/tests/test_quantile_head.py b/tests/test_quantile_head.py index 61111ae..d12348e 100644 --- a/tests/test_quantile_head.py +++ b/tests/test_quantile_head.py @@ -1051,3 +1051,170 @@ def test_audition_fixed_and_scaled_keep_the_head_booster(): raw = c[:, None] + s[:, None] * np.asarray( m._scaled_q_, dtype=np.float64)[None, :] np.testing.assert_array_equal(m.predict(Xt), _hand_rescaled(raw, m)) + + +# -------------------------------------------------------------------------- +# The full-data refit (Q12): off by default, audition-preserving. +# -------------------------------------------------------------------------- + + +def _aud_nosplit_fit(kind, scope="all", **kw): + """Fit on the kind's training rows WITHOUT an eval_set, so the library + carves its own split; return (model, Xtr, ytr, Xte, yte).""" + from sklearn.model_selection import train_test_split + X, y, sseed = _aud_data(kind) + Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, + random_state=sseed) + m = ChimeraBoostQuantileRegressor( + quantiles=_AUD_TAUS, n_estimators=300, early_stopping_rounds=50, + thread_count=1, random_state=0, **kw) + m._refit_scope = scope + return m.fit(Xtr, ytr), Xtr, ytr, Xte, yte + + +def test_no_eval_set_equals_suite_split_fit(): + """Without an eval_set the head carves the suite's own split: bit for + bit the same fit as training on the carved rows with that eval_set.""" + from sklearn.model_selection import train_test_split + rng = np.random.default_rng(60) + n = 1500 + num = rng.standard_normal((n, 3)) + cat = rng.integers(0, 6, n).astype(object) + X = np.column_stack([num, cat]) + y = (2.0 * num[:, 0] + (cat.astype(float) - 2.5) + + rng.standard_normal(n)) + Xf, Xv, yf, yv = train_test_split(X, y, test_size=0.2, random_state=0) + kw = dict(quantiles=_AUD_TAUS, n_estimators=300, + early_stopping_rounds=50, thread_count=1, random_state=0) + m_auto = ChimeraBoostQuantileRegressor(**kw).fit(X, y, cat_features=[3]) + m_split = ChimeraBoostQuantileRegressor(**kw).fit( + Xf, yf, cat_features=[3], eval_set=(Xv, yv)) + assert m_auto.audition_ is not None + assert m_auto.audition_ == m_split.audition_ + assert np.array_equal(m_auto.conformal_scale_, m_split.conformal_scale_) + assert np.array_equal(m_auto.predict(X), m_split.predict(X)) + + +def test_refit_full_false_equals_default(): + """refit_full=False is today's fit, exactly; the flag is a visible + constructor parameter defaulting to False.""" + assert (ChimeraBoostQuantileRegressor().get_params()["refit_full"] + is False) + X, y = _heteroscedastic(n=1200, seed=61) + kw = dict(random_state=0, n_estimators=100, thread_count=1) + m = ChimeraBoostQuantileRegressor(**kw).fit(X, y) + r = ChimeraBoostQuantileRegressor(refit_full=False, **kw).fit(X, y) + assert m.refit_ is None + assert r.refit_ is None + assert m.audition_ == r.audition_ + assert np.array_equal(m.conformal_scale_, r.conformal_scale_) + assert np.array_equal(m.predict(X), r.predict(X)) + + +def test_refit_full_needs_auto_split_and_no_honest_fold(): + """With a user eval_set, conformalize=True, or early stopping off, the + refit stays out: refit_full=True is exactly refit_full=False.""" + X, y = _heteroscedastic(n=1200, seed=62) + Xt, yt, Xv, yv = X[:900], y[:900], X[900:], y[900:] + base = dict(random_state=0, n_estimators=100, thread_count=1) + a = ChimeraBoostQuantileRegressor(**base).fit(Xt, yt, eval_set=(Xv, yv)) + b = ChimeraBoostQuantileRegressor(refit_full=True, **base).fit( + Xt, yt, eval_set=(Xv, yv)) + assert b.refit_ is None + assert a.audition_ == b.audition_ + assert np.array_equal(a.predict(Xt), b.predict(Xt)) + a = ChimeraBoostQuantileRegressor(conformalize=True, **base).fit(X, y) + b = ChimeraBoostQuantileRegressor( + conformalize=True, refit_full=True, **base).fit(X, y) + assert b.refit_ is None + assert np.array_equal(a.predict(X), b.predict(X)) + a = ChimeraBoostQuantileRegressor(early_stopping=False, **base).fit(X, y) + b = ChimeraBoostQuantileRegressor( + early_stopping=False, refit_full=True, **base).fit(X, y) + assert b.refit_ is None + assert np.array_equal(a.predict(X), b.predict(X)) + + +_REFIT_WANT = { + "head": (False, True), "bins": (False, True), + "recentred": (True, True), "fixed": (True, False), + "scaled": (True, False), +} + + +def test_refit_full_records_and_keeps_es_values(): + """Per winner: refit_ says what was retrained, at the replay round + count; audition_, factors, best round and curve keep the 80% values.""" + for kind in ("head", "bins", "recentred", "fixed", "scaled"): + mf, _, _, _, _ = _aud_nosplit_fit(kind) + mt, _, _, _, _ = _aud_nosplit_fit(kind, refit_full=True) + assert mf.audition_["selected"] == kind + assert mt.audition_ == mf.audition_ + want_c, want_h = _REFIT_WANT[kind] + assert mt.refit_["centre"] is want_c + assert mt.refit_["head"] is want_h + if want_h: + t_star = len(mf.model_.trees_) + want_rounds = min(int(np.ceil(t_star / 0.8)), 300) + assert mt.refit_["rounds"] == want_rounds + assert len(mt.model_.trees_) == want_rounds + else: + assert mt.refit_["rounds"] is None + assert np.array_equal(mt.conformal_scale_, mf.conformal_scale_) + assert mt.best_iteration_ == mf.best_iteration_ + assert mt.validation_history_ == mf.validation_history_ + + +def test_refit_full_centre_matches_full_row_regressor(): + """The delivered R/S/N centre is the same regressor fitted on all rows; + offsets, residual quantiles, floor and spread stay from the audition.""" + for kind in ("recentred", "fixed", "scaled"): + mt, Xtr, ytr, Xte, _ = _aud_nosplit_fit(kind, refit_full=True) + mf, _, _, _, _ = _aud_nosplit_fit(kind) + ref = ChimeraBoostRegressor( + n_estimators=300, early_stopping_rounds=50, thread_count=1, + random_state=0).fit(Xtr, ytr) + assert np.array_equal(mt._centre_model_.predict(Xte), + ref.predict(Xte)) + assert mt._centre_off_ == mf._centre_off_ + if kind == "fixed": + assert np.array_equal(mt._fixed_q_, mf._fixed_q_) + if kind == "scaled": + assert np.array_equal(mt._scaled_q_, mf._scaled_q_) + assert mt._spread_floor_ == mf._spread_floor_ + assert np.array_equal(mt._spread_model_.predict(Xte), + mf._spread_model_.predict(Xte)) + + +def test_refit_full_paths_serve_retrained_models(): + """Ordered grids, staged final equals predict, pickle round-trips.""" + for kind in ("head", "bins", "recentred", "fixed", "scaled"): + mt, _, _, Xte, _ = _aud_nosplit_fit(kind, refit_full=True) + Q = mt.predict(Xte) + assert np.all(np.diff(Q, axis=1) >= 0) + Xt = Xte[:50] + stages = list(mt.staged_predict(Xt)) + assert np.array_equal(stages[-1], mt.predict(Xt)) + back = pickle.loads(pickle.dumps(mt)) + assert np.array_equal(back.predict(Xte), Q) + assert back.refit_ == mt.refit_ + + +def test_refit_scope_centre_leaves_heads_untouched(): + """With _refit_scope = "centre", H and B winners come back exactly as + without the refit; R keeps its head and retrains only its centre.""" + for kind in ("head", "bins"): + mf, _, _, Xte, _ = _aud_nosplit_fit(kind) + mc, _, _, _, _ = _aud_nosplit_fit(kind, scope="centre", + refit_full=True) + assert mc.refit_ is None + assert np.array_equal(mc.predict(Xte), mf.predict(Xte)) + mr, Xtr, ytr, Xte, _ = _aud_nosplit_fit("recentred", scope="centre", + refit_full=True) + mh, _, _, _, _ = _aud_nosplit_fit("recentred") + assert mr.refit_ == {"centre": True, "head": False, "rounds": None} + assert np.array_equal(mr.model_.predict_raw(Xte), + mh.model_.predict_raw(Xte)) + ref = ChimeraBoostRegressor(n_estimators=300, early_stopping_rounds=50, + thread_count=1, random_state=0).fit(Xtr, ytr) + assert np.array_equal(mr._centre_model_.predict(Xte), ref.predict(Xte)) diff --git a/tests/test_quantile_shap.py b/tests/test_quantile_shap.py index 19a5c6a..28bcce4 100644 --- a/tests/test_quantile_shap.py +++ b/tests/test_quantile_shap.py @@ -322,3 +322,32 @@ def test_audition_winner_shap_is_locally_accurate(): assert not np.any(phiw), "fixed width attribution is zero" assert np.allclose(phiw[rows].sum(axis=1) + m.expected_value_, iv[rows, 1] - iv[rows, 0]) + + +def _aud_nosplit_shap(kind): + """The kind's training rows without an eval_set, refit on; (model, Xte).""" + from sklearn.model_selection import train_test_split + taus = [0.05, 0.25, 0.5, 0.75, 0.95] + X, y, sseed = _aud_data_shap(kind) + Xtr, Xte, ytr, _ = train_test_split(X, y, test_size=0.25, + random_state=sseed) + m = ChimeraBoostQuantileRegressor( + quantiles=taus, n_estimators=300, early_stopping_rounds=50, + thread_count=1, random_state=0, refit_full=True).fit(Xtr, ytr) + return m, Xte + + +def test_refit_full_shap_is_locally_accurate(): + """SHAP reconstructs the RETRAINED winner per level, for H/B/R/S/N.""" + for kind in ("head", "bins", "recentred", "fixed", "scaled"): + m, Xte = _aud_nosplit_shap(kind) + assert m.audition_["selected"] == kind + Xt = Xte[:20] + rows = slice(None) + if kind == "scaled": + rows = _unfloored_rows(m, Xt) + assert rows.any(), "no unfloored rows; test is moot" + phi = m.shap_values(Xt) + assert phi.shape == (20, Xt.shape[1], 5) + assert np.allclose(phi[rows].sum(axis=1) + m.expected_value_, + m.predict(Xt)[rows]) diff --git a/tests/test_quantile_suite.py b/tests/test_quantile_suite.py index 27c0363..b8247d8 100644 --- a/tests/test_quantile_suite.py +++ b/tests/test_quantile_suite.py @@ -233,7 +233,7 @@ def test_q0_probes_are_opt_in(monkeypatch): import quantile_synth as qsyn assert len(qs.ARMS) == 7 - assert len(qs.PROBES) == 17 + assert len(qs.PROBES) == 19 _stub_registry(monkeypatch, ["gr:reg_num/houses"], {"gr:reg_num/houses": "regression"}) @@ -686,3 +686,41 @@ def test_q9_audition_sn_library_matches_bench_oracle_bit_for_bit(monkeypatch): # records the spread model's round and the library the centre's. if kind != "scaled": assert m.best_iteration_ == best_b + + +def test_split_with_full_leaves_existing_arms_unchanged(): + """The .full-carrying split unpacks as the plain 4-tuple: an existing + arm's grid and stopping round are identical either way.""" + split, Xte, taus = _q0_split() + Xf, Xv, yf, yv = split + wrapped = qs._SplitWithFull(split, (np.concatenate([Xf, Xv]), + np.concatenate([yf, yv]))) + assert len(wrapped) == 4 + for a, b in zip(wrapped, split): + assert a is b + Qp, _, _, best_p = qs._fit_chimera_head(split, Xte, None, 1, taus) + Qw, _, _, best_w = qs._fit_chimera_head(wrapped, Xte, None, 1, taus) + assert best_w == best_p + assert np.array_equal(Qw, Qp) + + +def test_refit_probes_record_choice_and_refit(monkeypatch): + """The refit arms fit on split.full with no eval_set and record the + audition choice (0-4) plus what was retrained.""" + monkeypatch.setattr(rb, "MAX_ITERS", 300) + split, Xte, taus = _q0_split() + Xf, Xv, yf, yv = split + wrapped = qs._SplitWithFull(split, (np.concatenate([Xf, Xv]), + np.concatenate([yf, yv]))) + for name in ("ChimeraBoostQuantileRefitAll", + "ChimeraBoostQuantileRefitCentre"): + Q, _, _, _, extra = qs.PROBES[name](wrapped, Xte, None, 1, taus) + assert Q.shape == (Xte.shape[0], len(taus)) + assert np.all(np.diff(Q, axis=1) >= 0) + assert extra["audition_choice"] in (0, 1, 2, 3, 4) + assert isinstance(extra["refit_centre"], bool) + assert isinstance(extra["refit_head"], bool) + if extra["refit_head"]: + assert isinstance(extra["refit_rounds"], int) + else: + assert extra["refit_rounds"] is None From 6370e9a8bdec1c286f0e9cdd2d5fae451a81abb9 Mon Sep 17 00:00:00 2001 From: Nathan Walker Date: Fri, 25 Sep 2026 20:26:12 -0400 Subject: [PATCH 3/5] Plan: I075 verdict (the retrain passes, RefitAll wins); I076 pre-registers the default Co-Authored-By: Claude Opus 5.5 --- benchmarks/CAMPAIGN_PLAN.md | 68 +++++++++++++++++++++++++++++++++++-- 1 file changed, 66 insertions(+), 2 deletions(-) diff --git a/benchmarks/CAMPAIGN_PLAN.md b/benchmarks/CAMPAIGN_PLAN.md index 8a74f20..e5a30eb 100644 --- a/benchmarks/CAMPAIGN_PLAN.md +++ b/benchmarks/CAMPAIGN_PLAN.md @@ -186,7 +186,7 @@ fact: 2026-09-25 | standing quantile BASE = `results/quantile-20260925-172250.js | F3 | Classifier forced-cross | KILLED 2026-09-21 (S3, I021) | gr binary engaged 10W-13L, median −0.04%: the race earns its fee on the classifier. Knob stays opt-in (PR #117), no rung-1 pin | | F5 | hc-Brier gap vs CatBoost | **SHIPPED 2026-09-22 (I046, PR #145, d14bf38): `cat_count_features` on by default** (public 5W-1L, decide hc 6W-1L, bit-identical elsewhere, chart refreshed) — earlier: I038 library form, I039 S3 PASS | the published chart (`public_pareto.png`) refresh: DONE 2026-09-24 on 0.33.0 (`results/20260924-090541.json`, `--public --seeds 3`, 22 datasets, ~4 h: the default avg rank 1.94 [1.64-2.24] against CatBoost 1.84 [1.43-2.24] and LightGBM 2.23, at a median 5.6x against 47.2x and 1.0x; `docs/benchmarks.md` prose updated; `pareto.png` re-rendered from `results/20260924-130232.json` and `docs/PROJECT_STATUS.md` synced); S7 (multiclass CTR width) is the family's next idea if picked. History: **A per-categorical count column closes 47% of CatBoost's hc edge** (I035): 4W-1L on the gap sets, +0.43% Brier median, gains ordered by cardinality, sf-police and Traffic unanimous across seeds. The gap is the encoder (CatBoost on our TS keeps none of its edge); not the prior target, Counter, permutations or quantization (TS quantization kills on big sets, +1.7–3.6% on the two small controls — a small-data pointer, parked). `cat_count_features` (opt-in, card ≥ 256, invisible to the cross and linear-leaf races; I038) on the decision tier (I039): gr 0-0-59 exact ties, the 7 hc sets without a qualifying column exact ties, the engaged 7 **6W-1L** at +0.20% median (sf-police +0.73%, Traffic +0.86% Brier; employee_salaries +2.45%, wine-reviews +0.56% RMSE), hc@time 4-0, fit ×1.09 on hc (engaged median 1.165). PR up with the flag OFF. The random-effects alternative (per-column ANOVA λ for the TS, I040) KILLED: uncapped it collapses the small controls (−3.8 / −9.6%), capped at 10 it is a flat wash and still costs kick and eucalyptus; the count column keeps evidence the shrinkage deletes. Next: the maintainer's go on /experiment S4 for the default flip; meanwhile R4 S0 | | F6 | Ordered-TS train/test moment mismatch (shortlist R1) | KILLED 2026-09-21 (S1b, I025) — closed as barrier B18 | The defect is real (rare categories over-trusted, reliability 0.63–0.70) and two transform-side fixes both went 7W-5L against a bar of 8: the gain is sf-police (9 of 9 fits, +0.29% to +0.53%) and nothing else. Nothing ships; the open door is the Counter feature, which belongs to R3 | -| F7 | Multi-quantile head (`ChimeraBoostQuantileRegressor`) | ACTIVE again 2026-09-25 (the maintainer: "we can continue on the quantile thread"; PARKED 2026-09-24 for GitHub issues; ACTIVE from 2026-09-23). Phase 1, the bench, MERGED (I056, PR #156); Q0 DONE (I057, PR #157 merged): uncapped and validation-rescaled pay, depth 6 misses on its guard alone, recentred fails. Q3 CLOSED (I058, barrier B23: a larger rate loses); depth 6 + early-stopping-row calibration PAID (31W-5L, +0.33%, 90% error 3.43 → 0.47); PR #158 merged. **Q5 PASSED (I059): the head's default is now depth 6 + `conformalize="auto"`** (gr 31W-5L, +0.33%, 90% coverage error 3.43 → 0.47); PR #159 merged, snapshot rebaselined. **Q2 CLOSED (I060)**: no probe clears the tightened gate (sign p < 0.05); the five low-noise sets need resolution (bins) on two and a squared-error location on two; PR #160 merged. **Q6 T0 PASSED (I061)**: the validation-chosen head wins 23W-7L-6T (p 0.005), recovers 99% of the oracle, and tops the 59-key frontier above CatBoost MQ (0.6006 @ 7.2× against 0.5982 @ 129×); PR #161 merged. **Q7 PASSED (I062): the audition is the head's default** (identity 177/177 with the bench arm; gr 23W-7L-6T against the Q5 default, p 0.005; CatBoost MQ now 9W-27L against us, off the frontier); PR #162 merged, **released in 0.33.0**. **Q8 T0 PASSED (I070)**: a fixed-width (S) and a scaled-residual (N) audition candidate, gr 18W-1L-17T against the Q7 default (p 7.6e-5) at 1.11x the fit; S alone and uncapped rounds fail. **Q9 PASSED (I071): S and N in the library default** (identity 177/177; CatBoost MQ 29W-7L, RigidShift 35W-1L; the 59-key chart 0.6046 @ 8.3x); merged as PR #177 (a6a90d2); flag: hc `@time` 90% coverage error 2.12 -> 4.69 (Moneyball, a pointer). **Q10 T0 (I072) DIAGNOSED**: RoNGBa's lead on visualizing_soil and SGEMM is the S/N centre model's accuracy (our grid on its mean beats it); pol is a width gap. **Q11 T0-a (I073)**: centre settings on five sets: 254 bins matches RoNGBa's lead (visualizing_soil +24%, SGEMM +7%) but costs cpu_act 2.5%; 8000 rounds +5% on visualizing_soil only; bagging +30% (ceiling, not a default). **T0-b (I073b) FAILS**: as extra candidates the 254-bin centre still costs cpu_act 2.5% (validation cannot see the overfit); 8000 rounds closed unrun (the centre is cap-bound on 1 of 59 keys). The RoNGBa line is CLOSED. **Q4 CLOSED (I074)**: the N candidate already is the spread-aware categorical encoding; `catscale` at 10k now 16.0 against CatBoost MQ's 30.9 (excess CRPS x1000). **I063 DONE (PR #167)**: RoNGBa is the NGBoost opponent (issue #163) | next: **Q12 (I075), the maintainer's pick 2026-09-25 ("1 is fine")**: retrain the audition's winner on all rows, as `refit_full` does for the point model; bench probe arms, then the decide run. Q1 (P16) stays queued behind it. Program: `QUANTILE_PLAN.md` "Campaign 2026-09-23" | +| F7 | Multi-quantile head (`ChimeraBoostQuantileRegressor`) | ACTIVE again 2026-09-25 (the maintainer: "we can continue on the quantile thread"; PARKED 2026-09-24 for GitHub issues; ACTIVE from 2026-09-23). Phase 1, the bench, MERGED (I056, PR #156); Q0 DONE (I057, PR #157 merged): uncapped and validation-rescaled pay, depth 6 misses on its guard alone, recentred fails. Q3 CLOSED (I058, barrier B23: a larger rate loses); depth 6 + early-stopping-row calibration PAID (31W-5L, +0.33%, 90% error 3.43 → 0.47); PR #158 merged. **Q5 PASSED (I059): the head's default is now depth 6 + `conformalize="auto"`** (gr 31W-5L, +0.33%, 90% coverage error 3.43 → 0.47); PR #159 merged, snapshot rebaselined. **Q2 CLOSED (I060)**: no probe clears the tightened gate (sign p < 0.05); the five low-noise sets need resolution (bins) on two and a squared-error location on two; PR #160 merged. **Q6 T0 PASSED (I061)**: the validation-chosen head wins 23W-7L-6T (p 0.005), recovers 99% of the oracle, and tops the 59-key frontier above CatBoost MQ (0.6006 @ 7.2× against 0.5982 @ 129×); PR #161 merged. **Q7 PASSED (I062): the audition is the head's default** (identity 177/177 with the bench arm; gr 23W-7L-6T against the Q5 default, p 0.005; CatBoost MQ now 9W-27L against us, off the frontier); PR #162 merged, **released in 0.33.0**. **Q8 T0 PASSED (I070)**: a fixed-width (S) and a scaled-residual (N) audition candidate, gr 18W-1L-17T against the Q7 default (p 7.6e-5) at 1.11x the fit; S alone and uncapped rounds fail. **Q9 PASSED (I071): S and N in the library default** (identity 177/177; CatBoost MQ 29W-7L, RigidShift 35W-1L; the 59-key chart 0.6046 @ 8.3x); merged as PR #177 (a6a90d2); flag: hc `@time` 90% coverage error 2.12 -> 4.69 (Moneyball, a pointer). **Q10 T0 (I072) DIAGNOSED**: RoNGBa's lead on visualizing_soil and SGEMM is the S/N centre model's accuracy (our grid on its mean beats it); pol is a width gap. **Q11 T0-a (I073)**: centre settings on five sets: 254 bins matches RoNGBa's lead (visualizing_soil +24%, SGEMM +7%) but costs cpu_act 2.5%; 8000 rounds +5% on visualizing_soil only; bagging +30% (ceiling, not a default). **T0-b (I073b) FAILS**: as extra candidates the 254-bin centre still costs cpu_act 2.5% (validation cannot see the overfit); 8000 rounds closed unrun (the centre is cap-bound on 1 of 59 keys). The RoNGBa line is CLOSED. **Q4 CLOSED (I074)**: the N candidate already is the spread-aware categorical encoding; `catscale` at 10k now 16.0 against CatBoost MQ's 30.9 (excess CRPS x1000). **I063 DONE (PR #167)**: RoNGBa is the NGBoost opponent (issue #163) | next: **Q13 (I076)**: `refit_full=True` as the head's default (Q12, I075, PASSED: gr 35W-1L, +0.91%, p 1.1e-9, 1.28x fit); then Q1 (P16). Program: `QUANTILE_PLAN.md` "Campaign 2026-09-23" | ### F1 — Cross-feature cost trim v2 status: KILLED 2026-08-16 at S2 (I007 at k=6, I008 at k=12) — closed as barrier B16 @@ -453,6 +453,30 @@ Recommended pick: **R1, R2, R3, R4 + H(1)(4)(5)**. R1 and R2 have free probes an ## Iteration log (append-only) +#### I076 2026-09-25 F7 Q13 (`refit_full=True` becomes the head's default; LIBRARY change, pre-registered) +why now: I075's arm (b) passed the gate and beat arm (a) head to head. +change (`chimeraboost/quantile_api.py`): `refit_full` defaults to True +with arm (b)'s behaviour; the private `_refit_scope` and its "centre" +branch are removed (the probe answered); the retrained centre gets the +head's `validation_fraction` so its carve always equals the head's. +Benchmark: the field arm `ChimeraBoostQuantile` keeps passing the shared +split as an `eval_set`, which the retrain never touches, so every +comparison in the suite and the docs stays "every arm on the same rows" +(conservative: the table does not include the retrain's gain). The probe +arm `...RefitAll` stays, as the default called without an `eval_set`; +`...RefitCentre` is deleted with its switch. +forecast: (1) the default called without an `eval_set` equals I075's +RefitAll arm bit for bit, and with an `eval_set` equals today's default +bit for bit; (2) tests green, with the tests whose premise was no retrain +pinned to `refit_full=False` (listed); (3) the identity snapshot moves +only on quantile configs fitted without an `eval_set`; (4) the I075 +numbers stand for the new default without another run. +docs (Claude): `docs/quantiles.md` (a paragraph on the retrain: what it +does, what it keeps from the held-out fit, its measured gain and cost, +`refit_full=False`), `docs/parameters.md` (the `refit_full` row), +CHANGELOG (Unreleased, Changed). +verdict: PENDING(muse library task) + #### I075 2026-09-25 F7 Q12 T0 (the head retrains its winner on all rows after the audition, as the point regressor's `refit_full` does; BENCH-only probe arms, pre-registered) why now: the maintainer, 2026-09-25, picking option 1 of three: "1 is fine". Called without an `eval_set`, the head holds back @@ -493,7 +517,47 @@ median > 0, coverage guard ok) becomes the library rung Q13 (the Q7/Q9 pattern, the arm its exact oracle). Both pass: (b) only if it also beats (a) head to head on gr at p < 0.05; otherwise (a). Neither: the question closes. -verdict: PENDING(muse bench task, then the decide run) +muse (exit 0, `20260925-quantile-q12-refit-probe.md`): `refit_full=False` +on the head (the private `_refit_scope` for arm (a)); the centre retrained +as `ChimeraBoostRegressor().fit(all +rows)` (no `eval_set`: the same carve, its own replay refit), the head +from scratch at `min(ceil(len(trees_) / 0.8), n_estimators)` rounds with +`lr_` pinned; the early-stopped curves and `best_iteration_` carried over; +`refit_` records it. The suite passes `split.full` (a tuple subclass, the +original row order); arms `...RefitCentre`, `...RefitAll`; 11 tests; +1235 passed; ruff clean; identity snapshot 186/186 (off by default). +Review: correct; one nit for the library rung, the retrained centre uses +the regressor's own `validation_fraction` (0.2) rather than the head's. +Commit 7c2c9e9. +ran: `quantile_suite.py --decide --seeds 3 --jobs 5 --models +ChimeraBoostQuantile ...RefitCentre ...RefitAll --save` +(`results/quantile-20260925-202430.json`, 25 min). Against the default, +same run (median over decided sets; coverage = median |cov - nominal| +points at 90 / 80): + +| stratum | (a) RefitCentre | (b) RefitAll | +|---|---|---| +| gr (36) | 19W-0L-17T, +1.72%, p 3.8e-6; cov 0.32 -> 0.40 / 0.37 -> 0.50; fit 1.07x | 35W-1L, +0.91%, p 1.1e-9; cov 0.32 -> 0.47 / 0.37 -> 0.59; fit 1.28x | +| hc (6) | 5W-0L-1T, +1.20% | 6W-0L, +1.42% | +| gr@sus25 (7) | 6W-0L-1T, +1.01% | 7W-0L, +1.63% | +| gr@sus50 (4) | 2W-0L-2T | 3W-1L | +| hc@sus25 / hc@sus50 | 2-0 / 1-0 | 2-0 / 1-0 | +| hc@time (3) | 3W-0L, +4.10%; cov 4.69 -> 3.29 / 4.68 -> 2.62 | same | + +(b) against (a): gr 20W-1L-15T, +0.32%, p 2.1e-5, at 1.13x (a)'s fit. +Largest gains (both): Brazilian_houses +8.6% / +8.2%, visualizing_soil ++7.7%, pol +5.1%, sulfur +4.6 to +5.0%, superconduct +3.8%; (b)'s one loss +analcatdata_supreme -0.48%. Over 59 keys: skill 0.606 -> 0.609 (a) / +0.611 (b); mean 90% coverage 0.898 -> 0.900 / 0.901. Retrained: (b) the +centre on 99 fits (R 6, S 13, N 80), the head on 84 (H 38, B 40, R 6). +forecast: (1) HIT on every item (19 of 19 engaged, +1.72%, p 3.8e-6, +guard ok, fit 1.07x). (2) HALF: engages everywhere HIT, 35 of 36 HIT, +p HIT, fit 1.28x HIT, median +0.91% MISS (forecast +1% to +2%). (3) HIT +(visualizing_soil +7.7%, pol +5.1%). (4) HIT (hc 5-0 / 6-0). +unforecast: the retrain also repairs part of the hc `@time` coverage flag +(I071: 4.69 -> 3.29 points at 90%). +verdict: **both arms PASS; (b) RefitAll beats (a) at p 2.1e-5, so by the +pre-registered rule (b) is the library rung** (Q13, I076). #### I074 2026-09-25 F7 Q4 T0-a (does the five-candidate default still lose `catscale` to CatBoost? a synthetic recheck before any spread-TS work; bench-only, pre-registered) why now: Q4's premise (2026-09-23, `results/quantile-synth-20260923-090818.json`) From e418399ba720579be3a23315dc25cd7434a74ee6 Mon Sep 17 00:00:00 2001 From: Nathan Walker Date: Fri, 25 Sep 2026 20:51:43 -0400 Subject: [PATCH 4/5] Quantile head: retrain the winner on all rows by default (Q13, I076) refit_full now defaults to True: called without an eval_set, the head makes every choice on its carved fold as before (stopping round, calibration, candidate), then retrains the winner on all rows, as ChimeraBoostRegressor does. The retrained centre gets the head's validation_fraction so its carve always equals the head's. The probe-only _refit_scope switch and its centre-only branch are removed. Nothing changes with a user eval_set, early stopping off or conformalize=True. Benchmark: ChimeraBoostQuantileRefitAll becomes ChimeraBoostQuantileAllRows (the default on split.full, no eval_set); ChimeraBoostQuantileRefitCentre is deleted; the field arm keeps its eval_set. Tests: 7 changed, 1 added, 1 deleted; 1235 passed. Identity snapshot 182/186 (mq3, mq3_w_sub predictions and importances; calibration factors unchanged). Co-Authored-By: Claude Opus 5.5 --- benchmarks/quantile_suite.py | 29 +++------- chimeraboost/quantile_api.py | 42 +++++++------- tests/test_quantile_head.py | 104 +++++++++++++++++------------------ tests/test_quantile_shap.py | 5 +- tests/test_quantile_suite.py | 31 +++++------ 5 files changed, 97 insertions(+), 114 deletions(-) diff --git a/benchmarks/quantile_suite.py b/benchmarks/quantile_suite.py index 6b1c995..fbeeb98 100644 --- a/benchmarks/quantile_suite.py +++ b/benchmarks/quantile_suite.py @@ -845,18 +845,16 @@ def _refit_extras(m): "refit_rounds": refit.get("rounds")} -def _fit_chimera_refit(split, Xte, cat, threads, taus, scope): - """The Q12 probe: the head fitted on ALL training rows with - ``refit_full=True`` -- the audition runs on the library's own carve, - which reproduces the shared split, and the winner is then retrained - on every row. ``scope="centre"`` retrains only the R/S/N centre.""" +def _fit_chimera_all_rows(split, Xte, cat, threads, taus): + """The library default as a user without an ``eval_set`` gets it: + fitted on ALL training rows, so the audition runs on the library's + own carve (which reproduces the shared split) and the winner is + retrained on every row.""" Xtr, ytr = split.full m = ChimeraBoostQuantileRegressor( quantiles=taus, n_estimators=rb.MAX_ITERS, early_stopping_rounds=rb.PATIENCE, thread_count=threads, - random_state=0, refit_full=True) - if scope == "centre": - m._refit_scope = "centre" + random_state=0) t = time.time() m.fit(Xtr, ytr, cat_features=cat or None) fit_s = time.time() - t @@ -866,18 +864,6 @@ def _fit_chimera_refit(split, Xte, cat, threads, taus, scope): return Q, fit_s, pred_s, m.best_iteration_, _refit_extras(m) -def _fit_chimera_refit_all(split, Xte, cat, threads, taus): - """The head retrained on all rows: centre and head alike (Q12 arm b).""" - return _fit_chimera_refit(split, Xte, cat, threads, taus, scope="all") - - -def _fit_chimera_refit_centre(split, Xte, cat, threads, taus): - """The head retrained on all rows: the R/S/N centre only, every head - booster untouched (Q12 arm a).""" - return _fit_chimera_refit(split, Xte, cat, threads, taus, - scope="centre") - - PROBES = { "ChimeraBoostQuantileUncapped": _fit_chimera_uncapped, "ChimeraBoostQuantileDepth6": _fit_chimera_depth6, @@ -896,8 +882,7 @@ def _fit_chimera_refit_centre(split, Xte, cat, threads, taus): "ChimeraBoostQuantileAuditionS": _fit_chimera_audition_s, "ChimeraBoostQuantileAuditionSN": _fit_chimera_audition_sn, "ChimeraBoostQuantileDefaultUncapped": _fit_chimera_default_uncapped, - "ChimeraBoostQuantileRefitAll": _fit_chimera_refit_all, - "ChimeraBoostQuantileRefitCentre": _fit_chimera_refit_centre, + "ChimeraBoostQuantileAllRows": _fit_chimera_all_rows, } diff --git a/chimeraboost/quantile_api.py b/chimeraboost/quantile_api.py index f0643d7..5c0de27 100644 --- a/chimeraboost/quantile_api.py +++ b/chimeraboost/quantile_api.py @@ -2,8 +2,8 @@ Kept out of ``sklearn_api`` because it shares that module's input validation but almost none of its fit machinery: no loss family, no linear leaves, no -cross features, no bagging, and only an opt-in full-data refit. It imports -the validation helpers and keeps the same flat, module-function style. +cross features, no bagging, and a full-data refit that is on by default. It +imports the validation helpers and keeps the same flat, module-function style. """ import warnings @@ -347,21 +347,22 @@ class ChimeraBoostQuantileRegressor(BaseEstimator): calibrates the head and scored unweighted; ties go H, then B, then R, then S, then N. Needs ``conformalize="auto"`` and evaluation rows (the user's ``eval_set`` or the carved fold); otherwise the - single head is fitted. Costs about 2.6x a single head at the median: - S reuses R's centre fit, so the extra cost over the three-candidate - audition is the spread-model fit, about 10% of the fit at the - median. ``False`` fits a single head, as releases before 0.33.0 did. + single head is fitted. Costs about 2.6x a single head at the median + for fits with an ``eval_set``: S reuses R's centre fit, so the extra + cost over the three-candidate audition is the spread-model fit, + about 10% of the fit at the median. The full-data refit adds about + 30% at the median when it acts. ``False`` fits a single head, as + releases before 0.33.0 did. ``model_`` stays the fitted head booster whatever wins (the 254-bin booster for a bins win); for S and N it delivers nothing -- ``predict`` serves the centre/spread models -- but stays inspectable. - refit_full : bool, default False + refit_full : bool, default True Retrain the winner on all rows after early stopping chose the - budget, as ``ChimeraBoostRegressor`` does. Acts only when the fit - used the automatic early-stopping split (no user ``eval_set``, - early stopping on, the carve succeeded) and ``conformalize`` is - ``"auto"`` or False; with ``conformalize=True``, a user - ``eval_set``, or no carve the fit is exactly what it is without - it. An R, S or N winner's centre is then replaced by a default + budget, as ``ChimeraBoostRegressor`` does. Acts only on the + automatic early-stopping split: never with a user ``eval_set``, + with early stopping off, with ``conformalize=True``, or when the + carve failed -- those fits are exactly what they are without it. + An R, S or N winner's centre is then replaced by a default ``ChimeraBoostRegressor`` fitted on all rows without an ``eval_set`` -- offsets, spread model, residual quantiles, floor and calibration factors stay from the audition fit -- and the @@ -372,7 +373,8 @@ class ChimeraBoostQuantileRegressor(BaseEstimator): the early-stopped head kept. ``audition_``, ``conformal_scale_``, ``best_iteration_`` and ``validation_history_`` keep the early-stopped fit's values; - ``refit_`` records what was retrained. + ``refit_`` records what was retrained. ``False`` skips the + retrain. Attributes ---------- @@ -418,7 +420,7 @@ def __init__(self, quantiles=None, n_estimators=2000, learning_rate=None, validation_fraction=0.2, split_projection="rotate", exact_splits=False, conformalize="auto", calibration_fraction=0.2, audition=True, - refit_full=False): + refit_full=True): self.quantiles = quantiles self.n_estimators = n_estimators self.learning_rate = learning_rate @@ -445,10 +447,6 @@ def __init__(self, quantiles=None, n_estimators=2000, learning_rate=None, self.calibration_fraction = calibration_fraction self.audition = audition self.refit_full = refit_full - # Probe-only scope, set after construction (never a constructor - # parameter): "centre" retrains only the R/S/N centre and leaves - # every head booster untouched. - self._refit_scope = "all" def __sklearn_is_fitted__(self): return hasattr(self, "model_") @@ -532,8 +530,7 @@ def _maybe_refit_full(self, auto_split, taus, es_rounds, cat_features, y_full, sw_full, groups_full) centre_done = True head_done, rounds = False, None - if self._refit_scope != "centre" and selected in ( - None, "head", "bins", "recentred"): + if selected in (None, "head", "bins", "recentred"): rounds = self._refit_head_booster(taus, selected, cat_features, X_full, y_full, sw_full) head_done = True @@ -555,7 +552,8 @@ def _refit_audition_centre(self, es_rounds, cat_features, X_full, centre = ChimeraBoostRegressor( n_estimators=self.n_estimators, early_stopping_rounds=es_rounds, - thread_count=self.thread_count, random_state=self.random_state) + thread_count=self.thread_count, random_state=self.random_state, + validation_fraction=self.validation_fraction) centre.fit(X_full, y_full, cat_features=cat_features, sample_weight=sw_full, groups=groups_full) _keep_es_state(centre.model_, old.model_) diff --git a/tests/test_quantile_head.py b/tests/test_quantile_head.py index d12348e..315a889 100644 --- a/tests/test_quantile_head.py +++ b/tests/test_quantile_head.py @@ -1054,11 +1054,11 @@ def test_audition_fixed_and_scaled_keep_the_head_booster(): # -------------------------------------------------------------------------- -# The full-data refit (Q12): off by default, audition-preserving. +# The full-data refit (Q13): on by default, audition-preserving. # -------------------------------------------------------------------------- -def _aud_nosplit_fit(kind, scope="all", **kw): +def _aud_nosplit_fit(kind, **kw): """Fit on the kind's training rows WITHOUT an eval_set, so the library carves its own split; return (model, Xtr, ytr, Xte, yte).""" from sklearn.model_selection import train_test_split @@ -1068,13 +1068,13 @@ def _aud_nosplit_fit(kind, scope="all", **kw): m = ChimeraBoostQuantileRegressor( quantiles=_AUD_TAUS, n_estimators=300, early_stopping_rounds=50, thread_count=1, random_state=0, **kw) - m._refit_scope = scope return m.fit(Xtr, ytr), Xtr, ytr, Xte, yte def test_no_eval_set_equals_suite_split_fit(): """Without an eval_set the head carves the suite's own split: bit for - bit the same fit as training on the carved rows with that eval_set.""" + bit the same fit as training on the carved rows with that eval_set + (the refit pinned off, so the carve is what is compared).""" from sklearn.model_selection import train_test_split rng = np.random.default_rng(60) n = 1500 @@ -1085,7 +1085,8 @@ def test_no_eval_set_equals_suite_split_fit(): + rng.standard_normal(n)) Xf, Xv, yf, yv = train_test_split(X, y, test_size=0.2, random_state=0) kw = dict(quantiles=_AUD_TAUS, n_estimators=300, - early_stopping_rounds=50, thread_count=1, random_state=0) + early_stopping_rounds=50, thread_count=1, random_state=0, + refit_full=False) m_auto = ChimeraBoostQuantileRegressor(**kw).fit(X, y, cat_features=[3]) m_split = ChimeraBoostQuantileRegressor(**kw).fit( Xf, yf, cat_features=[3], eval_set=(Xv, yv)) @@ -1095,17 +1096,17 @@ def test_no_eval_set_equals_suite_split_fit(): assert np.array_equal(m_auto.predict(X), m_split.predict(X)) -def test_refit_full_false_equals_default(): - """refit_full=False is today's fit, exactly; the flag is a visible - constructor parameter defaulting to False.""" +def test_refit_full_true_equals_default(): + """refit_full=True is today's fit, exactly; the flag is a visible + constructor parameter defaulting to True.""" assert (ChimeraBoostQuantileRegressor().get_params()["refit_full"] - is False) + is True) X, y = _heteroscedastic(n=1200, seed=61) kw = dict(random_state=0, n_estimators=100, thread_count=1) m = ChimeraBoostQuantileRegressor(**kw).fit(X, y) - r = ChimeraBoostQuantileRegressor(refit_full=False, **kw).fit(X, y) - assert m.refit_ is None - assert r.refit_ is None + r = ChimeraBoostQuantileRegressor(refit_full=True, **kw).fit(X, y) + assert m.refit_ is not None + assert r.refit_ == m.refit_ assert m.audition_ == r.audition_ assert np.array_equal(m.conformal_scale_, r.conformal_scale_) assert np.array_equal(m.predict(X), r.predict(X)) @@ -1113,26 +1114,30 @@ def test_refit_full_false_equals_default(): def test_refit_full_needs_auto_split_and_no_honest_fold(): """With a user eval_set, conformalize=True, or early stopping off, the - refit stays out: refit_full=True is exactly refit_full=False.""" + refit stays out: the default is exactly refit_full=False.""" X, y = _heteroscedastic(n=1200, seed=62) Xt, yt, Xv, yv = X[:900], y[:900], X[900:], y[900:] base = dict(random_state=0, n_estimators=100, thread_count=1) - a = ChimeraBoostQuantileRegressor(**base).fit(Xt, yt, eval_set=(Xv, yv)) - b = ChimeraBoostQuantileRegressor(refit_full=True, **base).fit( + d = ChimeraBoostQuantileRegressor(**base).fit(Xt, yt, eval_set=(Xv, yv)) + f = ChimeraBoostQuantileRegressor(refit_full=False, **base).fit( Xt, yt, eval_set=(Xv, yv)) - assert b.refit_ is None - assert a.audition_ == b.audition_ - assert np.array_equal(a.predict(Xt), b.predict(Xt)) - a = ChimeraBoostQuantileRegressor(conformalize=True, **base).fit(X, y) - b = ChimeraBoostQuantileRegressor( - conformalize=True, refit_full=True, **base).fit(X, y) - assert b.refit_ is None - assert np.array_equal(a.predict(X), b.predict(X)) - a = ChimeraBoostQuantileRegressor(early_stopping=False, **base).fit(X, y) - b = ChimeraBoostQuantileRegressor( - early_stopping=False, refit_full=True, **base).fit(X, y) - assert b.refit_ is None - assert np.array_equal(a.predict(X), b.predict(X)) + assert d.refit_ is None + assert f.refit_ is None + assert d.audition_ == f.audition_ + assert np.array_equal(d.conformal_scale_, f.conformal_scale_) + assert np.array_equal(d.predict(Xt), f.predict(Xt)) + d = ChimeraBoostQuantileRegressor(conformalize=True, **base).fit(X, y) + f = ChimeraBoostQuantileRegressor( + conformalize=True, refit_full=False, **base).fit(X, y) + assert d.refit_ is None + assert f.refit_ is None + assert np.array_equal(d.predict(X), f.predict(X)) + d = ChimeraBoostQuantileRegressor(early_stopping=False, **base).fit(X, y) + f = ChimeraBoostQuantileRegressor( + early_stopping=False, refit_full=False, **base).fit(X, y) + assert d.refit_ is None + assert f.refit_ is None + assert np.array_equal(d.predict(X), f.predict(X)) _REFIT_WANT = { @@ -1146,9 +1151,9 @@ def test_refit_full_records_and_keeps_es_values(): """Per winner: refit_ says what was retrained, at the replay round count; audition_, factors, best round and curve keep the 80% values.""" for kind in ("head", "bins", "recentred", "fixed", "scaled"): - mf, _, _, _, _ = _aud_nosplit_fit(kind) - mt, _, _, _, _ = _aud_nosplit_fit(kind, refit_full=True) - assert mf.audition_["selected"] == kind + mf, _, _, _, _ = _aud_nosplit_fit(kind, refit_full=False) + mt, _, _, _, _ = _aud_nosplit_fit(kind) + assert mt.audition_["selected"] == kind assert mt.audition_ == mf.audition_ want_c, want_h = _REFIT_WANT[kind] assert mt.refit_["centre"] is want_c @@ -1169,8 +1174,8 @@ def test_refit_full_centre_matches_full_row_regressor(): """The delivered R/S/N centre is the same regressor fitted on all rows; offsets, residual quantiles, floor and spread stay from the audition.""" for kind in ("recentred", "fixed", "scaled"): - mt, Xtr, ytr, Xte, _ = _aud_nosplit_fit(kind, refit_full=True) - mf, _, _, _, _ = _aud_nosplit_fit(kind) + mt, Xtr, ytr, Xte, _ = _aud_nosplit_fit(kind) + mf, _, _, _, _ = _aud_nosplit_fit(kind, refit_full=False) ref = ChimeraBoostRegressor( n_estimators=300, early_stopping_rounds=50, thread_count=1, random_state=0).fit(Xtr, ytr) @@ -1189,7 +1194,7 @@ def test_refit_full_centre_matches_full_row_regressor(): def test_refit_full_paths_serve_retrained_models(): """Ordered grids, staged final equals predict, pickle round-trips.""" for kind in ("head", "bins", "recentred", "fixed", "scaled"): - mt, _, _, Xte, _ = _aud_nosplit_fit(kind, refit_full=True) + mt, _, _, Xte, _ = _aud_nosplit_fit(kind) Q = mt.predict(Xte) assert np.all(np.diff(Q, axis=1) >= 0) Xt = Xte[:50] @@ -1200,21 +1205,16 @@ def test_refit_full_paths_serve_retrained_models(): assert back.refit_ == mt.refit_ -def test_refit_scope_centre_leaves_heads_untouched(): - """With _refit_scope = "centre", H and B winners come back exactly as - without the refit; R keeps its head and retrains only its centre.""" - for kind in ("head", "bins"): - mf, _, _, Xte, _ = _aud_nosplit_fit(kind) - mc, _, _, _, _ = _aud_nosplit_fit(kind, scope="centre", - refit_full=True) - assert mc.refit_ is None - assert np.array_equal(mc.predict(Xte), mf.predict(Xte)) - mr, Xtr, ytr, Xte, _ = _aud_nosplit_fit("recentred", scope="centre", - refit_full=True) - mh, _, _, _, _ = _aud_nosplit_fit("recentred") - assert mr.refit_ == {"centre": True, "head": False, "rounds": None} - assert np.array_equal(mr.model_.predict_raw(Xte), - mh.model_.predict_raw(Xte)) - ref = ChimeraBoostRegressor(n_estimators=300, early_stopping_rounds=50, - thread_count=1, random_state=0).fit(Xtr, ytr) - assert np.array_equal(mr._centre_model_.predict(Xte), ref.predict(Xte)) +def test_refit_centre_uses_head_validation_fraction(): + """With validation_fraction=0.3 the retrained centre still carves the + head's split: it equals the same regressor with validation_fraction=0.3 + fitted on all rows.""" + for kind in ("recentred", "fixed", "scaled"): + mt, Xtr, ytr, Xte, _ = _aud_nosplit_fit( + kind, validation_fraction=0.3) + assert mt.audition_["selected"] in ("recentred", "fixed", "scaled") + ref = ChimeraBoostRegressor( + n_estimators=300, early_stopping_rounds=50, thread_count=1, + random_state=0, validation_fraction=0.3).fit(Xtr, ytr) + assert np.array_equal(mt._centre_model_.predict(Xte), + ref.predict(Xte)) diff --git a/tests/test_quantile_shap.py b/tests/test_quantile_shap.py index 28bcce4..b02409b 100644 --- a/tests/test_quantile_shap.py +++ b/tests/test_quantile_shap.py @@ -325,7 +325,8 @@ def test_audition_winner_shap_is_locally_accurate(): def _aud_nosplit_shap(kind): - """The kind's training rows without an eval_set, refit on; (model, Xte).""" + """The kind's training rows without an eval_set, under the default; + (model, Xte).""" from sklearn.model_selection import train_test_split taus = [0.05, 0.25, 0.5, 0.75, 0.95] X, y, sseed = _aud_data_shap(kind) @@ -333,7 +334,7 @@ def _aud_nosplit_shap(kind): random_state=sseed) m = ChimeraBoostQuantileRegressor( quantiles=taus, n_estimators=300, early_stopping_rounds=50, - thread_count=1, random_state=0, refit_full=True).fit(Xtr, ytr) + thread_count=1, random_state=0).fit(Xtr, ytr) return m, Xte diff --git a/tests/test_quantile_suite.py b/tests/test_quantile_suite.py index b8247d8..d2c035d 100644 --- a/tests/test_quantile_suite.py +++ b/tests/test_quantile_suite.py @@ -233,7 +233,7 @@ def test_q0_probes_are_opt_in(monkeypatch): import quantile_synth as qsyn assert len(qs.ARMS) == 7 - assert len(qs.PROBES) == 19 + assert len(qs.PROBES) == 18 _stub_registry(monkeypatch, ["gr:reg_num/houses"], {"gr:reg_num/houses": "regression"}) @@ -704,23 +704,22 @@ def test_split_with_full_leaves_existing_arms_unchanged(): assert np.array_equal(Qw, Qp) -def test_refit_probes_record_choice_and_refit(monkeypatch): - """The refit arms fit on split.full with no eval_set and record the - audition choice (0-4) plus what was retrained.""" +def test_all_rows_probe_records_choice_and_refit(monkeypatch): + """The AllRows arm fits the default on split.full with no eval_set and + records the audition choice (0-4) plus what was retrained.""" monkeypatch.setattr(rb, "MAX_ITERS", 300) split, Xte, taus = _q0_split() Xf, Xv, yf, yv = split wrapped = qs._SplitWithFull(split, (np.concatenate([Xf, Xv]), np.concatenate([yf, yv]))) - for name in ("ChimeraBoostQuantileRefitAll", - "ChimeraBoostQuantileRefitCentre"): - Q, _, _, _, extra = qs.PROBES[name](wrapped, Xte, None, 1, taus) - assert Q.shape == (Xte.shape[0], len(taus)) - assert np.all(np.diff(Q, axis=1) >= 0) - assert extra["audition_choice"] in (0, 1, 2, 3, 4) - assert isinstance(extra["refit_centre"], bool) - assert isinstance(extra["refit_head"], bool) - if extra["refit_head"]: - assert isinstance(extra["refit_rounds"], int) - else: - assert extra["refit_rounds"] is None + Q, _, _, _, extra = qs.PROBES["ChimeraBoostQuantileAllRows"]( + wrapped, Xte, None, 1, taus) + assert Q.shape == (Xte.shape[0], len(taus)) + assert np.all(np.diff(Q, axis=1) >= 0) + assert extra["audition_choice"] in (0, 1, 2, 3, 4) + assert isinstance(extra["refit_centre"], bool) + assert isinstance(extra["refit_head"], bool) + if extra["refit_head"]: + assert isinstance(extra["refit_rounds"], int) + else: + assert extra["refit_rounds"] is None From 2401f93edc3be48fd8b59a90fa28b133a35e31af Mon Sep 17 00:00:00 2001 From: Nathan Walker Date: Fri, 25 Sep 2026 20:51:43 -0400 Subject: [PATCH 5/5] Quantile docs and records for the all-rows retrain (I075, I076) Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 18 ++++++++++++++++++ benchmarks/CAMPAIGN_PLAN.md | 23 +++++++++++++++++++++-- benchmarks/QUANTILE_PLAN.md | 19 ++++++++++++++++++- docs/parameters.md | 1 + docs/quantiles.md | 30 ++++++++++++++++++++++++++++-- 5 files changed, 86 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 846ae1f..f999b77 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,24 @@ The format follows [Keep a Changelog](https://keepachangelog.com/). ## [Unreleased] ### Changed +- **The quantile head retrains its winner on all rows by default.** Called + without an `eval_set`, `ChimeraBoostQuantileRegressor` held back + `validation_fraction` of the rows for early stopping, calibration and + the candidate choice, and never trained on them. The new `refit_full`, + on by default, makes every choice on those rows as before and then + retrains the winner on all rows, as `ChimeraBoostRegressor` does. The + candidate, the calibration factors, the fixed and scaled candidates' + offsets, `best_iteration_` and `validation_history_` stay from the + held-out fit, and `refit_` records what was retrained. On the 36 + Grinsztajn regression datasets it improves CRPS on 35 and loses 1 by + 0.5% (median +0.9%; up to +8.6% on Brazilian_houses), and it improves all + 6 high-cardinality ones (median +1.4%). It also shrinks the time-split + caveat in the next entry: the median 90% coverage error there falls + from 4.7 to 3.3 points. It adds about 30% to the fit time where it acts. + Nothing changes with your own `eval_set`, with early stopping off, or + with `conformalize=True`; `refit_full=False` skips it. The comparisons + in `docs/quantiles.md` keep every model on the same rows, so they do not + include this gain. Record: `benchmarks/QUANTILE_PLAN.md`, Q12 and Q13. - **The quantile head's audition gains two candidates, and its CRPS improves most on low-noise targets.** `ChimeraBoostQuantileRegressor`'s default choice now also tries `"fixed"`, an ordinary squared-error diff --git a/benchmarks/CAMPAIGN_PLAN.md b/benchmarks/CAMPAIGN_PLAN.md index e5a30eb..18205ea 100644 --- a/benchmarks/CAMPAIGN_PLAN.md +++ b/benchmarks/CAMPAIGN_PLAN.md @@ -186,7 +186,7 @@ fact: 2026-09-25 | standing quantile BASE = `results/quantile-20260925-172250.js | F3 | Classifier forced-cross | KILLED 2026-09-21 (S3, I021) | gr binary engaged 10W-13L, median −0.04%: the race earns its fee on the classifier. Knob stays opt-in (PR #117), no rung-1 pin | | F5 | hc-Brier gap vs CatBoost | **SHIPPED 2026-09-22 (I046, PR #145, d14bf38): `cat_count_features` on by default** (public 5W-1L, decide hc 6W-1L, bit-identical elsewhere, chart refreshed) — earlier: I038 library form, I039 S3 PASS | the published chart (`public_pareto.png`) refresh: DONE 2026-09-24 on 0.33.0 (`results/20260924-090541.json`, `--public --seeds 3`, 22 datasets, ~4 h: the default avg rank 1.94 [1.64-2.24] against CatBoost 1.84 [1.43-2.24] and LightGBM 2.23, at a median 5.6x against 47.2x and 1.0x; `docs/benchmarks.md` prose updated; `pareto.png` re-rendered from `results/20260924-130232.json` and `docs/PROJECT_STATUS.md` synced); S7 (multiclass CTR width) is the family's next idea if picked. History: **A per-categorical count column closes 47% of CatBoost's hc edge** (I035): 4W-1L on the gap sets, +0.43% Brier median, gains ordered by cardinality, sf-police and Traffic unanimous across seeds. The gap is the encoder (CatBoost on our TS keeps none of its edge); not the prior target, Counter, permutations or quantization (TS quantization kills on big sets, +1.7–3.6% on the two small controls — a small-data pointer, parked). `cat_count_features` (opt-in, card ≥ 256, invisible to the cross and linear-leaf races; I038) on the decision tier (I039): gr 0-0-59 exact ties, the 7 hc sets without a qualifying column exact ties, the engaged 7 **6W-1L** at +0.20% median (sf-police +0.73%, Traffic +0.86% Brier; employee_salaries +2.45%, wine-reviews +0.56% RMSE), hc@time 4-0, fit ×1.09 on hc (engaged median 1.165). PR up with the flag OFF. The random-effects alternative (per-column ANOVA λ for the TS, I040) KILLED: uncapped it collapses the small controls (−3.8 / −9.6%), capped at 10 it is a flat wash and still costs kick and eucalyptus; the count column keeps evidence the shrinkage deletes. Next: the maintainer's go on /experiment S4 for the default flip; meanwhile R4 S0 | | F6 | Ordered-TS train/test moment mismatch (shortlist R1) | KILLED 2026-09-21 (S1b, I025) — closed as barrier B18 | The defect is real (rare categories over-trusted, reliability 0.63–0.70) and two transform-side fixes both went 7W-5L against a bar of 8: the gain is sf-police (9 of 9 fits, +0.29% to +0.53%) and nothing else. Nothing ships; the open door is the Counter feature, which belongs to R3 | -| F7 | Multi-quantile head (`ChimeraBoostQuantileRegressor`) | ACTIVE again 2026-09-25 (the maintainer: "we can continue on the quantile thread"; PARKED 2026-09-24 for GitHub issues; ACTIVE from 2026-09-23). Phase 1, the bench, MERGED (I056, PR #156); Q0 DONE (I057, PR #157 merged): uncapped and validation-rescaled pay, depth 6 misses on its guard alone, recentred fails. Q3 CLOSED (I058, barrier B23: a larger rate loses); depth 6 + early-stopping-row calibration PAID (31W-5L, +0.33%, 90% error 3.43 → 0.47); PR #158 merged. **Q5 PASSED (I059): the head's default is now depth 6 + `conformalize="auto"`** (gr 31W-5L, +0.33%, 90% coverage error 3.43 → 0.47); PR #159 merged, snapshot rebaselined. **Q2 CLOSED (I060)**: no probe clears the tightened gate (sign p < 0.05); the five low-noise sets need resolution (bins) on two and a squared-error location on two; PR #160 merged. **Q6 T0 PASSED (I061)**: the validation-chosen head wins 23W-7L-6T (p 0.005), recovers 99% of the oracle, and tops the 59-key frontier above CatBoost MQ (0.6006 @ 7.2× against 0.5982 @ 129×); PR #161 merged. **Q7 PASSED (I062): the audition is the head's default** (identity 177/177 with the bench arm; gr 23W-7L-6T against the Q5 default, p 0.005; CatBoost MQ now 9W-27L against us, off the frontier); PR #162 merged, **released in 0.33.0**. **Q8 T0 PASSED (I070)**: a fixed-width (S) and a scaled-residual (N) audition candidate, gr 18W-1L-17T against the Q7 default (p 7.6e-5) at 1.11x the fit; S alone and uncapped rounds fail. **Q9 PASSED (I071): S and N in the library default** (identity 177/177; CatBoost MQ 29W-7L, RigidShift 35W-1L; the 59-key chart 0.6046 @ 8.3x); merged as PR #177 (a6a90d2); flag: hc `@time` 90% coverage error 2.12 -> 4.69 (Moneyball, a pointer). **Q10 T0 (I072) DIAGNOSED**: RoNGBa's lead on visualizing_soil and SGEMM is the S/N centre model's accuracy (our grid on its mean beats it); pol is a width gap. **Q11 T0-a (I073)**: centre settings on five sets: 254 bins matches RoNGBa's lead (visualizing_soil +24%, SGEMM +7%) but costs cpu_act 2.5%; 8000 rounds +5% on visualizing_soil only; bagging +30% (ceiling, not a default). **T0-b (I073b) FAILS**: as extra candidates the 254-bin centre still costs cpu_act 2.5% (validation cannot see the overfit); 8000 rounds closed unrun (the centre is cap-bound on 1 of 59 keys). The RoNGBa line is CLOSED. **Q4 CLOSED (I074)**: the N candidate already is the spread-aware categorical encoding; `catscale` at 10k now 16.0 against CatBoost MQ's 30.9 (excess CRPS x1000). **I063 DONE (PR #167)**: RoNGBa is the NGBoost opponent (issue #163) | next: **Q13 (I076)**: `refit_full=True` as the head's default (Q12, I075, PASSED: gr 35W-1L, +0.91%, p 1.1e-9, 1.28x fit); then Q1 (P16). Program: `QUANTILE_PLAN.md` "Campaign 2026-09-23" | +| F7 | Multi-quantile head (`ChimeraBoostQuantileRegressor`) | ACTIVE again 2026-09-25 (the maintainer: "we can continue on the quantile thread"; PARKED 2026-09-24 for GitHub issues; ACTIVE from 2026-09-23). Phase 1, the bench, MERGED (I056, PR #156); Q0 DONE (I057, PR #157 merged): uncapped and validation-rescaled pay, depth 6 misses on its guard alone, recentred fails. Q3 CLOSED (I058, barrier B23: a larger rate loses); depth 6 + early-stopping-row calibration PAID (31W-5L, +0.33%, 90% error 3.43 → 0.47); PR #158 merged. **Q5 PASSED (I059): the head's default is now depth 6 + `conformalize="auto"`** (gr 31W-5L, +0.33%, 90% coverage error 3.43 → 0.47); PR #159 merged, snapshot rebaselined. **Q2 CLOSED (I060)**: no probe clears the tightened gate (sign p < 0.05); the five low-noise sets need resolution (bins) on two and a squared-error location on two; PR #160 merged. **Q6 T0 PASSED (I061)**: the validation-chosen head wins 23W-7L-6T (p 0.005), recovers 99% of the oracle, and tops the 59-key frontier above CatBoost MQ (0.6006 @ 7.2× against 0.5982 @ 129×); PR #161 merged. **Q7 PASSED (I062): the audition is the head's default** (identity 177/177 with the bench arm; gr 23W-7L-6T against the Q5 default, p 0.005; CatBoost MQ now 9W-27L against us, off the frontier); PR #162 merged, **released in 0.33.0**. **Q8 T0 PASSED (I070)**: a fixed-width (S) and a scaled-residual (N) audition candidate, gr 18W-1L-17T against the Q7 default (p 7.6e-5) at 1.11x the fit; S alone and uncapped rounds fail. **Q9 PASSED (I071): S and N in the library default** (identity 177/177; CatBoost MQ 29W-7L, RigidShift 35W-1L; the 59-key chart 0.6046 @ 8.3x); merged as PR #177 (a6a90d2); flag: hc `@time` 90% coverage error 2.12 -> 4.69 (Moneyball, a pointer). **Q10 T0 (I072) DIAGNOSED**: RoNGBa's lead on visualizing_soil and SGEMM is the S/N centre model's accuracy (our grid on its mean beats it); pol is a width gap. **Q11 T0-a (I073)**: centre settings on five sets: 254 bins matches RoNGBa's lead (visualizing_soil +24%, SGEMM +7%) but costs cpu_act 2.5%; 8000 rounds +5% on visualizing_soil only; bagging +30% (ceiling, not a default). **T0-b (I073b) FAILS**: as extra candidates the 254-bin centre still costs cpu_act 2.5% (validation cannot see the overfit); 8000 rounds closed unrun (the centre is cap-bound on 1 of 59 keys). The RoNGBa line is CLOSED. **Q4 CLOSED (I074)**: the N candidate already is the spread-aware categorical encoding; `catscale` at 10k now 16.0 against CatBoost MQ's 30.9 (excess CRPS x1000). **I063 DONE (PR #167)**: RoNGBa is the NGBoost opponent (issue #163) | next: **Q13 (I076) PASSED, PR for the maintainer**: `refit_full=True` as the head's default (gr 35W-1L, +0.91%, p 1.1e-9, 1.28x fit); after the merge, Q1 (P16). Program: `QUANTILE_PLAN.md` "Campaign 2026-09-23" | ### F1 — Cross-feature cost trim v2 status: KILLED 2026-08-16 at S2 (I007 at k=6, I008 at k=12) — closed as barrier B16 @@ -475,7 +475,26 @@ docs (Claude): `docs/quantiles.md` (a paragraph on the retrain: what it does, what it keeps from the held-out fit, its measured gain and cost, `refit_full=False`), `docs/parameters.md` (the `refit_full` row), CHANGELOG (Unreleased, Changed). -verdict: PENDING(muse library task) +muse (exit 0, `20260925-quantile-q13-refit-default.md`): `refit_full=True` +by default; `_refit_scope` and the centre-only branch removed; the +retrained centre gets the head's `validation_fraction` (a new test at 0.3 +fails against a 0.2 carve, so it matters off the default); docstrings; +`...RefitAll` renamed `...AllRows` (the default on `split.full`, no +`eval_set`), `...RefitCentre` deleted, probe count 18; tests: 7 changed +(one pinned to `refit_full=False`: the carve-premise test, whose two fits +now differ by design), 1 added, 1 deleted; 1235 passed; ruff clean. +checks (Claude): the default without an `eval_set` reproduces I075's +RefitAll CRPS exactly on 8 of 8 seed-0 fits covering H, B, R and N +winners (pol, visualizing_soil, cpu_act, colleges, analcatdata_supreme, +house_prices_nominal, sulfur@sus25, Mercedes_Benz), and the field arm +reproduces the standing BASE on the same 8. Identity snapshot 182/186: +the moved pins are `mq3` and `mq3_w_sub`'s predictions and importances +(quantile configs fitted without an `eval_set`); their calibration +factors do not move. Rebaseline after the merge. +forecast: (1) HIT; (2) HIT; (3) HIT; (4) HIT (I075's numbers, reproduced +exactly). +verdict: **PASS -> PR for the maintainer** (library, tests, docs). After +the merge: rebaseline the identity snapshot, delete the branch. #### I075 2026-09-25 F7 Q12 T0 (the head retrains its winner on all rows after the audition, as the point regressor's `refit_full` does; BENCH-only probe arms, pre-registered) why now: the maintainer, 2026-09-25, picking option 1 of three: "1 is diff --git a/benchmarks/QUANTILE_PLAN.md b/benchmarks/QUANTILE_PLAN.md index a3846e2..81883b4 100644 --- a/benchmarks/QUANTILE_PLAN.md +++ b/benchmarks/QUANTILE_PLAN.md @@ -424,6 +424,23 @@ Q0 reads moved it down to item 4. 1 of the 59 decide keys, so it cannot clear the gate. Bagging the centre ×5 is the ceiling (visualizing_soil +30%, pol +9%) and is an ensemble, so not a default. +3g. **Q12, retraining the winner on all rows (the maintainer's pick + 2026-09-25).** Without an `eval_set` the head's own carve equals the + suite's shared split bit for bit, so the probe is a retrain added to + today's fit: every choice and calibration quantity from the held-out + fit, then (a) the R/S/N centre retrained on all rows, or (b) that plus + the H/B/R head retrained from scratch at the replay-round rule. + **PASSED 2026-09-25 (I075, `results/quantile-20260925-202430.json`).** + Against the default: (a) gr 19W-0L-17T, +1.72%, p 3.8e-6, 1.07× fit; + (b) gr 35W-1L, +0.91%, p 1.1e-9, hc 6W-0L, 1.28× fit; (b) beats (a) + 20W-1L-15T (p 2.1e-5), so (b). Coverage guard fine (gr 90% error 0.32 → + 0.47 points); the hc `@time` flag shrinks (4.69 → 3.29). +3h. **Q13, `refit_full=True` as the default.** + **PASSED 2026-09-25 (I076), PR for the maintainer.** Without an + `eval_set` the default reproduces Q12's arm (b) bit for bit; with one, + nothing changes. The suite's field arm keeps its `eval_set`, so every + comparison stays "every model on the same rows" and leaves this gain + out. 4. **Q1, the narrow-interval defect (P16).** Leaf values are in-sample residual quantiles, so intervals over-narrow (0.869 at nominal 0.90 on 2026-08-30; coverage decays with rounds). Fit leaf quantiles @@ -533,7 +550,7 @@ offset on real data (a CRPS tie), which made it Q0's subject. against pinball. That is a default flip on a strength surface, so it needs its own pre-registration and the full `/experiment` protocol. Not attempted here. Recorded 2026-08-30. -- **The head never trains on its own early-stopping fold** (recorded +- RESOLVED 2026-09-25 (Q12 and Q13, I075 and I076): **The head never trains on its own early-stopping fold** (recorded 2026-09-25, I073b). Without an `eval_set` the head carves `validation_fraction` and neither it nor its S/N centre ever sees those rows, while `ChimeraBoostRegressor` refits on all rows by default diff --git a/docs/parameters.md b/docs/parameters.md index fdfd782..7ca0a95 100644 --- a/docs/parameters.md +++ b/docs/parameters.md @@ -74,6 +74,7 @@ A whole grid of conditional quantiles from one booster. See the User Guide: | `conformalize` | `"auto"` | How the intervals are calibrated. `"auto"` rescales them after the fit on the rows early stopping held out (your `eval_set` or the `validation_fraction` fold), so no extra rows are spent; with no such rows, or too few to certify the levels, you get the raw grid. `True` calibrates on a separate fold held out before the early-stopping split, which gives a distribution-free coverage guarantee and raises if that fold is too small. `False` returns the raw grid, which runs narrow. | | `calibration_fraction` | `0.2` | Rows reserved for calibration. Ignored unless `conformalize=True`. | | `audition` | `True` | Fit five candidates and keep the one with the best CRPS on the rows early stopping held out: the head as configured, the same head with 254 histogram bins, the head's shape moved onto the median of a squared-error `ChimeraBoostRegressor`, that regressor's prediction plus one offset per level (`"fixed"`, the same width for every row), and the same centre plus a second regressor's predicted spread times one factor per level (`"scaled"`). Every prediction method and `shap_values` serve the winner, recorded in `audition_`. About 2.6 times the fit time of one head. It needs early-stopping rows and `conformalize="auto"`; otherwise one head is fitted. `False` fits one head, as releases before 0.33.0 did. | +| `refit_full` | `True` | Once early stopping, calibration and the candidate choice are done on the held-out rows, retrain the winner on all rows, as `ChimeraBoostRegressor` does. Everything chosen on the held-out rows is kept: the candidate, the calibration factors, the offsets, `best_iteration_` and `validation_history_`; `refit_` records what was retrained. It acts only on the fold the model carves for itself: never with your own `eval_set`, with early stopping off, or with `conformalize=True`. About 30% more fit time when it acts. `False` skips it. | | `split_projection` | `"rotate"` | How the K gradient columns collapse for the split search. `"rotate"` measured best; `"sum"` is blind to changes in spread; `"gram"` measured no better than `"rotate"`. | | `exact_splits` | `False` | Score splits on the exact gain summed across levels instead of a projection, at the cost of K histogram channels per feature. Measured worse on real data (a worse CRPS on 27 of 36 Grinsztajn regression datasets), so it is a reference setting only. | diff --git a/docs/quantiles.md b/docs/quantiles.md index aab9060..8b3d6b7 100644 --- a/docs/quantiles.md +++ b/docs/quantiles.md @@ -156,6 +156,31 @@ with early stopping off and no `eval_set`, or with `conformalize` set to `True` `False`, one head is fitted. `audition=False` fits one head, as releases before 0.33.0 did. +## Retraining on all rows + +When you don't pass an `eval_set`, the model holds back `validation_fraction` of your +rows, 20% by default. It uses them to choose the stopping round, calibrate the intervals +and pick the candidate. Then it retrains the winner on all the rows, as +`ChimeraBoostRegressor` does, so the held-back rows still reach the final model. +Everything decided on them stays as it was: the candidate, the calibration factors, the +fixed and scaled candidates' offsets, `best_iteration_` and `validation_history_`. + +```python +model = ChimeraBoostQuantileRegressor().fit(X, y) +model.refit_ # what was retrained, e.g. {"centre": True, "head": False, "rounds": None} +``` + +On the 36 Grinsztajn regression datasets the retrain improves CRPS on 35 and loses on 1, +by 0.5%. The median gain is 0.9%, and six low-noise datasets gain between 3% and 9%. The +intervals come out very slightly wider: a nominal 90% interval covers 90.4% on average +instead of 90.1%. The retrain adds about 30% to the fit time at the median, and +`refit_full=False` skips it. + +It never trains on rows you hold out yourself. With your own `eval_set` nothing is +retrained, and with `conformalize=True` nothing is retrained either, which keeps the +coverage guarantee intact. Every model in the comparison below trains on the same rows, +with the same rows held out, so the table does not include this gain. + ## Scoring `chimeraboost.quantile_metrics` scores a predicted grid. @@ -336,8 +361,9 @@ calibration repairs the tails. Earlier releases used `depth=4` with no calibrati most extreme level on the grid, so a leaf estimating the 5% quantile keeps at least about 20 rows. -When fit time matters more than the last 3%, `audition=False` fits a single head in -about 40% of the time. +When fit time matters most, `refit_full=False` skips the retrain, which gives up a +median 0.9% of CRPS, and `audition=False` fits a single head instead of five candidates, +less than half the work, which gives up a median 3%. `split_projection` chooses how the K gradient columns collapse into the single vector the tree grower accepts. Leave it alone unless you are exploring: `"rotate"` measured