diff --git a/CHANGELOG.md b/CHANGELOG.md index 846ae1f..f999b77 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,24 @@ The format follows [Keep a Changelog](https://keepachangelog.com/). ## [Unreleased] ### Changed +- **The quantile head retrains its winner on all rows by default.** Called + without an `eval_set`, `ChimeraBoostQuantileRegressor` held back + `validation_fraction` of the rows for early stopping, calibration and + the candidate choice, and never trained on them. The new `refit_full`, + on by default, makes every choice on those rows as before and then + retrains the winner on all rows, as `ChimeraBoostRegressor` does. The + candidate, the calibration factors, the fixed and scaled candidates' + offsets, `best_iteration_` and `validation_history_` stay from the + held-out fit, and `refit_` records what was retrained. On the 36 + Grinsztajn regression datasets it improves CRPS on 35 and loses 1 by + 0.5% (median +0.9%; up to +8.6% on Brazilian_houses), and it improves all + 6 high-cardinality ones (median +1.4%). It also shrinks the time-split + caveat in the next entry: the median 90% coverage error there falls + from 4.7 to 3.3 points. It adds about 30% to the fit time where it acts. + Nothing changes with your own `eval_set`, with early stopping off, or + with `conformalize=True`; `refit_full=False` skips it. The comparisons + in `docs/quantiles.md` keep every model on the same rows, so they do not + include this gain. Record: `benchmarks/QUANTILE_PLAN.md`, Q12 and Q13. - **The quantile head's audition gains two candidates, and its CRPS improves most on low-noise targets.** `ChimeraBoostQuantileRegressor`'s default choice now also tries `"fixed"`, an ordinary squared-error diff --git a/benchmarks/CAMPAIGN_PLAN.md b/benchmarks/CAMPAIGN_PLAN.md index 9183530..18205ea 100644 --- a/benchmarks/CAMPAIGN_PLAN.md +++ b/benchmarks/CAMPAIGN_PLAN.md @@ -186,7 +186,7 @@ fact: 2026-09-25 | standing quantile BASE = `results/quantile-20260925-172250.js | F3 | Classifier forced-cross | KILLED 2026-09-21 (S3, I021) | gr binary engaged 10W-13L, median −0.04%: the race earns its fee on the classifier. Knob stays opt-in (PR #117), no rung-1 pin | | F5 | hc-Brier gap vs CatBoost | **SHIPPED 2026-09-22 (I046, PR #145, d14bf38): `cat_count_features` on by default** (public 5W-1L, decide hc 6W-1L, bit-identical elsewhere, chart refreshed) — earlier: I038 library form, I039 S3 PASS | the published chart (`public_pareto.png`) refresh: DONE 2026-09-24 on 0.33.0 (`results/20260924-090541.json`, `--public --seeds 3`, 22 datasets, ~4 h: the default avg rank 1.94 [1.64-2.24] against CatBoost 1.84 [1.43-2.24] and LightGBM 2.23, at a median 5.6x against 47.2x and 1.0x; `docs/benchmarks.md` prose updated; `pareto.png` re-rendered from `results/20260924-130232.json` and `docs/PROJECT_STATUS.md` synced); S7 (multiclass CTR width) is the family's next idea if picked. History: **A per-categorical count column closes 47% of CatBoost's hc edge** (I035): 4W-1L on the gap sets, +0.43% Brier median, gains ordered by cardinality, sf-police and Traffic unanimous across seeds. The gap is the encoder (CatBoost on our TS keeps none of its edge); not the prior target, Counter, permutations or quantization (TS quantization kills on big sets, +1.7–3.6% on the two small controls — a small-data pointer, parked). `cat_count_features` (opt-in, card ≥ 256, invisible to the cross and linear-leaf races; I038) on the decision tier (I039): gr 0-0-59 exact ties, the 7 hc sets without a qualifying column exact ties, the engaged 7 **6W-1L** at +0.20% median (sf-police +0.73%, Traffic +0.86% Brier; employee_salaries +2.45%, wine-reviews +0.56% RMSE), hc@time 4-0, fit ×1.09 on hc (engaged median 1.165). PR up with the flag OFF. The random-effects alternative (per-column ANOVA λ for the TS, I040) KILLED: uncapped it collapses the small controls (−3.8 / −9.6%), capped at 10 it is a flat wash and still costs kick and eucalyptus; the count column keeps evidence the shrinkage deletes. Next: the maintainer's go on /experiment S4 for the default flip; meanwhile R4 S0 | | F6 | Ordered-TS train/test moment mismatch (shortlist R1) | KILLED 2026-09-21 (S1b, I025) — closed as barrier B18 | The defect is real (rare categories over-trusted, reliability 0.63–0.70) and two transform-side fixes both went 7W-5L against a bar of 8: the gain is sf-police (9 of 9 fits, +0.29% to +0.53%) and nothing else. Nothing ships; the open door is the Counter feature, which belongs to R3 | -| F7 | Multi-quantile head (`ChimeraBoostQuantileRegressor`) | ACTIVE again 2026-09-25 (the maintainer: "we can continue on the quantile thread"; PARKED 2026-09-24 for GitHub issues; ACTIVE from 2026-09-23). Phase 1, the bench, MERGED (I056, PR #156); Q0 DONE (I057, PR #157 merged): uncapped and validation-rescaled pay, depth 6 misses on its guard alone, recentred fails. Q3 CLOSED (I058, barrier B23: a larger rate loses); depth 6 + early-stopping-row calibration PAID (31W-5L, +0.33%, 90% error 3.43 → 0.47); PR #158 merged. **Q5 PASSED (I059): the head's default is now depth 6 + `conformalize="auto"`** (gr 31W-5L, +0.33%, 90% coverage error 3.43 → 0.47); PR #159 merged, snapshot rebaselined. **Q2 CLOSED (I060)**: no probe clears the tightened gate (sign p < 0.05); the five low-noise sets need resolution (bins) on two and a squared-error location on two; PR #160 merged. **Q6 T0 PASSED (I061)**: the validation-chosen head wins 23W-7L-6T (p 0.005), recovers 99% of the oracle, and tops the 59-key frontier above CatBoost MQ (0.6006 @ 7.2× against 0.5982 @ 129×); PR #161 merged. **Q7 PASSED (I062): the audition is the head's default** (identity 177/177 with the bench arm; gr 23W-7L-6T against the Q5 default, p 0.005; CatBoost MQ now 9W-27L against us, off the frontier); PR #162 merged, **released in 0.33.0**. **Q8 T0 PASSED (I070)**: a fixed-width (S) and a scaled-residual (N) audition candidate, gr 18W-1L-17T against the Q7 default (p 7.6e-5) at 1.11x the fit; S alone and uncapped rounds fail. **Q9 PASSED (I071): S and N in the library default** (identity 177/177; CatBoost MQ 29W-7L, RigidShift 35W-1L; the 59-key chart 0.6046 @ 8.3x); merged as PR #177 (a6a90d2); flag: hc `@time` 90% coverage error 2.12 -> 4.69 (Moneyball, a pointer). **Q10 T0 (I072) DIAGNOSED**: RoNGBa's lead on visualizing_soil and SGEMM is the S/N centre model's accuracy (our grid on its mean beats it); pol is a width gap. **Q11 T0-a (I073)**: centre settings on five sets: 254 bins matches RoNGBa's lead (visualizing_soil +24%, SGEMM +7%) but costs cpu_act 2.5%; 8000 rounds +5% on visualizing_soil only; bagging +30% (ceiling, not a default). **T0-b (I073b) FAILS**: as extra candidates the 254-bin centre still costs cpu_act 2.5% (validation cannot see the overfit); 8000 rounds closed unrun (the centre is cap-bound on 1 of 59 keys). The RoNGBa line is CLOSED. **Q4 CLOSED (I074)**: the N candidate already is the spread-aware categorical encoding; `catscale` at 10k now 16.0 against CatBoost MQ's 30.9 (excess CRPS x1000). **I063 DONE (PR #167)**: RoNGBa is the NGBoost opponent (issue #163) | next: the maintainer's pick between Q1 (P16; coverage already solved by calibration, a small CRPS claim) and the no-`eval_set` refit design (`QUANTILE_PLAN.md`, Still open); until then F7 holds. Program: `QUANTILE_PLAN.md` "Campaign 2026-09-23" | +| F7 | Multi-quantile head (`ChimeraBoostQuantileRegressor`) | ACTIVE again 2026-09-25 (the maintainer: "we can continue on the quantile thread"; PARKED 2026-09-24 for GitHub issues; ACTIVE from 2026-09-23). Phase 1, the bench, MERGED (I056, PR #156); Q0 DONE (I057, PR #157 merged): uncapped and validation-rescaled pay, depth 6 misses on its guard alone, recentred fails. Q3 CLOSED (I058, barrier B23: a larger rate loses); depth 6 + early-stopping-row calibration PAID (31W-5L, +0.33%, 90% error 3.43 → 0.47); PR #158 merged. **Q5 PASSED (I059): the head's default is now depth 6 + `conformalize="auto"`** (gr 31W-5L, +0.33%, 90% coverage error 3.43 → 0.47); PR #159 merged, snapshot rebaselined. **Q2 CLOSED (I060)**: no probe clears the tightened gate (sign p < 0.05); the five low-noise sets need resolution (bins) on two and a squared-error location on two; PR #160 merged. **Q6 T0 PASSED (I061)**: the validation-chosen head wins 23W-7L-6T (p 0.005), recovers 99% of the oracle, and tops the 59-key frontier above CatBoost MQ (0.6006 @ 7.2× against 0.5982 @ 129×); PR #161 merged. **Q7 PASSED (I062): the audition is the head's default** (identity 177/177 with the bench arm; gr 23W-7L-6T against the Q5 default, p 0.005; CatBoost MQ now 9W-27L against us, off the frontier); PR #162 merged, **released in 0.33.0**. **Q8 T0 PASSED (I070)**: a fixed-width (S) and a scaled-residual (N) audition candidate, gr 18W-1L-17T against the Q7 default (p 7.6e-5) at 1.11x the fit; S alone and uncapped rounds fail. **Q9 PASSED (I071): S and N in the library default** (identity 177/177; CatBoost MQ 29W-7L, RigidShift 35W-1L; the 59-key chart 0.6046 @ 8.3x); merged as PR #177 (a6a90d2); flag: hc `@time` 90% coverage error 2.12 -> 4.69 (Moneyball, a pointer). **Q10 T0 (I072) DIAGNOSED**: RoNGBa's lead on visualizing_soil and SGEMM is the S/N centre model's accuracy (our grid on its mean beats it); pol is a width gap. **Q11 T0-a (I073)**: centre settings on five sets: 254 bins matches RoNGBa's lead (visualizing_soil +24%, SGEMM +7%) but costs cpu_act 2.5%; 8000 rounds +5% on visualizing_soil only; bagging +30% (ceiling, not a default). **T0-b (I073b) FAILS**: as extra candidates the 254-bin centre still costs cpu_act 2.5% (validation cannot see the overfit); 8000 rounds closed unrun (the centre is cap-bound on 1 of 59 keys). The RoNGBa line is CLOSED. **Q4 CLOSED (I074)**: the N candidate already is the spread-aware categorical encoding; `catscale` at 10k now 16.0 against CatBoost MQ's 30.9 (excess CRPS x1000). **I063 DONE (PR #167)**: RoNGBa is the NGBoost opponent (issue #163) | next: **Q13 (I076) PASSED, PR for the maintainer**: `refit_full=True` as the head's default (gr 35W-1L, +0.91%, p 1.1e-9, 1.28x fit); after the merge, Q1 (P16). Program: `QUANTILE_PLAN.md` "Campaign 2026-09-23" | ### F1 — Cross-feature cost trim v2 status: KILLED 2026-08-16 at S2 (I007 at k=6, I008 at k=12) — closed as barrier B16 @@ -453,6 +453,131 @@ Recommended pick: **R1, R2, R3, R4 + H(1)(4)(5)**. R1 and R2 have free probes an ## Iteration log (append-only) +#### I076 2026-09-25 F7 Q13 (`refit_full=True` becomes the head's default; LIBRARY change, pre-registered) +why now: I075's arm (b) passed the gate and beat arm (a) head to head. +change (`chimeraboost/quantile_api.py`): `refit_full` defaults to True +with arm (b)'s behaviour; the private `_refit_scope` and its "centre" +branch are removed (the probe answered); the retrained centre gets the +head's `validation_fraction` so its carve always equals the head's. +Benchmark: the field arm `ChimeraBoostQuantile` keeps passing the shared +split as an `eval_set`, which the retrain never touches, so every +comparison in the suite and the docs stays "every arm on the same rows" +(conservative: the table does not include the retrain's gain). The probe +arm `...RefitAll` stays, as the default called without an `eval_set`; +`...RefitCentre` is deleted with its switch. +forecast: (1) the default called without an `eval_set` equals I075's +RefitAll arm bit for bit, and with an `eval_set` equals today's default +bit for bit; (2) tests green, with the tests whose premise was no retrain +pinned to `refit_full=False` (listed); (3) the identity snapshot moves +only on quantile configs fitted without an `eval_set`; (4) the I075 +numbers stand for the new default without another run. +docs (Claude): `docs/quantiles.md` (a paragraph on the retrain: what it +does, what it keeps from the held-out fit, its measured gain and cost, +`refit_full=False`), `docs/parameters.md` (the `refit_full` row), +CHANGELOG (Unreleased, Changed). +muse (exit 0, `20260925-quantile-q13-refit-default.md`): `refit_full=True` +by default; `_refit_scope` and the centre-only branch removed; the +retrained centre gets the head's `validation_fraction` (a new test at 0.3 +fails against a 0.2 carve, so it matters off the default); docstrings; +`...RefitAll` renamed `...AllRows` (the default on `split.full`, no +`eval_set`), `...RefitCentre` deleted, probe count 18; tests: 7 changed +(one pinned to `refit_full=False`: the carve-premise test, whose two fits +now differ by design), 1 added, 1 deleted; 1235 passed; ruff clean. +checks (Claude): the default without an `eval_set` reproduces I075's +RefitAll CRPS exactly on 8 of 8 seed-0 fits covering H, B, R and N +winners (pol, visualizing_soil, cpu_act, colleges, analcatdata_supreme, +house_prices_nominal, sulfur@sus25, Mercedes_Benz), and the field arm +reproduces the standing BASE on the same 8. Identity snapshot 182/186: +the moved pins are `mq3` and `mq3_w_sub`'s predictions and importances +(quantile configs fitted without an `eval_set`); their calibration +factors do not move. Rebaseline after the merge. +forecast: (1) HIT; (2) HIT; (3) HIT; (4) HIT (I075's numbers, reproduced +exactly). +verdict: **PASS -> PR for the maintainer** (library, tests, docs). After +the merge: rebaseline the identity snapshot, delete the branch. + +#### I075 2026-09-25 F7 Q12 T0 (the head retrains its winner on all rows after the audition, as the point regressor's `refit_full` does; BENCH-only probe arms, pre-registered) +why now: the maintainer, 2026-09-25, picking option 1 of three: "1 is +fine". Called without an `eval_set`, the head holds back +`validation_fraction` of the rows and never trains on them; the point +regressor retrains on 100% after early stopping (`refit_full`). Checked +first (`scratchpad/carve_check.py`, visualizing_soil and pol, seed 0): +without an `eval_set` both the regressor and the head carve exactly the +suite's shared split (predictions bit-identical to the `eval_set` fits), +so today's field arm IS the library called without an `eval_set`, and the +regressor's own retrain cuts its test RMSE 4.8% and 5.1% there. +barriers: B2 (the rung-3 refit amplifies a bad audition) is about the +point model's audition budget; here the choice is untouched and only the +winner is retrained. B13 and the `refit_full` docstring's reason for +skipping `loss="Quantile"` (keep the conformal holdout honest): every +calibration quantity (factors, offsets, `q_s`, `q_n`, the spread model) +stays from the 80% fit, computed on rows it never trained on; only the +delivered centre (and in arm (b) the head) is retrained, so intervals may +run slightly wide, which the coverage guard reads. +arms (`quantile_suite.py`, muse task `20260925-quantile-q12-refit-probe.md`): +(a) `ChimeraBoostQuantileRefitCentre`: `_fit_audition(with_s, with_n)` +unchanged on the shared split; when R, S or N wins, its centre is replaced +by the same `ChimeraBoostRegressor` fitted on all training rows without an +`eval_set` (it carves the same split and retrains with its default replay +refit). Offsets, spread model, `q_s`, `q_n` and factors stay. H and B +winners are unchanged. (b) `ChimeraBoostQuantileRefitAll`: (a), plus H, B +and R's head retrained from scratch on all rows with early stopping off, +`n_estimators = min(ceil(t_star / 0.8), 2000)` (`_refit_on_full`'s rule, +t_star the trees the early-stopped head kept), rate pinned, the 80% fit's +factors applied. fit_s counts only what the library would add (the +retrain, not a repeat of the 80% fit). +forecast: (1) (a) engages where R, S or N wins (about half the gr fits), +wins >= 70% of the engaged sets (median +1% to +3%), gr sign p < 0.05, +coverage guard ok, fit +5% to +10%. (2) (b) engages everywhere, gr >= 26 +of 36 (median +1% to +2%), p < 0.05, fit +25% to +45%. (3) visualizing_soil +>= +5% and pol >= +3% under (a). (4) hc no worse than gr (a pointer). +decision: an arm passing the default gate (gr two-sided sign p < 0.05, +median > 0, coverage guard ok) becomes the library rung Q13 (the Q7/Q9 +pattern, the arm its exact oracle). Both pass: (b) only if it also beats +(a) head to head on gr at p < 0.05; otherwise (a). Neither: the question +closes. +muse (exit 0, `20260925-quantile-q12-refit-probe.md`): `refit_full=False` +on the head (the private `_refit_scope` for arm (a)); the centre retrained +as `ChimeraBoostRegressor().fit(all +rows)` (no `eval_set`: the same carve, its own replay refit), the head +from scratch at `min(ceil(len(trees_) / 0.8), n_estimators)` rounds with +`lr_` pinned; the early-stopped curves and `best_iteration_` carried over; +`refit_` records it. The suite passes `split.full` (a tuple subclass, the +original row order); arms `...RefitCentre`, `...RefitAll`; 11 tests; +1235 passed; ruff clean; identity snapshot 186/186 (off by default). +Review: correct; one nit for the library rung, the retrained centre uses +the regressor's own `validation_fraction` (0.2) rather than the head's. +Commit 7c2c9e9. +ran: `quantile_suite.py --decide --seeds 3 --jobs 5 --models +ChimeraBoostQuantile ...RefitCentre ...RefitAll --save` +(`results/quantile-20260925-202430.json`, 25 min). Against the default, +same run (median over decided sets; coverage = median |cov - nominal| +points at 90 / 80): + +| stratum | (a) RefitCentre | (b) RefitAll | +|---|---|---| +| gr (36) | 19W-0L-17T, +1.72%, p 3.8e-6; cov 0.32 -> 0.40 / 0.37 -> 0.50; fit 1.07x | 35W-1L, +0.91%, p 1.1e-9; cov 0.32 -> 0.47 / 0.37 -> 0.59; fit 1.28x | +| hc (6) | 5W-0L-1T, +1.20% | 6W-0L, +1.42% | +| gr@sus25 (7) | 6W-0L-1T, +1.01% | 7W-0L, +1.63% | +| gr@sus50 (4) | 2W-0L-2T | 3W-1L | +| hc@sus25 / hc@sus50 | 2-0 / 1-0 | 2-0 / 1-0 | +| hc@time (3) | 3W-0L, +4.10%; cov 4.69 -> 3.29 / 4.68 -> 2.62 | same | + +(b) against (a): gr 20W-1L-15T, +0.32%, p 2.1e-5, at 1.13x (a)'s fit. +Largest gains (both): Brazilian_houses +8.6% / +8.2%, visualizing_soil ++7.7%, pol +5.1%, sulfur +4.6 to +5.0%, superconduct +3.8%; (b)'s one loss +analcatdata_supreme -0.48%. Over 59 keys: skill 0.606 -> 0.609 (a) / +0.611 (b); mean 90% coverage 0.898 -> 0.900 / 0.901. Retrained: (b) the +centre on 99 fits (R 6, S 13, N 80), the head on 84 (H 38, B 40, R 6). +forecast: (1) HIT on every item (19 of 19 engaged, +1.72%, p 3.8e-6, +guard ok, fit 1.07x). (2) HALF: engages everywhere HIT, 35 of 36 HIT, +p HIT, fit 1.28x HIT, median +0.91% MISS (forecast +1% to +2%). (3) HIT +(visualizing_soil +7.7%, pol +5.1%). (4) HIT (hc 5-0 / 6-0). +unforecast: the retrain also repairs part of the hc `@time` coverage flag +(I071: 4.69 -> 3.29 points at 90%). +verdict: **both arms PASS; (b) RefitAll beats (a) at p 2.1e-5, so by the +pre-registered rule (b) is the library rung** (Q13, I076). + #### I074 2026-09-25 F7 Q4 T0-a (does the five-candidate default still lose `catscale` to CatBoost? a synthetic recheck before any spread-TS work; bench-only, pre-registered) why now: Q4's premise (2026-09-23, `results/quantile-synth-20260923-090818.json`) is that the ordered TS is a per-category MEAN, blind to a category that diff --git a/benchmarks/QUANTILE_PLAN.md b/benchmarks/QUANTILE_PLAN.md index e0499d6..81883b4 100644 --- a/benchmarks/QUANTILE_PLAN.md +++ b/benchmarks/QUANTILE_PLAN.md @@ -424,6 +424,23 @@ Q0 reads moved it down to item 4. 1 of the 59 decide keys, so it cannot clear the gate. Bagging the centre ×5 is the ceiling (visualizing_soil +30%, pol +9%) and is an ensemble, so not a default. +3g. **Q12, retraining the winner on all rows (the maintainer's pick + 2026-09-25).** Without an `eval_set` the head's own carve equals the + suite's shared split bit for bit, so the probe is a retrain added to + today's fit: every choice and calibration quantity from the held-out + fit, then (a) the R/S/N centre retrained on all rows, or (b) that plus + the H/B/R head retrained from scratch at the replay-round rule. + **PASSED 2026-09-25 (I075, `results/quantile-20260925-202430.json`).** + Against the default: (a) gr 19W-0L-17T, +1.72%, p 3.8e-6, 1.07× fit; + (b) gr 35W-1L, +0.91%, p 1.1e-9, hc 6W-0L, 1.28× fit; (b) beats (a) + 20W-1L-15T (p 2.1e-5), so (b). Coverage guard fine (gr 90% error 0.32 → + 0.47 points); the hc `@time` flag shrinks (4.69 → 3.29). +3h. **Q13, `refit_full=True` as the default.** + **PASSED 2026-09-25 (I076), PR for the maintainer.** Without an + `eval_set` the default reproduces Q12's arm (b) bit for bit; with one, + nothing changes. The suite's field arm keeps its `eval_set`, so every + comparison stays "every model on the same rows" and leaves this gain + out. 4. **Q1, the narrow-interval defect (P16).** Leaf values are in-sample residual quantiles, so intervals over-narrow (0.869 at nominal 0.90 on 2026-08-30; coverage decays with rounds). Fit leaf quantiles @@ -533,7 +550,7 @@ offset on real data (a CRPS tie), which made it Q0's subject. against pinball. That is a default flip on a strength surface, so it needs its own pre-registration and the full `/experiment` protocol. Not attempted here. Recorded 2026-08-30. -- **The head never trains on its own early-stopping fold** (recorded +- RESOLVED 2026-09-25 (Q12 and Q13, I075 and I076): **The head never trains on its own early-stopping fold** (recorded 2026-09-25, I073b). Without an `eval_set` the head carves `validation_fraction` and neither it nor its S/N centre ever sees those rows, while `ChimeraBoostRegressor` refits on all rows by default @@ -541,8 +558,10 @@ offset on real data (a CRPS tie), which made it Q0's subject. refit of the winner after the audition, keeping the calibration taken before it, is the candidate. `quantile_suite.py` cannot measure it: it passes the shared split as an `eval_set`, which the head must not train - on. Needs a no-`eval_set` protocol first; the maintainer's call whether - to open it. + on. Needs a no-`eval_set` protocol first. **OPENED 2026-09-25 as Q12** + (the maintainer picked it over Q1): the head's own carve equals the + suite's shared split bit for bit, so the probe is a retrain added to + today's arm (`CAMPAIGN_PLAN.md` I075). - RESOLVED 2026-09-23 (Q5, I059): `docs/quantiles.md` "How it compares" is re-measured against the new default, with the fixed-width baseline and NGBoost added. diff --git a/benchmarks/quantile_suite.py b/benchmarks/quantile_suite.py index 6dba34d..fbeeb98 100644 --- a/benchmarks/quantile_suite.py +++ b/benchmarks/quantile_suite.py @@ -96,6 +96,24 @@ RONGBA_LEAVES = 31 +class _SplitWithFull(tuple): + """The shared ``(Xf, Xv, yf, yv)`` split, carrying the training rows. + + Unpacks and indexes as the plain 4-tuple every arm already takes; + ``.full`` is the ``(Xtr, ytr)`` the shared split was carved from, in + their original order, so a probe arm can fit on all training rows + while the library's own carve reproduces the shared split. + """ + + def __new__(cls, split, full): + obj = super().__new__(cls, split) + obj.full = full + return obj + + def __init__(self, split, full): + pass + + def _fit_head_model(split, cat, threads, taus, n_estimators, **params): """Construct and fit the suite's head. @@ -811,6 +829,41 @@ def _fit_chimera_default_uncapped(split, Xte, cat, threads, taus): return Q, fit_s, time.time() - t, m.best_iteration_ +_AUDITION_CHOICE = {"head": 0, "bins": 1, "recentred": 2, + "fixed": 3, "scaled": 4} + + +def _refit_extras(m): + """The refit probes' extras: the audition choice as the audition arms + encode it (0-4 for H, B, R, S, N), whether the centre and the head + were retrained, and the head's retrained round count.""" + aud = m.audition_ or {} + refit = m.refit_ or {} + return {"audition_choice": _AUDITION_CHOICE.get(aud.get("selected")), + "refit_centre": bool(refit.get("centre", False)), + "refit_head": bool(refit.get("head", False)), + "refit_rounds": refit.get("rounds")} + + +def _fit_chimera_all_rows(split, Xte, cat, threads, taus): + """The library default as a user without an ``eval_set`` gets it: + fitted on ALL training rows, so the audition runs on the library's + own carve (which reproduces the shared split) and the winner is + retrained on every row.""" + Xtr, ytr = split.full + m = ChimeraBoostQuantileRegressor( + quantiles=taus, n_estimators=rb.MAX_ITERS, + early_stopping_rounds=rb.PATIENCE, thread_count=threads, + random_state=0) + t = time.time() + m.fit(Xtr, ytr, cat_features=cat or None) + fit_s = time.time() - t + t = time.time() + Q = m.predict(Xte) + pred_s = time.time() - t + return Q, fit_s, pred_s, m.best_iteration_, _refit_extras(m) + + PROBES = { "ChimeraBoostQuantileUncapped": _fit_chimera_uncapped, "ChimeraBoostQuantileDepth6": _fit_chimera_depth6, @@ -829,6 +882,7 @@ def _fit_chimera_default_uncapped(split, Xte, cat, threads, taus): "ChimeraBoostQuantileAuditionS": _fit_chimera_audition_s, "ChimeraBoostQuantileAuditionSN": _fit_chimera_audition_sn, "ChimeraBoostQuantileDefaultUncapped": _fit_chimera_default_uncapped, + "ChimeraBoostQuantileAllRows": _fit_chimera_all_rows, } @@ -944,7 +998,8 @@ def run_one(ds_name, seed, taus, threads, models): "regression") # One early-stopping split, shared by every arm, so no model is judged on # more data than another. Same carve the harness uses. - split = rb._val_split(Xtr, ytr, "regression", 0) + split = _SplitWithFull(rb._val_split(Xtr, ytr, "regression", 0), + (Xtr, ytr)) meta = {"task": "quantile", "n_train": int(Xtr.shape[0]), "n_total": int(X.shape[0]), "n_features": int(X.shape[1]), diff --git a/benchmarks/quantile_synth.py b/benchmarks/quantile_synth.py index 77fb8bb..72bacb5 100644 --- a/benchmarks/quantile_synth.py +++ b/benchmarks/quantile_synth.py @@ -219,7 +219,8 @@ def run_one(regime, n, seed, taus, threads, models): random_state=seed) Xtr, Xte, ytr, yte = X[idx_tr], X[idx_te], y[idx_tr], y[idx_te] Qo_te = Q_or[idx_te] - split = rb._val_split(Xtr, ytr, "regression", 0) + split = qs._SplitWithFull(rb._val_split(Xtr, ytr, "regression", 0), + (Xtr, ytr)) cat = [X.shape[1] - 1] if regime == "catscale" else None oracle_crps = float(qm.crps(yte, Qo_te, taus)) diff --git a/chimeraboost/quantile_api.py b/chimeraboost/quantile_api.py index 3b6646f..5c0de27 100644 --- a/chimeraboost/quantile_api.py +++ b/chimeraboost/quantile_api.py @@ -2,8 +2,8 @@ Kept out of ``sklearn_api`` because it shares that module's input validation but almost none of its fit machinery: no loss family, no linear leaves, no -cross features, no bagging, no full-data refit. It imports the validation -helpers and keeps the same flat, module-function style. +cross features, no bagging, and a full-data refit that is on by default. It +imports the validation helpers and keeps the same flat, module-function style. """ import warnings @@ -265,6 +265,15 @@ def _pick_audition_winner(ordered): return best_name +def _keep_es_state(new_booster, old_booster): + """Keep the early-stopped fit's visible state on a refit replacement, + as ``sklearn_api._refit_on_full`` does: the training and validation + curves, and the budget early stopping chose.""" + new_booster.train_history_ = old_booster.train_history_ + new_booster.valid_history_ = old_booster.valid_history_ + new_booster.best_iteration_ = old_booster.best_iteration_ + + def _phi_centre(phi, mi, mw): """Centre of attributions along the level axis, mirroring ``_centre``.""" if mw == 0.0: @@ -338,13 +347,34 @@ class ChimeraBoostQuantileRegressor(BaseEstimator): calibrates the head and scored unweighted; ties go H, then B, then R, then S, then N. Needs ``conformalize="auto"`` and evaluation rows (the user's ``eval_set`` or the carved fold); otherwise the - single head is fitted. Costs about 2.6x a single head at the median: - S reuses R's centre fit, so the extra cost over the three-candidate - audition is the spread-model fit, about 10% of the fit at the - median. ``False`` fits a single head, as releases before 0.33.0 did. + single head is fitted. Costs about 2.6x a single head at the median + for fits with an ``eval_set``: S reuses R's centre fit, so the extra + cost over the three-candidate audition is the spread-model fit, + about 10% of the fit at the median. The full-data refit adds about + 30% at the median when it acts. ``False`` fits a single head, as + releases before 0.33.0 did. ``model_`` stays the fitted head booster whatever wins (the 254-bin booster for a bins win); for S and N it delivers nothing -- ``predict`` serves the centre/spread models -- but stays inspectable. + refit_full : bool, default True + Retrain the winner on all rows after early stopping chose the + budget, as ``ChimeraBoostRegressor`` does. Acts only on the + automatic early-stopping split: never with a user ``eval_set``, + with early stopping off, with ``conformalize=True``, or when the + carve failed -- those fits are exactly what they are without it. + An R, S or N winner's centre is then replaced by a default + ``ChimeraBoostRegressor`` fitted on all rows without an + ``eval_set`` -- offsets, spread model, residual quantiles, floor + and calibration factors stay from the audition fit -- and the + head booster of an H, B or R winner is retrained from scratch on + all rows with early stopping off, at ``min(ceil(t_star / (1 - + validation_fraction)), n_estimators)`` rounds with the + early-stopped learning rate pinned, where t_star is the rounds + the early-stopped head kept. ``audition_``, + ``conformal_scale_``, ``best_iteration_`` and + ``validation_history_`` keep the early-stopped fit's values; + ``refit_`` records what was retrained. ``False`` skips the + retrain. Attributes ---------- @@ -364,6 +394,12 @@ class ChimeraBoostQuantileRegressor(BaseEstimator): "crps": {name: score}}`` for the candidates that ran, or ``None`` when no audition ran (no evaluation rows, ``conformalize`` True or False, or ``audition=False``). + refit_ : dict or None + ``{"centre": bool, "head": bool, "rounds": int or None}`` -- + whether the centre and the head booster were retrained on all + rows, and the head's retrained round count (None when the head + was not retrained). None when nothing was retrained + (``refit_full=False``, or a fit the refit does not cover). Notes ----- @@ -383,7 +419,8 @@ def __init__(self, quantiles=None, n_estimators=2000, learning_rate=None, quantize_gradients=True, early_stopping=True, validation_fraction=0.2, split_projection="rotate", exact_splits=False, conformalize="auto", - calibration_fraction=0.2, audition=True): + calibration_fraction=0.2, audition=True, + refit_full=True): self.quantiles = quantiles self.n_estimators = n_estimators self.learning_rate = learning_rate @@ -409,6 +446,7 @@ def __init__(self, quantiles=None, n_estimators=2000, learning_rate=None, self.conformalize = conformalize self.calibration_fraction = calibration_fraction self.audition = audition + self.refit_full = refit_full def __sklearn_is_fitted__(self): return hasattr(self, "model_") @@ -467,6 +505,82 @@ def _carve_es_split(self, X, y, sample_weight, eval_set, groups): es_rounds = 50 return X, y, sample_weight, eval_set, es_rounds + def _refit_full_active(self, auto_split): + """True when the full-data refit fires: asked for, on the automatic + split, with no honest holdout at stake (``conformalize=True`` keeps + its pristine fold instead).""" + return (bool(self.refit_full) and bool(auto_split) + and self.conformalize is not True) + + def _maybe_refit_full(self, auto_split, taus, es_rounds, cat_features, + X_full, y_full, sw_full, groups_full): + """Retrain the winner on all rows, after the audition (or the single + head's calibration) finished exactly as without the refit. The + centre of an R, S or N winner is replaced by a default + ``ChimeraBoostRegressor`` fitted on the full rows without an + ``eval_set``; the head booster of an H, B or R winner -- or of the + single head when no audition ran -- is retrained from scratch with + early stopping off. Records what happened in ``refit_``.""" + if not self._refit_full_active(auto_split): + return + selected = self._audition_selected() + centre_done = False + if selected in ("recentred", "fixed", "scaled"): + self._refit_audition_centre(es_rounds, cat_features, X_full, + y_full, sw_full, groups_full) + centre_done = True + head_done, rounds = False, None + if selected in (None, "head", "bins", "recentred"): + rounds = self._refit_head_booster(taus, selected, cat_features, + X_full, y_full, sw_full) + head_done = True + if not centre_done and not head_done: + return + self.refit_ = {"centre": centre_done, "head": head_done, + "rounds": rounds} + + def _refit_audition_centre(self, es_rounds, cat_features, X_full, + y_full, sw_full, groups_full): + """Replace the R/S/N centre with the same regressor fitted on the + full rows without an ``eval_set``, so it carves the same split and + retrains with its own default refit. Offset, spread model, + residual quantiles, floor and factors stay from the audition fit; + the early-stopped curves stay on the replacement, as the + regressor's own refit keeps them.""" + from .sklearn_api import ChimeraBoostRegressor + old = self._centre_model_ + centre = ChimeraBoostRegressor( + n_estimators=self.n_estimators, + early_stopping_rounds=es_rounds, + thread_count=self.thread_count, random_state=self.random_state, + validation_fraction=self.validation_fraction) + centre.fit(X_full, y_full, cat_features=cat_features, + sample_weight=sw_full, groups=groups_full) + _keep_es_state(centre.model_, old.model_) + self._centre_model_ = centre + + def _refit_head_booster(self, taus, selected, cat_features, X_full, + y_full, sw_full): + """Retrain the winner's head booster from scratch on the full rows: + the same construction (254 bins for a bins win), early stopping + off, ``min(ceil(t_star / (1 - validation_fraction)), + n_estimators)`` rounds with the early-stopped learning rate pinned. + Returns the retrained round count.""" + winner = self.model_ + t_star = len(winner.trees_) + frac = max(1.0 - float(self.validation_fraction), 1e-9) + rounds = min(int(np.ceil(t_star / frac)), int(self.n_estimators)) + booster = self._make_mq_booster( + taus, None, cat_features, len(y_full), + max_bins=254 if selected == "bins" else None) + booster.n_estimators = rounds + booster.learning_rate = float(winner.lr_) + booster.fit(X_full, y_full, cat_features=cat_features, + sample_weight=sw_full) + _keep_es_state(booster, winner) + self.model_ = booster + return rounds + def _make_mq_booster(self, taus, es_rounds, cat_features, n_rows, max_bins=None): """Resolve the auto defaults and construct the booster.""" @@ -530,6 +644,7 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None, self._median_idx_ = _median_index(taus) self.conformal_scale_ = np.ones(taus.shape[0]) self.audition_ = None + self.refit_ = None self._centre_model_ = None self._centre_off_ = None self._fixed_q_ = None @@ -544,8 +659,15 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None, cal, X, y, sample_weight, groups = self._carve_calibration_fold( X, y, sample_weight, groups) + # Kept for the optional full-data refit below: the auto split + # reassigns X/y, but the refit retrains on every row. + X_full, y_full, sw_full, groups_full = X, y, sample_weight, groups + had_user_eval = eval_set is not None X, y, sample_weight, eval_set, es_rounds = self._carve_es_split( X, y, sample_weight, eval_set, groups) + # The carve is the only path that sets eval_set when the user did + # not, so this is True exactly when the automatic split was used. + auto_split = not had_user_eval and eval_set is not None self.model_ = self._make_mq_booster(taus, es_rounds, cat_features, len(y)) @@ -573,6 +695,9 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None, except ValueError: pass # uncertifiable: keep the raw grid, never raise + self._maybe_refit_full(auto_split, taus, es_rounds, cat_features, + X_full, y_full, sw_full, groups_full) + return self def _wants_audition(self, eval_set): diff --git a/docs/parameters.md b/docs/parameters.md index fdfd782..7ca0a95 100644 --- a/docs/parameters.md +++ b/docs/parameters.md @@ -74,6 +74,7 @@ A whole grid of conditional quantiles from one booster. See the User Guide: | `conformalize` | `"auto"` | How the intervals are calibrated. `"auto"` rescales them after the fit on the rows early stopping held out (your `eval_set` or the `validation_fraction` fold), so no extra rows are spent; with no such rows, or too few to certify the levels, you get the raw grid. `True` calibrates on a separate fold held out before the early-stopping split, which gives a distribution-free coverage guarantee and raises if that fold is too small. `False` returns the raw grid, which runs narrow. | | `calibration_fraction` | `0.2` | Rows reserved for calibration. Ignored unless `conformalize=True`. | | `audition` | `True` | Fit five candidates and keep the one with the best CRPS on the rows early stopping held out: the head as configured, the same head with 254 histogram bins, the head's shape moved onto the median of a squared-error `ChimeraBoostRegressor`, that regressor's prediction plus one offset per level (`"fixed"`, the same width for every row), and the same centre plus a second regressor's predicted spread times one factor per level (`"scaled"`). Every prediction method and `shap_values` serve the winner, recorded in `audition_`. About 2.6 times the fit time of one head. It needs early-stopping rows and `conformalize="auto"`; otherwise one head is fitted. `False` fits one head, as releases before 0.33.0 did. | +| `refit_full` | `True` | Once early stopping, calibration and the candidate choice are done on the held-out rows, retrain the winner on all rows, as `ChimeraBoostRegressor` does. Everything chosen on the held-out rows is kept: the candidate, the calibration factors, the offsets, `best_iteration_` and `validation_history_`; `refit_` records what was retrained. It acts only on the fold the model carves for itself: never with your own `eval_set`, with early stopping off, or with `conformalize=True`. About 30% more fit time when it acts. `False` skips it. | | `split_projection` | `"rotate"` | How the K gradient columns collapse for the split search. `"rotate"` measured best; `"sum"` is blind to changes in spread; `"gram"` measured no better than `"rotate"`. | | `exact_splits` | `False` | Score splits on the exact gain summed across levels instead of a projection, at the cost of K histogram channels per feature. Measured worse on real data (a worse CRPS on 27 of 36 Grinsztajn regression datasets), so it is a reference setting only. | diff --git a/docs/quantiles.md b/docs/quantiles.md index aab9060..8b3d6b7 100644 --- a/docs/quantiles.md +++ b/docs/quantiles.md @@ -156,6 +156,31 @@ with early stopping off and no `eval_set`, or with `conformalize` set to `True` `False`, one head is fitted. `audition=False` fits one head, as releases before 0.33.0 did. +## Retraining on all rows + +When you don't pass an `eval_set`, the model holds back `validation_fraction` of your +rows, 20% by default. It uses them to choose the stopping round, calibrate the intervals +and pick the candidate. Then it retrains the winner on all the rows, as +`ChimeraBoostRegressor` does, so the held-back rows still reach the final model. +Everything decided on them stays as it was: the candidate, the calibration factors, the +fixed and scaled candidates' offsets, `best_iteration_` and `validation_history_`. + +```python +model = ChimeraBoostQuantileRegressor().fit(X, y) +model.refit_ # what was retrained, e.g. {"centre": True, "head": False, "rounds": None} +``` + +On the 36 Grinsztajn regression datasets the retrain improves CRPS on 35 and loses on 1, +by 0.5%. The median gain is 0.9%, and six low-noise datasets gain between 3% and 9%. The +intervals come out very slightly wider: a nominal 90% interval covers 90.4% on average +instead of 90.1%. The retrain adds about 30% to the fit time at the median, and +`refit_full=False` skips it. + +It never trains on rows you hold out yourself. With your own `eval_set` nothing is +retrained, and with `conformalize=True` nothing is retrained either, which keeps the +coverage guarantee intact. Every model in the comparison below trains on the same rows, +with the same rows held out, so the table does not include this gain. + ## Scoring `chimeraboost.quantile_metrics` scores a predicted grid. @@ -336,8 +361,9 @@ calibration repairs the tails. Earlier releases used `depth=4` with no calibrati most extreme level on the grid, so a leaf estimating the 5% quantile keeps at least about 20 rows. -When fit time matters more than the last 3%, `audition=False` fits a single head in -about 40% of the time. +When fit time matters most, `refit_full=False` skips the retrain, which gives up a +median 0.9% of CRPS, and `audition=False` fits a single head instead of five candidates, +less than half the work, which gives up a median 3%. `split_projection` chooses how the K gradient columns collapse into the single vector the tree grower accepts. Leave it alone unless you are exploring: `"rotate"` measured diff --git a/tests/test_quantile_head.py b/tests/test_quantile_head.py index 61111ae..315a889 100644 --- a/tests/test_quantile_head.py +++ b/tests/test_quantile_head.py @@ -1051,3 +1051,170 @@ def test_audition_fixed_and_scaled_keep_the_head_booster(): raw = c[:, None] + s[:, None] * np.asarray( m._scaled_q_, dtype=np.float64)[None, :] np.testing.assert_array_equal(m.predict(Xt), _hand_rescaled(raw, m)) + + +# -------------------------------------------------------------------------- +# The full-data refit (Q13): on by default, audition-preserving. +# -------------------------------------------------------------------------- + + +def _aud_nosplit_fit(kind, **kw): + """Fit on the kind's training rows WITHOUT an eval_set, so the library + carves its own split; return (model, Xtr, ytr, Xte, yte).""" + from sklearn.model_selection import train_test_split + X, y, sseed = _aud_data(kind) + Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, + random_state=sseed) + m = ChimeraBoostQuantileRegressor( + quantiles=_AUD_TAUS, n_estimators=300, early_stopping_rounds=50, + thread_count=1, random_state=0, **kw) + return m.fit(Xtr, ytr), Xtr, ytr, Xte, yte + + +def test_no_eval_set_equals_suite_split_fit(): + """Without an eval_set the head carves the suite's own split: bit for + bit the same fit as training on the carved rows with that eval_set + (the refit pinned off, so the carve is what is compared).""" + from sklearn.model_selection import train_test_split + rng = np.random.default_rng(60) + n = 1500 + num = rng.standard_normal((n, 3)) + cat = rng.integers(0, 6, n).astype(object) + X = np.column_stack([num, cat]) + y = (2.0 * num[:, 0] + (cat.astype(float) - 2.5) + + rng.standard_normal(n)) + Xf, Xv, yf, yv = train_test_split(X, y, test_size=0.2, random_state=0) + kw = dict(quantiles=_AUD_TAUS, n_estimators=300, + early_stopping_rounds=50, thread_count=1, random_state=0, + refit_full=False) + m_auto = ChimeraBoostQuantileRegressor(**kw).fit(X, y, cat_features=[3]) + m_split = ChimeraBoostQuantileRegressor(**kw).fit( + Xf, yf, cat_features=[3], eval_set=(Xv, yv)) + assert m_auto.audition_ is not None + assert m_auto.audition_ == m_split.audition_ + assert np.array_equal(m_auto.conformal_scale_, m_split.conformal_scale_) + assert np.array_equal(m_auto.predict(X), m_split.predict(X)) + + +def test_refit_full_true_equals_default(): + """refit_full=True is today's fit, exactly; the flag is a visible + constructor parameter defaulting to True.""" + assert (ChimeraBoostQuantileRegressor().get_params()["refit_full"] + is True) + X, y = _heteroscedastic(n=1200, seed=61) + kw = dict(random_state=0, n_estimators=100, thread_count=1) + m = ChimeraBoostQuantileRegressor(**kw).fit(X, y) + r = ChimeraBoostQuantileRegressor(refit_full=True, **kw).fit(X, y) + assert m.refit_ is not None + assert r.refit_ == m.refit_ + assert m.audition_ == r.audition_ + assert np.array_equal(m.conformal_scale_, r.conformal_scale_) + assert np.array_equal(m.predict(X), r.predict(X)) + + +def test_refit_full_needs_auto_split_and_no_honest_fold(): + """With a user eval_set, conformalize=True, or early stopping off, the + refit stays out: the default is exactly refit_full=False.""" + X, y = _heteroscedastic(n=1200, seed=62) + Xt, yt, Xv, yv = X[:900], y[:900], X[900:], y[900:] + base = dict(random_state=0, n_estimators=100, thread_count=1) + d = ChimeraBoostQuantileRegressor(**base).fit(Xt, yt, eval_set=(Xv, yv)) + f = ChimeraBoostQuantileRegressor(refit_full=False, **base).fit( + Xt, yt, eval_set=(Xv, yv)) + assert d.refit_ is None + assert f.refit_ is None + assert d.audition_ == f.audition_ + assert np.array_equal(d.conformal_scale_, f.conformal_scale_) + assert np.array_equal(d.predict(Xt), f.predict(Xt)) + d = ChimeraBoostQuantileRegressor(conformalize=True, **base).fit(X, y) + f = ChimeraBoostQuantileRegressor( + conformalize=True, refit_full=False, **base).fit(X, y) + assert d.refit_ is None + assert f.refit_ is None + assert np.array_equal(d.predict(X), f.predict(X)) + d = ChimeraBoostQuantileRegressor(early_stopping=False, **base).fit(X, y) + f = ChimeraBoostQuantileRegressor( + early_stopping=False, refit_full=False, **base).fit(X, y) + assert d.refit_ is None + assert f.refit_ is None + assert np.array_equal(d.predict(X), f.predict(X)) + + +_REFIT_WANT = { + "head": (False, True), "bins": (False, True), + "recentred": (True, True), "fixed": (True, False), + "scaled": (True, False), +} + + +def test_refit_full_records_and_keeps_es_values(): + """Per winner: refit_ says what was retrained, at the replay round + count; audition_, factors, best round and curve keep the 80% values.""" + for kind in ("head", "bins", "recentred", "fixed", "scaled"): + mf, _, _, _, _ = _aud_nosplit_fit(kind, refit_full=False) + mt, _, _, _, _ = _aud_nosplit_fit(kind) + assert mt.audition_["selected"] == kind + assert mt.audition_ == mf.audition_ + want_c, want_h = _REFIT_WANT[kind] + assert mt.refit_["centre"] is want_c + assert mt.refit_["head"] is want_h + if want_h: + t_star = len(mf.model_.trees_) + want_rounds = min(int(np.ceil(t_star / 0.8)), 300) + assert mt.refit_["rounds"] == want_rounds + assert len(mt.model_.trees_) == want_rounds + else: + assert mt.refit_["rounds"] is None + assert np.array_equal(mt.conformal_scale_, mf.conformal_scale_) + assert mt.best_iteration_ == mf.best_iteration_ + assert mt.validation_history_ == mf.validation_history_ + + +def test_refit_full_centre_matches_full_row_regressor(): + """The delivered R/S/N centre is the same regressor fitted on all rows; + offsets, residual quantiles, floor and spread stay from the audition.""" + for kind in ("recentred", "fixed", "scaled"): + mt, Xtr, ytr, Xte, _ = _aud_nosplit_fit(kind) + mf, _, _, _, _ = _aud_nosplit_fit(kind, refit_full=False) + ref = ChimeraBoostRegressor( + n_estimators=300, early_stopping_rounds=50, thread_count=1, + random_state=0).fit(Xtr, ytr) + assert np.array_equal(mt._centre_model_.predict(Xte), + ref.predict(Xte)) + assert mt._centre_off_ == mf._centre_off_ + if kind == "fixed": + assert np.array_equal(mt._fixed_q_, mf._fixed_q_) + if kind == "scaled": + assert np.array_equal(mt._scaled_q_, mf._scaled_q_) + assert mt._spread_floor_ == mf._spread_floor_ + assert np.array_equal(mt._spread_model_.predict(Xte), + mf._spread_model_.predict(Xte)) + + +def test_refit_full_paths_serve_retrained_models(): + """Ordered grids, staged final equals predict, pickle round-trips.""" + for kind in ("head", "bins", "recentred", "fixed", "scaled"): + mt, _, _, Xte, _ = _aud_nosplit_fit(kind) + Q = mt.predict(Xte) + assert np.all(np.diff(Q, axis=1) >= 0) + Xt = Xte[:50] + stages = list(mt.staged_predict(Xt)) + assert np.array_equal(stages[-1], mt.predict(Xt)) + back = pickle.loads(pickle.dumps(mt)) + assert np.array_equal(back.predict(Xte), Q) + assert back.refit_ == mt.refit_ + + +def test_refit_centre_uses_head_validation_fraction(): + """With validation_fraction=0.3 the retrained centre still carves the + head's split: it equals the same regressor with validation_fraction=0.3 + fitted on all rows.""" + for kind in ("recentred", "fixed", "scaled"): + mt, Xtr, ytr, Xte, _ = _aud_nosplit_fit( + kind, validation_fraction=0.3) + assert mt.audition_["selected"] in ("recentred", "fixed", "scaled") + ref = ChimeraBoostRegressor( + n_estimators=300, early_stopping_rounds=50, thread_count=1, + random_state=0, validation_fraction=0.3).fit(Xtr, ytr) + assert np.array_equal(mt._centre_model_.predict(Xte), + ref.predict(Xte)) diff --git a/tests/test_quantile_shap.py b/tests/test_quantile_shap.py index 19a5c6a..b02409b 100644 --- a/tests/test_quantile_shap.py +++ b/tests/test_quantile_shap.py @@ -322,3 +322,33 @@ def test_audition_winner_shap_is_locally_accurate(): assert not np.any(phiw), "fixed width attribution is zero" assert np.allclose(phiw[rows].sum(axis=1) + m.expected_value_, iv[rows, 1] - iv[rows, 0]) + + +def _aud_nosplit_shap(kind): + """The kind's training rows without an eval_set, under the default; + (model, Xte).""" + from sklearn.model_selection import train_test_split + taus = [0.05, 0.25, 0.5, 0.75, 0.95] + X, y, sseed = _aud_data_shap(kind) + Xtr, Xte, ytr, _ = train_test_split(X, y, test_size=0.25, + random_state=sseed) + m = ChimeraBoostQuantileRegressor( + quantiles=taus, n_estimators=300, early_stopping_rounds=50, + thread_count=1, random_state=0).fit(Xtr, ytr) + return m, Xte + + +def test_refit_full_shap_is_locally_accurate(): + """SHAP reconstructs the RETRAINED winner per level, for H/B/R/S/N.""" + for kind in ("head", "bins", "recentred", "fixed", "scaled"): + m, Xte = _aud_nosplit_shap(kind) + assert m.audition_["selected"] == kind + Xt = Xte[:20] + rows = slice(None) + if kind == "scaled": + rows = _unfloored_rows(m, Xt) + assert rows.any(), "no unfloored rows; test is moot" + phi = m.shap_values(Xt) + assert phi.shape == (20, Xt.shape[1], 5) + assert np.allclose(phi[rows].sum(axis=1) + m.expected_value_, + m.predict(Xt)[rows]) diff --git a/tests/test_quantile_suite.py b/tests/test_quantile_suite.py index 27c0363..d2c035d 100644 --- a/tests/test_quantile_suite.py +++ b/tests/test_quantile_suite.py @@ -233,7 +233,7 @@ def test_q0_probes_are_opt_in(monkeypatch): import quantile_synth as qsyn assert len(qs.ARMS) == 7 - assert len(qs.PROBES) == 17 + assert len(qs.PROBES) == 18 _stub_registry(monkeypatch, ["gr:reg_num/houses"], {"gr:reg_num/houses": "regression"}) @@ -686,3 +686,40 @@ def test_q9_audition_sn_library_matches_bench_oracle_bit_for_bit(monkeypatch): # records the spread model's round and the library the centre's. if kind != "scaled": assert m.best_iteration_ == best_b + + +def test_split_with_full_leaves_existing_arms_unchanged(): + """The .full-carrying split unpacks as the plain 4-tuple: an existing + arm's grid and stopping round are identical either way.""" + split, Xte, taus = _q0_split() + Xf, Xv, yf, yv = split + wrapped = qs._SplitWithFull(split, (np.concatenate([Xf, Xv]), + np.concatenate([yf, yv]))) + assert len(wrapped) == 4 + for a, b in zip(wrapped, split): + assert a is b + Qp, _, _, best_p = qs._fit_chimera_head(split, Xte, None, 1, taus) + Qw, _, _, best_w = qs._fit_chimera_head(wrapped, Xte, None, 1, taus) + assert best_w == best_p + assert np.array_equal(Qw, Qp) + + +def test_all_rows_probe_records_choice_and_refit(monkeypatch): + """The AllRows arm fits the default on split.full with no eval_set and + records the audition choice (0-4) plus what was retrained.""" + monkeypatch.setattr(rb, "MAX_ITERS", 300) + split, Xte, taus = _q0_split() + Xf, Xv, yf, yv = split + wrapped = qs._SplitWithFull(split, (np.concatenate([Xf, Xv]), + np.concatenate([yf, yv]))) + Q, _, _, _, extra = qs.PROBES["ChimeraBoostQuantileAllRows"]( + wrapped, Xte, None, 1, taus) + assert Q.shape == (Xte.shape[0], len(taus)) + assert np.all(np.diff(Q, axis=1) >= 0) + assert extra["audition_choice"] in (0, 1, 2, 3, 4) + assert isinstance(extra["refit_centre"], bool) + assert isinstance(extra["refit_head"], bool) + if extra["refit_head"]: + assert isinstance(extra["refit_rounds"], int) + else: + assert extra["refit_rounds"] is None