Skip to content

Quantile model: retrain the winner on all rows by default (refit_full) - #180

Merged
bbstats merged 5 commits into
mainfrom
campaign/quantile-q12-refit-probe
Sep 26, 2026
Merged

bbstats merged 5 commits into
mainfrom
campaign/quantile-q12-refit-probe

Conversation

@bbstats

@bbstats bbstats commented Sep 26, 2026

Copy link
Copy Markdown
Owner

The quantile model now learns from all of your training rows. Until now, fit(X, y) without an eval_set held back 20% of the rows to choose the stopping round, calibrate the intervals and pick its candidate, and those rows never reached the final model. ChimeraBoostRegressor has always retrained on all rows at the end; the quantile model now does too.

model = ChimeraBoostQuantileRegressor().fit(X, y)   # retrains the winner on all rows
model.refit_       # what was retrained, e.g. {"centre": True, "head": False, "rounds": None}
ChimeraBoostQuantileRegressor(refit_full=False)      # the previous behaviour

What changes for users

call before now
fit(X, y) final model trained on 80% of the rows choices made on the held-back 20% as before, then the winner retrained on all rows
fit(X, y, eval_set=...) unchanged unchanged: your held-out rows are never trained on
conformalize=True or early_stopping=False unchanged unchanged
name default meaning
refit_full True Retrain the winner on all rows once the choices are made on the held-out fold. False skips it.
refit_ What was retrained: {"centre": bool, "head": bool, "rounds": int or None}, or None.

Kept from the held-out fit: the chosen candidate, the calibration factors, the fixed and scaled candidates' offsets, best_iteration_ and validation_history_. Only the models that make the prediction are retrained. The squared-error centre of the recentred, fixed and scaled candidates goes through the regressor's own retrain. The quantile trees of the head, bins and recentred candidates are refitted from scratch at 1.25 times the rounds early stopping kept, which is the regressor's rule.

What it buys

On the 36 Grinsztajn regression datasets (3 seeds each), against today's default called the same way:

  • Better on 35, worse on 1 (analcatdata_supreme, by 0.5%), with a median CRPS gain of 0.9%. The chance of that by luck is about one in a billion.
  • The largest gains are on low-noise data: Brazilian_houses +8.6%, visualizing_soil +7.7%, pol +5.1%.
  • Better on all 6 high-cardinality datasets (median +1.4%).
  • Coverage stays on target: a nominal 90% interval covers 90.4% on average (90.1% before). On the three time-split datasets from Quantile model: fixed-width and scaled candidates in the default audition #177's caveat, the median coverage error falls from 4.7 to 3.3 points.
  • Cost: about 30% more fit time where it acts.

A cheaper version that retrains only the squared-error centre was measured in the same run: better on all 19 datasets it changed (median +1.7%) for 7% more time. The full retrain beat it head to head, 20 wins to 1, which was the rule set before the run.

Checks

  • Identity: without an eval_set, the new default reproduces the benchmark arm exactly on 8 of 8 fits covering the head, bins, recentred and scaled winners. With an eval_set, it reproduces today's numbers exactly.
  • Tests: 1235 passed, 1 skipped. New ones cover each winner's retrain: the centre equals the regressor fitted on all rows, the round rule holds, calibration is kept, and staged prediction, SHAP and pickling all work. Others cover a non-default validation_fraction and every case where nothing may be retrained.
  • Exact-output snapshot: 182 of 186 pins identical. The 4 that move are the predictions and importances of two quantile configurations fitted without an eval_set; their calibration factors don't move. I'll rebaseline after the merge.
  • Docs: docs/quantiles.md (a new "Retraining on all rows" section and a tuning note), docs/parameters.md (the refit_full row), CHANGELOG.
  • Touches chimeraboost/, tests/ and docs/, so it's yours to merge.

For reviewers

  • The retrained centre is ChimeraBoostRegressor(<the same arguments>, validation_fraction=<the head's>).fit(all rows). It carves exactly the rows the head carved (checked bit for bit), then retrains with its own default refit. That repeats the centre's 80% fit once; calling the regressor's refit directly would save a few percent of the fit.
  • The benchmark's comparison arm keeps passing the shared split as an eval_set, so the docs' comparison table still has every model on the same rows and does not include this gain. The docs say so.
  • Also on this branch: the benchmark passes each arm the original training rows (split.full), and the probe arm ChimeraBoostQuantileAllRows measures the default the way a user without an eval_set gets it.

🤖 Generated with Claude Code

bbstats and others added 5 commits September 25, 2026 19:20
…s (the maintainer's pick)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, I075)

ChimeraBoostQuantileRegressor gains refit_full=False. When True and the
fit used the automatic early-stopping split (no user eval_set, early
stopping on, conformalize "auto" or False), the winner is retrained on
all rows after the audition finishes: an R/S/N winner's centre becomes
the same ChimeraBoostRegressor fitted on all rows without an eval_set
(its own carve, identical to the audition's, and its own replay refit);
the head booster of an H/B/R winner is retrained from scratch at
min(ceil(t_star / 0.8), n_estimators) rounds with the rate pinned.
Calibration factors, offsets, residual quantiles and the spread model
stay from the held-out fit. refit_ records what was retrained. Off by
default: the identity snapshot is 186/186.

quantile_suite.py passes each arm the original training rows
(split.full) and adds the probe arms ChimeraBoostQuantileRefitCentre and
ChimeraBoostQuantileRefitAll; quantile_synth.py passes the rows too.
11 new tests; 1235 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sters the default

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
refit_full now defaults to True: called without an eval_set, the head
makes every choice on its carved fold as before (stopping round,
calibration, candidate), then retrains the winner on all rows, as
ChimeraBoostRegressor does. The retrained centre gets the head's
validation_fraction so its carve always equals the head's. The probe-only
_refit_scope switch and its centre-only branch are removed. Nothing
changes with a user eval_set, early stopping off or conformalize=True.

Benchmark: ChimeraBoostQuantileRefitAll becomes ChimeraBoostQuantileAllRows
(the default on split.full, no eval_set); ChimeraBoostQuantileRefitCentre
is deleted; the field arm keeps its eval_set. Tests: 7 changed, 1 added,
1 deleted; 1235 passed. Identity snapshot 182/186 (mq3, mq3_w_sub
predictions and importances; calibration factors unchanged).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bbstats
bbstats merged commit 180f765 into main Sep 26, 2026
7 checks passed
@bbstats
bbstats deleted the campaign/quantile-q12-refit-probe branch September 26, 2026 01:54
bbstats added a commit that referenced this pull request Sep 26, 2026
Plan: #180 merged; the quantile queue holds (I076)
This was referenced Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant