Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,24 @@ The format follows [Keep a Changelog](https://keepachangelog.com/).

## [Unreleased]
### Changed
- **The quantile head retrains its winner on all rows by default.** Called
without an `eval_set`, `ChimeraBoostQuantileRegressor` held back
`validation_fraction` of the rows for early stopping, calibration and
the candidate choice, and never trained on them. The new `refit_full`,
on by default, makes every choice on those rows as before and then
retrains the winner on all rows, as `ChimeraBoostRegressor` does. The
candidate, the calibration factors, the fixed and scaled candidates'
offsets, `best_iteration_` and `validation_history_` stay from the
held-out fit, and `refit_` records what was retrained. On the 36
Grinsztajn regression datasets it improves CRPS on 35 and loses 1 by
0.5% (median +0.9%; up to +8.6% on Brazilian_houses), and it improves all
6 high-cardinality ones (median +1.4%). It also shrinks the time-split
caveat in the next entry: the median 90% coverage error there falls
from 4.7 to 3.3 points. It adds about 30% to the fit time where it acts.
Nothing changes with your own `eval_set`, with early stopping off, or
with `conformalize=True`; `refit_full=False` skips it. The comparisons
in `docs/quantiles.md` keep every model on the same rows, so they do not
include this gain. Record: `benchmarks/QUANTILE_PLAN.md`, Q12 and Q13.
- **The quantile head's audition gains two candidates, and its CRPS improves
most on low-noise targets.** `ChimeraBoostQuantileRegressor`'s default
choice now also tries `"fixed"`, an ordinary squared-error
Expand Down
127 changes: 126 additions & 1 deletion benchmarks/CAMPAIGN_PLAN.md

Large diffs are not rendered by default.

25 changes: 22 additions & 3 deletions benchmarks/QUANTILE_PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -424,6 +424,23 @@ Q0 reads moved it down to item 4.
1 of the 59 decide keys, so it cannot clear the gate. Bagging the centre
×5 is the ceiling (visualizing_soil +30%, pol +9%) and is an ensemble,
so not a default.
3g. **Q12, retraining the winner on all rows (the maintainer's pick
2026-09-25).** Without an `eval_set` the head's own carve equals the
suite's shared split bit for bit, so the probe is a retrain added to
today's fit: every choice and calibration quantity from the held-out
fit, then (a) the R/S/N centre retrained on all rows, or (b) that plus
the H/B/R head retrained from scratch at the replay-round rule.
**PASSED 2026-09-25 (I075, `results/quantile-20260925-202430.json`).**
Against the default: (a) gr 19W-0L-17T, +1.72%, p 3.8e-6, 1.07× fit;
(b) gr 35W-1L, +0.91%, p 1.1e-9, hc 6W-0L, 1.28× fit; (b) beats (a)
20W-1L-15T (p 2.1e-5), so (b). Coverage guard fine (gr 90% error 0.32 →
0.47 points); the hc `@time` flag shrinks (4.69 → 3.29).
3h. **Q13, `refit_full=True` as the default.**
**PASSED 2026-09-25 (I076), PR for the maintainer.** Without an
`eval_set` the default reproduces Q12's arm (b) bit for bit; with one,
nothing changes. The suite's field arm keeps its `eval_set`, so every
comparison stays "every model on the same rows" and leaves this gain
out.
4. **Q1, the narrow-interval defect (P16).** Leaf values are in-sample
residual quantiles, so intervals over-narrow (0.869 at nominal 0.90 on
2026-08-30; coverage decays with rounds). Fit leaf quantiles
Expand Down Expand Up @@ -533,16 +550,18 @@ offset on real data (a CRPS tie), which made it Q0's subject.
against pinball. That is a default flip on a strength surface, so it needs
its own pre-registration and the full `/experiment` protocol. Not attempted
here. Recorded 2026-08-30.
- **The head never trains on its own early-stopping fold** (recorded
- RESOLVED 2026-09-25 (Q12 and Q13, I075 and I076): **The head never trains on its own early-stopping fold** (recorded
2026-09-25, I073b). Without an `eval_set` the head carves
`validation_fraction` and neither it nor its S/N centre ever sees those
rows, while `ChimeraBoostRegressor` refits on all rows by default
(`refit_full`), worth 8.8% RMSE on visualizing_soil and 6.2% on pol. A
refit of the winner after the audition, keeping the calibration taken
before it, is the candidate. `quantile_suite.py` cannot measure it: it
passes the shared split as an `eval_set`, which the head must not train
on. Needs a no-`eval_set` protocol first; the maintainer's call whether
to open it.
on. Needs a no-`eval_set` protocol first. **OPENED 2026-09-25 as Q12**
(the maintainer picked it over Q1): the head's own carve equals the
suite's shared split bit for bit, so the probe is a retrain added to
today's arm (`CAMPAIGN_PLAN.md` I075).
- RESOLVED 2026-09-23 (Q5, I059): `docs/quantiles.md` "How it compares" is
re-measured against the new default, with the fixed-width baseline and
NGBoost added.
Expand Down
57 changes: 56 additions & 1 deletion benchmarks/quantile_suite.py
Original file line number Diff line number Diff line change
Expand Up @@ -96,6 +96,24 @@
RONGBA_LEAVES = 31


class _SplitWithFull(tuple):
"""The shared ``(Xf, Xv, yf, yv)`` split, carrying the training rows.

Unpacks and indexes as the plain 4-tuple every arm already takes;
``.full`` is the ``(Xtr, ytr)`` the shared split was carved from, in
their original order, so a probe arm can fit on all training rows
while the library's own carve reproduces the shared split.
"""

def __new__(cls, split, full):
obj = super().__new__(cls, split)
obj.full = full
return obj

def __init__(self, split, full):
pass


def _fit_head_model(split, cat, threads, taus, n_estimators, **params):
"""Construct and fit the suite's head.

Expand Down Expand Up @@ -811,6 +829,41 @@ def _fit_chimera_default_uncapped(split, Xte, cat, threads, taus):
return Q, fit_s, time.time() - t, m.best_iteration_


_AUDITION_CHOICE = {"head": 0, "bins": 1, "recentred": 2,
"fixed": 3, "scaled": 4}


def _refit_extras(m):
"""The refit probes' extras: the audition choice as the audition arms
encode it (0-4 for H, B, R, S, N), whether the centre and the head
were retrained, and the head's retrained round count."""
aud = m.audition_ or {}
refit = m.refit_ or {}
return {"audition_choice": _AUDITION_CHOICE.get(aud.get("selected")),
"refit_centre": bool(refit.get("centre", False)),
"refit_head": bool(refit.get("head", False)),
"refit_rounds": refit.get("rounds")}


def _fit_chimera_all_rows(split, Xte, cat, threads, taus):
"""The library default as a user without an ``eval_set`` gets it:
fitted on ALL training rows, so the audition runs on the library's
own carve (which reproduces the shared split) and the winner is
retrained on every row."""
Xtr, ytr = split.full
m = ChimeraBoostQuantileRegressor(
quantiles=taus, n_estimators=rb.MAX_ITERS,
early_stopping_rounds=rb.PATIENCE, thread_count=threads,
random_state=0)
t = time.time()
m.fit(Xtr, ytr, cat_features=cat or None)
fit_s = time.time() - t
t = time.time()
Q = m.predict(Xte)
pred_s = time.time() - t
return Q, fit_s, pred_s, m.best_iteration_, _refit_extras(m)


PROBES = {
"ChimeraBoostQuantileUncapped": _fit_chimera_uncapped,
"ChimeraBoostQuantileDepth6": _fit_chimera_depth6,
Expand All @@ -829,6 +882,7 @@ def _fit_chimera_default_uncapped(split, Xte, cat, threads, taus):
"ChimeraBoostQuantileAuditionS": _fit_chimera_audition_s,
"ChimeraBoostQuantileAuditionSN": _fit_chimera_audition_sn,
"ChimeraBoostQuantileDefaultUncapped": _fit_chimera_default_uncapped,
"ChimeraBoostQuantileAllRows": _fit_chimera_all_rows,
}


Expand Down Expand Up @@ -944,7 +998,8 @@ def run_one(ds_name, seed, taus, threads, models):
"regression")
# One early-stopping split, shared by every arm, so no model is judged on
# more data than another. Same carve the harness uses.
split = rb._val_split(Xtr, ytr, "regression", 0)
split = _SplitWithFull(rb._val_split(Xtr, ytr, "regression", 0),
(Xtr, ytr))

meta = {"task": "quantile", "n_train": int(Xtr.shape[0]),
"n_total": int(X.shape[0]), "n_features": int(X.shape[1]),
Expand Down
3 changes: 2 additions & 1 deletion benchmarks/quantile_synth.py
Original file line number Diff line number Diff line change
Expand Up @@ -219,7 +219,8 @@ def run_one(regime, n, seed, taus, threads, models):
random_state=seed)
Xtr, Xte, ytr, yte = X[idx_tr], X[idx_te], y[idx_tr], y[idx_te]
Qo_te = Q_or[idx_te]
split = rb._val_split(Xtr, ytr, "regression", 0)
split = qs._SplitWithFull(rb._val_split(Xtr, ytr, "regression", 0),
(Xtr, ytr))
cat = [X.shape[1] - 1] if regime == "catscale" else None
oracle_crps = float(qm.crps(yte, Qo_te, taus))

Expand Down
139 changes: 132 additions & 7 deletions chimeraboost/quantile_api.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,8 @@

Kept out of ``sklearn_api`` because it shares that module's input validation
but almost none of its fit machinery: no loss family, no linear leaves, no
cross features, no bagging, no full-data refit. It imports the validation
helpers and keeps the same flat, module-function style.
cross features, no bagging, and a full-data refit that is on by default. It
imports the validation helpers and keeps the same flat, module-function style.
"""

import warnings
Expand Down Expand Up @@ -265,6 +265,15 @@ def _pick_audition_winner(ordered):
return best_name


def _keep_es_state(new_booster, old_booster):
"""Keep the early-stopped fit's visible state on a refit replacement,
as ``sklearn_api._refit_on_full`` does: the training and validation
curves, and the budget early stopping chose."""
new_booster.train_history_ = old_booster.train_history_
new_booster.valid_history_ = old_booster.valid_history_
new_booster.best_iteration_ = old_booster.best_iteration_


def _phi_centre(phi, mi, mw):
"""Centre of attributions along the level axis, mirroring ``_centre``."""
if mw == 0.0:
Expand Down Expand Up @@ -338,13 +347,34 @@ class ChimeraBoostQuantileRegressor(BaseEstimator):
calibrates the head and scored unweighted; ties go H, then B, then
R, then S, then N. Needs ``conformalize="auto"`` and evaluation
rows (the user's ``eval_set`` or the carved fold); otherwise the
single head is fitted. Costs about 2.6x a single head at the median:
S reuses R's centre fit, so the extra cost over the three-candidate
audition is the spread-model fit, about 10% of the fit at the
median. ``False`` fits a single head, as releases before 0.33.0 did.
single head is fitted. Costs about 2.6x a single head at the median
for fits with an ``eval_set``: S reuses R's centre fit, so the extra
cost over the three-candidate audition is the spread-model fit,
about 10% of the fit at the median. The full-data refit adds about
30% at the median when it acts. ``False`` fits a single head, as
releases before 0.33.0 did.
``model_`` stays the fitted head booster whatever wins (the 254-bin
booster for a bins win); for S and N it delivers nothing --
``predict`` serves the centre/spread models -- but stays inspectable.
refit_full : bool, default True
Retrain the winner on all rows after early stopping chose the
budget, as ``ChimeraBoostRegressor`` does. Acts only on the
automatic early-stopping split: never with a user ``eval_set``,
with early stopping off, with ``conformalize=True``, or when the
carve failed -- those fits are exactly what they are without it.
An R, S or N winner's centre is then replaced by a default
``ChimeraBoostRegressor`` fitted on all rows without an
``eval_set`` -- offsets, spread model, residual quantiles, floor
and calibration factors stay from the audition fit -- and the
head booster of an H, B or R winner is retrained from scratch on
all rows with early stopping off, at ``min(ceil(t_star / (1 -
validation_fraction)), n_estimators)`` rounds with the
early-stopped learning rate pinned, where t_star is the rounds
the early-stopped head kept. ``audition_``,
``conformal_scale_``, ``best_iteration_`` and
``validation_history_`` keep the early-stopped fit's values;
``refit_`` records what was retrained. ``False`` skips the
retrain.

Attributes
----------
Expand All @@ -364,6 +394,12 @@ class ChimeraBoostQuantileRegressor(BaseEstimator):
"crps": {name: score}}`` for the candidates that ran, or ``None``
when no audition ran (no evaluation rows, ``conformalize`` True or
False, or ``audition=False``).
refit_ : dict or None
``{"centre": bool, "head": bool, "rounds": int or None}`` --
whether the centre and the head booster were retrained on all
rows, and the head's retrained round count (None when the head
was not retrained). None when nothing was retrained
(``refit_full=False``, or a fit the refit does not cover).

Notes
-----
Expand All @@ -383,7 +419,8 @@ def __init__(self, quantiles=None, n_estimators=2000, learning_rate=None,
quantize_gradients=True, early_stopping=True,
validation_fraction=0.2, split_projection="rotate",
exact_splits=False, conformalize="auto",
calibration_fraction=0.2, audition=True):
calibration_fraction=0.2, audition=True,
refit_full=True):
self.quantiles = quantiles
self.n_estimators = n_estimators
self.learning_rate = learning_rate
Expand All @@ -409,6 +446,7 @@ def __init__(self, quantiles=None, n_estimators=2000, learning_rate=None,
self.conformalize = conformalize
self.calibration_fraction = calibration_fraction
self.audition = audition
self.refit_full = refit_full

def __sklearn_is_fitted__(self):
return hasattr(self, "model_")
Expand Down Expand Up @@ -467,6 +505,82 @@ def _carve_es_split(self, X, y, sample_weight, eval_set, groups):
es_rounds = 50
return X, y, sample_weight, eval_set, es_rounds

def _refit_full_active(self, auto_split):
"""True when the full-data refit fires: asked for, on the automatic
split, with no honest holdout at stake (``conformalize=True`` keeps
its pristine fold instead)."""
return (bool(self.refit_full) and bool(auto_split)
and self.conformalize is not True)

def _maybe_refit_full(self, auto_split, taus, es_rounds, cat_features,
X_full, y_full, sw_full, groups_full):
"""Retrain the winner on all rows, after the audition (or the single
head's calibration) finished exactly as without the refit. The
centre of an R, S or N winner is replaced by a default
``ChimeraBoostRegressor`` fitted on the full rows without an
``eval_set``; the head booster of an H, B or R winner -- or of the
single head when no audition ran -- is retrained from scratch with
early stopping off. Records what happened in ``refit_``."""
if not self._refit_full_active(auto_split):
return
selected = self._audition_selected()
centre_done = False
if selected in ("recentred", "fixed", "scaled"):
self._refit_audition_centre(es_rounds, cat_features, X_full,
y_full, sw_full, groups_full)
centre_done = True
head_done, rounds = False, None
if selected in (None, "head", "bins", "recentred"):
rounds = self._refit_head_booster(taus, selected, cat_features,
X_full, y_full, sw_full)
head_done = True
if not centre_done and not head_done:
return
self.refit_ = {"centre": centre_done, "head": head_done,
"rounds": rounds}

def _refit_audition_centre(self, es_rounds, cat_features, X_full,
y_full, sw_full, groups_full):
"""Replace the R/S/N centre with the same regressor fitted on the
full rows without an ``eval_set``, so it carves the same split and
retrains with its own default refit. Offset, spread model,
residual quantiles, floor and factors stay from the audition fit;
the early-stopped curves stay on the replacement, as the
regressor's own refit keeps them."""
from .sklearn_api import ChimeraBoostRegressor
old = self._centre_model_
centre = ChimeraBoostRegressor(
n_estimators=self.n_estimators,
early_stopping_rounds=es_rounds,
thread_count=self.thread_count, random_state=self.random_state,
validation_fraction=self.validation_fraction)
centre.fit(X_full, y_full, cat_features=cat_features,
sample_weight=sw_full, groups=groups_full)
_keep_es_state(centre.model_, old.model_)
self._centre_model_ = centre

def _refit_head_booster(self, taus, selected, cat_features, X_full,
y_full, sw_full):
"""Retrain the winner's head booster from scratch on the full rows:
the same construction (254 bins for a bins win), early stopping
off, ``min(ceil(t_star / (1 - validation_fraction)),
n_estimators)`` rounds with the early-stopped learning rate pinned.
Returns the retrained round count."""
winner = self.model_
t_star = len(winner.trees_)
frac = max(1.0 - float(self.validation_fraction), 1e-9)
rounds = min(int(np.ceil(t_star / frac)), int(self.n_estimators))
booster = self._make_mq_booster(
taus, None, cat_features, len(y_full),
max_bins=254 if selected == "bins" else None)
booster.n_estimators = rounds
booster.learning_rate = float(winner.lr_)
booster.fit(X_full, y_full, cat_features=cat_features,
sample_weight=sw_full)
_keep_es_state(booster, winner)
self.model_ = booster
return rounds

def _make_mq_booster(self, taus, es_rounds, cat_features, n_rows,
max_bins=None):
"""Resolve the auto defaults and construct the booster."""
Expand Down Expand Up @@ -530,6 +644,7 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None,
self._median_idx_ = _median_index(taus)
self.conformal_scale_ = np.ones(taus.shape[0])
self.audition_ = None
self.refit_ = None
self._centre_model_ = None
self._centre_off_ = None
self._fixed_q_ = None
Expand All @@ -544,8 +659,15 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None,

cal, X, y, sample_weight, groups = self._carve_calibration_fold(
X, y, sample_weight, groups)
# Kept for the optional full-data refit below: the auto split
# reassigns X/y, but the refit retrains on every row.
X_full, y_full, sw_full, groups_full = X, y, sample_weight, groups
had_user_eval = eval_set is not None
X, y, sample_weight, eval_set, es_rounds = self._carve_es_split(
X, y, sample_weight, eval_set, groups)
# The carve is the only path that sets eval_set when the user did
# not, so this is True exactly when the automatic split was used.
auto_split = not had_user_eval and eval_set is not None

self.model_ = self._make_mq_booster(taus, es_rounds, cat_features,
len(y))
Expand Down Expand Up @@ -573,6 +695,9 @@ def fit(self, X, y, cat_features=None, eval_set=None, groups=None,
except ValueError:
pass # uncertifiable: keep the raw grid, never raise

self._maybe_refit_full(auto_split, taus, es_rounds, cat_features,
X_full, y_full, sw_full, groups_full)

return self

def _wants_audition(self, eval_set):
Expand Down
Loading
Loading