Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,21 @@ All notable changes to ChimeraBoost are documented here.
The format follows [Keep a Changelog](https://keepachangelog.com/).

## [Unreleased]
### Added
- **`store_training_data` and `refresh(X, y)`** on the regressor and the
binary classifier (#131, first slice). With `store_training_data=True` a
model keeps its training rows in compact form (bins for plain numeric
columns, raw values only for the columns that feed cross features,
category codes), and `refresh` folds in new rows by replaying every
tree's splits on the old and new rows together and refitting only the
leaf values. It is the same replay the default full-data refit uses,
with the tree count, learning rate and splits pinned; refreshing with no
new rows reproduces the model bit for bit. Bagged models, multiclass,
`loss="Quantile"` and `random_effects=True` are not supported yet. The
default fit's results are unchanged. On a 3.3M-row regression, folding in
a day of new rows took 70 s against 232 s for a full refit, at the same
accuracy.

### Changed
- **The quantile head retrains its winner on all rows by default.** Called
without an `eval_set`, `ChimeraBoostQuantileRegressor` held back
Expand Down Expand Up @@ -49,6 +64,11 @@ The format follows [Keep a Changelog](https://keepachangelog.com/).
untouched: the exact-output snapshot moves only one quantile
configuration (183 of 186 pins identical). Record:
`benchmarks/QUANTILE_PLAN.md`, Phase 2 (Q8, Q9).
- **Fits with linear leaves are faster, with identical results.** The
per-leaf linear fit used to sum each leaf on one thread, and one leaf
often holds 40-60% of the rows. Large leaves now spread their sums over
several threads, each sum still adding the rows in the same order. A
500k-row fit went from 15.9 to 13.4 s, and a refresh from 5.0 to 3.6 s.
- **The quantile benchmark's NGBoost opponent now uses the RoNGBa
settings** (Ren, Sun and Wu 2019; #163): trees of up to 31 leaves, a
learning rate of 0.04 and at most 500 rounds, in place of NGBoost's stock
Expand Down
5 changes: 5 additions & 0 deletions benchmarks/CAMPAIGN_PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -1277,6 +1277,11 @@ PR #170 is PARKED OPEN (not merged, not closed); issue #131 stays open.
The loop moves on to the other issues; if he closes #170, the kernel
speedup (435e866, bit-identical, 1.19x on linear-leaf fits) is worth
salvaging as its own PR, since it helps every default fit.
2026-09-26, the maintainer: "let's move forward with 170, but clean up
the merge conflicts then i'll merge the PR". Un-parked: main merged into
the branch; the conflicts were CHANGELOG (entries on both sides, all
kept), this file and `REFRESH_PLAN.md` (log lines on both sides, all
kept); no code overlapped. Awaiting his merge.

#### I065 2026-09-24 issue #81 (research cascade: dead self-test anchor, stale `ideas.py` flags; BENCH tooling + test, pre-registered)
why now: the focus rule's second issue. PR #168 merged (3e02ab1), #84
Expand Down
5 changes: 5 additions & 0 deletions benchmarks/REFRESH_PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -131,3 +131,8 @@ Still open, for their slices:
implemented"). Slice 1 is PARKED; slices 2-7 are not started. If #170 is
closed, salvage the bit-identical replay-kernel speedup (435e866) as its
own PR: it speeds up every linear-leaf fit.
- 2026-09-26: un-parked. The maintainer: "let's move forward with 170,
but clean up the merge conflicts then i'll merge the PR". Main merged
into the branch (conflicts only in CHANGELOG and two plan files; no
code overlapped), so the salvage note above no longer applies. Slices
2-7 are not started.
25 changes: 25 additions & 0 deletions chimeraboost/booster.py
Original file line number Diff line number Diff line change
Expand Up @@ -824,6 +824,31 @@ def __init__(self, loss="RMSE", loss_kwargs=None, **kw):
self.loss_name = loss
self.loss_kwargs = loss_kwargs or {}

def replay_kwargs(self):
"""Constructor kwargs that rebuild an equivalent booster.

Every ``_BaseBooster.__init__`` parameter (each stored under its own
name, so read off ``inspect.signature`` rather than a hand-kept
list) plus ``loss`` (from ``loss_name``) and ``loss_kwargs``; lists
and dicts are copied.
"""
import inspect

kw = {}
params = inspect.signature(_BaseBooster.__init__).parameters
for name in params:
if name == "self":
continue
v = getattr(self, name)
if isinstance(v, list):
v = list(v)
elif isinstance(v, dict):
v = dict(v)
kw[name] = v
kw["loss"] = self.loss_name
kw["loss_kwargs"] = dict(self.loss_kwargs)
return kw

def _fit_impl(self, X, y, cat_features=None, eval_set=None,
sample_weight=None, callbacks=None, prep_cache=None):
"""Fit the additive model.
Expand Down
Loading
Loading