Skip to content

refresh(X, y): fold new rows into a fitted model (#131, slice 1) - #170

Merged
bbstats merged 5 commits into
mainfrom
campaign/issue131-refresh-slice1
Sep 26, 2026
Merged

bbstats merged 5 commits into
mainfrom
campaign/issue131-refresh-slice1

Conversation

@bbstats

@bbstats bbstats commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

refresh(): update a fitted model with new data

You trained a model, and new labelled rows keep arriving. Until now there were two options: throw the new rows away, or retrain from scratch. refresh() is a third. It keeps every tree the model already learned and recomputes the leaf values on the old and new rows together. For a daily update that's about 3x faster than a full refit, at the same accuracy.

from chimeraboost import ChimeraBoostRegressor

model = ChimeraBoostRegressor(store_training_data=True).fit(X, y)

# later, when new rows arrive
model.refresh(X_new, y_new)

Works on ChimeraBoostRegressor and on binary ChimeraBoostClassifier.

API

Constructor

Parameter Default Description
store_training_data False Keep a compact copy of the training rows so the model can be refreshed later.

Method: refresh(X, y, sample_weight=None) -> self

Argument Description
X New rows, with the same columns as at fit.
y Their targets. For the classifier, only classes seen at fit.
sample_weight Weights for the new rows. Required if the model was fit with weights, not allowed otherwise.

Attribute

Attribute Description
n_samples_trained_ Rows the leaf values are fit on. Set by fit, grown by refresh.

What a refresh updates

Updated Kept from fit
Leaf values and linear-leaf coefficients Every split of every tree
Categorical target statistics Number of trees and learning rate
Bin edges and selected features
Classifier calibration (temperature_)

The splits never change, so a refresh can't pick up a pattern that needs a new split. If your data shifts that way, refit.

Example: a daily update

Zurich public transport delays (regression, 8 features). A model is trained on 3.3M rows, then a day's worth of new rows (50k) arrives.

Time Test RMSE
Original fit on 3.3M rows 216 s 3.04147
refresh() with the day's 50k rows 70 s 3.04146
Full refit on all 3.35M rows 232 s 3.04145

A refresh costs about the same whatever the size of the new batch, because every tree is replayed over all the stored rows. That's what keeps it exact.

Guarantees

  • refresh with zero new rows returns the identical model, bit for bit. That's tested across 15 configurations: every refit_full mode, categorical and cross features, linear leaves, weights, missing values, custom objectives, subsampling, and the binary classifier.
  • Refreshing with batch A and then batch B gives the same model as one refresh with A and B together.
  • The default fit is unchanged. Every existing model output is bit-identical.

Storage

The stored copy is compact: numeric columns as 1-byte bin indices, categoricals as integer codes, and raw values only for the few columns that feed cross features. That's about 3x smaller than X for numeric data, and far smaller for text categories. It travels with the pickled model and grows with each refresh.

Not supported yet

With store_training_data=True, these raise NotImplementedError at fit:

  • bagged models (n_ensembles > 1, quality=4 or 5)
  • multiclass classification
  • loss="Quantile"
  • random_effects=True

These are the next slices of #131; the plan is in benchmarks/REFRESH_PLAN.md.

For reviewers

  • One code path. A refresh reuses the existing replay refit (the one behind refit_full="replay") unchanged. The stored rows are rebuilt into an equivalent X, which is exact because every stored bin maps back to a value in the same bin.
  • A faster replay kernel. Most of a refresh was the per-leaf linear fit. It summed each leaf on a single thread, and one leaf often holds 40-60% of the rows. Large leaves now spread their sums over several threads, and each sum still adds rows in the same order, so results are bit-identical (identity snapshot unchanged). Ordinary fits with linear leaves use the same kernel: a 500k-row fit went from 15.9 to 13.4 s.
  • Two departures from the issue's spec.
    • The store is private (_training_data_) and exposed through n_samples_trained_. In scikit-learn, X_train_ means the raw X.
    • The columns that feed cross features keep their raw values. That's why storage shrinks about 3x; the issue estimated 10 to 20x.
  • Tests: 36 in tests/test_refresh.py and 21 in tests/test_replay_kernel.py. The full suite passes: 1280 tests, 1 skipped.

Part of #131.

🤖 Generated with Claude Code

bbstats and others added 4 commits September 24, 2026 20:35
TrainingRows (chimeraboost/training_rows.py) keeps the rows a fitted
booster's leaves came from: plain numeric columns as bins of the fitted
binner, cross-feature parents as raw floats, categoricals as codes plus
their category values, y and raw weights. rebuild_X turns the store back
into a raw-equivalent X that the existing replay refit consumes
unchanged, so there is no second preprocessing path: every bin maps back
to a value that bins identically, and only cross parents are read raw.

GradientBoosting.replay_kwargs() returns the constructor kwargs of an
equivalent booster, read off the base signature.

Tests prove the round trip exact against the raw rows: the binned
matrix, the fitted preprocessor state, and the replayed booster's leaf
values and predictions, with and without appended rows. The public
store_training_data / refresh() API follows in pass 2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ChimeraBoostRegressor and the binary ChimeraBoostClassifier gain
store_training_data=False. When True, the fit keeps the rows the final
booster's leaves came from (TrainingRows, from pass 1), and
refresh(X, y, sample_weight=None) appends new rows, rebuilds a
raw-equivalent X from the store, and replays the fitted booster's
configuration on it: the same structure-transfer replay the default
full-data refit uses, with every tree's splits, the tree count and the
learning rate pinned, and no size-adaptive setting re-resolved.
n_samples_trained_ reports the rows the leaves come from.

Bagging, multiclass, loss="Quantile" and random_effects=True raise a
clear NotImplementedError at fit when the flag is on. refresh refuses
a model without a store, unseen class labels, and a weights mismatch.

Tests: a refresh with no new rows is bit-identical to the fitted model
in 15 configurations; a refresh equals a manual replay on the stacked
raw rows; two refreshes equal one combined refresh; pinned state never
moves; pickling and every error are covered. The default fit is
unchanged (identity snapshot 186/186).

Docs: a parameters row, a "Refreshing with new rows" recipe, and a
CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The per-leaf ridge sums in _linear_leaf_fit were bound by the largest
leaf, which often holds 40-60% of the rows and ran on one thread, after
a serial counting sort. Above _SMALL_N rows the kernel now sorts in
parallel, gathers each leaf's rows into contiguous buffers once, and
spreads the work over (leaf, accumulator) tasks, so a big leaf's sums
run on 1 + k threads. Every sum still adds its leaf's rows in increasing
row order with the same expressions, so results are bit-identical: the
identity snapshot is unchanged (186/186), and new tests compare the
kernel against a verbatim copy of the old one. replay_oblivious_tree
also uses the parallel leaf assignment.

On a 517k-row store: the linear-leaf fit drops from 9.9 to 7.8 ms per
tree, a daily refresh from 5.0 to 3.6 s, and a 500k-row fit with linear
leaves from 15.9 to 13.4 s, since the grow path shares the kernel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…131)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bbstats

bbstats commented Sep 25, 2026

Copy link
Copy Markdown
Owner Author

Pushed a faster replay kernel (435e866). In the daily-update case, most of a refresh was spent in the per-leaf linear fit, which summed each leaf on a single thread. That's slow when one leaf holds 40-60% of the rows. Large leaves now spread their sums across threads, and each sum still adds rows in the same order, so results are bit-identical: the identity snapshot is unchanged, and new tests check the kernel against a copy of the old one.

Before After
Daily refresh: 517k stored rows, 17k new 5.0 s 3.6 s
Ordinary 500k-row fit with linear leaves 15.9 s 13.4 s

I also replaced the example in the description with a daily update at scale. Refreshing a 3.3M-row model with a day of data takes 70 s, against 232 s for a full refit, at the same accuracy.

bbstats added a commit that referenced this pull request Sep 25, 2026
Plan: record the refresh rung (PR #170 parked) and queue #113
@bbstats

bbstats commented Sep 25, 2026

Copy link
Copy Markdown
Owner Author

Going to do more testing.

#170)

Conflicts were in CHANGELOG.md (Unreleased entries added on both sides;
all kept, main's two quantile entries first since one refers to the
next), benchmarks/CAMPAIGN_PLAN.md and benchmarks/REFRESH_PLAN.md (log
lines added on both sides; all kept, plus a 2026-09-26 line recording
the maintainer's go-ahead). No code overlapped. After the merge: 1292
passed, 1 skipped; ruff clean; identity snapshot 186/186 against main's
baseline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bbstats
bbstats merged commit fa577cc into main Sep 26, 2026
7 checks passed
@bbstats
bbstats deleted the campaign/issue131-refresh-slice1 branch September 26, 2026 14:47
bbstats added a commit that referenced this pull request Sep 26, 2026
This was referenced Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant