refresh(X, y): fold new rows into a fitted model (#131, slice 1) - #170
Merged
Merged
Conversation
TrainingRows (chimeraboost/training_rows.py) keeps the rows a fitted booster's leaves came from: plain numeric columns as bins of the fitted binner, cross-feature parents as raw floats, categoricals as codes plus their category values, y and raw weights. rebuild_X turns the store back into a raw-equivalent X that the existing replay refit consumes unchanged, so there is no second preprocessing path: every bin maps back to a value that bins identically, and only cross parents are read raw. GradientBoosting.replay_kwargs() returns the constructor kwargs of an equivalent booster, read off the base signature. Tests prove the round trip exact against the raw rows: the binned matrix, the fitted preprocessor state, and the replayed booster's leaf values and predictions, with and without appended rows. The public store_training_data / refresh() API follows in pass 2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ChimeraBoostRegressor and the binary ChimeraBoostClassifier gain store_training_data=False. When True, the fit keeps the rows the final booster's leaves came from (TrainingRows, from pass 1), and refresh(X, y, sample_weight=None) appends new rows, rebuilds a raw-equivalent X from the store, and replays the fitted booster's configuration on it: the same structure-transfer replay the default full-data refit uses, with every tree's splits, the tree count and the learning rate pinned, and no size-adaptive setting re-resolved. n_samples_trained_ reports the rows the leaves come from. Bagging, multiclass, loss="Quantile" and random_effects=True raise a clear NotImplementedError at fit when the flag is on. refresh refuses a model without a store, unseen class labels, and a weights mismatch. Tests: a refresh with no new rows is bit-identical to the fitted model in 15 configurations; a refresh equals a manual replay on the stacked raw rows; two refreshes equal one combined refresh; pinned state never moves; pickling and every error are covered. The default fit is unchanged (identity snapshot 186/186). Docs: a parameters row, a "Refreshing with new rows" recipe, and a CHANGELOG entry. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The per-leaf ridge sums in _linear_leaf_fit were bound by the largest leaf, which often holds 40-60% of the rows and ran on one thread, after a serial counting sort. Above _SMALL_N rows the kernel now sorts in parallel, gathers each leaf's rows into contiguous buffers once, and spreads the work over (leaf, accumulator) tasks, so a big leaf's sums run on 1 + k threads. Every sum still adds its leaf's rows in increasing row order with the same expressions, so results are bit-identical: the identity snapshot is unchanged (186/186), and new tests compare the kernel against a verbatim copy of the old one. replay_oblivious_tree also uses the parallel leaf assignment. On a 517k-row store: the linear-leaf fit drops from 9.9 to 7.8 ms per tree, a daily refresh from 5.0 to 3.6 s, and a 500k-row fit with linear leaves from 15.9 to 13.4 s, since the grow path shares the kernel. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…131) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Owner
Author
|
Pushed a faster replay kernel (435e866). In the daily-update case, most of a refresh was spent in the per-leaf linear fit, which summed each leaf on a single thread. That's slow when one leaf holds 40-60% of the rows. Large leaves now spread their sums across threads, and each sum still adds rows in the same order, so results are bit-identical: the identity snapshot is unchanged, and new tests check the kernel against a copy of the old one.
I also replaced the example in the description with a daily update at scale. Refreshing a 3.3M-row model with a day of data takes 70 s, against 232 s for a full refit, at the same accuracy. |
Owner
Author
|
Going to do more testing. |
#170) Conflicts were in CHANGELOG.md (Unreleased entries added on both sides; all kept, main's two quantile entries first since one refers to the next), benchmarks/CAMPAIGN_PLAN.md and benchmarks/REFRESH_PLAN.md (log lines added on both sides; all kept, plus a 2026-09-26 line recording the maintainer's go-ahead). No code overlapped. After the merge: 1292 passed, 1 skipped; ruff clean; identity snapshot 186/186 against main's baseline. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bbstats
added a commit
that referenced
this pull request
Sep 26, 2026
Plan: #170 merged, refresh slice 1 shipped
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
refresh(): update a fitted model with new dataYou trained a model, and new labelled rows keep arriving. Until now there were two options: throw the new rows away, or retrain from scratch.
refresh()is a third. It keeps every tree the model already learned and recomputes the leaf values on the old and new rows together. For a daily update that's about 3x faster than a full refit, at the same accuracy.Works on
ChimeraBoostRegressorand on binaryChimeraBoostClassifier.API
Constructor
store_training_dataFalseMethod:
refresh(X, y, sample_weight=None) -> selfXfit.yfit.sample_weightAttribute
n_samples_trained_fit, grown byrefresh.What a refresh updates
fittemperature_)The splits never change, so a refresh can't pick up a pattern that needs a new split. If your data shifts that way, refit.
Example: a daily update
Zurich public transport delays (regression, 8 features). A model is trained on 3.3M rows, then a day's worth of new rows (50k) arrives.
refresh()with the day's 50k rowsA refresh costs about the same whatever the size of the new batch, because every tree is replayed over all the stored rows. That's what keeps it exact.
Guarantees
refreshwith zero new rows returns the identical model, bit for bit. That's tested across 15 configurations: everyrefit_fullmode, categorical and cross features, linear leaves, weights, missing values, custom objectives, subsampling, and the binary classifier.fitis unchanged. Every existing model output is bit-identical.Storage
The stored copy is compact: numeric columns as 1-byte bin indices, categoricals as integer codes, and raw values only for the few columns that feed cross features. That's about 3x smaller than X for numeric data, and far smaller for text categories. It travels with the pickled model and grows with each refresh.
Not supported yet
With
store_training_data=True, these raiseNotImplementedErroratfit:n_ensembles > 1,quality=4or5)loss="Quantile"random_effects=TrueThese are the next slices of #131; the plan is in
benchmarks/REFRESH_PLAN.md.For reviewers
refit_full="replay") unchanged. The stored rows are rebuilt into an equivalent X, which is exact because every stored bin maps back to a value in the same bin._training_data_) and exposed throughn_samples_trained_. In scikit-learn,X_train_means the raw X.tests/test_refresh.pyand 21 intests/test_replay_kernel.py. The full suite passes: 1280 tests, 1 skipped.Part of #131.
🤖 Generated with Claude Code