Repository navigation
fix(generic): return NaN from expanding_mean/std when minp exceeds the length (engine parity) - #873
Merged
Conversation
…e length expanding_mean_nb and expanding_std_nb delegate to the rolling functions with window = len(a), so a minp larger than the number of rows raised 'minp must be <= window' in the Numba engine, while the Rust engine and pandas' expanding(min_periods=minp) return NaN. Use window = max(len(a), minp) so both engines agree with pandas.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
expanding_meanandexpanding_stdbehave differently across engines whenminpexceeds the number of rows:The Numba implementations delegate to the rolling functions with
window = len(a), so theminp <= windowvalidation that is right for a fixed window rejects an expanding window that has simply not grown tominpyet. pandas returns NaN for those rows, and so does the Rust engine. The discrepancy shows up in practice with short columns (e.g. the first chunks of a walk-forward split) whenengine="auto"picks a different engine depending on the input.Changes
expanding_mean_1d_nb,expanding_mean_nb,expanding_std_1d_nbandexpanding_std_nbusewindow = max(len(a), minp). A rolling window of at least the array length is an expanding window, so results are unchanged wheneverminp <= len(a), andminp > len(a)now yields NaN rows like pandas and the Rust engine.expanding_min/expanding_maxalready behaved correctly.test_expanding_minp_above_length, parametrized over both engines (Rust skipped when not installed), checks Series and DataFrame inputs against pandas.Found with a small parity harness that runs every
genericandreturnsdispatch function through both engines on clean, NaN-padded, constant and short inputs; this was the only difference that surfaced.