perf: opt-in validated tokenizer for per-atom lines (use_regex=False) - #38
Merged
Merged
Conversation
The per-line pcre2_match is the largest remaining read cost. Add an opt-in mode that skips it: extxyz_read_ll_opts(..., use_tokenizer) splits each atom line on whitespace and parses+validates each field by its column type, with no regex compile or match. extxyz_read_ll stays the regex default (thin wrapper); exposed on the Python side by wiring the existing `use_regex` flag through to the C backend (use_regex=False -> tokenizer), default unchanged. ABI-safe (new entry point; Fortran/test_C_main untouched, the latter gaining an optional "tok" arg). Validated, so it is safe: malformed numeric/bool fields or the wrong field count raise a clear error instead of silently parsing to 0 (the unvalidated prototype's flaw). Floats reuse parse_double_fast (already exact + validating) then strtod with a full-consume check (and an inf/nan/hex guard); ints use strtol; bools must be in the BOOL_RE set; field count must match. Bit-identical to the regex parser on valid input; marginally more lenient only on numeric edge cases (leading-zero ints, "1."/".5"), hence opt-in. ~1.5x on top of the existing read path (200k atoms 49.4 -> 32.9 ms; ~6.6x over ASE's built-in, and ~2x faster than extxyz-ng on the 20k single-frame bench) with the validation essentially free. New tests (equivalence + malformed rejection); benchmark + README show both modes; leaks clean in tokenizer mode. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The per-line
pcre2_matchis the single biggest remaining read cost. This adds an opt-in parser mode that replaces it with a whitespace tokenizer: each atom line is split on whitespace and each field is parsed and validated by its column type — no regex compile, no per-line match.No new argument — the existing
use_regexflag (already plumbed throughiread_dicts) is wired to the C backend, so it now means "regex or not" consistently for both backends. Default behaviour is unchanged.Validated — this is the difference from the throwaway experiment
The earlier prototype was ~1.5× but dropped validation (
H NOTANUM 0 0→ silent0). This version validates every field, so malformed input is a clear error, not a silent mis-parse:parse_double_fast(already exact + fully-consuming) thenstrtodwith an end-pointer check and aninf/nan/hex guard.strtolwith full-consume check (vsatoi).BOOL_REset; value computed as the regex fill does.It is bit-identical to the regex parser on valid input (incl. float bits, scientific, Fortran-
d, multi-column strings) — and the validation is essentially free (the float fast path already validates). The only difference is marginal extra leniency on numeric edge cases the grammar is picky about (leading-zero ints007,1./.5), which is why it's opt-in.ABI-safe
New
extxyz_read_ll_opts(..., int use_tokenizer);extxyz_read_llis a thin wrapper (regex). Fortran is untouched;test_C_maingains an optionaltok3rd arg for testing. In tokenizer mode the regex compile/JIT/match_dataare skipped entirely (also helps trajectories).Numbers (release build, same machine)
read_dictsregex →use_regex=False, a further ~1.5×:On the 20k species+pos single-frame bench it's 2.17 ms vs extxyz-ng's ~4.5 ms (~2×). README table + plot show both modes.
Verification
tests/test_fast_tokenizer.py: regex≡tokenizer bit-identical on a battery of valid files; malformed float/int/bool/field-count all raiseExtXYZErrorin fast mode.-W error::SyntaxWarningimport clean.leaks --atExit= 0 in tokenizer mode on valid / multi-column / malformed / read→write inputs.🤖 Generated with Claude Code