Skip to content

perf: opt-in validated tokenizer for per-atom lines (use_regex=False) - #38

Merged
jameskermode merged 1 commit into
masterfrom
perf/fast-tokenizer
Jun 10, 2026
Merged

jameskermode merged 1 commit into
masterfrom
perf/fast-tokenizer

Conversation

@jameskermode

Copy link
Copy Markdown
Member

What

The per-line pcre2_match is the single biggest remaining read cost. This adds an opt-in parser mode that replaces it with a whitespace tokenizer: each atom line is split on whitespace and each field is parsed and validated by its column type — no regex compile, no per-line match.

read_dicts("f.xyz", use_regex=False)     # tokenizer (C backend)
read_dicts("f.xyz")                       # regex, unchanged default

No new argument — the existing use_regex flag (already plumbed through iread_dicts) is wired to the C backend, so it now means "regex or not" consistently for both backends. Default behaviour is unchanged.

Validated — this is the difference from the throwaway experiment

The earlier prototype was ~1.5× but dropped validation (H NOTANUM 0 0 → silent 0). This version validates every field, so malformed input is a clear error, not a silent mis-parse:

  • float: parse_double_fast (already exact + fully-consuming) then strtod with an end-pointer check and an inf/nan/hex guard.
  • int: strtol with full-consume check (vs atoi).
  • bool: must be in the BOOL_RE set; value computed as the regex fill does.
  • field count: must equal the Properties column count.

It is bit-identical to the regex parser on valid input (incl. float bits, scientific, Fortran-d, multi-column strings) — and the validation is essentially free (the float fast path already validates). The only difference is marginal extra leniency on numeric edge cases the grammar is picky about (leading-zero ints 007, 1./.5), which is why it's opt-in.

ABI-safe

New extxyz_read_ll_opts(..., int use_tokenizer); extxyz_read_ll is a thin wrapper (regex). Fortran is untouched; test_C_main gains an optional tok 3rd arg for testing. In tokenizer mode the regex compile/JIT/match_data are skipped entirely (also helps trajectories).

Numbers (release build, same machine)

read_dicts regex → use_regex=False, a further ~1.5×:

atoms regex tokenizer ratio tokenizer / ASE built-in
16 000 3.96 ms 2.63 ms 1.50× 6.66×
64 000 15.6 ms 10.4 ms 1.50× 6.70×
200 000 49.4 ms 32.9 ms 1.50× 6.64×

On the 20k species+pos single-frame bench it's 2.17 ms vs extxyz-ng's ~4.5 ms (~2×). README table + plot show both modes.

Verification

  • New tests/test_fast_tokenizer.py: regex≡tokenizer bit-identical on a battery of valid files; malformed float/int/bool/field-count all raise ExtXYZError in fast mode.
  • Core suite (58) + ase-extxyz suite (38) green; -W error::SyntaxWarning import clean.
  • leaks --atExit = 0 in tokenizer mode on valid / multi-column / malformed / read→write inputs.

🤖 Generated with Claude Code

The per-line pcre2_match is the largest remaining read cost. Add an opt-in mode
that skips it: extxyz_read_ll_opts(..., use_tokenizer) splits each atom line on
whitespace and parses+validates each field by its column type, with no regex
compile or match. extxyz_read_ll stays the regex default (thin wrapper); exposed
on the Python side by wiring the existing `use_regex` flag through to the C
backend (use_regex=False -> tokenizer), default unchanged. ABI-safe (new entry
point; Fortran/test_C_main untouched, the latter gaining an optional "tok" arg).

Validated, so it is safe: malformed numeric/bool fields or the wrong field count
raise a clear error instead of silently parsing to 0 (the unvalidated prototype's
flaw). Floats reuse parse_double_fast (already exact + validating) then strtod
with a full-consume check (and an inf/nan/hex guard); ints use strtol; bools must
be in the BOOL_RE set; field count must match. Bit-identical to the regex parser
on valid input; marginally more lenient only on numeric edge cases (leading-zero
ints, "1."/".5"), hence opt-in.

~1.5x on top of the existing read path (200k atoms 49.4 -> 32.9 ms; ~6.6x over
ASE's built-in, and ~2x faster than extxyz-ng on the 20k single-frame bench) with
the validation essentially free. New tests (equivalence + malformed rejection);
benchmark + README show both modes; leaks clean in tokenizer mode.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@jameskermode
jameskermode merged commit 3f57fe9 into master Jun 10, 2026
26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant