Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 54 additions & 7 deletions .github/workflows/build-wheels.yml
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ jobs:
strategy:
matrix:
os: [ubuntu-22.04, macos-latest, macos-15-intel, windows-latest]
python-version: ['3.10', '3.11', '3.12', '3.13']
python-version: ['3.10', '3.11', '3.12', '3.13', '3.14']
fail-fast: false

steps:
Expand Down Expand Up @@ -61,7 +61,10 @@ jobs:
echo "PATH=${PCRE2_DIR}/bin;$PATH" >> $GITHUB_ENV

- name: Build wheels
uses: pypa/cibuildwheel@v2.22.0
# v4.x supports CPython 3.14 GA (cp314); CIBW_* env config below is
# still honoured. cp314-* matches the regular build, not cp314t
# (free-threaded), which we don't build.
uses: pypa/cibuildwheel@v4.1.0
env:
CIBW_BUILD: ${{ steps.set-py-version.outputs.cibw-python }}-*
CIBW_SKIP: pp* *musl*
Expand All @@ -85,23 +88,67 @@ jobs:
name: wheels-${{ matrix.os }}-py${{ matrix.python-version }}
path: ./wheelhouse/*.whl

build_sdist:
name: Build sdist
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@v4
with:
fetch-depth: 0

- name: Checkout submodules
run: git submodule update --init --recursive

- name: Install PCRE2
# meson-python configures the build (meson setup + meson dist) to
# produce the sdist, so the build deps incl. PCRE2 must be present.
run: |
sudo apt-get update -y
sudo apt-get install -y libpcre2-dev

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.12'

- name: Build sdist
run: |
python -m pip install --upgrade build
python -m build --sdist

- name: Upload sdist
uses: actions/upload-artifact@v4
with:
name: sdist
path: dist/*.tar.gz

publish:
name: Publish wheels to GitHub Release and PyPI
needs: build_wheels
name: Publish wheels and sdist to GitHub Release and PyPI
needs: [build_wheels, build_sdist]
if: startsWith(github.ref, 'refs/tags/v')
runs-on: ubuntu-latest
steps:
- name: Download all wheel artifacts
- name: Download wheel artifacts
uses: actions/download-artifact@v4
with:
path: dist
pattern: wheels-*
merge-multiple: true

- name: Attach wheels to GitHub Release
- name: Download sdist artifact
uses: actions/download-artifact@v4
with:
path: dist
pattern: sdist
merge-multiple: true

- name: Attach wheels and sdist to GitHub Release
uses: softprops/action-gh-release@v1
with:
files: dist/*.whl
files: |
dist/*.whl
dist/*.tar.gz
fail_on_unmatched_files: true
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
Expand Down
7 changes: 7 additions & 0 deletions .github/workflows/python-package.yml
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,13 @@ jobs:

- name: Test ase-extxyz (incl. fextxyz Fortran round-trip)
run: USE_FORTRAN=T pytest -v python/ase-extxyz/tests/

- name: Test with the first-char dispatch comment parser (use_cleri=False)
# Run the whole C-backend suite through the dispatcher so every fixture
# exercises it, in addition to the differential parity test.
run: |
USE_CLERI=false USE_CEXTXYZ=true pytest -v python/ase-extxyz/tests/
pytest -v tests/test_dispatch_parity.py
# Uncomment to get SSH access for testing
# - name: Setup tmate session
# if: failure()
Expand Down
74 changes: 53 additions & 21 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,21 +57,21 @@ forces, and a couple of `info` keys):

| atoms / frame | file size | ASE built-in `extxyz` | `cextxyz` plugin | `extxyz.read_dicts` (no Atoms) | speedup, plugin / built-in | speedup, parser / built-in |
|--:|--:|--:|--:|--:|--:|--:|
| 10 | 0.00 MB | 0.107 ms | 0.074 ms | 0.119 ms | 1.46× | 0.90× |
| 100 | 0.01 MB | 0.203 ms | 0.096 ms | 0.145 ms | 2.11× | 1.41× |
| 1 000 | 0.11 MB | 1.170 ms | 0.277 ms | 0.367 ms | 4.23× | 3.18× |
| 4 000 | 0.44 MB | 4.414 ms | 0.908 ms | 1.122 ms | 4.86× | 3.93× |
| 16 000 | 1.74 MB | 17.5 ms | 3.40 ms | 3.96 ms | 5.15× | 4.42× |
| 64 000 | 6.98 MB | 69.9 ms | 13.5 ms | 15.6 ms | 5.17× | 4.48× |
|200 000 | 21.80 MB |218.8 ms | 42.8 ms | 49.4 ms | 5.12× | 4.43× |
| 10 | 0.00 MB | 0.130 ms | 0.073 ms | 0.111 ms | 1.77× | 1.18× |
| 100 | 0.01 MB | 0.220 ms | 0.086 ms | 0.138 ms | 2.56× | 1.60× |
| 1 000 | 0.11 MB | 1.177 ms | 0.240 ms | 0.350 ms | 4.89× | 3.36× |
| 4 000 | 0.44 MB | 4.429 ms | 0.705 ms | 0.943 ms | 6.28× | 4.70× |
| 16 000 | 1.74 MB | 17.8 ms | 2.59 ms | 3.40 ms | 6.87× | 5.24× |
| 64 000 | 6.98 MB | 72.6 ms | 11.5 ms | 14.7 ms | 6.30× | 4.96× |
|200 000 | 21.80 MB |224.8 ms | 35.9 ms | 44.2 ms | 6.26× | 5.09× |

![Read-time benchmark](https://raw.githubusercontent.com/libAtoms/extxyz/master/benchmarks/read_speedup.png)

Below ~100 atoms per frame the per-call setup (file open, PCRE2 JIT
compile, libcleri grammar walk for the comment line) is larger than
the regex match itself, so the built-in is faster on tiny files. From
~1 000 atoms upwards the parser dominates and `cextxyz` runs at a
steady ~5× over the built-in end-to-end (~4.4× for the parser alone).
compile, libcleri grammar walk for the comment line) is a larger share
of the work, so the margin shrinks on tiny files. From ~1 000 atoms
upwards the parser dominates and `cextxyz` runs at a steady ~6× over
the built-in end-to-end (~5× for the regex parser alone).
The two cextxyz curves track each other closely: the `Frame → Atoms`
translation in the ASE plugin is kept cheap by aliasing the parser's
per-atom buffers directly into `atoms.arrays` (so `Atoms.__init__`
Expand All @@ -97,21 +97,51 @@ The single biggest remaining cost is the per-line `pcre2_match`. The default
per-atom lines are split on whitespace and each field is parsed and validated by
its column type, with no regex compile or match. It is **the default since
v0.4.2** (pass `use_regex=True` for the strict regex parser), and a further
**~1.5×** on top of everything above:
**~1.8×** on top of everything above:

| atoms / frame | `read_dicts` (regex) | `read_dicts` (`use_regex=False`) | tokenizer / regex | tokenizer / built-in |
|--:|--:|--:|--:|--:|
| 1 000 | 0.367 ms | 0.219 ms | 1.68× | 5.34× |
| 16 000 | 3.96 ms | 2.63 ms | 1.50× | 6.66× |
| 64 000 | 15.6 ms | 10.4 ms | 1.50× | 6.70× |
| 200 000 | 49.4 ms | 32.9 ms | 1.50× | 6.64× |
| 1 000 | 0.350 ms | 0.161 ms | 2.17× | 7.30× |
| 16 000 | 3.40 ms | 1.95 ms | 1.74× | 9.11× |
| 64 000 | 14.7 ms | 8.21 ms | 1.79× | 8.85× |
| 200 000 | 44.2 ms | 24.8 ms | 1.78× | 9.05× |

It validates each field (a malformed numeric/bool or the wrong field count is a
clear parse error, not a silent `0`) and is bit-identical to the regex parser on
valid input. The trade-off is that it is marginally more lenient than the grammar
on a few numeric edge cases (e.g. leading-zero integers `007`, `1.`/`.5`); pass
`use_regex=True` if you need the grammar enforced exactly.

### Comment-line parser (`use_cleri=False`)

The remaining per-frame cost is parsing the comment line. By default this walks
the libcleri grammar (PCRE2-backed). `use_cleri=False` (C backend only) instead
uses a hand-written first-char-dispatch parser that accepts the **same** language
— validated bit-identical against the grammar by a differential conformance test
(`tests/test_dispatch_parity.py`), with libcleri kept as the canonical grammar /
oracle / fallback — but builds the dicts in a single pass instead of constructing
and re-walking a generic parse tree.

Because the win is per comment line, it is amortised away on single huge frames
(the tables above are unchanged) and grows as frames get smaller. Sweeping
atoms-per-frame at a fixed ~1 M total atoms (C reader, whitespace tokenizer in
both; only the comment parser differs):

| atoms / frame | frames | `read_dicts` (cleri) | (`use_cleri=False`) | dispatch / cleri | full ASE read |
|--:|--:|--:|--:|--:|--:|
| 5 | 200 000 | 3.54 s | 2.00 s | 1.77× | 1.36× |
| 10 | 100 000 | 1.86 s | 1.07 s | 1.74× | 1.34× |
| 20 | 50 000 | 1.01 s | 0.613 s | 1.65× | 1.30× |
| 50 | 20 000 | 0.490 s | 0.332 s | 1.47× | 1.26× |
| 100 | 10 000 | 0.310 s | 0.231 s | 1.34× | 1.19× |
| 500 | 2 000 | 0.170 s | 0.153 s | 1.11× | 1.07× |
| 2 000 | 500 | 0.144 s | 0.135 s | 1.07× | 1.02× |

It is the libcleri grammar (`use_cleri=True`) by default for now; pass
`use_cleri=False` for the dispatch parser. ~75 % of its speedup is from not
building/walking a generic cleri parse tree (the matching itself is a small part
of the cost), so it stays grammar-faithful while skipping cleri's machinery.

### Marshalling in C

Once the C reader has parsed a frame it has to hand the `info`/`arrays` data
Expand Down Expand Up @@ -144,6 +174,8 @@ Reproduce locally (requires `extxyz`, `ase-extxyz`, `ase`, `matplotlib`):
```bash
python benchmarks/bench_read.py --max-atoms 200000 --repeats 3
python benchmarks/plot_bench.py
# comment-line parser, many small frames (use_cleri table above):
python benchmarks/bench_cleri_frames.py --total 1000000 --repeats 3
# writing (see below):
python benchmarks/bench_write.py --max-atoms 200000 --repeats 5
python benchmarks/plot_bench.py --in benchmarks/write_results.csv --out benchmarks/write_speedup.png
Expand All @@ -157,11 +189,11 @@ than `extxyz-ng`):

| atoms / frame | file size | ASE built-in `extxyz` | `cextxyz` plugin | `extxyz.write_dicts` (no Atoms) | speedup, plugin / built-in | speedup, writer / built-in |
|--:|--:|--:|--:|--:|--:|--:|
| 1 000 | 0.11 MB | 2.800 ms | 0.639 ms | 0.547 ms | 4.39× | 5.11× |
| 4 000 | 0.44 MB | 10.9 ms | 2.391 ms | 2.163 ms | 4.55× | 5.03× |
| 16 000 | 1.74 MB | 43.9 ms | 8.426 ms | 7.273 ms | 5.21× | 6.03× |
| 64 000 | 6.98 MB | 166.6 ms | 31.6 ms | 28.2 ms | 5.26× | 5.92× |
|200 000 | 21.80 MB | 521.3 ms | 106.7 ms | 92.5 ms | 4.88× | 5.64× |
| 1 000 | 0.11 MB | 2.794 ms | 0.630 ms | 0.565 ms | 4.44× | 4.95× |
| 4 000 | 0.44 MB | 10.8 ms | 2.104 ms | 1.957 ms | 5.11× | 5.50× |
| 16 000 | 1.74 MB | 41.4 ms | 8.681 ms | 7.667 ms | 4.77× | 5.41× |
| 64 000 | 6.98 MB | 167.3 ms | 33.5 ms | 28.8 ms | 4.99× | 5.80× |
|200 000 | 21.80 MB | 509.7 ms | 102.6 ms | 88.2 ms | 4.97× | 5.78× |

![Write-time benchmark](https://raw.githubusercontent.com/libAtoms/extxyz/master/benchmarks/write_speedup.png)

Expand Down
92 changes: 92 additions & 0 deletions benchmarks/bench_cleri_frames.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
"""Benchmark the comment-line parser: libcleri grammar (use_cleri=True, default)
vs the first-char dispatch parser (use_cleri=False).

The dispatch parser's win is on the PER-FRAME comment line, so it is amortised
away on single huge frames (see bench_read.py) and shows up on files with many
small frames. This sweeps atoms-per-frame at a fixed total atom count and times
the C reader both ways (per-atom tokenizer in both; only the comment parser
differs).

Run::

python benchmarks/bench_cleri_frames.py [--total 1000000] [--repeats 3]
"""
from __future__ import annotations

import argparse
import csv
import sys
import tempfile
import time
from pathlib import Path

import ase.io
import ase_extxyz.io # noqa: F401 (registers the 'cextxyz' format)
import extxyz

from bench_read import make_xyz # reuse the synthetic-frame generator


def _best(fn, repeats):
best = float('inf')
for _ in range(repeats):
t0 = time.perf_counter()
fn()
best = min(best, time.perf_counter() - t0)
return best


def main():
ap = argparse.ArgumentParser()
ap.add_argument('--out', type=Path, default=Path('benchmarks/cleri_results.csv'))
ap.add_argument('--total', type=int, default=1_000_000,
help='approx total atoms per file (held ~constant)')
ap.add_argument('--repeats', type=int, default=3)
args = ap.parse_args()

per_frame = [5, 10, 20, 50, 100, 500, 2000]

hdr = (f'{"atoms/frame":>11} {"frames":>7} {"MB":>6} '
f'{"cleri":>9} {"dispatch":>9} {"speedup":>8} {"full rd speedup":>15}')
print(hdr); print('-' * len(hdr))

rows = []
with tempfile.TemporaryDirectory() as tmp:
for nat in per_frame:
nframes = max(1, args.total // nat)
path = Path(tmp) / f'f_{nat}.xyz'
mb = make_xyz(path, nat, nframes) / 1e6
sp = str(path)

# parser only (read_dicts), per-atom tokenizer in both; comment parser differs
t_cleri = _best(lambda: extxyz.read_dicts(sp, use_cleri=True), args.repeats)
t_disp = _best(lambda: extxyz.read_dicts(sp, use_cleri=False), args.repeats)
# full ASE plugin read, both ways
t_cleri_full = _best(lambda: ase.io.read(sp, format='cextxyz', index=':',
use_cleri=True), args.repeats)
t_disp_full = _best(lambda: ase.io.read(sp, format='cextxyz', index=':',
use_cleri=False), args.repeats)

speedup = t_cleri / t_disp if t_disp else float('nan')
full_speedup = t_cleri_full / t_disp_full if t_disp_full else float('nan')
rows.append(dict(atoms_per_frame=nat, frames=nframes, file_mb=mb,
read_dicts_cleri_s=t_cleri, read_dicts_dispatch_s=t_disp,
dispatch_speedup=speedup,
full_read_cleri_s=t_cleri_full,
full_read_dispatch_s=t_disp_full,
full_read_speedup=full_speedup))
print(f'{nat:>11} {nframes:>7} {mb:>6.1f} '
f'{t_cleri:>8.3f}s {t_disp:>8.3f}s {speedup:>7.2f}x '
f'{full_speedup:>14.2f}x')

args.out.parent.mkdir(parents=True, exist_ok=True)
with args.out.open('w', newline='') as f:
w = csv.DictWriter(f, fieldnames=list(rows[0].keys()))
w.writeheader()
w.writerows(rows)
print(f'Wrote {args.out}')


if __name__ == '__main__':
sys.path.insert(0, str(Path(__file__).resolve().parent))
main()
8 changes: 5 additions & 3 deletions benchmarks/bench_read.py
Original file line number Diff line number Diff line change
Expand Up @@ -82,12 +82,14 @@ def time_read_ase(path: Path, fmt: str, repeats: int) -> float:


def time_read_dicts(path: Path, repeats: int) -> float:
"""Time the ASE-free dict-based reader — what extxyz.read_dicts costs."""
return _best_of(lambda: extxyz.read_dicts(str(path)), repeats)
"""Time the ASE-free dict-based reader with the strict regex per-atom parser
(use_regex=True) — the ``read_dicts`` baseline in the README tables. The
package default is now the tokenizer (use_regex=False), timed below."""
return _best_of(lambda: extxyz.read_dicts(str(path), use_regex=True), repeats)


def time_read_dicts_fast(path: Path, repeats: int) -> float:
"""As above but with the opt-in whitespace tokenizer (use_regex=False)."""
"""As above but with the default whitespace tokenizer (use_regex=False)."""
return _best_of(lambda: extxyz.read_dicts(str(path), use_regex=False), repeats)


Expand Down
8 changes: 8 additions & 0 deletions benchmarks/cleri_results.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
atoms_per_frame,frames,file_mb,read_dicts_cleri_s,read_dicts_dispatch_s,dispatch_speedup,full_read_cleri_s,full_read_dispatch_s,full_read_speedup
5,200000,147.8,3.54335541697219,2.001780458027497,1.7701019124063793,6.126253166003153,4.506593667087145,1.3593977222186266
10,100000,128.5,1.8616587079595774,1.0704289170680568,1.7391707924508686,3.1345056670252234,2.3414170830510557,1.3387216185083561
20,50000,118.75,1.0107304580742493,0.6128594169858843,1.6492044179481524,1.6579841249622405,1.2738971250364557,1.3015055080800133
50,20000,112.9,0.48988604196347296,0.33220262499526143,1.4746603581788698,0.7712598750367761,0.6124645000090823,1.2592727823822263
100,10000,110.96,0.3102357500465587,0.2310199160128832,1.3428961251516423,0.4610714999726042,0.3867935830494389,1.1920350289618722
500,2000,109.404,0.1701008330564946,0.15335649996995926,1.1091856757934313,0.22910087497439235,0.21320445893798023,1.0745594914646521
2000,500,109.1015,0.14380666706711054,0.1346410830738023,1.0680741998211958,0.1845930420095101,0.18073112494312227,1.0213683009365608
Binary file modified benchmarks/read_speedup.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
14 changes: 7 additions & 7 deletions benchmarks/results.csv
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
natoms,frames,file_mb,builtin_s,cextxyz_s,read_dicts_s,read_dicts_fast_s,speedup,parse_speedup,fast_speedup
10,1,0.001285,0.00010749998909886926,7.38749949960038e-05,0.00011920899851247668,5.129100463818759e-05,1.4551606955057577,0.9017774701598392,2.0958838661318113
100,1,0.011096,0.0002034579956671223,9.629200212657452e-05,0.0001446249953005463,6.579099863301963e-05,2.112927254328761,1.4067969042579032,3.0924898526317466
1000,1,0.109203,0.0011697909940266982,0.00027658400358632207,0.0003674159961519763,0.00021912499505560845,4.229423896026603,3.1838325121338245,5.33846444003266
4000,1,0.436203,0.004413624992594123,0.0009084579942282289,0.0011222080065635964,0.0007287080079549924,4.858369919837265,3.932982982459232,6.056781240788457
16000,1,1.744204,0.01751962500566151,0.0034028340014629066,0.003959790992666967,0.002632292002090253,5.148539422766341,4.424381245905565,6.65565408083507
64000,1,6.976204,0.06988954100233968,0.013509042008081451,0.015597417004755698,0.010424208012409508,5.173537913386456,4.480840704651941,6.70454205433541
200000,1,21.800211,0.21879487499245442,0.04276062498684041,0.04939795799145941,0.032940000004600734,5.116737069670721,4.429229140003777,6.642224497932463
10,1,0.001285,0.00012995791621506214,7.329101208597422e-05,0.00011058303061872721,3.7458958104252815e-05,1.7731767172571538,1.1752066794329092,3.4693414550766075
100,1,0.011096,0.00022045790683478117,8.595804683864117e-05,0.00013791699893772602,5.2332994528114796e-05,2.5647151714442815,1.5984824824554444,4.212598740481874
1000,1,0.109203,0.001177040976472199,0.00024049996864050627,0.00035045796539634466,0.0001612498890608549,4.894141912474061,3.3585796092294373,7.299483945864853
4000,1,0.436203,0.004429124994203448,0.0007049170089885592,0.0009425421012565494,0.0005173330428078771,6.28318644283888,4.699126954964413,8.561457760679518
16000,1,1.744204,0.017779374960809946,0.0025887920055538416,0.003395583014935255,0.0019517079927027225,6.867826740297067,5.2360301257865,9.109649100831469
64000,1,6.976204,0.07264966703951359,0.011525750043801963,0.014659667038358748,0.008212041924707592,6.303248531628651,4.955751508504059,8.84672383624009
200000,1,21.800211,0.22477775008883327,0.03588125004898757,0.044165458995848894,0.02484933298546821,6.264490500803376,5.089446712417505,9.04562509666889
10 changes: 5 additions & 5 deletions benchmarks/write_results.csv
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
natoms,frames,file_mb,builtin_s,cextxyz_s,write_dicts_s,pypython_s,speedup,write_speedup
1000,1,0.109203,0.002799833018798381,0.0006385000015143305,0.0005474580102600157,0.0030485409952234477,4.385016463834012,5.114242492257255
4000,1,0.436203,0.010870625003008172,0.002391167014138773,0.0021625419904012233,0.011302291008178145,4.546158816482104,5.026781006454034
16000,1,1.744204,0.04386191698722541,0.008425584004726261,0.007273458002600819,0.04454758297652006,5.2058013975792585,6.03040767837546
64000,1,6.976204,0.1665553750062827,0.031641958019463345,0.028150958998594433,0.1700303329853341,5.263750584076766,5.916508031381764
200000,1,21.800211,0.5213367920077872,0.10673779097851366,0.09247166602290235,0.5287190410017502,4.884275636852297,5.637800360152139
1000,1,0.109203,0.002793958061374724,0.0006298749940469861,0.0005647499347105622,0.0029531250474974513,4.435734213583189,4.947248135241741
4000,1,0.436203,0.010762332938611507,0.002104375045746565,0.001956957974471152,0.010912708006799221,5.114265615515972,5.499521747021634
16000,1,1.744204,0.04144945798907429,0.008681291015818715,0.007667165948078036,0.04373008303809911,4.774573034534458,5.40609897708874
64000,1,6.976204,0.16732429200783372,0.03351095807738602,0.028831499977968633,0.17090941697824746,4.993121701308446,5.803523650718598
200000,1,21.800211,0.5097258749883622,0.10255850001703948,0.08821750001516193,0.5198811669833958,4.970098771956242,5.778058490670849
Binary file modified benchmarks/write_speedup.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
2 changes: 2 additions & 0 deletions libextxyz/_extxyz.def
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,8 @@ EXPORTS
extxyz_write_ll_fmt
print_dict
free_dict
extxyz_dispatch_init
extxyz_dispatch_free
extxyz_fopen
extxyz_fclose
extxyz_ftell
Expand Down
Loading
Loading