1616//!
1717//! Every arm is asserted bit-identical on every frontier before timing.
1818//!
19- //! # Measured 2026-09-16 — Xeon @ 2.10 GHz, `avx512f=true`, release, median of 7
19+ //! # Measured 2026-09-16 — Xeon @ 2.10 GHz, `avx512f=true`, release, median of 7, 2 runs
2020//!
21- //! Cache build (32 × `sel_f`): **850 µs**, 256 KiB resident.
21+ //! Every output buffer is `black_box`ed inside its timed closure (Codex P1 on
22+ //! #79: with `()`-returning closures and outputs never read, LLVM may drop the
23+ //! scatter stores; re-measured with them observed — the numbers held).
2224//!
23- //! | frontier | \|src\| | \|dst\| | recompute µs | **cached µs** | gather µs | cached vs recompute | cached vs gather | break-even |
24- //! |---|---:|---:|---:|---:|---:|---:|---:|---:|
25- //! | rand7 | 7 | 11 | 756 | **15.5** | 0.8 | 48.8× | 0.05× | 1.1 hops |
26- //! | rand655 | 655 | 1 279 | 853 | **28.0** | 26.5 | 30.4× | 0.95× | 1.0 hops |
27- //! | classid | 3 933 | 6 943 | 773 | **78.4** | 395 | 9.9× | 5.0× | 1.2 hops |
28- //! | hop2 | 6 943 | 12 125 | 848 | **141** | 617 | 6.0× | 4.4× | 1.2 hops |
29- //! | all | 65 536 | 56 841 | 1 395 | **797** | 3 022 | 1.8× | 3.8× | 1.4 hops |
30- //!
31- //! **The cache pays for itself on the SECOND hop** (break-even 1.0–1.4 hops
32- //! at every frontier), because a store generation's predicates do not change
33- //! between hops and the shipped body recomputes 64 contiguous 256 KiB passes
34- //! per hop to re-derive them. What is left per hop is one 8 KiB `mask_and`
35- //! per facet (≤ 5 µs total) plus the scatter, which is O(frontier) and now the
36- //! whole cost above ~1 %.
37- //!
38- //! **Cached beats the gather from ~1 % up (0.95× at 655 rows, 5× at the
39- //! classid frontier, 3.8× at the full population)** and loses below it —
40- //! at 7 rows the gather is 0.8 µs against 15.5 µs, and those 15.5 µs are
41- //! 32 × (8 KiB `mask_and` + a 1 024-word scatter walk over a nearly empty
42- //! mask): the same word-walk floor `ternlogq_sparse_reapply_probe` measured.
43- //! The crossover is where R1's "any whole-plane op is O(population)" boundary
44- //! sits: below ~1 % the population serialization is cheaper in absolute
45- //! terms, and the doctrine still declines it (lgj `LATEST_STATE` 2026-08-27).
46- //!
47- //! What this does NOT measure: the invalidation cost (a write to any facet's
48- //! classid or hi32 lane of the store generation drops all 32 masks), and a
49- //! per-`edge_classid` cache multiplies the 256 KiB by the number of edge
50- //! classes queried. The tile is keyed `(generation, edge_classid)` per M1b.
51- //!
52- //! # Measured 2026-09-16 — Xeon @ 2.10 GHz, `avx512f=true`, release, median of 7
53- //!
54- //! Cache build (32 × `sel_f`): **850 µs**, 256 KiB resident.
25+ //! Cache build (32 × `sel_f`): **843–930 µs**, 256 KiB resident.
5526//!
5627//! | frontier | src | dst | recompute µs | **cached µs** | gather µs (scalar walk) | break-even |
5728//! |---|---:|---:|---:|---:|---:|---:|
58- //! | rand7 | 7 | 11 | 756 | **15.5 ** | 0.8 | 1.1 hops |
59- //! | rand655 | 655 | 1 279 | 853 | **28. 0** | 26.5 | 1.0 hops |
60- //! | classid | 3 933 | 6 943 | 773 | **78.4 ** | 395 | 1.2 hops |
61- //! | hop2 | 6 943 | 12 125 | 848 | **141 ** | 617 | 1.2 hops |
62- //! | all | 65 536 | 56 841 | 1 395 | **797 ** | 3 022 | 1.4 hops |
29+ //! | rand7 | 7 | 11 | 728–782 | **14.4– 15.0 ** | 0.8 | 1.2 hops |
30+ //! | rand655 | 655 | 1 279 | 855–864 | **31.8–33. 0** | 34.7–36.1 | 1.0–1.1 hops |
31+ //! | classid | 3 933 | 6 943 | 947–954 | **96–100 ** | 462–632 | 1.0–1.1 hops |
32+ //! | hop2 | 6 943 | 12 125 | 1 019–1 102 | **159–196 ** | 705–758 | 0.9–1.1 hops |
33+ //! | all | 65 536 | 56 841 | 1 630–1 779 | **903–1 056 ** | 3 583–3 863 | 1.2–1.3 hops |
6334//!
6435//! **The cached tile pays for itself on the second hop** (break-even 1.0–1.4
6536//! hops at every frontier): a store generation's predicates do not change
@@ -286,7 +257,8 @@ fn main() {
286257 & mut out_r,
287258 & mut sel,
288259 & mut st,
289- )
260+ ) ;
261+ black_box ( & out_r) ;
290262 } ) ;
291263 let t_c = time_us ( || {
292264 hop_cached (
@@ -295,9 +267,13 @@ fn main() {
295267 black_box ( src) ,
296268 & mut out_c,
297269 & mut sel,
298- )
270+ ) ;
271+ black_box ( & out_c) ;
272+ } ) ;
273+ let t_g = time_us ( || {
274+ hop_gather ( black_box ( & aos) , black_box ( src) , & mut out_g) ;
275+ black_box ( & out_g) ;
299276 } ) ;
300- let t_g = time_us ( || hop_gather ( black_box ( & aos) , black_box ( src) , & mut out_g) ) ;
301277 println ! (
302278 "{:<9} {:>7} {:>7} | {:>11.1} {:>10.1} {:>10.1} | {:.1}x, {:.2}x break-even {:.1} hops" ,
303279 name, n_src, n_dst, t_r, t_c, t_g, t_r / t_c, t_g / t_c, t_build / ( t_r - t_c) . max( 1e-9 )
0 commit comments