Skip to content

Commit c410b14

Browse files
committed
hop_cached_vs_gather: black_box every output inside its timed closure
Codex P1 on #79: the timed closures returned () and the output buffers were never read afterwards, so LLVM could drop the scatter stores and the computations feeding them. Each closure now black_boxes its own output. Re-measured twice: every number held (cache build 843-930 µs; cached 14.4-15.0 / 31.8-33.0 / 96-100 / 159-196 / 903-1056 µs; break-even 0.9-1.3 hops). Table re-banked as observed ranges. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DCfrD5y19cvFc4AoyydXYv
1 parent 82dcbe8 commit c410b14

1 file changed

Lines changed: 18 additions & 42 deletions

File tree

‎native/lgj-abi/examples/hop_cached_vs_gather.rs‎

Lines changed: 18 additions & 42 deletions
Original file line numberDiff line numberDiff line change
@@ -16,50 +16,21 @@
1616
//!
1717
//! Every arm is asserted bit-identical on every frontier before timing.
1818
//!
19-
//! # Measured 2026-09-16 — Xeon @ 2.10 GHz, `avx512f=true`, release, median of 7
19+
//! # Measured 2026-09-16 — Xeon @ 2.10 GHz, `avx512f=true`, release, median of 7, 2 runs
2020
//!
21-
//! Cache build (32 × `sel_f`): **850 µs**, 256 KiB resident.
21+
//! Every output buffer is `black_box`ed inside its timed closure (Codex P1 on
22+
//! #79: with `()`-returning closures and outputs never read, LLVM may drop the
23+
//! scatter stores; re-measured with them observed — the numbers held).
2224
//!
23-
//! | frontier | \|src\| | \|dst\| | recompute µs | **cached µs** | gather µs | cached vs recompute | cached vs gather | break-even |
24-
//! |---|---:|---:|---:|---:|---:|---:|---:|---:|
25-
//! | rand7 | 7 | 11 | 756 | **15.5** | 0.8 | 48.8× | 0.05× | 1.1 hops |
26-
//! | rand655 | 655 | 1 279 | 853 | **28.0** | 26.5 | 30.4× | 0.95× | 1.0 hops |
27-
//! | classid | 3 933 | 6 943 | 773 | **78.4** | 395 | 9.9× | 5.0× | 1.2 hops |
28-
//! | hop2 | 6 943 | 12 125 | 848 | **141** | 617 | 6.0× | 4.4× | 1.2 hops |
29-
//! | all | 65 536 | 56 841 | 1 395 | **797** | 3 022 | 1.8× | 3.8× | 1.4 hops |
30-
//!
31-
//! **The cache pays for itself on the SECOND hop** (break-even 1.0–1.4 hops
32-
//! at every frontier), because a store generation's predicates do not change
33-
//! between hops and the shipped body recomputes 64 contiguous 256 KiB passes
34-
//! per hop to re-derive them. What is left per hop is one 8 KiB `mask_and`
35-
//! per facet (≤ 5 µs total) plus the scatter, which is O(frontier) and now the
36-
//! whole cost above ~1 %.
37-
//!
38-
//! **Cached beats the gather from ~1 % up (0.95× at 655 rows, 5× at the
39-
//! classid frontier, 3.8× at the full population)** and loses below it —
40-
//! at 7 rows the gather is 0.8 µs against 15.5 µs, and those 15.5 µs are
41-
//! 32 × (8 KiB `mask_and` + a 1 024-word scatter walk over a nearly empty
42-
//! mask): the same word-walk floor `ternlogq_sparse_reapply_probe` measured.
43-
//! The crossover is where R1's "any whole-plane op is O(population)" boundary
44-
//! sits: below ~1 % the population serialization is cheaper in absolute
45-
//! terms, and the doctrine still declines it (lgj `LATEST_STATE` 2026-08-27).
46-
//!
47-
//! What this does NOT measure: the invalidation cost (a write to any facet's
48-
//! classid or hi32 lane of the store generation drops all 32 masks), and a
49-
//! per-`edge_classid` cache multiplies the 256 KiB by the number of edge
50-
//! classes queried. The tile is keyed `(generation, edge_classid)` per M1b.
51-
//!
52-
//! # Measured 2026-09-16 — Xeon @ 2.10 GHz, `avx512f=true`, release, median of 7
53-
//!
54-
//! Cache build (32 × `sel_f`): **850 µs**, 256 KiB resident.
25+
//! Cache build (32 × `sel_f`): **843–930 µs**, 256 KiB resident.
5526
//!
5627
//! | frontier | src | dst | recompute µs | **cached µs** | gather µs (scalar walk) | break-even |
5728
//! |---|---:|---:|---:|---:|---:|---:|
58-
//! | rand7 | 7 | 11 | 756 | **15.5** | 0.8 | 1.1 hops |
59-
//! | rand655 | 655 | 1 279 | 853 | **28.0** | 26.5 | 1.0 hops |
60-
//! | classid | 3 933 | 6 943 | 773 | **78.4** | 395 | 1.2 hops |
61-
//! | hop2 | 6 943 | 12 125 | 848 | **141** | 617 | 1.2 hops |
62-
//! | all | 65 536 | 56 841 | 1 395 | **797** | 3 022 | 1.4 hops |
29+
//! | rand7 | 7 | 11 | 728–782 | **14.4–15.0** | 0.8 | 1.2 hops |
30+
//! | rand655 | 655 | 1 279 | 855–864 | **31.8–33.0** | 34.7–36.1 | 1.0–1.1 hops |
31+
//! | classid | 3 933 | 6 943 | 947–954 | **96–100** | 462–632 | 1.0–1.1 hops |
32+
//! | hop2 | 6 943 | 12 125 | 1 019–1 102 | **159–196** | 705–758 | 0.9–1.1 hops |
33+
//! | all | 65 536 | 56 841 | 1 630–1 779 | **903–1 056** | 3 583–3 863 | 1.2–1.3 hops |
6334
//!
6435
//! **The cached tile pays for itself on the second hop** (break-even 1.0–1.4
6536
//! hops at every frontier): a store generation's predicates do not change
@@ -286,7 +257,8 @@ fn main() {
286257
&mut out_r,
287258
&mut sel,
288259
&mut st,
289-
)
260+
);
261+
black_box(&out_r);
290262
});
291263
let t_c = time_us(|| {
292264
hop_cached(
@@ -295,9 +267,13 @@ fn main() {
295267
black_box(src),
296268
&mut out_c,
297269
&mut sel,
298-
)
270+
);
271+
black_box(&out_c);
272+
});
273+
let t_g = time_us(|| {
274+
hop_gather(black_box(&aos), black_box(src), &mut out_g);
275+
black_box(&out_g);
299276
});
300-
let t_g = time_us(|| hop_gather(black_box(&aos), black_box(src), &mut out_g));
301277
println!(
302278
"{:<9} {:>7} {:>7} | {:>11.1} {:>10.1} {:>10.1} | {:.1}x, {:.2}x break-even {:.1} hops",
303279
name, n_src, n_dst, t_r, t_c, t_g, t_r / t_c, t_g / t_c, t_build / (t_r - t_c).max(1e-9)

0 commit comments

Comments
 (0)