Skip to content

Commit 7e5cfdf

Browse files
committed
board: LATEST_STATE entry for #79 — hop_cached_vs_gather, full table
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DCfrD5y19cvFc4AoyydXYv
1 parent c410b14 commit 7e5cfdf

1 file changed

Lines changed: 54 additions & 0 deletions

File tree

‎.claude/board/LATEST_STATE.md‎

Lines changed: 54 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,57 @@
1+
## 2026-09-16 — `hop_cached_vs_gather`: the M1b tile pays off on the second hop; the scatter walk is the access-shape question
2+
3+
**PR #79**, `native/lgj-abi/examples/hop_cached_vs_gather.rs` only — no ABI
4+
symbol, no Java change. Operator framing: not "is the gather faster" (R1 ruled
5+
the population serialization out) but *cached mask vs gather* as two shapes of
6+
the same selection, and then — for the random-access walk that remains — *"ob
7+
du bei kleinteiligen Aufgaben aus der Not eine Tugend machst und die
8+
kleinteilige Operation verwendest"*.
9+
10+
**Three arms, all bit-identical on every frontier before timing**
11+
(65 536 rows × 32 facets, `avx512f=true`, release, median of 7, two runs,
12+
every output `black_box`ed inside its timed closure — Codex P1 on #79):
13+
14+
| arm | per hop |
15+
|---|---|
16+
| `recompute` | the shipped `lgj_hop` body on FacetMajor: per facet two contiguous `eq_u32` passes + one `ternlog<AND3>` + scatter |
17+
| `cached` | `sel_f = class_f ∧ struct_f` built ONCE per store generation (32 × 8 KiB = **256 KiB**, the §9 M1b tile keyed `(generation, edge_classid)`); per hop one `mask_and` per facet + scatter |
18+
| `gather` | the scalar row-walk on AosRows (PR #40's shape) — the access-shape BASELINE, not a candidate |
19+
20+
Cache build: **843–930 µs**.
21+
22+
| frontier | src | dst | recompute µs | **cached µs** | gather µs | break-even |
23+
|---|---:|---:|---:|---:|---:|---:|
24+
| rand7 | 7 | 11 | 728–782 | **14.4–15.0** | 0.8 | 1.2 hops |
25+
| rand655 | 655 | 1 279 | 855–864 | **31.8–33.0** | 34.7–36.1 | 1.0–1.1 |
26+
| classid | 3 933 | 6 943 | 947–954 | **96–100** | 462–632 | 1.0–1.1 |
27+
| hop2 | 6 943 | 12 125 | 1 019–1 102 | **159–196** | 705–758 | 0.9–1.1 |
28+
| all | 65 536 | 56 841 | 1 630–1 779 | **903–1 056** | 3 583–3 863 | 1.2–1.3 |
29+
30+
- **The tile pays for itself on the second hop at every frontier.** A store
31+
generation's predicates do not change between hops; the shipped body
32+
re-derives them with 64 contiguous 256 KiB passes per hop. With the tile,
33+
what is left is ≤ 5 µs of `mask_and` plus the scatter, which is O(frontier)
34+
and the whole cost above ~1 %.
35+
- **Not measured here:** invalidation (a write to any classid/hi32 lane of the
36+
generation drops all 32 masks — the registry already does this wholesale for
37+
`cached_carving`), and the cache multiplies by the number of edge classes
38+
queried.
39+
- **The open question the gather column sets up** (the operator's actual
40+
point): inside the random-access scatter walk, use the op that matches the
41+
access granularity — one zmm `mask_cmpeq_epi32` per 4 facets deciding all 32
42+
facets of a visited row in 8 loads (the `simd_rowstore_facet_match` shape,
43+
plus the hi32 == 0 gate on lane 3), or one xmm compare per facet — instead
44+
of the scalar 64-load / 64-branch loop. Arms `gather_xmm` / `gather_zmm_row`
45+
are named in the module doc and NOT built. A facet is one xmm; a row is
46+
eight zmm; the shipped `emit` reads neither as such.
47+
48+
Cross-refs, same day: ndarray #311 (the `VPTERNLOGQ` tail descent 5–8× on
49+
1..7-word masks — `CallMask [u64; 3]` 18.1 → 2.3 ns — and the sparse
50+
re-apply probe showing a 64×2 rung is NOT the tool on a full-width 1 024-word
51+
mask, ≤ 1.24× from zmm chunk-skip only); lance-graph #1241 (the facet's
52+
per-axis LCP read straight off the `u128` masked to the axis bytes, 12.5 →
53+
5.8 ns for both axes — the `"{0}{1}" -f` done once at mint, never per call).
54+
155
## 2026-09-14 (6) — the lowering differential: one law, two implementations, finally compared
256

357
`plan_lower` lowers a FLAT OP LIST (a left fold seeded with all-ones,

0 commit comments

Comments
 (0)