|
| 1 | +## 2026-09-16 — `hop_cached_vs_gather`: the M1b tile pays off on the second hop; the scatter walk is the access-shape question |
| 2 | + |
| 3 | +**PR #79**, `native/lgj-abi/examples/hop_cached_vs_gather.rs` only — no ABI |
| 4 | +symbol, no Java change. Operator framing: not "is the gather faster" (R1 ruled |
| 5 | +the population serialization out) but *cached mask vs gather* as two shapes of |
| 6 | +the same selection, and then — for the random-access walk that remains — *"ob |
| 7 | +du bei kleinteiligen Aufgaben aus der Not eine Tugend machst und die |
| 8 | +kleinteilige Operation verwendest"*. |
| 9 | + |
| 10 | +**Three arms, all bit-identical on every frontier before timing** |
| 11 | +(65 536 rows × 32 facets, `avx512f=true`, release, median of 7, two runs, |
| 12 | +every output `black_box`ed inside its timed closure — Codex P1 on #79): |
| 13 | + |
| 14 | +| arm | per hop | |
| 15 | +|---|---| |
| 16 | +| `recompute` | the shipped `lgj_hop` body on FacetMajor: per facet two contiguous `eq_u32` passes + one `ternlog<AND3>` + scatter | |
| 17 | +| `cached` | `sel_f = class_f ∧ struct_f` built ONCE per store generation (32 × 8 KiB = **256 KiB**, the §9 M1b tile keyed `(generation, edge_classid)`); per hop one `mask_and` per facet + scatter | |
| 18 | +| `gather` | the scalar row-walk on AosRows (PR #40's shape) — the access-shape BASELINE, not a candidate | |
| 19 | + |
| 20 | +Cache build: **843–930 µs**. |
| 21 | + |
| 22 | +| frontier | src | dst | recompute µs | **cached µs** | gather µs | break-even | |
| 23 | +|---|---:|---:|---:|---:|---:|---:| |
| 24 | +| rand7 | 7 | 11 | 728–782 | **14.4–15.0** | 0.8 | 1.2 hops | |
| 25 | +| rand655 | 655 | 1 279 | 855–864 | **31.8–33.0** | 34.7–36.1 | 1.0–1.1 | |
| 26 | +| classid | 3 933 | 6 943 | 947–954 | **96–100** | 462–632 | 1.0–1.1 | |
| 27 | +| hop2 | 6 943 | 12 125 | 1 019–1 102 | **159–196** | 705–758 | 0.9–1.1 | |
| 28 | +| all | 65 536 | 56 841 | 1 630–1 779 | **903–1 056** | 3 583–3 863 | 1.2–1.3 | |
| 29 | + |
| 30 | +- **The tile pays for itself on the second hop at every frontier.** A store |
| 31 | + generation's predicates do not change between hops; the shipped body |
| 32 | + re-derives them with 64 contiguous 256 KiB passes per hop. With the tile, |
| 33 | + what is left is ≤ 5 µs of `mask_and` plus the scatter, which is O(frontier) |
| 34 | + and the whole cost above ~1 %. |
| 35 | +- **Not measured here:** invalidation (a write to any classid/hi32 lane of the |
| 36 | + generation drops all 32 masks — the registry already does this wholesale for |
| 37 | + `cached_carving`), and the cache multiplies by the number of edge classes |
| 38 | + queried. |
| 39 | +- **The open question the gather column sets up** (the operator's actual |
| 40 | + point): inside the random-access scatter walk, use the op that matches the |
| 41 | + access granularity — one zmm `mask_cmpeq_epi32` per 4 facets deciding all 32 |
| 42 | + facets of a visited row in 8 loads (the `simd_rowstore_facet_match` shape, |
| 43 | + plus the hi32 == 0 gate on lane 3), or one xmm compare per facet — instead |
| 44 | + of the scalar 64-load / 64-branch loop. Arms `gather_xmm` / `gather_zmm_row` |
| 45 | + are named in the module doc and NOT built. A facet is one xmm; a row is |
| 46 | + eight zmm; the shipped `emit` reads neither as such. |
| 47 | + |
| 48 | +Cross-refs, same day: ndarray #311 (the `VPTERNLOGQ` tail descent 5–8× on |
| 49 | +1..7-word masks — `CallMask [u64; 3]` 18.1 → 2.3 ns — and the sparse |
| 50 | +re-apply probe showing a 64×2 rung is NOT the tool on a full-width 1 024-word |
| 51 | +mask, ≤ 1.24× from zmm chunk-skip only); lance-graph #1241 (the facet's |
| 52 | +per-axis LCP read straight off the `u128` masked to the axis bytes, 12.5 → |
| 53 | +5.8 ns for both axes — the `"{0}{1}" -f` done once at mint, never per call). |
| 54 | + |
1 | 55 | ## 2026-09-14 (6) — the lowering differential: one law, two implementations, finally compared |
2 | 56 |
|
3 | 57 | `plan_lower` lowers a FLAT OP LIST (a left fold seeded with all-ones, |
|
0 commit comments