diff --git a/.claude/agents/jdk-toolchain-warden.md b/.claude/agents/jdk-toolchain-warden.md new file mode 100644 index 0000000..fc804ee --- /dev/null +++ b/.claude/agents/jdk-toolchain-warden.md @@ -0,0 +1,112 @@ +--- +name: jdk-toolchain-warden +description: Guards against the single most repeated failure in this repo's history — shipping Java work UNVERIFIED because a session concluded "no JDK here" without checking where one actually is. Fires BEFORE any commit message, PR body, status line, or board entry that says the Java side is unverified/unavailable/untested; BEFORE adding a `--enable-preview` flag; and BEFORE any claim about FFM, Vector API or Valhalla preview status. Also fires when a session is about to skip `AllTests` for any reason. +tools: Read, Glob, Grep, Bash +--- + +# JDK toolchain warden + +**THE RULE: "the Java side could not be verified" is a falsifiable claim, and +in this container it has always been false. Falsify it before writing it.** + +Four rungs, four minutes. Report "no JDK" only after all four fail, and say +which ones you ran. + +## The incident this card exists for + +Two commits on the ABI-minor-11 branch are titled **"UNVERIFIED: JDK 26 not +available in this sandbox"**, and the Java half of a 7,793-line change went to +PR unverified on that basis. + +Re-checked 2026-09-16: **`/opt/jdks/jdk-26.0.2` was present the whole time**, +and the suite runs `ALL PASSED (409 checks)` on it. `.claude/knowledge/ +jdk-toolchain-facts.md` had already named that exact path. + +The session looked at `java -version` and `/usr/lib/jvm` — which show only the +system OpenJDK 21 — and inferred absence. **Neither of those sees +`/opt/jdks`.** + +Two distinct failure shapes, and both recur: + +1. **Absence inferred from the wrong instrument.** `which java` answers "what + is on PATH", never "what is installed". It is not evidence about the second + question. +2. **A cached view reporting absence.** The apt index named a withdrawn + `openjdk-25` version, so the install died on a bare `404 Not Found` — which + reads as "no such package" rather than "your index is stale". One + `apt-get update` fixed it. Same shape as (1): a stale view is not a fact. + +## The check, in order + +```sh +ls -d /opt/jdks/*/ && for d in /opt/jdks/*/; do "$d/bin/java" -version; done # 1 +ls /usr/lib/jvm/ # 2 +sudo apt-get update && apt-cache policy openjdk-25-jdk-headless # 3 +# 4: Adoptium source for 26; jdk.java.net/valhalla for the JEP 401 EA build +``` + +Full ladder, with the verified-by-execution table and the commands: +`.claude/knowledge/jdk-toolchain-facts.md` § ACQUISITION LADDER. + +## Verdicts + +- **VERIFIABLE-NOW** — a rung succeeded. The Java half is a gate, not an + aspiration; run `AllTests` and report the 409 line. Naming the rung you used + is part of the report. +- **NEEDS-FETCH** — only the Valhalla EA is missing (`value class` / + `value record`). That is rung 4 and genuinely not on apt. Say so precisely + and name `jdk.java.net/valhalla`; do NOT generalize it to "no JDK". +- **GENUINELY-ABSENT** — all four rungs failed. Report which, with output. + This has never been observed; treat it as a finding worth recording, not as + a routine excuse. + +## What this card does NOT license + +`valhalla-lab/` and `bench/` are **measurement arms, not gates**. Their absence +never blocks a merge and must not be reported as if it did. The merge gate is +the Rust suite plus the 409-check `AllTests` run. + +And `--enable-preview` is a **classfile-poisoning flag**: every class compiled +with it can only run with it, transitively. Never let a preview-compiled class +reach the path a production consumer loads — that is why the Valhalla-flavoured +sources are physically separate from `java/`. + +## Why the Java half is worth this much guarding + +Operator, 2026-09-16: *"java is the low code intake glove around the +lance-graph spine — lance-graph-java just happens to offer the menu to the +table in a pleasing way, using masking ops, offering 5 star for the price of a +blink."* + +The menu is the whole product. `view.where(..).hop(..).count()` reads as +ordinary Java and costs a blink because **the work and the data are both +somewhere else**, and Java is never told. That fluency is exactly what a +409-check run protects: it is the only thing standing between "a pleasing +menu" and "a pleasing menu that lies about what the kitchen did." Shipping it +unverified forfeits the product, not a test. + +> **⊘ CORRECTED, same day — the menu is BORING, and the example above is the +> wrong one.** Operator: *"Java doesnt use masking ops. `Mask.minus()`, +> `RowStore.hop()`. Lance-graph does. Java just sees boring `sql()` handed to +> duckdb."* and *"nobody should ever start trying to optimize Java (except +> making it boring front)."* So `view.where(..).hop(..).count()` is the mask +> algebra on the wrong side of the wall, not the product. Read the paragraph +> above with `sql("select …")` in its place; every word of it still holds, and +> the reason the 409-check run matters is unchanged — it is what keeps a +> boring surface honest instead of a disguise, since **the boringness is handed +> down zero-copy**. The endgame: low-code *"Bring your own software"* against +> Palantir Foundry, where novel API is lock-in-by-learning-curve and therefore +> the enemy. Full ruling: `CLAUDE.md` § "THE JAVA SURFACE IS `sql()`, NOT THE +> MASK ALGEBRA". + +The allocation the menu sits on (operator, same day): + +| what | lives in | membrane that keeps Java out of it | +|---|---|---| +| **thinking** | lance-graph | **Panama** — computation never lives in Java | +| **SIMD** | ndarray | the `ndarray::simd` facade (see `simd-savant`) | +| **storage** | lance-graph | **Valhalla** — storage never lives in Java | + +Panama and Valhalla are ORTHOGONAL guarantees, not a layer and not a pair — +see `CLAUDE.md` § "the simd.rs isomorphism", whose middle row is explicitly +marked wrong for exactly this reason. diff --git a/.claude/board/AGENT_LOG.md b/.claude/board/AGENT_LOG.md index 8d6b063..4e0e026 100644 --- a/.claude/board/AGENT_LOG.md +++ b/.claude/board/AGENT_LOG.md @@ -1,3 +1,13 @@ +## 2026-09-14 — lgj-abi missing mask ops (one Opus worker, orchestrator-gated, ABI minor 11) + +**D-ids:** D-MRL-1a (TERNLOG / TERNARY_MATCH ops — now REACHABLE at the ABI; +the `ogar_loco::TERNLOG = 0x86` consumption is the lowering plan's, not this +change's). **Output:** `native/lgj-abi/src/{abi.rs,kernels.rs,exports.rs}`, +`docs/abi.md` §19. Worker ran `cargo test`/`clippy`/`fmt` in the crate only; +the orchestrator re-ran all three (164+3 green, clippy clean, fmt clean) and +`nm -D` (29 symbols) before this entry. One writer: this log and +`LATEST_STATE.md` written by the main thread. No commit by the worker. + ## 2026-08-28 — W1.1 epoch-recheck 5+3 council (8 agents, orchestrator-consolidated) **D-ids:** D-LGJ-MMV-1a (council half). **Output:** diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index 0f33359..cd6da7a 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -1,3 +1,305 @@ +## 2026-09-14 (6) — the lowering differential: one law, two implementations, finally compared + +`plan_lower` lowers a FLAT OP LIST (a left fold seeded with all-ones, +`acc &= p` / `acc |= p`) to a `mask_risc::Program`. `lance-graph-quack`'s +`lower` lowers a Boolean TREE to the same `Program`. They implement the same +law twice — the accumulator gate, the AND/OR gating asymmetry, and the drop +of ops before the first AND — and until now nothing compared them, because +each has its OWN oracle: this crate's is `pr4_equivalence.rs`'s frozen +pre-PR4 loop; quack's is a per-row reading in its own crate that has never +heard of `plan_lower`. Neither can see the other drift. + +`src/exports/tests/lowering_convergence.rs` closes that. Three tests, 603 +lines, all green. + +**The conversion is where the prefix rewrite turns out not to be a special +case.** `fold_to_tree` reads the op list as the tree it denotes: +`all_ones & p == p` makes the first AND the accumulator; `all_ones | p == +all_ones` makes an OR before that first AND simply never become a node. +`plan_lower` reaches the identical answer by scanning for the least AND +index. Same predicate, two routes — and +`the_all_rows_shortcut_is_the_same_condition_on_both_sides` pins exactly +that, two-sided: 3 of the 28 combine vectors are `AllRows` (one all-OR +vector per arity), 25 are not. + +**Deliberately a differential, NOT a delegation.** `lance-graph-quack` is a +**dev**-dependency, and the manifest comment says why: the membrane — the +`cdylib` Java's `Linker` loads — must not depend on a CONSUMER of the IR it +serves. Sharing the LAW between two independent implementations is the +point; sharing a runtime dependency in that direction would invert the +layering. `plan_lower` stays, and now it is pinned. + +**Both anti-vacuity bounds were floors, and both floors were hiding +something.** The worker's spec said `>= 15` of 28 and `>= 7` of 9. Measured, +both passed EXACTLY at their bound — which is what a floor looks like when +the fixture is dead: + +- the arity-4 arm appended `LT_I32(500)`, and the values lane is + `-150..=361`, so the op was always-true. The whole 16-vector arm collapsed + to eight saturated 1000s plus eight verbatim copies of the arity-3 row: + zero additional discriminating power, and it cleared the floor by sitting + on it. `LT_I32(200)` (676 of 1000) took the count **15 -> 21**. +- `LT_I32`/`LE_I32` were given `500` in the per-opcode arm too, so both + selected every row. Two arms agreeing that EVERY row survives cannot + separate `LtI32` from `LeI32` from any other always-true reading — the + exact swap this file exists to catch would have passed on those two rows. + `LT_I32(200)` / `LE_I32(300)` took it **7 of 9 -> 9 of 9**, and the four + ordered comparisons now carry distinct counts on purpose: this is a + differential between two ARMS, so a mis-map is visible only when it moves + one arm's count, and two opcodes selecting the same number of rows would + hide a swap between exactly those two. + +Both are now `assert_eq!`, measured, with the incident recorded at the +assertion. Zero `TODO`s left in the file. + +**And the fixture fix is itself measured, not asserted.** Mis-mapping +`LGJ_OP_LE_I32` to `Pred::LtI32` in `plan_lower` — a one-token change, and +exactly the defect class this file exists to catch — is **RED at operand +300** (`LeI32(300)` = 866, `LtI32(300)` = 865; one row is enough) and +**GREEN at operand 500**, where both readings select all 1000 and the two +arms agree on an answer neither computed correctly. So the operand is the +difference between a test and a decoration, and the pre-fix version of this +file would have shipped blind to a real mis-map on two of its nine opcodes. + +**Disable table — five arms, all red:** + +| arm | disable | +|---|---| +| the AND/OR asymmetry | `plan_lower` gates an OR on the accumulator | +| the prefix rewrite | `plan_lower` takes `k := 0` always | +| a one-opcode mis-map | `LGJ_OP_LE_I32` -> `Pred::LtI32` | +| the fold's OR node | the fold flattens an OR into the enclosing AND | +| the fold's dead prefix | an OR before the first AND becomes a leaf | + +Numbers, from `--nocapture`: per-opcode seeds 65 / 935 / 2 / 998 / 498 / +676 / 866 / 500 / 65; the 28-vector sweep runs 39, 498, 524, 1000 (n=2), +0, 40, 40, 73, 112, 531, 557, 1000 (n=3) — both identical to +`pr4_matrix.rs`'s own recorded figures, which is the cross-reference the +shared `(n=1000, seed=33)` fixture buys — and 0, 20, 20, 53, 64, 207, 233, +676, 676, 696, 696, 696, 724, 1000, 1000, 1000 (n=4). + +lgj-abi 182 lib tests (179 + 3), every integration binary green, clippy +`-D warnings` and fmt clean. + +## 2026-09-14 (5) — PR4 C3: the allocation gate, and what it actually measures + +Two new integration binaries, both disable-verified. + +**`tests/plan_eval_no_alloc.rs`** — one `#[test]`, a pass-through +`GlobalAlloc` over `System` counting bytes. Rust-side deliberately: this +repo's other allocation gates use `getThreadAllocatedBytes`, which measures the +JAVA heap and cannot see a Rust `vec!` at all, so citing them for this property +would have been an overclaim. + +**Measured, not predicted: 160 B per call at 64, 999, 1_000, 8_192 and 65_536 +rows — identical.** A 32-op plan costs 2_016 B. The old loop allocated +`2 * n_words * 8` per call: 16 KiB at 65_536 rows, growing without bound with +the population. So the claim the gate pins is not "zero" — the lowering's own +`Vec` is real and proportional to the OP count — it is the honest one: +**per-call allocation is independent of `n_rows`.** + +**A defect in the gate, found by its own disable run.** `measure()` originally +did one un-measured warm-up call per row count. Replacing the monotonic +`resize` with a per-size `*buf = vec![…]` — a cache keyed by row count, the +exact thing C-ALLOC denies — left the gate GREEN, because the per-measurement +warm-up absorbed the first call at each size, which is the only place such a +cache allocates. The plan predicted the shape ("a gate that warms up over a +fixed sweep and then measures that same sweep is green while the property is +false") and the gate had it anyway. Fixed: the arena is warmed ONCE, globally, +at the largest row count, and every arm — including `UNSEEN = 999`, which is +smaller and never seen before — is measured cold. D9/D10/D11 then all go red. + +D9 is worth its own line: routing the gate through `lgj_plan_eval_scalar` +reddens it at 592 → 278_848 B per call across the sweep. That is Fork A's cost +made visible — the row-at-a-time oracle allocates one `bool` per row per slot +by design, which is what lets it falsify a bit-packing bug — and it is why the +gate names `lgj_plan_eval` and only it. + +**`tests/c_one_evaluator.rs`** — the structural guard that there is one plan +evaluator. The invariant is an allowlist over MODULES: only `abi.rs` (the +definition), `exports.rs` (the `extern "C"` signatures and `validate_plan`, +which rejects rather than evaluates) and `plan_lower.rs` (the one lowering) may +name `LgjOpDesc` in code. Three narrower rules were tried and rejected against +the tree, each recorded in the file: "only `plan_lower` may call +`eval_predicate`" is false (the unfused single-predicate exports call it with +constant opcodes, correctly); "the opcode argument must be a literal" fires on +`lgj_mask_combine`, which legitimately takes a runtime combine; "no production +function may iterate `&[LgjOpDesc]`" is the right property and not textually +checkable. + +Two scanner defects its own assertions caught. Splitting production from test +at the first `#[cfg(test)]` silently discarded most of `exports.rs` — that file +and `registry.rs` carry test-only ITEMS long before their test module — so the +split now anchors on the trailing `#[cfg(test)] mod` and brace-matches to +prove it really is the file's last item. And scanning raw text flagged +`kernels.rs`, which names the type four times in PROSE and never touches it; +allowlisting it would have been the wrong repair, since kernels is exactly +where a second evaluator would most plausibly grow, so comments are stripped +instead. + +Disable table: D12a (a second production module holding plan ops) RED; D12b +(widening the allowlist to a module that does NOT hold them) RED — the guard +is a genuine equality, not the permissive `<=` the plan warned about; D12c +(the type renamed out from under the scanner) RED on the anti-vacuity arm. + +Gate: 174 lib + 1 + 3 + 1 integration, clippy `-D warnings` and fmt clean. + +## 2026-09-14 (4) — PR4 C2: `plan_eval` stops being a second evaluator + +`plan_eval_impl` no longer holds an opcode loop. It lowers +`&[LgjOpDesc]` to a `lance_graph_mask_risc::Program` (new module +`src/plan_lower.rs`) and runs it — `execute` on the SIMD path, +`reference_execute` on the scalar one. New path dep +`lance-graph-mask-risc`; `ndarray` remains the only source of SIMD in this +crate, because mask-risc names no ISA at all (a grep test in that crate +enforces it) and delegates every op to the same `ndarray::simd` facade +`kernels.rs` uses. What moved is who SEQUENCES the ops, not who computes them. + +**What the lowering is.** Let `k` be the least index whose combine is AND. The +old loop seeded an all-ones accumulator and folded each op in, so an `|=` +before the first AND cannot shrink anything and `all_ones & p == p` — every op +in `0..k` is dead. So ops `0..k` are dropped; op `k` becomes a bare `Pred` +writing slot 0 (it IS the accumulator); each later op writes slot 1 and folds +into slot 0. If no op combines with AND, the answer is every row and **no +program is built at all**. That prefix rewrite is not an optimisation bolted +on: it is what lets the lowering work without a fill/constant op, which the +mask-RISC deliberately does not have. + +**The asymmetry that is the whole correctness question.** A later AND-combined +op is gated `under` slot 0 — `acc & p` depends on `p` only where `acc` already +survives, so the predicate runs over the accumulator's live 64-row words. A +later OR-combined op is NOT gated: `acc | p` depends on `p` exactly where +`acc` is ZERO, so gating it would discard precisely the bits that matter and +quietly answer `acc`. `pr4_matrix`'s full `{AND, OR}^n` sweep against the +frozen oracle is what holds this. + +**Allocation.** The per-call `acc` and `scratch` vecs (`2 * n_words * 8` bytes +— 16 KiB at 65,536 rows) are gone. The accumulator lives in a thread-local +arena grown monotonically; a smaller call carves a strict prefix of the same +buffer, so allocation is a function of the thread's MAXIMUM row count, never +of its history. `acc` is deleted outright rather than kept: after +`validate_plan` returns `Ok` there is no error path left before the single +write, so the copy-then-publish dance was protecting against an empty set of +points. Both properties it bought survive — `dst_mask` is written exactly +once, and every error returns before `publish` is reached. + +**Fork A, and a rename that is mandatory.** +`lgj_plan_eval_scalar` runs mask-risc's row-at-a-time oracle rather than +`kernels::scalar_*`. So `simd_and_scalar_plans_agree_bit_for_bit` is renamed +`the_executor_and_the_row_oracle_agree_bit_for_bit`: under the old +implementation the name was accurate (two backends of one evaluator), and it +is not any more. The comparison is now executor-against-oracle — strictly +stronger, and a different claim. Backend parity is `ndarray::simd`'s business; +an ABI-level test could only ever have reached it through a proxy. A test +whose name claims a property it no longer checks is worse than no test. + +**Two guards worth naming.** Three `const _: () = assert!(…)` lines pin +`LANE_IDS/CLASSES/VALUES == 0/1/2`, because the lowering uses `lane_id` +DIRECTLY as an index into `Planes::lanes` — renumber the fixture and every +lowered predicate silently reads a different column, of the right kind, on a +plan the validator accepts. And `Planes::masks` is `&[]` on purpose: handing +`dst_mask` in as an input plane would turn a caller's dirty prior tail into +`ExecError::PlaneTail`, a spurious failure on a destination about to be +overwritten wholesale. + +**Honest about the error map.** `exec_error_to_status` exists so a bug in this +file becomes a status rather than a panic, and **every arm is unreachable +through the ABI** — `validate_plan` rejects unknown opcodes, bad combines, +out-of-range lanes and kind mismatches before the lowering is built, and the +lowering names no plane, no sum terminal and no blend. Its arms are therefore +NOT claimed to be individually falsifiable, and the doc comment says so rather +than leaving a future session hunting for the disable run that pins each one. + +Gate: 174/174 lib tests, `clippy --all-targets -D warnings` clean, fmt clean. +The 10 PR4 C1 falsifiers — written against the OLD loop and green before this +change — pass UNCHANGED against the new one, which is the equivalence claim. + +## 2026-09-14 (3) — minor 11 Java side VERIFIED under JDK 26: 409/409; the 11 `ApiSurfaceTest` failures were on `main` already, fence narrowed with a disable run + +JDK 26 obtained the documented way's equivalent: OpenJDK 26.0.2.1 GA tarball from +`jdk.java.net/26` (download.java.net) into `/opt/jdks/jdk-26.0.2.1`, symlinked to +the path `java/README.md` names (`/opt/jdks/jdk-26.0.2`). The `.so` rebuilt into +the Java-consumption `target/` reports `abi 0.11`. `AllTests` under the +documented `javac`/`java` commands: **ALL PASSED (409 checks)** — every suite, +including the new `MaskingOpCompletionTest` (48) and the three `OldAbiCompatTest` +gates. The "UNVERIFIED" caveat of the previous entry is discharged. + +**One pre-existing defect surfaced and fixed on the way.** The first run was 408 +passed / 11 failed, all in `ApiSurfaceTest`'s unnamed-materialiser fence. Baseline +check (origin/main `8720d1d`'s `java/` tree archived, compiled with JDK 26, same +`.so`): the identical 11 breaches — five exception types × `Throwable.getStackTrace` +/`getSuppressed`, plus `Carving.values()`. None is a project-authored crossing; the +fence excluded only `Object.class` and so held JDK-declared members and the +compiler-generated enum `values()` to the `materialize*/import*` naming law. Fixed +by exempting members whose declaring class is outside `com.adaworldapi.lancegraph` +and enum `values()`. **Disable run:** in a scratch copy, renaming +`Mask.materializeRows` → `rows` makes the narrowed fence report exactly one breach +(`Mask.rows returns long[]`) — it still fires on the thing it exists for. So the +CI lint gate (fmt/clippy/rust-test) never saw this: the Java suite is not in CI, +and nobody had run it under JDK 26 since the fence's allowlist was tightened. + +## 2026-09-14 (2) — Java side of minor 11 landed UNVERIFIED under the documented JDK 26 — read this before trusting it + +`Downcalls.Minor11` (lazy holder mirroring `Minor10`), `Engine.{maskTernlog, +ternaryMatch, reduceI32}` each behind `Abi.requireMinor(11)` FIRST (the shape +minors 2–4 still violate), `Mask.ternlog(b, c, imm)`, `RowStore. +maskOfFacetTernaryMatch(..)`, `View.{minOf,maxOf}` / `Lens.{min,max}` → +`OptionalLong`, `Layouts.REDUCE_OP_*` + `FACET_REGISTER_BYTES` (derived, never a +literal 12), `MaskingOpCompletionTest` (truth-table fixture for 8 immediates, +ternary-match parity against an independent Java loop with anti-vacuity, the +"extreme UNSELECTED row must not win" min/max falsifier), three +`OldAbiCompatTest` gates on the existing `-Dlgj.oldlibrary=` path. 367 +insertions, 0 deletions. + +**Verification status, stated plainly:** the documented toolchain is JDK 26 +(`/opt/jdks/jdk-26.0.2`) and it is NOT installed in this sandbox. The full tree +compiled with 0 errors under the only available JDK (21, preview FFM) in a +SCRATCH copy that patched three pre-existing JDK-22+ `Arena.allocate` overloads +the codebase already used before this change; at runtime that same JDK 21 build +fails in the untouched `SmokeTest` at `Engine.rowCount` (`WrongMethodTypeException`, +a VarHandle/FFM API difference) — so NO Java test in this change has been RUN. +The `.so` was rebuilt into the Java-consumption `target/` per +`valhalla-lab/README.md` and reports `abi 0.11`. This repo's CI runs fmt + +clippy + rust-test only; nothing gates the Java side. **The next session with +JDK 26 runs `AllTests` before anything else cites this entry as done.** Landed +on the branch rather than left uncommitted in a live checkout (container-loss +insurance), with this caveat in the commit message as well. + +## 2026-09-14 — ABI minor 10 → 11: the ndarray masking facade reaches the ABI in THREE symbols, not fifteen + +**Branch `claude/clone-repositories-71a5sw`** (post ndarray #306). Every +`ndarray::simd` mask primitive now has a `kernels.rs` wrapper; the ABI grew +`lgj_mask_ternlog` (runtime `u8` immediate → the const-generic facade; xor/not/ +maj3 and 252 more ride it, no `lgj_mask_xor`/`lgj_mask_not`), `lgj_op_ternary_match` +(the TCAM / prefix-ancestry op — the one genuinely new 24-byte operand), and +`lgj_reduce_i32` (op-coded: sum 0 / min 1 / max 2, per `abi.md` §15's +pre-commitment). The i32 comparison family (`eq/ne/lt/le/ge`, `ne_u32`) landed as +**op-codes 3–8 on the existing compare symbol — zero new symbols**. Measured +`nm -D`: 26 → **29** `T lgj_` (abi.md §1/§7 previously disagreed 26 vs 24; both now +read 29). Deliberate non-exports with revisit conditions (`abi.md` §19.5): +`mask_any`/`mask_all` (= `lgj_mask_count > 0` / `== n_rows` at identical cost), +`blend_i32` (vacuous: no resource carries two `I32` lanes), `ternary_match_u64` +(needs 128 bits of operand). One correction to the brief's census: +`eq_u32_strided_to_mask` was already `lgj_op_eq_classid` (minor 2). + +Gates (orchestrator, in `native/lgj-abi`): `cargo test` 164 + 3 (was 138 + 3), +`clippy --all-targets -D warnings` clean, `fmt --check` clean, release `.so` +800 KB. Seven disable runs (`abi.md` §19.7) — two were vacuous on first try +(an `if false {}` beside the real call; a knob that did not bind) and were +re-done to assert the guard TEXTUALLY absent; a real bug caught by an +anti-vacuity assertion (`FACET_PAYLOAD_HI32_OFFSET` is facet-relative, not +register-relative). + +**Java side NOT touched** (not a one-line addition): follow-up = `Downcalls.Minor11` +mirroring `Minor10` (`Downcalls.java:461`) + `Engine` wrappers behind +`Abi.requireMinor(11)` + `Mask`/`View` facade methods. Nothing breaks meanwhile +(`Layouts.LGJ_ABI_MINOR = 1`, `requireMinor` tests `>=`). Pre-existing E2/E3 +residue noticed, not fixed: `kernels.rs` private `FACET_CLASSID_BYTES` duplicates +`rowstore::FACET_CLASSID_BYTES`; `scalar_rowstore_{classid_mask,facet_match}` +are `pub fn` in `src/main` reachable only from tests. Measurement caveat: the +`lance-graph` path dep was a live checkout with concurrent uncommitted edits +during the run (its `8e5eb8f` + WIP); ndarray resolved to the local checkout. + ## 2026-09-05 — storno: the gate's own first run corrected two claims above Corrects the entry immediately below, which is left in place per the diff --git a/.claude/board/exec-runs/pr4-abi-membrane.md b/.claude/board/exec-runs/pr4-abi-membrane.md new file mode 100644 index 0000000..885740d --- /dev/null +++ b/.claude/board/exec-runs/pr4-abi-membrane.md @@ -0,0 +1,144 @@ +# abi-membrane-warden — ruling on PR4 (proposed, pre-implementation) + +**Scope:** `docs/abi.md` contract only. Not Java ergonomics (java-surface-warden), +not SIMD provenance (simd-savant), not registry internals (handle-lifecycle-auditor). +Read-only pass: no code written, no cargo run. + +**Subject:** replace `native/lgj-abi/src/exports.rs::plan_eval_impl`'s allocating +evaluator loop with the shared `lance-graph-mask-risc` / `ndarray::simd` floor. + +## Verdicts + +| Q | Verdict | +|---|---| +| Q1 path dep on `lance-graph-mask-risc` | **PERMITTED, GATED** — fence extension owed in the same commit | +| Q2 ABI surface change | **NO surface change** (29 symbols, same struct, same manifest, no minor bump) — but **BLOCKED on a semantic gap** | +| Q3 where scratch lives | (B) resource-owned **BLOCKED**, (C) scratch handle **BLOCKED**, (E) Java-owned **BLOCKED**, (D) thread-local admissible, **(F) substrate-first is the ruling** | + +## Q1 evidence + +- `tests/g11_contract_import_fence.rs` `const CRATE = "lance_graph_contract"`, + `const ALLOWED = ["canonical_node","class_view","facet","ontology"]`. It scans + `lgj-abi/src` for that literal ONLY. A `lance_graph_mask_risc::` import is + invisible to it — **the fence does not speak to a new crate dep at all.** +- mask-risc is contract-shaped: its `Cargo.toml` has EXACTLY ONE dependency, + `ndarray { default-features = false, features = ["std"] }` — identical + coordinates to lgj-abi's own line. Zero contract dep ("`Planes::masks` is + `&[&[u64]]`"). `#![forbid(unsafe_code)]` (lib.rs:79). No `cfg(target_feature)` + (law L3). Workspace MEMBER (`lance-graph/Cargo.toml:8`), so CI-covered. +- No new `[patch]` risk: the existing `[patch."…/lance-graph"]` exists because + `ogar-class-view` pulls the contract by git; mask-risc pulls nothing. +- **Gap:** G11 scans `src/` only, so a future mask-risc contract import is + invisible from here and the engine could arrive transitively. + +**Keeping it honest:** extend `g11_contract_import_fence.rs` with a CRATE +allowlist — parse `[dependencies]` from `lgj-abi/Cargo.toml` (and mask-risc's) +and assert equality with a named `ALLOWED_CRATES`, reusing the existing +`documented_allowlist` / `g11_documented_allowlists_equal_the_enforced_one` +shape. Disable-verify with a bogus dep line. + +## Q2 evidence + +- Surface is exactly **29** `pub [unsafe] extern "C" fn` in exports.rs, matching + abi.md §7/§2 minor 11. (`lgj_rowstore_open_handle` at :3741 is a `cfg(test)` + helper, not an export.) +- `plan_eval_impl` (:1680) is PRIVATE; `lgj_plan_eval` (:1778) and + `lgj_plan_eval_scalar` (:1803) signatures unchanged. +- `LgjOpDesc` unchanged: opcodes 1..9 (abi.rs:263-306) map **1:1 and + exhaustively** onto `mask_risc::Pred` (ir.rs:62-83). Manifest unchanged + (`size_of_op_desc`/`align_of_op_desc` derive from an unchanged struct). + **No minor bump owed** — a body swap adds nothing observable. +- PR4 must NOT take a new symbol for scratch: abi.md §1 (growth is a smell), + §19.4 precedent (15 capabilities / 3 symbols), §4 ("Rust. Always."). + +**BLOCKER — the all-set accumulator is inexpressible in mask-risc's IR.** +abi.md §7 normative `V0 = all`; exports.rs:1718-1723 implements it. +`MaskOp` (ir.rs:95-120) = {Pred, And, Or, Xor, AndNot, Not, Ternlog} — **no +Fill/Const**. Pre-filling a `Scratch` slot is refused by design: +`ExecError::ScratchReadBeforeWrite` (value.rs:56), enforced in `validate` +(reference.rs:106) called by BOTH `execute` (exec.rs:376) and +`reference_execute` (reference.rs:348); its doc says "pre-filled scratch is a +named PR5 gap, not a supported input." +The tempting rewrite (start from `op0`) holds only for `combine == AND` +(`all ∩ op0 = op0`, abi.md §19.4). For `ops[0].combine == OR` the answer is +`all ∪ op0 = all` — every row. **And no test pins it:** `LGJ_COMBINE_OR` +appears at exports.rs:2792 and :2958, both at op index >= 1. + +Unblock: `MaskOp::Fill { value, dst }` lands in mask-risc FIRST (+ reference +arm, differential case, no_alloc case, tail-clear semantics). The +`Pred; Not; Or` fabrication is available without an IR change but burns a full +O(n_rows) sweep to synthesise a constant — reject it. + +**Hazard 2 — `Path::Scalar` != `reference_execute`.** kernels.rs:814-819 selects +between two word-parallel implementations; `reference_execute` is a row-at-a-time +oracle that never touches ndarray (law L4) and allocates (reference.rs:396). +Re-pointing `lgj_plan_eval_scalar` at it strengthens the claim but changes what +abi.md §7's membrane parity test compares. Rule deliberately; do not let it drift. + +**Hazard 3 — ternlog owner (mask-risc lib.rs:19-23).** lgj carries +`kernels::simd_mask_ternlog_assign_dyn` (kernels.rs:360-362); mask-risc carries +`ternlog_dispatch` (758 lines). Real duplication, but its call sites are +`lgj_mask_ternlog` (:816) and `lgj_hop` (:2162, const-generic `::`, not +even the dyn table) — NOT plan_eval. File separately; do not bundle. + +## Q3 evidence + +- **(B) resource-owned — BLOCKED.** registry.rs:123-127: `Payload::Pattern` is + "read-only … no lock needed, because no ABI path mutates it"; RowStore + "likewise lock-free". Only `Mask` has a `RwLock` (:129), and :120 says why + ("so bulk ops on distinct masks do not serialize"). Mutable scratch on the + pattern needs a lock (first same-resource serialization point in the ABI; + today two `plan_eval`s on one pattern are fully parallel) or unsound interior + mutability across the cloned `Arc` (registry.rs:133-134). + abi.md §4 already admits contention is unbenchmarked. **No precedent:** grep + of `lgj-abi/src` finds only `OnceLock` (registry.rs:226, + class_view_provider.rs:84/:298 — init-once) and the mask `RwLock`. No Cell / + RefCell / Mutex / static mut / thread_local anywhere. +- **(C) scratch handle — BLOCKED.** Needs >=2 lifecycle symbols (Q2 rules that + out) AND a parameter on an existing symbol = **major** bump (abi.md §2). Drags + scratch into the sole-closer contract and PARENT_CLOSED for zero capability. +- **(E) Java-owned buffer — BLOCKED.** Signature change (major) + hands Java a + writable native buffer for general use, against CLAUDE.md "mutation crosses as + verbs, not writable memory" whose one exception is `RowStore.importRows`. +- **(D) thread-local — admissible, not first.** Invisible at the ABI, ownership + stays Rust, no lock added. Honest costs: per-thread retention for process life, + `catch_unwind` (§9) must not leave an inconsistent arena, and it would be the + first `thread_local!` in this crate. +- **(F) substrate-first — THE RULING.** Two measured reasons: + 1. Allocation census, exports.rs non-test: `lgj_row_facet_match` :1558 (1); + `lgj_rowstore_facet_match_count` :1627 (1); `plan_eval_impl` :1718+:1724 + (2); **`lgj_hop` :2090 `to_vec()` + :2097 + :2133 + :2134 (4)**; + `mask_ternlog_impl` :736/:740/:750 (up to 3). plan_eval is 2 of ~11 and + **not the worst site**. A bespoke fix here gets invented twice. + 2. Missing-capability STOP rule: the lacked capability is "reusable + non-allocating scratch for `extern "C"` bodies". mask-risc's `Scratch` + (exec.rs:46-102) + `tests/no_alloc.rs` IS that capability — but it lacks + (i) `Fill`, (ii) any contract for reuse across differently-shaped programs. + +**Trap that must be named:** `Scratch::for_program` (exec.rs:63-75) allocates +`scratch_slots × words` PER CALL. A `plan_eval_impl` that builds a `Program` +(`Vec`, ir.rs:169) + fresh `Scratch` each downcall replaces two 8 KB +`vec!`s with a `Vec` PLUS the same two 8 KB buffers — strictly worse +while looking like the fix. mask-risc's `no_alloc` test builds both ONCE outside +its 1000-iteration loop: **the zero-alloc law is a property of `execute`, not of +getting to `execute`**, and the per-call path is PR4's entire premise. + +## Falsifiers owed (each disable-verified red-then-green) + +1. `ops[0].combine == OR` → `out_count == n_rows`, dst all-set clean tail. + **Land BEFORE the swap** — currently untested, and it is the exact semantic + the likely rewrite breaks silently. +2. Native allocation gate on `lgj_plan_eval` (counting-global-allocator shape + mask-risc already uses), population-size independent. No current gate + measures the native side. +3. Existing `lgj_plan_eval` vs `lgj_plan_eval_scalar` parity suite green, + unchanged — if it moves, Hazard 2 fired. +4. G11 crate-allowlist test, disabled by a bogus `[dependencies]` line. + +## Measurement caveat + +"16 KB/call" is arithmetic (`n_words = 1024` × 2 × 8 B), **not a measurement**. +Repo rule is measure-then-pin; R8's recorded finding is that bulk crossings cost +nothing. PR4 must carry a before/after number or drop the performance framing — +the architecture argument (ONE evaluator, not two) is the strong one and stands +alone. diff --git a/.claude/board/exec-runs/pr4-kernel-membrane.md b/.claude/board/exec-runs/pr4-kernel-membrane.md new file mode 100644 index 0000000..0963f69 --- /dev/null +++ b/.claude/board/exec-runs/pr4-kernel-membrane.md @@ -0,0 +1,260 @@ +# PR4 pre-ruling — kernel-membrane-warden (T1/T2) + +**Scope:** a PROPOSED change, not written code. Read-only run; no cargo. +**Question:** what tier is `lance-graph-mask-risc`, and does `lgj_plan_eval` +consuming it cross T1/T2 in the legal direction? + +**Verdicts:** (a) **T2-COMPOSER — the card's T2 row is too narrow, split the +row, do NOT add a rung.** (b) **HAND-COMPOSED ×2, and the violating tier is +T1, not T2 — the one table belongs in ndarray; the filed TECH_DEBT entry's +two options are both wrong.** (c) **NAMED / legal as composition, BLOCKED on +an absent fence — G11 cannot see this dependency edge.** + +--- + +## (a) mask-risc is T2 (a composer), not a second T1 + +`execute` (`exec.rs:357-512`) runs an arbitrary-length `Program` under an +opcode loop, resolves operands and aliasing (`two_input`, `exec.rs:175-208`), +manages caller-owned scratch slots, and remaps ternlog immediates +(`remap_imm`, `:149`). A T1 word has fixed arity and one pass; this is a +composer of T1 words by construction — the same shape as `plan_eval_impl`. + +It sits correctly ABOVE the T0/T1 membrane: `exec.rs:25-34` imports ~30 +`ndarray::simd` facade words and nothing else, and L2 is structurally gated +by `the_crate_names_no_isa` (`exec.rs:566-603`), which greps every production +module for `target_feature` / `core::arch` / `std::arch` / `cfg(target_` / +`is_x86_feature_detected!`. + +**The card cannot classify it, and that is the card's defect.** My T2 row +enumerates "`lgj_hop`, `where`, `plan_eval`, the ABI exports" — i.e. T2 is +implicitly *the membrane crate*. mask-risc is T2-shaped but exports no ABI +symbol, crosses no Panama boundary, holds no handle: it is a **reusable T2 +library**, a T2 whose consumer is another T2 rather than T3. + +**Fix: split the T2 row into two ROLES; do not add a T1.5.** +- **T2-composer** — IR + evaluator: opcode loop, operand resolution, scratch. + (`mask_risc::execute`, `plan_eval_impl`.) +- **T2-membrane** — the crossing: handles, status codes, generation checks. + (`exports.rs`.) + +A new rung would be actively wrong: the T1/T2 law (no hand-composed T1 op, no +computed geometry) must apply to mask-risc UNCHANGED, and it does bite there +(`ternlog_self`'s `fill(0)` / `fill(u64::MAX)`, `exec.rs:215-225`; `clear_tail`, +`:135`) — each licensed by a NAMED T1 gap, which is the correct handling. A +rung would create a tier where those rules are weaker. + +Precedent for splitting a row rather than adding a rung is in the doctrine +itself (`membrane-tiers.md:47-49`): *"The ladder does not need a sixth tier. +**T1 was described too narrowly.**"* Same move, one tier up. + +## (b) The duplicate ternlog table — T1 is the violator + +**What is duplicated is not an algorithm.** Both tables call the identical +`ndarray::simd` word. What is duplicated is the **runtime-`u8` → +const-generic-`i32` bridge**: + +| | lgj-abi | mask-risc | +|---|---|---| +| shape | `static TERNLOG_ASSIGN: [[fn; 16]; 16]`, `ternlog_row16!` (`kernels.rs:282-342`) | two 256-arm `match`es, generated (`ternlog_dispatch.rs`, 758 lines) | +| forms | in-place only (`:360`) | in-place `:320` **and** out-of-place `:58` | +| gate | none | `--check` regenerate-and-diff (`rust-test.yml:126`) | +| consumers | 1 (`exports.rs:768`) | `exec.rs:38` | + +**Root cause, measured:** `ndarray::simd` exports ONLY the const-generic form +— `mask_ternlog` (`simd_masking_ops.rs:595`) and +`mask_ternlog_assign` (`:630`), re-exported at +`simd.rs:805-806`. **There is no runtime-immediate form at T1.** A const +generic instantiates only from a literal, so *every* consumer holding a +runtime immediate is forced into a 256-way fan-out. Both consumers were +forced; neither did anything wrong locally. + +So: **a T1 gap, surfacing as a T2 duplication.** Per my card's own rule — +*"if none exists, say 'the T1 primitive is missing; add it at T1 first (W1a +discipline), then call it,' never 'compose it at T2 for now'"* — the verdict +is HAND-COMPOSED at both sites, with the primitive owed at T1. + +**The one table belongs in ndarray.** Three reasons, two of them already +written down in this tree: +1. `kernels.rs:305-319` argues it for itself: *"`ndarray`'s POLYFILL LAW puts + truth-table specialisation inside the backend file… a locally-written SIMD + abstraction, the thing `abi.md` §8 forbids outright."* The identical + reasoning condemns hosting the *dispatch* outside ndarray — the fan-out + exists solely to reach ndarray's specialization, so it is part of ndarray's + realization, not of anyone's behavior. +2. A backend may beat a 256-way branch: `VPTERNLOGQ` takes its immediate as an + instruction byte, so AVX-512 could offer a genuinely dynamic form. A + consumer-side fan-out forecloses that — a consumer making a realization + decision, which the polyfill law forbids. +3. lgj `CLAUDE.md` E5 / the STOP rule: *"A facade method that cannot be one + delegation is the signal the substrate is missing a word."* + `simd_mask_ternlog_assign_dyn` is verbatim that. + +**The filed TECH_DEBT entry is right on the facts and wrong on the options.** +`TECH_DEBT.md:3-9` correctly identifies the duplication, the agreement on index +convention, and the cheap-now/expensive-later economics. But its option set — +*"either lgj-abi delegates to the crate, or the crate is scoped to the +evaluator"* — omits the only resolution that REMOVES the duplication: +- option A makes the membrane crate depend on a lance-graph workspace crate + solely to obtain a fan-out, and drags an unfenced edge across G11 (see (c)); +- option B **preserves** the duplication and relabels it. +Storno the entry with the third option. + +**Migration order — additive only, three commits, each green:** +1. **ndarray**: add `mask_ternlog_dyn(imm: u8, …)` + `mask_ternlog_assign_dyn` + to `simd_masking_ops.rs`, re-export from `simd.rs`. Purely additive; the + const-generic forms STAY (correct when the immediate is static, e.g. + `exports.rs:2162`'s `AND3`). Falsifier: all 256 immediates agree with the + const-generic form. Nothing downstream moves. +2. **mask-risc**: bodies become two delegations; delete + `tools/gen_ternlog_dispatch.py` **and** `rust-test.yml:126` in the same + commit (a generator with no output is the stale-gate shape). Keep the public + names `ternlog_dispatch`/`ternlog_dispatch_assign` as one-line wrappers so + `exec.rs:38` and `lib.rs:96` do not move. +3. **lgj-abi**: `simd_mask_ternlog_assign_dyn` keeps its name and signature and + becomes one delegation; `TERNLOG_ASSIGN` + `ternlog_row16!` deleted. + `exports.rs:768` does not move — **no ABI minor bump** (implementation + change behind an existing symbol). + +1 strictly before 2 and 3. 2 and 3 are independent. Doing 2-or-3 first is the +only red window. + +### The generalization — worth more than the table + +The tail clear has the SAME shape and **three** spellings: +- `ndarray/src/simd_masking_ops.rs:857` — `fn clear_mask_tail` — **private** +- `lance-graph-mask-risc/src/exec.rs:135` — `fn clear_tail`, whose own doc says + it is *"Narrower than `ndarray`'s private `clear_mask_tail`"* +- `lgj-abi/src/abi.rs:526` — `pub fn clear_tail_bits` + +Three tiers, one function, because T1 keeps it private. Same root cause as the +ternlog table; nobody filed it. **Rule to add: if two consumers of +`ndarray::simd` independently write the same helper, that helper is a missing +T1 export, not a consumer concern.** Fixing ternlog alone leaves the pattern +live — file `clear_tail` as its own TECH_DEBT entry. + +**Not charged as GEOMETRY-LEAK, deliberately.** `clear_tail` computes +`n_rows % 64` / `n_rows / 64`, but the bit order is NORMATIVE and PUBLISHED +(`kernels.rs:29-32`; `lib.rs:26-28` "`ndarray::simd`'s normative mask order"). +A published contract constant is not a computed geometry — same distinction lgj +`CLAUDE.md` draws for the 512×32×(4+12) row contract. It is a duplication with +a different fix (export it), and over-charging it would be wrong. + +## (c) `plan_eval_impl` consuming mask-risc + +**Same shape, confirmed:** + +| | `plan_eval_impl` (`exports.rs:1680-1756`) | `execute` (`exec.rs:357-512`) | +|---|---|---| +| input | flat `&[LgjOpDesc]` | `&Program { ops, terminal }` | +| per op | `eval_predicate` → `combine_into` (`:1731-1738`) | `run_pred` / `two_input` / ternlog arm | +| combine | AND, OR only | full Boolean + 256 ternlog | +| terminal | popcount → `out_count` | `Terminal::{Count,Any,All,MaskedSum,…}` | + +So it is **T2 calling T2**, with the callee strictly more expressive. + +**No rule exists.** The doctrine states the vertical law only — +*"A tier may only know the vocabulary of the membrane directly beneath it. +Nothing crosses a membrane except by NAME"* (`membrane-tiers.md:15-16`) — and +every row of the table is about crossing DOWN or UP. The horizontal case is +unaddressed. It must be ALLOWED (otherwise every helper is forbidden; +`plan_eval_impl` already calls its peers `validate_plan` and +`resolve_pattern_and_mask`). + +**Minimal proposed rule — the no-laundering rule:** + +> **T2 may call a T2 peer, provided the call does not re-cross a membrane.** A +> T2 composer may delegate to another T2 composer only if the callee reaches T1 +> by the same route the caller would have: the same `ndarray::simd` facade, no +> second SIMD provenance, no second geometry spelling, and no widening of the +> caller's own import fence. **A tier may never obtain vocabulary through a peer +> that its fence would have refused directly.** + +Vertical analog already enforced: `Cargo.toml`'s fence comment — the cognitive +modules *"compile UNCONDITIONALLY… so the fence, not `default-features`, is the +only barrier."* + +**Applying it: composition LEGAL.** mask-risc's only dependency is `ndarray`, +`default-features = false, features = ["std"]` — the SAME coordinates lgj-abi +uses; its manifest says outright *"No contract dep."* Plus +`#![forbid(unsafe_code)]` (`lib.rs:79`) and the structural no-ISA test. No +laundering. + +**But the EDGE is unfenced — this is the blocker:** +1. **G11 structurally cannot see it.** `ALLOWED = ["canonical_node", + "class_view", "facet", "ontology"]` and `const CRATE: &str = + "lance_graph_contract"`; `references()` scans for that one string. A + `use lance_graph_mask_risc::…` is invisible. A test named "the + contract-import fence" would give false assurance of coverage — the exact + `ISS-LGJ-G11-FENCE-WAS-PROSE` failure mode, one crate over. +2. **Policy has no verdict for this category.** lgj `CLAUDE.md`: *"the full + lance-graph ENGINE is never a dependency."* mask-risc is neither engine nor + contract — it is a third thing: a lance-graph **workspace member** + (`lance-graph/Cargo.toml:8`). +3. **First non-leaf dependency.** Today: `ndarray`, `lance-graph-contract` + (zero-dep by design), optional `ogar-class-view`. mask-risc is zero-risk + *today* (one dep, std-only) and nothing pins it there. + +**Requirement: the fence moves in the SAME commit as the dependency.** +Generalize `g11_contract_import_fence.rs` from one `CRATE` constant to a +per-crate policy table admitting `lance_graph_mask_risc` by name; update +`CLAUDE.md` § Enforcement and the `Cargo.toml` comment together (the test +parses both and fails on drift, in either direction). + +### Two sub-findings PR4 must decide, in neither doc + +- **`lgj_plan_eval_scalar` would be orphaned.** `kernels::eval_predicate` / + `combine_into`'s `Path::Scalar` arms are a SHIPPED symbol whose whole value + is independence (`kernels.rs:20-27`: *"if the reference shared code with the + SIMD path, a parity test between them would be checking that a function + agrees with itself"*). mask-risc has its own oracle (`reference.rs`, L4) but + it is not reachable across the ABI. Delegating naively either kills that + symbol's independence or silently leaves two evaluators. +- **Not a like-for-like swap.** `LgjOpDesc[]` (AND/OR) → `Program` (full + Boolean + ternlog + terminals) is a NEW lowering pass. Correct per + *"Extend the plan language, NOT the ABI surface"* (`membrane-tiers.md:105-118`), + but materially more than "pick one owner". + +## Doc defect found in passing — the law numbering does not agree with itself + +`lib.rs:40` — *"## The four laws this crate is built to make structural"* (4 +laws, unnumbered-as-L). `exec.rs:4-23` — L1..L5 (5 laws). Not the same set: +lib.rs's law 2 (*"Masks choose admissibility; magnitude is a separate reduction +over survivors"*) has no L-number; exec.rs's **L3** (one delegation per op) and +**L5** (one materialiser) are absent from lib.rs's four. In lib.rs, "no ISA" is +law **3**; in exec.rs it is **L2**. Anyone citing "L2" from `lib.rs` cites the +wrong law. One-line fix; worth doing before PR4 cites a law number. + +**L3 is correctly labelled `[claimed, unverified]`** (`exec.rs:6-7`, +`TECH_DEBT.md:15-18`) — no instrument counts facade calls. Note that if step (b) +lands, L3 gets *easier* to instrument, because the fan-out stops being a +delegation the counter would have to special-case. + +--- + +## Orchestrator verification (2026-09-14, post-ruling) + +The ruling's root-cause claim for finding (b) — that the duplicate 256-arm +ternlog tables exist only because `ndarray::simd` exports the const-generic +form ALONE — was checked against the source rather than taken on the agent's +reading. Confirmed: + + src/simd_masking_ops.rs:595 pub fn mask_ternlog(...) + src/simd_masking_ops.rs:630 pub fn mask_ternlog_assign(...) + src/simd.rs:805-806 the only two names re-exported + +There is no runtime-immediate (`_dyn`) form at any level of the facade. A +const generic instantiates only from a literal, so a consumer holding a +runtime `u8` immediate has no option but to fan out 256 ways. Both `lgj-abi` +(`ternlog_row16!`) and `lance-graph-mask-risc` (`ternlog_dispatch.rs`) did +exactly that, independently, and neither erred locally. + +So the migration order the ruling gives is sound and its first step is +genuinely additive: `mask_ternlog_dyn` / `mask_ternlog_assign_dyn` land in +ndarray with the const-generic forms untouched (still correct wherever the +immediate IS static, e.g. `lgj_hop`'s `AND3`), nothing downstream moves, and +the falsifier is that all 256 immediates agree with the const-generic form. + +NOT started. It is PR4-scope and the arc runs one PR at a time; #1226 was +mid-landing when this was verified, and opening a third in-flight PR to save +twenty minutes is not a trade worth making. diff --git a/.claude/knowledge/jdk-toolchain-facts.md b/.claude/knowledge/jdk-toolchain-facts.md index a3a71a8..db5be3c 100644 --- a/.claude/knowledge/jdk-toolchain-facts.md +++ b/.claude/knowledge/jdk-toolchain-facts.md @@ -63,3 +63,103 @@ and ran on `/opt/jdks/jdk-27` with `--enable-preview`, and Any doc or code comment asserting "value classes are final in JDK 28" or "Vector API is finalized" without re-verifying against a real build is wrong until re-checked — both were incubating/preview at last verification. + +--- + +## ⊘ ACQUISITION LADDER — added 2026-09-16, because this doc was RIGHT and a session still concluded the opposite + +This file already named `/opt/jdks/jdk-26.0.2` as the production path. A later +session nonetheless shipped two commits titled **"UNVERIFIED: JDK 26 not +available in this sandbox"** and left the Java half of an ABI-minor-11 change +unverified through to PR. + +Re-checked 2026-09-16: **`/opt/jdks/jdk-26.0.2` was present the whole time** +(26.0.2.1, 2026-08-18), and the full suite runs `ALL PASSED (409 checks)` on +it. Nothing was missing. The session looked at `java -version` / `/usr/lib/jvm` +— which show only the system OpenJDK 21 — and inferred absence. + +**The generalizable defect is in the SHAPE of what this doc recorded, not in +its accuracy.** It recorded a LOCATION. A location is environment-specific and +evaporates when a container is rebuilt, and a reader who does not find it has +no next move. A METHOD survives. So the ladder below is the method; check the +rungs in order and only report "no JDK" after all four fail. + +### Rung 1 — the pre-provisioned path (check this FIRST, always) + +```sh +ls -d /opt/jdks/*/ && for d in /opt/jdks/*/; do "$d/bin/java" -version; done +``` + +`which java` and `/usr/lib/jvm` do **not** see `/opt/jdks`. A session that only +runs those two will conclude "21 only" while a 26 sits one directory away. That +is exactly what happened. + +### Rung 2 — the system archive (Ubuntu noble → JDK 25) + +```sh +sudo apt-get update # NOT optional, see the 404 trap +sudo apt-get install -y openjdk-25-jdk-headless +``` + +Panama FFM is **final since 22**, so 25 runs the whole `internal/ffm` membrane +and the entire suite. Verified 2026-09-16: `ALL PASSED (409 checks)` on +`openjdk 25.0.4`, identical to 26. + +**The 404 trap.** The pre-seeded index named `openjdk-25 25.0.2+10-1~24.04`, a +version already withdrawn from the pool, so the install died on a bare +`404 Not Found` — which reads as *"no such package"* rather than *"your index +is stale"*. `apt-get update` moved the candidate to `25.0.4+7-1~24.04` and it +installed immediately. **A cached view reporting absence is not evidence of +absence** — the same failure shape as rung 1. + +### Rung 3 — an extra apt source (Adoptium → JDK 26, and 8..26 generally) + +Ubuntu noble stops at 25. Adoptium carries 8 through 26: + +```sh +curl -fsSL https://packages.adoptium.net/artifactory/api/gpg/key/public \ + | sudo tee /etc/apt/keyrings/adoptium.asc >/dev/null +. /etc/os-release && echo "deb [signed-by=/etc/apt/keyrings/adoptium.asc] \ +https://packages.adoptium.net/artifactory/deb $VERSION_CODENAME main" \ + | sudo tee /etc/apt/sources.list.d/adoptium.list +sudo apt-get update && sudo apt-get install -y temurin-26-jdk +``` + +Verified 2026-09-16: `temurin-26-jdk` → `26.0.2.1`, `ALL PASSED (409 checks)`. +`apt-cache search temurin` shows **no 27+**, so this rung tops out at 26. + +### Rung 4 — `jdk.java.net` for the Valhalla EA, and this one IS now needed + +**`/opt/jdks/jdk-27` is GONE from this container** — the JEP 401 EA build this +doc's table names for `value class` / `value record` is no longer +pre-provisioned, and no apt rung can replace it (Adoptium stops at 26, and +mainline 27 would not help anyway: `value record` is a Valhalla EA feature, not +a mainline one). + +`https://jdk.java.net/valhalla/` is reachable through this environment's proxy +(verified `HTTP 200`). Re-fetch there when `valhalla-lab` is the actual work. + +### Which rung for which need + +| need | rung | status 2026-09-16 | +|---|---|---| +| the 16-suite `AllTests` run (409 checks) | 1, 2 or 3 — any | **verified on all three**: `/opt/jdks/jdk-26.0.2`, `openjdk-25`, `temurin-26` | +| Panama FFM / the whole `internal/ffm` membrane | 1, 2 or 3 | final since 22, so any of them | +| `valhalla-lab/src/stable` (records, JDK 26) | 1 or 3 | fine | +| `valhalla-lab/src/valhalla` (`value record`, JEP 401 EA) | **4 only** | **pre-provisioned build GONE; refetch** | + +`bench/` and `valhalla-lab/` are measurement arms, **not gates**. The merge gate +is the Rust suite plus the 409-check Java run, and every rung that satisfies it +is reachable here. + +### The standing consequence + +**"The Java side could not be verified" is not an acceptable status line for +this repo.** Four rungs, four minutes. A PR that claims the Java surface is +unverified is claiming something falsifiable by `ls /opt/jdks`. + +## Falsifier for this section + +If a future container has none of rungs 1-3, the claim "always reachable" is +dead and this section must be re-graded rather than trusted. Check by running +rung 1 and rung 2 and reporting both outcomes — not by recalling this table. diff --git a/.claude/plans/pr4-plan-eval-mask-risc-equivalence-v1.md b/.claude/plans/pr4-plan-eval-mask-risc-equivalence-v1.md new file mode 100644 index 0000000..f9c5dd1 --- /dev/null +++ b/.claude/plans/pr4-plan-eval-mask-risc-equivalence-v1.md @@ -0,0 +1,749 @@ +# pr4-plan-eval-mask-risc-equivalence-v1 — the falsifier set for PR4 + +> **Status:** PROPOSAL (2026-09-14). Read-only design pass, authored by an Opus +> planner and persisted by the orchestrator. Scope: the falsifier set only — the +> tests, gates, disable table and commit order that make PR4's claims +> refutable. It does not design the lowering; it designs what would catch a +> wrong one. Nothing here was run; nothing is claimed to compile or pass. +> +> Companion rulings, same session: +> `.claude/board/exec-runs/pr4-abi-membrane.md` (ABI surface + scratch +> ownership) and `.claude/board/exec-runs/pr4-kernel-membrane.md` (T1/T2 tier, +> the duplicate ternlog table). Read all three before writing PR4 code. + +## §0 — What PR4 changes, and the three separate claims it makes + +`native/lgj-abi/src/exports.rs::plan_eval_impl` (`:1680`-`:1756`) today: +validate the plan (`validate_plan`, `:1651`), allocate `acc` and `scratch` +(`:1718`, `:1724`), seed `acc` to all-ones + `clear_tail_bits`, then per op +`kernels::eval_predicate` followed by `kernels::combine_into`, then a final +`clear_tail_bits`, `kernels::popcount`, and one `copy_from_slice` publish. + +PR4 replaces the opcode loop with the shared mask-risc floor and removes the +per-call allocation. That is **three claims needing three kinds of evidence**: + +| # | Claim | Evidence that can settle it | +|---|---|---| +| **C-BEHAVE** | the new `plan_eval` answers **identically to the old one** | a differential against a **frozen copy of the old loop** — nothing else | +| **C-ALLOC** | no per-call allocation, **at any row count** | a Rust-side counting global allocator, swept over n and both directions | +| **C-ONE** | there is now **ONE** evaluator, not two | a **structural** guard, plus an instrument for the cost half | + +Conflating these is the standard failure. C-ONE is not falsifiable by any +differential: two evaluators that agree are exactly what a differential +certifies. + +## §1 — The behaviour-equivalence falsifier + +### 1.1 The reference side: **both**, and the old loop is load-bearing + +Three answers per input: + +1. **`legacy_plan_eval_impl`** — a byte-for-byte frozen copy of today's body, + moved into `#[cfg(test)]`, with a header comment pinning the pre-PR4 sha. +2. **the new `lgj_plan_eval`** (mask-risc `execute`). +3. **`lgj_plan_eval_scalar`** (see §1.2). + +**Why mask-risc's own oracle is not enough, and this is the central point.** +PR4 introduces a component that did not exist before: the lowering +`&[LgjOpDesc] -> mask_risc::Program`. It sits ABOVE both `execute` and +`reference_execute`. A bug in it (a swapped opcode, a dropped all-ones seed, a +mis-packed TCAM operand) produces the same wrong `Program` for both, both +evaluate it faithfully, and mask-risc's differential stays green. Only an +oracle that never sees the `Program` catches it. That oracle is the old loop. + +**Freeze discipline (non-negotiable).** The frozen copy must not be refactored +to share a helper with the new path — a shared helper is a shared bug and turns +the differential into a tautology. Two checks: the copy calls only +`kernels::{eval_predicate, combine_into, popcount}` + `clear_tail_bits`; and +disable row **D7** breaks the new path only, then the oracle only, asserting +the differential reddens both times. One direction alone leaves "the oracle is +inert" indistinguishable from "the guard works". + +### 1.2 The `lgj_plan_eval_scalar` fork — decide explicitly + +`kernels::Path::{Simd, Scalar}` exists so SIMD/scalar parity is falsifiable +through the membrane. mask-risc has no `Path`. + +- **Fork A — scalar -> `reference_execute`.** Elegant, but `reference_scratch` + **allocates by design**, so **C-ALLOC's gate must be scoped to + `lgj_plan_eval` only**. A gate accidentally written against the scalar symbol + is green-for-the-wrong-reason; row **D9** disables exactly that. +- **Fork B — scalar keeps `kernels::scalar_*`.** Two evaluators remain, so + C-ONE must be narrowed in the PR text to "one *SIMD* evaluator", with the + structural guard narrowed to match or it is decoration. + +State the fork in the PR description. Do not let it be settled by whichever +code compiles first. + +### 1.3 The input space + +Compare **status, `out_count`, and `dst_mask` words bit-for-bit including the +tail word**, across: + +- **row counts** `[0, 1, 63, 64, 65, 127, 128, 999, 1000, 4097, 65_536]`. + `999` is deliberate — not in the allocation gate's warm-up set (§2.4) and in + no existing test. +- **seeds** `[0, 7, 0xFEEDFACE, 0xABCD]`. +- **op counts** `{1, 2, 3, 4, 8}`. Nothing today exceeds 4. +- **combine modes** — the full `{AND, OR}^n` product at `n <= 3`, sampled at 4 + and 8. **The most important axis, and currently 1-dimensional** (trap V1). +- **opcodes** — all nine, x **three positions**: first, non-first after AND, + non-first after OR. Today the seven minor-11 opcodes appear only in + single-op plans. +- **operands** — per opcode, one selecting ~nothing, ~half, ~everything. +- **TCAM** (op 9): `care in {0, 0b111, 0b1000, u32::MAX}` x `pattern in + {5, 7, 0xDEADBEEF}`, always including `pattern != care`. +- **error inputs** — the five from `a_bad_plan_leaves_dst_mask_untouched`, each + ALSO at a non-first and a last position, plus `MASK_LENGTH_MISMATCH` and + `EMPTY_PLAN` with and without a null `ops`. +- **dst_mask prior state** — `{fresh-empty, fresh-all, re-used, poisoned via + set_words with a dirty tail}`. **Entirely absent today** (trap V4). +- **scratch prior state** — `{cold, primed dense, primed sparse}` (trap V3). + +### 1.4 The vacuity traps — named BEFORE they are written + +Every one of these is a fixture that a wrong PR4 would pass today. + +**V1 — every existing plan begins with AND.** Verified by reading every +`LGJ_COMBINE_OR` site: `:2792` and `:2958` are both the *second* op. No test +anywhere has a leading OR. Old semantics are "acc starts as all rows set", so +a plan whose FIRST op is OR must return **all rows**, regardless of the +predicate. A lowering taking the natural shortcut — *"with an all-ones +accumulator the first AND is a no-op, so let `acc := pred_0`"* — and applying +it unconditionally is **wrong only for a leading OR**, and passes 100% of the +current suite. +> Falsifier **F-PR4-SEED**: for every opcode and operand, a 1-op plan with +> `combine = LGJ_COMBINE_OR` yields `LGJ_OK`, `out_count == n_rows`, and +> `dst_mask` all-ones with a clean tail. Asserted against the frozen oracle, +> not a hand-derived expectation. + +**V1b — leading-OR and the tail trap compose.** `n = 65`, single OR op: correct +answer `count == 65`. A whole-word all-ones seed without `clear_tail_bits` +gives `128`. This combination exists in no test today. Assert the count AND +`words.last() >> (n % 64) == 0`, so a failure names which defect fired. + +**V2 — the "then == els" analogue is an exhausted accumulator.** Once +`acc == 0` every subsequent AND is a no-op and ops `k+1..n` are unexercised; +once all-ones, every subsequent OR is. The existing monotonicity test asserts +non-degeneracy only at the END, and only for one plan. +> Anti-vacuity **A1**: for every plan designated non-degenerate, at **every +> prefix k**, `0 < count_k < n_rows` (guarded at `n_rows >= 3`). Degenerate +> plans get explicitly-named cases and are excluded by name, never by an +> unstated `if`. + +**V3 — the scratch-reuse trap, which mask-risc structurally cannot see.** +`reference.rs`'s own doc says it: *"a divergence no fixture built from +`Scratch::for_program` can express, because every such fixture starts +zeroed."* PR4's whole allocation win is exactly that reuse. Worse: a gated +predicate that SKIPS gated-off words instead of WRITING zero is +indistinguishable from a correct one under a fresh zeroed arena, and becomes a +stale-bit leak the moment the buffer is reused. [unverified: whether ndarray's +`pack_under` writes or skips is the open question — this falsifier answers it +rather than assuming.] +> Falsifier **F-PR4-REUSE**, through the ABI with no new API: (1) run a plan +> selecting everything so scratch ends dense; (2) run a selective multi-op plan +> into a DIFFERENT dst; (3) assert (2) is bit-identical to `legacy` for (2). +> Repeat inverted, with alternating `n_rows` and patterns. Plus idempotence. + +**V4 — `dst_mask` prior content and its dirty tail.** The most tempting +allocation saving is to make the destination's own words the accumulator. +Nothing today catches accumulating INTO dst rather than overwriting: the +monotonicity test reuses one mask and an AND-into-prior-content still produces +a shrinking chain — it passes. +> Falsifier **F-PR4-DST**: for each of `{fresh-empty, fresh-all, re-used, +> poisoned}`, the same plan yields identical status, `out_count` and words; and +> the published tail word is zero even when the prior tail was `u64::MAX`. +> Separate row: `dst_mask` must NOT be passed to `Planes::masks` — `validate` +> refuses a plane with a dirty tail (`ExecError::PlaneTail`), so a poisoned dst +> would turn a clean overwrite into a spurious error. `Planes::masks` is `&[]`. + +**V5 — arms no `LgjOpDesc` can reach.** The lowering emits only `Pred`, `And`, +`Or` (+ possibly one `Ternlog{0xFF}` seed). So `AndNot`, `Not`, the +non-commutative in-place remap, `remap_imm`'s 3-cycles, and 255 of 256 +immediates are unreachable from any admissible plan. **Do not claim PR4's +differential covers them** — that coverage lives in +`lance-graph-mask-risc/tests/differential.rs`, inherited as a dependency, not +as evidence. If PR4 seeds via `Ternlog{0xFF}`, that is `ternlog_self`'s +fill-ones arm — the arm the mask-risc council found vacuous — and F-PR4-SEED +becomes its second consumer and must say so. + +**V6 — the TCAM operand halves.** `(care << 32) | pattern`. A swap is invisible +whenever `pattern == care`, and the existing partial-care case uses +`tcam_operand(0b1000, 0b1000)` — exactly that. The `care=0`/`care=MAX` cases do +pin the halves, but **only at position 0**. +> Falsifier **F-PR4-TCAM**: `tcam_operand(5, 0b111)` at positions 0, +> 1-after-AND and 1-after-OR. `pattern != care` in every one. + +**V7 — opcode identity under the gate.** A mis-map is caught at position 0 +(with an all-ones accumulator a 1-op AND plan IS the predicate). It is NOT +caught if the mis-map exists only in the gated (`under: Some`) arm — a +plausible shape, since AND- and OR-combines take different routes. Hence the +opcode x position matrix. + +**V8 — the error path at a non-first position.** The existing test puts the bad +op second in all five cases, never first or last in a 4-op plan. Extend to +positions `{first, middle, last}` and `n_ops = 4`. An `ExecError -> LGJ_ERR_*` +map that collapses two codes passes a test that only checks "dst unchanged". + +**V9 — `out_count` not written on failure.** Under mask-risc the count arrives +as `Value::Count`; a `let mut count = 0; ... *out_count = count;` written +before the error check restores the bug. Keep the `12345` sentinel shape. + +## §2 — The allocation falsifier + +### 2.1 What is removed + +`:1718` + `:1724` = `2 * n_words * 8` bytes per call. At `n_rows = 65_536`, +`n_words = 1024` -> 16 KiB/call. **State the number with its row count +attached; it is not a constant, and it is arithmetic, not a measurement.** + +### 2.2 The counter, and where it lives + +**Rust side, not Java side.** `getThreadAllocatedBytes` — the counter this +repo's existing allocation gates use — measures **Java heap**. It cannot see a +Rust `vec!`. Citing the existing Java gates as evidence for C-ALLOC would be an +overclaim. Say so in the PR body. + +The instrument is the one `lance-graph-mask-risc/tests/no_alloc.rs` uses: a +pass-through `#[global_allocator]` over `System` incrementing an +`AtomicUsize`. Placement: a new integration test +`native/lgj-abi/tests/plan_eval_no_alloc.rs` — a separate binary, so the global +allocator does not perturb the in-crate suite. lgj-abi is `["cdylib","rlib"]`, +so the rlib links and the exports are callable directly. + +**Header constraint:** this binary contains **exactly one `#[test]`**. The +counter is process-global; a second concurrent test adds bytes and flakes. Do +not "fix" that by relaxing `== 0` to a tolerance — that is how the gate dies. + +### 2.3 BLOCKING design finding: unreachable with today's mask-risc `Scratch` + +`execute` requires `scratch.words == words_for(planes.n_rows)` **exactly**, +returning `ExecError::ScratchWords` otherwise, and `Scratch`'s only +constructors allocate. Consequences: + +- one cached `Scratch` sized for 65_536 rows **cannot serve a 64-row call**; +- a cache keyed by `words` allocates the first time each distinct `n_rows` is + seen — allocation becomes a function of the population's HISTORY, the exact + property C-ALLOC denies. A gate that warms up over a fixed sweep and then + measures the same sweep is **green while the property is false**. The + `n = 999` arm is what reddens it; +- relaxing the check to `>=` is **not** safe: `clear_tail` clears only the tail + WORD while ndarray's `mask_not` zeroes every word PAST it, and they agree + only on exactly-sized buffers. An oversized scratch reopens the tail law. + +**Clean resolution: a mask-risc co-change that lands FIRST** (same rule that +put D-MRX-0 in ndarray before PR3): a borrowing constructor, shape +`Scratch::over(buf: &mut [u64], slots, words)`, carving exactly +`words_for(n_rows)` per slot from one caller-owned growing buffer. Upstream +falsifier: pre-poison the buffer past `slots * words_for(n)` with `u64::MAX` +and assert every result and every `slot()` read is unchanged. + +**If PR4 ships without it**, narrow C-ALLOC in the PR text to "zero per-call +allocation for row counts already seen by this process", write the `n = 999` +arm anyway marked `#[ignore]`, and record the STOP in `ISSUES.md` — never +silently omit it, because an omitted arm reads as a passing property. + +### 2.4 The exact assertion + +One test, one binary. Handles opened and fixture generated BEFORE the baseline. + +``` +SWEEP = [0, 1, 63, 64, 65, 127, 128, 1000, 4097, 65_536] +UNSEEN = 999 // NOT in SWEEP, never warmed +PLAN = 3 ops, mixed AND/OR/AND, mixed opcodes and lane kinds + +// warm-up: one call at every n in SWEEP +let before = BYTES.load(Relaxed); +for _round in 0..64 { + for &n in SWEEP.iter() { plan_eval(pattern[n], PLAN, dst[n]) } // ascending + for &n in SWEEP.iter().rev() { plan_eval(pattern[n], PLAN, dst[n]) } // descending +} +assert_eq!(BYTES.load(Relaxed) - before, 0, "... per-n: {per_n:?}"); + +let b = BYTES.load(Relaxed); +plan_eval(pattern[UNSEEN], PLAN, dst[UNSEEN]); +assert_eq!(BYTES.load(Relaxed) - b, 0, + "a row count this process has not seen before must still allocate nothing"); + +// can-it-fire, without which the file proves only that the counter is broken +let a = BYTES.load(Relaxed); +let probe = vec![0u8; 4096]; +assert!(BYTES.load(Relaxed) - a >= probe.len()); +``` + +**Why this is population-size independent, precisely.** Three mechanisms, and +the failure message names which: (1) the sweep spans 0..65_536, so a gate that +holds only at one n cannot be green; (2) both directions in the same round — a +cache reallocating on GROW is caught ascending, on SHRINK descending; (3) an +unseen n — a per-n cache is green under (1) and (2) after warm-up and red only +here. + +**Not evidence:** RSS, `mallinfo`, wall-clock, or any Java-side counter. + +## §3 — The disable-run table + +Standing rule: a guard is not verified until its disable run is +**red-then-green**. Two traps designed around here: + +> **(a) Zeroing a constant is not a disable when the guarded quantity can reach +> the same outcome by another route.** Every row names WHICH test must go red, +> and the operator records the FULL set that went red — a disable reddening +> twenty tests localizes nothing. +> +> **(b) An edit whose anchor has moved silently no-ops, and a green run then +> reads as "the guard is inert."** Every disable is applied by a script that +> asserts `git status --porcelain` is EMPTY first, asserts the anchor matched +> **exactly once**, asserts the result differs from the input, runs, then +> restores with `git checkout `. Anchors are authored AFTER `cargo fmt` +> and a re-read. (Recorded failure: anchors written against pre-`fmt` text, three +> edits silently no-op'd, three guards reported GREEN having never landed. +> Recorded a second time on the PR3 arc: a pin added mid-pass was destroyed by +> the restore, because the restore reverts to the last COMMIT.) + +| # | Guard | Mutation that must turn it red | Must go red | Trap note | +|---|---|---|---|---| +| **D1** | F-PR4-SEED (leading OR => all rows) | make the first op's result the accumulator **unconditionally** | leading-OR arms only; all-AND stays green | (a): seeding to `0` reaches the same wrong answer by another route — that is D2. Record each separately | +| **D2** | the all-ones seed | seed to all-**zero** | all-AND multi-op arms + `fused_plan_equals_the_unfused_composition` | if this reddens the SAME set as D1, one guard is redundant — say so rather than counting both | +| **D3** | seed tail clear | delete `clear_tail_bits` on the seed | `n in {63,65,127,999,4097}` leading-OR arms | (a): running at `n % 64 == 0` is NOT a disable — the quantity is unreachable there | +| **D4** | final `clear_tail_bits` | delete it | F-PR4-DST tail assertion at `n = 65`, plus `out_count` there | as D3 | +| **D5** | F-PR4-DST (overwrite, not accumulate) | change publish from `copy_from_slice` to `\|=` | re-used and poisoned arms; NOT fresh-empty | this row proves the fresh-mask tests were never evidence for this property | +| **D6** | tail repair on publish | skip the tail word | the poisoned-dirty-tail arm | | +| **D7a** | differential independence, new side | in the NEW path only, map `LE_I32 -> Pred::LtI32` | the whole opcode x position matrix | (a): if done where both paths read it, nothing reddens — and the correct reading is "the oracle is not independent", not "the guard is inert" | +| **D7b** | differential independence, oracle side | in the FROZEN oracle only, swap AND and OR | the same matrix | both directions required | +| **D8** | F-PR4-TCAM | swap `pattern` and `care` in the unpack | the `pattern != care` arms at all three positions | (a): the existing `tcam_operand(0b1000,0b1000)` survives this — it is vacuous fixture V6 | +| **D9** | C-ALLOC gate, correct symbol | route the gate through `lgj_plan_eval_scalar` | Fork A: must go RED. Fork B: stays green, and THAT is the tell the gate does not distinguish the symbols — narrow the claim | the gate names its symbol in its own failure message | +| **D10** | C-ALLOC gate, sizing | put `let _ = vec![0u64; n_words];` back in the new path | `plan_eval_no_alloc` | (a): a `vec![0u8;1]` proves only the counter moves; `vec![_;0]` allocates NOTHING — run at `n_rows = 65_536`, never at `n = 0` | +| **D11** | C-ALLOC population independence | key the scratch cache on `words` | the `UNSEEN = 999` arm ONLY | if the sweep arms also redden, the cache was never shared and the gate measured something else | +| **D12** | C-ONE structural guard | add a second opcode-dispatch `match` in a production module | the structural guard | (a): if written `<= 1`, deleting BOTH sites also passes. Must be `== 1` with the allowed site NAMED, as `exactly_one_materialiser` names `reference_scratch`. Second disable: delete the one allowed site, assert red for "zero" too | +| **D13** | the `ExecError -> LGJ_ERR_*` map | delete one arm (fold two codes into one) | `..._rejects_a_lane_of_the_wrong_kind` + extended `a_bad_plan_leaves_dst_mask_untouched` | once per arm, not once for the map | +| **D14** | F-PR4-REUSE | stop zeroing the ping-pong slot between calls / route the gated predicate to SKIP rather than write | reuse arms; NOT fresh-call arms | if no code change makes this red, the gated predicate writes totally and V3 closes as **verified negative** — record that, it is a result | +| **D15** | `out_count` not written on failure | move the write above the error check | the `12345` sentinel | | + +**Deliberately excluded, with the reason stated:** A1, the `lt + ge == n` / +`eq + ne == n` partitions, and the fixture-fraction pins. These are **fixture +anti-vacuity checks, not code guards** — no production line's deletion reddens +them, and pretending otherwise inflates the table. They get a different check: +mutate the FIXTURE (move the `GT_I32` threshold to `-1000`) and assert A1 +fires. Run once, record once, exclude from the code-disable table. + +## §4 — The ordering, and where the commit goes + +**COMMIT, THEN DISABLE, THEN RESTORE.** The restore is `git checkout `, +which reverts to the last COMMIT; uncommitted work is destroyed by it. **No +disable run begins while `git status --porcelain` is non-empty**, and that check +is the first line of the script, not a habit. Anything added mid-pass restarts +the clock. + +| Commit | Content | Tree | Disable runs | +|---|---|---|---| +| **C0** *(sibling repo, if §2.3's co-change is taken)* | `lance-graph`: the borrowing `Scratch` constructor + poisoned-buffer falsifier. Merges FIRST | green | its own, after C0 is committed | +| **C1** | the frozen oracle + ALL behaviour falsifiers of §1 — **written against the OLD implementation, still in place** | **green** — the load-bearing property of the whole ordering | — | +| | *Gate at C1:* every new test green BEFORE any implementation change. A test red here has found a **pre-existing defect in the old loop**. Report it; do not fix it silently in the same PR — a silently fixed old bug changes what "equivalent to the old" means | | | +| **C2** | the lowering + `plan_eval_impl` routed through `execute`; the error map; the path dep; the scalar fork decision; **board hygiene in the same commit** | green — **C1's tests pass UNCHANGED.** Editing a C1 test in C2 is how an equivalence claim dies; if one must change, that is a finding | after C2: D1-D8, D13-D15 | +| **C3** | `tests/plan_eval_no_alloc.rs` + the C-ONE structural guard | green. Both were RED at C2^ and that run IS their red-then-green evidence — record the sha | after C3: D9-D12 | +| **C4** | only if §2.3's STOP is taken without C0: the `#[ignore]`d arm + the `ISSUES.md` entry | green | — | + +ABI surface: behaviour unchanged, therefore **`abi_minor` must not move**. Add +`lgj_abi_manifest().abi_minor == ` to the manifest test — a weak guard, +but it catches the reflexive bump. `docs/abi.md` needs no section. + +## §5 — What I would NOT test, and why + +**5.1 Which routing the AND-combines take.** Gated (`Pred{under: Some(acc)}`) +vs `Pred{under: None}` + `mask_and_assign` is **semantically identical** — this +is F-X1 from the mask-risc arc, inherited verbatim. Its instrument there was +`count_probe`'s sparse-gate pair (3572 ns vs 10729 ns at count 1741). PR4's +analogous claim needs the same: **a probe, never a threshold assertion.** +Measure-then-pin means a number in the board entry, not an `assert!` in CI. + +**5.2 Whether the code physically goes through `execute`.** No behavioural test +distinguishes "delegates" from "reimplemented identically inline". That is +C-ONE, and its only honest evidence is a **structural guard** in the family of +`the_crate_names_no_isa` / `exactly_one_materialiser`: exactly one production +opcode-dispatch site, allowed site NAMED, `== 1` never `<= 1`. + +**5.3 Where the scratch buffer lives.** Registry-owned vs thread-local vs +carved-from-dst is invisible to the allocation counter — all three read zero +once warm. The CONSEQUENCES are testable and are rows above: reuse-staleness +(V3/D14), per-n allocation (D11), dst accumulation (V4/D5). + +**5.4 Memory-footprint deltas.** If scratch moves onto the mask resource, two +masks over one pattern each carry a buffer — a real regression, but RSS is not +a gate. A probe printing the total across `mask_create`, recorded as a number. + +**5.5 The mask-risc arms PR4 cannot reach** (V5). Do not test them here, do not +claim them. Cite the sibling suite by path. + +**5.6 Backend parity across ISAs.** `execute` carries no ISA (law L2). CI builds +`x86-64-v3`; `-v4` is local-only and there is **no v4 CI arm**. Do not write +"verified on all backends." + +## Open questions PR4 must close before writing code + +- **OQ-1 (blocking).** §2.3: the borrowing `Scratch` co-change (C0), or the + narrowed C-ALLOC claim + `ISSUES.md` entry. Picking neither means the + allocation gate is green for the wrong reason. +- **OQ-2.** §1.2: Fork A or Fork B for `lgj_plan_eval_scalar`. +- **OQ-3.** Does ndarray's `pack_under` WRITE zero into gated-off words or SKIP + them? [unverified.] V3/D14 answers it, and the answer decides whether scratch + reuse is safe at all. +- **OQ-4.** The G11 fence covers `lance_graph_contract::` module names only. A + new crate dependency is outside it — see `exec-runs/pr4-abi-membrane.md` Q1 + and `exec-runs/pr4-kernel-membrane.md` (c), which both require the fence to + grow a crate allowlist **in the same commit as the dep**. +- **OQ-5 (from the ABI ruling, not this pass).** `mask_risc::MaskOp` has **no + Fill/Const op**, so the all-set accumulator `V0 = all` that `abi.md` §7 + mandates has no representation. See `exec-runs/pr4-abi-membrane.md` — this is + a second substrate-first blocker, and V1/F-PR4-SEED is the falsifier for the + rewrite that would paper over it. + +--- + +## OQ-5 — RESOLVED (2026-09-14, main thread). No `Fill` op is needed. + +The all-set accumulator is not a missing IR primitive. It is a statement about +a **prefix of the op list**, decidable at lowering time. + +### The algebra + +`abi.md` §7 is `acc_0 = ALL`, then `acc_i = acc_{i-1} ⊕_{combine_i} pred_i`. +Two facts settle it: + +- `ALL & p = p` — an AND against the all-set accumulator **is** the predicate. +- `ALL | p = ALL` — an OR against it is a **no-op that discards `pred_i`**. + +So let `k` = the least index with `ops[k].combine == AND` (`k = n` if none). +Every op before `k` ORs into `ALL` and leaves it `ALL`; op `k` reduces to +`acc = pred_k`. Therefore: + +- **`k < n`** — drop ops `0..k`; lower op `k` as + `MaskOp::Pred { under: None, dst: 0 }` (a direct write, no seed); lower + `k+1..n` as their own combine into slot 0. Exact. +- **`k = n`** (every combine is OR) — the result is `ALL`, tail-cleared, and + `count = n_rows`. A constant; no program runs. + +### Why dropping ops `0..k` cannot change ERROR behaviour + +This is the half that had to be checked rather than assumed, and it holds by +construction rather than by test. After `validate_plan` returns `Ok`, +`kernels::eval_predicate` is **infallible** for every op in the plan: its only +two error arms are `LGJ_ERR_LANE_KIND_MISMATCH` and `LGJ_ERR_UNKNOWN_OPCODE` +(`kernels.rs:928-938`), and `validate_plan` already rejects both, for every +op, up front (`exports.rs:1662-1672` — `opcode_required_kind` and the +`lane.kind() != required` test). The kernels below it return `()`, never a +`Result`. The in-loop `lane_view` is the same argument: `validate_plan` called +it for the same `lane_id` already. + +So a skipped op has no observable effect of any kind. That is what makes the +prefix rewrite a **rewrite** rather than a behaviour change. + +### The trap the rewrite must NOT fall into + +`k == 0` for every plan the Java facade can currently build — `View` exposes +no `or`, `PlanOp::narrowing` is documented as "the only kind reachable from +`View.where`", and `COMBINE_OR` is "deliberately unreachable from here" +(`View.java:23-32`, `PlanOp.java:22-24`). So the simple rewrite — *op 0 always +writes directly* — passes every test that exists, and is **wrong for a C +caller**, which `validate_plan` explicitly permits to send `LGJ_COMBINE_OR` in +any position. Implement the general prefix rule; do not special-case `k = 0`. + +F-PR4-SEED must therefore carry a `combine = OR` leading op built at the ABI +level, not through the Java facade, which cannot express one. + +### An ABI observation, recorded not fixed + +A plan whose ops are **all** `OR` returns every row, whatever the predicates +say. That is the §7 contract read literally, it is reachable from C today, and +it is not PR4's to change — PR4 owes behaviour equivalence, not correction. +Filed here so the lowering's `k = n` arm reads as deliberate rather than as a +degenerate case someone should "clean up" later. + +### Consequence for the OQ list + +OQ-5 no longer blocks: it needs no substrate change, no new `MaskOp` variant, +and no ndarray primitive. It becomes a lowering requirement plus one falsifier. +**OQ-1 remains the blocking one** — and note its shape is narrower than the +plan states: `n_rows` is a property of the *pattern resource* and immutable for +its lifetime, so there is exactly ONE row count per pattern and no +"first sight of each n" cache is required at all. What that leaves open is +purely **where a per-pattern scratch lives and how concurrent +`lgj_plan_eval` calls on one handle share it** — a different and smaller +question than the one recorded above. + +--- + +## OQ-1 — the obvious home is DISQUALIFIED, and one of the two buffers can just go (2026-09-14, main thread) + +Not yet closed, but narrowed twice and with the tempting answer ruled out. + +### Per-pattern scratch would serialize a path that is parallel today + +`n_rows` is immutable per pattern resource, so "cache a scratch per row +count" collapses to "cache a scratch per pattern" — one buffer, no keying, +no first-sight-of-each-n problem. That is why the plan's framing was too +wide. But the narrower version is **wrong for a different reason**, and the +registry says so in its own words. + +`registry.rs:22-34`: a call takes a **short read lock**, clones the +`Arc`, and **drops the registry lock before touching the +payload**; masks then carry a per-resource `RwLock` specifically "so bulk ops +on distinct masks do not serialize". A pattern's payload is a fixture whose +buffer is "immutable for the resource's whole life" (`:126`) — it takes no +lock at all, which is exactly what lets N threads evaluate N plans against +one pattern with zero contention. + +Hanging a mutable scratch off `ResourceEntry` ends that. It would need its +own `RwLock`, and two concurrent `lgj_plan_eval` calls on the same pattern — +the common shape, since a pattern is the thing you query repeatedly — would +then block on each other. **Trading a per-call allocation for a per-call +lock on the hot path is not an optimisation**, and it converts the one +resource kind that is deliberately lock-free into a contended one. + +So per-pattern scratch is out. Remaining candidates, in preference order: +a thread-local keyed by word count (no contention, allocates once per +(thread, size), and the number of live sizes is the number of live +patterns); or a scratch RESOURCE with its own handle, which is ABI-clean in +principle — a handle is a name, not a byte position — but costs a new symbol +and owes the membrane warden a reason. + +### `acc` needs no home at all — it can be deleted + +`plan_eval_impl` allocates TWO buffers. The second one, `acc`, exists for a +stated reason (`exports.rs`): "dst_mask is written exactly once, and an error +at any point leaves it byte-for-byte as it was." + +That reason is discharged by the OQ-5 finding above. After `validate_plan` +returns `Ok`, **there is no error path left**: `eval_predicate` is infallible +for every op in the plan, and the kernels beneath it return `()`. The +in-loop `lane_view` re-check is the same argument. So "an error at any +point" describes a set of points that is empty, and the copy-then-publish +dance is protecting against nothing. + +With validation total and up front, the loop can accumulate directly into +the `dst_mask` words under their existing write guard. One allocation +disappears with no new storage, no new lock, and no ABI change — and it is +the one that is exactly `n_words` of the caller's own destination. + +That leaves a single buffer to home, which is a materially smaller question +than the one this OQ started as. **Falsifier owed:** a test that an error +returned by `plan_eval` leaves `dst_mask` unchanged must still pass — the +errors that remain (`EMPTY_PLAN`, `NULL_ARGUMENT`, the resolve failures, +`validate_plan`'s own rejections, `MASK_LENGTH_MISMATCH`) all fire before +the first write, and that ordering is now load-bearing rather than +incidental. It should be pinned as such, not assumed. + +--- + +## OQ-2 — RESOLVED: Fork A. The mask-risc oracle is MORE independent, not less (2026-09-14, main thread) + +The question read as a trade — keep `kernels::scalar_*` (Fork B, two +evaluators, C-ONE narrowed to "one SIMD evaluator") or route +`lgj_plan_eval_scalar` at `reference_execute` (Fork A, C-ONE stays strong but +the allocation gate must be scoped). It is not a trade, because both sides +were being judged against the wrong property. + +`kernels.rs` states the property in its own words: *"`scalar_*` below is +written in plain Rust loops with **no ndarray at all**. That independence is +the entire value of `lgj_plan_eval_scalar`: if the reference shared code with +the SIMD path, a parity test between them would be checking that a function +agrees with itself."* + +Measured against that property, `reference.rs` wins outright: *"the same +Program evaluated ONE ROW AT A TIME in plain Rust, with no SIMD facade +anywhere in this file (law L4: a test greps the source)"* — and L4 is +**enforced by a grep test**, where `kernels::scalar_*`'s independence rests on +the discipline of living in the same file as the SIMD wrappers and not calling +them. Fork A therefore **strengthens** the exact property the symbol exists +for, and it upgrades the comparison from kernel-vs-kernel to +whole-evaluator-vs-whole-evaluator. + +**Decision: Fork A.** `lgj_plan_eval_scalar` lowers the same plan and runs +`reference_execute`. + +### The two costs, stated rather than absorbed + +**1. `reference_scratch` allocates by design** — it materialises one `bool` +per row per slot, which is its whole reason for existing (an oracle sharing +the executor's bit packing could not falsify a bit-packing bug). So the +C-ALLOC gate is scoped to `lgj_plan_eval` **only**, and must name its symbol +in its own failure message. That is row **D9**, and under Fork A D9 must go +RED — a gate accidentally written against the scalar symbol is green for the +wrong reason. + +**2. `simd_and_scalar_plans_agree_bit_for_bit` changes meaning and must be +renamed.** It stops being a SIMD-vs-scalar backend-parity test and becomes +executor-vs-oracle. That is not a loss: per §5.6 and the `lib.rs` correction +in lance-graph #1230, **backend parity is `ndarray::simd`'s business**, and an +ABI-level test could only reach it through a proxy. Renaming it is mandatory — +a test whose name claims backend parity it no longer checks is worse than no +test, because a future session reads the name. + +### The consequence that makes C1 permanent + +Both symbols now share the **lowering**. A lowering bug is therefore shared, +and no ABI-level differential between the two can see it — which is precisely +why the frozen oracle exists. **`legacy_plan_eval_impl` is not PR4 scaffolding +to be deleted after the merge; it is the only remaining independent check on +the lowering, and it stays.** Its cost is one frozen function that nothing +else may call. + +--- + +## OQ-4 — the G11 fence must grow a CRATE allowlist in the same commit as the dep + +The fence today (`native/lgj-abi/tests/g11_contract_import_fence.rs`) walks +`src/` and rejects any `lance_graph_contract::` module outside four named +ones. It says nothing about which CRATES `lgj-abi` may depend on, so adding +`lance-graph-mask-risc` passes it silently — and the fence's own history is +that it *was prose for months while already false*. + +PR4 adds the crate allowlist in the same commit as the dep, spelled once in +the test and once in `CLAUDE.md`, with the same red-then-green discipline the +module allowlist got. + +--- + +## OQ-1 — CLOSED: a thread-local growing arena + +The blocker is gone: `Scratch::over(buf, words, slots)` and +`scratch_words_for` landed in lance-graph #1230 and are on `main`. What +remained was only *where the buffer lives*, with per-pattern already +disqualified (it would give the one deliberately lock-free resource kind a +`RwLock` and serialise concurrent `lgj_plan_eval` calls on one pattern — the +common shape). + +**Decision: a thread-local `RefCell>`, grown monotonically to +`scratch_words_for(words_for(n_rows), SLOTS)`.** + +- No contention by construction — it is not shared, so there is no lock to + take and nothing for two threads to serialise on. +- Allocation becomes a function of the **maximum** row count a thread has + seen, not of the history: a smaller call carves a strict prefix of the same + buffer. That is what makes the `UNSEEN = 999` arm (D11) pass rather than + being warmed into passing. +- It allocates once per `(thread, high-water mark)`. The honest bound to state + in the PR body: a thread that has evaluated one 65,536-row plan holds + `SLOTS * 1024 * 8` bytes until it exits. That is a real cost, and it is + bounded, and it is not per-call. + +`acc` is deleted outright, per the finding above: after `validate_plan` +returns `Ok` there is no error path left, so the copy-then-publish dance +protects against an empty set of points. Slot 0 of the scratch is the +accumulator, and one `copy_from_slice` publishes it under the existing write +guard — the same single write the old loop did. + +**Falsifier owed and not optional:** the ordering that makes this safe — +every remaining error fires BEFORE the first write — is now load-bearing +rather than incidental. T4/T5 of `pr4_dst_reuse` pin it at every position in +a plan, with a pre-poisoned destination so "unchanged" is an observation and +not "still zero". + +### The lowering, stated so the implementation has nothing to invent + +Per OQ-5, `k` = the least index whose combine is AND (`k = n` if none). + +- **`k = n`** (every combine is OR): the answer is ALL rows, tail cleared, + `count = n_rows`. A constant — no program is built and none runs. +- **`k < n`**: drop ops `0..k`; lower op `k` as + `MaskOp::Pred { under: None, dst: 0 }`; lower each later op `i` as + `Pred { under: .., dst: 1 }` followed by `And`/`Or { a: 0, b: 1, dst: 0 }`. +- Terminal: `Count { mask: Scratch(0) }`; the words are read back out of + `scratch.slot(0)`. +- `SLOTS = 2`. + +Opcode map, one arm each, exhaustive over the nine: +`EQ_U32 -> EqU32`, `GT_I32 -> GtI32`, `NE_U32 -> NeU32`, `EQ_I32 -> EqI32`, +`NE_I32 -> NeI32`, `LT_I32 -> LtI32`, `LE_I32 -> LeI32`, `GE_I32 -> GeI32`, +`TERNARY_MATCH_U32 -> MatchU32 { pattern, care }` unpacked from +`(care << 32) | pattern`. + +--- + +## C2 disable run — RESULTS (2026-09-14, commit `93da288`) + +Eleven arms run against the NEW implementation, each patched, tested and +restored by a script that asserts its anchor matched exactly once first (the +recorded failure this discipline exists for: anchors written against pre-`fmt` +text, three edits silently no-op'd, three guards reported GREEN having never +landed). + +**Nine load-bearing.** D1, D2, D4, D6, D7a, D7b, D8, plus two arms not in the +table above that cover the two decisions C2 actually had to make: + +- **D-GATE** — gate an OR-combined op `under` the accumulator (i.e. delete the + `is_and` condition). RED, including `or_plans_widen` and the whole combine + sweep. This is the asymmetry the lowering's doc calls the correctness + question, and it is now measured rather than argued. +- **D-LANE** — shift the lane index by one inside `pred_of`. RED across 17 + tests. The `const _` lane-id assertions guard the OTHER direction (the + fixture renumbering under a correct lowering), which no runtime test can + reach; this arm covers the lowering getting it wrong directly. + +**D3 has no guard left to disable, and that is a finding rather than a gap.** +The row reads "seed tail clear: delete `clear_tail_bits` on the seed". There +is no seed. The prefix rewrite means op `k` writes the accumulator directly, so +the all-ones seed the old loop built — and had to tail-clear — does not exist +in the new path. The one remaining `clear_tail_bits` is the `AllRows` arm, and +that is D4, which is RED. + +**Two came back VACUOUS, and both are exactly what the table predicted.** + +| arm | result | the table's own trap note | +|---|---|---| +| **D5** — publish accumulates (`\|=`) instead of overwriting | GREEN | *"this row proves the fresh-mask tests were never evidence for this property"* | +| **D15** — `out_count` written above the error check | GREEN | the `12345` sentinel arm | + +Neither is a weak guard: every test that exists today publishes into a FRESHLY +CREATED empty mask, and `|=` into an all-zero destination is `copy_from_slice`. +Nothing in the suite has a non-empty prior to be polluted, and nothing triggers +an error and then reads the sentinel. The arms that would redden both live in +`pr4_dst_reuse.rs` (F-PR4-DST's re-used/poisoned destinations; the +error-at-every-position sweep with a pre-poisoned `out_count`), which is why +that file is C1's last owed piece rather than a nice-to-have. + +So the run did the thing a disable table is for: it converted "we think these +two properties are untested" from a prediction into a measurement, BEFORE the +file that fixes it landed. D5 and D15 are re-run once it does, and a green +result there would then be the real defect. + +### D5 and D15 re-run once `pr4_dst_reuse.rs` landed — the receipt + +The two arms recorded VACUOUS above were re-run against the same C2 +implementation with the falsifier file in place. Both are now covered, and the +second one is a correction to the disable, not to the code. + +**D5 — RED.** Changing `publish` from `copy_from_slice` to `|=` now reddens +`the_same_plan_lands_identically_whatever_the_destination_held`, +`a_poisoned_tail_is_published_clean` and +`a_dirty_destination_is_not_read_as_an_input_plane`. Red-then-green complete, +and the prediction the table made in advance — *"this row proves the fresh-mask +tests were never evidence for this property"* — is now a measurement at both +ends. + +**D15 — still GREEN as the table spells it, and that is a defect in the +DISABLE.** The row says "move the write above the error check". Under the old +loop the write and the error checks shared one function. Under C2 they do not: +`out_count` is written in exactly ONE place (`publish`), and `publish` is +reached only after every error has already returned — so moving the write to +the top of `publish` cannot break anything, because on a failing plan `publish` +is never called at all. + +The error check that is actually reachable is `validate_plan`, one level up. +Re-targeted there, both directions redden: + +| arm | mutation | result | +|---|---|---| +| **D15a** | write `*out_count` before `validate_plan` | **RED** — `a_bad_plan_leaves_dst_mask_untouched`, `an_error_at_any_position_in_a_plan_leaves_the_destination_untouched` | +| **D15b** | zero the destination before `validate_plan` | **RED** — the two above plus `every_minor_11_opcode_rejects_a_lane_of_the_wrong_kind` | + +So the property holds and is load-bearing; what changed is WHERE its guard +lives. The lesson is the one this table's own header already carries in a +different form: a disable written against the old structure can pass against +the new one for a reason that has nothing to do with the guard. **The right +reading of a green disable is "find out why", never "the guard is inert".** + +Both `publish`-internal error paths (`write_mask()` returning `None`, and the +length mismatch) are unreachable for the same reason the `ExecError` map's arms +are: `resolve_pattern_and_mask` rejects a wrong-kind handle and a +wrong-length mask before the plan runs. They stay as defensive returns and are +not claimed to be falsifiable. diff --git a/CLAUDE.md b/CLAUDE.md index fa2dc48..3c24bea 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -230,9 +230,24 @@ build): not row count — verified fixed-size on both the Rust and Java sides); `Abi.java`'s `readCarvings` (bounded by `CARVING_SLOTS`, a manifest constant, not n_rows); `Engine.facetSumResolved`'s fixed `long[2]` - result pair. None of the five is a hidden proportional-to-n_rows - population copy — keep this list exhaustive when a sixth site is added, - rather than letting the enumeration silently go stale again. + result pair; `View.where()`'s `List.copyOf` of the PREDICATE chain; and + `NativePattern.plan()`'s `predicates.stream()…toList()`. None of the + seven is a hidden proportional-to-n_rows population copy — keep this list + exhaustive when an eighth site is added, rather than letting the + enumeration silently go stale again. + + > **⊘ RE-AUDITED 2026-09-16 and the list WAS stale — it claimed five and + > the tree had seven.** `View.where()` and `NativePattern.plan()` were + > never listed. Both are legitimate, and *why* is the sharper statement of + > the rule: they are bounded by the number of predicates the developer + > CHAINED, i.e. by the size of the query they typed — never by the data. + > So the invariant is not "no allocation on the Java side"; it is + > **nothing proportional to ROWS**, and a query-shaped allocation is the + > boring front's own size, not the substrate's. The audit that found this + > is one grep (`long[]`, `toArray`, `copyOf`, `.stream()` across + > `java/src/main`) and it is the check to re-run before citing the list — + > the enumeration had already gone stale once under a sentence telling the + > next session not to let it. Temporary kernel scratch (SIMD scratch masks, decode buffers) is allowed and is NOT the same claim as a second canonical copy. - **Layout parity is independently derived, not self-compared.** @@ -274,6 +289,204 @@ simd_{amx,avx512,avx2, Rust: lgj-abi kernels → ndarray::simd (backends) ndarray's top) ``` +> **⊘ THE MIDDLE ROW IS WRONG, AND IT MISLEADS (operator-corrected +> 2026-09-14).** `Valhalla + Panama` is NOT this side's analog of ndarray's +> `cfg dispatch`. There is **no analog**, because there is nothing to +> dispatch: `ndarray` IS the SIMD polyfill, there is exactly one +> implementation of every word, and it is in Rust. Reading the row as an +> analogy invites treating Panama as a dispatch or compute layer — which is +> precisely the confabulation it produced (a "T0 owns backend realization" +> tier story invented to justify a conclusion that needed no tiers; see +> lance-graph `TECH_DEBT.md`'s ternlog storno). +> +> Valhalla and Panama are **two ORTHOGONAL guarantees**, not a pair and not +> a layer: +> +> - **Panama — computation never lives in Java.** The crossing mechanism. +> Java hands the question across and receives the projection; the +> decomposition of an answer never crosses. This is what E1 enforces. +> - **Valhalla — storage never lives in Java.** The orthogonal axis. Value +> classes carry SHAPE without identity or heap storage, so the Java side +> holds names, handles and addresses — never the bytes. This is what the +> one-copy law, the `materialize*` naming rule, and +> `java-surface-warden`'s "no Stream-over-hydrated-elements" all enforce +> from different directions. +> +> **What they are FOR, stated positively** (operator, same day): *Java is +> the low-code thin surface over zero-copy, with methods that look so +> natural and still compute in lance-graph — Java just thinks it's on +> steroids without knowing why.* +> +> Sharpened 2026-09-16, and this is the sentence to keep: *"java is the low +> code intake GLOVE around the lance-graph spine — lance-graph-java just +> happens to offer the MENU TO THE TABLE in a pleasing way, using masking ops, +> offering 5 star for the price of a blink."* The menu is the product. The +> kitchen is elsewhere and the diner never sees it; what makes the menu honest +> rather than a disguise is that the work and the data genuinely never cross. +> The whole allocation, same ruling: +> +> | what | lives in | membrane that keeps Java out of it | +> |---|---|---| +> | **thinking** | lance-graph | **Panama** | +> | **SIMD** | ndarray | the `ndarray::simd` facade | +> | **storage** | lance-graph | **Valhalla** | +> +> Note both membranes name the SAME home for two different things — thinking +> and storage are both lance-graph's, and Panama and Valhalla are two +> orthogonal ways of keeping Java out of them. That is why the middle row of +> the table above is wrong: they are not one layer. The fluency of `view.where(..).hop(..) +> .count()` is real; the work and the data are both elsewhere; and Java is +> never told. Zero-copy is what makes the naturalness honest rather than a +> disguise — there is no hidden hydration behind the nice method name. +> +> Corollary for any design that reaches for this table: the isomorphism +> holds at the TOP row (facade ↔ facade) and the BOTTOM row (backends ↔ the +> Rust floor). It BREAKS in the middle. Do not reason from the middle row. + +## THE JAVA SURFACE IS `sql()`, NOT THE MASK ALGEBRA (operator-ruled, 2026-09-16) + +**Verbatim, in four parts:** + +> *"Java doesnt use masking ops. `Mask.minus()`, `RowStore.hop()`. Lance-graph +> does. Java just sees boring `sql()` handed to duckdb (Example)."* +> +> *"nobody should ever start trying to optimize Java (except making it boring +> front)."* +> +> *"the boringness is then handed to lancegraph as zero copy."* +> +> *"and handled akin to ndarray polyfill — java doesnt know why there is +> `sql()` polyfill, we just make sure there is."* + +### What this corrects, by name + +⊘ The sentence three paragraphs up — *"offers the menu to the table in a +pleasing way, **using masking ops**"* — is **struck on those three words**. +The menu framing survives; the mechanism named in it does not. lance-graph +uses masking ops. Java does not, must not, and is never told they exist. + +⊘ The isomorphism table's top row previously read `Java (View / Mask / +RowStore / consumers)` as the facade. **`Mask` and `RowStore` are NOT the +facade** — they are the substrate's algebra standing on the Java side of the +wall. `Mask.minus()` and `RowStore.hop()` are the operator's own two examples +of the wrong shape. The facade is the boring call a Java developer already +knows how to write. + +### The polyfill relation is EXACT, and it is the whole design + +This is the part that makes the ruling operational rather than stylistic. +`ndarray::simd` is a facade with **37 functions and zero shipping +instructions** — a consumer crate calls `U8x64::cmpeq_mask` and has no idea +whether AVX-512, NEON, wasm or scalar answered it. The consumer does not +know why the polyfill exists. It only knows it is there. + +`sql()` is that, one tier up: + +| tier | the caller writes | the caller does not know | +|---|---|---| +| consumer crate → ndarray | `U8x64::cmpeq_mask(..)` | which of six backends ran | +| Java → lance-graph | `sql("select …")` | that masks, ternlog, hops or popcounts exist at all | + +**"We just make sure there is one"** is the standing obligation, and it falls +on the Rust side, never on Java. A missing `sql()` capability is a +lance-graph/ABI gap to close — exactly as a missing SIMD primitive is an +`ndarray::simd` gap to close, never a licence for a consumer to write +intrinsics. This is the MISSING-CAPABILITY STOP RULE already in this file, +now with its Java-side spelling. + +### The three operational consequences + +1. **Nobody optimizes Java. Ever.** The only sanctioned work on the Java side + is making it **more boring** — more ordinary, more familiar, closer to what + a Java developer would have written without us. A PR whose stated goal is a + faster Java path is rejected on its goal, before its diff is read. "Boring" + is the performance strategy: the speed is elsewhere. +2. **The boringness is handed down as zero copy.** The ordinary-looking call + is not translated, re-encoded, or marshalled. It crosses as a NAME and the + bytes stay put — which is what makes the boring surface honest rather than + a disguise. Zero-copy is not an optimization applied to the glove; it is + the condition under which a boring glove is allowed to exist. +3. **A mask concept on a public Java signature is a defect**, regardless of + how well it works. `java-surface-warden` already blocks byte positions and + arithmetic; this ruling adds the population algebra by name — `Mask`, + `minus`, `hop`, `ternlog`, `popcount`, lane ids, opcodes. Those are + T1/T2 words (see `kernel-membrane-warden`), and T3 does not speak them. + +4. **Java never does materialization — and the row count is what proves it** + (operator, same day): *"Java never does materialization. A billion rows op + is handed behind the front and handled as a masking hattrick nobody knew + what's coming."* This is the boring front's load-bearing property, not a + performance note. A billion-row operation crosses as a NAME and returns a + scalar; the billion never exists on the Java side, in any form, at any + moment. That is why the surface can afford to be boring: `sql()` looks + identical at 10 rows and at 10^9, because Java's work is the same at both — + none. + + **The checkable form of "never":** every public Java call must be + **O(1) in rows**, and the ONLY exit is a method whose name begins with + `materialize` (today: exactly one, `Mask.materializeRows()`, O(n), stated + in its javadoc). The naming rule is not a loophole in "never" — it is the + mechanism that makes "never" auditable, because anything proportional to + row count that is NOT named that way is a defect by inspection, with no + judgement call required. Keep the exhaustive materialization-site list in + the zero-copy section current for exactly this reason: an unnamed sixth + site is the failure this rule exists to catch. + + **The corollary that bites hardest in review:** a Java-side loop is not + merely slow, it is *proof that something was materialized to loop over*. + E1 forbids the loop; this forbids the thing that made a loop expressible. + Two statements of one rule from opposite ends. + +### THE ENDGAME — why "boring" is strategy, not taste (operator, same day) + +> *"endgame is make Java low code **'Bring your own software'** to beat +> palantir foundry at its own game."* + +Foundry's proposition is *bring your data to our platform* — and then learn +its ontology tooling, its pipeline builder, its workshop, its idioms. The +learning curve is not a cost of the product; **it IS the product**, because +every hour a customer spends learning it is an hour of lock-in that a +competitor must refund before they can switch. + +BYOS inverts exactly that: **bring your own software.** The customer keeps +their Java, their SQL, their JDBC-shaped habits, their existing build — and +the substrate is underneath without them learning a new vocabulary to reach +it. Nothing to port, nothing to adopt, nothing to unlearn if they leave. + +**This is what makes each of the three consequences above non-negotiable +rather than stylistic:** + +- **Novel API is the enemy, not slow API.** `Mask.minus()` is something a + customer must LEARN; `sql()` is something they already know. Every unit of + novelty on the Java surface is a unit of Foundry-shaped lock-in-by-curve — + built by us, against our own pitch. That is why optimizing Java loses the + game even when it succeeds: optimization on that side produces novel API, + and novel API is precisely the thing BYOS promises not to require. +- **"Boring" is therefore the competitive moat, measured as absence.** The + win condition is that a Java developer writes what they would have written + anyway and the result arrives at substrate speed. There is no demo of this + that looks impressive — it looks like nothing happened, which is the point. +- **The advantage must live where the customer does not have to look.** Speed + comes from lance-graph and ndarray; the customer gets it by writing ordinary + code. A surface that has to be studied to be fast has already conceded the + argument, however fast it then is. + +**Practical test for any proposed Java-side addition:** *would a customer who +has never read our documentation write this line by accident, from habit +alone?* If yes, it is a candidate. If it needs a paragraph of explanation, it +is Foundry's business model with our name on it — reject it and close the gap +on the Rust side instead. + +### Scope — what this does NOT do + +It does not delete `Mask` or `RowStore` today. They exist, they are tested, +and the four leak-on-throw fixes that landed with this ruling make existing +code correct rather than expanding it. What the ruling settles is the +DIRECTION: the boring `sql()`-shaped surface is the target, the mask algebra +is not to be grown on the Java side, and no future session may cite the +struck sentence above as licence to add another `Mask.*` verb. + + Measured grounding (2026-08-27, in-tree): `simd.rs` is 37 functions and ZERO shipping instructions — every raw intrinsic in it sits inside `#[cfg(test)]` as the wrapper's oracle; `simd_avx512.rs` alone carries 488. @@ -320,6 +533,126 @@ the signal the substrate is missing a word. `internal.*` in any public signature — the exact analog of "all SIMD from `ndarray::simd`, never `simd_{arch}`, never raw intrinsics". +## Never skip the Java half — check FOUR places before saying "no JDK" (2026-09-16) + +**Operator ruling: document this so no future session runs without Panama at +all.** This repo's history contains commits titled *"UNVERIFIED: JDK 26 not +available in this sandbox"*, and the Java half of an ABI-minor-11 change went +to PR on that basis. + +> ⊘ **The first version of this section, written earlier the same day, led with +> the apt route and was itself misleading.** Re-checked: **`/opt/jdks/jdk-26.0.2` +> was present the whole time** (26.0.2.1), the suite runs `ALL PASSED (409 +> checks)` on it, and `.claude/knowledge/jdk-toolchain-facts.md` had already +> named that exact path. Nothing needed installing. The session looked at +> `java -version` and `/usr/lib/jvm` — **neither of which sees `/opt/jdks`** — +> and inferred absence. So the apt routes below are a RECOVERY path, not the +> answer; the answer is `ls /opt/jdks`. + +**The rungs, in order. Report "no JDK" only after all four fail.** + +```sh +ls -d /opt/jdks/*/ && for d in /opt/jdks/*/; do "$d/bin/java" -version; done # 1 +ls /usr/lib/jvm/ # 2 +``` + +Rung 1 is the production path this repo pins, and it is the one that was +missed. `which java` answers "what is on PATH", never "what is installed" — it +is not evidence about the second question. Full ladder with the +verified-by-execution table: `.claude/knowledge/jdk-toolchain-facts.md` +§ ACQUISITION LADDER; the guard is `.claude/agents/jdk-toolchain-warden.md`. + +Rungs 3 and 4 are apt, and matter when `/opt/jdks` is absent — which it may +well be in a rebuilt container, since **`/opt/jdks/jdk-27` (the JEP 401 +Valhalla EA build) is already GONE from this one.** + +**JDK 26 — the version this repo's own commits targeted — needs ONE extra apt +source, and Ubuntu's stock archive alone will never offer it.** Add Adoptium +and it is an ordinary `apt-get install`: + +```sh +# the extra source — Ubuntu noble stops at openjdk-25; Adoptium carries 8..26 +curl -fsSL https://packages.adoptium.net/artifactory/api/gpg/key/public \ + | sudo tee /etc/apt/keyrings/adoptium.asc >/dev/null +. /etc/os-release && echo "deb [signed-by=/etc/apt/keyrings/adoptium.asc] \ +https://packages.adoptium.net/artifactory/deb $VERSION_CODENAME main" \ + | sudo tee /etc/apt/sources.list.d/adoptium.list +sudo apt-get update +sudo apt-get install -y temurin-26-jdk # -> /usr/lib/jvm/temurin-26-jdk-amd64 +``` + +Ubuntu's own archive serves **openjdk-25** (also fine — Panama is final since +22, and the suite passes on it too): + +```sh +sudo apt-get update # NOT optional — see the trap below +sudo apt-get install -y openjdk-25-jdk-headless +/usr/lib/jvm/java-25-openjdk-amd64/bin/java -version # openjdk 25.0.4 +``` + +**The trap that made this look impossible.** The pre-seeded apt index pointed +at `openjdk-25 25.0.2+10-1~24.04`, a version already withdrawn from the pool, +so the install died on a bare `404 Not Found` — which reads as *"this package +does not exist"* rather than *"your index is stale"*. `apt-get update` moved +the candidate to `25.0.4+7-1~24.04` and the install succeeded immediately. +Same shape as this workspace's other freshness traps: **a cached view reporting +absence is not evidence of absence.** + +### Running the Java suite — there is no build tool, and that is deliberate + +`java/src/test/.../AllTests.java` is a plain `main` that runs every suite in +one JVM and exits 0 / 1 / 2 (passed / failed / **native artifact absent, so +nothing ran** — a missing library reported as a failure would send the reader +hunting a bug that is not there). So: build the `.so`, `javac`, `java`. + +```sh +# 1. the native artifact +cd native/lgj-abi && cargo build --release # -> target/release/liblgj_abi.so + +# 2. compile main + test together (56 files) +J=/usr/lib/jvm/temurin-26-jdk-amd64/bin # or java-25-openjdk-amd64 +find java/src/main java/src/test -name '*.java' > /tmp/srcs.txt +$J/javac -d /tmp/jout @/tmp/srcs.txt + +# 3. run. `-Dlgj.library` is an EXPLICIT request and `Abi.locateLibrary` +# refuses to fall back to a search path if it cannot be honoured — by +# design, so a run can never silently measure a different artifact. +$J/java --enable-native-access=ALL-UNNAMED \ + -Dlgj.library=$PWD/target/release/liblgj_abi.so \ + -cp /tmp/jout com.adaworldapi.lancegraph.AllTests +``` + +**Measured 2026-09-16 on BOTH `Temurin 26.0.2.1` and `openjdk 25.0.4`: +`ALL PASSED (409 checks)`** across all 16 suites, against `abi 0.11, simd +ndarray::simd avx512, profile release` — identical on the two JDKs. The +runtime line the suite prints at the end names the ABI minor, the SIMD backend +and the library path — read it, because it is what tells you the run exercised +the artifact you meant. + +### What JDK 25 does and does NOT give you + +| need | reachable here? | source | +|---|---|---| +| **Panama FFM** (`java.lang.foreign`, the whole `internal/ffm` membrane) | **YES** — final since 22 | Ubuntu `openjdk-25-jdk-headless`, or Adoptium `temurin-26-jdk` | +| the 16-suite `AllTests` run (409 checks) | **YES**, verified on 25 AND 26 | either of the above | +| `valhalla-lab/src/stable` (records, JDK 26) | **YES** | Adoptium `temurin-26-jdk` | +| `valhalla-lab/src/valhalla` (**value** records, JDK 27 EA) | not via apt | `https://jdk.java.net/valhalla/` — reachable, `HTTP 200` | + +**Only the Valhalla EA arm needs a non-apt fetch, and it is not a gate.** +Adoptium's noble repo carries 8 through 26 and stops there (`apt-cache search +temurin` shows no 27+), and mainline 27 would not help anyway — `value record` +is a Valhalla EA feature, not a mainline one. `https://jdk.java.net/valhalla/` +**is reachable through this environment's proxy** (verified `HTTP 200`), so +that build is a download away rather than a blocker. `bench/` and +`valhalla-lab/` are measurement arms, NOT gates; the merge gate is the Rust +suite plus the 409-check Java run, and both are reachable here. + +**Consequence for every future session:** "the Java side could not be verified" +is no longer an acceptable status line for this repo. The Rust half +(`cargo test` in `native/lgj-abi`) and the Java half (409 checks) are BOTH +runnable in this container, and a PR that claims the Java surface is unverified +is claiming something that takes about four minutes to falsify. + ## Missing-capability STOP rule A consumer or facade that needs a capability the substrate lacks does diff --git a/docs/abi.md b/docs/abi.md index 3364a18..613f0e4 100644 --- a/docs/abi.md +++ b/docs/abi.md @@ -62,7 +62,9 @@ cannot disagree with itself. The ABI is a **machine membrane**. It is not the product. The product is the Java semantic API (see `architecture.md`). Therefore: -- It is **small** — currently 26 symbols (minor 10's one addition — the +- It is **small** — currently 29 symbols (minor 11's three additions are argued + in §19, and the same section argues the SEVEN capabilities it deliberately did + NOT spend a symbol on; 26 at minor 10, whose one addition — the columnar constructor — is argued in §18; 25 at minor 9, whose one addition is argued in §11: a reduction Java was performing on the wrong side of the membrane, moved to where the data is; 24 at minor 8, which adds @@ -83,7 +85,7 @@ semantic API (see `architecture.md`). Therefore: ``` LGJ_ABI_MAJOR = 0 // incompatible change ⇒ bump; Java refuses to load -LGJ_ABI_MINOR = 8 // additive change ⇒ bump; older Java may still load +LGJ_ABI_MINOR = 11 // additive change ⇒ bump; older Java may still load LGJ_MAGIC = 0x4C_47_4A_5F_41_42_49_00 // "LGJ_ABI\0" big-endian-read ``` @@ -143,6 +145,20 @@ required — a gate that rejected everything would satisfy a rejection-only test ### Minor version history +- **Minor 11** (2026-09-14) — the masking-op completion (§19): the ndarray + masking facade finished growing, and this minor consumes what it grew. + **Three symbols and seven op-codes for fifteen capabilities**, which is the + whole argument: `lgj_mask_ternlog` (§19.1) makes every 3-input Boolean + function of three masks an IMMEDIATE rather than a symbol, so XOR, NOT and + MAJ3 arrive without one each; `lgj_op_ternary_match` (§19.2) is the first + bitwise-over-payload predicate this ABI has carried, and the only new + capability whose parameter (24 bytes) genuinely does not fit an op-code; + `lgj_reduce_i32` (§19.3) is the PARAMETERISED reduction §15 mandated in + advance rather than the second and third reduce symbols §1 warns about. The + comparison family (`ne_u32`, `eq`/`ne`/`lt`/`le`/`ge_i32`) and the `u32` TCAM + cost **no symbol at all** — `LgjOpDesc.op` is the generalisation vehicle this + ABI already carries for predicates. No new status; no manifest growth; a + minor-10 Java loads and sees none of it. - **Minor 10** (2026-08-27) — `lgj_rowstore_open_columnar` (§18): the facet-major columnar store. A layout is a SCHEMA over the same 512 bytes per row (R11), so it is a CONSTRUCTOR, not a resource kind: every mask, @@ -411,7 +427,7 @@ predicates or rows are involved. The unfused per-predicate ops are retained only so the fused path can be benchmarked *against* something and so parity can be checked predicate-by-predicate. -## 7. The function surface (24 symbols) +## 7. The function surface (29 symbols) All symbols are prefixed `lgj_`. All return `i32` status except the manifest getter. `out_*` parameters are written only on `OK`. @@ -453,6 +469,7 @@ i32 lgj_mask_describe(u64 mask, LgjLaneDesc* out) // MASK_WORD lan i32 lgj_mask_and(u64 a, u64 b, u64 dst) i32 lgj_mask_or(u64 a, u64 b, u64 dst) i32 lgj_mask_andnot(u64 a, u64 b, u64 dst) // ABI minor ≥ 4, see §13 +i32 lgj_mask_ternlog(u64 a, u64 b, u64 c, u64 dst, u8 imm) // ABI minor ≥ 11, see §19.1 i32 lgj_mask_count(u64 mask, u64* out_count) ``` @@ -471,6 +488,12 @@ i32 lgj_op_gt_i32(u64 res, u32 lane_id, i32 threshold, u64 dst_mask) Each *overwrites* `dst_mask` with the predicate's result. Composition is the caller's job via `lgj_mask_and`. +**Minor 11 adds no symbol here.** The six further comparisons (`ne_u32`, +`eq`/`ne`/`lt`/`le`/`ge_i32`) and the `u32` TCAM are `LgjOpDesc` OP-CODES, so +they are reached through the fused plan below — and `n_ops = 1` is the unfused +form, because the accumulator starts all-set and `all ∩ op0 = op0`. §19.4 is +the argument for spending op-codes rather than symbols. + ### Fused plan (N predicates — ONE crossing) ``` @@ -499,6 +522,8 @@ i32 lgj_reduce_facet_sum(u64 res, u32 facet, u32 carving, u64 mask, i64* out_sum) // ABI minor >= 5, see §14 i32 lgj_reduce_facet_sum_resolved(u64 res, u32 facet, u64 mask, i64* out_sum, u32* out_carving) // minor >= 6, §15 +i32 lgj_reduce_i32(u64 res, u32 lane_id, u32 reduce_op, u64 mask, + i64* out_value, u32* out_present) // minor >= 11, §19.3 ``` Sums the `I32` lane over set mask bits into a widened `i64` (no overflow for @@ -522,6 +547,17 @@ i32 lgj_hop(u64 store, u32 edge_classid, u64 facet_mask, u32 decode_mode, u64 src_mask, u64 dst_mask) ``` +### Row-store register predicate (ABI minor ≥ 11) + +``` +i32 lgj_op_ternary_match(u64 res, u32 facet, + const u8* pattern, const u8* care, u64 dst_mask) +``` + +TCAM over a facet's 12-byte V3 register: `((register ^ pattern) & care) == 0`, +one bit per row. `AosRows` only (`UNSUPPORTED_LAYOUT` otherwise). §19.2 is the +full normative statement. + Overwrites `dst_mask` with the one-hop reachable set from `src_mask` over `store`'s `edge_classid`-matching facets, gated by the `lance-graph-contract` `ClassView`/`FieldMask` LAW. §13 is the full normative statement (effective @@ -1309,3 +1345,325 @@ unaligned loads either way. The carvings' own contract is untouched: every and `512 = 8 × 64` keeps the row stride cache-line-quantised (both pinned in `rowstore.rs`, the substrate half of what R4/R10 measured from the Valhalla side). + +--- + +## 19. The masking-op completion (ABI minor ≥ 11) + +`ndarray`'s masking facade finished growing (`src/simd_masking_ops.rs`, the +ergonomic tier of the three-layer SIMD contract). This minor consumes what it +grew — and the interesting half of the change is what it did **not** spend a +symbol on. + +Fifteen facade capabilities landed in this crate. **Three became symbols** +(29 total, up from 26); **seven became `LgjOpDesc` op-codes**, which cost +nothing; **five are substrate-tier kernels with no membrane exposure at all**, +each with a stated reason and a stated condition for revisiting. §1 makes +symbol growth "a design smell to be argued for, not a default", so every row +below is the argument rather than a changelog. + +| facade capability | how it lands | why | +|---|---|---| +| `mask_ternlog` / `mask_ternlog_assign` | **symbol** `lgj_mask_ternlog` | the family's general member (§19.1) | +| `mask_xor` / `mask_xor_assign` | immediate `0x3C` to that symbol | a 3-input table subsumes it | +| `mask_not` / `mask_not_assign` | immediate `0x0F` to that symbol | likewise | +| `ternary_match_strided_to_mask` | **symbol** `lgj_op_ternary_match` | 24 bytes of parameter; no op-code fits (§19.2) | +| `masked_min_i32` / `masked_max_i32` | **symbol** `lgj_reduce_i32` | §15 mandated ONE parameterised reduce symbol (§19.3) | +| `ne_u32_to_mask` | op-code `3` | `LgjOpDesc.op` already generalises predicates (§19.4) | +| `eq_i32` / `ne_i32` / `lt_i32` / `le_i32` / `ge_i32` | op-codes `4`–`8` | likewise | +| `ternary_match_u32_to_mask` | op-code `9`, packed operand | two `u32`s are exactly 64 bits | +| `ternary_match_u64_to_mask` | kernel only | 128 bits of parameter (§19.5) | +| `mask_any` / `mask_all` | kernel only | exactly derivable from `lgj_mask_count` (§19.5) | +| `blend_i32` | kernel only | no resource carries two `I32` lanes (§19.5) | + +Every one of the fifteen has a kernel wrapper in `kernels.rs` regardless of +exposure. That is the Missing-capability STOP rule read in its own direction: +the substrate tier is where a capability lands *first*, and a symbol is a NAME +for something a backend already does (root `CLAUDE.md`, E5). A capability that +exists in the kernel and not at the membrane is a decision; one that has to be +invented at the membrane is the failure the rule prevents. + +### 19.1 `lgj_mask_ternlog` — the mask-op family's general member + +``` +i32 lgj_mask_ternlog(u64 a, u64 b, u64 c, u64 dst, u8 imm) +``` + +`dst = ternlog::(a, b, c)`, word-wise. `imm` is the 8-bit truth table in +the VPTERNLOG convention (index `(a<<2)|(b<<1)|c`, result bit +`(imm >> index) & 1`). **Every value `0..=255` is legal**, so there is no +unknown-immediate rejection path and no status for one. + +| function | `imm` | +|---|---| +| `a & b` | `0xC0` | +| `a \| b` | `0xFC` | +| `a ^ b` | `0x3C` | +| `a & !b` | `0x30` | +| `!a` | `0x0F` | +| `a & b & c` | `0x80` | +| majority of three | `0xE8` | + +**Why this instead of `lgj_mask_xor` + `lgj_mask_not` + `lgj_mask_maj3` + …** +Because the alternative has no stopping point. XOR and NOT are the two the +mask algebra visibly lacked; MAJ3, `(a&b)|c`, `(a|b)&c` and 250 others are +equally real Boolean functions of three masks, and a membrane that mints a +symbol per truth table is a membrane that grows without bound. One symbol whose +parameter IS the truth table closes the family. This is also the op +`.claude/plans/mask-risc-lowering-v1.md` D-MRL-1a names, consuming +`ogar_loco::TERNLOG = FnIndex(0x86)`'s semantics — **no mint**, per that plan's +own v4.1 amendment. + +**The immediate is runtime at the membrane and compile-time at the facade**, +and closing that gap is the implementation's only interesting decision. The +crate carries a 256-entry table of function pointers, one per monomorphisation +of `ndarray::simd::mask_ternlog_assign::`. The tempting alternative — +decompose the immediate into its eight minterms and OR them word-wise — is +wrong on both counts that matter: it evaluates the truth table HERE, in ~24 word +ops instead of one `VPTERNLOGQ` per 512 bits, and it is a locally-written SIMD +abstraction, which §8 forbids outright. `ndarray`'s POLYFILL LAW puts +truth-table specialisation *inside the backend file*; reaching the const generic +is what keeps it there. + +**Aliasing.** `dst` may be any of `a`, `b`, `c`, or none, and every operand is +read as it was BEFORE the call. The mechanism is `lgj_hop`'s, not +`lgj_mask_and`'s: each operand is snapshotted under a READ lock that is fully +RELEASED before the next lock is taken, and only then is the single WRITE lock +on `dst` acquired. Never more than one lock at a time, so every aliasing case is +deadlock-free **by construction rather than by case analysis** — which is what +four handles need, where `lgj_mask_andnot`'s three could be enumerated. +`registry::lock_masks_ordered` is deliberately not used: it holds at most three +entries, and more decisively it takes a WRITE guard on every one, which is how a +mask's memoised carving (§15) is invalidated — reading `a`/`b`/`c` through it +would silently discard three memos this operation does not touch. + +**The snapshots are a real cost, stated rather than hidden:** up to three +mask-sized copies plus one pass. `lgj_mask_and`/`or`/`andnot` are therefore NOT +removed and NOT reimplemented on top of this — removal would not be an additive +minor, and each keeps a two-mask analysis that is cheaper for its own function. +Call them for AND/OR/ANDNOT; call this for everything else. + +**Tail.** Bits at row index `>= n_rows` are zero on return, and here that clear +is load-bearing rather than defensive. The facade's contract is exact: the +result's tail is `imm & 1` replicated, because every conforming input's tail is +zero and index `0` is the all-zero-inputs row of the table. So an **odd** `imm` +— `NOT_A = 0x0F` among them — sets every tail bit before the clear. A test +asserts that it does, so the clear is a measured requirement and not a +precaution that could be removed without anything going red. + +**Compatibility.** All four masks must share the same parent and row count, or +`MASK_LENGTH_MISMATCH` — the same reading `lgj_mask_and` uses, extended from +three handles to four. The rule is unchanged; only the arity is. + +### 19.2 `lgj_op_ternary_match` — TCAM over the V3 register + +``` +i32 lgj_op_ternary_match(u64 res, u32 facet, + const u8* pattern, const u8* care, u64 dst_mask) +``` + +**Overwrites** `dst_mask` with `((register(row, facet) ^ pattern) & care) == 0` +over the 12-byte content-blind register that follows each facet's classid. + +**This is the first bitwise-over-payload predicate this ABI has ever carried.** +Every predicate before it compares a whole scalar field — a classid, a value, an +id — and `.claude/plans/mask-risc-lowering-v1.md` §1 records that absence +explicitly ("all scalar-valued, none bitwise over the payload"). `care` is what +makes the difference structural rather than incremental: it selects which BITS +must agree, so setting the leading bytes and clearing the rest turns the op into +a PREFIX test, and prefix containment is ancestry (OGAR `CLAUDE.md`'s 3×4 canon; +`rail_geometry.rs:178`'s `is_ancestor_of`). One call, one mask, the whole +subtree — which is the vertical axis that plan's mask trie stands on. + +`care` all-zero matches every row (a deliberate total, not an error); `care` +all-ones is exact register equality. + +**Why a symbol and not an op-code.** Its parameter is 24 bytes — a 12-byte +pattern and a 12-byte care mask. `LgjOpDesc.operand` is 64 bits. Widening the +descriptor would not be an additive change, and packing a pointer into an +integer operand would put a lifetime across a struct the plan evaluator copies. +It is the only capability in this minor whose parameter genuinely does not fit. + +**Layout.** `AosRows` only. A facet's register is 12 CONTIGUOUS bytes there, +which is what makes one strided pass possible; `FacetMajor` (§18) deliberately +splits it into a lo64 region and a hi32 region far apart in the buffer. +Gathering it back per row would be the serialization this ABI exists to forbid, +so the answer is `UNSUPPORTED_LAYOUT` (`-18`) — the same deferral-stated-as-a- +status the register-sweep family already uses. Checked BEFORE the mask is +resolved, so `dst_mask` is provably untouched on a refusal. + +**Mask compatibility is row-count only**, matching `lgj_op_eq_classid` — the +closest sibling, a predicate that writes a mask — rather than the register +sweeps' stricter parent-identity check. abi.md names only the row-count +condition for a predicate's destination (§3), and inventing a second reading for +one new predicate is exactly the drift this membrane exists to prevent. + +**Kernel.** `ndarray::simd::ternary_match_strided_to_mask`, which lowers the +test to one `ternlog::` per 16 records plus a compare-against-zero — +never an XOR, then an AND, then a compare. The primitive owns bounds checking +(overflow-safe, up front) and the trailing-bits-zero guarantee, exactly as +`eq_u32_strided_to_mask` does for `lgj_op_eq_classid`. + +### 19.3 `lgj_reduce_i32` — the parameterised reduction §15 mandated in advance + +``` +i32 lgj_reduce_i32(u64 res, u32 lane_id, u32 reduce_op, u64 mask, + i64* out_value, u32* out_present) +``` + +`reduce_op`: `0 = SUM`, `1 = MIN`, `2 = MAX`. Anything else is +`UNKNOWN_OPCODE`, never a silent fall back to `SUM` — an unknown reduction must +not alias a known one, the same argument §14 makes for an unknown carving. + +**§15 wrote this shape down before the need arose**, and naming that is the +point rather than a flourish: + +> if a second reduction is ever needed (min/max/count-distinct/histogram), do +> NOT add a second symbol. That is the point at which the operation should be +> generalised — an op-code parameter on one reduce symbol, mirroring how +> `lgj_plan_eval`'s `LgjOpDesc` already generalises predicates — and `sum` +> becomes op-code 0. Two reduction symbols would be the smell §1 warns about; +> one parameterised symbol is the shape this ABI already uses elsewhere. + +Min and max ARE that second reduction. So this is one symbol, `sum` is op-code +`0`, and a future `count-distinct` or `histogram` is one more arm rather than +one more symbol. + +**`lgj_reduce_sum_i32` is retained unchanged.** Removing it would not be an +additive minor and a minor-1 Java is still entitled to it; `reduce_op = 0` +answers identically, pinned by a test so the two cannot drift. The §15 rule's +letter ("`sum` becomes op-code 0") is honoured; its one impossible clause +(retroactively deleting a shipped symbol) is not, and that substitution is +recorded here rather than left to be noticed. + +**`out_present` is not a formality.** Min and max over an EMPTY population have +no answer, and every in-band sentinel is wrong: `i32::MAX` is a value a real +population can contain, and `0` is worse. `out_present = 0` means "nothing +selected; `*out_value` is `0` and means nothing". `SUM` is always present, +because the sum of an empty population is `0` and that IS the answer — the same +distinction §15 draws when it refuses to report a zero-fallback carving for an +empty population. + +Both outputs are written on `LGJ_OK` and neither on any failure, matching +`lgj_reduce_facet_sum_resolved`'s two-output precedent. Widening to `i64` is +the SUM's contract and exact for an `i32` extremum, so one return type carries +all three without a lossy case. + +**Cost:** `O(mask_words + popcount)` for every op. The mask-word scan is +unconditional, so an empty mask costs one pass rather than nothing — the cost +shape §14 already states for the register sweep, and §6's bulk rule holds by +the same reading. + +### 19.4 Seven op-codes that cost no symbol + +``` +LGJ_OP_NE_U32 = 3 // U32 lane +LGJ_OP_EQ_I32 = 4 // I32 lane +LGJ_OP_NE_I32 = 5 +LGJ_OP_LT_I32 = 6 +LGJ_OP_LE_I32 = 7 +LGJ_OP_GE_I32 = 8 +LGJ_OP_TERNARY_MATCH_U32 = 9 // U32 lane +``` + +Each is an `LgjOpDesc.op` value, reached through `lgj_plan_eval` / +`lgj_plan_eval_scalar`. Two consequences, both wanted: + +- **Fused for free.** Six new predicates compose with each other and with the + existing two in ONE crossing, which is the property §6 exists to protect. Six + new `lgj_op_*` symbols would have given the unfused form only, and a caller + chaining them would pay a crossing each. +- **Unfused for free too.** `n_ops = 1` IS the unfused form: the accumulator + starts all-set, so `all ∩ op0 = op0` exactly. Nothing is lost by not minting + the symbols. + +`LGJ_OP_TERNARY_MATCH_U32` **packs both halves into `operand`**: +`operand as u64` is `(care << 32) | pattern`. That is exact — two `u32`s are +exactly 64 bits — and it is a new reading of an existing field for a NEW op-code +only; every pre-existing op-code's reading of `operand` is untouched. + +**Both plan symbols gained scalar arms.** An op-code with no scalar +implementation would make `lgj_plan_eval_scalar` answer `UNKNOWN_OPCODE` for +exactly the ops the plan evaluator was extended to cover — an asymmetry between +two symbols §7 says have "identical semantics", and it would quietly hollow out +the parity escape hatch. Each scalar arm is written as the predicate's own +definition (`v <= t`), never as a complement of its neighbour, so a facade that +lowered `ge` as a buggy complement is caught rather than agreed with. + +### 19.5 What is NOT a symbol, and why + +Named so their absence is a decision on record rather than an oversight — §10's +own convention, applied to this minor. + +**`mask_any` / `mask_all` — exactly derivable, so they buy nothing.** +`any` is `lgj_mask_count > 0`; `all` is `lgj_mask_count == n_rows`, and a Java +`Mask` already holds its row count. Both cost ONE crossing either way, and both +are `O(mask_words)` either way — the facade's `mask_any` does not short-circuit, +it ORs every word, exactly as `popcount_batch_u64` sums every word. A symbol +that is a rename of an existing answer at identical cost is the growth §1 names +as a smell. The kernels ship (a future fused plan's `EXISTS` terminal will want +the early-out shape); the membrane does not. + +**`blend_i32` — no resource has two `I32` lanes.** A blend needs two of them. +The pattern fixture has exactly one (`LANE_VALUES`); the row store has none +(`U8` raw, `U32` classid, `U64` lo64, `U32` hi32). Every legal call would be +`blend(m, values, values)`, which is `values` — a symbol whose only reachable +invocation is the identity. **Condition to revisit, stated so it can fire:** a +resource that carries two `I32` lanes, or a measured need for the +constant-fallback form (`dst[i] = mask[i] ? lane[i] : k`), which is a different +primitive and would be named as one. + +**`ternary_match_u64_to_mask` — 128 bits of parameter.** Its `pattern` + `care` +needs twice what `LgjOpDesc.operand` carries, so it cannot be an op-code without +growing the descriptor, which would not be additive. Nor does it earn the symbol +`lgj_op_ternary_match` earns: that one answers over the 12-byte V3 register, +which is where the prefix/ancestry question lives; a `u64` TCAM would answer over +the `ids` lane, for which no caller exists. **Condition to revisit:** a caller +with a real care-masked query over a `U64` lane — at which point the honest shape +is a pointer-taking symbol like §19.2's, not a widened descriptor. + +**`mask_xor` / `mask_not` — already reachable.** Immediates `0x3C` and `0x0F` to +§19.1. A dedicated symbol for each would save one snapshot on the XOR path and +nothing at all on NOT, against two more symbols forever. + +### 19.6 Bulk-rule conformance (§6, applied) + +None of the three is lifecycle, and none is fixed-cost. A caller that doubles +`n_rows` observes roughly double the work in each: `lgj_mask_ternlog` twice the +mask words (and twice the snapshot bytes); `lgj_op_ternary_match` twice the +records scanned; `lgj_reduce_i32` twice the mask-word scan and, on average, +twice the popcount. The seven op-codes inherit `lgj_plan_eval`'s own +conformance unchanged. Same test every existing symbol passes. + +### 19.7 The disable table (red-then-green, or it is not evidence) + +Every guard this minor adds was verified by breaking it and watching the named +test fail, then restoring. Seven runs: + +| what was disabled | test that went red | observed | +|---|---|---| +| the tail clear in `mask_ternlog_impl` | `an_odd_immediate_leaves_no_tail_bit_behind` | counted **127** of 130 rows instead of **65** — the 62 dead bits of the last word | +| the 256-table index, transposed (`[imm & 15][imm >> 4]`) | `the_ternlog_dispatch_table_agrees_with_the_truth_table_for_all_256_immediates` | FAILED | +| the `AosRows` gate in `lgj_op_ternary_match` | `ternary_match_rejects_null_bad_facet_wrong_kind_and_a_facet_major_layout` | FAILED | +| `out_present` hardcoded to `1` | `reduce_i32_reports_an_empty_population_instead_of_inventing_an_extremum` | FAILED | +| one op-code's scalar arm (`le_i32` → `UNKNOWN_OPCODE`) | `every_minor_11_opcode_runs_on_both_plan_paths_and_they_agree` | FAILED | +| the packed TCAM operand's halves, swapped | `the_tcam_opcodes_packed_operand_decodes_both_halves` | FAILED | +| the `a` seed in `mask_ternlog_impl` (never snapshot `a`) | `mask_ternlog_answers_the_truth_table_in_every_aliasing_shape` | FAILED at `imm=0x01, distinct dst` | + +**Two of the seven passed on the first attempt, and both were defects in the +CHECK rather than evidence about the guard.** Recorded because the repo's own +rules name this trap and it fired anyway: + +1. The tail disable inserted `if false { clear_tail_bits(…) }` *beside* the real + call instead of removing it. The `assert count == 1` fired, so the EDIT + happened — but the edit did not remove the guard. *"A disable that does not + apply is indistinguishable from a guard that is not load-bearing."* +2. The aliasing disable flipped `if Arc::ptr_eq(&ed, &ea)` to `if false`, which + makes the code snapshot `a` ALWAYS — strictly more conservative, so nothing + could break. **The knob did not bind.** + +Both were corrected to disables that provably apply (each now asserts the guard +is textually ABSENT from the function body afterwards, not merely that a +replacement occurred), and both then went red. The lesson is the one already on +record and worth one more instance: assert what the disable REMOVED, never only +that an edit landed. diff --git a/java/src/main/java/com/adaworldapi/lancegraph/Lens.java b/java/src/main/java/com/adaworldapi/lancegraph/Lens.java index fc25fd7..a18552e 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/Lens.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/Lens.java @@ -39,6 +39,21 @@ public long sum() { return view.sumOf(field); } + /** + * Minimum of this column over the selected rows, widened to 64 bits (docs/abi.md §19.3, ABI + * minor ≥ 11) — the first of this class's own documented growth: min/max. Absent when the + * underlying view selects no rows: unlike {@link #sum} (always present, since the sum of an + * empty population is {@code 0}), the minimum of an empty population has no answer. + */ + public java.util.OptionalLong min() { + return view.minOf(field); + } + + /** {@link #min}'s maximum sibling — same shape, same absence-on-empty-selection reading. */ + public java.util.OptionalLong max() { + return view.maxOf(field); + } + /** How many rows the underlying view selects. */ public long count() { return view.count(); diff --git a/java/src/main/java/com/adaworldapi/lancegraph/Mask.java b/java/src/main/java/com/adaworldapi/lancegraph/Mask.java index 3b5466b..d3cdfb7 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/Mask.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/Mask.java @@ -1,5 +1,6 @@ package com.adaworldapi.lancegraph; +import com.adaworldapi.lancegraph.internal.ffm.Abi; import com.adaworldapi.lancegraph.internal.ffm.Engine; /** @@ -90,8 +91,81 @@ public Mask minus(Mask other) { java.util.Objects.requireNonNull(other, "other"); requireUsable("minus()"); other.requireUsable("minus()'s argument"); + // Same reasoning as ternlog() below: a pure manifest check, no handle touched, run + // BEFORE dst is allocated rather than left solely to Engine.maskAndNot's own copy of the + // same guard, so a too-old library never leaves an orphaned mask behind. + Abi.requireMinor(4); long dst = Engine.createMask(resourceHandleOf(parent), false); - Engine.maskAndNot(handle, other.handle, dst); + try { + Engine.maskAndNot(handle, other.handle, dst); + } catch (LanceGraphException e) { + // The version gate above already passed; a failure here is the actual op + // (MASK_LENGTH_MISMATCH if other belongs to a mismatched population). dst was + // allocated for a Mask this method never got to construct, so nothing else owns it + // -- release it before the failure propagates, or it and its registry slot leak for + // the life of the process. + closeOnFailure(dst, e); + throw e; + } + return new Mask(parent, dst); + } + + /** + * A new selection: {@code ternlog::(this, b, c)}, word-wise — the mask-op family's + * general member (docs/abi.md §19.1; {@code lgj_mask_ternlog}). {@code imm} is the 8-bit + * VPTERNLOG truth table, index {@code (a<<2)|(b<<1)|c} where {@code a} is {@code this}, result + * bit {@code (imm >> index) & 1}; every value {@code 0..255} is legal, so there is no + * unknown-immediate rejection path. + * + *

This generalises {@link #minus} and the raw AND/OR/XOR/NOT this facade otherwise omits + * (the {@code minus()} javadoc's "public and/or composition stays out of scope" scoped a + * DIFFERENT wave's plan, not this symbol — abi.md §19.1's whole argument is that one + * parameterised member closes the family a per-truth-table method set never would): common + * immediates are {@code 0xC0} = {@code a & b}, {@code 0xFC} = {@code a | b}, {@code 0x3C} = + * {@code a ^ b}, {@code 0x0F} = {@code !a} ({@code b}/{@code c} unused), {@code 0x80} = + * {@code a & b & c}, {@code 0xE8} = majority of three. + * + *

{@code b}/{@code c} need not share this selection's parent resource; if either does not, + * or the row counts differ, the ABI's own {@code MASK_LENGTH_MISMATCH} surfaces as a + * {@link NativeCallException} — this method performs no redundant Java-side parent/row-count + * check of its own, matching {@link #minus}'s own reading. {@code b} and/or {@code c} may be + * {@code this} (or each other) — every operand is read as it stood BEFORE the call, so no + * arrangement of aliasing changes the answer. + * + * @param b the second operand + * @param c the third operand + * @param imm the 8-bit truth table, {@code 0..255} + * @throws IllegalArgumentException if {@code imm} is outside {@code 0..255} + * @throws ClosedResourceException if this selection, {@code b}, {@code c}, or its resource, + * is closed + * @throws AbiMismatchException if the loaded library reports ABI minor < 11 + */ + public Mask ternlog(Mask b, Mask c, int imm) { + java.util.Objects.requireNonNull(b, "b"); + java.util.Objects.requireNonNull(c, "c"); + if (imm < 0 || imm > 255) { + throw new IllegalArgumentException("imm must be in 0..255, was " + imm); + } + requireUsable("ternlog()"); + b.requireUsable("ternlog()'s b argument"); + c.requireUsable("ternlog()'s c argument"); + // A pure manifest check, no handle touched -- run BEFORE dst is allocated rather than + // left solely to Engine.maskTernlog's own copy of the same guard, so a too-old library + // never leaves an orphaned mask behind (there is nothing yet to leak). Engine.maskTernlog + // still carries its own requireMinor(11) too, for any caller that reaches it directly. + Abi.requireMinor(11); + long dst = Engine.createMask(resourceHandleOf(parent), false); + try { + Engine.maskTernlog(handle, b.handle, c.handle, dst, (byte) imm); + } catch (LanceGraphException e) { + // The version gate above already passed; a failure here is the actual op (e.g. + // MASK_LENGTH_MISMATCH if b/c belong to a mismatched population). dst was allocated + // for a Mask this method never got to construct, so nothing else owns it -- release + // it before the failure propagates, or it and its registry slot leak for the life of + // the process. + closeOnFailure(dst, e); + throw e; + } return new Mask(parent, dst); } @@ -285,6 +359,30 @@ private static long resourceHandleOf(NativeResource resource) { + " implementation: " + resource.getClass()); } + /** + * Release a mask allocated as a would-be result, after the fallible native call meant to + * populate it failed before a {@link Mask} could be constructed to own it — the shared + * close-on-failure step for every {@code Engine.createMask(...)} + fallible-op pair on this + * facade ({@link #ternlog}; {@link RowStore#maskOfFacetTernaryMatch} reaches this too, since + * it is package-private "for peers", exactly like {@link #handle()} just above per the + * section comment introducing it). + * + *

Without this, {@code handleToClose}'s native allocation and registry slot would leak + * silently: there is no {@link Mask} object anywhere for a caller to close, because the + * constructor that would have handed one out never ran. The original failure is always what + * propagates to the caller — a failure while releasing {@code handleToClose} is attached to + * it as a {@linkplain Throwable#addSuppressed suppressed} exception rather than replacing it, + * so diagnosing "why did the operation fail" is never hijacked by "why did the cleanup of the + * failed operation fail". + */ + static void closeOnFailure(long handleToClose, LanceGraphException primary) { + try { + Engine.close(handleToClose); + } catch (LanceGraphException cleanupFailure) { + primary.addSuppressed(cleanupFailure); + } + } + private void requireUsable(String what) { if (closed) { throw new ClosedResourceException(what + " was called on a closed selection"); diff --git a/java/src/main/java/com/adaworldapi/lancegraph/NativePattern.java b/java/src/main/java/com/adaworldapi/lancegraph/NativePattern.java index 180612f..f1bd991 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/NativePattern.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/NativePattern.java @@ -1,9 +1,11 @@ package com.adaworldapi.lancegraph; import com.adaworldapi.lancegraph.internal.ffm.Engine; +import com.adaworldapi.lancegraph.internal.ffm.Layouts; import com.adaworldapi.lancegraph.internal.ffm.PlanOp; import java.util.List; +import java.util.OptionalLong; /** * A set of rows held natively, opened once and closed once. @@ -191,6 +193,39 @@ long sumOf(List predicates, I32Field field) { } } + /** + * Evaluate a plan, then reduce one lane over the resulting selection under {@code reduceOp} + * (docs/abi.md §19.3), absent when the selection is empty. Unlike {@link #sumOf} — always + * present, since the sum of an empty population is {@code 0} — MIN/MAX have no answer over + * an empty population, so an {@link OptionalLong} is the honest shape rather than inventing + * one. + */ + private OptionalLong reduceI32(List predicates, I32Field field, int reduceOp) { + requireOpen("reduce()"); + synchronized (lock) { + requireOpen("reduce()"); + long mask; + if (predicates.isEmpty()) { + mask = all(); + } else { + mask = scratch(); + Engine.evaluateFused(handle, plan(predicates), mask); + } + Engine.ReduceOutcome r = Engine.reduceI32(handle, field.lane().index(), reduceOp, mask); + return r.present() ? OptionalLong.of(r.value()) : OptionalLong.empty(); + } + } + + /** {@link #reduceI32}, fixed to {@link Layouts#REDUCE_OP_MIN}. */ + OptionalLong minOf(List predicates, I32Field field) { + return reduceI32(predicates, field, Layouts.REDUCE_OP_MIN); + } + + /** {@link #reduceI32}, fixed to {@link Layouts#REDUCE_OP_MAX}. */ + OptionalLong maxOf(List predicates, I32Field field) { + return reduceI32(predicates, field, Layouts.REDUCE_OP_MAX); + } + /** Materialise a selection the caller owns and closes. */ Mask selectInto(List predicates) { requireOpen("select()"); diff --git a/java/src/main/java/com/adaworldapi/lancegraph/RowStore.java b/java/src/main/java/com/adaworldapi/lancegraph/RowStore.java index 5334145..3d34ead 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/RowStore.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/RowStore.java @@ -1,5 +1,6 @@ package com.adaworldapi.lancegraph; +import com.adaworldapi.lancegraph.internal.ffm.Abi; import com.adaworldapi.lancegraph.internal.ffm.Engine; import com.adaworldapi.lancegraph.internal.ffm.Layouts; @@ -160,6 +161,63 @@ public Mask maskOfFacetClass(FacetId facet, int classId) { return new Mask(this, mask); } + /** + * A selection of the rows whose {@code facet}'s 12-byte content-blind register matches + * {@code ((register ^ pattern) & care) == 0} — TCAM over the V3 register (docs/abi.md §19.2; + * {@code lgj_op_ternary_match}). {@code care} all-zero matches every row (a deliberate total, + * not an error); {@code care} all-ones is exact register equality. Setting the leading bytes + * of {@code care} and clearing the rest turns this into a PREFIX test — one call, one mask, + * a whole ancestry subtree (OGAR's 3×4 canon; prefix containment is ancestry). + * + *

{@code pattern} and {@code care} are each the 12-byte register split as low 64 / high 32 + * bits, matching {@link #payloadLow64At}/{@link #payloadHi32At}'s own convention. + * + *

{@code AosRows} layout only: a store opened with {@link #openColumnar} splits the + * register into per-field regions no strided pass can gather without the serialization this + * ABI exists to forbid, and this method's {@code UNSUPPORTED_LAYOUT} surfaces as a + * {@link NativeCallException} there — checked before the mask is even resolved, matching the + * refusal {@link #facetSumAs} already gives a facet-major store. + * + *

One native crossing ({@code lgj_mask_create} then {@code lgj_op_ternary_match} — two, + * matching {@link #maskOfFacetClass}'s own accounting), a flat, per-call cost proportional to + * the row count, never to any prior result. + * + * @param facet which of the 32 facet lanes to read — a facet index, not a lane id + * @param patternLo64 the low 64 bits of the 12-byte pattern to match + * @param patternHi32 the high 32 bits of the 12-byte pattern to match + * @param careLo64 the low 64 bits of the 12-byte care mask + * @param careHi32 the high 32 bits of the 12-byte care mask + * @throws AbiMismatchException if the loaded library reports ABI minor < 11 + */ + public Mask maskOfFacetTernaryMatch(FacetId facet, long patternLo64, int patternHi32, + long careLo64, int careHi32) { + java.util.Objects.requireNonNull(facet, "facet"); + requireOpen("maskOfFacetTernaryMatch()"); + // A pure manifest check, no handle touched -- run BEFORE mask is allocated rather than + // left solely to Engine.ternaryMatch's own copy of the same guard, so a too-old library + // never leaves an orphaned mask behind (there is nothing yet to leak). Engine.ternaryMatch + // still carries its own requireMinor(11) too, for any caller that reaches it directly. + // Mirrors Mask.ternlog's own reasoning (see that method). + Abi.requireMinor(11); + long mask = Engine.createMask(handle, false); + try { + Engine.ternaryMatch(handle, facet.index(), patternLo64, patternHi32, careLo64, + careHi32, mask); + } catch (LanceGraphException e) { + // The version gate above already passed; a failure here is the actual op -- + // UNSUPPORTED_LAYOUT on a columnar store (see this method's own javadoc), or a + // MASK_LENGTH_MISMATCH. mask was allocated for a Mask this method never got to + // construct, so nothing else owns it -- release it before the failure propagates, or + // it and its registry slot leak for the life of the process. Mask.closeOnFailure is + // package-private "for peers" (the same standing Mask.handle() already grants + // RowStore) precisely so this site does not need to reinvent the same + // close-then-suppress shape. + Mask.closeOnFailure(mask, e); + throw e; + } + return new Mask(this, mask); + } + /** * For every facet, which register grouping this selection's rows carry (abi.md §16) — the * whole-row alignment answer, in ONE crossing. @@ -306,8 +364,21 @@ public Mask hop(int edgeClassid, WideFieldMask facets, Mask src) { java.util.Objects.requireNonNull(facets, "facets"); java.util.Objects.requireNonNull(src, "src"); requireOpen("hop()"); + // Same reasoning as maskOfFacetTernaryMatch() above: a pure manifest check, no handle + // touched, run BEFORE dst is allocated rather than left solely to Engine.hop's own copy + // of the same guard, so a too-old library never leaves an orphaned mask behind. + Abi.requireMinor(4); long dst = Engine.createMask(handle, false); - Engine.hop(handle, edgeClassid, facets.bits(), 0, src.handle(), dst); + try { + Engine.hop(handle, edgeClassid, facets.bits(), 0, src.handle(), dst); + } catch (LanceGraphException e) { + // The version gate above already passed; a failure here is the actual op. dst was + // allocated for a Mask this method never got to construct, so nothing else owns it + // -- release it before the failure propagates, or it and its registry slot leak for + // the life of the process. Reuses Mask.closeOnFailure rather than a second helper. + Mask.closeOnFailure(dst, e); + throw e; + } return new Mask(this, dst); } diff --git a/java/src/main/java/com/adaworldapi/lancegraph/View.java b/java/src/main/java/com/adaworldapi/lancegraph/View.java index 0bea08d..bdd2d04 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/View.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/View.java @@ -2,6 +2,7 @@ import java.util.ArrayList; import java.util.List; +import java.util.OptionalLong; /** * An immutable, lazy description of a set of rows. @@ -86,6 +87,33 @@ public long sumOf(I32Field field) { return owner.sumOf(predicates, field); } + /** + * Minimum of a signed 32-bit column over the rows this view selects, widened to 64 bits — + * the parameterised reduce §15 mandated in advance (docs/abi.md §19.3). Absent when this view + * selects no rows: unlike {@link #sumOf} (always present, since the sum of an empty + * population is {@code 0}), the minimum of an empty population has no answer, so this is the + * honest shape rather than an invented sentinel. + * + *

Two crossings: evaluate the chain, then reduce. Still independent of the row count. + * + * @throws AbiMismatchException if the loaded library reports ABI minor < 11 + */ + public OptionalLong minOf(I32Field field) { + java.util.Objects.requireNonNull(field, "field"); + return owner.minOf(predicates, field); + } + + /** + * {@link #minOf}'s maximum sibling — same shape, same absence-on-empty-selection reading, + * same two-crossing cost. + * + * @throws AbiMismatchException if the loaded library reports ABI minor < 11 + */ + public OptionalLong maxOf(I32Field field) { + java.util.Objects.requireNonNull(field, "field"); + return owner.maxOf(predicates, field); + } + /** * A projection of one column through this view. * diff --git a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Downcalls.java b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Downcalls.java index df03cb3..8e8e38e 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Downcalls.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Downcalls.java @@ -467,6 +467,30 @@ private static final class Minor10 { private Minor10() {} } + /** + * ABI minor 11 symbols (docs/abi.md §19, the masking-op completion): the mask-op family's + * general member, TCAM over the V3 register, and the parameterised min/max reduce §15 + * mandated in advance. Lazy per the minor-2..10 rule. + */ + private static final class Minor11 { + static final MethodHandle MASK_TERNLOG = mh("lgj_mask_ternlog", + FunctionDescriptor.of(ValueLayout.JAVA_INT, ValueLayout.JAVA_LONG, + ValueLayout.JAVA_LONG, ValueLayout.JAVA_LONG, ValueLayout.JAVA_LONG, + ValueLayout.JAVA_BYTE)); + + static final MethodHandle OP_TERNARY_MATCH = mh("lgj_op_ternary_match", + FunctionDescriptor.of(ValueLayout.JAVA_INT, ValueLayout.JAVA_LONG, + ValueLayout.JAVA_INT, ValueLayout.ADDRESS, ValueLayout.ADDRESS, + ValueLayout.JAVA_LONG)); + + static final MethodHandle REDUCE_I32 = mh("lgj_reduce_i32", + FunctionDescriptor.of(ValueLayout.JAVA_INT, ValueLayout.JAVA_LONG, + ValueLayout.JAVA_INT, ValueLayout.JAVA_INT, ValueLayout.JAVA_LONG, + ValueLayout.ADDRESS, ValueLayout.ADDRESS)); + + private Minor11() {} + } + /** * Sum one facet's 12-byte register, under {@code carving}, over the rows a mask selects. * @@ -552,6 +576,69 @@ public static void rowstoreFacetMatchCount(long res, int classId, MemorySegment Status.check("lgj_rowstore_facet_match_count", st); } + /** + * {@code dst = ternlog::(a, b, c)}, word-wise — the mask-op family's general member + * (docs/abi.md §19.1). {@code imm} is the 8-bit VPTERNLOG truth table (index + * {@code (a<<2)|(b<<1)|c}, result bit {@code (imm >> index) & 1}); every value + * {@code 0..=255} is legal, so there is no unknown-immediate rejection path or status for + * one. Each of {@code a}/{@code b}/{@code c} is read as it was BEFORE the call, so + * {@code dst} may alias any of them. + */ + public static void maskTernlog(long a, long b, long c, long dst, byte imm) { + crossed(); + int st; + try { + st = (int) Minor11.MASK_TERNLOG.invokeExact(a, b, c, dst, imm); + } catch (Throwable t) { + throw wrap("lgj_mask_ternlog", t); + } + Status.check("lgj_mask_ternlog", st); + } + + /** + * Overwrites {@code dstMask} with {@code ((register(row, facet) ^ pattern) & care) == 0} + * over the 12-byte content-blind register that follows {@code facet}'s classid — TCAM over + * the V3 register (docs/abi.md §19.2). {@code AosRows} layout only; a facet-major store + * answers {@code UNSUPPORTED_LAYOUT}, checked before {@code dstMask} is even resolved, so it + * is provably untouched on that status. {@code pattern}/{@code care} must each point at + * {@link Layouts#FACET_REGISTER_BYTES} readable bytes; neither is written. + */ + public static void opTernaryMatch(long res, int facet, MemorySegment pattern, + MemorySegment care, long dstMask) { + crossed(); + int st; + try { + st = (int) Minor11.OP_TERNARY_MATCH.invokeExact(res, facet, pattern, care, dstMask); + } catch (Throwable t) { + throw wrap("lgj_op_ternary_match", t); + } + Status.check("lgj_op_ternary_match", st); + } + + /** + * The parameterised reduce §15 mandated in advance (docs/abi.md §19.3): {@code reduceOp} + * {@code 0 = SUM}, {@code 1 = MIN}, {@code 2 = MAX} over an {@code I32} lane, widened to + * {@code i64} in {@code outValue}. {@code outPresent} is written {@code 0} exactly when the + * selected population is empty and {@code reduceOp} is MIN or MAX — SUM is always present. + * An unrecognised {@code reduceOp} is {@code UNKNOWN_OPCODE}, never a silent fallback to + * SUM. Both outputs are written on {@code LGJ_OK} and neither on any failure. + * + * @return {@code outValue}'s widened result, read back after the call + */ + public static long reduceI32(long res, int laneId, int reduceOp, long mask, + MemorySegment outValue, MemorySegment outPresent) { + crossed(); + int st; + try { + st = (int) Minor11.REDUCE_I32.invokeExact(res, laneId, reduceOp, mask, outValue, + outPresent); + } catch (Throwable t) { + throw wrap("lgj_reduce_i32", t); + } + Status.check("lgj_reduce_i32", st); + return outValue.get(ValueLayout.JAVA_LONG, 0); + } + // ── row store (docs/abi.md §11, ABI minor 2) ───────────────────────────────────────────── // // Callers above this class are expected to have already checked Abi.requireMinor(2) — these diff --git a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Engine.java b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Engine.java index acd0894..f7d59a1 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Engine.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Engine.java @@ -323,6 +323,62 @@ public static long rowstoreFacetMatchCount(long store, int classId) { } } + // ── the masking-op completion (docs/abi.md §19, ABI minor ≥ 11) ───────────────────────── + // + // Same requireMinor-before-any-downcall discipline as every section above: a Java build + // compiled against this feature never reaches a missing symbol inside Downcalls, it fails + // here, first, naming the version gap. + + /** + * {@code dst = ternlog::(a, b, c)}, word-wise — the mask-op family's general member + * (docs/abi.md §19.1). Requires ABI minor >= 11. + */ + public static void maskTernlog(long a, long b, long c, long dst, byte imm) { + Abi.requireMinor(11); + Downcalls.maskTernlog(a, b, c, dst, imm); + } + + /** + * Overwrites {@code dstMask} with the TCAM match {@code ((register(row, facet) ^ pattern) & + * care) == 0} over the 12-byte content-blind register that follows {@code facet}'s classid + * (docs/abi.md §19.2). {@code pattern}/{@code care} are each the register split as low 64 / + * high 32 bits, matching {@link RowStore#payloadLow64At}/{@link RowStore#payloadHi32At}'s own + * convention — the two halves are marshalled into scratch buffers here so no caller above + * this class ever touches a {@link MemorySegment}. {@code AosRows} layout only; a facet-major + * store answers {@code UNSUPPORTED_LAYOUT}. Requires ABI minor >= 11. + */ + public static void ternaryMatch(long res, int facet, long patternLo64, int patternHi32, + long careLo64, int careHi32, long dstMask) { + Abi.requireMinor(11); + try (Arena a = Arena.ofConfined()) { + MemorySegment pattern = a.allocate(Layouts.FACET_REGISTER_BYTES, 8); + MemorySegment care = a.allocate(Layouts.FACET_REGISTER_BYTES, 8); + pattern.set(ValueLayout.JAVA_LONG, 0, patternLo64); + pattern.set(ValueLayout.JAVA_INT, 8, patternHi32); + care.set(ValueLayout.JAVA_LONG, 0, careLo64); + care.set(ValueLayout.JAVA_INT, 8, careHi32); + Downcalls.opTernaryMatch(res, facet, pattern, care, dstMask); + } + } + + /** The two outputs of {@link #reduceI32}: a widened value and whether it is meaningful. */ + public record ReduceOutcome(long value, boolean present) {} + + /** + * The parameterised reduce §15 mandated in advance (docs/abi.md §19.3): {@code reduceOp} + * {@link Layouts#REDUCE_OP_SUM}/{@link Layouts#REDUCE_OP_MIN}/{@link Layouts#REDUCE_OP_MAX} + * over an {@code I32} lane, widened to {@code i64}. {@link ReduceOutcome#present} is + * {@code false} exactly when the selected population is empty and {@code reduceOp} is MIN or + * MAX — SUM is always present. Requires ABI minor >= 11. + */ + public static ReduceOutcome reduceI32(long resource, int laneId, int reduceOp, long mask) { + Abi.requireMinor(11); + Scratch s = SCRATCH.get(); + long value = Downcalls.reduceI32(resource, laneId, reduceOp, mask, s.out, s.out2); + boolean present = s.out2.get(ValueLayout.JAVA_INT, 0) != 0; + return new ReduceOutcome(value, present); + } + // ── mask complement + hop (docs/abi.md §13, ABI minor ≥ 4) ───────────────────────────── // // Same requireMinor-before-any-downcall discipline as the row store section above: a Java diff --git a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Layouts.java b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Layouts.java index bb41155..6866f1a 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Layouts.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Layouts.java @@ -90,6 +90,26 @@ private Layouts() {} /** {@code lane[i] > (i32) operand}, signed. Requires an {@code I32} lane. */ public static final int OP_GT_I32 = 2; + // ── lgj_reduce_i32's reduce_op (docs/abi.md §19.3, ABI minor ≥ 11) ─────────────────────── + // + // The parameterised reduce §15 mandated in advance: "if a second reduction is ever needed + // (min/max/count-distinct/histogram), do NOT add a second symbol ... an op-code parameter on + // one reduce symbol ... and sum becomes op-code 0." lgj_reduce_sum_i32 (no reduce_op + // parameter, this file's REDUCE_SUM_I32 handle) is retained unchanged; these three values are + // reduce_op's own numeric convention, confirmed against the compiled artifact. + + /** + * Sum the selected elements, widened to {@code i64}. Always present — the sum of an empty + * population is {@code 0}, a real answer rather than a missing one. + */ + public static final int REDUCE_OP_SUM = 0; + + /** Minimum of the selected elements. Not present over an empty population. */ + public static final int REDUCE_OP_MIN = 1; + + /** Maximum of the selected elements. Not present over an empty population. */ + public static final int REDUCE_OP_MAX = 2; + /** Intersect into the accumulator — narrowing. */ public static final int COMBINE_AND = 0; @@ -321,6 +341,14 @@ private static long off(String name) { public static final long FACET_PAYLOAD_HI32_OFFSET = FACET_PAYLOAD_OFFSET + ValueLayout.JAVA_LONG.byteSize(); + /** + * Width, in bytes, of the 12-byte content-blind register that follows a facet's classid — + * {@code lgj_op_ternary_match}'s {@code pattern}/{@code care} operand width (docs/abi.md + * §19.2; mirrors Rust's {@code kernels::FACET_REGISTER_BYTES}). Derived from the facet's own + * total size minus its classid field, never a literal {@code 12}. + */ + public static final long FACET_REGISTER_BYTES = FACET_BYTES - ValueLayout.JAVA_INT.byteSize(); + /** * Compile-time-ish self check: the byte sizes the ABI document states in prose, checked against * the sizes these layouts actually derive. If a layout is edited wrongly this fails at class diff --git a/java/src/test/java/com/adaworldapi/lancegraph/AllTests.java b/java/src/test/java/com/adaworldapi/lancegraph/AllTests.java index 8b89c67..054c4a9 100644 --- a/java/src/test/java/com/adaworldapi/lancegraph/AllTests.java +++ b/java/src/test/java/com/adaworldapi/lancegraph/AllTests.java @@ -33,6 +33,7 @@ public static void main(String[] args) { suites.put("FacetSumParityTest", FacetSumParityTest::run); suites.put("CarvingTableTest", CarvingTableTest::run); suites.put("ColumnarStoreTest", ColumnarStoreTest::run); + suites.put("MaskingOpCompletionTest", MaskingOpCompletionTest::run); if (!NativeRuntime.isAvailable()) { // ApiSurfaceTest and DoctrineFenceTest need no native library — the API's shape is a diff --git a/java/src/test/java/com/adaworldapi/lancegraph/ApiSurfaceTest.java b/java/src/test/java/com/adaworldapi/lancegraph/ApiSurfaceTest.java index 94e59d7..044a6ef 100644 --- a/java/src/test/java/com/adaworldapi/lancegraph/ApiSurfaceTest.java +++ b/java/src/test/java/com/adaworldapi/lancegraph/ApiSurfaceTest.java @@ -95,7 +95,16 @@ public static void run(Checks c) { List unnamedArrays = new ArrayList<>(); for (Class type : types) { for (Method m : type.getMethods()) { - if (!isPublicApi(m) || m.getDeclaringClass() == Object.class) { + // Members the JDK declares (Object, Throwable's getStackTrace/getSuppressed, + // ...) and the compiler-generated enum values() are not project + // materialisers: their array returns are mandated shapes this API cannot + // rename. Only methods DECLARED in this project are held to the naming law. + // (Measured 2026-09-14: without this the fence reported 11 breaches on + // origin/main itself — five exception types x 2 Throwable members, plus + // Carving.values() — none of them a project-authored crossing.) + boolean jdkDeclared = !m.getDeclaringClass().getName().startsWith("com.adaworldapi.lancegraph"); + boolean enumValues = type.isEnum() && m.getName().equals("values") && m.getParameterCount() == 0; + if (!isPublicApi(m) || m.getDeclaringClass() == Object.class || jdkDeclared || enumValues) { continue; } if (m.getReturnType().isArray() && !isNamedBreach(m.getName())) { diff --git a/java/src/test/java/com/adaworldapi/lancegraph/MaskingOpCompletionTest.java b/java/src/test/java/com/adaworldapi/lancegraph/MaskingOpCompletionTest.java new file mode 100644 index 0000000..bcd471c --- /dev/null +++ b/java/src/test/java/com/adaworldapi/lancegraph/MaskingOpCompletionTest.java @@ -0,0 +1,349 @@ +package com.adaworldapi.lancegraph; + +/** + * Falsifiers for the ABI minor 11 facade additions (docs/abi.md §19, "the masking-op + * completion"): {@link Mask#ternlog}, {@link RowStore#maskOfFacetTernaryMatch}, and + * {@link View#minOf}/{@link View#maxOf} (surfaced again through {@link Lens#min}/{@link Lens#max}, + * the growth point that class's own javadoc names in advance). + * + *

Every "expected" value below is either hand-derived from a truth table written out in this + * file (the {@code ternlog} fixture — eight rows, one per {@code (a,b,c)} combination, so every + * truth-table bit is independently checkable), computed by a Java-side loop over the ALREADY + * public, already-tested per-row accessors {@link RowStore#payloadLow64At}/{@link + * RowStore#payloadHi32At} (the {@code ternary-match} fixture), or transcribed from the row-store + * generator's own published SplitMix64 description (the {@code min}/{@code max} fixture, + * mirroring {@link FixtureParityTest}'s method) — never derived by calling the operation under + * test a second time. + * + *

{@code Abi.requireMinor(11)} itself — the "old library reports minor 10" gate — is exercised + * by {@link OldAbiCompatTest}'s established mechanism (a real older {@code .so}, via + * {@code -Dlgj.oldlibrary}), not here: there is no way in this codebase to fake the loaded + * manifest's minor from within one JVM (it is a {@code static final} read once at class init from + * the real library), so a real older artifact is the only genuine falsifier for that gate. + */ +public final class MaskingOpCompletionTest { + + private MaskingOpCompletionTest() {} + + public static void main(String[] args) { + System.out.println("MaskingOpCompletionTest"); + if (!NativeRuntime.isAvailable()) { + System.exit(Checks.reportUnavailable("MaskingOpCompletionTest")); + } + Checks c = new Checks("MaskingOpCompletionTest"); + run(c); + System.exit(c.report()); + } + + public static void run(Checks c) { + maskTernlog(c); + ternaryMatch(c); + minMax(c); + } + + // ── Mask.ternlog (docs/abi.md §19.1) ───────────────────────────────────────────────────── + + /** + * Eight rows, one per {@code (a,b,c) in {0,1}^3} combination (row index IS the truth-table + * index {@code (a<<2)|(b<<1)|c}), built via {@link RowStore#importRows} so every bit is + * hand-placed rather than fixture-derived. This lets every immediate in docs/abi.md §19.1's + * own table be checked against its EXACT expected row set, by hand, with no RNG in the way. + */ + private static void maskTernlog(Checks c) { + c.section("Mask.ternlog(): the full truth table, row index = (a<<2)|(b<<1)|c"); + final int n = 8; + try (RowStore s = RowStore.open(n, 0x1234L)) { + // a=1 at rows 4..7, b=1 at rows {2,3,6,7}, c=1 at rows {1,3,5,7} — exactly the eight + // (a,b,c) combinations, one per row. + try (Mask a = s.importRows(4, 5, 6, 7); + Mask b = s.importRows(2, 3, 6, 7); + Mask cc = s.importRows(1, 3, 5, 7)) { + + c.eq("a (bit index 2) selects rows 4..7", 4, a.count()); + c.eq("b (bit index 1) selects rows {2,3,6,7}", 4, b.count()); + c.eq("c (bit index 0) selects rows {1,3,5,7}", 4, cc.count()); + + assertTernlog(c, "0x80 = a & b & c (AND3) narrows to exactly row 7", + a, b, cc, 0x80, new long[] {7}); + assertTernlog(c, "0xFF leaves the mask FULL regardless of a/b/c" + + " (anti-vacuity twin of AND3: a different imm, a different answer)", + a, b, cc, 0xFF, new long[] {0, 1, 2, 3, 4, 5, 6, 7}); + assertTernlog(c, "0x00 selects NOTHING regardless of a/b/c (declines, non-trivially:" + + " the same three real operand masks as every other case here)", + a, b, cc, 0x00, new long[] {}); + assertTernlog(c, "0xC0 = a & b -> rows {6,7}", + a, b, cc, 0xC0, new long[] {6, 7}); + assertTernlog(c, "0xFC = a | b -> rows {2,3,4,5,6,7}", + a, b, cc, 0xFC, new long[] {2, 3, 4, 5, 6, 7}); + assertTernlog(c, "0x3C = a ^ b -> rows {2,3,4,5}", + a, b, cc, 0x3C, new long[] {2, 3, 4, 5}); + assertTernlog(c, "0xE8 = majority(a,b,c) -> rows {3,5,6,7}", + a, b, cc, 0xE8, new long[] {3, 5, 6, 7}); + assertTernlog(c, "0x0F = !a (b/c unused) -> rows {0,1,2,3}", + a, b, cc, 0x0F, new long[] {0, 1, 2, 3}); + + c.section("aliasing: b and c may be the SAME Mask object as the receiver"); + try (Mask maj3OfA = a.ternlog(a, a, 0xE8)) { + // majority(x,x,x) == x for any x, so this must equal `a` exactly. + assertSameSet(c, "a.ternlog(a, a, MAJ3) == a (majority of three equal copies)", + maj3OfA.materializeRows(), new long[] {4, 5, 6, 7}); + } + try (Mask notA = a.ternlog(a, a, 0x0F)) { + assertSameSet(c, "a.ternlog(a, a, NOT_A) == !a (rows where a is 0)", + notA.materializeRows(), new long[] {0, 1, 2, 3}); + } + + c.section("imm validation"); + c.throwsUp("imm = 256 is rejected before any crossing", IllegalArgumentException.class, + () -> a.ternlog(b, cc, 256)); + c.throwsUp("imm = -1 is rejected before any crossing", IllegalArgumentException.class, + () -> a.ternlog(b, cc, -1)); + + c.section("closed-selection guard"); + Mask closedSoon = s.importRows(0); + closedSoon.close(); + c.throwsUp("ternlog() on a closed receiver throws", ClosedResourceException.class, + () -> closedSoon.ternlog(a, b, 0xFC)); + c.throwsUp("ternlog() with a closed b/c argument throws", + ClosedResourceException.class, () -> a.ternlog(closedSoon, b, 0xFC)); + } + } + } + + private static void assertTernlog(Checks c, String what, Mask a, Mask b, Mask cc, int imm, + long[] expected) { + try (Mask d = a.ternlog(b, cc, imm)) { + c.eq(what + " (count)", expected.length, d.count()); + assertSameSet(c, what + " (row set)", d.materializeRows(), expected); + } + } + + // ── RowStore.maskOfFacetTernaryMatch (docs/abi.md §19.2) ───────────────────────────────── + + private static void ternaryMatch(Checks c) { + c.section("RowStore.maskOfFacetTernaryMatch(): care=all-ones is exact register equality" + + " (\"equals eq\") — parity against a Java loop over the ALREADY public per-row" + + " accessors, never against the same op called twice"); + final int n = 500; + final int facetIdx = 3; + FacetId facet = FacetId.of(facetIdx); + try (RowStore s = RowStore.open(n, 0xC0FFEEL)) { + long patLo = s.payloadLow64At(0, facet); + int patHi = s.payloadHi32At(0, facet); + + java.util.SortedSet expected = new java.util.TreeSet<>(); + for (long row = 0; row < n; row++) { + long lo = s.payloadLow64At(row, facet); + int hi = s.payloadHi32At(row, facet); + if (lo == patLo && hi == patHi) { + expected.add(row); + } + } + c.that("row 0 matches its own register, reflexively", expected.contains(0L)); + c.that("anti-vacuity: exact 96-bit equality over pseudo-random noise does NOT match" + + " every row", expected.size() < n); + + try (Mask m = s.maskOfFacetTernaryMatch(facet, patLo, patHi, -1L, -1)) { + c.eq("care=all-ones count matches the independent loop", expected.size(), m.count()); + java.util.SortedSet got = new java.util.TreeSet<>(); + for (long row : m.materializeRows()) { + got.add(row); + } + c.eq("care=all-ones row set matches the independent loop", expected, got); + } + + c.section("care=0 is a deliberate TOTAL — matches every row regardless of pattern"); + try (Mask allRows = s.maskOfFacetTernaryMatch(facet, 0x1122334455667788L, 0x2468ACE0, + 0L, 0)) { + c.eq("care=0 selects every row", n, allRows.count()); + } + + c.section("a pattern no row's register attains — declines, not vacuously (a real" + + " search over the actual fixture, not an assumed-absent magic constant)"); + long absentLo = -1; + int absentHi = -1; + boolean found = false; + for (long candidate = 1; candidate <= 64 && !found; candidate++) { + long tryLo = patLo ^ candidate; + int tryHi = patHi; + boolean collides = false; + for (long row = 0; row < n; row++) { + if (s.payloadLow64At(row, facet) == tryLo + && s.payloadHi32At(row, facet) == tryHi) { + collides = true; + break; + } + } + if (!collides) { + absentLo = tryLo; + absentHi = tryHi; + found = true; + } + } + c.that("found a pattern this fixture's own registers never attain (search precondition)", + found); + try (Mask none = s.maskOfFacetTernaryMatch(facet, absentLo, absentHi, -1L, -1)) { + c.eq("a genuinely absent exact pattern selects nothing", 0, none.count()); + } + } + + c.section("AosRows layout only: a facet-major (columnar) store refuses with the LAYOUT" + + " status, checked before dstMask is even resolved"); + try (RowStore col = RowStore.openColumnar(64, 0xC0FFEEL)) { + boolean threw = false; + try { + col.maskOfFacetTernaryMatch(facet, 0L, 0, -1L, -1); + } catch (LanceGraphException e) { + threw = e.getMessage().contains("layout"); + } + c.that("maskOfFacetTernaryMatch on a columnar store names the layout", threw); + } + try (RowStore aos = RowStore.open(64, 0xC0FFEEL)) { + try (Mask m = aos.maskOfFacetTernaryMatch(facet, 0L, 0, 0L, 0)) { + c.that("the same call still works on AoS (the gate discriminates, not a" + + " blanket refusal)", m.count() == 64); + } + } + } + + // ── View.minOf / View.maxOf, and Lens.min / Lens.max (docs/abi.md §19.3) ──────────────── + + /** + * SplitMix64, transcribed from the row-store generator's own published description + * (identical constants and shifts to {@link FixtureParityTest}'s transcription of the + * PATTERN fixture's generator — a coincidence of the shared algorithm, not of the two + * fixtures sharing a schema). + */ + private static final class SplitMix64 { + private long state; + + SplitMix64(long seed) { + this.state = seed; + } + + long next() { + state += 0x9E3779B97F4A7C15L; + long z = state; + z = (z ^ (z >>> 30)) * 0xBF58476D1CE4E5B9L; + z = (z ^ (z >>> 27)) * 0x94D049BB133111EBL; + return z ^ (z >>> 31); + } + } + + private record Lanes(int[] classes, int[] values) {} + + private static Lanes generate(int nRows, long seed) { + SplitMix64 rng = new SplitMix64(seed); + int[] classes = new int[nRows]; + int[] values = new int[nRows]; + for (int i = 0; i < nRows; i++) { + long a = rng.next(); + long b = rng.next(); + classes[i] = (int) ((a >>> 33) & 0xF); + values[i] = (int) (((b >>> 40) & 0x1FF) - 150); + } + return new Lanes(classes, values); + } + + private static void minMax(Checks c) { + c.section("View.minOf/maxOf: absent over an empty selection, unlike sumOf (always" + + " present, since the sum of an empty population is 0)"); + final int rows = 4096; + final long seed = 0xC0FFEEL; + Lanes expected = generate(rows, seed); + try (NativePattern p = NativePattern.open(rows, seed)) { + c.eq("sumOf() over a selection that matches nothing is 0 (SUM's own, unchanged," + + " always-present shape)", 0L, p.view().where(Pattern.CLASS.eq(999)).sumOf( + Pattern.VALUE)); + c.that("minOf() over that same empty selection is ABSENT, not an invented sentinel", + p.view().where(Pattern.CLASS.eq(999)).minOf(Pattern.VALUE).isEmpty()); + c.that("maxOf() likewise", + p.view().where(Pattern.CLASS.eq(999)).maxOf(Pattern.VALUE).isEmpty()); + + c.section("parity: the unrestricted view's min/max matches the independently" + + " computed global extremes"); + int globalMin = Integer.MAX_VALUE; + int globalMax = Integer.MIN_VALUE; + for (int v : expected.values()) { + globalMin = Math.min(globalMin, v); + globalMax = Math.max(globalMax, v); + } + c.eq("unrestricted minOf() == the independently computed global minimum", + java.util.OptionalLong.of(globalMin), p.view().minOf(Pattern.VALUE)); + c.eq("unrestricted maxOf() == the independently computed global maximum", + java.util.OptionalLong.of(globalMax), p.view().maxOf(Pattern.VALUE)); + + c.section("a masked population with a deliberately extreme UNSELECTED row that must" + + " NOT win: pick a class whose own rows never attain the global min/max at" + + " all, so a scan that ignored the mask would answer differently"); + int chosenClass = -1; + int expectedMinK = 0; + int expectedMaxK = 0; + long expectedCountK = 0; + for (int k = 0; k < 16 && chosenClass < 0; k++) { + boolean minLeaks = false; + boolean maxLeaks = false; + int mn = Integer.MAX_VALUE; + int mx = Integer.MIN_VALUE; + long cnt = 0; + for (int i = 0; i < rows; i++) { + if (expected.classes()[i] == k) { + cnt++; + mn = Math.min(mn, expected.values()[i]); + mx = Math.max(mx, expected.values()[i]); + if (expected.values()[i] == globalMin) { + minLeaks = true; + } + if (expected.values()[i] == globalMax) { + maxLeaks = true; + } + } + } + if (cnt > 0 && !minLeaks && !maxLeaks) { + chosenClass = k; + expectedMinK = mn; + expectedMaxK = mx; + expectedCountK = cnt; + } + } + c.that("found a class none of whose rows attain the true global extremes (setup" + + " precondition — the fixture must actually offer this shape)", + chosenClass >= 0); + c.that("the chosen class' own minimum is strictly greater than the true global" + + " minimum (the excluded row's value, by construction)", + expectedMinK > globalMin); + c.that("the chosen class' own maximum is strictly less than the true global maximum", + expectedMaxK < globalMax); + + View restricted = p.view().where(Pattern.CLASS.eq(chosenClass)); + c.eq("restricted minOf() matches the independent per-class computation, NOT the" + + " global minimum", java.util.OptionalLong.of(expectedMinK), + restricted.minOf(Pattern.VALUE)); + c.eq("restricted maxOf() matches the independent per-class computation, NOT the" + + " global maximum", java.util.OptionalLong.of(expectedMaxK), + restricted.maxOf(Pattern.VALUE)); + c.that("restricted selection is non-trivial: neither empty nor the whole population", + expectedCountK > 0 && expectedCountK < rows); + + c.section("Lens.min()/max() (the growth point Lens's own class javadoc names in" + + " advance) agree with View.minOf()/maxOf()"); + Lens lens = restricted.lens(Pattern.VALUE); + c.eq("Lens.min() == View.minOf()", restricted.minOf(Pattern.VALUE), lens.min()); + c.eq("Lens.max() == View.maxOf()", restricted.maxOf(Pattern.VALUE), lens.max()); + c.eq("Lens.sum() is unaffected (still the pre-existing SUM path)", + restricted.sumOf(Pattern.VALUE), lens.sum()); + } + } + + private static void assertSameSet(Checks c, String what, long[] a, long[] b) { + java.util.SortedSet sa = new java.util.TreeSet<>(); + for (long x : a) { + sa.add(x); + } + java.util.SortedSet sb = new java.util.TreeSet<>(); + for (long x : b) { + sb.add(x); + } + c.eq(what, sb, sa); + } +} diff --git a/java/src/test/java/com/adaworldapi/lancegraph/OldAbiCompatTest.java b/java/src/test/java/com/adaworldapi/lancegraph/OldAbiCompatTest.java index 16d67d3..00a08af 100644 --- a/java/src/test/java/com/adaworldapi/lancegraph/OldAbiCompatTest.java +++ b/java/src/test/java/com/adaworldapi/lancegraph/OldAbiCompatTest.java @@ -88,6 +88,26 @@ public static void run(Checks c) { } }); + // Minor 11 — the parameterised min/max reduction. Gated here, UNCONDITIONALLY, rather + // than nested inside the `if (loaded >= 2)` block below: this feature is built entirely + // on NativePattern, the same minor-1 base surface already exercised in section (1) + // above, and has no RowStore dependency at all — unlike every other minor-11+ gate + // below, each of which needs a minor-2 RowStore just to construct its own fixture (a + // Mask, a facet sum, ...). Nesting this gate the same way would have SKIPPED it + // entirely against a genuinely minor-1 library — the exact library this compatibility + // suite exists to check against, and precisely the case a review caught: an + // eager-resolution or missing-symbol regression in lgj_reduce_i32 would have passed + // this run clean. + gate(c, loaded, 11, "View.minOf/maxOf", () -> { + try (NativePattern p = NativePattern.open(64, 0x1234L)) { + java.util.OptionalLong mn = p.view().minOf(Pattern.VALUE); + java.util.OptionalLong mx = p.view().maxOf(Pattern.VALUE); + if (mn.isEmpty() || mx.isEmpty()) { + throw new IllegalStateException("64 rows must select something"); + } + } + }); + // Minor 4 — mask complement. Needs a minor-2 store to build masks on, so it is only // meaningful once the library has minor 2 as well. if (loaded >= 2) { @@ -148,6 +168,32 @@ public static void run(Checks c) { } } }); + + // Minor 11 — the masking-op completion: mask_ternlog and TCAM + // ternary-match over the V3 register. Against an older library each + // gate must name minor 11, never a missing symbol and never a + // failure of some OTHER minor's feature. (The min/max reduce sibling + // of this trio has no RowStore dependency and is gated unconditionally + // above, before this block — see its own comment for why.) + gate(c, loaded, 11, "Mask.ternlog", () -> { + try (RowStore s = RowStore.open(64, 0x1234L); + Mask a = s.maskOfFacetClass(FacetId.of(0), 3); + Mask b = s.maskOfFacetClass(FacetId.of(0), 4); + Mask d = a.ternlog(a, b, 0xFC)) { // OR + if (d.count() > 64) { + throw new IllegalStateException("impossible count"); + } + } + }); + + gate(c, loaded, 11, "RowStore.maskOfFacetTernaryMatch", () -> { + try (RowStore s = RowStore.open(64, 0x1234L); + Mask m = s.maskOfFacetTernaryMatch(FacetId.of(0), 0L, 0, 0L, 0)) { + if (m.count() != 64) { + throw new IllegalStateException("care=0 must match every row"); + } + } + }); } else { c.note("minors 4 and 5 need a minor-2 row store to build a mask on; skipped here" + " because this library predates it"); diff --git a/native/lgj-abi/Cargo.lock b/native/lgj-abi/Cargo.lock index 1c6dd7c..37977ba 100644 --- a/native/lgj-abi/Cargo.lock +++ b/native/lgj-abi/Cargo.lock @@ -51,11 +51,27 @@ dependencies = [ "serde_yaml", ] +[[package]] +name = "lance-graph-mask-risc" +version = "0.1.0" +dependencies = [ + "ndarray", +] + +[[package]] +name = "lance-graph-quack" +version = "0.1.0" +dependencies = [ + "lance-graph-mask-risc", +] + [[package]] name = "lgj-abi" version = "0.1.0" dependencies = [ "lance-graph-contract", + "lance-graph-mask-risc", + "lance-graph-quack", "ndarray", "ogar-class-view", ] diff --git a/native/lgj-abi/Cargo.toml b/native/lgj-abi/Cargo.toml index f1ab669..bacfde9 100644 --- a/native/lgj-abi/Cargo.toml +++ b/native/lgj-abi/Cargo.toml @@ -45,12 +45,42 @@ ndarray = { path = "../../../ndarray", default-features = false, features = ["st # this crate consumes, and the crate compiles fine without them. lance-graph-contract = { path = "../../../lance-graph/crates/lance-graph-contract", default-features = false } +# The ONE mask evaluator. `plan_eval` lowers `&[LgjOpDesc]` to a +# `mask_risc::Program` and runs it; this crate keeps no second opcode loop. +# That is the whole point of PR4 — before it, `plan_eval_impl` held its own +# accumulate-and-combine loop over `kernels::`, a second evaluator beside the +# one in lance-graph, each able to drift from the other. +# +# It is NOT a second source of SIMD (abi.md §8 is unchanged): mask-risc names +# no ISA at all — a test in that crate greps its own source for one — and +# delegates every op to `ndarray::simd`, the same facade `kernels.rs` uses. +# So the SIMD floor is still ndarray and only ndarray; what moved is who +# sequences the ops. +lance-graph-mask-risc = { path = "../../../lance-graph/crates/lance-graph-mask-risc" } + # The REAL ClassView provider (OGAR Core). `FixtureClassView` answers the same # 32 facets for every classid; `OgarClassView` walks `ogar_vocab`'s promoted # classes and gives a genuine per-class field basis -- attributes AND # associations -- which is what makes `edge_participation` discriminate. ogar-class-view = { path = "../../../OGAR/crates/ogar-class-view", optional = true } +[dev-dependencies] +# TEST-ONLY, and the boundary is the point. `lance-graph-quack` is the +# DuckDB-shaped query surface in the lance-graph repo; it lowers a Boolean +# TREE to the same `mask_risc::Program` this crate's `plan_lower` lowers a +# flat OP LIST to. The two therefore implement the same law twice — the +# accumulator gate, the AND/OR gating asymmetry, the drop of ops before the +# first AND — and nothing compared them, because each has its own oracle: +# this crate's is the frozen pre-PR4 loop, quack's is a per-row reading. +# Neither can see the other drift. +# +# `src/exports/tests/lowering_convergence.rs` closes that, and it is a +# dev-dependency rather than a real one ON PURPOSE: the membrane must not +# depend on a consumer of the IR it serves. Sharing the LAW is the goal; +# sharing a DEPENDENCY in that direction would invert the layering. What +# ships is the differential, not a delegation. +lance-graph-quack = { path = "../../../lance-graph/crates/lance-graph-quack" } + [profile.release] opt-level = 3 lto = "thin" diff --git a/native/lgj-abi/src/abi.rs b/native/lgj-abi/src/abi.rs index 73a0c9e..12e6de5 100644 --- a/native/lgj-abi/src/abi.rs +++ b/native/lgj-abi/src/abi.rs @@ -69,7 +69,20 @@ pub const LGJ_ABI_MAJOR: u32 = 0; /// require only the base 104-byte prefix rather than the full layout — without /// that, every future manifest field would be a hard incompatibility with every /// older artifact. -pub const LGJ_ABI_MINOR: u32 = 10; +/// +/// Minor **11** (2026-09-14): the masking-op completion (`docs/abi.md` §19). +/// THREE symbols — [`crate::exports::lgj_mask_ternlog`] (the mask-op family's +/// general member: every 3-input Boolean function of three masks, so XOR, NOT, +/// MAJ3 and 252 others arrive as immediates rather than as symbols), +/// [`crate::exports::lgj_op_ternary_match`] (TCAM over the V3 12-byte facet +/// register — the first bitwise-over-payload predicate this ABI has ever +/// carried), and [`crate::exports::lgj_reduce_i32`] (the PARAMETERISED +/// reduction §15 mandated in advance: `sum`/`min`/`max` behind one op-code +/// argument, never a second and third reduce symbol) — plus SEVEN new +/// [`LgjOpDesc`] op-codes that cost no symbol at all +/// ([`LGJ_OP_NE_U32`] … [`LGJ_OP_TERNARY_MATCH_U32`]). No new status and no +/// manifest growth, so a minor-10 Java loads and sees none of it. +pub const LGJ_ABI_MINOR: u32 = 11; /// `"LGJ_ABI\0"` read big-endian. /// @@ -251,6 +264,62 @@ pub const LGJ_OP_EQ_U32: u32 = 1; /// `lane[i] > operand as i32` (signed), over an `I32` lane. pub const LGJ_OP_GT_I32: u32 = 2; +// ── ABI minor >= 11 (docs/abi.md §19) ─────────────────────────────────────── +// +// These cost NO symbol. `LgjOpDesc.op` is the generalisation vehicle this ABI +// already carries for predicates (§1: "growth is a design smell"), so the +// comparison family lands as op-codes and is reachable BOTH fused +// (`lgj_plan_eval`, N predicates one crossing) and unfused (`n_ops = 1`, whose +// accumulator is all-set so the result is exactly `op0`). + +/// `lane[i] != operand as u32`, over a `U32` lane. ABI minor >= 11. +pub const LGJ_OP_NE_U32: u32 = 3; +/// `lane[i] == operand as i32` (signed lanes, exact), over an `I32` lane. +/// ABI minor >= 11. +pub const LGJ_OP_EQ_I32: u32 = 4; +/// `lane[i] != operand as i32`, over an `I32` lane. ABI minor >= 11. +pub const LGJ_OP_NE_I32: u32 = 5; +/// `lane[i] < operand as i32` (signed, strict), over an `I32` lane. +/// ABI minor >= 11. +pub const LGJ_OP_LT_I32: u32 = 6; +/// `lane[i] <= operand as i32` (signed), over an `I32` lane. ABI minor >= 11. +pub const LGJ_OP_LE_I32: u32 = 7; +/// `lane[i] >= operand as i32` (signed), over an `I32` lane. ABI minor >= 11. +pub const LGJ_OP_GE_I32: u32 = 8; + +/// `((lane[i] ^ pattern) & care) == 0` — care-masked (ternary/TCAM) equality +/// over a `U32` lane. ABI minor >= 11. +/// +/// **The operand packs BOTH halves**: `operand as u64` is +/// `(care << 32) | pattern`, i.e. `pattern` in the low 32 bits and `care` in +/// the high 32. That is a new reading of an existing field for a NEW op-code +/// only — every pre-existing op-code's reading of `operand` is untouched — and +/// it is exact, because two `u32`s are exactly 64 bits. `care == 0` matches +/// every element; `care == u32::MAX` is [`LGJ_OP_EQ_U32`]. +/// +/// The `u64` sibling (`ndarray::simd::ternary_match_u64_to_mask`) is a NAMED +/// GAP, not an oversight: its pattern+care is 128 bits and `LgjOpDesc.operand` +/// carries 64, so it cannot be an op-code without growing the descriptor — +/// which would not be additive. It exists as a kernel +/// ([`crate::kernels::simd_ternary_match_u64_to_mask`]) and is stated as +/// unexposed in `docs/abi.md` §19 rather than half-exposed here. +pub const LGJ_OP_TERNARY_MATCH_U32: u32 = 9; + +// ── Reduction op-codes for `lgj_reduce_i32` (ABI minor >= 11, §19) ────────── +// +// abi.md §15 wrote the rule this obeys BEFORE the need arose: "if a second +// reduction is ever needed (min/max/count-distinct/histogram), do NOT add a +// second symbol ... an op-code parameter on one reduce symbol ... and `sum` +// becomes op-code 0." + +/// Sum the selected elements, widened to `i64`. Always "present" — the sum of +/// an empty population is `0`, a real answer rather than a missing one. +pub const LGJ_REDUCE_SUM: u32 = 0; +/// Minimum of the selected elements. Not present over an empty population. +pub const LGJ_REDUCE_MIN: u32 = 1; +/// Maximum of the selected elements. Not present over an empty population. +pub const LGJ_REDUCE_MAX: u32 = 2; + /// Narrow the accumulator: `acc &= op_result`. pub const LGJ_COMBINE_AND: u32 = 0; /// Widen the accumulator: `acc |= op_result`. @@ -259,8 +328,9 @@ pub const LGJ_COMBINE_OR: u32 = 1; /// The element kind an opcode requires of its lane. `None` ⇒ unknown opcode. pub(crate) const fn opcode_required_kind(op: u32) -> Option { match op { - LGJ_OP_EQ_U32 => Some(LgjElemKind::U32), - LGJ_OP_GT_I32 => Some(LgjElemKind::I32), + LGJ_OP_EQ_U32 | LGJ_OP_NE_U32 | LGJ_OP_TERNARY_MATCH_U32 => Some(LgjElemKind::U32), + LGJ_OP_GT_I32 | LGJ_OP_EQ_I32 | LGJ_OP_NE_I32 | LGJ_OP_LT_I32 | LGJ_OP_LE_I32 + | LGJ_OP_GE_I32 => Some(LgjElemKind::I32), _ => None, } } diff --git a/native/lgj-abi/src/exports.rs b/native/lgj-abi/src/exports.rs index dcc002a..59e7c93 100644 --- a/native/lgj-abi/src/exports.rs +++ b/native/lgj-abi/src/exports.rs @@ -22,13 +22,50 @@ //! `out_*` parameters are written **only on `OK`**, so a failed call cannot //! leave Java reading a half-filled descriptor. +use std::cell::RefCell; use std::panic::{catch_unwind, AssertUnwindSafe}; +use lance_graph_mask_risc::{ + execute, reference_execute, reference_scratch, scratch_words_for, ExecError, LaneRef, Planes, + Scratch, Value, +}; + use crate::abi::*; use crate::fixture::PATTERN_LANE_COUNT; use crate::kernels::{self, LaneView, Path}; +use crate::plan_lower::{self, Lowered}; use crate::registry::{self, ResourceEntry}; +thread_local! { + /// The plan evaluator's scratch arena. + /// + /// Thread-local and grown monotonically, which is the whole of OQ-1's answer. + /// Per-pattern was disqualified: it would give the one deliberately lock-free + /// resource kind a `RwLock` and serialise concurrent `lgj_plan_eval` calls on + /// one pattern — the common shape. A thread-local has no lock to take. + /// + /// Allocation is a function of the MAXIMUM row count a thread has seen, never + /// of its history: a smaller call carves a strict prefix of the same buffer, + /// because `Scratch::over` is handed a length and uses exactly that prefix. + /// That is what lets an unseen, smaller row count stay allocation-free rather + /// than be warmed into passing by an earlier call of the same size. + /// + /// The cost, stated rather than absorbed: a thread that has evaluated one + /// 65,536-row plan holds `scratch_words_for(1024, 2) * 8` bytes until it + /// exits. Bounded, and not per call. + static PLAN_SCRATCH: RefCell> = const { RefCell::new(Vec::new()) }; +} + +// The lowering addresses lanes by using `LgjOpDesc.lane_id` DIRECTLY as an +// index into `Planes::lanes`. That is correct only while the fixture's lane +// ids are 0, 1, 2 in this order — renumber them and every lowered predicate +// silently reads a different column, of the right kind, on a plan the +// validator happily accepts. The cheapest possible guard, and it fails the +// build rather than the run. +const _: () = assert!(crate::fixture::LANE_IDS == 0); +const _: () = assert!(crate::fixture::LANE_CLASSES == 1); +const _: () = assert!(crate::fixture::LANE_VALUES == 2); + /// Run `f`, converting a panic into [`LGJ_ERR_PANIC`]. /// /// `AssertUnwindSafe` is required because the closures below capture raw @@ -681,6 +718,142 @@ pub extern "C" fn lgj_mask_andnot(a: u64, b: u64, dst: u64) -> i32 { guard(|| mask_andnot_impl(a, b, dst)) } +/// The body behind [`lgj_mask_ternlog`]. +/// +/// # Why this does not use `lock_masks_ordered` +/// +/// That helper solves the multi-mask lock problem for up to THREE entries, and +/// this call has FOUR handles. Two things stop it being the tool here, and +/// neither is the count: +/// +/// 1. it takes a WRITE guard on every entry, and a write guard is how a mask's +/// memoised carving is invalidated — so reading `a`/`b`/`c` through it would +/// silently throw away three memos this operation does not touch; +/// 2. the deadlock-freedom it buys is unnecessary once no two locks are ever +/// held at once. +/// +/// So this follows [`lgj_hop`]'s discipline instead — snapshot each operand +/// under a READ lock that is fully RELEASED before the next lock is taken, then +/// take the single WRITE lock. One lock at a time means aliasing carries zero +/// deadlock risk **by construction rather than by case analysis**, and the +/// snapshots give the aliasing SEMANTICS for free: every operand is the value +/// it had BEFORE the call, whichever of the four handles `dst` happens to be. +fn mask_ternlog_impl(a: u64, b: u64, c: u64, dst: u64, imm: u8) -> i32 { + let (ea, pa) = match registry::resolve_mask_with_parent(a) { + Ok(t) => t, + Err(e) => return e, + }; + let (eb, _) = match registry::resolve_mask_with_parent(b) { + Ok(t) => t, + Err(e) => return e, + }; + let (ec, _) = match registry::resolve_mask_with_parent(c) { + Ok(t) => t, + Err(e) => return e, + }; + let (ed, _) = match registry::resolve_mask_with_parent(dst) { + Ok(t) => t, + Err(e) => return e, + }; + + // Same "share the same parent and row count" reading as `mask_binop` and + // `mask_andnot_impl` (abi.md §7/§13), extended from three masks to four — + // the rule is unchanged, only the arity is. + if ea.n_rows != eb.n_rows || ea.n_rows != ec.n_rows || ea.n_rows != ed.n_rows { + return LGJ_ERR_MASK_LENGTH_MISMATCH; + } + if ea.parent != eb.parent || ea.parent != ec.parent || ea.parent != ed.parent { + return LGJ_ERR_MASK_LENGTH_MISMATCH; + } + let n_rows = pa.n_rows; + + // Snapshot `b` and `c`, each under its own read lock, each released before + // the next is taken. + let snap_b: Vec = match eb.read_mask() { + Some(g) => g.words.to_vec(), + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + let snap_c: Vec = match ec.read_mask() { + Some(g) => g.words.to_vec(), + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + // `a` is snapshotted ONLY when it is not already `dst`: in the aliasing + // case the write guard's own buffer already holds `a`'s pre-call value, + // which is the first truth-table operand, so the copy would be pure cost. + let snap_a: Option> = if std::sync::Arc::ptr_eq(&ed, &ea) { + None + } else { + match ea.read_mask() { + Some(g) => Some(g.words.to_vec()), + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + } + }; + + let mut g = match ed.write_mask() { + Some(g) => g, + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + if g.words.len() != snap_b.len() || g.words.len() != snap_c.len() { + return LGJ_ERR_MASK_LENGTH_MISMATCH; + } + if let Some(sa) = &snap_a { + if g.words.len() != sa.len() { + return LGJ_ERR_MASK_LENGTH_MISMATCH; + } + g.words.copy_from_slice(sa); + } + kernels::simd_mask_ternlog_assign_dyn(imm, &mut g.words, &snap_b, &snap_c); + // MANDATORY, not defensive-in-the-usual-sense: the result's tail is + // `imm & 1` replicated, so every ODD table — `NOT_A = 0x0F` among them — + // sets every bit past the population. `lgj_mask_count` reads that tail. + clear_tail_bits(&mut g.words, n_rows); + LGJ_OK +} + +/// `dst = ternlog::(a, b, c)` — ANY 3-input Boolean function of three +/// masks, in one pass (ABI minor >= 11, `docs/abi.md` §19). +/// +/// This is the mask-op family's GENERAL MEMBER, and adding it is why XOR, NOT, +/// MAJ3 and 252 other functions arrive without a symbol each. `imm` is the +/// 8-bit truth table in the VPTERNLOG convention (index `(a<<2)|(b<<1)|c`, +/// result bit `(imm >> index) & 1`); every value `0..=255` is legal, so there +/// is no unknown-immediate rejection path. The spellings this ABI's own prose +/// names: +/// +/// | function | `imm` | +/// |---|---| +/// | `a & b` | `0xC0` | +/// | `a \| b` | `0xFC` | +/// | `a ^ b` | `0x3C` | +/// | `a & !b` | `0x30` | +/// | `!a` | `0x0F` | +/// | `a & b & c` | `0x80` | +/// | majority of three | `0xE8` | +/// +/// The three existing algebra symbols ([`lgj_mask_and`], [`lgj_mask_or`], +/// [`lgj_mask_andnot`]) are NOT removed and NOT reimplemented on top of this: +/// removal would not be an additive minor, and each keeps a two-mask aliasing +/// analysis that is cheaper than this symbol's snapshot discipline. They remain +/// the ones to call for their own function; this is the one to call for +/// everything else. +/// +/// **Aliasing:** `dst` may be any of `a`, `b`, `c`, or none of them, and every +/// operand is read as it was BEFORE the call — see [`mask_ternlog_impl`] for +/// the mechanism. All four must share the same parent and row count, or +/// `MASK_LENGTH_MISMATCH`. +/// +/// **Tail:** bits at row index `>= n_rows` are zero on return. That is load +/// bearing rather than incidental here: unlike AND/OR/ANDNOT, whose tables are +/// all even, an ODD `imm` is true of `(0,0,0)` and therefore sets every tail +/// bit before the clear. +/// +/// **Bulk (§6):** work is `n_rows / 64` words for the pass, plus up to three +/// snapshots of the same size — doubling `n_rows` doubles both. +#[no_mangle] +pub extern "C" fn lgj_mask_ternlog(a: u64, b: u64, c: u64, dst: u64, imm: u8) -> i32 { + guard(|| mask_ternlog_impl(a, b, c, dst, imm)) +} + /// Population count of a mask — how many rows are selected. /// # Safety /// @@ -864,6 +1037,126 @@ pub extern "C" fn lgj_op_eq_classid(res: u64, facet: u32, needle: u32, dst_mask: }) } +/// TCAM over a facet's 12-byte V3 register: **overwrites** `dst_mask` with +/// `((register(row, facet) ^ pattern) & care) == 0` (ABI minor >= 11, +/// `docs/abi.md` §19). +/// +/// # Why this is a symbol and the comparison family is not +/// +/// Every other predicate this minor adds is an `LgjOpDesc` op-code, because +/// `operand: i64` carries all the parameter each of them has. This one carries +/// **24 bytes** — a 12-byte pattern and a 12-byte care mask — so it cannot be +/// an op-code without growing the descriptor, which would not be additive. It +/// is also the first BITWISE-over-payload predicate in this ABI at all: every +/// predicate before it compared a whole scalar field (a classid, a value, an +/// id), and `.claude/plans/mask-risc-lowering-v1.md` §1 records that absence as +/// the gap its vertical axis needs closed. +/// +/// # What `care` buys +/// +/// `care` selects which BITS must agree; the rest are "don't care". That makes +/// this a PREFIX test — set the leading bytes, clear the rest, and the answer +/// is "which rows carry this prefix", which is HHTL ancestry expressed as one +/// mask. `care` all-zero matches every row (a deliberate total, not an error); +/// `care` all-ones is exact register equality. +/// +/// # Layout +/// +/// `AosRows` only. A facet's register is 12 CONTIGUOUS bytes there — the +/// property that makes one strided pass possible — and `FacetMajor` +/// deliberately splits it into a lo64 region and a hi32 region far apart in the +/// buffer. Gathering it back together per row would be the serialization this +/// ABI exists to forbid, so the answer is `UNSUPPORTED_LAYOUT` (`-18`), the +/// same deferral-stated-as-a-status the register-sweep family already uses +/// (§14/§15/§16, abi.md §18). +/// +/// # Mask compatibility +/// +/// Row-count only, matching [`lgj_op_eq_classid`] — the closest sibling, a +/// predicate that WRITES a mask — rather than the register sweeps' stricter +/// parent-identity check. abi.md names only the row-count condition for a +/// predicate's destination, and inventing a second reading for one new +/// predicate would be the drift this membrane is built to prevent. +/// +/// # Bulk (§6) +/// +/// One strided pass over `n_rows` records. Doubling `n_rows` doubles the work. +/// +/// # Safety +/// +/// A null `pattern` or `care` is *handled*, not UB: `NULL_ARGUMENT`. Otherwise +/// each must point at 12 readable bytes that stay valid for the call. Neither +/// is written. `dst_mask` is written only on `LGJ_OK`. +#[no_mangle] +pub unsafe extern "C" fn lgj_op_ternary_match( + res: u64, + facet: u32, + pattern: *const u8, + care: *const u8, + dst_mask: u64, +) -> i32 { + guard(|| { + if pattern.is_null() || care.is_null() { + return LGJ_ERR_NULL_ARGUMENT; + } + let store_entry = match registry::resolve_kind(res, LGJ_RESOURCE_ROWSTORE) { + Ok(e) => e, + Err(e) => return e, + }; + let store = match store_entry.rowstore() { + Some(s) => s, + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + // Checked BEFORE the mask is resolved and long before anything is + // written, so `dst_mask` is provably untouched on a refusal — the same + // ordering §14's carving check and §13's decode_mode check use. + if store.layout != crate::rowstore::RowLayout::AosRows { + return LGJ_ERR_UNSUPPORTED_LAYOUT; + } + if facet >= crate::rowstore::ROW_FACETS { + return LGJ_ERR_INVALID_LANE; + } + let (mask, _parent) = match registry::resolve_mask_with_parent(dst_mask) { + Ok(t) => t, + Err(e) => return e, + }; + if mask.n_rows != store_entry.n_rows { + return LGJ_ERR_MASK_LENGTH_MISMATCH; + } + + const N: usize = kernels::FACET_REGISTER_BYTES; + // SAFETY: both pointers are non-null (checked above) and the caller + // states each points at `N` readable bytes. Java builds them from a + // MemorySegment whose size it derived from the manifest's facet + // geometry. Copied out immediately; neither pointer outlives this call. + let (pat, car): ([u8; N], [u8; N]) = + unsafe { (*(pattern as *const [u8; N]), *(care as *const [u8; N])) }; + + let mut g = match mask.write_mask() { + Some(g) => g, + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + // ONE spelling of the geometry (E3): under `AosRows` the register + // BEGINS at the lo64 lane's base and the hi32 field is the four bytes + // that follow it, contiguously — which is exactly the property + // `FacetMajor` breaks and the refusal above exists for. Reading the + // pair off `RowLayout` rather than writing `facet * 16 + 4` here keeps + // the row geometry in the one place that owns it. + let (off, stride) = store.layout.lo64_lane(store_entry.n_rows, facet); + kernels::simd_rowstore_ternary_match_mask( + store.as_bytes(), + off, + stride, + store_entry.n_rows as usize, + &pat, + &car, + &mut g.words, + ); + clear_tail_bits(&mut g.words, store_entry.n_rows); + LGJ_OK + }) +} + /// Sum one facet's 12-byte register, under a caller-supplied CARVING, over /// the rows a mask selects — the mask-native sweep (ABI minor >= 5, /// `docs/abi.md` §14). @@ -1421,6 +1714,55 @@ fn validate_plan(pattern: &ResourceEntry, ops: &[LgjOpDesc]) -> Result<(), i32> /// The body behind both `lgj_plan_eval` and `lgj_plan_eval_scalar` — one code /// path, two symbols, so the parity test compares two *paths* rather than a /// function against itself. +/// `mask_risc::ExecError` → an ABI status. +/// +/// **Every arm here is unreachable through the ABI**, and saying so is worth +/// more than implying otherwise. `validate_plan` runs first and rejects an +/// unknown opcode, a bad combine, an out-of-range lane and a kind mismatch +/// before the lowering is even built; the lowering names no input plane +/// (`Planes::masks` is `&[]`), no sum terminal and no blend. What is left — +/// the scratch-sizing family, `ScratchReadBeforeWrite`, `GateAliasesDst` — +/// would be a bug in THIS file, not in a caller's plan. +/// +/// So the map exists to turn such a bug into a status a caller can see +/// instead of a panic, and its arms are deliberately NOT claimed to be +/// individually falsifiable through the ABI. A future session looking for the +/// disable run that pins each arm will not find one; this is why. +fn exec_error_to_status(e: ExecError) -> i32 { + match e { + ExecError::LaneOutOfRange(_) => LGJ_ERR_INVALID_LANE, + ExecError::LaneKind { .. } => LGJ_ERR_LANE_KIND_MISMATCH, + ExecError::LenMismatch { .. } => LGJ_ERR_MASK_LENGTH_MISMATCH, + ExecError::PlaneOutOfRange(_) | ExecError::PlaneTail(_) => LGJ_ERR_INVALID_HANDLE, + ExecError::SumRowBound { .. } => LGJ_ERR_SUM_OVERFLOW, + ExecError::ScratchTooSmall { .. } + | ExecError::ScratchBufferTooSmall { .. } + | ExecError::ScratchSlotUndeclared { .. } + | ExecError::ScratchSlotsUnaddressable { .. } + | ExecError::ScratchReadBeforeWrite { .. } + | ExecError::ScratchWords { .. } + | ExecError::BlendNeedsOut + | ExecError::GateAliasesDst { .. } => LGJ_ERR_ALLOCATION_FAILED, + } +} + +/// Write the answer into `dst_mask` and `out_count` — the ONE place either is +/// written, reached only after every error has already returned. +fn publish(mask: &ResourceEntry, words: &[u64], count: u64, out_count: *mut u64) -> i32 { + let mut g = match mask.write_mask() { + Some(g) => g, + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + if g.words.len() != words.len() { + return LGJ_ERR_MASK_LENGTH_MISMATCH; + } + g.words.copy_from_slice(words); + drop(g); + // SAFETY: non-null (checked by the caller); written only on success. + unsafe { *out_count = count }; + LGJ_OK +} + fn plan_eval_impl( res: u64, ops: *const LgjOpDesc, @@ -1456,47 +1798,122 @@ fn plan_eval_impl( let n_rows = pattern.n_rows; let n_words = mask_words_for(n_rows) as usize; - // Accumulate into scratch, then publish. Two consequences, both wanted: - // dst_mask is written exactly once, and an error at any point leaves it - // byte-for-byte as it was. - let mut acc = vec![0u64; n_words]; - // "the accumulator starts as all rows set" (§7). - for w in acc.iter_mut() { - *w = u64::MAX; - } - clear_tail_bits(&mut acc, n_rows); - let mut scratch = vec![0u64; n_words]; + let lowered = match plan_lower::lower_plan(ops) { + Some(l) => l, + // Unreachable: `validate_plan` already rejected every unknown opcode. + // Reported rather than asserted — see `exec_error_to_status`. + None => return LGJ_ERR_UNKNOWN_OPCODE, + }; - for op in ops { - let lane = match lane_view(&pattern, op.lane_id) { - Ok(l) => l, - Err(e) => return e, - }; - if let Err(e) = - kernels::eval_predicate(path, op.op, op.operand, &lane, n_rows, &mut scratch) - { - return e; - } - if let Err(e) = kernels::combine_into(path, op.combine, &mut acc, &scratch) { - return e; + // The old loop allocated an accumulator and a predicate buffer per call + // and copied the accumulator into the destination at the end, so that + // `dst_mask` was written exactly once and an error anywhere left it + // byte-for-byte as it was. Both properties survive without the + // allocation: the accumulator now lives in the thread-local arena, and + // every error path below returns before `publish` is ever reached. + let program = match lowered { + Lowered::AllRows => { + // No program at all. An accumulator that starts all-ones and is + // only ever OR-ed cannot shrink, so the answer is every row. + let mut g = match mask.write_mask() { + Some(g) => g, + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + if g.words.len() != n_words { + return LGJ_ERR_MASK_LENGTH_MISMATCH; + } + for w in g.words.iter_mut() { + *w = u64::MAX; + } + clear_tail_bits(&mut g.words, n_rows); + drop(g); + // SAFETY: non-null (checked above); written only on this success path. + unsafe { *out_count = n_rows }; + return LGJ_OK; } - } - clear_tail_bits(&mut acc, n_rows); - let count = kernels::popcount(path, &acc); + Lowered::Program(p) => p, + }; - let mut g = match mask.write_mask() { - Some(g) => g, + let fixture = match pattern.fixture() { + Some(f) => f, None => return LGJ_ERR_WRONG_RESOURCE_KIND, }; - if g.words.len() != acc.len() { - return LGJ_ERR_MASK_LENGTH_MISMATCH; - } - g.words.copy_from_slice(&acc); - drop(g); + let rows = match usize::try_from(n_rows) { + Ok(r) => r, + Err(_) => return LGJ_ERR_LENGTH_OVERFLOW, + }; + // Index order is LANE_IDS, LANE_CLASSES, LANE_VALUES — pinned by the + // `const _` assertions at the top of this file, because the lowering uses + // `lane_id` directly as the index. + let lanes = [ + LaneRef::U64(fixture.ids()), + LaneRef::U32(fixture.classes()), + LaneRef::I32(fixture.values()), + ]; + // `masks` is EMPTY on purpose. Handing `dst_mask` in as an input plane + // would turn a caller's dirty prior tail into `ExecError::PlaneTail` — a + // spurious failure on a destination that is about to be overwritten + // wholesale. The destination is an output here and nothing else. + let planes = Planes { + n_rows: rows, + masks: &[], + lanes: &lanes, + }; - // SAFETY: non-null (checked above); written only on the success path. - unsafe { *out_count = count }; - LGJ_OK + match path { + Path::Simd => { + let need = match scratch_words_for(n_words, plan_lower::SLOTS as usize) { + Some(n) => n, + None => return LGJ_ERR_LENGTH_OVERFLOW, + }; + PLAN_SCRATCH.with(|cell| { + let mut buf = cell.borrow_mut(); + if buf.len() < need { + buf.resize(need, 0); + } + let mut scratch = + match Scratch::over(&mut buf[..need], n_words, plan_lower::SLOTS as usize) { + Ok(s) => s, + Err(e) => return exec_error_to_status(e), + }; + let count = match execute(&program, &planes, &mut scratch, None) { + Ok(Value::Count(c)) => c as u64, + // The lowering emits exactly one terminal and it is + // `Count`; any other value means this file built a + // program it did not intend to. + Ok(_) => return LGJ_ERR_ALLOCATION_FAILED, + Err(e) => return exec_error_to_status(e), + }; + let words = match scratch.slot(plan_lower::ACC_SLOT) { + Some(w) => w, + None => return LGJ_ERR_ALLOCATION_FAILED, + }; + publish(&mask, words, count, out_count) + }) + } + // Fork A: the scalar symbol runs mask-risc's row-at-a-time oracle, so + // what it proves is whole-evaluator-against-whole-evaluator rather + // than kernel-against-kernel. The oracle ALLOCATES by design — one + // `bool` per row per slot is exactly what lets it falsify a + // bit-packing bug — which is why the allocation gate names + // `lgj_plan_eval` and only it. + Path::Scalar => { + let count = match reference_execute(&program, &planes, None) { + Ok(Value::Count(c)) => c as u64, + Ok(_) => return LGJ_ERR_ALLOCATION_FAILED, + Err(e) => return exec_error_to_status(e), + }; + let slots = match reference_scratch(&program, &planes) { + Ok(s) => s, + Err(e) => return exec_error_to_status(e), + }; + let words = match slots.get(plan_lower::ACC_SLOT as usize) { + Some(w) => w.as_slice(), + None => return LGJ_ERR_ALLOCATION_FAILED, + }; + publish(&mask, words, count, out_count) + } + } } /// Evaluate `n_ops` predicates in **one** crossing. @@ -1606,6 +2023,96 @@ pub unsafe extern "C" fn lgj_reduce_sum_i32( }) } +/// The PARAMETERISED masked reduction over an `I32` lane — `sum` / `min` / +/// `max` behind ONE symbol (ABI minor >= 11, `docs/abi.md` §19). +/// +/// # This shape was mandated in advance +/// +/// `abi.md` §15 set the condition and the answer before the need arose: +/// *"if a second reduction is ever needed (min/max/count-distinct/histogram), +/// do NOT add a second symbol ... an op-code parameter on one reduce symbol, +/// mirroring how `lgj_plan_eval`'s `LgjOpDesc` already generalises predicates — +/// and `sum` becomes op-code 0. Two reduction symbols would be the smell §1 +/// warns about; one parameterised symbol is the shape this ABI already uses +/// elsewhere."* Min and max are that second reduction, so this is that symbol, +/// and `count-distinct` or `histogram` is a future op-code rather than a future +/// symbol. +/// +/// [`lgj_reduce_sum_i32`] is retained unchanged: removing it would not be an +/// additive minor, and a minor-1 Java is still entitled to it. `reduce_op = 0` +/// here answers identically — pinned by a test, so the two cannot drift. +/// +/// # `out_present` is not a formality +/// +/// Min and max over an EMPTY population have no answer, and every in-band +/// sentinel is wrong: `i32::MAX` is a value a real population can contain, and +/// `0` is worse. So `out_present` carries the distinction out — `0` means +/// "nothing selected; `*out_value` is 0 and means nothing". `LGJ_REDUCE_SUM` is +/// always present, because the sum of an empty population is `0` and that IS +/// the answer. +/// +/// Both outputs are written on `LGJ_OK` and neither on any failure, matching +/// [`lgj_reduce_facet_sum_resolved`]'s two-output precedent. +/// +/// An unknown `reduce_op` is `UNKNOWN_OPCODE`, never a silent fall back to +/// `SUM` — an unknown reduction must not alias a known one, exactly as §14 +/// argues for an unknown carving. +/// +/// # Bulk (§6) +/// +/// `O(mask_words + popcount)` for every op — the mask-word scan is +/// unconditional, so an empty mask costs one pass rather than nothing, the same +/// cost shape §14 states for the register sweep. +/// +/// # Safety +/// +/// A null `out_value`/`out_present` is *handled*, not UB: `NULL_ARGUMENT`. +/// Otherwise each must point at one writable, correctly-aligned slot (`i64` +/// and `u32`); both are written only on `LGJ_OK`. +#[no_mangle] +pub unsafe extern "C" fn lgj_reduce_i32( + res: u64, + lane_id: u32, + reduce_op: u32, + mask: u64, + out_value: *mut i64, + out_present: *mut u32, +) -> i32 { + guard(|| { + if out_value.is_null() || out_present.is_null() { + return LGJ_ERR_NULL_ARGUMENT; + } + let (pattern, maskr) = match resolve_pattern_and_mask(res, mask) { + Ok(t) => t, + Err(e) => return e, + }; + let lane = match lane_view(&pattern, lane_id) { + Ok(l) => l, + Err(e) => return e, + }; + let values = match lane { + LaneView::I32(v) => v, + _ => return LGJ_ERR_LANE_KIND_MISMATCH, + }; + let g = match maskr.read_mask() { + Some(g) => g, + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + let answer = kernels::reduce_i32(Path::Simd, reduce_op, values, &g.words); + drop(g); + let answer = match answer { + Ok(v) => v, + Err(e) => return e, + }; + // SAFETY: both non-null, checked above; written only on success. + unsafe { + *out_value = answer.unwrap_or(0); + *out_present = u32::from(answer.is_some()); + } + LGJ_OK + }) +} + // ─────────────────────────────────────────────────────────────────────────── // Graph traversal (ABI minor ≥ 4) — the first symbol gated by the // lance-graph-contract ClassView/FieldMask LAW (docs/abi.md §13). @@ -1862,6 +2369,19 @@ mod tests { use super::*; use crate::fixture::{Fixture, LANE_CLASSES, LANE_IDS, LANE_VALUES}; + /// The lowering differential: this crate's `plan_lower` (a flat op list) + /// against `lance-graph-quack`'s `lower` (a Boolean tree), both producing + /// a `mask_risc::Program`. Two implementations of one law; until this + /// file, nothing compared them. + mod lowering_convergence; + mod pr4_dst_reuse; + /// PR4's behaviour-equivalence oracle and falsifier matrix — the frozen + /// pre-PR4 loop, kept in its own file so the freeze is visible as a file + /// boundary rather than as a convention inside a 2000-line module. + mod pr4_equivalence; + mod pr4_matrix; + mod pr4_seed; + /// Safe wrappers over the pointer-taking exports. /// /// These are not a second implementation: each one is a direct call into the @@ -2431,11 +2951,30 @@ mod tests { } } - /// The headline parity property, through the same code path the Java tests - /// exercise: SIMD and the independent scalar reference must agree exactly, - /// including at row counts that are not multiples of 64. + /// FAILS IF: the bit-packed executor and the row-at-a-time oracle + /// disagree on the same lowered plan, including at row counts that are + /// not multiples of 64. + /// + /// **Renamed in PR4, and the rename is mandatory rather than cosmetic.** + /// It was `simd_and_scalar_plans_agree_bit_for_bit`, and under the old + /// implementation that name was accurate: `lgj_plan_eval_scalar` ran + /// `kernels::scalar_*`, so the two symbols were two BACKENDS of one + /// evaluator and this test was backend parity. + /// + /// It is not that any more. Both symbols now lower the same plan and the + /// scalar one runs `mask_risc::reference_execute` — an oracle that + /// evaluates one ROW at a time in plain Rust with no SIMD facade anywhere + /// in its file, a property a grep test in that crate enforces. So what + /// this compares is executor-against-oracle, which is strictly stronger + /// (whole evaluator, not kernel) and is a DIFFERENT claim. + /// + /// Backend parity itself is `ndarray::simd`'s business — an ABI-level + /// test could only reach it through a proxy — so it is not lost here; it + /// was never really held here. A test whose name claims a property it no + /// longer checks is worse than no test, because the next session reads + /// the name and not the body. #[test] - fn simd_and_scalar_plans_agree_bit_for_bit() { + fn the_executor_and_the_row_oracle_agree_bit_for_bit() { for n in [0u64, 1, 63, 64, 65, 127, 1000, 4097] { for seed in [0u64, 7, 0xFEED_FACE] { let p = open(n, seed); @@ -3397,4 +3936,702 @@ mod tests { assert_eq!(unsafe { lgj_rowstore_open(n, seed, &mut h) }, LGJ_OK); h } + + // ═══════════════════════════════════════════════════════════════════════ + // ABI minor 11 — the masking-op completion, through the membrane + // (docs/abi.md §19). Each symbol gets: a CAN-FIRE arm, a CAN-STAY-SILENT + // arm, and the handle/status arms every export in this ABI owes. + // ═══════════════════════════════════════════════════════════════════════ + + mod call11 { + use super::*; + + pub fn op_ternary_match( + res: u64, + facet: u32, + pattern: *const u8, + care: *const u8, + dst: u64, + ) -> i32 { + unsafe { lgj_op_ternary_match(res, facet, pattern, care, dst) } + } + pub fn reduce_i32( + res: u64, + lane: u32, + reduce_op: u32, + m: u64, + out_value: *mut i64, + out_present: *mut u32, + ) -> i32 { + unsafe { lgj_reduce_i32(res, lane, reduce_op, m, out_value, out_present) } + } + } + + /// Pack a `LGJ_OP_TERNARY_MATCH_U32` operand the way `abi.rs` documents: + /// pattern in the low 32 bits, care in the high 32. ONE spelling in the + /// tests, so a hand-packing typo cannot make a passing arm mean something + /// other than it reads. + fn tcam_operand(pattern: u32, care: u32) -> i64 { + ((u64::from(care) << 32) | u64::from(pattern)) as i64 + } + + /// Bit-serial truth-table oracle — deliberately NOT the kernel, so the two + /// cannot agree by sharing code. + fn ternlog_oracle(imm: u8, a: &[u64], b: &[u64], c: &[u64], n_rows: u64) -> Vec { + let mut out = vec![0u64; a.len()]; + for w in 0..a.len() { + for bit in 0..64u32 { + let row = (w as u64) * 64 + u64::from(bit); + if row >= n_rows { + continue; // the tail is normative zero + } + let idx = + (((a[w] >> bit) & 1) << 2) | (((b[w] >> bit) & 1) << 1) | ((c[w] >> bit) & 1); + if (u64::from(imm) >> idx) & 1 == 1 { + out[w] |= 1u64 << bit; + } + } + } + out + } + + // ── lgj_mask_ternlog ─────────────────────────────────────────────────── + + /// CAN-FIRE, across the whole immediate space AND every aliasing shape. + /// + /// The aliasing matrix is the point: with FOUR handles there are more + /// overlap cases than `lgj_mask_andnot`'s three could be enumerated into, + /// which is why the implementation snapshots instead of case-analysing. + /// Every shape must give the value each operand had BEFORE the call. + #[test] + fn mask_ternlog_answers_the_truth_table_in_every_aliasing_shape() { + let n = 130u64; // three words, the last one with 2 live bits + let p = open(n, 7); + let (a, b, c, d) = ( + mask(p, LGJ_MASK_INIT_EMPTY), + mask(p, LGJ_MASK_INIT_EMPTY), + mask(p, LGJ_MASK_INIT_EMPTY), + mask(p, LGJ_MASK_INIT_EMPTY), + ); + let va = [0x0F0F_0F0F_0F0F_0F0Fu64, 0xFFFF_0000_FFFF_0000, 0b01]; + let vb = [0x3333_3333_3333_3333u64, 0x00FF_00FF_00FF_00FF, 0b10]; + let vc = [0x5555_5555_5555_5555u64, 0xF0F0_F0F0_F0F0_F0F0, 0b11]; + + // Every immediate, distinct destination. + for imm in 0u8..=255 { + set_words(a, &va); + set_words(b, &vb); + set_words(c, &vc); + assert_eq!(lgj_mask_ternlog(a, b, c, d, imm), LGJ_OK); + assert_eq!( + read_words(d), + ternlog_oracle(imm, &va, &vb, &vc, n), + "imm={imm:#04x}, distinct dst" + ); + } + + // dst aliases each operand in turn, plus the all-same degenerate case. + // MAJ3 is used because it depends on all three operands, so an + // implementation that silently dropped one would be caught. + const MAJ3: u8 = 0xE8; + for (label, dst) in [("dst==a", a), ("dst==b", b), ("dst==c", c)] { + set_words(a, &va); + set_words(b, &vb); + set_words(c, &vc); + assert_eq!(lgj_mask_ternlog(a, b, c, dst, MAJ3), LGJ_OK); + assert_eq!( + read_words(dst), + ternlog_oracle(MAJ3, &va, &vb, &vc, n), + "{label}: operands must be read as they were BEFORE the call" + ); + } + // a == b == c == dst: majority of x with itself is x. + set_words(a, &va); + assert_eq!(lgj_mask_ternlog(a, a, a, a, MAJ3), LGJ_OK); + assert_eq!(read_words(a), va.to_vec(), "maj(x,x,x) == x"); + + for h in [a, b, c, d, p] { + assert_eq!(lgj_close(h), LGJ_OK); + } + } + + /// CAN-STAY-SILENT: the tail. An ODD immediate is true of `(0,0,0)`, so + /// every bit past `n_rows` is set by the kernel and must be cleared before + /// return — otherwise `lgj_mask_count` reads rows that do not exist. + /// + /// Two-sided, and the second half is what makes it a falsifier rather than + /// a re-statement: `NOT` of a mask must count `n_rows - popcount`, a number + /// a leaked tail cannot produce. + #[test] + fn an_odd_immediate_leaves_no_tail_bit_behind() { + const NOT_A: u8 = 0x0F; + let n = 130u64; + let p = open(n, 11); + let (a, d) = (mask(p, LGJ_MASK_INIT_EMPTY), mask(p, LGJ_MASK_INIT_EMPTY)); + set_words(a, &[u64::MAX, 0, 0b01]); + let before = count(a); + assert_eq!(before, 65, "fixture must be non-trivial: 64 + 1 rows set"); + + assert_eq!(lgj_mask_ternlog(a, a, a, d, NOT_A), LGJ_OK); + assert_eq!( + count(d), + n - before, + "NOT must partition the population — a leaked tail would over-count" + ); + let w = read_words(d); + assert_eq!(w.last().unwrap() >> (n % 64), 0, "tail bits survived"); + + for h in [a, d, p] { + assert_eq!(lgj_close(h), LGJ_OK); + } + } + + /// The handle/status arms. Stale, fabricated, cross-parent and + /// parent-closed each answer with a status — never a crash, never silently. + #[test] + fn mask_ternlog_rejects_every_incompatible_handle_shape() { + let p = open(128, 1); + let q = open(128, 2); // same size, DIFFERENT parent + let (a, b) = (mask(p, LGJ_MASK_INIT_ALL), mask(p, LGJ_MASK_INIT_ALL)); + let other = mask(q, LGJ_MASK_INIT_ALL); + let short = { + let r = open(64, 3); + let m = mask(r, LGJ_MASK_INIT_ALL); + (r, m) + }; + + for bogus in [0u64, 1, 0xDEAD_BEEF, u64::MAX] { + assert_eq!( + lgj_mask_ternlog(bogus, a, b, a, 0xC0), + LGJ_ERR_INVALID_HANDLE + ); + assert_eq!( + lgj_mask_ternlog(a, bogus, b, a, 0xC0), + LGJ_ERR_INVALID_HANDLE + ); + assert_eq!( + lgj_mask_ternlog(a, b, bogus, a, 0xC0), + LGJ_ERR_INVALID_HANDLE + ); + assert_eq!( + lgj_mask_ternlog(a, b, a, bogus, 0xC0), + LGJ_ERR_INVALID_HANDLE + ); + } + // A pattern handle where a mask is required. + assert_eq!( + lgj_mask_ternlog(p, a, b, a, 0xC0), + LGJ_ERR_WRONG_RESOURCE_KIND + ); + // Same row count, different parent — "a different population wearing + // the right size", the same rejection lgj_mask_and gives. + assert_eq!( + lgj_mask_ternlog(a, other, b, a, 0xC0), + LGJ_ERR_MASK_LENGTH_MISMATCH + ); + // Different row count. + assert_eq!( + lgj_mask_ternlog(a, short.1, b, a, 0xC0), + LGJ_ERR_MASK_LENGTH_MISMATCH + ); + // Parent closed. + assert_eq!(lgj_close(q), LGJ_OK); + assert_eq!( + lgj_mask_ternlog(other, other, other, other, 0xC0), + LGJ_ERR_PARENT_CLOSED + ); + + for h in [a, b, other, short.1, short.0, p] { + let _ = lgj_close(h); + } + } + + // ── lgj_op_ternary_match ─────────────────────────────────────────────── + + /// CAN-FIRE + CAN-STAY-SILENT over a real row store, through the membrane. + /// + /// The three arms are the TCAM's whole contract: `care == 0` matches every + /// row (total), the exact register matches at least its own row, and + /// flipping a CARED byte of that same pattern drops it while flipping an + /// UNCARED one does not. The third arm is what stops a "compare the whole + /// register" implementation from passing. + #[test] + fn ternary_match_honours_care_and_is_two_sided_through_the_membrane() { + use crate::rowstore::{FACET_BYTES, FACET_CLASSID_BYTES, ROW_BYTES}; + const N: usize = kernels::FACET_REGISTER_BYTES; + let n = 300u64; + let store = lgj_rowstore_open_handle(n, 0xBEEF); + let m = mask(store, LGJ_MASK_INIT_EMPTY); + let facet = 5u32; + + // Row 0's own register, read through the registry (the same back door + // `set_words` uses) so the pattern is guaranteed to have a hit. + let pattern: [u8; N] = { + let e = registry::resolve(store).unwrap(); + let rs = e.rowstore().unwrap(); + let off = (u64::from(facet) * FACET_BYTES + FACET_CLASSID_BYTES) as usize; + rs.as_bytes()[off..off + N].try_into().unwrap() + }; + + // care == 0: every row matches. Total, not an error. + assert_eq!( + call11::op_ternary_match(store, facet, pattern.as_ptr(), [0u8; N].as_ptr(), m), + LGJ_OK + ); + assert_eq!(count(m), n, "care == 0 must match every row"); + + // Exact: at least row 0, and strictly fewer than all — otherwise the + // predicate is not discriminating and the arms below prove nothing. + assert_eq!( + call11::op_ternary_match(store, facet, pattern.as_ptr(), [0xFFu8; N].as_ptr(), m), + LGJ_OK + ); + let exact = count(m); + assert!(exact >= 1, "row 0 must match its own register"); + assert!( + exact < n, + "an exact register match must not select everything" + ); + assert_eq!(read_words(m)[0] & 1, 1, "and row 0 specifically"); + + // Flip a byte, CARE about it → row 0 drops out. + let mut flipped = pattern; + flipped[0] ^= 0xFF; + assert_eq!( + call11::op_ternary_match(store, facet, flipped.as_ptr(), [0xFFu8; N].as_ptr(), m), + LGJ_OK + ); + assert_eq!( + read_words(m)[0] & 1, + 0, + "a cared-byte mismatch must exclude" + ); + + // Same flip, DON'T care about it → row 0 comes back. + let mut care_rest = [0xFFu8; N]; + care_rest[0] = 0; + assert_eq!( + call11::op_ternary_match(store, facet, flipped.as_ptr(), care_rest.as_ptr(), m), + LGJ_OK + ); + assert_eq!( + read_words(m)[0] & 1, + 1, + "an uncared byte must not exclude — this is what makes it a TCAM" + ); + + // The register really is 12 bytes at facet*16+4, stride 512: pin the + // geometry the export reads off RowLayout rather than trusting it. + assert_eq!(N, 12); + assert_eq!(ROW_BYTES, 512); + + assert_eq!(lgj_close(m), LGJ_OK); + assert_eq!(lgj_close(store), LGJ_OK); + } + + /// Every rejection path, and the layout gate PINNED TWO-SIDED: the same + /// call succeeds on AoS and refuses on facet-major, so the gate is proven + /// to discriminate rather than merely to exist. + #[test] + fn ternary_match_rejects_null_bad_facet_wrong_kind_and_a_facet_major_layout() { + const N: usize = kernels::FACET_REGISTER_BYTES; + let n = 128u64; + let aos = lgj_rowstore_open_handle(n, 5); + let m = mask(aos, LGJ_MASK_INIT_EMPTY); + let pat = [0u8; N]; + let care = [0u8; N]; + + // Null pointers are HANDLED, not UB. + assert_eq!( + call11::op_ternary_match(aos, 0, std::ptr::null(), care.as_ptr(), m), + LGJ_ERR_NULL_ARGUMENT + ); + assert_eq!( + call11::op_ternary_match(aos, 0, pat.as_ptr(), std::ptr::null(), m), + LGJ_ERR_NULL_ARGUMENT + ); + // Facet out of range. + assert_eq!( + call11::op_ternary_match( + aos, + crate::rowstore::ROW_FACETS, + pat.as_ptr(), + care.as_ptr(), + m + ), + LGJ_ERR_INVALID_LANE + ); + // A pattern resource where a row store is required. + let p = open(n, 1); + assert_eq!( + call11::op_ternary_match(p, 0, pat.as_ptr(), care.as_ptr(), m), + LGJ_ERR_WRONG_RESOURCE_KIND + ); + // A mask of the wrong row count. + let small = open(64, 1); + let small_m = mask(small, LGJ_MASK_INIT_EMPTY); + assert_eq!( + call11::op_ternary_match(aos, 0, pat.as_ptr(), care.as_ptr(), small_m), + LGJ_ERR_MASK_LENGTH_MISMATCH + ); + // Fabricated handles. + for bogus in [0u64, 0xDEAD_BEEF, u64::MAX] { + assert_eq!( + call11::op_ternary_match(bogus, 0, pat.as_ptr(), care.as_ptr(), m), + LGJ_ERR_INVALID_HANDLE + ); + assert_eq!( + call11::op_ternary_match(aos, 0, pat.as_ptr(), care.as_ptr(), bogus), + LGJ_ERR_INVALID_HANDLE + ); + } + + // THE TWO-SIDED LAYOUT GATE. Same arguments, same facet, same care. + assert_eq!( + call11::op_ternary_match(aos, 0, pat.as_ptr(), care.as_ptr(), m), + LGJ_OK, + "AosRows must succeed — or the refusal below proves nothing" + ); + let mut col = 0u64; + assert_eq!( + unsafe { lgj_rowstore_open_columnar(n, 5, 0, 0x0, 25, &mut col) }, + LGJ_OK + ); + let col_m = mask(col, LGJ_MASK_INIT_EMPTY); + assert_eq!( + call11::op_ternary_match(col, 0, pat.as_ptr(), care.as_ptr(), col_m), + LGJ_ERR_UNSUPPORTED_LAYOUT, + "a facet-major store splits the register; refuse rather than gather" + ); + + for h in [m, small_m, small, p, aos, col_m, col] { + let _ = lgj_close(h); + } + } + + // ── lgj_reduce_i32 ───────────────────────────────────────────────────── + + /// CAN-FIRE: each reduction answers, `SUM` agrees BIT-FOR-BIT with the + /// symbol it generalises, and the mask genuinely selects (a reduction over + /// a subset differs from one over everything). + #[test] + fn reduce_i32_answers_every_op_and_agrees_with_the_symbol_it_generalises() { + let n = 1000u64; + let p = open(n, 21); + let all = mask(p, LGJ_MASK_INIT_ALL); + let some = mask(p, LGJ_MASK_INIT_EMPTY); + set_rows(some, &[0, 1, 2, 500, 999]); + + let mut v = 0i64; + let mut present = 0u32; + + // SUM through the new symbol == SUM through the old one, on both masks. + for m in [all, some] { + assert_eq!( + call11::reduce_i32(p, LANE_VALUES, LGJ_REDUCE_SUM, m, &mut v, &mut present), + LGJ_OK + ); + assert_eq!(present, 1, "a sum is always present"); + let mut legacy = 0i64; + assert_eq!(call::reduce_sum_i32(p, LANE_VALUES, m, &mut legacy), LGJ_OK); + assert_eq!(v, legacy, "reduce_op 0 must BE lgj_reduce_sum_i32"); + } + + // MIN/MAX bracket the sum's population, and the subset's bracket is + // inside the full one — so the mask is demonstrably load-bearing. + let read = |m: u64, op: u32| { + let (mut x, mut pr) = (0i64, 0u32); + assert_eq!( + call11::reduce_i32(p, LANE_VALUES, op, m, &mut x, &mut pr), + LGJ_OK + ); + (x, pr) + }; + let (min_all, p1) = read(all, LGJ_REDUCE_MIN); + let (max_all, p2) = read(all, LGJ_REDUCE_MAX); + let (min_some, p3) = read(some, LGJ_REDUCE_MIN); + let (max_some, p4) = read(some, LGJ_REDUCE_MAX); + assert_eq!((p1, p2, p3, p4), (1, 1, 1, 1)); + assert!(min_all < max_all, "the fixture must not be constant"); + assert!(min_all <= min_some && max_some <= max_all); + assert!( + min_all < min_some || max_some < max_all, + "a 5-row subset of 1000 must be strictly narrower somewhere" + ); + + for h in [all, some, p] { + assert_eq!(lgj_close(h), LGJ_OK); + } + } + + /// CAN-STAY-SILENT: an empty population has NO min and NO max, and that is + /// reported rather than encoded as a sentinel. The paired half — `SUM` is + /// present over the same empty mask — is what stops "always absent" from + /// passing. + #[test] + fn reduce_i32_reports_an_empty_population_instead_of_inventing_an_extremum() { + let p = open(256, 4); + let empty = mask(p, LGJ_MASK_INIT_EMPTY); + let (mut v, mut present) = (-1i64, 9u32); + + for op in [LGJ_REDUCE_MIN, LGJ_REDUCE_MAX] { + assert_eq!( + call11::reduce_i32(p, LANE_VALUES, op, empty, &mut v, &mut present), + LGJ_OK + ); + assert_eq!(present, 0, "no rows selected ⇒ no extremum"); + assert_eq!(v, 0, "and the value slot carries nothing meaningful"); + } + assert_eq!( + call11::reduce_i32(p, LANE_VALUES, LGJ_REDUCE_SUM, empty, &mut v, &mut present), + LGJ_OK + ); + assert_eq!((present, v), (1, 0), "the sum of nothing IS zero"); + + assert_eq!(lgj_close(empty), LGJ_OK); + assert_eq!(lgj_close(p), LGJ_OK); + } + + /// Rejections: unknown reduction, wrong lane kind, null outputs, bad + /// handles — and NOTHING is written on any of them. + #[test] + fn reduce_i32_rejects_an_unknown_reduction_rather_than_defaulting_to_sum() { + let p = open(128, 6); + let m = mask(p, LGJ_MASK_INIT_ALL); + let (mut v, mut present) = (0xAAAA_AAAAi64, 0xAAAA_AAAAu32); + + for bogus in [3u32, 4, 99, u32::MAX] { + assert_eq!( + call11::reduce_i32(p, LANE_VALUES, bogus, m, &mut v, &mut present), + LGJ_ERR_UNKNOWN_OPCODE + ); + } + // A U32 lane cannot be reduced as I32. + assert_eq!( + call11::reduce_i32(p, LANE_CLASSES, LGJ_REDUCE_MIN, m, &mut v, &mut present), + LGJ_ERR_LANE_KIND_MISMATCH + ); + assert_eq!( + call11::reduce_i32(p, LANE_IDS, LGJ_REDUCE_MIN, m, &mut v, &mut present), + LGJ_ERR_LANE_KIND_MISMATCH + ); + // Out of range lane. + assert_eq!( + call11::reduce_i32(p, 99, LGJ_REDUCE_MIN, m, &mut v, &mut present), + LGJ_ERR_INVALID_LANE + ); + // Nulls are handled, not UB. + assert_eq!( + call11::reduce_i32( + p, + LANE_VALUES, + LGJ_REDUCE_MIN, + m, + std::ptr::null_mut(), + &mut present + ), + LGJ_ERR_NULL_ARGUMENT + ); + assert_eq!( + call11::reduce_i32( + p, + LANE_VALUES, + LGJ_REDUCE_MIN, + m, + &mut v, + std::ptr::null_mut() + ), + LGJ_ERR_NULL_ARGUMENT + ); + for bogus in [0u64, 0xDEAD_BEEF, u64::MAX] { + assert_eq!( + call11::reduce_i32(bogus, LANE_VALUES, LGJ_REDUCE_MIN, m, &mut v, &mut present), + LGJ_ERR_INVALID_HANDLE + ); + } + // Not one of those calls wrote either output. + assert_eq!((v, present), (0xAAAA_AAAA, 0xAAAA_AAAA), "out_* on failure"); + + assert_eq!(lgj_close(m), LGJ_OK); + assert_eq!(lgj_close(p), LGJ_OK); + } + + // ── the seven op-codes ───────────────────────────────────────────────── + + /// The new predicates through BOTH plan symbols, and the two must agree — + /// which is the parity escape hatch's whole reason to exist, extended to + /// the ops that were added to it. An op-code with no scalar arm would + /// answer `UNKNOWN_OPCODE` here and fail. + #[test] + fn every_minor_11_opcode_runs_on_both_plan_paths_and_they_agree() { + let n = 1000u64; + let p = open(n, 33); + let (d1, d2) = (mask(p, LGJ_MASK_INIT_EMPTY), mask(p, LGJ_MASK_INIT_EMPTY)); + + let cases: [(u32, u32, i64); 7] = [ + (LGJ_OP_NE_U32, LANE_CLASSES, 7), + (LGJ_OP_EQ_I32, LANE_VALUES, 100), + (LGJ_OP_NE_I32, LANE_VALUES, 100), + (LGJ_OP_LT_I32, LANE_VALUES, 100), + (LGJ_OP_LE_I32, LANE_VALUES, 100), + (LGJ_OP_GE_I32, LANE_VALUES, 100), + // pattern 5, care 0b111 — a genuine TCAM, not an exact compare. + ( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + tcam_operand(5, 0b111), + ), + ]; + for (opcode, lane, operand) in cases { + let ops = [op(opcode, lane, operand, LGJ_COMBINE_AND)]; + let (mut c1, mut c2) = (0u64, 0u64); + assert_eq!( + call::plan_eval(p, ops.as_ptr(), 1, d1, &mut c1), + LGJ_OK, + "op {opcode}" + ); + assert_eq!( + call::plan_eval_scalar(p, ops.as_ptr(), 1, d2, &mut c2), + LGJ_OK, + "op {opcode} has no scalar arm — the parity hatch is hollow" + ); + assert_eq!( + read_words(d1), + read_words(d2), + "SIMD/scalar parity, op {opcode}" + ); + assert_eq!(c1, c2); + // Anti-vacuity: each predicate must select SOMETHING and not + // EVERYTHING, or the parity above compares two empty answers. + assert!( + c1 > 0 && c1 < n, + "op {opcode} selected {c1} of {n} — not discriminating" + ); + } + + // A complementary pair must partition the population exactly. This is + // the arm that catches a `lt` lowered as a buggy complement of `ge`. + let mut lt = 0u64; + let mut ge = 0u64; + let a = [op(LGJ_OP_LT_I32, LANE_VALUES, 100, LGJ_COMBINE_AND)]; + let b = [op(LGJ_OP_GE_I32, LANE_VALUES, 100, LGJ_COMBINE_AND)]; + assert_eq!(call::plan_eval(p, a.as_ptr(), 1, d1, &mut lt), LGJ_OK); + assert_eq!(call::plan_eval(p, b.as_ptr(), 1, d2, &mut ge), LGJ_OK); + assert_eq!(lt + ge, n, "< and >= must partition"); + + // …and so must eq/ne, on a DIFFERENT lane kind. + let mut eq = 0u64; + let mut ne = 0u64; + let c = [op(LGJ_OP_EQ_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND)]; + let d = [op(LGJ_OP_NE_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND)]; + assert_eq!(call::plan_eval(p, c.as_ptr(), 1, d1, &mut eq), LGJ_OK); + assert_eq!(call::plan_eval(p, d.as_ptr(), 1, d2, &mut ne), LGJ_OK); + assert_eq!(eq + ne, n, "== and != must partition"); + + for h in [d1, d2, p] { + assert_eq!(lgj_close(h), LGJ_OK); + } + } + + /// The op-code/lane validator must reject every new op against a lane of + /// the wrong kind BEFORE anything is written — the same "validate the whole + /// plan first" property the existing ops have. + #[test] + fn every_minor_11_opcode_rejects_a_lane_of_the_wrong_kind() { + let p = open(128, 2); + let d = mask(p, LGJ_MASK_INIT_ALL); + let before = read_words(d); + + // U32 ops against the I32 lane, and I32 ops against the U32 lane. + for (opcode, wrong_lane) in [ + (LGJ_OP_NE_U32, LANE_VALUES), + (LGJ_OP_TERNARY_MATCH_U32, LANE_VALUES), + (LGJ_OP_EQ_I32, LANE_CLASSES), + (LGJ_OP_NE_I32, LANE_CLASSES), + (LGJ_OP_LT_I32, LANE_IDS), + (LGJ_OP_LE_I32, LANE_CLASSES), + (LGJ_OP_GE_I32, LANE_IDS), + ] { + let ops = [op(opcode, wrong_lane, 0, LGJ_COMBINE_AND)]; + let mut c = 0u64; + assert_eq!( + call::plan_eval(p, ops.as_ptr(), 1, d, &mut c), + LGJ_ERR_LANE_KIND_MISMATCH, + "op {opcode} on lane {wrong_lane}" + ); + } + // An op-code past the known set is still unknown. + for bogus in [10u32, 99, u32::MAX] { + let ops = [op(bogus, LANE_CLASSES, 0, LGJ_COMBINE_AND)]; + let mut c = 0u64; + assert_eq!( + call::plan_eval(p, ops.as_ptr(), 1, d, &mut c), + LGJ_ERR_UNKNOWN_OPCODE + ); + } + assert_eq!(read_words(d), before, "a rejected plan must write nothing"); + + assert_eq!(lgj_close(d), LGJ_OK); + assert_eq!(lgj_close(p), LGJ_OK); + } + + /// The TCAM op-code's packed operand is the one place this minor gives an + /// existing field a new reading, so it gets its own pin: `care` all-ones + /// must reproduce `LGJ_OP_EQ_U32` exactly, and `care` all-zero must select + /// every row. Between them they prove both halves of the packing land + /// where the documentation says. + #[test] + fn the_tcam_opcodes_packed_operand_decodes_both_halves() { + let n = 512u64; + let p = open(n, 77); + let (d1, d2) = (mask(p, LGJ_MASK_INIT_EMPTY), mask(p, LGJ_MASK_INIT_EMPTY)); + let (mut c1, mut c2) = (0u64, 0u64); + + // care = u32::MAX (high half) ⇒ exact equality with pattern 7. + let exact = tcam_operand(7, u32::MAX); + let a = [op( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + exact, + LGJ_COMBINE_AND, + )]; + let b = [op(LGJ_OP_EQ_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND)]; + assert_eq!(call::plan_eval(p, a.as_ptr(), 1, d1, &mut c1), LGJ_OK); + assert_eq!(call::plan_eval(p, b.as_ptr(), 1, d2, &mut c2), LGJ_OK); + assert_eq!(read_words(d1), read_words(d2), "care=MAX must BE eq_u32"); + assert!( + c1 > 0, + "the fixture must contain class 7 or this proves nothing" + ); + + // care = 0 ⇒ everything matches, whatever the pattern is. + let anything = tcam_operand(0xDEAD_BEEF, 0); + let c = [op( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + anything, + LGJ_COMBINE_AND, + )]; + assert_eq!(call::plan_eval(p, c.as_ptr(), 1, d1, &mut c1), LGJ_OK); + assert_eq!(c1, n, "care = 0 is 'don't care about anything'"); + + // A partial care mask is strictly between the two — the property that + // makes this a TCAM and not a renamed equality. + let partial = tcam_operand(0b1000, 0b1000); + let e = [op( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + partial, + LGJ_COMBINE_AND, + )]; + assert_eq!(call::plan_eval(p, e.as_ptr(), 1, d1, &mut c1), LGJ_OK); + assert!( + c1 > c2 && c1 < n, + "a partial care set must be strictly between" + ); + + for h in [d1, d2, p] { + assert_eq!(lgj_close(h), LGJ_OK); + } + } } diff --git a/native/lgj-abi/src/exports/tests/lowering_convergence.rs b/native/lgj-abi/src/exports/tests/lowering_convergence.rs new file mode 100644 index 0000000..f4a2631 --- /dev/null +++ b/native/lgj-abi/src/exports/tests/lowering_convergence.rs @@ -0,0 +1,646 @@ +//! The differential between this crate's own plan lowering and +//! `lance-graph-quack`'s — two lowerings of the SAME LAW, compared directly +//! for the first time. +//! +//! [`crate::plan_lower`] lowers a FLAT OP LIST (`&[LgjOpDesc]`, a left fold +//! seeded with all-ones and folded with `&=`/`|=`) to a `mask_risc::Program`. +//! `lance_graph_quack::lower` lowers a Boolean [`Filter`] TREE to the same +//! `Program` IR. Read side by side, both implement one algebra: the +//! accumulator starts all-ones; an AND narrows it and gates every later +//! comparison under whatever survives so far; an OR widens it and must NOT +//! be gated, because `acc | p` depends on `p` exactly where `acc` is zero — +//! gating there would discard precisely the bits an OR exists to admit; and +//! any op sitting before the plan's first AND is dead, since +//! `all_ones | p == all_ones` regardless of what `p` selects. `plan_lower` +//! drops the dead prefix by scanning for the least AND index; the tree +//! reading below drops it for free, because [`fold_to_tree`]'s fold never +//! establishes an accumulator until the first AND arrives — an `Or` +//! encountered with no accumulator yet is not merely un-selective, it is +//! genuinely absent from the tree. +//! +//! Each side already has its OWN oracle, and neither can see the other +//! drift: `plan_lower`'s is `pr4_equivalence.rs`'s frozen pre-PR4 loop, a +//! per-op accumulate-and-combine that never builds a `Program` at all; +//! `lance-graph-quack`'s is its own per-row reading, in its own crate, that +//! has never heard of `plan_lower`. A defect the two lowerings happen to +//! SHARE — the same wrong reading of the AND/OR asymmetry, say — would pass +//! both of those oracles and would ALSO pass this file, since "the two +//! lowerings agree with each other" is a different claim from "either one +//! agrees with its own oracle". What this file adds is orthogonal to both: +//! it is the only place that would notice the two independently-written +//! lowerings have started to DISAGREE with each other, which a same-crate +//! refactor in either one could otherwise introduce with nothing here to +//! catch it. +//! +//! This file is therefore a DIFFERENTIAL between the two lowerings, and +//! explicitly NOT a delegation from one to the other — the two sections +//! below say what that means and what it costs. +//! +//! # What this does NOT prove +//! +//! This is a differential over ANSWERS — row counts a `Program` selects — +//! never over the emitted op lists themselves, so two lowerings that reach +//! the same count via structurally different `Program`s (different slot +//! counts, different op ordering) are indistinguishable here, and that is +//! fine: nothing in this file claims otherwise. It also cannot, by +//! construction, catch a defect both lowerings share (see above). Neither +//! gap is a hole in THIS file specifically — `pr4_equivalence.rs`'s frozen +//! loop and `lance-graph-quack`'s own per-row oracle each independently pin +//! their own side's correctness against ground truth, and the three arms +//! (plan_lower-vs-frozen-loop, quack-vs-its-own-oracle, and +//! plan_lower-vs-quack here) are complementary, not redundant. +//! +//! # A differential, deliberately NOT a delegation +//! +//! This crate could, in principle, retire [`crate::plan_lower`] entirely and +//! have `plan_eval_impl` build [`fold_to_tree`] and hand it to +//! `lance_graph_quack::lower`. It does not, and `lance-graph-quack` is wired +//! as a DEV-dependency (see this crate's `Cargo.toml`, whose own comment +//! says so) specifically to make that impossible by construction: the +//! membrane (`lgj-abi`, which ships as the `cdylib` Java's `Linker` loads) +//! must not depend on a CONSUMER of the IR it serves — that dependency +//! points the wrong way. Sharing the LAW between two independent +//! implementations is the whole point of this file; sharing a runtime +//! dependency in that direction would invert the layering the membrane +//! exists to keep. What ships from this crate is the differential above, +//! never a delegation to the crate it is differential against. + +use super::*; + +use lance_graph_mask_risc::Program; +use lance_graph_quack::{Agg, Cmp, Col, Filter, Query}; + +/// `n = 1000, seed = 33` — the identical fixture +/// `pr4_matrix.rs::combine_mode_products_agree_with_the_oracle` sweeps, so +/// its recorded per-vector counts (in that file's own module doc) are a +/// usable cross-reference. SplitMix64 is deterministic: the same `(n, seed)` +/// always regenerates the identical lanes. +const SWEEP_N: u64 = 1000; +const SWEEP_SEED: u64 = 33; + +/// `&[LgjOpDesc]` (a left fold seeded with all-ones) read as the Boolean +/// TREE it denotes. `None` means the fold is still sitting on the all-ones +/// seed — no op has combined with AND yet. +/// +/// This is where the prefix rewrite `plan_lower`'s own module doc describes +/// falls out for free, as a PROPERTY of the tree construction rather than +/// as a special case anyone had to remember to write: an `Or` encountered +/// while `acc` is still `None` returns `(None, false)` below, so the op +/// simply never becomes a node — `all_ones | p == all_ones`, so a leaf that +/// contributes nothing is a leaf that is never built, not a leaf that is +/// built and then discarded. `plan_lower` reaches the identical answer by +/// scanning for the least AND index up front; this fold reaches it by +/// construction, one op at a time. +fn fold_to_tree(ops: &[LgjOpDesc]) -> Option { + let mut acc: Option = None; + for o in ops { + let leaf = Filter::Cmp(Col(o.lane_id as u16), cmp_of(o)); + acc = match (acc, o.combine == LGJ_COMBINE_AND) { + // all_ones & p == p — the first AND IS the accumulator. + (None, true) => Some(leaf), + // all_ones | p == all_ones — the op is DEAD. This is exactly + // `plan_lower`'s prefix rewrite, arriving here as a property of + // the tree construction rather than as a special case. + (None, false) => None, + // AND is associative, so consecutive ANDs flatten into one + // n-ary junction instead of left-nesting. + (Some(Filter::And(mut v)), true) => { + v.push(leaf); + Some(Filter::And(v)) + } + (Some(t), true) => Some(Filter::and([t, leaf])), + (Some(Filter::Or(mut v)), false) => { + v.push(leaf); + Some(Filter::Or(v)) + } + (Some(t), false) => Some(Filter::or([t, leaf])), + }; + } + acc +} + +/// One `LgjOpDesc` opcode + operand -> one `Cmp`, kept TEXTUALLY PARALLEL to +/// `plan_lower::pred_of` (every cast below is the identical cast, in the +/// identical order) rather than merely equivalent to it: a different cast +/// here is exactly the class of bug this differential exists to catch, and +/// two spellings that happen to agree today are less trustworthy than two +/// spellings that are visibly the same read. +fn cmp_of(o: &LgjOpDesc) -> Cmp { + match o.op { + LGJ_OP_EQ_U32 => Cmp::EqU32(o.operand as u32), + LGJ_OP_NE_U32 => Cmp::NeU32(o.operand as u32), + LGJ_OP_EQ_I32 => Cmp::EqI32(o.operand as i32), + LGJ_OP_NE_I32 => Cmp::NeI32(o.operand as i32), + LGJ_OP_GT_I32 => Cmp::GtI32(o.operand as i32), + LGJ_OP_LT_I32 => Cmp::LtI32(o.operand as i32), + LGJ_OP_LE_I32 => Cmp::LeI32(o.operand as i32), + LGJ_OP_GE_I32 => Cmp::GeI32(o.operand as i32), + LGJ_OP_TERNARY_MATCH_U32 => { + let packed = o.operand as u64; + Cmp::MatchU32 { + pattern: packed as u32, + care: (packed >> 32) as u32, + } + } + other => panic!("opcode {other} is not in the nine this file maps"), + } +} + +/// Execute an already-lowered [`Program`] against `planes`, reading its +/// `Terminal::Count` result. +/// +/// Both arms funnel through this ONE runner, and — unlike the Planes +/// construction each arm builds for itself below — sharing the EXECUTOR +/// here is correct rather than risky: `execute` is third-party ground truth +/// neither lowering is allowed to reimplement, so calling the identical +/// function on both sides is the whole point of comparing LOWERINGS rather +/// than comparing EXECUTIONS. +fn run(p: &Program, planes: &Planes) -> usize { + let words = planes.n_rows.div_ceil(64); + let slots = p.scratch_slots as usize; + let mut buf = vec![0u64; scratch_words_for(words, slots).expect("sized")]; + let mut scratch = Scratch::over(&mut buf, words, slots).expect("carves"); + match execute(p, planes, &mut scratch, None).expect("runs") { + Value::Count(c) => c, + other => panic!("not a count: {other:?}"), + } +} + +/// Arm A: this crate's own flat-op-list lowering. +/// +/// `lanes`/`planes` are built HERE, independently of [`eval_quack`]'s own +/// copy below, exactly the way `exports.rs`'s `plan_eval_impl` builds them +/// (`LaneRef::U64(ids)`, `LaneRef::U32(classes)`, `LaneRef::I32(values)`, in +/// that order — pinned by the same `const _` index assertions +/// `plan_eval_impl` relies on). Duplicating this handful of lines rather +/// than sharing a builder between the two arms is deliberate: a shared +/// Planes-construction bug would corrupt both arms identically and the +/// differential could not see it, where duplicated construction turns such +/// a bug into an ordinary mismatch between the two arms. +fn eval_plan_lower(fixture: &Fixture, n_rows: u64, ops: &[LgjOpDesc]) -> usize { + let lanes = [ + LaneRef::U64(fixture.ids()), + LaneRef::U32(fixture.classes()), + LaneRef::I32(fixture.values()), + ]; + let planes = Planes { + n_rows: n_rows as usize, + masks: &[], + lanes: &lanes, + }; + match plan_lower::lower_plan(ops).expect("opcode is mapped") { + plan_lower::Lowered::AllRows => n_rows as usize, + plan_lower::Lowered::Program(p) => run(&p, &planes), + } +} + +/// Arm B: `lance-graph-quack`'s Boolean-tree lowering, fed the SAME plan +/// read as the tree it denotes (see [`fold_to_tree`]). See [`eval_plan_lower`] +/// for why the `lanes`/`planes` construction is duplicated rather than +/// shared between the two arms. +fn eval_quack(fixture: &Fixture, n_rows: u64, ops: &[LgjOpDesc]) -> usize { + let lanes = [ + LaneRef::U64(fixture.ids()), + LaneRef::U32(fixture.classes()), + LaneRef::I32(fixture.values()), + ]; + let planes = Planes { + n_rows: n_rows as usize, + masks: &[], + lanes: &lanes, + }; + match fold_to_tree(ops) { + None => n_rows as usize, + Some(filter) => { + let q = Query { + filter, + agg: Agg::Count, + }; + run(&lance_graph_quack::lower(&q).expect("lowers"), &planes) + } + } +} + +/// Run both arms over the identical `ops` and assert they select the +/// identical row count, returning that count so callers can also fold it +/// into an anti-vacuity check. +fn assert_lowerings_agree(fixture: &Fixture, n_rows: u64, ops: &[LgjOpDesc], label: &str) -> usize { + let from_plan_lower = eval_plan_lower(fixture, n_rows, ops); + let from_quack = eval_quack(fixture, n_rows, ops); + assert_eq!( + from_plan_lower, from_quack, + "{label}: plan_lower's flat-op-list lowering and lance-graph-quack's \ + tree lowering select different row counts ({from_plan_lower} vs {from_quack})" + ); + from_plan_lower +} + +// --------------------------------------------------------------------------- +// The 28-vector combine sweep, shared by the first two tests below. +// --------------------------------------------------------------------------- + +/// A FIXED (opcode, lane, operand) pair per arity, so that across the whole +/// sweep only the COMBINE VECTOR varies — any disagreement is therefore +/// attributable to combine dispatch alone, the same discipline +/// `pr4_matrix.rs::combine_mode_products_agree_with_the_oracle` uses for its +/// own 2- and 3-op sweeps. +const TWO_OP: [(u32, u32, i64); 2] = [ + (LGJ_OP_EQ_U32, LANE_CLASSES, 7), + (LGJ_OP_GT_I32, LANE_VALUES, 100), +]; +const THREE_OP: [(u32, u32, i64); 3] = [ + (LGJ_OP_EQ_U32, LANE_CLASSES, 7), + (LGJ_OP_GT_I32, LANE_VALUES, 100), + (LGJ_OP_EQ_U32, LANE_CLASSES, 3), +]; +/// The arity-3 set above, plus a fourth op — this repo's own combine sweep +/// stops at 3; this file widens it to 4 (16 vectors) because a 3-op sweep +/// cannot exercise a SECOND consecutive AND landing after an OR-then-AND +/// pair, a shape only a 4-op vector can produce. +const FOUR_OP: [(u32, u32, i64); 4] = [ + (LGJ_OP_EQ_U32, LANE_CLASSES, 7), + (LGJ_OP_GT_I32, LANE_VALUES, 100), + (LGJ_OP_EQ_U32, LANE_CLASSES, 3), + (LGJ_OP_LT_I32, LANE_VALUES, 200), +]; + +/// Render a combine value as the word its constant names, for failure +/// messages — the same convention `pr4_matrix.rs::combine_name` uses (a +/// private helper of the same name in a sibling module; the two cannot +/// collide, they simply agree). +fn combine_name(c: u32) -> &'static str { + if c == LGJ_COMBINE_AND { + "AND" + } else { + "OR" + } +} + +/// Bit `i` of `bits` selects op `i`'s combiner: `0` = AND, `1` = OR. +fn ops_for_combine(base: &[(u32, u32, i64)], bits: u32) -> Vec { + base.iter() + .enumerate() + .map(|(i, &(opcode, lane, operand))| { + let combine = if bits & (1 << i) == 0 { + LGJ_COMBINE_AND + } else { + LGJ_COMBINE_OR + }; + op(opcode, lane, operand, combine) + }) + .collect() +} + +/// A human-readable label for one combine vector, e.g. `"n=3 [AND, OR, AND]"`. +fn combine_label(base: &[(u32, u32, i64)], bits: u32) -> String { + let names: Vec<&str> = (0..base.len()) + .map(|i| { + combine_name(if bits & (1 << i) == 0 { + LGJ_COMBINE_AND + } else { + LGJ_COMBINE_OR + }) + }) + .collect(); + format!("n={} [{}]", base.len(), names.join(", ")) +} + +/// Drive `visit` over every combine vector this file's first two tests both +/// need — the full `{AND, OR}^n` product for `n = 2, 3, 4` (4 + 8 + 16 = 28 +/// vectors) — factored ONCE so the two tests provably sweep the SAME 28 +/// rather than two independently-typed lookalikes that could silently drift +/// apart from each other. +fn for_each_combine_vector(mut visit: impl FnMut(String, Vec)) { + for bits in 0u32..4 { + visit(combine_label(&TWO_OP, bits), ops_for_combine(&TWO_OP, bits)); + } + for bits in 0u32..8 { + visit( + combine_label(&THREE_OP, bits), + ops_for_combine(&THREE_OP, bits), + ); + } + for bits in 0u32..16 { + visit( + combine_label(&FOUR_OP, bits), + ops_for_combine(&FOUR_OP, bits), + ); + } +} + +/// FAILS IF: `plan_lower` and `lance-graph-quack` select a different row +/// count for ANY of the 28 combine vectors this test sweeps over two fixed +/// (opcode, lane, operand) sets — the differential's headline claim, over +/// the widest combine-dispatch surface either implementation is exercised +/// against anywhere in this crate. +#[test] +fn the_two_lowerings_agree_on_every_combine_vector() { + let fixture = Fixture::generate(SWEEP_N, SWEEP_SEED).expect("n_rows fits usize on this target"); + let mut results: Vec<(String, usize)> = Vec::new(); + + for_each_combine_vector(|label, ops| { + let count = assert_lowerings_agree(&fixture, SWEEP_N, &ops, &label); + results.push((label, count)); + }); + + // `--nocapture` shows the real numbers; the anti-vacuity check below is + // computed from this SAME vector, not asserted separately from it. + for (label, count) in &results { + println!("{label} -> {count}"); + } + assert_eq!( + results.len(), + 28, + "the sweep must cover all 4 + 8 + 16 = 28 combine vectors, got {}", + results.len() + ); + + // Anti-vacuity for the sweep as a whole (this workspace's own + // falsifiability rule): agreement between two lowerings that both + // answered "0" or "n" on every vector would prove nothing about combine + // dispatch — it would just be two implementations of the identity + // function agreeing with each other. + // + // MEASURED, not chosen: 21 of 28. The seven degenerate vectors are + // `n=2 [OR,OR]`, `n=3 [AND,AND,AND]` and `[OR,OR,OR]`, and + // `n=4 [AND,AND,AND,AND]`, `[OR,AND,OR,OR]`, `[AND,OR,OR,OR]`, + // `[OR,OR,OR,OR]`. Four of those are the structural all-OR-tail + // saturations the `assert_eq!`s above already pin separately; the + // all-AND zeros are structural too (`class == 7` and `class == 3` + // cannot both hold). + // + // The exact form is what made a real fixture defect visible. A first + // version of the arity-4 arm appended `LT_I32(500)` — past the values + // lane's own maximum of 361, so an always-true op — and the whole + // 16-vector arm collapsed to eight saturated 1000s plus eight verbatim + // copies of the arity-3 row, contributing exactly ZERO discriminating + // power. It still cleared a `>= 15` floor, by sitting on it. Measuring + // the lane and replacing the operand with `LT_I32(200)` (676 of 1000) + // took the count 15 -> 21. A floor would have shipped the dead arm. + let non_degenerate = results + .iter() + .filter(|(_, count)| *count > 0 && *count < SWEEP_N as usize) + .count(); + assert_eq!( + non_degenerate, + 21, + "the combine sweep's discriminating power moved: {non_degenerate} of \ + {} vectors were non-degenerate (neither 0 nor {SWEEP_N}), expected 21. \ + Re-measure with --nocapture and re-pin; do NOT relax this to a floor", + results.len() + ); +} + +/// FAILS IF: `plan_lower`'s `Lowered::AllRows` shortcut and the tree fold's +/// `None` sentinel are reached under DIFFERENT conditions. They are meant to +/// be the SAME predicate — "no op in this plan combines with AND" — reached +/// by two different routes: a linear scan for the least AND index on one +/// side, a fold that never establishes an accumulator until the first AND +/// arrives on the other. A plan that flips one without flipping the other +/// would silently change which arm of a caller's `match` fires. +#[test] +fn the_all_rows_shortcut_is_the_same_condition_on_both_sides() { + let mut all_rows = 0usize; + let mut not_all_rows = 0usize; + + for_each_combine_vector(|label, ops| { + let is_all_rows = matches!( + plan_lower::lower_plan(&ops).expect("opcode is mapped"), + plan_lower::Lowered::AllRows + ); + let is_none = fold_to_tree(&ops).is_none(); + assert_eq!( + is_all_rows, is_none, + "{label}: plan_lower reports AllRows={is_all_rows} but the tree fold's \ + None-ness is {is_none} — the two must agree on whether any op in this \ + plan combines with AND" + ); + if is_all_rows { + all_rows += 1; + } else { + not_all_rows += 1; + } + }); + + // Two-sided anti-vacuity, both counts checked EXACTLY: a version that + // always agreed by both sides unconditionally returning `false` (or + // both unconditionally returning `true`) would pass every assertion + // above while proving nothing about the shared condition. The all-OR + // vector is EXACTLY one per arity (every bit set: 0b11, 0b111, 0b1111), + // so the split is known in advance rather than merely "some of each". + assert_eq!( + all_rows, 3, + "expected exactly 3 all-OR vectors (one per arity) to read AllRows/None \ + on both sides, got {all_rows}" + ); + assert_eq!( + not_all_rows, 25, + "expected the remaining 25 of 28 vectors to build a real program on both \ + sides, got {not_all_rows}" + ); +} + +// --------------------------------------------------------------------------- +// Per-opcode predicate agreement. +// --------------------------------------------------------------------------- + +/// The nine mapped opcodes, each paired with a (lane, operand) that is a +/// well-formed predicate over that lane's documented domain (`fixture.rs`: +/// classes `0..=15`, values `-150..=361`). +/// +/// Every operand is MEASURED against the real lane rather than chosen from +/// the documented domain bounds, and that distinction cost a revision: a +/// first version gave `LT_I32`/`LE_I32` the operand `500`, past the values +/// lane's maximum of `361`, on the reasoning that a structurally always-true +/// predicate still exercises opcode-to-predicate AGREEMENT. It does not +/// exercise it usefully. Two arms agreeing that EVERY row survives cannot +/// separate `LtI32` from `LeI32` from any other predicate that is also +/// always-true on this lane — so a swapped entry in either mapping table +/// would have passed, which is precisely the defect this file exists to +/// catch. Measured (n = 1000, seed = 33): values span `-150..=361` with +/// median 100, so `LT_I32(200)` selects 676 and `LE_I32(300)` selects 866. +/// +/// The four ordered comparisons are also given DISTINCT counts on purpose. +/// This is a differential between two arms, not against ground truth, so a +/// mis-map is visible only when it moves ONE arm's count — and two opcodes +/// that happen to select the same number of rows would hide a swap between +/// exactly those two. +/// +/// That is not a theoretical worry; it was MEASURED on this exact file. +/// Mis-mapping `LGJ_OP_LE_I32` to `Pred::LtI32` in `plan_lower` — a +/// one-token change, and the precise defect class this test exists to catch +/// — is RED at operand `300` (`LeI32(300)` selects 866, `LtI32(300)` selects +/// 865: one row of difference is all it takes) and **GREEN at operand +/// `500`**, where both readings select all 1000 and the arms agree on an +/// answer neither of them computed correctly. The operand is what makes this +/// a test rather than a decoration; do not move it back inside the "the +/// domain doc says 361, so 500 is safely past it" reasoning that put it +/// there. +const OPCODE_CASES: [(&str, u32, u32, i64); 9] = [ + ("EQ_U32", LGJ_OP_EQ_U32, LANE_CLASSES, 7), + ("NE_U32", LGJ_OP_NE_U32, LANE_CLASSES, 7), + ("EQ_I32", LGJ_OP_EQ_I32, LANE_VALUES, 100), + ("NE_I32", LGJ_OP_NE_I32, LANE_VALUES, 100), + ("GT_I32", LGJ_OP_GT_I32, LANE_VALUES, 100), + ("LT_I32", LGJ_OP_LT_I32, LANE_VALUES, 200), + ("LE_I32", LGJ_OP_LE_I32, LANE_VALUES, 300), + ("GE_I32", LGJ_OP_GE_I32, LANE_VALUES, 100), + ( + "TERNARY_MATCH", + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + ((0x0000_000Fu64 << 32) | 7) as i64, + ), +]; + +/// Two DISTINCT AND predicates on the lane OPPOSITE `lane`, used to sandwich +/// a subject opcode at position 1 of a genuine 3-op AND chain. Opposite-lane +/// so the chain never repeats the subject's own (opcode, lane) pair; +/// distinct from EACH OTHER so a lowering that silently reused position 0's +/// fields for position 2 (or vice versa) cannot hide behind two identical +/// entries. +fn flanking_ands(lane: u32) -> [LgjOpDesc; 2] { + if lane == LANE_VALUES { + [ + op(LGJ_OP_NE_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND), + op(LGJ_OP_EQ_U32, LANE_CLASSES, 0, LGJ_COMBINE_AND), + ] + } else { + [ + op(LGJ_OP_LT_I32, LANE_VALUES, 200, LGJ_COMBINE_AND), + op(LGJ_OP_GT_I32, LANE_VALUES, -100, LGJ_COMBINE_AND), + ] + } +} + +/// FAILS IF: any of the nine mapped opcodes translate to a DIFFERENT +/// predicate in `plan_lower`'s lowering than in `lance-graph-quack`'s — +/// checked both as the bare seed (position 0, ungated by anything) and +/// sandwiched at position 1 of a genuine 3-op AND chain, where BOTH +/// lowerings must gate the subject's own evaluation under an accumulator +/// two DIFFERENT flanking predicates have already narrowed. Also checks +/// that the TCAM opcode's packed `pattern`/`care` halves are not +/// interchangeable in either lowering, and that a half-swap moves both +/// lowerings identically. +#[test] +fn every_opcode_maps_to_the_same_predicate_in_both_lowerings() { + let fixture = Fixture::generate(SWEEP_N, SWEEP_SEED).expect("n_rows fits usize on this target"); + let mut seed_counts: Vec<(&str, usize)> = Vec::new(); + + for &(label, opcode, lane, operand) in &OPCODE_CASES { + let subject = op(opcode, lane, operand, LGJ_COMBINE_AND); + + // Position 0: the bare seed. A lone AND op against the all-ones + // accumulator IS its own predicate on both sides — no gate to get + // wrong yet. + let seed_count = assert_lowerings_agree( + &fixture, + SWEEP_N, + &[subject], + &format!("{label} at position 0"), + ); + seed_counts.push((label, seed_count)); + + // Position 1 of 3: genuinely sandwiched, gated by an accumulator + // TWO other predicates have already narrowed — see `flanking_ands`. + let [before, after] = flanking_ands(lane); + let three_op = [before, subject, after]; + assert_lowerings_agree( + &fixture, + SWEEP_N, + &three_op, + &format!("{label} at position 1 of 3"), + ); + } + + for (label, count) in &seed_counts { + println!("{label} at position 0 -> {count}"); + } + assert_eq!( + seed_counts.len(), + 9, + "expected all 9 mapped opcodes to be checked, got {}", + seed_counts.len() + ); + + // Anti-vacuity: an opcode whose seed selects everything or nothing + // cannot discriminate ANYTHING checked above for that opcode — two + // lowerings agreeing that "everything" or "nothing" survives is not + // evidence either translated the predicate correctly. + // + // MEASURED, and it is ALL NINE: 65, 935, 2, 998, 498, 676, 866, 500, 65 + // at n = 1000. A first version gave `LT_I32`/`LE_I32` the operand `500` + // and both selected every row, which made the floor `7` — and a floor of + // 7 is exactly what let two structurally untested opcode mappings ship. + // `LtI32`, `LeI32`, and any other always-true reading are + // indistinguishable when every row survives, so the swap this file + // exists to catch would have passed on those two rows. See + // `OPCODE_CASES`'s doc comment for the measurement that replaced them. + let proper_subset = seed_counts + .iter() + .filter(|(_, c)| *c > 0 && *c < SWEEP_N as usize) + .count(); + assert_eq!( + proper_subset, 9, + "{proper_subset} of 9 opcode seeds selected a proper subset of \ + {SWEEP_N} rows, expected all 9. An opcode that selects everything or \ + nothing is not being compared — re-measure the lane and pick an \ + operand inside it; do NOT relax this to a floor" + ); + + // The TCAM half-swap: pattern and care must not be interchangeable in + // EITHER lowering. `pattern=7, care=0x0000_000F` (`OPCODE_CASES`'s own + // TCAM row) already has `pattern != care`, which is the precondition a + // half-swap needs to be observable at all (per `pr4_matrix.rs`'s own + // note: the bug this guards against "is invisible whenever + // `pattern == care`"). + let original_operand = ((0x0000_000Fu64 << 32) | 7) as i64; + let swapped_operand = ((7u64 << 32) | 0x0000_000F) as i64; + assert_ne!( + original_operand, swapped_operand, + "the swap must actually change the packed operand" + ); + + let original = op( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + original_operand, + LGJ_COMBINE_AND, + ); + let swapped = op( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + swapped_operand, + LGJ_COMBINE_AND, + ); + + let original_count = + assert_lowerings_agree(&fixture, SWEEP_N, &[original], "TCAM original halves"); + let swapped_count = + assert_lowerings_agree(&fixture, SWEEP_N, &[swapped], "TCAM swapped halves"); + // The point of this half-swap is that it must move BOTH sides + // IDENTICALLY, never just one — `assert_lowerings_agree` above already + // proved that for each operand separately. This second assertion is the + // other half: the swap must move them somewhere DIFFERENT from where + // the original sat, or a lowering that silently ignored the swap + // entirely (reading the operand identically before and after) would + // still pass every check above by construction. + assert_ne!( + original_count, swapped_count, + "swapping pattern and care selected the identical {original_count} rows — \ + pattern=7 and care=0x0000_000F are not equal, so a genuine half-swap must \ + select a DIFFERENT population; an identical count means the swap had no \ + observable effect on at least one lowering" + ); + + // And the swapped operand must ALSO agree once sandwiched, same as + // every other opcode above — a half-swap bug could plausibly be scoped + // to the ungated arm alone. + let [before, after] = flanking_ands(LANE_CLASSES); + assert_lowerings_agree( + &fixture, + SWEEP_N, + &[before, swapped, after], + "TCAM swapped halves at position 1 of 3", + ); +} diff --git a/native/lgj-abi/src/exports/tests/pr4_dst_reuse.rs b/native/lgj-abi/src/exports/tests/pr4_dst_reuse.rs new file mode 100644 index 0000000..c7c6ac0 --- /dev/null +++ b/native/lgj-abi/src/exports/tests/pr4_dst_reuse.rs @@ -0,0 +1,445 @@ +//! F-PR4-DST and F-PR4-REUSE: destination-state and scratch-reuse +//! falsifiers for PR4's fused-plan evaluator. +//! +//! `pr4_equivalence.rs` and `pr4_seed.rs` both run every plan into a FRESH, +//! empty destination created moments before the call — so neither file, nor +//! anything in the pre-existing suite, has ever asked what happens when +//! `dst_mask` already holds something, or what happens to a scratch buffer +//! that outlives one call. Two properties follow directly from `abi.md`'s +//! own contract ("dst_mask is written exactly once... an error at any point +//! leaves it byte-for-byte as it was") that no existing test can see fail: +//! +//! - **F-PR4-DST** — the published result must not depend on `dst_mask`'s +//! PRIOR content, whether that prior content is a fresh all-ones mask, a +//! genuinely different population left there by an earlier call, or raw +//! poison written straight through the described segment. A dirty prior +//! tail must never be read as an input plane. And a rejected plan must +//! leave that prior content untouched at EVERY position an invalid op +//! could occupy, not only the one position the pre-existing error-path +//! test happens to use. +//! - **F-PR4-REUSE** — PR4's whole allocation win comes from reusing a +//! scratch buffer across calls instead of allocating one per call. A +//! scratch buffer left dense (or sparse) by one call must not leak into +//! the next, unrelated call's answer, in either order, and repeating the +//! identical call must be idempotent. +//! +//! # Freeze discipline (same rule as `pr4_equivalence.rs`) +//! +//! Every test below is written against the OLD implementation, still in +//! place, and must be GREEN before any implementation change. A red result +//! here has found either a pre-existing defect in the old loop or a bug in +//! this test — either way, report it rather than silently reconciling it, +//! because a silent fix on either side changes what "equivalent to the old" +//! means. + +use super::pr4_equivalence::{assert_agrees, eval_legacy, eval_live, Outcome}; +use super::*; + +/// Overwrite mask `m`'s own words with `fill`, through the exact segment +/// [`read_words`] reads back — the write-side mirror of that helper, needed +/// only in this file because F-PR4-DST specifically needs a destination +/// that already held *something* before the call, and there is no ABI call +/// that seeds a fresh mask with arbitrary content directly. +fn poison_words(m: u64, fill: u64) { + let mut d = LgjLaneDesc::default(); + assert_eq!(call::mask_describe(m, &mut d), LGJ_OK); + assert_eq!(d.elem_kind, LgjElemKind::MaskWord as u32); + assert_ne!(d.flags & LGJ_FLAG_WRITABLE, 0); + // SAFETY: the descriptor is exactly the contract Java relies on to build + // a writable `MemorySegment` over this lane (confirmed writable by the + // flag check above); the lane is alive because `m` has not been closed + // in this scope; and this call is the sole writer of the slice for the + // duration of the write below, so there is no aliasing with any other + // live reference to it. + unsafe { + let words = std::slice::from_raw_parts_mut(d.addr as *mut u64, d.len_elems as usize); + words.fill(fill); + } +} + +/// Create a fresh mask over `pattern` and poison every one of its words to +/// `fill`, returning the handle. Several tests below need a destination +/// that already holds deliberate garbage rather than nothing; this pairs +/// [`mask`] with [`poison_words`] at every call site that needs one. +fn poisoned_mask(pattern: u64, fill: u64) -> u64 { + let m = mask(pattern, LGJ_MASK_INIT_EMPTY); + poison_words(m, fill); + m +} + +/// True iff no bit at row index `>= n_rows` survives in the last word. +/// +/// Copied from `pr4_seed.rs` rather than shared as a common helper (that +/// file's own doc comment explains why it is written as its own +/// independent formula rather than delegating to `clear_tail_bits`, and the +/// same reasoning applies here: this must be able to catch a bug at the +/// call site that matters, not merely re-confirm what a shared helper +/// always computes). +/// +/// `n_rows % 64 == 0` needs its own branch: when the last word is exactly +/// full there is no tail region at all, and shifting by zero would +/// (wrongly) demand the whole word be zero. +fn tail_is_clean(words: &[u64], n_rows: u64) -> bool { + let rem = n_rows % 64; + if rem == 0 { + return true; + } + match words.last() { + Some(&last) => last >> rem == 0, + None => true, + } +} + +/// FAILS IF: `dst_mask`'s prior content changes the published result. Every +/// arm below runs the IDENTICAL plan into a destination that started in a +/// different state — fresh-empty, fresh-saturated, holding a genuinely +/// different prior population, or holding raw poison — and all four must +/// publish byte-for-byte identically. This is what catches an +/// implementation that ACCUMULATES into its destination (reading +/// `dst_mask`'s own words as part of the computation) rather than +/// overwriting it outright: three of the four priors below are non-empty, +/// and one has every bit set, so an accumulate bug has ample non-trivial +/// state to leak through. +#[test] +fn the_same_plan_lands_identically_whatever_the_destination_held() { + let n = 1000u64; + let p = open(n, 33); + let plan = [ + op(LGJ_OP_EQ_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND), + op(LGJ_OP_GT_I32, LANE_VALUES, 100, LGJ_COMBINE_AND), + ]; + + let fresh_empty = mask(p, LGJ_MASK_INIT_EMPTY); + let fresh_all = mask(p, LGJ_MASK_INIT_ALL); + + // A genuinely different, non-trivial prior: `class == 7 AND value <= + // 100` is disjoint from `class == 7 AND value > 100` (the real plan + // below) BY CONSTRUCTION — `LE_I32` and `GT_I32` at the same threshold + // partition every value exactly — so this prior can never coincide + // with the real plan's own answer, whatever the fixture's seed happens + // to draw. + let re_used = mask(p, LGJ_MASK_INIT_EMPTY); + let prior = eval_live( + p, + &[ + op(LGJ_OP_EQ_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND), + op(LGJ_OP_LE_I32, LANE_VALUES, 100, LGJ_COMBINE_AND), + ], + re_used, + ); + assert!( + prior.count > 0 && prior.count < n, + "the re-used destination's prior population must itself be \ + non-trivial, or this arm cannot tell an accumulate bug apart from \ + a genuine overwrite (got {})", + prior.count + ); + + let poisoned = poisoned_mask(p, u64::MAX); + + let out_empty = eval_live(p, &plan, fresh_empty); + let out_all = eval_live(p, &plan, fresh_all); + let out_reused = eval_live(p, &plan, re_used); + let out_poisoned = eval_live(p, &plan, poisoned); + + assert_eq!(out_empty.status, LGJ_OK); + assert!( + out_empty.count > 0 && out_empty.count < n, + "the plan under test must itself select a proper subset, or none of \ + these four arms can be told apart (got {})", + out_empty.count + ); + + assert_eq!( + out_empty, out_all, + "a fresh-empty and a fresh-saturated destination must publish \ + identically" + ); + assert_eq!( + out_empty, out_reused, + "a fresh-empty destination and one already holding a disjoint prior \ + population must publish identically" + ); + assert_eq!( + out_empty, out_poisoned, + "a fresh-empty destination and one poisoned to all-ones (including \ + a dirty tail) must publish identically" + ); + + assert_eq!(lgj_close(p), LGJ_OK); +} + +/// FAILS IF: the FINAL `clear_tail_bits` before publish is skipped, is +/// applied to the wrong buffer, or only clears a freshly-zeroed word rather +/// than genuinely masking off whatever bits were already there. Every `n` +/// below except `128` has a real tail region (bits at index `>= n % 64` in +/// the last word); the destination is poisoned to `u64::MAX` — dirtying +/// EXACTLY that region — before the call, so survival of even one of those +/// bits is directly attributable to the publish step, not to some other +/// source of garbage. +#[test] +fn a_poisoned_tail_is_published_clean() { + for n in [1u64, 63, 65, 127, 128, 999, 4097] { + let p = open(n, 0xDA57_0000 | n); + let m = poisoned_mask(p, u64::MAX); + + let rem = n % 64; + if rem != 0 { + // Anti-vacuity: prove the poison actually reached the tail + // region BEFORE the call runs, so a `poison_words` that + // silently no-oped (or wrote the wrong lane) could not make + // this test pass by leaving nothing to clean up. Skipped only + // at `rem == 0` (n = 128 here): a whole final word has no tail + // bits at all, so there is nothing above `n_rows` to dirty. + let before = read_words(m); + assert!( + before.last().copied().unwrap_or(0) >> rem != 0, + "n={n}: poison_words did not actually dirty the tail region \ + above row {n} (last word = {:#x})", + before.last().copied().unwrap_or(0) + ); + } + + let ops = [op(LGJ_OP_GT_I32, LANE_VALUES, 100, LGJ_COMBINE_AND)]; + let out = eval_live(p, &ops, m); + assert_eq!( + out.status, LGJ_OK, + "n={n}: a well-formed single-op plan must validate" + ); + assert!( + tail_is_clean(&out.words, n), + "n={n}: a poisoned prior tail survived publish (last word = \ + {:#x})", + out.words.last().copied().unwrap_or(0) + ); + + assert_eq!(lgj_close(p), LGJ_OK); + } +} + +/// FAILS IF: a future lowering hands `dst_mask` itself to `mask_risc`'s +/// `Planes::masks` — the input-plane list. `validate` refuses ANY plane +/// whose tail is dirty (`ExecError::PlaneTail`), and `Planes::masks` must +/// be `&[]` for `plan_eval`. If `dst_mask` were ever fed back in as an +/// input plane, a poisoned prior tail would turn what should be an +/// ordinary, successful overwrite into a spurious validation error. This +/// poisons the WHOLE destination (interior and tail alike) and asserts the +/// call still succeeds, which distinguishes "this mask is merely where the +/// answer will be written" from "this mask is also read as input" — the +/// exact seam a future lowering could get wrong. +#[test] +fn a_dirty_destination_is_not_read_as_an_input_plane() { + let n = 777u64; // not a multiple of 64: a real tail region exists to dirty + let p = open(n, 0xD1A7_D57A); + let ops = [ + op(LGJ_OP_GT_I32, LANE_VALUES, 100, LGJ_COMBINE_AND), + op(LGJ_OP_EQ_U32, LANE_CLASSES, 7, LGJ_COMBINE_OR), + ]; + + let dirty = poisoned_mask(p, u64::MAX); + let out_dirty = eval_live(p, &ops, dirty); + assert_eq!( + out_dirty.status, LGJ_OK, + "a destination whose tail was poisoned before the call must still \ + evaluate successfully — its dirty tail describes only what will be \ + overwritten, and must never be validated as if it were an input \ + plane" + ); + + let fresh = mask(p, LGJ_MASK_INIT_EMPTY); + let out_fresh = eval_live(p, &ops, fresh); + assert_eq!( + out_dirty, out_fresh, + "the same plan into a poisoned-tail destination and a fresh \ + destination must publish identically" + ); + + assert_eq!(lgj_close(p), LGJ_OK); +} + +/// FAILS IF: validation and evaluation are interleaved rather than fully +/// separated — the OQ-1 ordering concern. Today `validate_plan` walks the +/// WHOLE plan before any op is evaluated, so a bad op anywhere leaves +/// `dst_mask` untouched regardless of its position. A future lowering that +/// validates op-by-op WHILE evaluating (rather than validating the whole +/// plan up front) could let earlier, valid ops' effects reach `dst_mask` +/// before a later op's invalidity is discovered. Sweeping every position of +/// a 4-op plan, with the destination pre-poisoned to a value neither +/// `LGJ_MASK_INIT_EMPTY` nor `LGJ_MASK_INIT_ALL` could ever produce, makes +/// "unchanged" an OBSERVATION rather than "still zero, which a half-applied +/// write could also produce by accident." +#[test] +fn an_error_at_any_position_in_a_plan_leaves_the_destination_untouched() { + const POISON: u64 = 0xDEAD_BEEF_DEAD_BEEF; + let n_rows = 500u64; + let p = open(n_rows, 0xBAD0_5EED); + let expected_poison = vec![POISON; mask_words_for(n_rows) as usize]; + + let good = [ + op(LGJ_OP_EQ_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND), + op(LGJ_OP_GT_I32, LANE_VALUES, 100, LGJ_COMBINE_OR), + op(LGJ_OP_LT_I32, LANE_VALUES, 200, LGJ_COMBINE_AND), + op(LGJ_OP_EQ_U32, LANE_CLASSES, 3, LGJ_COMBINE_OR), + ]; + // Out of range on ANY pattern: `PATTERN_LANE_COUNT` is 3, so 9999 is + // rejected by `validate_plan`'s own lane check regardless of `n_rows`. + let bad = op(LGJ_OP_GT_I32, 9999, 0, LGJ_COMBINE_AND); + + for k in 0..good.len() { + let mut plan = good; + plan[k] = bad; + + let out_live = eval_live(p, &plan, poisoned_mask(p, POISON)); + let out_legacy = eval_legacy(p, &plan, poisoned_mask(p, POISON)); + + assert_eq!( + out_live.status, LGJ_ERR_INVALID_LANE, + "k={k}: a plan carrying an out-of-range lane_id must be \ + rejected with LGJ_ERR_INVALID_LANE" + ); + assert_eq!( + out_live.status, out_legacy.status, + "k={k}: the live path and the frozen oracle must reject the \ + same malformed plan the same way" + ); + assert_eq!( + out_live.words, expected_poison, + "k={k}: a rejected plan must leave the poisoned destination \ + byte-for-byte untouched, not merely zeroed (live path)" + ); + assert_eq!( + out_legacy.words, expected_poison, + "k={k}: a rejected plan must leave the poisoned destination \ + byte-for-byte untouched, not merely zeroed (oracle side)" + ); + assert_eq!( + out_live.count, 12345, + "k={k}: out_count must not be written on a rejected plan (live \ + path)" + ); + assert_eq!( + out_legacy.count, 12345, + "k={k}: out_count must not be written on a rejected plan \ + (oracle side)" + ); + } + + assert_eq!(lgj_close(p), LGJ_OK); +} + +/// FAILS IF: a scratch buffer shared or reused across separate `plan_eval` +/// calls (PR4's whole allocation-avoidance strategy) is left DENSE or +/// SPARSE by one call and that residue leaks into a LATER, unrelated +/// call's answer instead of being fully re-established by that later call. +/// +/// [`assert_agrees`] is used for every measurement below, not only the +/// final one, and that is deliberate: its two sub-calls run on the SAME +/// thread as everything before them, so its LIVE half inherits whatever +/// the shared scratch buffer looked like when the PREVIOUS `assert_agrees` +/// call finished, while its LEGACY half allocates its own fresh scratch on +/// every call (the frozen copy's own `let mut scratch = vec![0u64; +/// n_words];`) and is therefore immune to this leak by construction — +/// exactly what makes it a valid oracle for detecting one. Each call also +/// creates its own brand-new destination internally, which is what gives +/// every measurement below "a DIFFERENT destination" for free. +#[test] +fn scratch_left_dense_by_one_call_does_not_leak_into_the_next() { + // NOT a single all-OR op. `plan_lower::lower_plan` special-cases a plan + // with no AND-combined op at all into `Lowered::AllRows`: it writes + // `dst_mask` directly (all-ones, tail cleared) and returns WITHOUT ever + // building or running a `Program` — so it never touches the reusable + // scratch slots this test exists to exercise. An earlier version of this + // fixture (`[op(LGJ_OP_EQ_U32, LANE_CLASSES, 999, LGJ_COMBINE_OR)]`) was + // exactly that: a single OR-combined op, so `lower_plan` found no AND + // position and took the `AllRows` shortcut every time — the test never + // reached the code it claimed to be testing. + // + // This plan is TWO AND-combined ops instead, which `lower_plan` can only + // lower to a real `Program` (the shortcut applies solely when zero ops + // combine with AND — see the module doc in `plan_lower.rs`): op 0 writes + // scratch slot `plan_lower::ACC_SLOT` (0) directly, op 1 writes the + // predicate scratch slot (1) and ANDs it back into slot 0. Both + // thresholds sit strictly outside the fixture's own documented `values` + // range (`-150..=361`, `fixture.rs`'s `Fixture::generate` doc), so BOTH + // predicates are true for every row regardless of seed or row count — + // the plan executes for real and still leaves every scratch slot fully + // dense, exactly the state this test needs to seed before the selective + // call that follows. + let everything = [ + op(LGJ_OP_GE_I32, LANE_VALUES, -1000, LGJ_COMBINE_AND), + op(LGJ_OP_LE_I32, LANE_VALUES, 1000, LGJ_COMBINE_AND), + ]; + let selective = [ + op(LGJ_OP_EQ_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND), + op(LGJ_OP_GT_I32, LANE_VALUES, 100, LGJ_COMBINE_AND), + ]; + + // Round A: a call whose scratch ends DENSE — a two-op AND plan whose + // predicates are both unconditionally true saturates the accumulator to + // all-ones through a REAL program, not through the `AllRows` shortcut — + // immediately followed by a selective call — and then the identical + // selective call again, to pin idempotence on the same shared state. + let n1 = 1000u64; + let p1 = open(n1, 33); + + let dense_a: Outcome = assert_agrees(p1, &everything, "round A: dense seeding call"); + assert_eq!( + dense_a.count, n1, + "round A's seeding call must actually saturate every row through a \ + real program (both AND'd predicates are unconditionally true), or \ + it never leaves scratch dense in the first place" + ); + + let selective_a: Outcome = assert_agrees( + p1, + &selective, + "round A: selective call right after a dense-leaving one", + ); + assert!( + selective_a.count > 0 && selective_a.count < n1, + "round A: the selective plan must select a proper subset, or a \ + stale dense bit leaking in from the seeding call would be \ + invisible (got {})", + selective_a.count + ); + + // Idempotence: the identical selective plan, run again right after, on + // the same pattern (so it shares whatever scratch state the two calls + // above already left behind), must publish identically. + let selective_a2: Outcome = assert_agrees(p1, &selective, "round A: idempotence re-run"); + assert_eq!( + selective_a, selective_a2, + "idempotence: the identical selective plan, run twice in a row, \ + must publish identically" + ); + + assert_eq!(lgj_close(p1), LGJ_OK); + + // Round B, INVERTED order and a second (n_rows, seed) pair: the + // selective call runs FIRST (leaving scratch mostly zero), the + // saturating one SECOND. + let n2 = 2500u64; + let p2 = open(n2, 0x51DE_5EED); + + let selective_b: Outcome = assert_agrees(p2, &selective, "round B: selective seeding call"); + assert!( + selective_b.count > 0 && selective_b.count < n2, + "round B's seeding call must itself be a proper subset, or it \ + never leaves scratch sparse in the first place (got {})", + selective_b.count + ); + + let dense_b: Outcome = assert_agrees( + p2, + &everything, + "round B: saturating call right after a sparse-leaving one", + ); + assert_eq!( + dense_b.count, n2, + "round B: an always-true two-op AND plan must still select every \ + row, whatever scratch state the preceding selective call left \ + behind" + ); + + assert_eq!(lgj_close(p2), LGJ_OK); +} diff --git a/native/lgj-abi/src/exports/tests/pr4_equivalence.rs b/native/lgj-abi/src/exports/tests/pr4_equivalence.rs new file mode 100644 index 0000000..3bd475d --- /dev/null +++ b/native/lgj-abi/src/exports/tests/pr4_equivalence.rs @@ -0,0 +1,203 @@ +//! PR4's behaviour-equivalence oracle and its falsifier matrix. +//! +//! # Why a frozen copy, and not mask-risc's own oracle +//! +//! PR4 routes [`plan_eval_impl`] through `lance_graph_mask_risc::execute` in +//! place of its own opcode loop. That introduces a component which did not +//! exist before: the lowering `&[LgjOpDesc] -> mask_risc::Program`. It sits +//! ABOVE both `execute` and mask-risc's `reference_execute`, so a bug in it — +//! a swapped opcode, a dropped all-ones seed, a mis-packed TCAM operand — +//! produces the same wrong `Program` for both, both evaluate it faithfully, +//! and mask-risc's own differential stays green. +//! +//! Only an oracle that never sees the `Program` can catch that. This is that +//! oracle: a frozen copy of the pre-PR4 evaluation loop. +//! +//! # Freeze discipline (non-negotiable) +//! +//! [`legacy_plan_eval_impl`] is a byte-for-byte copy of `plan_eval_impl`'s +//! body as of **`b78c4f356a747e975b796a357d37a7c94f22967a`**. It must NOT be +//! refactored to share a helper with the new path — a shared helper is a +//! shared bug, and turns the differential into a tautology. Its evaluation +//! core calls only `kernels::{eval_predicate, combine_into, popcount}` and +//! `clear_tail_bits`. +//! +//! The *validation* prologue is deliberately shared (`validate_plan`, +//! `resolve_pattern_and_mask`, `lane_view`): PR4 does not change it, and +//! duplicating it would mean the differential could not see a regression in +//! the one part both paths genuinely have in common. +//! +//! # Ordering +//! +//! Everything in this file is written against the OLD implementation, still +//! in place, and is GREEN before any implementation change. A test red here +//! has found a pre-existing defect in the old loop — report it, do not fix it +//! silently, because a silently fixed old bug changes what "equivalent to the +//! old" means. + +use super::*; + +/// The pre-PR4 evaluation loop, frozen. +/// +/// See the module doc for the freeze discipline. The only differences from +/// the original are the name and this comment; every line of the body is the +/// original's. +pub(super) fn legacy_plan_eval_impl( + res: u64, + ops: *const LgjOpDesc, + n_ops: u32, + dst_mask: u64, + out_count: *mut u64, + path: Path, +) -> i32 { + // EMPTY_PLAN is checked before the null test: `(null, 0)` is a caller + // describing an empty plan, which has its own dedicated code, and reporting + // NULL_ARGUMENT there would send a Java author looking for the wrong bug. + if n_ops == 0 { + return LGJ_ERR_EMPTY_PLAN; + } + if ops.is_null() || out_count.is_null() { + return LGJ_ERR_NULL_ARGUMENT; + } + let (pattern, mask) = match resolve_pattern_and_mask(res, dst_mask) { + Ok(t) => t, + Err(e) => return e, + }; + // SAFETY: `ops` is non-null (checked) and the caller states it points at + // `n_ops` contiguous `LgjOpDesc` — the same contract the live export + // relies on, exercised here only from in-crate tests that build a real + // slice. + let ops: &[LgjOpDesc] = unsafe { std::slice::from_raw_parts(ops, n_ops as usize) }; + + if let Err(e) = validate_plan(&pattern, ops) { + return e; + } + + let n_rows = pattern.n_rows; + let n_words = mask_words_for(n_rows) as usize; + + // Accumulate into scratch, then publish. Two consequences, both wanted: + // dst_mask is written exactly once, and an error at any point leaves it + // byte-for-byte as it was. + let mut acc = vec![0u64; n_words]; + // "the accumulator starts as all rows set" (§7). + for w in acc.iter_mut() { + *w = u64::MAX; + } + clear_tail_bits(&mut acc, n_rows); + let mut scratch = vec![0u64; n_words]; + + for op in ops { + let lane = match lane_view(&pattern, op.lane_id) { + Ok(l) => l, + Err(e) => return e, + }; + if let Err(e) = + kernels::eval_predicate(path, op.op, op.operand, &lane, n_rows, &mut scratch) + { + return e; + } + if let Err(e) = kernels::combine_into(path, op.combine, &mut acc, &scratch) { + return e; + } + } + clear_tail_bits(&mut acc, n_rows); + let count = kernels::popcount(path, &acc); + + let mut g = match mask.write_mask() { + Some(g) => g, + None => return LGJ_ERR_WRONG_RESOURCE_KIND, + }; + if g.words.len() != acc.len() { + return LGJ_ERR_MASK_LENGTH_MISMATCH; + } + g.words.copy_from_slice(&acc); + drop(g); + + // SAFETY: non-null (checked above); written only on the success path. + unsafe { *out_count = count }; + LGJ_OK +} + +/// One evaluation's full observable result: everything the differential +/// compares. Bundling them means a test cannot silently check two of three. +#[derive(Debug, PartialEq, Eq)] +pub(super) struct Outcome { + pub status: i32, + pub count: u64, + pub words: Vec, +} + +/// Run the LIVE export over `ops`, publishing into a freshly-created mask. +/// +/// `out_count` is pre-seeded with a sentinel so "not written on failure" is +/// observable rather than assumed — a zero would be indistinguishable from a +/// legitimately-empty result. +pub(super) fn eval_live(pattern: u64, ops: &[LgjOpDesc], dst: u64) -> Outcome { + let mut count = 12345u64; + let status = call::plan_eval( + pattern, + ops.as_ptr(), + ops.len() as u32, + dst, + &mut count as *mut u64, + ); + Outcome { + status, + count, + words: read_words(dst), + } +} + +/// Run the FROZEN oracle over `ops`, publishing into `dst`. +pub(super) fn eval_legacy(pattern: u64, ops: &[LgjOpDesc], dst: u64) -> Outcome { + let mut count = 12345u64; + let status = legacy_plan_eval_impl( + pattern, + ops.as_ptr(), + ops.len() as u32, + dst, + &mut count as *mut u64, + Path::Simd, + ); + Outcome { + status, + count, + words: read_words(dst), + } +} + +/// The differential itself: live and frozen must agree on status, count, and +/// every word INCLUDING the tail, over two independently-created masks. +/// +/// Two masks rather than one is deliberate: sharing a destination would let a +/// path that accumulates INTO its destination agree with one that overwrites, +/// because the second run would start from the first's answer. +pub(super) fn assert_agrees(pattern: u64, ops: &[LgjOpDesc], what: &str) -> Outcome { + let dst_live = mask(pattern, LGJ_MASK_INIT_EMPTY); + let dst_legacy = mask(pattern, LGJ_MASK_INIT_EMPTY); + let live = eval_live(pattern, ops, dst_live); + let legacy = eval_legacy(pattern, ops, dst_legacy); + assert_eq!( + live, legacy, + "{what}: the live path and the frozen pre-PR4 oracle disagree" + ); + live +} + +#[test] +fn the_frozen_oracle_agrees_with_the_live_path_on_a_plain_and_plan() { + let p = open(1000, 7); + let ops = [ + op(LGJ_OP_GT_I32, LANE_VALUES, 0, LGJ_COMBINE_AND), + op(LGJ_OP_EQ_U32, LANE_CLASSES, 1, LGJ_COMBINE_AND), + ]; + let out = assert_agrees(p, &ops, "two-op AND plan"); + // Anti-vacuity: an oracle that agreed because both answered "nothing" + // would prove nothing at all. + assert!( + out.count > 0 && out.count < 1000, + "the fixture must select a proper subset, got {}", + out.count + ); +} diff --git a/native/lgj-abi/src/exports/tests/pr4_matrix.rs b/native/lgj-abi/src/exports/tests/pr4_matrix.rs new file mode 100644 index 0000000..02b99e8 --- /dev/null +++ b/native/lgj-abi/src/exports/tests/pr4_matrix.rs @@ -0,0 +1,373 @@ +//! PR4's C1 falsifier matrix. +//! +//! `pr4_equivalence.rs` proves the frozen oracle agrees with the live path on +//! ONE two-op AND plan. That is not the matrix: PR4 replaces `plan_eval_impl`'s +//! opcode loop with a lowering (`&[LgjOpDesc] -> mask_risc::Program`) plus +//! `mask_risc::execute`, and a bug in the lowering can be scoped to exactly the +//! shapes this file checks and nothing the single example above would ever +//! reach: +//! +//! - an opcode wired correctly only for the UNGATED (position-0) case, wrong +//! once it sits behind a real AND-narrowed or OR-widened accumulator +//! (`every_opcode_agrees_with_the_oracle_at_every_position_in_a_plan`); +//! - the TCAM op's packed `(care << 32) | pattern` operand with its two +//! halves swapped — invisible whenever `pattern == care` +//! (`the_tcam_operand_halves_are_not_interchangeable`); +//! - a plan whose accumulator quietly bottoms out or saturates partway +//! through, so ops after that point run unexercised even though the FINAL +//! count still looks ordinary +//! (`a_non_degenerate_plan_stays_non_degenerate_at_every_prefix`); +//! - a combine vector this crate's existing plans never happened to use — the +//! full `{AND, OR}^n` product rather than just all-AND or AND-then-OR +//! (`combine_mode_products_agree_with_the_oracle`). +//! +//! Every plan here is still run through [`assert_agrees`], so the oracle +//! comparison is the same one `pr4_equivalence.rs` establishes — this file +//! only widens WHICH plans get compared. + +use super::pr4_equivalence::assert_agrees; +use super::*; + +/// Render a combine value as the word its constant names, for failure +/// messages in [`combine_mode_products_agree_with_the_oracle`]. +fn combine_name(c: u32) -> &'static str { + if c == LGJ_COMBINE_AND { + "AND" + } else { + "OR" + } +} + +/// FAILS IF: an opcode's `LgjOpDesc.op -> predicate kernel` mapping is wired +/// correctly only for the UNGATED, position-0 case. Because `plan_eval_impl` +/// now routes through `mask_risc::execute` and AND- vs. OR-narrowing may take +/// different branches inside `kernels::combine_into`, a mis-map could be +/// scoped to exactly one gated arm — passing at position 0 (a single op +/// against the all-ones seed IS the predicate) and disagreeing with the +/// oracle only once a real accumulator sits in front of it. +#[test] +fn every_opcode_agrees_with_the_oracle_at_every_position_in_a_plan() { + let n = 1000u64; + let p = open(n, 33); + + // (label, opcode, lane, operand). Every triple except `GT_I32` is the + // exact one `every_minor_11_opcode_runs_on_both_plan_paths_and_they_agree` + // (and its eq/ne-partition arm, both in this same file) already measured + // as a proper subset on THIS (n=1000, seed=33) fixture — SplitMix64 is + // deterministic, so the same (n, seed) always regenerates the identical + // lanes. `GT_I32` predates ABI minor 11 and is checked fresh below. + let cases: [(&str, u32, u32, i64); 9] = [ + ("EQ_U32", LGJ_OP_EQ_U32, LANE_CLASSES, 7), + ("GT_I32", LGJ_OP_GT_I32, LANE_VALUES, 100), + ("NE_U32", LGJ_OP_NE_U32, LANE_CLASSES, 7), + ("EQ_I32", LGJ_OP_EQ_I32, LANE_VALUES, 100), + ("NE_I32", LGJ_OP_NE_I32, LANE_VALUES, 100), + ("LT_I32", LGJ_OP_LT_I32, LANE_VALUES, 100), + ("LE_I32", LGJ_OP_LE_I32, LANE_VALUES, 100), + ("GE_I32", LGJ_OP_GE_I32, LANE_VALUES, 100), + ( + "TERNARY_MATCH_U32", + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + tcam_operand(5, 0b111), + ), + ]; + + for (label, opcode, lane, operand) in cases { + // `other` is a different opcode on the OTHER lane, so a plan that + // only ever evaluated one of the two ops would still disagree with + // one that genuinely evaluated both. + let other = if lane == LANE_VALUES { + op(LGJ_OP_NE_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND) + } else { + op(LGJ_OP_LT_I32, LANE_VALUES, 100, LGJ_COMBINE_AND) + }; + + // Position 0: the op alone. `combine` MUST be AND here — seeded with + // OR, the all-ones accumulator stays all-ones regardless of the + // predicate, which would fail the anti-vacuity check below for every + // opcode instead of measuring anything about this one. + let pos0 = assert_agrees( + p, + &[op(opcode, lane, operand, LGJ_COMBINE_AND)], + &format!("{label} at position 0"), + ); + assert!( + pos0.count > 0 && pos0.count < n, + "{label} at position 0 selected {} of {n} rows — not a proper subset, \ + so its gated arms below cannot be discriminating either", + pos0.count + ); + + // Position 1, gated by an AND: [AND other, AND subject]. + assert_agrees( + p, + &[other, op(opcode, lane, operand, LGJ_COMBINE_AND)], + &format!("{label} at position 1, AND-gated"), + ); + + // Position 1, gated by an OR: [AND other, OR subject]. + assert_agrees( + p, + &[other, op(opcode, lane, operand, LGJ_COMBINE_OR)], + &format!("{label} at position 1, OR-gated"), + ); + } + + assert_eq!(lgj_close(p), LGJ_OK); +} + +/// FAILS IF: `LGJ_OP_TERNARY_MATCH_U32`'s packed `operand` can have its two +/// 32-bit halves swapped in the mask-risc lowering (`(pattern << 32) | care` +/// instead of the documented `(care << 32) | pattern`) without changing the +/// observable result. That bug is invisible whenever `pattern == care` — +/// swapping two equal halves is a no-op — which is exactly the shape of the +/// pre-existing `the_tcam_opcodes_packed_operand_decodes_both_halves` fixture +/// (`tcam_operand(0b1000, 0b1000)`). Every arm here uses `pattern != care`. +#[test] +fn the_tcam_operand_halves_are_not_interchangeable() { + let n = 1000u64; + let p = open(n, 33); + let other = op(LGJ_OP_LT_I32, LANE_VALUES, 100, LGJ_COMBINE_AND); + + // (label, pattern, care) — both pairs asymmetric, so a half-swap changes + // the selected population instead of silently reproducing it. + let cases: [(&str, u32, u32); 2] = [ + ("pattern=5 care=0b111", 5, 0b111), + ("pattern=7 care=0b1000", 7, 0b1000), + ]; + + for (label, pattern, care) in cases { + assert_ne!( + pattern, care, + "{label}: a pattern==care fixture cannot see a half-swap" + ); + let operand = tcam_operand(pattern, care); + + let pos0 = assert_agrees( + p, + &[op( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + operand, + LGJ_COMBINE_AND, + )], + &format!("{label} at position 0"), + ); + assert!( + pos0.count > 0 && pos0.count < n, + "{label} at position 0 selected {} of {n} rows — not a proper subset", + pos0.count + ); + + assert_agrees( + p, + &[ + other, + op( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + operand, + LGJ_COMBINE_AND, + ), + ], + &format!("{label} at position 1, AND-gated"), + ); + + assert_agrees( + p, + &[ + other, + op( + LGJ_OP_TERNARY_MATCH_U32, + LANE_CLASSES, + operand, + LGJ_COMBINE_OR, + ), + ], + &format!("{label} at position 1, OR-gated"), + ); + } + + assert_eq!(lgj_close(p), LGJ_OK); +} + +/// FAILS IF: an implementation only guards against a fully-degenerate FINAL +/// accumulator (all-zero or all-ones), so ops sitting after the accumulator +/// bottoms out at zero (making every later AND a no-op) or saturates to +/// all-ones (making every later OR a no-op) go unexercised without the final +/// count ever looking wrong — a mid-plan collapse can still land the last op +/// on a perfectly ordinary-looking count. +#[test] +fn a_non_degenerate_plan_stays_non_degenerate_at_every_prefix() { + let n = 1000u64; + let p = open(n, 33); + + // A 4-op mixed AND/OR plan. Op 0 must be AND — an OR at position 0 is + // unconditionally degenerate against the all-ones seed (see the + // position-0 comment above) — and from there the combine vector + // genuinely alternates. + let ops = [ + op(LGJ_OP_EQ_U32, LANE_CLASSES, 7, LGJ_COMBINE_AND), + op(LGJ_OP_GT_I32, LANE_VALUES, 100, LGJ_COMBINE_OR), + op(LGJ_OP_LT_I32, LANE_VALUES, 200, LGJ_COMBINE_AND), + op(LGJ_OP_EQ_U32, LANE_CLASSES, 3, LGJ_COMBINE_OR), + ]; + + for k in 1..=ops.len() { + let prefix = &ops[..k]; + let out = assert_agrees(p, prefix, &format!("prefix length {k}")); + assert!( + out.count > 0 && out.count < n, + "prefix length {k} selected {} of {n} rows — a degenerate prefix means \ + op(s) {k}..{} of the full 4-op plan would run with a bottomed-out or \ + saturated accumulator and never really be exercised", + out.count, + ops.len() + ); + } + + assert_eq!(lgj_close(p), LGJ_OK); +} + +/// FAILS IF: combine dispatch is only correct for the two shapes this crate's +/// existing plans have ever used — all-AND, or AND-then-OR — and mishandles +/// any other member of the `{AND, OR}^n` product: most plausibly an OR that +/// is not in the final position, or two ORs in a row (which is provably +/// `n_rows` regardless of predicate content, checked directly below, so +/// getting it wrong would show up as a SHRUNK count rather than merely a +/// mismatch against the oracle). +#[test] +fn combine_mode_products_agree_with_the_oracle() { + let n = 1000u64; + let p = open(n, 33); + + // A FIXED opcode/lane/operand set per arity: only the combine VECTOR + // varies across the sweep, so any divergence is attributable to combine + // dispatch alone. + let two_op: [(u32, u32, i64); 2] = [ + (LGJ_OP_EQ_U32, LANE_CLASSES, 7), + (LGJ_OP_GT_I32, LANE_VALUES, 100), + ]; + let three_op: [(u32, u32, i64); 3] = [ + (LGJ_OP_EQ_U32, LANE_CLASSES, 7), + (LGJ_OP_GT_I32, LANE_VALUES, 100), + (LGJ_OP_EQ_U32, LANE_CLASSES, 3), + ]; + + let mut degenerate: Vec = Vec::new(); + let mut non_degenerate = 0u32; + + for bits in 0u32..4 { + // bit i of `bits` selects op i's combiner: 0 = AND, 1 = OR. + let combine: [u32; 2] = [ + if bits & 1 == 0 { + LGJ_COMBINE_AND + } else { + LGJ_COMBINE_OR + }, + if bits & 2 == 0 { + LGJ_COMBINE_AND + } else { + LGJ_COMBINE_OR + }, + ]; + let label = format!( + "n=2 [{}, {}]", + combine_name(combine[0]), + combine_name(combine[1]) + ); + let ops = [ + op(two_op[0].0, two_op[0].1, two_op[0].2, combine[0]), + op(two_op[1].0, two_op[1].1, two_op[1].2, combine[1]), + ]; + let out = assert_agrees(p, &ops, &label); + if out.count == 0 || out.count == n { + degenerate.push(format!("{label} -> {}", out.count)); + } else { + non_degenerate += 1; + } + // Structurally guaranteed regardless of what the two predicates + // select: once the accumulator starts all-ones, `acc |= x` can never + // shrink it — an all-OR vector is always `n_rows`, full stop. + if combine == [LGJ_COMBINE_OR, LGJ_COMBINE_OR] { + assert_eq!( + out.count, n, + "{label}: an all-OR vector must stay saturated at all-ones" + ); + } + } + + for bits in 0u32..8 { + let combine: [u32; 3] = [ + if bits & 1 == 0 { + LGJ_COMBINE_AND + } else { + LGJ_COMBINE_OR + }, + if bits & 2 == 0 { + LGJ_COMBINE_AND + } else { + LGJ_COMBINE_OR + }, + if bits & 4 == 0 { + LGJ_COMBINE_AND + } else { + LGJ_COMBINE_OR + }, + ]; + let label = format!( + "n=3 [{}, {}, {}]", + combine_name(combine[0]), + combine_name(combine[1]), + combine_name(combine[2]) + ); + let ops = [ + op(three_op[0].0, three_op[0].1, three_op[0].2, combine[0]), + op(three_op[1].0, three_op[1].1, three_op[1].2, combine[1]), + op(three_op[2].0, three_op[2].1, three_op[2].2, combine[2]), + ]; + let out = assert_agrees(p, &ops, &label); + if out.count == 0 || out.count == n { + degenerate.push(format!("{label} -> {}", out.count)); + } else { + non_degenerate += 1; + } + if combine == [LGJ_COMBINE_OR, LGJ_COMBINE_OR, LGJ_COMBINE_OR] { + assert_eq!( + out.count, n, + "{label}: an all-OR vector must stay saturated at all-ones" + ); + } + } + + // Anti-vacuity for the sweep as a whole: the fixed opcode/operand set + // must exercise SOME real composition, or every oracle-agreement check + // above would have passed by comparing two identically-degenerate + // answers rather than by exercising combine dispatch. + // + // The bound is MEASURED, not chosen. On this (n=1000, seed=33) fixture + // the 12 sweeps are: n=2 [AND,AND] 39, [OR,AND] 498, [AND,OR] 524, + // [OR,OR] 1000; n=3 [AND,AND,AND] 0, [OR,AND,AND] 40, [AND,OR,AND] 40, + // [OR,OR,AND] 73, [AND,AND,OR] 112, [OR,AND,OR] 531, [AND,OR,OR] 557, + // [OR,OR,OR] 1000. Exactly THREE are degenerate and two of those are + // degenerate by construction rather than by fixture — an all-OR vector + // saturates at `n` no matter what the predicates select, which the two + // structural `assert_eq!`s above already pin separately. So the real + // discriminating power of this sweep is 9 of 12, and that is the bound. + // + // `> 0` was the original, and it is the weak form this repo's own + // falsifiability rule names: it passes when ELEVEN of twelve sweeps have + // gone degenerate, which is precisely the state where the oracle + // comparisons stop comparing anything. A fixture drift that hollowed the + // sweep out would read as green under it and red under this. + assert_eq!( + non_degenerate, 9, + "the combine sweep's discriminating power moved: expected 9 of 12 \ + non-degenerate, got {non_degenerate} (degenerate: {degenerate:?}). \ + Two of the three degenerate sweeps are the structural all-OR \ + saturations; a third is fixture-dependent. Re-measure before re-pinning" + ); + + assert_eq!(lgj_close(p), LGJ_OK); +} diff --git a/native/lgj-abi/src/exports/tests/pr4_seed.rs b/native/lgj-abi/src/exports/tests/pr4_seed.rs new file mode 100644 index 0000000..5917477 --- /dev/null +++ b/native/lgj-abi/src/exports/tests/pr4_seed.rs @@ -0,0 +1,328 @@ +//! PR4 C1 seed: the leading-OR falsifier matrix. +//! +//! `abi.md` §7's contract is unambiguous: "the accumulator starts as all rows +//! set." A lowering can still get this wrong by taking a shortcut that reads +//! naturally but ignores the FIRST op's own `combine` field — something like +//! "the accumulator starts all-ones, so the first AND is a no-op; just let +//! `acc := pred_0` directly." Applied unconditionally, that shortcut is +//! correct for `combine = AND` and silently WRONG for `combine = OR`, because +//! `ALL | pred_0 == ALL` regardless of what `pred_0` selects, while +//! `acc := pred_0` gives `pred_0`'s own — generally smaller — selection. +//! +//! Every plan in this crate's pre-existing test suite begins with `AND` (the +//! only two `LGJ_COMBINE_OR` sites are both the SECOND op — see +//! `the_executor_and_the_row_oracle_agree_bit_for_bit` (renamed in PR4 from +//! `simd_and_scalar_plans_agree_bit_for_bit`) and `or_plans_widen` in +//! `exports.rs`), so that wrong shortcut passes the entire suite as it stands +//! today. These tests exist to make it fail: they are pure ADDITIONS, and per +//! this module's freeze discipline they must be GREEN against the current +//! (pre-PR4) implementation, which does not take the shortcut — `plan_eval_impl` +//! seeds `acc` to all-ones and clears its tail unconditionally, before +//! looking at any op at all. +//! +//! T1-T3 pin the OR side across opcodes, operands, plan lengths and word +//! boundaries. T4 pins the algebra of OR followed by AND. T5 is the paired +//! silent half — without it, an implementation that returned every row for +//! EVERY single-op plan (AND included) would pass T1-T4 while being wrong. + +use super::pr4_equivalence::assert_agrees; +use super::*; + +/// The bit pattern for "every one of `n_rows` rows is selected": every word +/// full of ones except the last, whose bits at row index `>= n_rows` are +/// cleared. Built from the two primitives every write path in this crate +/// already uses for exactly this purpose (`plan_eval_impl`'s own seeding +/// step is `for w in acc.iter_mut() { *w = u64::MAX; } clear_tail_bits(&mut +/// acc, n_rows);`) — reusing them here means this expected value can never +/// silently drift from the tail-bit convention the implementation itself is +/// built on. +fn all_rows_words(n_rows: u64) -> Vec { + let mut w = vec![u64::MAX; mask_words_for(n_rows) as usize]; + clear_tail_bits(&mut w, n_rows); + w +} + +/// True iff no bit at row index `>= n_rows` survives in the last word. +/// +/// Written as its own small, independent formula (not by delegating to +/// [`clear_tail_bits`] and diffing) so it can catch a bug at the ONE call +/// site that matters — the plan's published `dst_mask` — rather than merely +/// re-confirming that the helper above computes what it always computes. +/// +/// `n_rows % 64 == 0` needs its own branch: when the last word is exactly +/// full there is no tail region at all, and shifting by zero would (wrongly) +/// demand the whole word be zero — the same edge case `clear_tail_bits`'s own +/// `if used != 0` guards against (`abi.rs`'s `tail_bits_are_cleared` test +/// pins it: `clear_tail_bits(&mut w2, 128)` is a no-op, not a zeroing). +fn tail_is_clean(words: &[u64], n_rows: u64) -> bool { + let rem = n_rows % 64; + if rem == 0 { + return true; + } + match words.last() { + Some(&last) => last >> rem == 0, + None => true, + } +} + +/// The 8 non-ternary comparison opcodes, paired with the lane +/// `opcode_required_kind` (`abi.rs`) requires of them: the three `U32` +/// opcodes read `LANE_CLASSES` (`0..=15`), the five `I32` opcodes read +/// `LANE_VALUES` (`-150..=361`) — both ranges are `fixture.rs`'s own +/// documented generation contract. +/// +/// `LGJ_OP_TERNARY_MATCH_U32` is deliberately excluded: its `operand` packs +/// two 32-bit halves (`(care << 32) | pattern`) rather than a plain needle, +/// which needs its own construction orthogonal to this seed. +const LEADING_OR_OPCODES: [(u32, u32); 8] = [ + (LGJ_OP_EQ_U32, LANE_CLASSES), + (LGJ_OP_GT_I32, LANE_VALUES), + (LGJ_OP_NE_U32, LANE_CLASSES), + (LGJ_OP_EQ_I32, LANE_VALUES), + (LGJ_OP_NE_I32, LANE_VALUES), + (LGJ_OP_LT_I32, LANE_VALUES), + (LGJ_OP_LE_I32, LANE_VALUES), + (LGJ_OP_GE_I32, LANE_VALUES), +]; + +/// Three operands per opcode, chosen against the documented domains +/// (`fixture.rs`: classes `0..=15`, values `-150..=361`) to select roughly +/// nothing, roughly half, and roughly everything of the fixture's rows — so +/// a single coincidentally-safe operand cannot hide the shortcut. +/// +/// The four inequality opcodes over `LANE_VALUES` (`GT`/`LT`/`LE`/`GE`) hit +/// all three bands EXACTLY: the domain has 512 values, and the midpoint +/// operands below (105/106) split it into two halves of 256 values each. +/// `EQ`/`NE` cannot naturally reach "half" over a 16- or 512-valued domain (a +/// single (in)equality selects at most ~1/16 or ~1/512, never ~50%), so their +/// middle entry is a second, different in-range value instead of a forced +/// "half" — still a genuinely distinct predicate, which is what the property +/// under test actually needs. +fn leading_or_operands(op: u32) -> [i64; 3] { + match op { + // class == X: 999 matches no row (no class is that high); 7 and 0 + // each match ~1/16 of rows, from opposite ends of the domain. + LGJ_OP_EQ_U32 => [999, 7, 0], + // class != X: 999 matches EVERY row (no class equals it); 7 and 0 + // each exclude ~1/16, matching ~15/16 from opposite ends. + LGJ_OP_NE_U32 => [999, 7, 0], + // value == X: 1000 is out of range (matches nothing); 105 and -150 + // (the domain's own minimum) each match only a handful of rows. + LGJ_OP_EQ_I32 => [1000, 105, -150], + // value != X: 1000 is out of range (matches EVERY row); 105 and + // -150 each exclude only a handful, matching nearly every row. + LGJ_OP_NE_I32 => [1000, 105, -150], + // value > X: 1000 matches nothing (max is 361); 105 splits the + // domain exactly in half (values 106..=361, 256 of 512); -1000 + // matches every row (min is -150). + LGJ_OP_GT_I32 => [1000, 105, -1000], + // value < X: mirror of GT_I32, with the halves and extremes swapped. + LGJ_OP_LT_I32 => [-1000, 106, 1000], + // value <= X: same shape as LT_I32 (the boundary shifts inclusion + // by one, which is irrelevant to "roughly nothing/half/everything"). + LGJ_OP_LE_I32 => [-1000, 105, 1000], + // value >= X: mirror of LE_I32. + LGJ_OP_GE_I32 => [1000, 106, -1000], + _ => unreachable!("leading_or_operands: opcode {op} is not in LEADING_OR_OPCODES"), + } +} + +/// FAILS IF: a lowering takes the shortcut "the accumulator starts all-ones, +/// so the first AND is a no-op; let `acc := pred_0` directly" and applies it +/// WITHOUT checking that op's own `combine` field. Every plan below has +/// exactly one op, with `combine = OR` — per `abi.md` §7 ("the accumulator +/// starts as all rows set"), `ALL | pred_0 == ALL` regardless of what +/// `pred_0` selects, so the correct answer is `n` for every operand tried. +/// The wrong shortcut instead publishes `pred_0` itself, which is wrong for +/// every operand below except the handful chosen to already select +/// everything — the spread of selectivities is what stops those from +/// masking the rest. +#[test] +fn a_leading_or_returns_every_row_whatever_the_predicate_says() { + let n = 1000u64; + let p = open(n, 0xA11C_E000); + let expected = all_rows_words(n); + + for &(opcode, lane) in &LEADING_OR_OPCODES { + for operand in leading_or_operands(opcode) { + let plan = [op(opcode, lane, operand, LGJ_COMBINE_OR)]; + let out = assert_agrees( + p, + &plan, + &format!("leading OR, opcode={opcode} lane={lane} operand={operand}"), + ); + assert_eq!( + out.status, LGJ_OK, + "opcode={opcode} operand={operand}: a well-formed single-op plan must validate" + ); + assert_eq!( + out.count, n, + "opcode={opcode} operand={operand}: a leading OR must select every row \ + (got {}), whatever the predicate itself matched", + out.count + ); + assert_eq!( + out.words, expected, + "opcode={opcode} operand={operand}: a leading OR must publish an all-ones, \ + tail-clean mask, not the predicate's own bits" + ); + } + } + + lgj_close(p); +} + +/// FAILS IF: the seed used ahead of a leading OR is a raw memset of every +/// WORD to all-ones that skips the tail-clearing the normal path always +/// applies before publishing — the composed defect this whole seed matrix +/// exists to catch. At a word boundary that shows up as two DIFFERENT wrong +/// numbers: the popcount over-counts by the tail width (`n_words * 64` +/// instead of `n`), and the raw published word carries garbage bits past row +/// `n`. Asserting both means a failure names which one actually fired. +/// +/// The predicate (`value > 1000`) is chosen to select LITERALLY ZERO rows for +/// any seed — no fixture value ever exceeds 361 (`fixture.rs`'s documented +/// range) — so a DIFFERENT shortcut (`acc := pred_0`, T1's target) would also +/// be caught here, with its own distinct wrong number (`0` instead of `n`). +#[test] +fn a_leading_or_at_a_word_boundary_keeps_its_tail_clean() { + for n in [1u64, 63, 65, 127, 128, 999, 4097] { + let p = open(n, 0x5EED_0000 | n); + let ops = [op(LGJ_OP_GT_I32, LANE_VALUES, 1000, LGJ_COMBINE_OR)]; + let out = assert_agrees(p, &ops, &format!("word-boundary leading OR, n={n}")); + + assert_eq!( + out.count, n, + "n={n}: a leading OR must select every row, not a rounded-up-to-64 word count \ + (got {})", + out.count + ); + assert!( + tail_is_clean(&out.words, n), + "n={n}: dirty tail bits survived past row {n} (last word = {:#x})", + out.words.last().copied().unwrap_or(0) + ); + + lgj_close(p); + } +} + +/// FAILS IF: an all-OR plan is evaluated as if only its first op existed — +/// each op below uses a DIFFERENT opcode, lane and operand, so a bug that +/// silently reused op[0]'s fields for every later position (rather than +/// genuinely reading each one) cannot hide behind identical entries. The +/// contract's own algebra says the answer must be `n` regardless of length: +/// the first OR forces the accumulator to ALL, and `ALL | anything == ALL` +/// for every op after it. +#[test] +fn an_all_or_plan_returns_every_row_however_many_ops_it_has() { + let n = 3000u64; + let p = open(n, 0x0A11_0A11); + let expected = all_rows_words(n); + + // Four distinct, individually well-formed predicates spanning both + // required lane kinds, so a 2/3/4-op prefix of this pool never repeats + // an (opcode, lane, operand) triple. + let pool = [ + op(LGJ_OP_EQ_U32, LANE_CLASSES, 3, LGJ_COMBINE_OR), + op(LGJ_OP_GT_I32, LANE_VALUES, 100, LGJ_COMBINE_OR), + op(LGJ_OP_NE_U32, LANE_CLASSES, 9, LGJ_COMBINE_OR), + op(LGJ_OP_LT_I32, LANE_VALUES, -50, LGJ_COMBINE_OR), + ]; + + for len in 2..=4usize { + let plan = &pool[..len]; + let out = assert_agrees(p, plan, &format!("all-OR plan, len={len}")); + assert_eq!( + out.count, n, + "len={len}: an all-OR plan must select every row regardless of its length \ + (got {})", + out.count + ); + assert_eq!( + out.words, expected, + "len={len}: an all-OR plan must publish an all-ones, tail-clean mask" + ); + } + + lgj_close(p); +} + +/// FAILS IF: the leading OR contributes anything beyond forcing the +/// accumulator to ALL before the AND narrows it. `[OR p0, AND p1]` and +/// `[AND p1]` alone must be bit-for-bit indistinguishable — status, count, +/// AND every published word — because a leading OR that reaches ALL (T1) +/// followed by an AND (T5) is, by the contract's own algebra, exactly the +/// AND alone. `p0` is deliberately arbitrary: ANY well-formed predicate +/// there must give the identical final answer, since a leading OR erases +/// whatever it selects. +#[test] +fn an_or_then_and_plan_reduces_to_the_and_predicate_alone() { + let n = 5000u64; + let p = open(n, 0x0FF1_CE00); + + let p0 = op(LGJ_OP_EQ_U32, LANE_CLASSES, 5, LGJ_COMBINE_OR); + let p1 = op(LGJ_OP_GT_I32, LANE_VALUES, 100, LGJ_COMBINE_AND); + + let plan_or_then_and = [p0, p1]; + let plan_and_alone = [p1]; + + let out_combined = assert_agrees(p, &plan_or_then_and, "[OR p0, AND p1]"); + let out_alone = assert_agrees(p, &plan_and_alone, "[AND p1] alone"); + + assert_eq!( + out_combined, out_alone, + "a leading OR followed by an AND must equal the AND predicate alone" + ); + // Anti-vacuity: both sides agreeing because both selected everything (or + // nothing) would prove nothing about the ALGEBRA — p1 must genuinely cut + // the population down to a proper subset, or this passes for a broken + // implementation that ignores every op and always returns everything. + assert!( + out_combined.count > 0 && out_combined.count < n, + "fixture must select a proper subset, got {}", + out_combined.count + ); + + lgj_close(p); +} + +/// FAILS IF: a single-op plan treats its op's `combine` field as irrelevant +/// and always publishes "the whole answer" regardless of it — the ONE case +/// where that reading happens to be correct is a single AND, and this test +/// pins it against the fixture's OWN ground truth (computed by direct +/// iteration over `Fixture::values()`, not via any plan machinery) rather +/// than merely against another run of the same code path. +/// +/// This is the paired silent half of T1-T4: without it, an implementation +/// that unconditionally returned every row for EVERY single-op plan — AND +/// included — would satisfy every assertion above while being wrong exactly +/// here, where "every row" is not the answer. +#[test] +fn a_leading_and_is_still_exactly_the_predicate() { + let n = 2500u64; + let seed = 0xDEAD_10CC; + let p = open(n, seed); + let f = Fixture::generate(n, seed).expect("n_rows fits usize on this target"); + + let predicate = op(LGJ_OP_LT_I32, LANE_VALUES, 50, LGJ_COMBINE_AND); + let out = assert_agrees(p, &[predicate], "single AND op"); + + let want = f.values().iter().filter(|&&v| v < 50).count() as u64; + assert_eq!( + out.count, want, + "a single AND op must equal the predicate's own selection exactly, \ + computed independently from the raw fixture values (got {}, want {want})", + out.count + ); + // Anti-vacuity: without this, a broken implementation that returns ALL + // rows unconditionally for every single-op plan would pass by accident — + // this is precisely the case where "all rows" is the WRONG answer. + assert!( + out.count > 0 && out.count < n, + "fixture must select a proper subset, got {}", + out.count + ); + + lgj_close(p); +} diff --git a/native/lgj-abi/src/kernels.rs b/native/lgj-abi/src/kernels.rs index 6ba1f5b..59c7003 100644 --- a/native/lgj-abi/src/kernels.rs +++ b/native/lgj-abi/src/kernels.rs @@ -53,6 +53,95 @@ pub fn simd_gt_i32_to_mask(values: &[i32], threshold: i32, out_words: &mut [u64] ndarray::simd::gt_i32_to_mask(values, threshold, out_words); } +// ── The comparison family (ABI minor >= 11, docs/abi.md §19) ──────────────── +// +// Six more predicates, ZERO new ABI symbols: each is an `LgjOpDesc` op-code, +// so they arrive through `lgj_plan_eval` — fused with any number of siblings +// in one crossing, or alone (`n_ops = 1`, whose all-set accumulator makes the +// result exactly `op0`). `LgjOpDesc.op` is the generalisation vehicle this ABI +// already carries for predicates, and using it is why §1's "growth is a design +// smell" is respected rather than argued around. +// +// Every one is ONE delegation. None re-derives a complement or a tail: the +// facade owns both (`ge` is `lt` complemented WITH the tail re-cleared, `ne_i32` +// is a single two-compare pass so its tail is zero by construction), and +// duplicating that reasoning here would be a second surface to keep in step. + +/// `out_words[i-th bit] = (values[i] != needle)`, fully overwriting. +#[inline] +pub fn simd_ne_u32_to_mask(values: &[u32], needle: u32, out_words: &mut [u64]) { + ndarray::simd::ne_u32_to_mask(values, needle, out_words); +} + +/// `out_words[i-th bit] = (values[i] == needle)`, signed lanes, exact. +#[inline] +pub fn simd_eq_i32_to_mask(values: &[i32], needle: i32, out_words: &mut [u64]) { + ndarray::simd::eq_i32_to_mask(values, needle, out_words); +} + +/// `out_words[i-th bit] = (values[i] != needle)`, signed lanes. +#[inline] +pub fn simd_ne_i32_to_mask(values: &[i32], needle: i32, out_words: &mut [u64]) { + ndarray::simd::ne_i32_to_mask(values, needle, out_words); +} + +/// `out_words[i-th bit] = (values[i] < threshold)`, signed, strict. +/// +/// Exact at `i32::MIN`: the facade lowers it as `threshold > values[i]`, never +/// as `values[i] > threshold - 1`, so the bottom of the range does not wrap. +#[inline] +pub fn simd_lt_i32_to_mask(values: &[i32], threshold: i32, out_words: &mut [u64]) { + ndarray::simd::lt_i32_to_mask(values, threshold, out_words); +} + +/// `out_words[i-th bit] = (values[i] <= threshold)`, signed. +#[inline] +pub fn simd_le_i32_to_mask(values: &[i32], threshold: i32, out_words: &mut [u64]) { + ndarray::simd::le_i32_to_mask(values, threshold, out_words); +} + +/// `out_words[i-th bit] = (values[i] >= threshold)`, signed. +#[inline] +pub fn simd_ge_i32_to_mask(values: &[i32], threshold: i32, out_words: &mut [u64]) { + ndarray::simd::ge_i32_to_mask(values, threshold, out_words); +} + +/// TCAM over a contiguous `U32` lane: `((values[i] ^ pattern) & care) == 0` — +/// equality on the bits `care` selects, "don't care" elsewhere. +/// +/// `care == 0` matches every element; `care == u32::MAX` is exact equality. +/// The facade lowers it as ONE `ternlog::` per 16 lanes plus a +/// compare-against-zero, which is why this is a genuine primitive rather than +/// an XOR followed by an AND followed by a compare. +#[inline] +pub fn simd_ternary_match_u32_to_mask( + values: &[u32], + pattern: u32, + care: u32, + out_words: &mut [u64], +) { + ndarray::simd::ternary_match_u32_to_mask(values, pattern, care, out_words); +} + +/// The 64-bit sibling of [`simd_ternary_match_u32_to_mask`]. +/// +/// **Deliberately not reachable across the membrane, and that is a NAMED GAP +/// rather than an oversight** (`docs/abi.md` §19): its `pattern` + `care` is +/// 128 bits and `LgjOpDesc.operand` carries 64, so it cannot become an op-code +/// without growing the descriptor — which would not be an additive change. +/// It lands here, tested, so the capability exists at the substrate tier the +/// moment a caller earns it, which is the direction the Missing-capability +/// STOP rule requires. +#[inline] +pub fn simd_ternary_match_u64_to_mask( + values: &[u64], + pattern: u64, + care: u64, + out_words: &mut [u64], +) { + ndarray::simd::ternary_match_u64_to_mask(values, pattern, care, out_words); +} + /// `dst = a & b`. `dst` must not alias `a` or `b` (Rust's borrow rules enforce /// it here; the aliasing ABI cases route to the `_assign` forms instead). #[inline] @@ -94,6 +183,67 @@ pub fn simd_mask_andnot_assign(dst: &mut [u64], src: &[u64]) { ndarray::simd::mask_andnot_assign(dst, src); } +// ── XOR / NOT / ANY / ALL (ABI minor >= 11, docs/abi.md §19) ──────────────── +// +// All four are substrate-tier primitives here. Only XOR and NOT are reachable +// across the membrane, and neither as its own symbol: both are immediates to +// `lgj_mask_ternlog` (`XOR2 = 0x3C`, `NOT_A = 0x0F`). ANY and ALL are +// deliberately NOT exported — see `docs/abi.md` §19's "what is not a symbol, +// and why": both are exact functions of `lgj_mask_count`'s answer, cost the +// same one crossing and the same one pass over the mask words, and so buy a +// caller nothing that it does not already have. + +/// `dst = a ^ b` — symmetric difference. +/// +/// XOR preserves the trailing-zero guarantee iff both inputs conform +/// (`0 ^ 0 = 0`); this crate re-establishes it defensively anyway at every +/// export, exactly as `lgj_mask_andnot` documents for the complement. +#[inline] +pub fn simd_mask_xor(a: &[u64], b: &[u64], dst: &mut [u64]) { + ndarray::simd::mask_xor(a, b, dst); +} + +/// `dst ^= src`. +#[inline] +pub fn simd_mask_xor_assign(dst: &mut [u64], src: &[u64]) { + ndarray::simd::mask_xor_assign(dst, src); +} + +/// `dst = !src` over `n_rows` rows — the TAIL-AWARE complement. +/// +/// `n_rows` is a parameter and not an inference, because it has to be: a plain +/// `!` over the words sets every bit past the population, and the tail's zero +/// is normative here (`lgj_mask_count` reads it). This is the one mask-algebra +/// primitive whose correctness cannot be stated without the row count. +#[inline] +pub fn simd_mask_not(src: &[u64], n_rows: usize, dst: &mut [u64]) { + ndarray::simd::mask_not(src, n_rows, dst); +} + +/// `dst = !dst` over `n_rows` rows, in place. Same tail contract as +/// [`simd_mask_not`]. +#[inline] +pub fn simd_mask_not_assign(dst: &mut [u64], n_rows: usize) { + ndarray::simd::mask_not_assign(dst, n_rows); +} + +/// `true` iff any row is selected — the `EXISTS` terminal. +/// +/// Takes no `n_rows`: it relies on the normative tail-zero contract every +/// writer in this crate re-establishes, which is exactly why that contract is +/// defended defensively rather than inherited. +#[inline] +pub fn simd_mask_any(words: &[u64]) -> bool { + ndarray::simd::mask_any(words) +} + +/// `true` iff every one of the first `n_rows` rows is selected. A zero-row +/// population is vacuously `true`. +#[inline] +pub fn simd_mask_all(words: &[u64], n_rows: usize) -> bool { + ndarray::simd::mask_all(words, n_rows) +} + /// The truth-table immediates for [`simd_mask_ternlog_assign`] — re-exported /// so a call site in `exports` names `kernels::ternlog::AND3`, never reaching /// past this module for its SIMD vocabulary (abi.md §8). @@ -112,12 +262,145 @@ pub fn simd_mask_ternlog_assign(a: &mut [u64], b: &[u64], c: &[u ndarray::simd::mask_ternlog_assign::(a, b, c); } +/// `dst = ternlog::(a, b, c)` — the out-of-place form, `dst` distinct +/// from all three operands. +/// +/// Not reached by any export today (`lgj_mask_ternlog` snapshots its operands +/// and runs the in-place form, which is what makes every aliasing case one +/// code path — see that symbol's doc). It lands here because the substrate +/// tier is where a primitive belongs, and because the parity tests below need +/// the non-aliasing spelling to check the aliasing one against. +#[inline] +pub fn simd_mask_ternlog(a: &[u64], b: &[u64], c: &[u64], dst: &mut [u64]) { + ndarray::simd::mask_ternlog::(a, b, c, dst); +} + +/// One monomorphisation of [`simd_mask_ternlog_assign`], as a value. +type TernlogAssignFn = fn(&mut [u64], &[u64], &[u64]); + +/// Expand one row of sixteen consecutive immediates. +macro_rules! ternlog_row16 { + ($base:expr) => { + [ + simd_mask_ternlog_assign::<{ $base }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 1 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 2 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 3 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 4 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 5 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 6 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 7 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 8 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 9 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 10 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 11 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 12 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 13 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 14 }> as TernlogAssignFn, + simd_mask_ternlog_assign::<{ $base + 15 }> as TernlogAssignFn, + ] + }; +} + +/// Every immediate, as a function pointer to its own monomorphisation. +/// +/// # Why a TABLE and not a runtime truth-table evaluation +/// +/// `IMM` is a *runtime* `u8` at the membrane and a *compile-time* `i32` at the +/// facade, and closing that gap is this table's entire job. The tempting +/// alternative — decompose the immediate into its eight minterms and OR them +/// word-wise — is wrong here for a stated reason, not a preference: `ndarray`'s +/// POLYFILL LAW puts truth-table specialisation *inside the backend file* +/// (`VPTERNLOGQ` is one instruction per 512 bits on AVX-512, and each backend +/// owns its own realization). A minterm decomposition would evaluate the table +/// HERE, in this crate, in ~24 word ops instead of one — which is both slower +/// and a locally-written SIMD abstraction, the thing `abi.md` §8 forbids +/// outright. Reaching the const generic is what keeps the specialisation where +/// the law puts it. +/// +/// Indexed `[imm >> 4][imm & 15]`, so it is a `[[_; 16]; 16]` rather than a +/// flat 256 purely to keep the source a 16-line list instead of a 256-line one. +/// A `static` rather than a `const` so the 256 instantiations exist once in the +/// artifact instead of being re-materialised at every use site. +static TERNLOG_ASSIGN: [[TernlogAssignFn; 16]; 16] = [ + ternlog_row16!(0), + ternlog_row16!(16), + ternlog_row16!(32), + ternlog_row16!(48), + ternlog_row16!(64), + ternlog_row16!(80), + ternlog_row16!(96), + ternlog_row16!(112), + ternlog_row16!(128), + ternlog_row16!(144), + ternlog_row16!(160), + ternlog_row16!(176), + ternlog_row16!(192), + ternlog_row16!(208), + ternlog_row16!(224), + ternlog_row16!(240), +]; + +/// `a = ternlog::(a, b, c)` with `imm` chosen at RUNTIME. +/// +/// Total over `0..=255`: every 3-input Boolean function of three masks has an +/// entry, so this one dispatch is the whole mask-op family — `AND2` (`0xC0`), +/// `OR2` (`0xFC`), `XOR2` (`0x3C`), `AND2_ANDNOT` (`0x30`), `NOT_A` (`0x0F`), +/// `AND3` (`0x80`), `MAJ3` (`0xE8`), and the 249 others. There is no unknown +/// immediate and therefore no rejection path. +/// +/// # Tail +/// +/// The caller MUST clear the tail afterwards. The facade's contract is precise +/// about why: the result's tail is `IMM & 1` replicated, so an ODD table — and +/// `NOT_A` is odd — sets every bit past the population. Every export in this +/// crate calls [`crate::abi::clear_tail_bits`] on the way out for exactly this +/// class of reason. +#[inline] +pub fn simd_mask_ternlog_assign_dyn(imm: u8, a: &mut [u64], b: &[u64], c: &[u64]) { + TERNLOG_ASSIGN[(imm >> 4) as usize][(imm & 0x0F) as usize](a, b, c); +} + /// Sum of `values[i]` over set mask bits, widened to `i64`. #[inline] pub fn simd_masked_sum_i32(values: &[i32], mask_words: &[u64]) -> i64 { ndarray::simd::masked_sum_i32(values, mask_words) } +/// Minimum of `values[i]` over set mask bits; `None` when nothing is selected. +/// +/// `None` is the honest answer, not a sentinel: `i32::MAX` would be +/// indistinguishable from a population that really does contain `i32::MAX`. +/// The membrane carries the distinction out as a separate `out_present` flag +/// (`lgj_reduce_i32`) rather than collapsing it. +#[inline] +pub fn simd_masked_min_i32(values: &[i32], mask_words: &[u64]) -> Option { + ndarray::simd::masked_min_i32(values, mask_words) +} + +/// Maximum of `values[i]` over set mask bits; `None` when nothing is selected. +#[inline] +pub fn simd_masked_max_i32(values: &[i32], mask_words: &[u64]) -> Option { + ndarray::simd::masked_max_i32(values, mask_words) +} + +/// `dst[i] = if mask bit i { a[i] } else { b[i] }` — conditional select +/// (`CASE WHEN`) over a row mask, no compaction. +/// +/// **Substrate-tier only; deliberately NOT an ABI symbol**, and the reason is +/// arithmetic rather than taste: a blend needs TWO `I32` lanes, and neither +/// resource kind this ABI carries has two. The pattern fixture has exactly one +/// (`LANE_VALUES`); the row store has none at all (its lanes are `U8` raw, +/// `U32` classid, `U64` lo64, `U32` hi32). Every legal call would therefore be +/// `blend(m, values, values)`, which is `values` — a symbol whose only +/// reachable invocation is the identity. `docs/abi.md` §19 records the +/// condition that would earn it: a resource carrying two `I32` lanes, or a +/// measured need for the constant-fallback form. +#[inline] +pub fn simd_blend_i32(mask_words: &[u64], a: &[i32], b: &[i32], dst: &mut [i32]) { + ndarray::simd::blend_i32(mask_words, a, b, dst); +} + /// Population count over mask words. /// /// **Reused, not reimplemented** — this already exists in `ndarray` @@ -191,6 +474,53 @@ pub fn simd_rowstore_classid_mask( simd_rowstore_u32_eq_mask(bytes, first_offset, stride_bytes, n_rows, needle, out_words); } +/// Bytes of the V3 content-blind register that follows a facet's classid. +/// +/// DERIVED from the row store's own two constants, never written as `12`. +/// E3 — the geometry has one spelling — is the rule; this is the spelling it +/// points at, and a third literal here is exactly what that rule forbids. +pub const FACET_REGISTER_BYTES: usize = + (crate::rowstore::FACET_BYTES - crate::rowstore::FACET_CLASSID_BYTES) as usize; + +/// THE strided TCAM over the row store: `out_words[row-th bit] = +/// (((register(row, facet) ^ pattern) & care) == 0)`, over the 12-byte V3 +/// content-blind register. +/// +/// This is the first BITWISE-over-payload predicate this crate has ever +/// carried. Every predicate before it compared a whole scalar field — a +/// classid, a value, an id. `care` makes it a PREFIX test: set the bytes that +/// must match, clear the rest, and the answer is "which rows carry this +/// prefix", which is the vertical (HHTL ancestry) axis +/// `.claude/plans/mask-risc-lowering-v1.md` §0 puts the whole mask trie on. +/// `care` all-zero matches every row; `care` all-ones is exact register +/// equality. +/// +/// Routes through `ndarray::simd::ternary_match_strided_to_mask`, which lowers +/// the test to ONE `ternlog::` per 16 records over the lo64 halves +/// plus a hi32 pass — never an XOR, then an AND, then a compare. The primitive +/// owns bounds checking (overflow-safe, up front) and the trailing-bits-zero +/// guarantee, exactly as [`simd_rowstore_u32_eq_mask`] does. +#[inline] +pub fn simd_rowstore_ternary_match_mask( + bytes: &[u8], + first_offset: usize, + stride_bytes: usize, + n_rows: usize, + pattern: &[u8; FACET_REGISTER_BYTES], + care: &[u8; FACET_REGISTER_BYTES], + out_words: &mut [u64], +) { + ndarray::simd::ternary_match_strided_to_mask( + bytes, + first_offset, + stride_bytes, + n_rows, + pattern, + care, + out_words, + ); +} + /// Per-row facet-match: `out[row]` gets bit `f` set iff facet `f`'s classid /// in that row equals `needle` — "which facets of this node carry class X", /// one `u32` answer per row, written into the caller's buffer. @@ -270,6 +600,134 @@ pub fn scalar_gt_i32_to_mask(values: &[i32], threshold: i32, out_words: &mut [u6 } } +// ── The minor-11 comparison family's scalar half ──────────────────────────── +// +// These are NOT test-only oracles sitting in `src/main`: `lgj_plan_eval_scalar` +// is a shipped symbol whose whole purpose is to run the SAME plan down the +// independent path, so an op-code with no scalar arm would make the parity +// escape hatch answer `UNKNOWN_OPCODE` for exactly the ops it was extended to +// cover — an asymmetry between two symbols that `abi.md` §7 says have +// "identical semantics". Each is written as the predicate's own definition, so +// a facade that lowered `ge` as a buggy complement would be caught rather than +// agreed with. + +/// Reference `ne_u32` → mask. +pub fn scalar_ne_u32_to_mask(values: &[u32], needle: u32, out_words: &mut [u64]) { + for w in out_words.iter_mut() { + *w = 0; + } + for (i, &v) in values.iter().enumerate() { + if v != needle { + out_words[i / 64] |= 1u64 << (i % 64); + } + } +} + +/// Reference signed `eq_i32` → mask. +pub fn scalar_eq_i32_to_mask(values: &[i32], needle: i32, out_words: &mut [u64]) { + for w in out_words.iter_mut() { + *w = 0; + } + for (i, &v) in values.iter().enumerate() { + if v == needle { + out_words[i / 64] |= 1u64 << (i % 64); + } + } +} + +/// Reference signed `ne_i32` → mask. +pub fn scalar_ne_i32_to_mask(values: &[i32], needle: i32, out_words: &mut [u64]) { + for w in out_words.iter_mut() { + *w = 0; + } + for (i, &v) in values.iter().enumerate() { + if v != needle { + out_words[i / 64] |= 1u64 << (i % 64); + } + } +} + +/// Reference signed `lt_i32` → mask. +pub fn scalar_lt_i32_to_mask(values: &[i32], threshold: i32, out_words: &mut [u64]) { + for w in out_words.iter_mut() { + *w = 0; + } + for (i, &v) in values.iter().enumerate() { + if v < threshold { + out_words[i / 64] |= 1u64 << (i % 64); + } + } +} + +/// Reference signed `le_i32` → mask. +pub fn scalar_le_i32_to_mask(values: &[i32], threshold: i32, out_words: &mut [u64]) { + for w in out_words.iter_mut() { + *w = 0; + } + for (i, &v) in values.iter().enumerate() { + if v <= threshold { + out_words[i / 64] |= 1u64 << (i % 64); + } + } +} + +/// Reference signed `ge_i32` → mask. +pub fn scalar_ge_i32_to_mask(values: &[i32], threshold: i32, out_words: &mut [u64]) { + for w in out_words.iter_mut() { + *w = 0; + } + for (i, &v) in values.iter().enumerate() { + if v >= threshold { + out_words[i / 64] |= 1u64 << (i % 64); + } + } +} + +/// Reference care-masked (TCAM) `u32` match → mask. +pub fn scalar_ternary_match_u32_to_mask( + values: &[u32], + pattern: u32, + care: u32, + out_words: &mut [u64], +) { + for w in out_words.iter_mut() { + *w = 0; + } + for (i, &v) in values.iter().enumerate() { + if (v ^ pattern) & care == 0 { + out_words[i / 64] |= 1u64 << (i % 64); + } + } +} + +/// Reference masked minimum. `None` when nothing is selected. +pub fn scalar_masked_min_i32(values: &[i32], mask_words: &[u64]) -> Option { + let mut acc: Option = None; + for (i, &v) in values.iter().enumerate() { + if (mask_words[i / 64] >> (i % 64)) & 1 == 1 { + acc = Some(match acc { + Some(a) if a <= v => a, + _ => v, + }); + } + } + acc +} + +/// Reference masked maximum. `None` when nothing is selected. +pub fn scalar_masked_max_i32(values: &[i32], mask_words: &[u64]) -> Option { + let mut acc: Option = None; + for (i, &v) in values.iter().enumerate() { + if (mask_words[i / 64] >> (i % 64)) & 1 == 1 { + acc = Some(match acc { + Some(a) if a >= v => a, + _ => v, + }); + } + } + acc +} + /// Reference `dst &= src`. pub fn scalar_mask_and_assign(dst: &mut [u64], src: &[u64]) { for (d, &s) in dst.iter_mut().zip(src.iter()) { @@ -413,7 +871,70 @@ pub fn eval_predicate( Path::Scalar => scalar_gt_i32_to_mask(v, threshold, out_words), } } - (LGJ_OP_EQ_U32, _) | (LGJ_OP_GT_I32, _) => return Err(LGJ_ERR_LANE_KIND_MISMATCH), + (LGJ_OP_NE_U32, LaneView::U32(v)) => { + let needle = operand as u32; + match path { + Path::Simd => simd_ne_u32_to_mask(v, needle, out_words), + Path::Scalar => scalar_ne_u32_to_mask(v, needle, out_words), + } + } + (LGJ_OP_EQ_I32, LaneView::I32(v)) => { + let needle = operand as i32; + match path { + Path::Simd => simd_eq_i32_to_mask(v, needle, out_words), + Path::Scalar => scalar_eq_i32_to_mask(v, needle, out_words), + } + } + (LGJ_OP_NE_I32, LaneView::I32(v)) => { + let needle = operand as i32; + match path { + Path::Simd => simd_ne_i32_to_mask(v, needle, out_words), + Path::Scalar => scalar_ne_i32_to_mask(v, needle, out_words), + } + } + (LGJ_OP_LT_I32, LaneView::I32(v)) => { + let threshold = operand as i32; + match path { + Path::Simd => simd_lt_i32_to_mask(v, threshold, out_words), + Path::Scalar => scalar_lt_i32_to_mask(v, threshold, out_words), + } + } + (LGJ_OP_LE_I32, LaneView::I32(v)) => { + let threshold = operand as i32; + match path { + Path::Simd => simd_le_i32_to_mask(v, threshold, out_words), + Path::Scalar => scalar_le_i32_to_mask(v, threshold, out_words), + } + } + (LGJ_OP_GE_I32, LaneView::I32(v)) => { + let threshold = operand as i32; + match path { + Path::Simd => simd_ge_i32_to_mask(v, threshold, out_words), + Path::Scalar => scalar_ge_i32_to_mask(v, threshold, out_words), + } + } + (LGJ_OP_TERNARY_MATCH_U32, LaneView::U32(v)) => { + // The operand packs BOTH halves: pattern low, care high. Exact — + // two u32s are exactly 64 bits — and confined to this op-code, so + // no existing opcode's reading of `operand` changes. `abi.rs` + // documents the packing beside the constant. + let packed = operand as u64; + let pattern = packed as u32; + let care = (packed >> 32) as u32; + match path { + Path::Simd => simd_ternary_match_u32_to_mask(v, pattern, care, out_words), + Path::Scalar => scalar_ternary_match_u32_to_mask(v, pattern, care, out_words), + } + } + (LGJ_OP_EQ_U32, _) + | (LGJ_OP_GT_I32, _) + | (LGJ_OP_NE_U32, _) + | (LGJ_OP_EQ_I32, _) + | (LGJ_OP_NE_I32, _) + | (LGJ_OP_LT_I32, _) + | (LGJ_OP_LE_I32, _) + | (LGJ_OP_GE_I32, _) + | (LGJ_OP_TERNARY_MATCH_U32, _) => return Err(LGJ_ERR_LANE_KIND_MISMATCH), _ => return Err(LGJ_ERR_UNKNOWN_OPCODE), } // Both primitives already zero the tail, but re-establishing it here means @@ -458,6 +979,49 @@ pub fn masked_sum_i32(path: Path, values: &[i32], mask_words: &[u64]) -> i64 { } } +/// The PARAMETERISED masked reduction over an `I32` lane (ABI minor >= 11, +/// `docs/abi.md` §19) — `sum` / `min` / `max` behind one op-code argument. +/// +/// `abi.md` §15 wrote this shape down before the need arose: *"if a second +/// reduction is ever needed (min/max/count-distinct/histogram), do NOT add a +/// second symbol ... an op-code parameter on one reduce symbol, mirroring how +/// `lgj_plan_eval`'s `LgjOpDesc` already generalises predicates — and `sum` +/// becomes op-code 0."* This is that, and a future `count-distinct` or +/// `histogram` is one more arm rather than one more symbol. +/// +/// `Ok(None)` is "no answer", not "zero": min and max over an empty population +/// have none, and `i32::MAX` / `i32::MIN` as sentinels would be +/// indistinguishable from a population that genuinely contains them. `sum` is +/// always `Ok(Some(_))` — the sum of nothing is `0`, which IS the answer. +/// +/// Widening to `i64` is the SUM's contract, not a reinterpretation of min/max: +/// an `i32` extremum is exact in `i64`, so one return type carries all three +/// without a lossy case. +pub fn reduce_i32( + path: Path, + reduce_op: u32, + values: &[i32], + mask_words: &[u64], +) -> Result, i32> { + match reduce_op { + LGJ_REDUCE_SUM => Ok(Some(masked_sum_i32(path, values, mask_words))), + LGJ_REDUCE_MIN => Ok(match path { + Path::Simd => simd_masked_min_i32(values, mask_words), + Path::Scalar => scalar_masked_min_i32(values, mask_words), + } + .map(i64::from)), + LGJ_REDUCE_MAX => Ok(match path { + Path::Simd => simd_masked_max_i32(values, mask_words), + Path::Scalar => scalar_masked_max_i32(values, mask_words), + } + .map(i64::from)), + // Not in abi.md's reduction set. Rejected like an unknown opcode rather + // than defaulted to SUM, so a future reduction cannot be silently + // misread by an older build. + _ => Err(LGJ_ERR_UNKNOWN_OPCODE), + } +} + /// The three readings of the V3 content-blind 12-byte facet register. /// /// **This is a THIN WIRE ADAPTER over the contract's own @@ -1423,4 +1987,622 @@ mod tests { None ); } + + // ═══════════════════════════════════════════════════════════════════════ + // ABI minor 11 — the masking-op completion (docs/abi.md §19) + // + // Every op below carries the same three-part gate: an EQUIVALENCE arm + // against an independent scalar oracle across the row-count shapes that + // straddle every boundary in the packing (0/1 degenerate, 15/16/17 the + // 16-lane group, 63/64/65 the word, 4095 one short of a power of two), a + // CAN-FIRE arm proving the op changes something on non-trivial input, and + // a CAN-STAY-SILENT arm proving it does NOT touch what it must not — in + // this file's currency, the tail bits past `n_rows`. + // ═══════════════════════════════════════════════════════════════════════ + + /// The row-count shapes every minor-11 arm sweeps. 4095 is deliberately + /// `64 * 64 - 1`: it exercises a final word with 63 live bits, the widest + /// possible tail, which is the shape a "clear the tail" bug survives on a + /// round row count. + const SHAPES: [u64; 14] = [ + 0, 1, 15, 16, 17, 63, 64, 65, 127, 128, 129, 1000, 4095, 4096, + ]; + + /// SplitMix64 — the generator `Fixture`/`RowStore` already use, repeated + /// here rather than imported so the test data is independent of whatever + /// the fixture happens to generate. + fn mix(state: &mut u64) -> u64 { + *state = state.wrapping_add(0x9E37_79B9_7F4A_7C15); + let mut z = *state; + z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9); + z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB); + z ^ (z >> 31) + } + + /// `n_words` of pseudo-random mask words with the tail cleared — a + /// CONFORMING mask, which is what every op's contract is stated against. + fn random_mask(n_rows: u64, seed: u64) -> Vec { + let mut st = seed; + let mut w = words_for(n_rows); + for x in w.iter_mut() { + *x = mix(&mut st); + } + clear_tail_bits(&mut w, n_rows); + w + } + + /// Assert every bit at or past `n_rows` is zero, naming the offender. + fn assert_tail_clear(words: &[u64], n_rows: u64, what: &str) { + let used = (n_rows % 64) as u32; + if used != 0 { + if let Some(&last) = words.last() { + assert_eq!( + last >> used, + 0, + "{what}: tail bits past n_rows={n_rows} survived (last word {last:#018x})" + ); + } + } + } + + // ── The comparison family ─────────────────────────────────────────────── + + /// EQUIVALENCE. Each of the six new comparisons, and the `u32` TCAM, + /// against its independently-written scalar definition at every shape. + #[test] + fn minor_11_predicates_match_their_scalar_references_at_every_row_count_shape() { + for n in SHAPES { + for seed in [0u64, 1, 0xDEAD_BEEF] { + let f = Fixture::generate(n, seed).unwrap(); + let (mut a, mut b) = (words_for(n), words_for(n)); + + simd_ne_u32_to_mask(f.classes(), 7, &mut a); + scalar_ne_u32_to_mask(f.classes(), 7, &mut b); + assert_eq!(a, b, "ne_u32 at n={n} seed={seed}"); + assert_tail_clear(&a, n, "ne_u32"); + + for needle in [-1000i32, 0, 100, i32::MIN, i32::MAX] { + simd_eq_i32_to_mask(f.values(), needle, &mut a); + scalar_eq_i32_to_mask(f.values(), needle, &mut b); + assert_eq!(a, b, "eq_i32({needle}) at n={n} seed={seed}"); + assert_tail_clear(&a, n, "eq_i32"); + + simd_ne_i32_to_mask(f.values(), needle, &mut a); + scalar_ne_i32_to_mask(f.values(), needle, &mut b); + assert_eq!(a, b, "ne_i32({needle}) at n={n} seed={seed}"); + assert_tail_clear(&a, n, "ne_i32"); + + simd_lt_i32_to_mask(f.values(), needle, &mut a); + scalar_lt_i32_to_mask(f.values(), needle, &mut b); + assert_eq!(a, b, "lt_i32({needle}) at n={n} seed={seed}"); + assert_tail_clear(&a, n, "lt_i32"); + + simd_le_i32_to_mask(f.values(), needle, &mut a); + scalar_le_i32_to_mask(f.values(), needle, &mut b); + assert_eq!(a, b, "le_i32({needle}) at n={n} seed={seed}"); + assert_tail_clear(&a, n, "le_i32"); + + simd_ge_i32_to_mask(f.values(), needle, &mut a); + scalar_ge_i32_to_mask(f.values(), needle, &mut b); + assert_eq!(a, b, "ge_i32({needle}) at n={n} seed={seed}"); + assert_tail_clear(&a, n, "ge_i32"); + } + + for (pat, care) in [(0u32, 0u32), (7, 0xFFFF_FFFF), (5, 0b101), (0, 0b1)] { + simd_ternary_match_u32_to_mask(f.classes(), pat, care, &mut a); + scalar_ternary_match_u32_to_mask(f.classes(), pat, care, &mut b); + assert_eq!(a, b, "tcam_u32({pat},{care}) at n={n} seed={seed}"); + assert_tail_clear(&a, n, "ternary_match_u32"); + } + } + } + } + + /// CAN-FIRE + CAN-STAY-SILENT, on the boundaries the signed comparisons + /// get wrong when they are lowered as a complement or an off-by-one. + /// + /// `lt` at `i32::MIN` must select NOTHING (silence), and `ge` at + /// `i32::MIN` must select EVERYTHING (fire) — the pair that catches a + /// `x > t - 1` lowering, which wraps at the bottom of the range and + /// inverts both answers. + #[test] + fn the_signed_comparisons_are_exact_at_the_extremes() { + let v: Vec = vec![i32::MIN, -1000, -1, 0, 1, 100, i32::MAX]; + let n = v.len() as u64; + let all = (1u64 << v.len()) - 1; + let mut w = words_for(n); + + simd_lt_i32_to_mask(&v, i32::MIN, &mut w); + assert_eq!(w[0], 0, "nothing is < i32::MIN"); + simd_ge_i32_to_mask(&v, i32::MIN, &mut w); + assert_eq!(w[0], all, "everything is >= i32::MIN"); + simd_le_i32_to_mask(&v, i32::MAX, &mut w); + assert_eq!(w[0], all, "everything is <= i32::MAX"); + simd_lt_i32_to_mask(&v, 0, &mut w); + assert_eq!(w[0], 0b000_0111, "i32::MIN, -1000 and -1 are the negatives"); + simd_eq_i32_to_mask(&v, i32::MIN, &mut w); + assert_eq!(w[0], 0b000_0001, "exactly one element equals i32::MIN"); + simd_ne_i32_to_mask(&v, i32::MIN, &mut w); + assert_eq!(w[0], all & !0b1, "and exactly the others differ from it"); + } + + /// CAN-STAY-SILENT for the whole family: a predicate true of EVERY row + /// must still leave the tail zero. 70 rows leaves word 1 with 6 live bits + /// and 58 that `popcount` would otherwise count. + #[test] + fn minor_11_predicates_leave_no_tail_bit_set_even_when_every_row_matches() { + let n = 70u64; + let u: Vec = vec![7; n as usize]; + let i: Vec = vec![7; n as usize]; + let mut w = words_for(n); + + // `!=` against a needle no element carries: true of every row. + simd_ne_u32_to_mask(&u, 9, &mut w); + assert_eq!(simd_popcount(&w), n, "the predicate really does select all"); + assert_tail_clear(&w, n, "ne_u32"); + + for f in [ + simd_ne_i32_to_mask as fn(&[i32], i32, &mut [u64]), + simd_le_i32_to_mask, + simd_lt_i32_to_mask, + ] { + f(&i, 9, &mut w); + assert_eq!(simd_popcount(&w), n); + assert_tail_clear(&w, n, "i32 predicate"); + } + simd_ge_i32_to_mask(&i, 7, &mut w); + assert_eq!(simd_popcount(&w), n); + assert_tail_clear(&w, n, "ge_i32"); + simd_eq_i32_to_mask(&i, 7, &mut w); + assert_eq!(simd_popcount(&w), n); + assert_tail_clear(&w, n, "eq_i32"); + // care == 0 is "don't care about anything" — true of every row. + simd_ternary_match_u32_to_mask(&u, 0xDEAD_BEEF, 0, &mut w); + assert_eq!(simd_popcount(&w), n); + assert_tail_clear(&w, n, "ternary_match_u32"); + } + + /// The care mask is the whole point of a TCAM, so it gets its own + /// two-sided arm: a bit OUTSIDE `care` may differ freely (fire), and a bit + /// INSIDE `care` must exclude (silence). Without the second half a + /// "matches everything" implementation would pass. + #[test] + fn ternary_match_ignores_uncared_bits_and_honours_cared_ones() { + let v: Vec = vec![0b1010, 0b1110, 0b0010, 0b1011]; + let mut w = words_for(4); + // care clears bit 2: elements 0 and 1 differ ONLY there, so both match. + simd_ternary_match_u32_to_mask(&v, 0b1010, 0b1011, &mut w); + assert_eq!(w[0], 0b0011, "an uncared bit must not exclude"); + // care now includes bit 2: element 1 differs there and drops out. + simd_ternary_match_u32_to_mask(&v, 0b1010, 0b1111, &mut w); + assert_eq!(w[0], 0b0001, "a cared bit must exclude"); + // The 64-bit sibling agrees on the same shape, one bit higher. + let v64: Vec = v.iter().map(|&x| u64::from(x) | (1u64 << 40)).collect(); + let mut w64 = words_for(4); + simd_ternary_match_u64_to_mask(&v64, 0b1010 | (1 << 40), 0b1011 | (1 << 40), &mut w64); + assert_eq!(w64[0], 0b0011, "u64 TCAM must see the high bit too"); + simd_ternary_match_u64_to_mask(&v64, 0b1010, 0b1011 | (1 << 40), &mut w64); + assert_eq!(w64[0], 0, "a cared HIGH bit must exclude every element"); + } + + // ── XOR / NOT / ANY / ALL ─────────────────────────────────────────────── + + /// EQUIVALENCE + the falsifier pair for XOR, at every shape. + #[test] + fn mask_xor_is_symmetric_difference_and_keeps_a_conforming_tail() { + for n in SHAPES { + let a = random_mask(n, 0x1234); + let b = random_mask(n, 0x5678); + let mut dst = words_for(n); + simd_mask_xor(&a, &b, &mut dst); + let expect: Vec = a.iter().zip(&b).map(|(x, y)| x ^ y).collect(); + assert_eq!(dst, expect, "xor at n={n}"); + assert_tail_clear(&dst, n, "mask_xor"); + + let mut inplace = a.clone(); + simd_mask_xor_assign(&mut inplace, &b); + assert_eq!(inplace, dst, "xor_assign disagrees with xor at n={n}"); + + // CAN-STAY-SILENT: `x ^ x` is empty for every x. + let mut self_xor = a.clone(); + simd_mask_xor_assign(&mut self_xor, &a); + assert!( + self_xor.iter().all(|&w| w == 0), + "x ^ x must be empty at n={n}" + ); + } + // CAN-FIRE, with the anti-vacuity half explicit: the inputs really do + // differ, so an implementation that wrote zeros could not pass. + let a = [0b0110u64]; + let b = [0b0011u64]; + let mut dst = [0u64; 1]; + simd_mask_xor(&a, &b, &mut dst); + assert_ne!(dst[0], a[0]); + assert_eq!(dst[0], 0b0101); + } + + /// NOT is the ONE mask-algebra op whose correctness cannot be stated + /// without `n_rows`, so its falsifier is the tail itself — and it is + /// two-sided: the live bits must all flip (fire) and the tail must be + /// zero where a plain `!` would set every bit (silence). + #[test] + fn mask_not_flips_the_live_bits_and_clears_the_tail_a_plain_complement_would_set() { + for n in SHAPES { + let src = random_mask(n, 0xABCD); + let mut dst = words_for(n); + simd_mask_not(&src, n as usize, &mut dst); + assert_tail_clear(&dst, n, "mask_not"); + // Every LIVE bit flipped: the two popcounts must partition n. + assert_eq!( + simd_popcount(&src) + simd_popcount(&dst), + n, + "complement must partition the population at n={n}" + ); + let mut inplace = src.clone(); + simd_mask_not_assign(&mut inplace, n as usize); + assert_eq!(inplace, dst, "not_assign disagrees with not at n={n}"); + } + // The silence half, made concrete: at 4 rows a raw `!` leaves 60 stray + // bits in the word, and the test asserts the naive answer is NOT what + // we got — so a version without the clear fails here, not merely + // elsewhere. + let src = [0b0011u64]; + let mut dst = [0u64; 1]; + simd_mask_not(&src, 4, &mut dst); + assert_eq!(dst[0], 0b1100); + assert_ne!(dst[0], !src[0], "a plain complement must not be the answer"); + } + + /// ANY and ALL, both directions each. The non-trivial inputs matter: an + /// empty-vs-empty pair would pass for a function that always answered the + /// same way. + #[test] + fn mask_any_and_mask_all_both_fire_and_both_stay_silent() { + let n = 70u64; + let empty = words_for(n); + assert!(!simd_mask_any(&empty), "an empty mask selects nothing"); + assert!( + !simd_mask_all(&empty, n as usize), + "and is certainly not full" + ); + + let mut one = words_for(n); + one[1] |= 1 << 5; // row 69 — the last live bit + assert!(simd_mask_any(&one), "one set row is 'any'"); + assert!(!simd_mask_all(&one, n as usize), "one set row is not 'all'"); + + let mut full = vec![u64::MAX; mask_words_for(n) as usize]; + clear_tail_bits(&mut full, n); + assert!(simd_mask_any(&full)); + assert!( + simd_mask_all(&full, n as usize), + "a conforming full mask is 'all'" + ); + + // One row short of full is NOT all — the off-by-one this predicate + // invites, checked at the boundary between the two words. + let mut almost = full.clone(); + almost[0] &= !(1u64 << 63); + assert!( + !simd_mask_all(&almost, n as usize), + "63 of 64 in word 0 is not 'all'" + ); + // A zero-row population is vacuously full, and that must not read as a + // bug in the row-count arithmetic. + assert!(simd_mask_all(&[], 0)); + } + + // ── TERNLOG ───────────────────────────────────────────────────────────── + + /// The 256-entry dispatch table's own falsifier: for EVERY immediate, the + /// runtime dispatch must equal the truth table evaluated bit by bit. + /// + /// This is what makes the table's indexing (`[imm >> 4][imm & 15]`) + /// checkable rather than asserted — a transposed index, an off-by-one row, + /// or a `$base` typo in the macro shifts some immediates and this fails on + /// exactly those. + #[test] + fn the_ternlog_dispatch_table_agrees_with_the_truth_table_for_all_256_immediates() { + let n = 129u64; // three words, the last one short + let a0 = random_mask(n, 0x11); + let b = random_mask(n, 0x22); + let c = random_mask(n, 0x33); + for imm in 0u8..=255 { + let mut got = a0.clone(); + simd_mask_ternlog_assign_dyn(imm, &mut got, &b, &c); + // Bit-serial oracle: index (a<<2)|(b<<1)|c into the immediate. + let mut want = vec![0u64; got.len()]; + for (w, out) in want.iter_mut().enumerate() { + for bit in 0..64u32 { + let av = (a0[w] >> bit) & 1; + let bv = (b[w] >> bit) & 1; + let cv = (c[w] >> bit) & 1; + let idx = (av << 2) | (bv << 1) | cv; + if (u64::from(imm) >> idx) & 1 == 1 { + *out |= 1u64 << bit; + } + } + } + assert_eq!(got, want, "ternlog imm={imm:#04x}"); + } + } + + /// CAN-FIRE for the general member: the named immediates must reproduce + /// the dedicated ops exactly, which is the claim that lets XOR and NOT + /// arrive as immediates instead of as symbols. + #[test] + fn the_named_immediates_reproduce_the_dedicated_mask_ops() { + const AND2: u8 = 0xC0; + const OR2: u8 = 0xFC; + const XOR2: u8 = 0x3C; + const ANDNOT2: u8 = 0x30; + const NOT_A: u8 = 0x0F; + + let n = 200u64; + let a = random_mask(n, 0xAAAA); + let b = random_mask(n, 0xBBBB); + let mut want = words_for(n); + let run = |imm: u8| { + let mut got = a.clone(); + // `c` is ignored by every two-input table; feeding it a mask that + // is NOT all-zero is what proves it really is ignored. + simd_mask_ternlog_assign_dyn(imm, &mut got, &b, &random_mask(n, 0xCCCC)); + got + }; + + simd_mask_and(&a, &b, &mut want); + assert_eq!(run(AND2), want, "AND2"); + simd_mask_or(&a, &b, &mut want); + assert_eq!(run(OR2), want, "OR2"); + simd_mask_xor(&a, &b, &mut want); + assert_eq!(run(XOR2), want, "XOR2"); + simd_mask_andnot(&a, &b, &mut want); + assert_eq!(run(ANDNOT2), want, "AND2_ANDNOT"); + + // NOT is the interesting one: its table is ODD, so before the tail + // clear the result has every bit past n_rows set. That is the state + // the export's `clear_tail_bits` exists for, and asserting it here + // makes the export's tail step a measured requirement rather than a + // precaution. + let mut not_got = run(NOT_A); + assert_ne!( + not_got.last().copied().unwrap_or(0) >> (n % 64), + 0, + "an odd immediate must set the tail — otherwise the export's clear is untested" + ); + clear_tail_bits(&mut not_got, n); + simd_mask_not(&a, n as usize, &mut want); + assert_eq!(not_got, want, "NOT_A after the tail clear"); + } + + /// The in-place and out-of-place spellings must agree — the property the + /// export relies on when it copies `a` into `dst` and then assigns. + #[test] + fn ternlog_assign_and_out_of_place_agree() { + let n = 1000u64; + let a = random_mask(n, 1); + let b = random_mask(n, 2); + let c = random_mask(n, 3); + let mut out = words_for(n); + simd_mask_ternlog::<{ ternlog::MAJ3 }>(&a, &b, &c, &mut out); + let mut inplace = a.clone(); + simd_mask_ternlog_assign::<{ ternlog::MAJ3 }>(&mut inplace, &b, &c); + assert_eq!(out, inplace); + // …and the runtime dispatch reaches the same monomorphisation. + let mut dynamic = a.clone(); + simd_mask_ternlog_assign_dyn(ternlog::MAJ3 as u8, &mut dynamic, &b, &c); + assert_eq!(out, dynamic); + } + + // ── Reductions + blend ────────────────────────────────────────────────── + + /// EQUIVALENCE for min/max against their scalar definitions, plus the + /// empty-population arm that `None` exists for. + #[test] + fn masked_min_and_max_match_their_scalar_references_and_report_an_empty_population() { + for n in SHAPES { + for seed in [0u64, 7, 0xFEED] { + let f = Fixture::generate(n, seed).unwrap(); + let m = random_mask(n, seed ^ 0x5A5A); + assert_eq!( + simd_masked_min_i32(f.values(), &m), + scalar_masked_min_i32(f.values(), &m), + "min at n={n} seed={seed}" + ); + assert_eq!( + simd_masked_max_i32(f.values(), &m), + scalar_masked_max_i32(f.values(), &m), + "max at n={n} seed={seed}" + ); + // CAN-STAY-SILENT: an empty mask has no extremum at all. + let empty = words_for(n); + assert_eq!(simd_masked_min_i32(f.values(), &empty), None); + assert_eq!(simd_masked_max_i32(f.values(), &empty), None); + } + } + // CAN-FIRE, and the mask must actually SELECT: a min over everything + // and a min over a subset must differ, or the mask is decorative. + let v = [10i32, -5, 30, 2]; + assert_eq!(simd_masked_min_i32(&v, &[0b1111]), Some(-5)); + assert_eq!(simd_masked_min_i32(&v, &[0b1101]), Some(2)); + assert_eq!(simd_masked_max_i32(&v, &[0b1111]), Some(30)); + assert_eq!(simd_masked_max_i32(&v, &[0b1011]), Some(10)); + } + + /// `reduce_i32` dispatch: every op on both paths, and an unknown op + /// rejected rather than defaulted to SUM. + #[test] + fn reduce_i32_dispatches_every_op_on_both_paths_and_rejects_an_unknown_one() { + let f = Fixture::generate(500, 3).unwrap(); + let m = random_mask(500, 0x99); + for path in [Path::Simd, Path::Scalar] { + assert_eq!( + reduce_i32(path, LGJ_REDUCE_SUM, f.values(), &m), + Ok(Some(masked_sum_i32(path, f.values(), &m))) + ); + assert_eq!( + reduce_i32(path, LGJ_REDUCE_MIN, f.values(), &m), + Ok(scalar_masked_min_i32(f.values(), &m).map(i64::from)) + ); + assert_eq!( + reduce_i32(path, LGJ_REDUCE_MAX, f.values(), &m), + Ok(scalar_masked_max_i32(f.values(), &m).map(i64::from)) + ); + for bogus in [3u32, 4, 99, u32::MAX] { + assert_eq!( + reduce_i32(path, bogus, f.values(), &m), + Err(LGJ_ERR_UNKNOWN_OPCODE), + "an unknown reduction must not alias a known one" + ); + } + } + // SUM is present over an empty population; MIN/MAX are not. That is + // the distinction `out_present` carries across the membrane. + let empty = words_for(500); + assert_eq!( + reduce_i32(Path::Simd, LGJ_REDUCE_SUM, f.values(), &empty), + Ok(Some(0)) + ); + assert_eq!( + reduce_i32(Path::Simd, LGJ_REDUCE_MIN, f.values(), &empty), + Ok(None) + ); + } + + /// `blend_i32` — substrate-tier only (no export; see its doc). Falsifier + /// pair: the mask really selects between the two sources, and a bit set + /// nowhere leaves `b` untouched. + #[test] + fn blend_i32_selects_from_both_sources() { + let a = [1i32, 2, 3, 4]; + let b = [10i32, 20, 30, 40]; + let mut dst = [0i32; 4]; + simd_blend_i32(&[0b0101], &a, &b, &mut dst); + assert_eq!(dst, [1, 20, 3, 40]); + simd_blend_i32(&[0], &a, &b, &mut dst); + assert_eq!(dst, b, "an empty mask must take everything from b"); + simd_blend_i32(&[0b1111], &a, &b, &mut dst); + assert_eq!(dst, a, "a full mask must take everything from a"); + } + + // ── The strided TCAM over a row store ─────────────────────────────────── + + /// EQUIVALENCE + can-fire/can-stay-silent for the register TCAM, over the + /// REAL row store, on every facet. + /// + /// The scalar oracle is local to this test rather than a `scalar_*` in + /// `src/main`: nothing on `Path::Scalar` reaches this kernel, so shipping + /// an oracle for it would put a scalar path in the artifact for the + /// benefit of the test suite alone (E2 licenses scalar code as a TEST + /// oracle, which is where this one lives). + #[test] + fn the_register_tcam_matches_a_byte_wise_oracle_on_every_facet() { + use crate::rowstore::{RowStore, FACET_BYTES, FACET_CLASSID_BYTES, ROW_BYTES, ROW_FACETS}; + const N: usize = FACET_REGISTER_BYTES; + + for n in [1u64, 63, 64, 65, 300] { + let store = RowStore::generate(n, 0xC0FFEE).unwrap(); + let bytes = store.as_bytes(); + for facet in 0..ROW_FACETS { + let off = (u64::from(facet) * FACET_BYTES + FACET_CLASSID_BYTES) as usize; + // Read row 0's own register and use it as the pattern, so the + // "exact" arm is guaranteed to have at least one hit — the + // anti-vacuity half of the can-fire assertion below. + let mut pattern = [0u8; N]; + pattern.copy_from_slice(&bytes[off..off + N]); + + for care in [[0u8; N], [0xFFu8; N], { + let mut c = [0u8; N]; + c[0] = 0xFF; + c + }] { + let mut got = words_for(n); + simd_rowstore_ternary_match_mask( + bytes, + off, + ROW_BYTES as usize, + n as usize, + &pattern, + &care, + &mut got, + ); + let mut want = words_for(n); + for row in 0..n as usize { + let o = off + row * ROW_BYTES as usize; + let hit = (0..N).all(|k| (bytes[o + k] ^ pattern[k]) & care[k] == 0); + if hit { + want[row / 64] |= 1u64 << (row % 64); + } + } + assert_eq!(got, want, "tcam n={n} facet={facet} care={care:?}"); + assert_tail_clear(&got, n, "register tcam"); + } + + // CAN-FIRE: the exact pattern selects at least row 0 … + let mut exact = words_for(n); + simd_rowstore_ternary_match_mask( + bytes, + off, + ROW_BYTES as usize, + n as usize, + &pattern, + &[0xFF; N], + &mut exact, + ); + assert_eq!(exact[0] & 1, 1, "row 0 must match its own register"); + // … and CAN-STAY-SILENT: flipping a CARED byte drops it. + let mut miss = pattern; + miss[0] ^= 0xFF; + let mut none = words_for(n); + simd_rowstore_ternary_match_mask( + bytes, + off, + ROW_BYTES as usize, + n as usize, + &miss, + &[0xFF; N], + &mut none, + ); + assert_eq!(none[0] & 1, 0, "a cared-byte mismatch must exclude row 0"); + // … while flipping it with that byte UN-cared keeps it. + let mut care_rest = [0xFFu8; N]; + care_rest[0] = 0; + let mut still = words_for(n); + simd_rowstore_ternary_match_mask( + bytes, + off, + ROW_BYTES as usize, + n as usize, + &miss, + &care_rest, + &mut still, + ); + assert_eq!(still[0] & 1, 1, "an uncared byte must not exclude row 0"); + } + } + } + + /// The register geometry has ONE spelling (E3). This pins that the + /// constant is DERIVED — a literal `12` here would survive a change to + /// `FACET_BYTES` and this does not. + #[test] + fn the_register_size_is_derived_from_the_row_geometry() { + assert_eq!( + FACET_REGISTER_BYTES, + (crate::rowstore::FACET_BYTES - crate::rowstore::FACET_CLASSID_BYTES) as usize + ); + // The same 12 bytes read the other way: the register is the lo64 span + // (hi32's offset measured from the REGISTER base, not the facet base) + // plus the 4-byte hi32 field. Getting this relation wrong is how a + // "12" ends up hand-written somewhere, which is what E3 forbids — the + // first version of this assertion added `4` to the facet-relative + // offset and measured 16, the facet's own size. + assert_eq!( + FACET_REGISTER_BYTES, + (crate::rowstore::FACET_PAYLOAD_HI32_OFFSET - crate::rowstore::FACET_CLASSID_BYTES) + as usize + + 4, + "register = lo64 span (8) + hi32 (4); the two spellings must agree" + ); + } } diff --git a/native/lgj-abi/src/lib.rs b/native/lgj-abi/src/lib.rs index 2e719d7..f342cfe 100644 --- a/native/lgj-abi/src/lib.rs +++ b/native/lgj-abi/src/lib.rs @@ -52,6 +52,7 @@ pub mod class_view_provider; pub mod exports; pub mod fixture; pub mod kernels; +pub mod plan_lower; pub mod registry; pub mod rowstore; diff --git a/native/lgj-abi/src/plan_lower.rs b/native/lgj-abi/src/plan_lower.rs new file mode 100644 index 0000000..0518ff8 --- /dev/null +++ b/native/lgj-abi/src/plan_lower.rs @@ -0,0 +1,160 @@ +//! The plan lowering — `&[LgjOpDesc]` becomes one `mask_risc::Program`. +//! +//! This is the component PR4 introduces that did not exist before, and the +//! reason `pr4_equivalence.rs` freezes a copy of the old loop rather than +//! reusing mask-risc's own oracle: the lowering sits ABOVE both `execute` and +//! `reference_execute`, so a bug here produces the same wrong `Program` for +//! both and mask-risc's own differential stays green. +//! +//! # The prefix rewrite +//! +//! The old loop seeded its accumulator with all rows and then folded each op +//! in with `&=` or `|=`. An `|=` against an all-ones accumulator cannot shrink +//! it, so **every op before the first AND is dead**: whatever it selects, the +//! accumulator is still all rows when the first AND arrives, and +//! `all_ones & p == p`. +//! +//! So let `k` be the least index whose combine is AND. Ops `0..k` are dropped; +//! op `k` becomes a bare `Pred` writing slot 0 (it IS the accumulator, no +//! seed needed); each later op writes slot 1 and folds into slot 0. If no op +//! combines with AND (`k == n`) the answer is every row and **no program is +//! built at all** — there is nothing left to evaluate. +//! +//! That is not an optimisation bolted on: it is what lets the lowering avoid +//! needing a fill/constant op, which the mask-RISC deliberately does not have. +//! +//! # The survivor skip, and where it is NOT sound +//! +//! A later op whose combine is AND is gated `under` slot 0: `acc & p` depends +//! on `p` only where `acc` already has a survivor, so the predicate runs over +//! the accumulator's live 64-row words and nothing else. +//! +//! A later op whose combine is OR is **not** gated, and the asymmetry is the +//! whole correctness question in this file. `acc | p` depends on `p` exactly +//! where `acc` is ZERO — the rows a gate under `acc` would discard. Gating an +//! OR would zero precisely the bits that matter and quietly shrink the answer +//! to `acc` itself. The combine-vector sweep in `pr4_matrix.rs` is what holds +//! this: it runs the full `{AND, OR}^n` product against the frozen oracle. + +use crate::abi::{ + LgjOpDesc, LGJ_COMBINE_AND, LGJ_OP_EQ_I32, LGJ_OP_EQ_U32, LGJ_OP_GE_I32, LGJ_OP_GT_I32, + LGJ_OP_LE_I32, LGJ_OP_LT_I32, LGJ_OP_NE_I32, LGJ_OP_NE_U32, LGJ_OP_TERNARY_MATCH_U32, +}; +use lance_graph_mask_risc::{MaskOp, Operand, Pred, Program, Terminal}; + +/// Scratch slots the lowering ever names: slot 0 is the accumulator (the old +/// `acc`, now living in the caller's arena rather than in a per-call `vec!`), +/// slot 1 the per-op predicate destination (the old `scratch`). +pub(crate) const SLOTS: u32 = 2; + +/// The scratch slot the final mask lands in — the accumulator. +pub(crate) const ACC_SLOT: u16 = 0; + +/// What a plan lowers to. +pub(crate) enum Lowered { + /// Every combine is OR. The accumulator starts all-ones and `|=` can never + /// shrink it, so the answer is every row with a clean tail — a constant, + /// reached without building or running a program. + AllRows, + /// The surviving suffix, as one straight-line program whose terminal + /// counts slot [`ACC_SLOT`]. + Program(Program), +} + +/// Lower a validated plan. +/// +/// `ops` must already have passed `validate_plan`, which is what makes the +/// opcode map total. It is still written as a `None` rather than a panic: an +/// unreachable arm that aborts the process is a worse failure than one that +/// returns the error the validator would have returned anyway. +pub(crate) fn lower_plan(ops: &[LgjOpDesc]) -> Option { + let Some(k) = ops.iter().position(|o| o.combine == LGJ_COMBINE_AND) else { + return Some(Lowered::AllRows); + }; + + let mut program_ops = Vec::with_capacity(2 * (ops.len() - k) - 1); + program_ops.push(MaskOp::Pred { + pred: pred_of(&ops[k])?, + under: None, + dst: ACC_SLOT, + }); + for op in &ops[k + 1..] { + let is_and = op.combine == LGJ_COMBINE_AND; + program_ops.push(MaskOp::Pred { + pred: pred_of(op)?, + // See the module doc: an AND reads its predicate only where the + // accumulator already survives; an OR reads it exactly where the + // accumulator does not, so gating one would answer `acc`. + under: is_and.then_some(Operand::Scratch(ACC_SLOT)), + dst: 1, + }); + let (a, b, dst) = (Operand::Scratch(ACC_SLOT), Operand::Scratch(1), ACC_SLOT); + program_ops.push(if is_and { + MaskOp::And { a, b, dst } + } else { + MaskOp::Or { a, b, dst } + }); + } + + Some(Lowered::Program(Program::new( + program_ops, + Terminal::Count { + mask: Operand::Scratch(ACC_SLOT), + }, + ))) +} + +/// One `LgjOpDesc` opcode + operand → one `Pred`. +/// +/// Every operand conversion is the one `kernels::eval_predicate` performs, and +/// deliberately spelled the same way: the operand arrives sign-extended in an +/// `i64`, a `u32` needle is its low 32 bits as an exact bit pattern, and the +/// TCAM operand packs `(care << 32) | pattern`. A different cast here is +/// exactly the class of bug the frozen oracle exists to catch, so the two +/// readings are kept textually parallel rather than merely equivalent. +fn pred_of(op: &LgjOpDesc) -> Option { + let lane = u16::try_from(op.lane_id).ok()?; + Some(match op.op { + LGJ_OP_EQ_U32 => Pred::EqU32 { + lane, + v: op.operand as u32, + }, + LGJ_OP_NE_U32 => Pred::NeU32 { + lane, + v: op.operand as u32, + }, + LGJ_OP_EQ_I32 => Pred::EqI32 { + lane, + v: op.operand as i32, + }, + LGJ_OP_NE_I32 => Pred::NeI32 { + lane, + v: op.operand as i32, + }, + LGJ_OP_GT_I32 => Pred::GtI32 { + lane, + t: op.operand as i32, + }, + LGJ_OP_LT_I32 => Pred::LtI32 { + lane, + t: op.operand as i32, + }, + LGJ_OP_LE_I32 => Pred::LeI32 { + lane, + t: op.operand as i32, + }, + LGJ_OP_GE_I32 => Pred::GeI32 { + lane, + t: op.operand as i32, + }, + LGJ_OP_TERNARY_MATCH_U32 => { + let packed = op.operand as u64; + Pred::MatchU32 { + lane, + pattern: packed as u32, + care: (packed >> 32) as u32, + } + } + _ => return None, + }) +} diff --git a/native/lgj-abi/tests/c_one_evaluator.rs b/native/lgj-abi/tests/c_one_evaluator.rs new file mode 100644 index 0000000..e9884a1 --- /dev/null +++ b/native/lgj-abi/tests/c_one_evaluator.rs @@ -0,0 +1,231 @@ +//! C-ONE — there is exactly one plan evaluator, and this is the structural +//! guard that says so. +//! +//! # What "one evaluator" means here, precisely +//! +//! Before PR4, `plan_eval_impl` held its own loop over `&[LgjOpDesc]`: +//! resolve the lane, dispatch the opcode, combine into an accumulator. That +//! is an evaluator, and it stood beside the one in `lance-graph-mask-risc`, +//! each free to drift. PR4 deletes it — the plan is lowered to a +//! `mask_risc::Program` and run there. +//! +//! Nothing stops a future change from growing a second one back. It would not +//! look like a mistake at the time: a "fast path for two-op plans", a +//! "specialised loop for the all-AND case", a "scalar fallback" — each is one +//! small loop over `&[LgjOpDesc]` that answers the same question a different +//! way, which is the whole failure mode. +//! +//! # The invariant this actually checks, and why it is this one +//! +//! **Only an allowlisted set of production modules may name `LgjOpDesc` in +//! CODE.** Naming the type is the minimum any second evaluator must do: it +//! cannot loop over plan ops without a plan op in scope. Comments are +//! stripped first — see [`without_comments`] for the file that forced that +//! and why allowlisting it instead would have been the wrong repair. +//! +//! Three narrower rules were considered and rejected, each for a measured +//! reason: +//! +//! - *"only `plan_lower` may call `kernels::eval_predicate`"* — wrong, and +//! the tree says so: `lgj_op_eq_u32` and `lgj_op_gt_i32` call it, and +//! correctly. Those are the UNFUSED single-predicate exports that exist so +//! the fused plan has something to be benchmarked against; each passes a +//! compile-time-constant opcode and evaluates exactly one predicate. One +//! predicate is not an evaluator. +//! - *"the opcode argument must be a literal"* — would catch those two, but +//! `lgj_mask_combine` passes a RUNTIME `combine` to `kernels::combine_into` +//! and is a legitimate single bulk op. The rule would fire on a +//! non-violation. +//! - *"no production function may iterate `&[LgjOpDesc]`"* — the right +//! property, and not textually checkable: `for op in ops` carries no type +//! in its text, so a scanner would have to be a type checker. +//! +//! So the allowlist is over MODULES, in the same shape as the G11 +//! contract-import fence: widening it is a deliberate act, in one commit, and +//! a new module holding plan ops is exactly the event that deserves a human +//! reading the diff. +//! +//! # Scope, stated rather than implied +//! +//! This guard is structural. It cannot tell a second evaluator from a second +//! legitimate USE of the type inside an already-allowed module — `exports.rs` +//! is allowlisted, so a new loop added there passes this test. What catches +//! that is the differential in `pr4_equivalence.rs`, which compares whatever +//! the live path does against a frozen copy of the old loop. The two are +//! complementary and neither subsumes the other: the differential cannot see +//! a second evaluator that happens to agree, and this cannot see one hidden +//! in an allowed file. + +use std::collections::BTreeSet; +use std::fs; +use std::path::{Path, PathBuf}; + +/// Production modules permitted to name `LgjOpDesc`, each with the reason. +/// +/// - `abi.rs` — declares the type and its `LGJ_OP_*` / `LGJ_COMBINE_*` +/// constants. The definition site. +/// - `exports.rs` — the `extern "C"` signatures that receive a plan, plus +/// `validate_plan`, which reads every field to REJECT a bad plan. A +/// validator is not an evaluator: it computes no mask and returns no rows. +/// - `plan_lower.rs` — the one lowering, `&[LgjOpDesc]` to one +/// `mask_risc::Program`. The only production place a runtime opcode is +/// dispatched. +const ALLOWED: &[&str] = &["abi.rs", "exports.rs", "plan_lower.rs"]; + +const TYPE: &str = "LgjOpDesc"; + +/// Every `.rs` file under `dir`, recursively. +/// +/// Recursive rather than a flat `read_dir`: the G11 fence shipped with a flat +/// scan and codex found the nested-directory bypass standing beside the path +/// its disable run had walked. Same scanner, same bypass, so the same fix. +fn rust_files(dir: &Path, out: &mut Vec) { + let entries = fs::read_dir(dir).unwrap_or_else(|e| panic!("read_dir {dir:?}: {e}")); + for entry in entries.flatten() { + let path = entry.path(); + if path.is_dir() { + rust_files(&path, out); + } else if path.extension().is_some_and(|e| e == "rs") { + out.push(path); + } + } +} + +/// `text` with `//`-comments removed, line by line. +/// +/// The invariant is about naming the type in CODE. `kernels.rs` mentions +/// `LgjOpDesc` four times in prose — explaining that its six extra predicates +/// are op-codes rather than new ABI symbols — and never touches the type. A +/// first version scanned raw text and flagged it, which would have forced a +/// choice between a false failure and allowlisting `kernels.rs`. The second +/// is the worse of the two: kernels is precisely where a second evaluator +/// would most plausibly grow, so exempting it because of a doc comment would +/// hollow the guard out at exactly the point it matters. +/// +/// Limitation, stated because it is real: a `//` inside a string literal +/// truncates the rest of that line, so an occurrence AFTER such a literal on +/// the SAME line would be missed. No line in this crate has that shape, and +/// the anti-vacuity assertion catches a scanner that stops seeing the type +/// altogether — but a full lexer is what would make this airtight. +fn without_comments(text: &str) -> String { + text.lines() + .map(|line| match line.find("//") { + Some(i) => &line[..i], + None => line, + }) + .collect::>() + .join("\n") +} + +/// `text` with its trailing `#[cfg(test)] mod` removed. +/// +/// The split anchors on the test MODULE, not on the attribute. A first +/// version split at the first `#[cfg(test)]` and was immediately red on two +/// files, correctly: `registry.rs` and `exports.rs` both carry test-only +/// ITEMS — a `slot_count()` helper, a `RESOLUTIONS` counter, and one +/// statement-level attribute inside a production function — long before their +/// test module. Splitting there would have silently discarded most of +/// `exports.rs` as "test code", which for this scanner means a violation in +/// the discarded region reads as compliance. +/// +/// Keeping those items on the PRODUCTION side is the conservative direction: +/// a test-only helper that names the type can only ADD a holder, never hide +/// one, so the guard errs toward firing rather than toward silence. +/// +/// Returns `None` when the module is not the file's last item, which the +/// brace match establishes rather than assumes — the caller fails loudly +/// instead of scanning less than it claims to. +fn production_half(text: &str) -> Option<&str> { + let Some(rel) = text.rfind("#[cfg(test)]\nmod ") else { + return Some(text); + }; + let open = text[rel..].find('{')? + rel; + let mut depth = 0usize; + let mut close = None; + let mut chars = text[open..].char_indices(); + for (off, ch) in &mut chars { + match ch { + '{' => depth += 1, + '}' => { + depth -= 1; + if depth == 0 { + close = Some(open + off); + break; + } + } + _ => {} + } + } + // Nothing but whitespace may follow: that is what makes "the module is + // last" a fact rather than a convention this scanner relies on. + if !text[close? + 1..].trim().is_empty() { + return None; + } + Some(&text[..rel]) +} + +/// FAILS IF: a production module outside the allowlist names `LgjOpDesc` — +/// the minimum any second plan evaluator must do — or if the scan found +/// nothing, which would make the first half pass for a scanner pointed at the +/// wrong directory. +#[test] +fn only_the_lowering_and_the_validator_hold_plan_ops() { + let src = Path::new(env!("CARGO_MANIFEST_DIR")).join("src"); + let mut files = Vec::new(); + rust_files(&src, &mut files); + assert!( + files.len() >= 5, + "found {} files under {src:?} — the scanner is pointed at the wrong tree", + files.len() + ); + + let mut holders = BTreeSet::new(); + let mut interleaved = Vec::new(); + for path in &files { + // The PR4 falsifier files live under `src/exports/tests/` and are + // test code by location; the frozen oracle names the type by + // necessity, which is the point of freezing it. + if path.components().any(|c| c.as_os_str() == "tests") { + continue; + } + let text = fs::read_to_string(path).unwrap_or_else(|e| panic!("read {path:?}: {e}")); + let Some(prod) = production_half(&text) else { + interleaved.push(path.clone()); + continue; + }; + if without_comments(prod).contains(TYPE) { + let name = path + .file_name() + .and_then(|n| n.to_str()) + .unwrap_or_default() + .to_string(); + holders.insert(name); + } + } + assert!( + interleaved.is_empty(), + "in these files the `#[cfg(test)] mod` is not the last item, so the \ + production/test split cannot be made by brace-matching it: \ + {interleaved:?}. Move the test module to the end of the file, or \ + teach this scanner to handle what follows it" + ); + + // Anti-vacuity: the type must actually have been SEEN. A typo in `TYPE`, + // or a rename, would otherwise make an empty holder set look like + // perfect compliance. + assert!( + !holders.is_empty(), + "no production file names `{TYPE}` — either the type was renamed (in \ + which case rename it here too) or this scanner reads nothing" + ); + + let allowed: BTreeSet = ALLOWED.iter().map(|s| (*s).to_string()).collect(); + assert_eq!( + holders, allowed, + "the set of production modules holding plan ops moved.\n found: \ + {holders:?}\n allowed: {allowed:?}\nA NEW module here is how a \ + second plan evaluator arrives — the one PR4 deleted was a loop over \ + these exact ops. If it is legitimate, widen ALLOWED and say in its \ + doc comment why that module needs plan ops, in the same commit." + ); +} diff --git a/native/lgj-abi/tests/plan_eval_no_alloc.rs b/native/lgj-abi/tests/plan_eval_no_alloc.rs new file mode 100644 index 0000000..55f8ac2 --- /dev/null +++ b/native/lgj-abi/tests/plan_eval_no_alloc.rs @@ -0,0 +1,283 @@ +//! C-ALLOC — what `lgj_plan_eval` allocates per call, measured rather than +//! asserted. +//! +//! # Why this is its own binary, with exactly one `#[test]` +//! +//! The counter is a process-global `#[global_allocator]`. A second test +//! running concurrently adds its own bytes and the gate flakes. The fix for a +//! flake here is never to relax the bound — that is how the gate dies — it is +//! to keep this binary at one test. A separate integration binary also keeps +//! the counting allocator out of the in-crate suite, which it would otherwise +//! perturb. +//! +//! `lgj-abi` is `["cdylib", "rlib"]`, so the rlib links here and the exported +//! symbols are callable directly. +//! +//! # The instrument measures Rust, and the existing gates do not +//! +//! This repo's other allocation gates use `getThreadAllocatedBytes`, which +//! measures the JAVA heap. It cannot see a Rust `vec!` at all. Citing those +//! gates as evidence for this property would be an overclaim, so this is a +//! Rust-side instrument: a pass-through `GlobalAlloc` over `System` +//! incrementing an `AtomicUsize`, the same shape +//! `lance-graph-mask-risc/tests/no_alloc.rs` uses. +//! +//! # The property, stated exactly +//! +//! The old loop allocated `2 * n_words * 8` bytes per call — an accumulator +//! and a predicate buffer, both sized by the ROW COUNT. At 65,536 rows that +//! is 16 KiB per call, and it grew without bound as the population did. +//! +//! PR4 removes that. What it does NOT remove is the lowering's own +//! `Vec`: `Program` owns its op list, so building one allocates +//! bytes proportional to the number of OPS. That is a different quantity +//! entirely — a handful of ops against arbitrarily many rows — and the +//! honest claim is the one this test pins: +//! +//! **per-call allocation is independent of `n_rows`.** +//! +//! Which is why the two halves below are not decoration. The first measures +//! the same plan at row counts spanning three orders of magnitude and +//! requires the per-call bytes to be IDENTICAL, not merely small: a +//! row-proportional allocation cannot survive that, however cheap it looks at +//! one size. The second scales the OP count and requires the bytes to move, +//! so "identical across row counts" cannot be passing because the counter is +//! inert. +//! +//! # The unseen-row-count arm +//! +//! A thread-local arena grown monotonically means a SMALLER call carves a +//! prefix of a buffer that already exists. A gate that warms up over a fixed +//! sweep and then measures that same sweep is green while the property is +//! false — it would pass for a cache keyed by row count, which allocates the +//! first time it sees each distinct size and therefore depends on the +//! population's HISTORY, the exact thing C-ALLOC denies. `UNSEEN` is a row +//! count no earlier call in this test has used, measured after the warm-up. + +use std::alloc::{GlobalAlloc, Layout, System}; +use std::sync::atomic::{AtomicUsize, Ordering}; + +use lgj_abi::abi::*; +use lgj_abi::exports::{lgj_close, lgj_mask_create, lgj_pattern_open, lgj_plan_eval}; +use lgj_abi::fixture::{LANE_CLASSES, LANE_VALUES}; + +struct Counting; + +static BYTES: AtomicUsize = AtomicUsize::new(0); + +// SAFETY: a pure pass-through to `System`; the counter is the only addition. +unsafe impl GlobalAlloc for Counting { + unsafe fn alloc(&self, layout: Layout) -> *mut u8 { + BYTES.fetch_add(layout.size(), Ordering::Relaxed); + // SAFETY: same layout, same contract as the caller's. + unsafe { System.alloc(layout) } + } + unsafe fn dealloc(&self, ptr: *mut u8, layout: Layout) { + // SAFETY: `ptr` came from `alloc` above with this `layout`. + unsafe { System.dealloc(ptr, layout) } + } +} + +#[global_allocator] +static A: Counting = Counting; + +/// Row counts the sweep measures. `UNSEEN` is deliberately not among them and +/// deliberately smaller than the largest — see the module doc. +const SWEEP: [u64; 4] = [64, 1_000, 8_192, 65_536]; +const UNSEEN: u64 = 999; +const REPS: usize = 100; + +fn open(n: u64, seed: u64) -> u64 { + let mut h = 0u64; + assert_eq!(unsafe { lgj_pattern_open(n, seed, &mut h) }, LGJ_OK); + h +} + +fn mask_of(p: u64) -> u64 { + let mut h = 0u64; + assert_eq!( + unsafe { lgj_mask_create(p, LGJ_MASK_INIT_EMPTY, &mut h) }, + LGJ_OK + ); + h +} + +fn desc(op: u32, lane_id: u32, operand: i64, combine: u32) -> LgjOpDesc { + LgjOpDesc { + op, + lane_id, + operand, + combine, + _reserved: 0, + } +} + +/// A three-op plan: one AND-seeded predicate, one gated AND, one OR. +fn three_ops() -> Vec { + vec![ + desc(LGJ_OP_GT_I32, LANE_VALUES, -1000, LGJ_COMBINE_AND), + desc(LGJ_OP_LT_I32, LANE_VALUES, 1000, LGJ_COMBINE_AND), + desc(LGJ_OP_EQ_U32, LANE_CLASSES, 3, LGJ_COMBINE_OR), + ] +} + +/// Run `ops` against `(p, m)` `REPS` times and return the bytes allocated, +/// with the status checked so a call that silently failed cannot be measured +/// as a cheap one. +/// +/// **There is deliberately no warm-up call inside this function**, and that +/// is load-bearing rather than tidy. A per-measurement warm-up absorbs the +/// FIRST call at each row count — which is exactly and only where a cache +/// keyed by row count allocates. Measured: with one un-measured call here, +/// replacing the monotonic `resize` with a per-size `*buf = vec![…]` left the +/// whole gate GREEN. The warm-up was hiding the property the gate exists to +/// check, so every arm below is COLD at its own row count and the arena is +/// warmed exactly once, globally, at the largest size in the sweep. +fn measure(p: u64, m: u64, ops: &[LgjOpDesc]) -> usize { + let before = BYTES.load(Ordering::Relaxed); + for _ in 0..REPS { + let mut c = 0u64; + let s = unsafe { lgj_plan_eval(p, ops.as_ptr(), ops.len() as u32, m, &mut c) }; + assert_eq!(s, LGJ_OK); + } + BYTES.load(Ordering::Relaxed) - before +} + +/// Grow the thread-local arena to its high-water mark, un-measured. +/// +/// The very first plan a thread evaluates allocates the arena. That is the +/// allocation this test BOUNDS (once per thread, at the largest row count it +/// ever sees) rather than the per-call one it forbids, so it happens here, +/// outside every measured region, and exactly once. +fn warm_the_arena(p: u64, m: u64, ops: &[LgjOpDesc]) { + let mut c = 0u64; + assert_eq!( + unsafe { lgj_plan_eval(p, ops.as_ptr(), ops.len() as u32, m, &mut c) }, + LGJ_OK + ); +} + +/// FAILS IF: per-call allocation depends on the ROW COUNT — the property the +/// old loop's `2 * n_words * 8` bytes per call violated by construction — or +/// if the counter is inert, which would make the first half pass for the +/// wrong reason. +#[test] +fn per_call_allocation_does_not_depend_on_the_row_count() { + // Everything that legitimately allocates happens before any measurement: + // opening a pattern builds its fixture lanes, and creating a mask + // allocates its words. Neither is per-call work. + let handles: Vec<(u64, u64, u64)> = SWEEP + .iter() + .map(|&n| { + let p = open(n, 0xC0FFEE ^ n); + (n, p, mask_of(p)) + }) + .collect(); + let unseen = { + let p = open(UNSEEN, 0xBEEF); + (UNSEEN, p, mask_of(p)) + }; + let ops = three_ops(); + + // ONE global warm-up, at the LARGEST row count in the sweep, so the arena + // is already at its high-water mark before anything is measured. Every + // later call — including `UNSEEN`, which is smaller — must therefore + // carve a prefix of a buffer that already exists. + let largest = handles + .iter() + .max_by_key(|&&(n, _, _)| n) + .expect("the sweep is not empty"); + warm_the_arena(largest.1, largest.2, &ops); + + let mut per_call = Vec::new(); + for &(n, p, m) in &handles { + let bytes = measure(p, m, &ops); + assert_eq!( + bytes % REPS, + 0, + "n={n}: {bytes} bytes is not a clean per-call figure over {REPS} reps" + ); + per_call.push((n, bytes / REPS)); + } + + // The unseen arm runs AFTER the sweep, at a row count no earlier call + // used and smaller than the largest already seen. A monotonically grown + // arena carves a prefix and allocates nothing extra; a cache keyed by row + // count allocates here and only here. + // + // The divisibility check below is not decoration — skipping it (as an + // earlier version of this arm did, dividing straight into `unseen_bytes`) + // is a real gap: a ONE-OFF allocation smaller than `REPS` bytes on the + // very first `UNSEEN` call — exactly what a row-count-keyed cache would + // produce, populating its entry once and never again — truncates to zero + // under plain integer division and reads as identical to `baseline`, + // hiding precisely the defect this arm exists to catch. + let (_, up, um) = unseen; + let unseen_raw_bytes = measure(up, um, &ops); + assert_eq!( + unseen_raw_bytes % REPS, + 0, + "n={UNSEEN}: {unseen_raw_bytes} bytes is not a clean per-call figure \ + over {REPS} reps" + ); + let unseen_bytes = unseen_raw_bytes / REPS; + + let baseline = per_call[0].1; + for &(n, bytes) in &per_call { + assert_eq!( + bytes, baseline, + "per-call allocation moved with the row count: {per_call:?} — \ + {bytes} B at n={n} against {baseline} B at n={}. The old loop \ + allocated 2 * n_words * 8 per call and this is what forbids it", + per_call[0].0 + ); + } + assert_eq!( + unseen_bytes, baseline, + "n={UNSEEN} was never seen before and cost {unseen_bytes} B against \ + {baseline} B for every warmed size — allocation is following the \ + population's HISTORY, which is a cache keyed by row count, not a \ + monotonically grown arena" + ); + + // Can-it-fire. Scaling the OP count must move the figure, or "identical + // across row counts" above would hold just as well for a counter that + // never increments — the failure mode that makes an allocation gate + // decorative. + let (_, p, m) = handles[0]; + let mut wide = three_ops(); + for i in 0..29 { + wide.push(desc(LGJ_OP_NE_I32, LANE_VALUES, i, LGJ_COMBINE_AND)); + } + let wide_bytes = measure(p, m, &wide) / REPS; + assert!( + wide_bytes > baseline, + "a 32-op plan allocated {wide_bytes} B against {baseline} B for a \ + 3-op one: the instrument cannot see the lowering's own Vec, \ + so the row-count invariance above proves nothing" + ); + + // What remains is proportional to the OP count, not the row count, and it + // is small. The bound is measured, not chosen: re-measure before moving + // it. Its purpose is to catch a future change that reintroduces a + // row-sized buffer on a path this sweep happens not to reach. + assert!( + baseline <= 512, + "a 3-op plan allocates {baseline} B per call; the lowering's op list \ + is the only thing that should be there" + ); + + // Printed, never silently pinned: the numbers above are the evidence for + // every claim in this file, and `--nocapture` is how a future session + // re-reads them instead of trusting this comment. + println!("per-call bytes by row count: {per_call:?}"); + println!("n={UNSEEN} (never seen before): {unseen_bytes} B"); + println!("32-op plan: {wide_bytes} B (3-op: {baseline} B)"); + + for &(_, p, m) in &handles { + lgj_close(m); + lgj_close(p); + } + lgj_close(unseen.2); + lgj_close(unseen.1); +}