Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .claude/board/LATEST_STATE.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,13 @@
## 2026-10-04 — read-mode resolved once per population (branch `ccr-f6094d67-h6ulb3`, unmerged)

### Current Contract Inventory — net delta (`nan_projection.rs`, `soa_graph.rs`)
- `project_energy_nonfinite_resolved(rows, ValueSchema)` and
`energy_all_finite_resolved(rows, ValueSchema)`: no per-row lookup. The
mixed-batch API resolves once per classid run. Entry `D-HPS-2`.
- `soa_graph` helpers take the domain's resolved `TailVariant`.
- `lance-graph-mask-risc`: `Pred::MatchFacet16Strided` (16-byte strided
ternary match, over `ndarray::simd::ternary_match_strided16_to_mask`).

## 2026-10-04 — Register128 slab reading + bounded power sums (branch `ccr-1d39fce9-gdgy6k`, unmerged, D-LXC-29)

### Current Contract Inventory — net delta (`register128.rs`, `hotplug.rs`, `canonical_node.rs`)
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# D-HPS-2 — resolve once, then run the population (2026-10-04)

**STATUS:** measured (branch `ccr-f6094d67-h6ulb3`, unmerged)

Follows `D-HPS-1`. The metadata is resolved once per population, and the row
loops do no registry / ClassView / read-mode lookup.

- `soa_graph`: `project_snapshot` and `nearest_anchor` resolve the domain's
`TailVariant` once (`domain_tail`) and pass it to `hhtl_path` / `family_of` /
`identity_of`. Before: one `classid_read_mode` per helper call per row.
- `nan_projection`: new `project_energy_nonfinite_resolved(rows, schema)` and
`energy_all_finite_resolved(rows, schema)` take the population's
`ValueSchema` (e.g. `ResolvedReading::read_mode.value_schema`) and do no
lookup. The mixed API resolves once per run of equal classid and delegates
to them; its per-row residue is one classid compare. The schema gate and the
exponent-mask test are unchanged.
- Falsifiers: a test-only counter in `classid_read_mode` pins lookups at 1 for
both 1 and 1000 rows (soa_graph, mixed wrapper) and at 0 for the resolved
path. Disable runs, all red: per-row lookup in `hhtl_path`; per-row lookup in
the resolved loop; schema gate dropped; mixed wrapper reduced to runs of 1.
- Measured (release, `target-cpu=native`, 100k homogeneous rows):
per-row lookup 23.7–24.6 ns/row, resolved 5.0 ns/row, mixed wrapper
10.0–10.3 ns/row. soa_graph was not timed.
- The 128-bit matcher lives in ndarray (`ternary_match_strided16_to_mask`,
ndarray #340). 12 B vs 16 B at stride 512, 65,536 rows: 5.1–5.3 ns/row for
both; no measurable cost.

- mask-risc: `Pred::MatchFacet16Strided {lane, pattern: [u8; 16], care: [u8; 16]}`
lowers to `ternary_match_strided16_to_mask` (ndarray #340, merged). The view
is validated 16 bytes wide; the oracle compares all 16 bytes. Disable runs,
all red: executor on the 12 B kernel; oracle on 12 bytes; width check at 12.
No new IR shape; Quack unchanged (no caller needs a spelling yet).

**OPEN:**
- The bake paths were not changed: q2 `osint-bake/src/bin/fma.rs` does a
homogeneous per-node `classid_read_mode(CLASSID_FMA)` (another repo), and
deepnsm-v2 `promote.rs` `key_at` does one per call.
- symbiont `domino.rs` is not migrated: its boards are classid 0, so the mixed
wrapper already costs it one lookup.
3 changes: 2 additions & 1 deletion .claude/board/entries/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,12 +25,13 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately
opposite directions; the stranding this convention prevents shows up in
exactly one of them, never both.

212 entries, 2026-08-06 .. 2026-10-04.
213 entries, 2026-08-06 .. 2026-10-04.

| date | entry id | finding | file |
|---|---|---|---|
| 2026-10-04 | `D-LXC-22` | | [2026-10-04-wordnet-clam-chaoda-scope-and-grammar-read-params.md](2026-10-04-wordnet-clam-chaoda-scope-and-grammar-read-params.md) |
| 2026-10-04 | `D-HPS-1` | | [2026-10-04-spog-slab-hotplug-resolution.md](2026-10-04-spog-slab-hotplug-resolution.md) |
| 2026-10-04 | `D-HPS-2` | | [2026-10-04-resolve-once-population-execution.md](2026-10-04-resolve-once-population-execution.md) |
| 2026-10-04 | `D-LXC-29-R` | | [2026-10-04-register128-bounded-power-sums.md](2026-10-04-register128-bounded-power-sums.md) |
| 2026-10-04 | `D-LXC-25` | | [2026-10-04-deepnsm-v2-wechsel-lane-quorum.md](2026-10-04-deepnsm-v2-wechsel-lane-quorum.md) |
| 2026-10-04 | `D-LXC-27` | | [2026-10-04-deepnsm-v2-unseen-noun-gender.md](2026-10-04-deepnsm-v2-unseen-noun-gender.md) |
Expand Down
24 changes: 24 additions & 0 deletions crates/lance-graph-contract/src/canonical_node.rs
Original file line number Diff line number Diff line change
Expand Up @@ -1724,12 +1724,36 @@ static BUILTIN_READ_MODES: LazyLock<HashMap<u32, ReadMode>> = LazyLock::new(|| {
m
});

#[cfg(test)]
thread_local! {
/// Test-only count of [`classid_read_mode`] calls on this thread. Lets a
/// test assert that a population path resolves its reading a constant
/// number of times, not once per row. Thread-local, so parallel tests do
/// not see each other's lookups.
pub(crate) static READ_MODE_LOOKUPS: core::cell::Cell<usize> =
const { core::cell::Cell::new(0) };
}

/// Number of [`classid_read_mode`] calls on this thread since the last reset.
#[cfg(test)]
pub(crate) fn read_mode_lookups() -> usize {
READ_MODE_LOOKUPS.with(core::cell::Cell::get)
}

/// Reset this thread's [`classid_read_mode`] call count.
#[cfg(test)]
pub(crate) fn reset_read_mode_lookups() {
READ_MODE_LOOKUPS.with(|c| c.set(0));
}

/// Resolve a `classid` to its [`ReadMode`] — the single source both consumers
/// and OGAR inherit. Reads the [`BUILTIN_READ_MODES`] registry, falling through
/// to [`ReadMode::DEFAULT`] for any unconfigured classid (the key's own
/// zero-fallback ladder). [`NodeGuid::read_mode`] is the carrier-method form.
#[inline]
pub fn classid_read_mode(classid: u32) -> ReadMode {
#[cfg(test)]
READ_MODE_LOOKUPS.with(|c| c.set(c.get() + 1));
BUILTIN_READ_MODES
.get(&classid)
.copied()
Expand Down
198 changes: 159 additions & 39 deletions crates/lance-graph-contract/src/nan_projection.rs
Original file line number Diff line number Diff line change
Expand Up @@ -24,17 +24,20 @@
//! happened to zero out. Each row is therefore gated on its OWN resolved
//! `[ValueSchema::has]` before its `Energy` bytes are read at all.
//!
//! **What "branchless" still means after the gate, precisely (codex review,
//! 2026-07-30).** The FINITENESS TEST — the exponent-mask compare on the
//! four already-loaded bytes — is unchanged: still zero branches on the
//! value. That is not the same claim as "the sweep costs what it did
//! before." [`row_has_energy`] calls [`NodeGuid::read_mode`], which resolves
//! through [`classid_read_mode`] — a `HashMap` lookup behind a `LazyLock`,
//! not a bitmask. That lookup is real per-row work, added on top of the old
//! four-byte load, and can plausibly dominate it for an in-cache homogeneous
//! batch. Not benchmarked; do not read "branchless" below as "free" — if
//! this projection lands on a genuinely hot path, that lookup is the first
//! place to look before assuming the schema gate is costless.
//! **Where the schema is resolved.** The finiteness test — the exponent-mask
//! compare on four loaded bytes — has no branch on the value. The schema
//! lookup ([`classid_read_mode`], a `HashMap` behind a `LazyLock`) is kept
//! out of the per-row loop:
//!
//! - [`project_energy_nonfinite_resolved`] / [`energy_all_finite_resolved`]
//! take the population's already-resolved [`ValueSchema`] (e.g.
//! `ResolvedReading::read_mode.value_schema`). One schema check per call;
//! no registry lookup at all.
//! - [`project_energy_nonfinite`] / [`energy_all_finite`] accept a mixed
//! batch. They resolve once per RUN of equal classid and hand each run to
//! the resolved path, so a homogeneous batch costs one lookup, and a batch
//! alternating classids costs one per change. The only per-row metadata
//! work left there is comparing the row's classid with the run's.
//!
//! [`NanReport::skipped`] makes the gate's effect observable rather than a
//! silent no-op, per the workspace's can-it-fire testing rule.
Expand All @@ -48,7 +51,7 @@
//! [`NodeGuid::read_mode`]: crate::canonical_node::NodeGuid::read_mode
//! [`classid_read_mode`]: crate::canonical_node::classid_read_mode

use crate::canonical_node::{NodeRow, ValueTenant};
use crate::canonical_node::{classid_read_mode, NodeRow, ValueSchema, ValueTenant};

/// `true` iff an `f32` bit pattern is non-finite (Inf or NaN): the exponent
/// field is all-ones. No float materialised.
Expand Down Expand Up @@ -92,8 +95,8 @@ impl NanReport {
}

/// Read one board's `Energy` tenant as a raw `f32` bit pattern (no float load).
/// Caller MUST have already confirmed the row's schema materialises `Energy`
/// ([`row_has_energy`]) — this function does not gate.
/// Caller MUST have already confirmed the population's schema materialises
/// `Energy` — this function does not gate.
#[inline]
fn energy_bits(row: &NodeRow) -> u32 {
let off = ValueTenant::Energy.value_offset();
Expand All @@ -105,45 +108,93 @@ fn energy_bits(row: &NodeRow) -> u32 {
])
}

/// Does this row's OWN resolved schema materialise `Energy`? The one branch
/// this module adds — on schema presence, never on the float value.
#[inline]
fn row_has_energy(row: &NodeRow) -> bool {
row.key.read_mode().value_schema.has(ValueTenant::Energy)
}

/// Project a batch of canonical boards onto the NaN-detection surface by reading
/// each one's `Energy` tenant — schema-gated per row (see module docs). Read-only;
/// returns the indices of non-finite boards among those actually inspected.
/// This is the demoted singleton BindSpace — a projection, never a carrier.
pub fn project_energy_nonfinite(rows: &[NodeRow]) -> NanReport {
let mut total = 0usize;
let mut skipped = 0usize;
/// Project a population whose reading is ALREADY RESOLVED onto the
/// NaN-detection surface. `schema` is the population's value schema (for a
/// [`crate::hotplug::ResolvedReading`], its `read_mode.value_schema`); every
/// row is read under it. No registry or read-mode lookup happens here.
///
/// If `schema` does not materialise `Energy`, no `Energy` bytes are read: every
/// row is `skipped` and the report is clean. Otherwise each row's `Energy` is
/// tested with the integer exponent mask.
///
/// The caller owns the homogeneity claim. For a batch that may mix classids,
/// use [`project_energy_nonfinite`].
pub fn project_energy_nonfinite_resolved(rows: &[NodeRow], schema: ValueSchema) -> NanReport {
if !schema.has(ValueTenant::Energy) {
return NanReport {
total: 0,
nonfinite: Vec::new(),
skipped: rows.len(),
};
}
let mut nonfinite = Vec::new();
for (i, row) in rows.iter().enumerate() {
if !row_has_energy(row) {
skipped += 1;
continue;
}
total += 1;
if f32_bits_nonfinite(energy_bits(row)) {
nonfinite.push(i as u32);
}
}
NanReport {
total,
total: rows.len(),
nonfinite,
skipped,
skipped: 0,
}
}

/// Clean/dirty answer for a population whose reading is already resolved —
/// the sibling of [`project_energy_nonfinite_resolved`]. Early-outs on the
/// first non-finite board; `true` without reading anything when `schema` has
/// no `Energy`.
pub fn energy_all_finite_resolved(rows: &[NodeRow], schema: ValueSchema) -> bool {
!schema.has(ValueTenant::Energy) || rows.iter().all(|row| !f32_bits_nonfinite(energy_bits(row)))
}

/// Split `rows` into maximal runs of equal classid, resolving each run's
/// value schema once. Yields `(start index, run, schema)`.
fn schema_runs(rows: &[NodeRow]) -> impl Iterator<Item = (usize, &[NodeRow], ValueSchema)> {
let mut start = 0usize;
core::iter::from_fn(move || {
if start >= rows.len() {
return None;
}
let classid = rows[start].key.classid();
let len = rows[start..]
.iter()
.take_while(|r| r.key.classid() == classid)
.count();
let run = &rows[start..start + len];
let at = start;
start += len;
Some((at, run, classid_read_mode(classid).value_schema))
})
}

/// Project a batch of canonical boards onto the NaN-detection surface by reading
/// each one's `Energy` tenant — schema-gated per row (see module docs). Read-only;
/// returns the indices of non-finite boards among those actually inspected.
/// This is the demoted singleton BindSpace — a projection, never a carrier.
///
/// Accepts a batch mixing classids. Each run of equal classid is resolved once
/// and handed to [`project_energy_nonfinite_resolved`]; a caller that already
/// knows its population's reading should call that directly.
pub fn project_energy_nonfinite(rows: &[NodeRow]) -> NanReport {
let mut report = NanReport::default();
for (at, run, schema) in schema_runs(rows) {
let r = project_energy_nonfinite_resolved(run, schema);
report.total += r.total;
report.skipped += r.skipped;
report
.nonfinite
.extend(r.nonfinite.into_iter().map(|i| i + at as u32));
}
report
}

/// Fast clean/dirty answer without materialising the index list — the cheapest
/// projection (early-outs on the first non-finite board). Rows whose schema
/// omits `Energy` are skipped, not treated as a violation.
/// omits `Energy` are skipped, not treated as a violation. Mixed batches are
/// resolved once per run of equal classid.
pub fn energy_all_finite(rows: &[NodeRow]) -> bool {
rows.iter()
.filter(|row| row_has_energy(row))
.all(|row| !f32_bits_nonfinite(energy_bits(row)))
schema_runs(rows).all(|(_, run, schema)| energy_all_finite_resolved(run, schema))
}

#[cfg(test)]
Expand Down Expand Up @@ -256,4 +307,73 @@ mod tests {
"energy_all_finite must agree with project_energy_nonfinite"
);
}

// ── Resolved path: schema resolved once by the caller ─────────────────────

#[test]
fn resolved_path_does_no_read_mode_lookup() {
use crate::canonical_node::{read_mode_lookups, reset_read_mode_lookups};
let cognitive = classid_read_mode(NodeGuid::CLASSID_OSINT).value_schema;
assert!(cognitive.has(ValueTenant::Energy));
for n in [1usize, 1000] {
let mut rows: Vec<NodeRow> = (0..n).map(|i| board_with(i as f32)).collect();
rows[n - 1] = board_with(f32::NAN);
reset_read_mode_lookups();
let r = project_energy_nonfinite_resolved(&rows, cognitive);
let clean = energy_all_finite_resolved(&rows, cognitive);
assert_eq!(read_mode_lookups(), 0, "n = {n}");
// anti-vacuity: the rows were actually read
assert_eq!(r.total, n);
assert_eq!(r.nonfinite, vec![(n - 1) as u32]);
assert!(!clean);
}
}

#[test]
fn resolved_path_keeps_the_schema_gate() {
let compressed = classid_read_mode(NodeGuid::CLASSID_FMA).value_schema;
assert!(!compressed.has(ValueTenant::Energy));
let rows = vec![
board_with_classid(NodeGuid::CLASSID_FMA, f32::NAN),
board_with_classid(NodeGuid::CLASSID_FMA, f32::INFINITY),
];
// the bytes really are poisoned, so a missing gate would report them
assert!(f32_bits_nonfinite(energy_bits(&rows[0])));
let r = project_energy_nonfinite_resolved(&rows, compressed);
assert_eq!((r.total, r.skipped), (0, 2));
assert!(r.nonfinite.is_empty());
assert!(energy_all_finite_resolved(&rows, compressed));
}

#[test]
fn mixed_wrapper_resolves_once_per_classid_run() {
use crate::canonical_node::{read_mode_lookups, reset_read_mode_lookups};
for n in [1usize, 1000] {
let rows: Vec<NodeRow> = (0..n).map(|i| board_with(i as f32)).collect();
reset_read_mode_lookups();
let r = project_energy_nonfinite(&rows);
assert_eq!(read_mode_lookups(), 1, "homogeneous batch, n = {n}");
assert_eq!(r.total, n);
reset_read_mode_lookups();
assert!(energy_all_finite(&rows));
assert_eq!(read_mode_lookups(), 1, "homogeneous batch, n = {n}");
}
}

#[test]
fn mixed_wrapper_reports_indices_into_the_whole_batch() {
let rows = vec![
board_with_classid(NodeGuid::CLASSID_FMA, f32::NAN),
board_with_classid(NodeGuid::CLASSID_OSINT, f32::NAN),
board_with_classid(NodeGuid::CLASSID_FMA, f32::NAN),
board_with_classid(NodeGuid::CLASSID_OSINT, 1.0),
board_with_classid(NodeGuid::CLASSID_OSINT, f32::INFINITY),
];
let r = project_energy_nonfinite(&rows);
assert_eq!(r.nonfinite, vec![1, 4]);
assert_eq!((r.total, r.skipped), (3, 2));
assert!(!energy_all_finite(&rows));
// the FMA rows alone are skipped, not read
assert!(energy_all_finite(&[rows[0], rows[2]]));
}
}
Loading
Loading