|
| 1 | +## 2026-08-23 — E-ONE-RECEIPT-MANY-BORROWED-CONSUMERS-1 — the integration half of #1012: one versioned tokenization drives Tantivy, DeepNSM-v2 and a forward surface with zero re-tokenization, and the four gaps that stand between that and a carrier |
| 2 | + |
| 3 | +**Status:** FINDING — [MEASURED] (`PROBE-TOKEN-SEAM-1`, 37 gates, **13 |
| 4 | +disable-runs each verified red-then-green**; two committed real corpora — the |
| 5 | +in-tree KJV Genesis scene `PROBE-TOKEN-BPE-GEOMETRY-1` used, carried verbatim so |
| 6 | +the numbers are comparable, and Project Gutenberg's *Alice* split into 300 |
| 7 | +paragraphs, 75 514 B). The probe lives in `AdaWorldAPI/paperless-rs` |
| 8 | +(`crates/paperless-token`, `docs/TOKEN-SEAM-ARCHITECTURE.md`) because it needs a |
| 9 | +Tantivy dependency this workspace does not carry; this entry records the |
| 10 | +**lance-graph-side** findings. |
| 11 | +**This closes the integration half** of |
| 12 | +`E-TOKEN-BPE-CAN-FIT-NOT-YET-BUY-1` (#1012). That verdict measured that BPE FITS |
| 13 | +the `6×(8:8)` geometry and refused to buy a carrier. It did not ask whether one |
| 14 | +tokenization can SERVE several consumers at once. It can. |
| 15 | +**Confidence:** High for what is measured at these two corpus scales; the 8-bit |
| 16 | +lane saturated at 75 KB, so nothing here is a scale claim. |
| 17 | + |
| 18 | +**Headline.** ONE tokenization per span drove all three consumers and each added |
| 19 | +**zero** further tokenizations — not by discipline but by construction. Totals: |
| 20 | +313 source tokenizations for 308 spans plus 5 deliberate fixtures, and 1 QUERY |
| 21 | +tokenization on a deliberately separate counter (a query is different bytes; |
| 22 | +folding it into one number would make the claim a lie). |
| 23 | + |
| 24 | +**The two facts that make it adoptable, and neither was designed for this:** |
| 25 | + |
| 26 | +1. **DeepNSM-v2's library is already tokenizer-free.** `parse_to_spo(&[Tagged])` |
| 27 | + consumes `(WordId, Pos)` pairs and touches no string; the |
| 28 | + `split_whitespace`/`normalise` logic lives ONLY in `examples/bible_wave.rs` |
| 29 | + and `examples/genre_shapes.rs`. The seam needed **no change to the crate**. |
| 30 | + The `(WordId, Pos)` boundary is the shipped seam and nobody had used it as |
| 31 | + one. |
| 32 | +2. **Tantivy structurally cannot own offsets.** Its indexer reads `Token::text` |
| 33 | + and `Token::position`, uses `position_length` transiently, and reads |
| 34 | + `offset_from`/`offset_to` NOWHERE outside its own tests. Offsets are consumed |
| 35 | + only by snippet generation, which re-tokenizes the STORED text at query time. |
| 36 | + An index cannot become the ABI here even by accident. |
| 37 | + |
| 38 | +**Measured, and the numbers are the point rather than the verdict:** |
| 39 | + |
| 40 | +| | KJV scene | Alice | |
| 41 | +|---|---|---| |
| 42 | +| bytes / spans / tokens | 1 126 / 8 / 354 | 75 514 / 300 / 37 149 | |
| 43 | +| compression | 3.18× | **2.03×** | |
| 44 | +| distinct ids used | 137 | **247 of 255** | |
| 45 | +| resident lane bytes | 832 (74 % of source) | 55 572 (74 %) | |
| 46 | +| receipt share of resident | 54 % | 30 % | |
| 47 | +| particles/span p50/p95/max | 4 / 8 / 8 | 8 / 30 / 43 | |
| 48 | +| tokens per lexical unit p50/max | 1 / 7 | 2 / 15 | |
| 49 | +| tokens straddling a word boundary | 30 | 58 | |
| 50 | + |
| 51 | +- **The 8-bit lane SATURATES.** 247 of 255 ids on 75 KB of ordinary English, with |
| 52 | + compression already fallen from 3.18× to 2.03×. The canon's answer — the hi |
| 53 | + byte of each `(8:8)` pair as a PAGE lane, two separate bytes, never a widened |
| 54 | + `u16` — is **untested**. Until it is measured no scale claim about token BPE |
| 55 | + should be made, and this supersedes any reading of #1012's 3.35× as a |
| 56 | + corpus-independent figure. |
| 57 | +- **The resident lane is ~74 % of the source text, not a fraction of it**, and |
| 58 | + at these span sizes **framing is 30–54 % of it**. A 56-byte receipt against |
| 59 | + 12-byte particles means the RECEIPT's column layout matters more than the |
| 60 | + particle's. #1012 could not see this — it had no receipt. |
| 61 | +- **Cardinality is not 1:1 in either direction**, so a BPE↔`WordId` projection |
| 62 | + is a real function, not a relabelling. Nothing in the seam assigns a `WordId` |
| 63 | + to a BPE token; that would be a second vocabulary wearing DeepNSM's |
| 64 | + coordinate system. |
| 65 | +- **Byte offsets are DERIVED**, by prefix sum over a per-id decoded-length |
| 66 | + table. The receipt stores no offset column at all. |
| 67 | + |
| 68 | +**Four lance-graph-side gaps, each named rather than worked around:** |
| 69 | + |
| 70 | +1. **No shipped token continuation mechanism anywhere.** The nearest precedent |
| 71 | + in SHAPE is `rail_geometry::RailCarving::AxisSlab { reg, cont: Option<usize> }`, |
| 72 | + which chains one register to one continuation and caps at `RAIL_MAX_DEPTH = 24` |
| 73 | + levels — **below the measured p50 of 4 particles**, so it does not fit. The |
| 74 | + probe uses a contiguous run (`first_particle + particle_count + token_count`). |
| 75 | + Stated honestly there are TWO lawful framings and the trade is exact: |
| 76 | + `particle_count` alone bounds the run and a PAD scan inside that bound is |
| 77 | + already exact BECAUSE PAD is reserved (cost: one vocabulary slot, which at a |
| 78 | + 255-cap that saturates is not free); or `token_count` costs 4 bytes and frees |
| 79 | + the slot for a full 256-id alphabet. What is unlawful is inferring the end |
| 80 | + from padding with no bound — measured, a lane-wide PAD scan overshoots |
| 81 | + receipt 0 by 10 tokens straight into receipt 1. |
| 82 | +2. **`ValueTenant` has no token variant** (16 discriminants, none for text), so a |
| 83 | + lawful resident lane must implement `SoaEnvelope` or land as a new tenant. |
| 84 | + The probe's lane is a probe-local `Vec` and says so. |
| 85 | +3. **There is no callable part-of-speech surface, and the reason is a decision |
| 86 | + already taken.** `coca_pos`/`archaic_pos`/`normalise` are byte-identical in |
| 87 | + BOTH deepnsm-v2 examples, above a comment stating `deepnsm_v2::lexicon` was |
| 88 | + DELETED after an audit found `lance-graph-planner`'s `insight_coca_read` |
| 89 | + already grounds it. That grounding does not reach a lean consumer: |
| 90 | + `insight_coca_read` is itself an **example binary**, in a crate carrying |
| 91 | + `serde`/`serde_yml`/`tokio`/`ndarray`, and its master `lexicon.tsv` is absent |
| 92 | + from this checkout. The probe restated the twenty-line tagger rather than |
| 93 | + re-litigate the deletion — recorded so the next consumer has the evidence the |
| 94 | + audit did not. |
| 95 | +4. **The semantic half is unexercised.** `cam96_codebook.bin` / `cam96_codes.bin` |
| 96 | + are release assets, absent here, so palette256² DISTANCE never ran. Only the |
| 97 | + lexical/grammar half was measured. |
| 98 | + |
| 99 | +**Polars: refuted, and the honest form is weaker than the question invited.** A |
| 100 | +sweep of nine checkouts found **zero** `polars` occurrences in any manifest or |
| 101 | +source; every `DataFrame` mention is prose. `paperless-rs` and `tesseract-rs` |
| 102 | +declare none of arrow/datafusion/lance/lancedb. There was nothing to remove. The |
| 103 | +structured-evidence path is likewise not tabular algebra: |
| 104 | +`lance-graph-arm-discovery` takes `Dataset { spec: FeatureSpec, rows: |
| 105 | +Vec<Vec<u32>> }` — category-index rows against a schema. |
| 106 | + |
| 107 | +**Method note, and it is the transferable part.** An independent vacuity audit of |
| 108 | +the finished probe found FIVE holes: a gate asserting byte counts but never span |
| 109 | +counts (the CRLF bug that collapsed 300 spans into 1 would have re-passed it), a |
| 110 | +threshold true by construction, an assertion about a type signature rather than |
| 111 | +behaviour, an unconditional prefix check, and an unexercised ASCII-vs-Unicode |
| 112 | +whitespace divergence. All five are fixed; four gained their own disable-runs; |
| 113 | +the fifth is bounded by a measured count of 0. Separately, TWO disable-runs were |
| 114 | +themselves wrong first — one relaxed a knob that does not bind on the fixture, |
| 115 | +one targeted a mechanism the gate did not actually rest on — and an early batch |
| 116 | +reported "no failure" six times in a row because the probe binary path was wrong |
| 117 | +and nothing ran. **A knob that does not bind is not a disable; a fixture's SHAPE |
| 118 | +is part of a test's coverage; and a null result is a claim about the apparatus |
| 119 | +until proven otherwise.** |
| 120 | + |
| 121 | +**⚠ SELF-CORRECTION, same session, after being pointed at `ogar-doc-ir`.** Two |
| 122 | +claims above were wrong and are corrected here rather than left standing: |
| 123 | + |
| 124 | +1. **The seam invented an identity that already existed.** The first cut minted |
| 125 | + `source_id`/`span_id` integers. `ogar_doc_ir::DocIr` already answers all |
| 126 | + three questions a tokenization receipt asks — `content_sha256` for WHICH |
| 127 | + document, `(DocPage::number, Region::reading_order)` for WHICH span, and |
| 128 | + `Region::text` for the span's canonical text. The probe was re-cut to read |
| 129 | + them (`docir.rs`; gates `T-DOCIR` / `T-DOCIR-KEY` / `T-DOCIR-SPANS`, 41 |
| 130 | + total, 18 disable-runs). Note the crate's own docs CORRECT its plan's first |
| 131 | + sketch on what that hash is: a **per-acquisition dedup key**, not a |
| 132 | + cross-retina identity — which is exactly the right reading for a receipt, |
| 133 | + because you tokenize bytes. |
| 134 | +2. **"The OCR boundary supplies no byte offsets" is RETIRED as a gap.** It |
| 135 | + supplies no PAGE-wide offset and does not need to: a region owns its text, |
| 136 | + so an offset is region-local, and `ogar-from-docv1::region_text` is where the |
| 137 | + `leading_space`-aware join already happens. What remains is far smaller — a |
| 138 | + sub-region span needs a non-zero `byte_from`, which the receipt already |
| 139 | + carries and no producer emits. |
| 140 | + |
| 141 | +Also corrected: the `247 of 255` figure quoted above is the count of ids |
| 142 | +APPEARING in the lane, not the vocabulary size. Measured, the trained table is |
| 143 | +**full at 255/255** on Alice (and on the whole 170 KB file), and **180 of 255** |
| 144 | +on the KJV fixture — where, as #1016's own record of that fixture says, the |
| 145 | +CORPUS rather than the cap set it. The saturation conclusion holds and is |
| 146 | +stronger; the number was the wrong quantity. |
| 147 | + |
| 148 | +**Untouched by this.** HHTL is address geometry and BPE is tokenization — |
| 149 | +#1012's measured refutation of the merge tree as a radix prefix partition |
| 150 | +stands. Content never travels in classid: the contract id is a FIELD on the |
| 151 | +receipt, gated by a grep of the library's own non-comment source. |
1 | 152 | ## 2026-08-23 — E-THE-SEVEN-OPCODE-PROJECTION-IS-NOT-X86-AND-THE-CHAIN-CARRIER-WINS-1 — four wave probes: the chain carrier confirmed, the vocabulary survives optimization, and the boundary that qualifies all of it |
2 | 153 |
|
3 | 154 | **Status:** FINDING — [MEASURED] × 4 (`PROBE-R2IL-OPTIMIZATION-TRANSFER-1` 5/5, |
|
0 commit comments