Skip to content

Commit 3d113c2

Browse files
authored
Merge pull request #1017 from AdaWorldAPI/claude/bpe-tokenization-architecture-3xd4eh
the integration half of #1012: one receipt, three borrowed consumers
2 parents 6fc8507 + e7a5f44 commit 3d113c2

6 files changed

Lines changed: 5896 additions & 0 deletions

File tree

‎.claude/board/AGENT_LOG.md‎

Lines changed: 51 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,54 @@
1+
## 2026-08-23 — the token-seam arc: four read-only research lanes, one probe, one vacuity audit
2+
3+
- **Why:** an operator brief asked whether ONE versioned BPE tokenization can
4+
simultaneously serve Tantivy lexical indexing, DeepNSM-v2 lexical/grammar
5+
projection and an LSTM forward-prediction input surface, without
6+
retokenizing, without rebuilding a DataFrame, and without a second cognitive
7+
population — the integration half `E-TOKEN-BPE-CAN-FIT-NOT-YET-BUY-1` (#1012)
8+
explicitly did not ask. Tiering: Opus on the main thread for the architecture,
9+
the probe and every gate; Sonnet for the bounded read-only lanes; no worker
10+
ran cargo (the orchestrator compiled centrally, per the shared-target rule).
11+
- **Lane A — DeepNSM-v2 lexical contract.** The decisive finding: the LIBRARY is
12+
already tokenizer-free. `parse_to_spo(&[Tagged])` takes `(WordId, Pos)` and no
13+
string; `split_whitespace`/`normalise` live only in the two examples. The seam
14+
therefore needed no change to the crate. Also: `academic_20k.csv` is present
15+
(20 845 rows, 18 559 distinct surface forms); `bible_vocab.txt` and the cam96
16+
codebook/codes are ABSENT.
17+
- **Lane B — Polars / online-path falsifier.** Zero `polars` occurrences across
18+
nine checkouts; every `DataFrame` mention is prose. `paperless-rs` and
19+
`tesseract-rs` declare no arrow/datafusion/lance/lancedb. Also established
20+
that `doc.v1` carries bbox/conf/leading_space and NO offset or span field —
21+
which is the gap that blocks the seam on real scanned documents.
22+
- **Lane C — Tantivy indexing-path audit.** The indexer never reads
23+
`offset_from`/`offset_to` outside its own tests; snippets re-tokenize STORED
24+
text at query time; `PreTokenizedString` costs ≈ `4 + 2N` allocations because
25+
`segment_writer` deep-clones the boxed value. That last number is why the seam
26+
uses a custom tokenizer with one reused `Token` buffer.
27+
- **Lane D — SoA lane / continuation precedent.** No shipped token continuation
28+
mechanism anywhere; the nearest in shape, `RailCarving::AxisSlab`, caps at 24
29+
levels — below the measured p50 of 4 particles. `ValueTenant` has 16 variants
30+
and none for text.
31+
- **The probe (main thread, Opus).** `PROBE-TOKEN-SEAM-1` — 37 gates, 13
32+
disable-runs verified red-then-green, in `AdaWorldAPI/paperless-rs`
33+
(`crates/paperless-token`), with `docs/TOKEN-SEAM-ARCHITECTURE.md` as the
34+
bounded architecture. Result: `E-ONE-RECEIPT-MANY-BORROWED-CONSUMERS-1`.
35+
- **Lane E — vacuity audit (Sonnet, read-only, adversarial).** Found FIVE holes
36+
in the finished probe, all real: a gate checking byte counts but not span
37+
counts (the CRLF bug that collapsed 300 spans into 1 would have re-passed), a
38+
threshold true by construction, an assertion about a type signature rather
39+
than behaviour, an unconditional prefix check, and an unexercised
40+
ASCII-vs-Unicode whitespace divergence. All five fixed; four gained their own
41+
disable-runs; the fifth is bounded by a measured count of zero.
42+
- **Two method failures worth the entry.** (1) Two disable-runs were themselves
43+
wrong first — one relaxed a constant that does not bind on the fixture, one
44+
targeted a mechanism the gate did not rest on — so both "passed" while proving
45+
nothing. (2) An early disable batch reported "no failure" six times in a row
46+
because the probe binary path was wrong and nothing ran at all. A null result
47+
is a claim about the apparatus until proven otherwise.
48+
- **Outcome:** #1017 (board), plus a paperless-rs commit that is **committed
49+
locally and BLOCKED from pushing** — the GitHub App has no access to
50+
`AdaWorldAPI/paperless-rs` for this org, verified through both the session
51+
proxy and a proxy-bypassed attempt.
152
## 2026-08-23 — autoattended R2IL wave: 4 Sonnet probe workers + 1 Sonnet scribe + 1 Opus synthesis + 1 Haiku guarded executor
253

354
- **Why:** operator directive to run the pattern autonomously — "sonnet agents

‎.claude/board/EPIPHANIES.md‎

Lines changed: 151 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,154 @@
1+
## 2026-08-23 — E-ONE-RECEIPT-MANY-BORROWED-CONSUMERS-1 — the integration half of #1012: one versioned tokenization drives Tantivy, DeepNSM-v2 and a forward surface with zero re-tokenization, and the four gaps that stand between that and a carrier
2+
3+
**Status:** FINDING — [MEASURED] (`PROBE-TOKEN-SEAM-1`, 37 gates, **13
4+
disable-runs each verified red-then-green**; two committed real corpora — the
5+
in-tree KJV Genesis scene `PROBE-TOKEN-BPE-GEOMETRY-1` used, carried verbatim so
6+
the numbers are comparable, and Project Gutenberg's *Alice* split into 300
7+
paragraphs, 75 514 B). The probe lives in `AdaWorldAPI/paperless-rs`
8+
(`crates/paperless-token`, `docs/TOKEN-SEAM-ARCHITECTURE.md`) because it needs a
9+
Tantivy dependency this workspace does not carry; this entry records the
10+
**lance-graph-side** findings.
11+
**This closes the integration half** of
12+
`E-TOKEN-BPE-CAN-FIT-NOT-YET-BUY-1` (#1012). That verdict measured that BPE FITS
13+
the `6×(8:8)` geometry and refused to buy a carrier. It did not ask whether one
14+
tokenization can SERVE several consumers at once. It can.
15+
**Confidence:** High for what is measured at these two corpus scales; the 8-bit
16+
lane saturated at 75 KB, so nothing here is a scale claim.
17+
18+
**Headline.** ONE tokenization per span drove all three consumers and each added
19+
**zero** further tokenizations — not by discipline but by construction. Totals:
20+
313 source tokenizations for 308 spans plus 5 deliberate fixtures, and 1 QUERY
21+
tokenization on a deliberately separate counter (a query is different bytes;
22+
folding it into one number would make the claim a lie).
23+
24+
**The two facts that make it adoptable, and neither was designed for this:**
25+
26+
1. **DeepNSM-v2's library is already tokenizer-free.** `parse_to_spo(&[Tagged])`
27+
consumes `(WordId, Pos)` pairs and touches no string; the
28+
`split_whitespace`/`normalise` logic lives ONLY in `examples/bible_wave.rs`
29+
and `examples/genre_shapes.rs`. The seam needed **no change to the crate**.
30+
The `(WordId, Pos)` boundary is the shipped seam and nobody had used it as
31+
one.
32+
2. **Tantivy structurally cannot own offsets.** Its indexer reads `Token::text`
33+
and `Token::position`, uses `position_length` transiently, and reads
34+
`offset_from`/`offset_to` NOWHERE outside its own tests. Offsets are consumed
35+
only by snippet generation, which re-tokenizes the STORED text at query time.
36+
An index cannot become the ABI here even by accident.
37+
38+
**Measured, and the numbers are the point rather than the verdict:**
39+
40+
| | KJV scene | Alice |
41+
|---|---|---|
42+
| bytes / spans / tokens | 1 126 / 8 / 354 | 75 514 / 300 / 37 149 |
43+
| compression | 3.18× | **2.03×** |
44+
| distinct ids used | 137 | **247 of 255** |
45+
| resident lane bytes | 832 (74 % of source) | 55 572 (74 %) |
46+
| receipt share of resident | 54 % | 30 % |
47+
| particles/span p50/p95/max | 4 / 8 / 8 | 8 / 30 / 43 |
48+
| tokens per lexical unit p50/max | 1 / 7 | 2 / 15 |
49+
| tokens straddling a word boundary | 30 | 58 |
50+
51+
- **The 8-bit lane SATURATES.** 247 of 255 ids on 75 KB of ordinary English, with
52+
compression already fallen from 3.18× to 2.03×. The canon's answer — the hi
53+
byte of each `(8:8)` pair as a PAGE lane, two separate bytes, never a widened
54+
`u16` — is **untested**. Until it is measured no scale claim about token BPE
55+
should be made, and this supersedes any reading of #1012's 3.35× as a
56+
corpus-independent figure.
57+
- **The resident lane is ~74 % of the source text, not a fraction of it**, and
58+
at these span sizes **framing is 30–54 % of it**. A 56-byte receipt against
59+
12-byte particles means the RECEIPT's column layout matters more than the
60+
particle's. #1012 could not see this — it had no receipt.
61+
- **Cardinality is not 1:1 in either direction**, so a BPE↔`WordId` projection
62+
is a real function, not a relabelling. Nothing in the seam assigns a `WordId`
63+
to a BPE token; that would be a second vocabulary wearing DeepNSM's
64+
coordinate system.
65+
- **Byte offsets are DERIVED**, by prefix sum over a per-id decoded-length
66+
table. The receipt stores no offset column at all.
67+
68+
**Four lance-graph-side gaps, each named rather than worked around:**
69+
70+
1. **No shipped token continuation mechanism anywhere.** The nearest precedent
71+
in SHAPE is `rail_geometry::RailCarving::AxisSlab { reg, cont: Option<usize> }`,
72+
which chains one register to one continuation and caps at `RAIL_MAX_DEPTH = 24`
73+
levels — **below the measured p50 of 4 particles**, so it does not fit. The
74+
probe uses a contiguous run (`first_particle + particle_count + token_count`).
75+
Stated honestly there are TWO lawful framings and the trade is exact:
76+
`particle_count` alone bounds the run and a PAD scan inside that bound is
77+
already exact BECAUSE PAD is reserved (cost: one vocabulary slot, which at a
78+
255-cap that saturates is not free); or `token_count` costs 4 bytes and frees
79+
the slot for a full 256-id alphabet. What is unlawful is inferring the end
80+
from padding with no bound — measured, a lane-wide PAD scan overshoots
81+
receipt 0 by 10 tokens straight into receipt 1.
82+
2. **`ValueTenant` has no token variant** (16 discriminants, none for text), so a
83+
lawful resident lane must implement `SoaEnvelope` or land as a new tenant.
84+
The probe's lane is a probe-local `Vec` and says so.
85+
3. **There is no callable part-of-speech surface, and the reason is a decision
86+
already taken.** `coca_pos`/`archaic_pos`/`normalise` are byte-identical in
87+
BOTH deepnsm-v2 examples, above a comment stating `deepnsm_v2::lexicon` was
88+
DELETED after an audit found `lance-graph-planner`'s `insight_coca_read`
89+
already grounds it. That grounding does not reach a lean consumer:
90+
`insight_coca_read` is itself an **example binary**, in a crate carrying
91+
`serde`/`serde_yml`/`tokio`/`ndarray`, and its master `lexicon.tsv` is absent
92+
from this checkout. The probe restated the twenty-line tagger rather than
93+
re-litigate the deletion — recorded so the next consumer has the evidence the
94+
audit did not.
95+
4. **The semantic half is unexercised.** `cam96_codebook.bin` / `cam96_codes.bin`
96+
are release assets, absent here, so palette256² DISTANCE never ran. Only the
97+
lexical/grammar half was measured.
98+
99+
**Polars: refuted, and the honest form is weaker than the question invited.** A
100+
sweep of nine checkouts found **zero** `polars` occurrences in any manifest or
101+
source; every `DataFrame` mention is prose. `paperless-rs` and `tesseract-rs`
102+
declare none of arrow/datafusion/lance/lancedb. There was nothing to remove. The
103+
structured-evidence path is likewise not tabular algebra:
104+
`lance-graph-arm-discovery` takes `Dataset { spec: FeatureSpec, rows:
105+
Vec<Vec<u32>> }` — category-index rows against a schema.
106+
107+
**Method note, and it is the transferable part.** An independent vacuity audit of
108+
the finished probe found FIVE holes: a gate asserting byte counts but never span
109+
counts (the CRLF bug that collapsed 300 spans into 1 would have re-passed it), a
110+
threshold true by construction, an assertion about a type signature rather than
111+
behaviour, an unconditional prefix check, and an unexercised ASCII-vs-Unicode
112+
whitespace divergence. All five are fixed; four gained their own disable-runs;
113+
the fifth is bounded by a measured count of 0. Separately, TWO disable-runs were
114+
themselves wrong first — one relaxed a knob that does not bind on the fixture,
115+
one targeted a mechanism the gate did not actually rest on — and an early batch
116+
reported "no failure" six times in a row because the probe binary path was wrong
117+
and nothing ran. **A knob that does not bind is not a disable; a fixture's SHAPE
118+
is part of a test's coverage; and a null result is a claim about the apparatus
119+
until proven otherwise.**
120+
121+
**⚠ SELF-CORRECTION, same session, after being pointed at `ogar-doc-ir`.** Two
122+
claims above were wrong and are corrected here rather than left standing:
123+
124+
1. **The seam invented an identity that already existed.** The first cut minted
125+
`source_id`/`span_id` integers. `ogar_doc_ir::DocIr` already answers all
126+
three questions a tokenization receipt asks — `content_sha256` for WHICH
127+
document, `(DocPage::number, Region::reading_order)` for WHICH span, and
128+
`Region::text` for the span's canonical text. The probe was re-cut to read
129+
them (`docir.rs`; gates `T-DOCIR` / `T-DOCIR-KEY` / `T-DOCIR-SPANS`, 41
130+
total, 18 disable-runs). Note the crate's own docs CORRECT its plan's first
131+
sketch on what that hash is: a **per-acquisition dedup key**, not a
132+
cross-retina identity — which is exactly the right reading for a receipt,
133+
because you tokenize bytes.
134+
2. **"The OCR boundary supplies no byte offsets" is RETIRED as a gap.** It
135+
supplies no PAGE-wide offset and does not need to: a region owns its text,
136+
so an offset is region-local, and `ogar-from-docv1::region_text` is where the
137+
`leading_space`-aware join already happens. What remains is far smaller — a
138+
sub-region span needs a non-zero `byte_from`, which the receipt already
139+
carries and no producer emits.
140+
141+
Also corrected: the `247 of 255` figure quoted above is the count of ids
142+
APPEARING in the lane, not the vocabulary size. Measured, the trained table is
143+
**full at 255/255** on Alice (and on the whole 170 KB file), and **180 of 255**
144+
on the KJV fixture — where, as #1016's own record of that fixture says, the
145+
CORPUS rather than the cap set it. The saturation conclusion holds and is
146+
stronger; the number was the wrong quantity.
147+
148+
**Untouched by this.** HHTL is address geometry and BPE is tokenization —
149+
#1012's measured refutation of the merge tree as a radix prefix partition
150+
stands. Content never travels in classid: the contract id is a FIELD on the
151+
receipt, gated by a grep of the library's own non-comment source.
1152
## 2026-08-23 — E-THE-SEVEN-OPCODE-PROJECTION-IS-NOT-X86-AND-THE-CHAIN-CARRIER-WINS-1 — four wave probes: the chain carrier confirmed, the vocabulary survives optimization, and the boundary that qualifies all of it
2153

3154
**Status:** FINDING — [MEASURED] × 4 (`PROBE-R2IL-OPTIMIZATION-TRANSFER-1` 5/5,

‎.claude/board/LATEST_STATE.md‎

Lines changed: 29 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,32 @@
1+
## 2026-08-23 — #1017 OPEN — the token seam: #1012's integration half answered, and the 8-bit lane's ceiling found
2+
3+
- **What exists now:** `PROBE-TOKEN-SEAM-1` (37 gates, 13 disable-runs) in
4+
`AdaWorldAPI/paperless-rs crates/paperless-token`, with
5+
`docs/TOKEN-SEAM-ARCHITECTURE.md` as its bounded architecture. It answers the
6+
question #1012 left open: ONE versioned BPE tokenization of a span drives
7+
Tantivy, DeepNSM-v2 and a forward-prediction input surface simultaneously,
8+
each consumer BORROWING, none re-tokenizing. Neither Tantivy nor DeepNSM-v2
9+
needed a line changed.
10+
- **Standing laws banked:** ONE SOURCE SPAN → ONE TOKENIZATION RECEIPT;
11+
TOKENIZE ONCE, PROJECT MANY TIMES; AN INDEX MAY ACCELERATE THE ABI, IT MUST
12+
NEVER BECOME THE ABI; BPE SEQUENCE IDENTITY IS NOT A SEMANTIC WORD
13+
COORDINATE; and (method) A KNOB THAT DOES NOT BIND IS NOT A DISABLE — a
14+
fixture's SHAPE is part of a test's coverage.
15+
- **The number that bounds prior work:** the 8-bit id lane saturates at 75 KB
16+
(247/255 ids, compression 3.18× → 2.03×). #1012's 3.35× is a 1 KB figure and
17+
must not be read as corpus-independent. The hi-byte PAGE lane is the next
18+
probe and no scale claim survives without it.
19+
- **Open, in order:** (1) the paged vocabulary; (2) a real retina — this probe
20+
builds its `DocIr` from text, so the next one should take
21+
`ogar-from-docv1` on an actual scan and a `spider_doc_ir` crawl of the same
22+
content and check both present the same span-population shape;
23+
(3) a lawful resident lane (`SoaEnvelope` or a new `ValueTenant`, designed
24+
against the measured 30–54 % framing overhead); (4) a real forward arm —
25+
the seam supplies the input, but which representation a trained model prefers
26+
is untested; (5) a structured-evidence corpus to exercise the parallel typed
27+
path (`arm-discovery`'s `FeatureSpec` + category-index rows) and the shared
28+
span identity where the two meet.
29+
130
## 2026-08-23 — #1006..#1014 MERGED — the belief-ABI arc: Step 1 audit, root-law probes, Step 2 ruling request, frontier Phases 1+2, the real-episode measurement
231

332
Nine PRs (#1010/#1011 landing stacked via #1009's merge) closed the arc:

‎.claude/board/PR_ARC_INVENTORY.md‎

Lines changed: 39 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,42 @@
1+
## 2026-08-23 — lance-graph #1017 (OPEN) — the integration half of #1012: one receipt, three borrowed consumers
2+
3+
- **Added:** `E-ONE-RECEIPT-MANY-BORROWED-CONSUMERS-1` — the board record of
4+
`PROBE-TOKEN-SEAM-1` (37 gates, 13 disable-runs red-then-green; probe code in
5+
`AdaWorldAPI/paperless-rs crates/paperless-token` + `docs/TOKEN-SEAM-ARCHITECTURE.md`,
6+
which lives there because it needs a Tantivy dep this workspace does not carry).
7+
ONE tokenization per span drove Tantivy, DeepNSM-v2 and a forward-prediction
8+
surface; each added ZERO further tokenizations (313 source for 308 spans + 5
9+
fixtures; 1 query on a separate counter).
10+
- **Locked:** DeepNSM-v2's library is ALREADY the seam — `parse_to_spo(&[Tagged])`
11+
takes `(WordId, Pos)` and no string, so the crate needed no change; a BPE token
12+
is never assigned a `WordId` (different id spaces, cardinality measured
13+
non-1:1 in both directions); byte offsets are DERIVED by prefix sum over a
14+
per-id length table, so a receipt stores no offset column; Tantivy cannot own
15+
offsets (its indexer never reads them).
16+
- **Measured, and it bounds #1012's headline:** the 8-bit lane SATURATES — 247
17+
of 255 ids on 75 KB, compression 3.18× → **2.03×**. The resident lane is 74 %
18+
of source and framing is 30–54 % of THAT (56-byte receipt vs 12-byte
19+
particles), so the receipt's layout outranks the particle's.
20+
- **Self-corrected in-session:** the first cut minted `source_id`/`span_id`;
21+
re-cut onto `ogar_doc_ir::DocIr` (`content_sha256` + `(page, reading_order)` +
22+
`Region::text`), so the receipt mints nothing. That RETIRES the
23+
"no byte offsets at the OCR boundary" gap — offsets are region-local — and
24+
corrects `247 of 255` (ids appearing in the lane) to a table that is FULL at
25+
255/255 on Alice and 180/255 on the KJV fixture.
26+
- **Deferred / named:** no shipped token continuation mechanism (the
27+
`RailCarving::AxisSlab` precedent caps at 24 levels, under the measured p50 of
28+
4 particles); `ValueTenant` has no token variant; no callable PoS surface —
29+
`deepnsm_v2::lexicon` was deliberately deleted and the `insight_coca_read`
30+
grounding cited for it is an example binary outside a lean consumer's
31+
dependency barrier; cam96 codebook/codes ABSENT so the semantic half never
32+
ran; the hi-byte PAGE lane untested, which is the next probe.
33+
- **Refuted:** Polars in the online path — zero occurrences across nine
34+
checkouts; `paperless-rs`/`tesseract-rs` declare no arrow/datafusion/lance/lancedb
35+
at all. There was nothing to remove.
36+
- **Docs:** `E-ONE-RECEIPT-MANY-BORROWED-CONSUMERS-1`; LATEST_STATE updated.
37+
- **Confidence:** High for the two measured corpora; no scale claim — the
38+
vocabulary was full at 75 KB.
39+
140
## 2026-08-23 — GAP FILLED — arc rows for #976..#1005 (consolidated by the orchestrator from a wave scribe's primary-source reconstruction)
241

342
The gap marker recorded below this block is now DISCHARGED for #977..#1005.

0 commit comments

Comments
 (0)