feat(dataset): write compaction's spilled lineage into the fragment's data file - #9347
Merged
Merged
Conversation
This was referenced Sep 17, 2026
BubbleCal
force-pushed
the
yang/oss-2269-5-row-lineage-update
branch
from
September 17, 2026 13:39
cd2243c to
be46048
Compare
BubbleCal
force-pushed
the
yang/oss-2269-7-row-lineage-main-file
branch
from
September 17, 2026 13:39
c2ad5b8 to
d57cbaa
Compare
BubbleCal
force-pushed
the
yang/oss-2269-5-row-lineage-update
branch
from
September 18, 2026 17:36
be46048 to
7a5fb44
Compare
BubbleCal
force-pushed
the
yang/oss-2269-7-row-lineage-main-file
branch
from
September 18, 2026 17:36
d57cbaa to
4186975
Compare
BubbleCal
force-pushed
the
yang/oss-2269-5-row-lineage-update
branch
from
September 19, 2026 07:49
7a5fb44 to
ad708da
Compare
BubbleCal
force-pushed
the
yang/oss-2269-7-row-lineage-main-file
branch
from
September 19, 2026 07:49
4186975 to
32d10f9
Compare
BubbleCal
force-pushed
the
yang/oss-2269-5-row-lineage-update
branch
from
September 22, 2026 04:46
ad708da to
052f595
Compare
BubbleCal
force-pushed
the
yang/oss-2269-7-row-lineage-main-file
branch
from
September 22, 2026 04:46
32d10f9 to
78fe92f
Compare
BubbleCal
added a commit
that referenced
this pull request
Sep 23, 2026
First of a stack for §5.3 of the Stable Row ID GA design (#8931, tracked as #9250): let a fragment's row lineage sequences -- its row ids and its created-at and last-updated-at versions -- leave the manifest and live as hidden `uint64` columns of a Lance data file. This PR defines the format; the stack continues with readers and the spill primitives, then compaction, the commit read-ahead, the update path, and compaction writing the columns into its own output files. Builds on the prototype in #8953 (Will is co-author). ## Stack 1. this PR feat(format): define hidden row lineage columns 2. #9336 feat(dataset): read and write spilled row lineage columns 3. #9337 feat(dataset): spill row lineage at compaction 4. #9338 feat(table): load spilled row lineage ahead of a commit 5. #9339 feat(dataset): spill row lineage when updating rows 6. #9347 feat(dataset): write compaction's spilled lineage into the fragment's data file 7. #9340 test(bench): row lineage spill benchmark ## Problem Each sequence is stored inline in the fragment's manifest entry. An appended fragment's sequences are single runs and cost a few dozen bytes, but once compaction merges fragments whose rows came from many places the row id sequence degrades to 4-8 bytes per row and the version sequences to a run per row. The manifest then grows with the table's row count and every commit rewrites all of it (#8621). ## Format - Every negative field id is reserved for system use and never names a schema field; readers skip any negative id in a data file's `fields`. `-1` and `-2` keep their meaning; `-3`, `-4`, `-5` are the hidden `_rowid`, `_row_created_at_version` and `_row_last_updated_at_version` columns. - `DataFragment` gains an empty `RowLineageColumn` marker arm on each lineage oneof: `column_row_ids = 12`, `column_last_updated_at_versions = 13`, `column_created_at_versions = 14`. The marker carries no file reference: the column lives in one of the fragment's `files`, the single entry whose `fields` carry the reserved id, and its `column_indices` locates it like a user column. Zero or several carriers is corruption. The three sequences may share one file with each other or with the fragment's user data. - The columns have an executable schema: three non-nullable `uint64` fields, each holding exactly `physical_rows` values in physical row order. A null or a length mismatch is corruption and is rejected, never defaulted. - The marker is valid only in a fragment whose data files are Lance v2 files: a legacy v1 `DataFile` has no `column_indices` to locate the column, and a fragment cannot mix v1 and v2 files. A writer on a v1 dataset leaves every sequence inline. - New feature flag `FLAG_UNSTABLE_SPILLED_ROW_LINEAGE = 1 << 11` (value 2048), the next free bit after the two fragment-reuse flags; `FLAG_UNKNOWN` moves to `1 << 12`. Every earlier build has its unknown boundary at or below bit 11 (v11.0.0 at 256, the v12/v13 pre-releases at 512, main before this PR at 2048), so each already refuses such a dataset. Like data overlay files, release builds understand the bit only with `LANCE_ENABLE_UNSTABLE_SPILLED_ROW_LINEAGE=1`; debug builds always do. - The `external_*` arms (field numbers 6, 8, 10) are retired: no writer ever emitted them and the column arms replace that design. Their numbers and names are reserved in the proto, and the `External` variants of `RowIdMeta` and `RowDatasetVersionMeta` go away with them. `ExternalFile` itself stays; the fragment reuse index uses it. Placement rule, which the writers in later PRs follow: a value the commit assigns -- an appended fragment's row ids, an inserted row's created-at, every row's last-updated-at -- can change when a commit conflict is retried, so it stays inline where the retry can rewrite it. A value carried over from existing rows is fixed before the commit and may go to a data file. ## Code - `RowIdMeta::Column` and `RowDatasetVersionMeta::Column` unit variants, with proto and JSON round trips (`{"column": {}}`, mirroring the empty proto message) and manifest interning. `Fragment::row_lineage_file(field_id)` finds the carrier among `files` and reports more than one as corruption. Because the carrier is an ordinary entry of `files`, cleanup, shallow-clone `base_id` rewriting, file listing and validation already cover it; validation skips the negative ids in a file's `fields`. - `apply_feature_flags` sets the flag when any fragment uses a column arm. - Nothing writes the columns yet. Reading one -- in the row id loader, the fragment reader, and the commit-time paths that resolve an update's lineage from existing fragments -- returns `NotSupported` instead of falling back to defaults. - `assign_row_ids` treats a spilled sequence as covering every physical row, as it always does. ## Validation - `cargo test -p lance-table` - `cargo test -p lance --lib -- rowid row_version stable_row optimize::tests dataset_transactions fragment::tests feature_flag update::tests merge_insert::tests dataset_io` - `cargo clippy --all --tests --benches -- -D warnings`, `cargo fmt --all` - `cargo check --manifest-path python/Cargo.toml`, `cargo check --manifest-path java/lance-jni/Cargo.toml` Refs #8931, #9250 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Will Jones <willjones127@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
BubbleCal
force-pushed
the
yang/oss-2269-5-row-lineage-update
branch
2 times, most recently
from
September 23, 2026 17:05
4c075ac to
7442882
Compare
BubbleCal
added a commit
that referenced
this pull request
Sep 24, 2026
Second of the §5.3 stack (#8931, #9250), on top of #9253. The previous PR defined the hidden row lineage columns; this one reads them and adds the primitives that write them. Nothing calls the writer yet; compaction does in the next PR. ## Stack 1. #9253 feat(format): define hidden row lineage columns 2. this PR feat(dataset): read and write spilled row lineage columns 3. #9337 feat(dataset): spill row lineage at compaction 4. #9338 feat(table): load spilled row lineage ahead of a commit 5. #9339 feat(dataset): spill row lineage when updating rows 6. #9347 feat(dataset): write compaction's spilled lineage into the fragment's data file 7. #9340 test(bench): row lineage spill benchmark ## Read - `load_row_id_sequence` gains a `Column` arm that reads the `_rowid` column back through the ordinary file reader, projected by field id, into the same per-fragment cache as inline sequences. - New `load_row_version_sequence(dataset, fragment, RowVersionKind)` loads either version sequence wherever it is stored; a spilled one is cached per fragment and file, an inline one decodes from the manifest bytes as before. - `FileFragment::open` loads the version sequences asynchronously alongside the row ids and hands them to the reader, instead of the reader builder decoding them synchronously and silently falling back to version 1 on any failure. The `NotSupported` guard from the previous PR goes away with it. - `Dataset::validate` checks the length of spilled version sequences too. ## Write primitives `place_row_lineage(dataset, &RowLineage)` encodes a fragment's three sequences and, for each one whose encoding exceeds the table's inline budget, writes it as a column of one new lineage file per fragment. It returns a `PlacedRowLineage`: the three arms to put on the fragment plus the lineage file, which `apply` adds to the fragment's `files` so the marker can be resolved. It is only correct for lineage a commit conflict cannot change, which is what its callers carry over from existing rows. Spilling is opt-in per table: `lance.row_lineage.spill=true`, with `lance.row_lineage.inline_max_bytes` overriding the 200 KiB default. A table that never sets it is unchanged, and a build that does not understand the feature flag never spills either, so it cannot write a dataset it then refuses to open. A legacy v1 dataset never spills regardless of its config, since the format only allows the columns in v2 files. The primitives are public so a writer outside this crate that assembles its own transactions can spill at write time. ## Validation - `cargo test -p lance --lib rowids::` (round trip of all three columns through one shared file, found by field id among the fragment's files; only the sequences over the budget spill; a table that has not opted in never spills; a v1 table never spills) - `cargo test -p lance --lib -- rowid stable_row fragment::tests dataset_transactions` - `cargo clippy --all --tests --benches -- -D warnings`, `cargo fmt --all` - `cargo check --manifest-path python/Cargo.toml`, `cargo check --manifest-path java/lance-jni/Cargo.toml` Refs #8931, #9250 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Will Jones <willjones127@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
BubbleCal
force-pushed
the
yang/oss-2269-5-row-lineage-update
branch
from
September 24, 2026 09:18
7442882 to
4d8ca88
Compare
BubbleCal
added a commit
that referenced
this pull request
Sep 25, 2026
Third of the §5.3 stack (#8931, #9250), on top of #9336. Compaction becomes the first writer of the hidden row lineage columns. ## Stack 1. #9253 feat(format): define hidden row lineage columns 2. #9336 feat(dataset): read and write spilled row lineage columns 3. this PR feat(dataset): spill row lineage at compaction 4. #9338 feat(table): load spilled row lineage ahead of a commit 5. #9339 feat(dataset): spill row lineage when updating rows 6. #9347 feat(dataset): write compaction's spilled lineage into the fragment's data file 7. #9340 test(bench): row lineage spill benchmark ## Change Compaction used to rechunk the row ids and the two version sequences in two separate passes and write each back inline. It now computes all three together, then places each one through `place_row_lineage`: inline when its encoding fits the table's budget, otherwise as a column of one lineage file per output fragment, listed among the fragment's files after its data file. The sixth PR of the stack moves those columns into the data file itself. Compaction's output is retry-stable -- every value is carried over from the input fragments -- so spilling it needs nothing from the commit. Only tables that set `lance.row_lineage.spill=true` are affected; a table that never opts in compacts exactly as before. ## Not in this PR Updating rows of a table whose compaction has spilled still returns `NotSupported` from the commit, which cannot read a data file; the next PR lifts that. With this PR alone, a user who opts in must not update the table until then, which is why the flag stays unstable and the config defaults off. ## Benchmarks `LANCE_ENABLE_UNSTABLE_SPILLED_ROW_LINEAGE=1 cargo bench --bench rowid_spill` (the benchmark lands as the last PR of the stack), 8 fragments x 1,000,000 rows, local NVMe (macOS), one run per arm, measured on the pre-split branch whose compaction and read code is what this stack carries. The two arms differ only in the table config: the inline arm never opts in (today's behavior), the spilled arm sets `lance.row_lineage.spill=true` with the default 200 KiB budget. Byte counts are deterministic; latencies are single samples (cold open averaged over 10, sequence loads over 3), so treat differences under about 2x as noise. Lower is better for every row. **deleted**: 30% of rows deleted, then compacted. The row id sequences still run-encode as range plus bitmap, so this is the workload where spilling is a bad trade on bytes. | Scenario / metric | Baseline (inline) | This PR (spilled) | Benefit | | --- | ---: | ---: | ---: | | manifest size | 5.73 MiB | < 0.01 MiB | >500x smaller | | compaction transaction file | 2.86 MiB | < 0.01 MiB | >500x smaller | | cold dataset open | 1.45 ms | 0.55 ms | 2.6x speedup | | append commit (mean of 5) | 3.65 ms | 1.46 ms | 2.5x speedup | | load one fragment's sequence (cold) | 0.15 ms | 4.88 ms | 33x slowdown | | row id index build (cold) | 0.11 ms | 16.34 ms | 150x slowdown | | take by row id, index built | 0.54 ms | 4.55 ms | 8x slowdown (first touch, see below) | | compaction | 221 ms | 258 ms | 1.2x slowdown | | data files on disk | 74.4 MiB | 89.2 MiB | 1.2x larger | **shuffled**: every row rewritten in random order, then compacted. No run structure survives, which is the workload the design is for. | Scenario / metric | Baseline (inline) | This PR (spilled) | Benefit | | --- | ---: | ---: | ---: | | manifest size | 30.52 MiB | < 0.01 MiB | >3000x smaller | | compaction transaction file | 61.04 MiB | 30.52 MiB | 2x smaller | | cold dataset open | 5.54 ms | 0.20 ms | 28x speedup | | append commit (mean of 5) | 15.73 ms | 1.34 ms | 12x speedup | | load one fragment's sequence (cold) | 2.35 ms | 2.30 ms | 1.0x | | row id index build (cold) | 142 ms | 165 ms | 1.2x slowdown | | take by row id, index built | 2.65 ms | 2.75 ms | 1.0x | | compaction | 305 ms | 296 ms | 1.0x | | data files on disk | 136.0 MiB | 158.1 MiB | 1.2x larger | The `deleted` take row: `FileFragment::open` loads the fragment's row id sequence whenever the table uses stable row ids, whether or not `_rowid` is projected, and the index build reads sequences uncached, so the first take after it pays one spilled-file read per fragment. That eager load predates this stack; making it conditional on the projection is a follow-up. ## Validation - `cargo test -p lance --lib rowids::` (compaction spills all three sequences into one lineage file among the fragment's files and reads them back through the scan and the loaders; cleanup keeps the live lineage file; a cold reopen serves the columns) - `cargo test -p lance --lib -- rowid stable_row optimize::tests cleanup::tests dataset_transactions` - `cargo clippy --all --tests --benches -- -D warnings`, `cargo fmt --all` Refs #8931, #9250 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Will Jones <willjones127@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
BubbleCal
force-pushed
the
yang/oss-2269-5-row-lineage-update
branch
from
September 25, 2026 15:22
4d8ca88 to
281ddfb
Compare
BubbleCal
added a commit
that referenced
this pull request
Sep 28, 2026
Fourth of the §5.3 stack (#8931, #9250), on top of #9337. Updates on a table with spilled lineage keep every row's lineage, the way they do on an inline table. ## Stack 1. #9253 feat(format): define hidden row lineage columns 2. #9336 feat(dataset): read and write spilled row lineage columns 3. #9337 feat(dataset): spill row lineage at compaction 4. this PR feat(table): load spilled row lineage ahead of a commit 5. #9339 feat(dataset): spill row lineage when updating rows 6. #9347 feat(dataset): write compaction's spilled lineage into the fragment's data file 7. #9340 test(bench): row lineage spill benchmark ## Problem Building a manifest is synchronous and has no object store. Two commit-time paths need existing lineage: resolving which rows an update rewrote, to carry each row's created-at version over (`resolve_update_version_metadata`), and overlaying a partial column rewrite's patched offsets onto the fragment's last-updated-at sequence (`refresh_row_latest_update_meta_for_partial_frag_rewrite_cols`). Once a fragment's sequences are spilled, neither can read them, and the previous PRs made them refuse with `NotSupported`. ## Change The commit path in `lance` reads every spilled sequence of the current manifest ahead of each build attempt (`load_spilled_row_lineage`, served from the same caches the readers use) and hands them over in `ManifestBuildConfig::spilled_row_lineage`. The two paths consult that map for a spilled fragment and still refuse if a sequence they need is missing, so a caller of `build_manifest` that skips the read-ahead gets an error rather than defaulted lineage. Loaded per attempt, so a rebase onto a newer manifest sees that manifest's fragments; only `Update` and `DataOverlay` operations pay for it. `UpdateBuilder`, `merge_insert` in both write modes, and externally assembled `Operation::Update`s therefore work on a spilled table unchanged. The lineage they produce for the rewritten rows is inline; a refreshed last-updated-at sequence goes back inline as well. An update-heavy table thus regrows its manifest between compactions and the next compaction spills it again. ## Validation - `cargo test -p lance-table row_version` (spilled source lineage resolves from the config; missing lineage still refuses) - `cargo test -p lance --lib rowids::` (update keeps the rewritten row's id and created-at; merge_insert keeps matched rows' lineage and stamps inserted ones; a partial column rewrite on a spilled fragment stamps only the patched rows) - `cargo test -p lance --lib -- rowid row_version stable_row update::tests merge_insert::tests dataset_transactions` - `cargo clippy --all --tests --benches -- -D warnings`, `cargo fmt --all` Refs #8931, #9250 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
BubbleCal
force-pushed
the
yang/oss-2269-5-row-lineage-update
branch
from
September 28, 2026 15:59
281ddfb to
be440e2
Compare
BubbleCal
force-pushed
the
yang/oss-2269-7-row-lineage-main-file
branch
from
September 29, 2026 05:48
78fe92f to
492b887
Compare
… data file Compaction computes the output fragments' row lineage before writing: the row ids and versions carry over from the inputs and the output file sizes are fixed in advance. When the table's policy spills a sequence type, its values ride along as hidden uint64 columns of the batches being written, under the reserved field ids, and each output fragment's metadata marks them as spilled into its own data file. A compacted fragment on a table that opts in has one file instead of two, and a scan that projects _rowid reads it from the file it already has open. The write path lets the columns through: write_fragments sets the three hidden fields aside before the schema is checked against the dataset's and puts them back on the written schema under their reserved ids, and Schema::validate admits exactly those ids under those names. The field id constants move to lance-core for that. Binary-copy compaction cannot add columns to the files it copies, so it keeps writing a separate lineage file, as does the update path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
read_spilled_column took the projection field from the file schema's top-level field at the lineage column's physical index. That only holds for lineage-only files. Compaction now writes the lineage columns after the user columns, and a column index counts physical columns: one per leaf from 2.1 on, and one per list or struct as well in 2.0. After a struct or a 2.0 list, the lookup picked another lineage field or none at all, so every reader of a spilled sequence (scan, take, the row id index, commit read-ahead, validate, the next compaction) failed on that fragment. Build the projection field from the reserved id instead, as a non-nullable UInt64 with the reserved name, and keep the column index from the data file. RowLineageSpill::schema_fields builds the written fields through the same helper, so the read and write sides share one definition. A lookup by id in the file schema would not work: a lineage-only file stores its fields under ids 0..2. The reader now opens through open_projected_reader. For a wide compaction output it fetches only the lineage column's metadata instead of decoding every user column's, using the fragment reader's threshold. Lineage-only files and 2.0 files still load the full metadata. Either way, the column index is checked against the file's column count before a reader is built on it, so an entry past the file's columns is still reported as a corrupt file naming the data file, as the old schema lookup did, rather than as invalid input from the reader. compaction_spills_and_reads_back_row_lineage becomes an rstest with a struct column on 2.2 and a list column on 2.0. With all three sequences spilled, both cases failed before this change. spilled_lineage_shares_one_file_and_round_trips also checks the error for an out-of-range column index. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
versions::write_fragments runs for every writer: insert, update, both merge_insert paths and compaction. It used to treat any top-level field named _rowid, _row_created_at_version or _row_last_updated_at_version as a hidden lineage column, overwrite its id with the reserved one and write it where readers skip it. A user column with one of those names, from alter_columns or an external caller of write_fragments_internal, would have been dropped from the dataset's view without an error. Matching by name was only needed because the 2.2+ blob promotion's set_field_id renumbers every negative id before the split ran. Split before the promotion instead, and take only fields that already carry the reserved id of their name. RowLineageSpill::schema_fields is the only code that assigns those ids. A field with a reserved id that is not a non-nullable UInt64 is rejected with an internal error, because the readers decode these columns as exactly that type. The unwrap and the second name lookup go away. Schema::validate now exempts reserved ids only for top-level fields, where a data file stores the lineage columns. A nested field that reuses one fails the negative-id or duplicate-id check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Compaction computed its output fragments' row lineage from the planned file row counts before writing, then paired it with the fragments the writer produced and checked only how many there were. A byte limit (max_bytes_per_file) makes the writer close a file early and spread the remaining rows over files of other sizes, so a stable-row-id compaction failed with a fragment count mismatch or, when the count happened to match, placed row ids shifted onto the wrong rows. This held for every stable-row-id table, including ones that never opted in to spilling. Only a table that can spill now plans its lineage before the write. Every other table carries its lineage over after the write, sized by the written physical rows, as it did before lineage could go into the data file. Placement follows the written sizes: when they differ from the plan, the inline sequences are cut again at them, the written total must equal the planned one, and every inline sequence must match its fragment's physical rows. The hidden columns need nothing, since they were appended row by row. The spill plan is now decided per sequence type. A type goes into the data files only when it is over the budget in every output fragment. When it is over in only some, the task writes no hidden columns and places each fragment's lineage on its own after the write, so a Range sequence that fits in a few manifest bytes no longer costs a read per row. The spilled sequences are released as soon as their column values are built, and each batch gets its own copy of those values. A slice shared the whole task's buffer, which the writer counts when sizing pages, so it cut a page for every batch. RowLineageSpill and the plan are no longer public API, and one table of column names and field ids drives the field ids, the write schema and the stream schema. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Compaction writes spilled row lineage into the data file of each fragment it writes, and schema-only commits keep that file because it holds the lineage's only copy. Once every user column in it is dropped, cast or replaced, the file carries nothing but lineage and dead user bytes, and two things went wrong. FileFragment::validate treated any file with a non-negative field id as user data and opened it. After drop_columns the carrier still lists the dropped ids, shares no field with the schema, and validate() reported the table as corrupt. A file with no schema field that carries one of the fragment's spilled sequences is now left unopened and checked through the sequences it carries. Every other file is opened as before, so a stale file of user fields the schema no longer has is still reported. Nothing rewrote the file either: the planner picks fragments by size, deletions and overlays only, so the dead bytes stayed referenced for good. A fragment with a file that holds no schema field, holds a dropped or tombstoned user field, and carries lineage now compacts on its own, which moves the lineage next to the live columns and lets the file go. Neither layout compaction writes matches -- its data files hold live fields, and the lineage-only files of binary copy and update hold no user field -- so the rewrite cannot repeat. Binary copy cannot rewrite such a file, so under ForceBinaryCopy the planner leaves the fragment alone rather than plan a task that fails the whole run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…llow-ups Compaction now writes spilled lineage into the data file of each fragment it writes, but the tests only covered one output fragment with no deleted rows, never compacted such a fragment again, and the cleanup test had quietly stopped covering a lineage-only file. - compaction_spills_and_reads_back_row_lineage adds a task that writes two fragments, so the hidden columns have to break where the writer ends a file, and a compaction of deleted rows. Every case checks that each fragment's sequences hold its physical rows, and takes sampled row ids through the row id index after a cold reopen. Four appends of 250 under a 300-row target plan two single-output tasks, so the two-output case compacts three appends. - compact_twice_reads_back_in_file_lineage compacts a fragment whose lineage is in its data file: the scan must not return the hidden columns the task appends itself. - cleanup_keeps_a_live_spilled_file runs for the in-file layout and for binary copy's lineage-only file, and pins each layout before cleaning up. - schema_change_keeps_spilled_lineage also casts a column after its binary-copy compaction, and points to the test covering a data file that carries the lineage next to user columns. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
BubbleCal
added a commit
that referenced
this pull request
Sep 29, 2026
Part of the §5.3 stack (#8931, #9250). #9253, #9336, #9337 and #9338 are merged; this PR is now based on `main`. Updates on a table that opts into spilling move the lineage of the rows they rewrite out of the manifest, so an update-heavy table no longer regrows its manifest between compactions. It also fixes schema-only commits dropping the file a spilled sequence lives in, which has been reachable since #9337. ## Stack 1. #9253 feat(format): define hidden row lineage columns (merged) 2. #9336 feat(dataset): read and write spilled row lineage columns (merged) 3. #9337 feat(dataset): spill row lineage at compaction (merged) 4. #9338 feat(table): load spilled row lineage ahead of a commit (merged) 5. this PR: feat(dataset): spill row lineage when updating rows 6. #9347 feat(dataset): write compaction's spilled lineage into the fragment's data file 7. #9340 test(bench): row lineage spill benchmark ## Change **Update places the lineage it carries over, only when it spills.** - `UpdateBuilder` resolves the table's inline budget before any IO. A malformed `lance.row_lineage.inline_max_bytes` now fails the update before anything is written. - Only on a table that can spill (opted in, v2 files, a build that understands the flag) does the scan project `_row_created_at_version` next to `_rowid`. `make_rowid_capture_stream` captures it run-length encoded, so a rewrite of N rows written at a few versions costs a few runs rather than 8 bytes per row. - After the write, each new fragment's row ids and created-at versions are split to its size. They leave the manifest only when one of them is over the budget, into a lineage-only file added to the fragment's `files`. Last-updated-at is never placed by the writer; it is the commit's. - A table that never opted in, or an update whose sequences fit inline, leaves created-at to the commit exactly as `main` does. **Commit contract** (`resolve_update_version_metadata`, `has_writer_placed_lineage`): - A new fragment whose row ids or created-at versions are `Column` is writer-placed. It keeps its created-at and gets last-updated-at stamped with the commit version. - Placed inline created-at must have one value per physical row, or the commit is refused. - Every other fragment is resolved from the existing fragments as before. A caller-supplied inline created-at is still recomputed, which keeps the documented contract for fragments committed from Python. - The commit skips the created-at lookup for writer-placed fragments, and the read-ahead from #9338 is skipped for an update whose new fragments are all writer-placed. **Spilled lineage survives schema-only commits.** `Operation::Project` (drop, rename, nullability), `DataReplacement` and the cast in `alter_columns` kept only files with a live schema field, so they dropped the only carrier of a spilled sequence and left a `Column` arm pointing at nothing. After cleanup, those row ids and versions were gone. All three now keep a file that carries one of the fragment's spilled ids. `build_manifest` also refuses to commit a fragment whose `Column` arm has no carrier, or a carrier in a v1 file, so a future path cannot commit that state silently. `merge_insert` is unchanged: its rewritten rows still get inline lineage resolved at commit time. ## Validation Run on an x86 build host (r8i.8xlarge): - `cargo fmt --all -- --check`, `cargo clippy --all --tests --benches -- -D warnings` - `cargo test -p lance-table`: 573 passed, including `project_keeps_file_carrying_spilled_row_lineage`, `data_replacement_keeps_lineage_carrier`, `build_manifest_rejects_spilled_arm_without_carrier` and the new `row_version` contract tests. - `cargo test -p lance --lib -- rowids:: row_version stable_row rowid optimize::tests dataset_transactions fragment::tests cleanup::tests feature_flag update::tests merge_insert::tests io::commit dataset_io schema_evolution versions:: write::tests utils::tests`: 1545 passed, including `schema_change_keeps_spilled_lineage` (update or compaction followed by rename, drop or cast), `update_leaves_inline_created_at_to_the_commit`, `update_rejects_malformed_inline_max_bytes_before_writing` and `place_rewritten_lineage_splits_lineage_by_output_fragment`. - `cargo check --manifest-path python/Cargo.toml`, `cargo check --manifest-path java/lance-jni/Cargo.toml` Refs #8931, #9250 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
BubbleCal
force-pushed
the
yang/oss-2269-7-row-lineage-main-file
branch
from
September 29, 2026 10:01
492b887 to
ea175d5
Compare
BubbleCal
marked this pull request as ready for review
September 29, 2026 10:01
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.
Tip: disable this comment in your organization's Code Review settings.
Contributor
There was a problem hiding this comment.
✅ Gate recommendation: approve.
Writing retry-stable lineage into re-encoded data files removes a separate lineage object for each compacted fragment that spills. The existing separate-file path remains for binary copy, which cannot add columns. The focused readback, re-splitting, deletion, schema-change, repeated-compaction, and cleanup tests passed; I found no blocker.
Xuanwo
approved these changes
Sep 29, 2026
BubbleCal
added a commit
that referenced
this pull request
Sep 29, 2026
…#9340) Last of the §5.3 stack (#8931, #9250). The benchmark for row lineage placement, inline in the manifest versus spilled to data file columns. Everything it measures is on `main`: this PR only adds `rust/lance/benches/rowid_spill.rs` and its `[[bench]]` entry in `rust/lance/Cargo.toml`. ## Stack 1. #9253, #9336, #9337, #9338, #9339, #9347 (merged) 2. this PR: test(bench): measure row lineage placement against a checked workload ## What it measures ```bash LANCE_ENABLE_UNSTABLE_SPILLED_ROW_LINEAGE=1 cargo bench -p lance --profile release-with-debug --bench rowid_spill ``` It is sized with `BENCH_FRAGMENTS`, `BENCH_ROWS_PER_FRAGMENT`, `BENCH_DELETE_PERCENT`, `BENCH_APPENDS`, `BENCH_SCENARIOS` and `BENCH_INLINE_MAX_BYTES`. Bad values are rejected up front. There are two workloads: - `deleted`: a share of rows is deleted, then compacted. The sequences still run-encode. - `shuffled`: every row is rewritten in random order, with its real created-at and last-updated-at versions, then compacted. No run structure survives. Within each workload, the inline arm never opts in, and the spilled arm sets `lance.row_lineage.spill=true` with the given budget. Per arm it reports: - steady-state manifest bytes, split by sequence; - the compaction manifest and transaction file; - cold open; - the commit of a small append, timed apart from the data write; - one cold sequence load; - the row id index build; - a take by row id with the fragment's sequence preloaded; - compaction time and the bytes compaction wrote. Timed rows report the median and the minimum of their samples. The bench checks itself outside the timers: - the spilled arm must actually have spilled, and the inline arm must not; - a full scan asserts every row's `_rowid`, created-at and last-updated-at against the workload; - every take probe must return the row it asked for. ## Smoke run This is one small run, not a performance claim. It ran on an x86 r8i.8xlarge (local NVMe) with 2 fragments x 200,000 rows, `BENCH_INLINE_MAX_BYTES=0` so everything spills, the `release-with-debug` profile, and each arm once. Lower is better for every row, and the ratio is inline / spilled. | Scenario / metric | Inline | Spilled | Ratio | | --- | ---: | ---: | ---: | | shuffled: steady-state manifest | 7.49 MB | < 0.01 MB | 10367x smaller | | shuffled: cold dataset open (median) | 4.32 ms | 0.12 ms | 36.5x faster | | shuffled: append commit (median) | 6.27 ms | 0.78 ms | 8.1x faster | | shuffled: row id index build, cold (median) | 36.11 ms | 35.53 ms | 1.0x | | shuffled: one sequence load, cold (median) | 0.18 ms | 1.97 ms | 0.09x (spilled is slower) | | shuffled: compaction output (data files) | 1.93 MB | 2.96 MB | 0.65x (spilled is larger) | | deleted: steady-state manifest | 0.14 MB | < 0.01 MB | 237x smaller | | deleted: row id index build, cold (median) | 2.53 ms | 5.48 ms | 0.46x (spilled is slower) | The trade-off is the one the design expects. Spilling takes the per-row lineage out of every manifest read and write, so open and commit stop scaling with the table. The cost moves to a column read the first time a sequence is needed, and 2 to 3 bytes per row in the compacted data files. On the `deleted` workload the sequences are small either way, which is why the default 200 KiB budget keeps them inline. ## Validation - `cargo fmt --all -- --check`, `cargo clippy --all --tests --benches -- -D warnings` on the head of this branch - The smoke run above completed with every self-check passing. Refs #8931, #9250 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Will Jones <willjones127@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of the §5.3 stack (#8931, #9250); based on
mainnow that #9339 has merged. Compaction writes the lineage sequences it spills into each output fragment's own data file, as hidden columns next to the user columns, instead of a separate lineage file. It also fixes three bugs in reading, placing and reclaiming spilled lineage that the review of this stack found.Stack
Change
Lineage in the data file.
uint64column of the batches being written. The fragment's metadata then marks it as spilled into that same data file.RowLineagePlan::PerFragment) writes no hidden columns. After the write, each fragment is placed on its own, so a cheapRangesequence never turns into a column read.versions::write_fragmentssets the hidden fields aside by their reserved ids before the schema check and puts them back on the written schema.Schema::validateadmits the reserved ids only on top-level fields with those names.Fixes.
max_bytes_per_file) can close a file early and re-split the rows, so the plan's file sizes are not the written ones. Placement used to zip the plan with the written fragments. That made a stable-row-id compaction fail outright, or shift row ids by one when the fragment count happened to match. It now rechunks the planned lineage to each fragment'sphysical_rowsand checks every inline sequence's length.read_spilled_columntook the projected field from the file schema by column index. Once lineage columns follow nested or list user columns, the column index is no longer the top-level position, so the read errored or took the wrong field. The field is now built from the reserved id. Only the needed column's metadata is opened.FileFragment::validateno longer tries to open it as user data. The compaction planner picks such a fragment up, so the lineage moves into a fresh file and the dead bytes are reclaimed. Lineage-only files never match, so this cannot loop.Memory. Spilled sequences are moved into their column values rather than copied. Each batch gets its own buffer, so the encoder no longer cuts a page per batch from a shared slice. Pre-write planning only happens on tables that can spill.
Validation
Run on an x86 build host (r8i.8xlarge):
cargo fmt --all -- --check,cargo clippy --all --tests --benches -- -D warningscargo test -p lance-table: 573 passed.cargo test -p lance --lib -- rowids:: row_version stable_row rowid optimize::tests dataset_transactions fragment::tests cleanup::tests feature_flag update::tests merge_insert::tests io::commit dataset_io schema_evolution versions:: write::tests utils::tests: 1576 passed.compaction_spills_and_reads_back_row_lineagecases:flat,nested_v2_2,list_v2_0,multi_output,with_deletions.compact_twice_reads_back_in_file_lineage.place_row_lineage_in_files_follows_written_sizes: extra file, shifted split, mismatched totals.plan_row_lineage_spill_decides_per_kindandper_fragment_plan_spills_only_the_fragments_over_budget.append_row_lineage_columns_gives_each_batch_its_own_buffer.dropping_every_column_of_a_lineage_carrier_keeps_lineage_until_compaction: drop, cast, replace.compaction_reclaims_lineage_carriers_of_dead_user_columns.cleanup_keeps_a_live_spilled_file: in-file and lineage-file layouts.split_row_lineage_fields_keys_on_reserved_ids.cargo check --manifest-path python/Cargo.toml,cargo check --manifest-path java/lance-jni/Cargo.tomlRefs #8931, #9250
🤖 Generated with Claude Code