Skip to content

feat(dataset): spill row lineage when updating rows - #9339

Merged
BubbleCal merged 5 commits into
mainfrom
yang/oss-2269-5-row-lineage-update
Sep 29, 2026
Merged

BubbleCal merged 5 commits into
mainfrom
yang/oss-2269-5-row-lineage-update

Conversation

@BubbleCal

@BubbleCal BubbleCal commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Part of the §5.3 stack (#8931, #9250). #9253, #9336, #9337 and #9338 are merged; this PR is now based on main. Updates on a table that opts into spilling move the lineage of the rows they rewrite out of the manifest, so an update-heavy table no longer regrows its manifest between compactions. It also fixes schema-only commits dropping the file a spilled sequence lives in, which has been reachable since #9337.

Stack

  1. feat(format): define hidden row lineage columns #9253 feat(format): define hidden row lineage columns (merged)
  2. feat(dataset): read and write spilled row lineage columns #9336 feat(dataset): read and write spilled row lineage columns (merged)
  3. feat(dataset): spill row lineage at compaction #9337 feat(dataset): spill row lineage at compaction (merged)
  4. feat(table): load spilled row lineage ahead of a commit #9338 feat(table): load spilled row lineage ahead of a commit (merged)
  5. this PR: feat(dataset): spill row lineage when updating rows
  6. feat(dataset): write compaction's spilled lineage into the fragment's data file #9347 feat(dataset): write compaction's spilled lineage into the fragment's data file
  7. test(bench): measure row lineage placement against a checked workload #9340 test(bench): row lineage spill benchmark

Change

Update places the lineage it carries over, only when it spills.

  • UpdateBuilder resolves the table's inline budget before any IO. A malformed lance.row_lineage.inline_max_bytes now fails the update before anything is written.
  • Only on a table that can spill (opted in, v2 files, a build that understands the flag) does the scan project _row_created_at_version next to _rowid. make_rowid_capture_stream captures it run-length encoded, so a rewrite of N rows written at a few versions costs a few runs rather than 8 bytes per row.
  • After the write, each new fragment's row ids and created-at versions are split to its size. They leave the manifest only when one of them is over the budget, into a lineage-only file added to the fragment's files. Last-updated-at is never placed by the writer; it is the commit's.
  • A table that never opted in, or an update whose sequences fit inline, leaves created-at to the commit exactly as main does.

Commit contract (resolve_update_version_metadata, has_writer_placed_lineage):

  • A new fragment whose row ids or created-at versions are Column is writer-placed. It keeps its created-at and gets last-updated-at stamped with the commit version.
  • Placed inline created-at must have one value per physical row, or the commit is refused.
  • Every other fragment is resolved from the existing fragments as before. A caller-supplied inline created-at is still recomputed, which keeps the documented contract for fragments committed from Python.
  • The commit skips the created-at lookup for writer-placed fragments, and the read-ahead from feat(table): load spilled row lineage ahead of a commit #9338 is skipped for an update whose new fragments are all writer-placed.

Spilled lineage survives schema-only commits. Operation::Project (drop, rename, nullability), DataReplacement and the cast in alter_columns kept only files with a live schema field, so they dropped the only carrier of a spilled sequence and left a Column arm pointing at nothing. After cleanup, those row ids and versions were gone. All three now keep a file that carries one of the fragment's spilled ids. build_manifest also refuses to commit a fragment whose Column arm has no carrier, or a carrier in a v1 file, so a future path cannot commit that state silently.

merge_insert is unchanged: its rewritten rows still get inline lineage resolved at commit time.

Validation

Run on an x86 build host (r8i.8xlarge):

  • cargo fmt --all -- --check, cargo clippy --all --tests --benches -- -D warnings
  • cargo test -p lance-table: 573 passed, including project_keeps_file_carrying_spilled_row_lineage, data_replacement_keeps_lineage_carrier, build_manifest_rejects_spilled_arm_without_carrier and the new row_version contract tests.
  • cargo test -p lance --lib -- rowids:: row_version stable_row rowid optimize::tests dataset_transactions fragment::tests cleanup::tests feature_flag update::tests merge_insert::tests io::commit dataset_io schema_evolution versions:: write::tests utils::tests: 1545 passed, including schema_change_keeps_spilled_lineage (update or compaction followed by rename, drop or cast), update_leaves_inline_created_at_to_the_commit, update_rejects_malformed_inline_max_bytes_before_writing and place_rewritten_lineage_splits_lineage_by_output_fragment.
  • cargo check --manifest-path python/Cargo.toml, cargo check --manifest-path java/lance-jni/Cargo.toml

Refs #8931, #9250

🤖 Generated with Claude Code

@github-actions github-actions Bot added the enhancement New feature or request label Sep 17, 2026
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from 5f4feb7 to 4601068 Compare September 17, 2026 08:39
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from 4601068 to f54507f Compare September 17, 2026 08:47
@BubbleCal
BubbleCal added this pull request to stack #9343 September 17, 2026 08:56
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from f54507f to fff8cd3 Compare September 17, 2026 10:03
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from fff8cd3 to cd2243c Compare September 17, 2026 10:52
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from cd2243c to be46048 Compare September 17, 2026 13:39
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from be46048 to 7a5fb44 Compare September 18, 2026 17:36
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from 7a5fb44 to ad708da Compare September 19, 2026 07:49
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from ad708da to 052f595 Compare September 22, 2026 04:46
BubbleCal added a commit that referenced this pull request Sep 23, 2026
First of a stack for §5.3 of the Stable Row ID GA design (#8931, tracked
as #9250): let a fragment's row lineage sequences -- its row ids and its
created-at and last-updated-at versions -- leave the manifest and live
as hidden `uint64` columns of a Lance data file. This PR defines the
format; the stack continues with readers and the spill primitives, then
compaction, the commit read-ahead, the update path, and compaction
writing the columns into its own output files.

Builds on the prototype in #8953 (Will is co-author).

## Stack

1. this PR feat(format): define hidden row lineage columns
2. #9336 feat(dataset): read and write spilled row lineage columns
3. #9337 feat(dataset): spill row lineage at compaction
4. #9338 feat(table): load spilled row lineage ahead of a commit
5. #9339 feat(dataset): spill row lineage when updating rows
6. #9347 feat(dataset): write compaction's spilled lineage into the
fragment's data file
7. #9340 test(bench): row lineage spill benchmark

## Problem

Each sequence is stored inline in the fragment's manifest entry. An
appended fragment's sequences are single runs and cost a few dozen
bytes, but once compaction merges fragments whose rows came from many
places the row id sequence degrades to 4-8 bytes per row and the version
sequences to a run per row. The manifest then grows with the table's row
count and every commit rewrites all of it (#8621).

## Format

- Every negative field id is reserved for system use and never names a
schema field; readers skip any negative id in a data file's `fields`.
`-1` and `-2` keep their meaning; `-3`, `-4`, `-5` are the hidden
`_rowid`, `_row_created_at_version` and `_row_last_updated_at_version`
columns.
- `DataFragment` gains an empty `RowLineageColumn` marker arm on each
lineage oneof: `column_row_ids = 12`, `column_last_updated_at_versions =
13`, `column_created_at_versions = 14`. The marker carries no file
reference: the column lives in one of the fragment's `files`, the single
entry whose `fields` carry the reserved id, and its `column_indices`
locates it like a user column. Zero or several carriers is corruption.
The three sequences may share one file with each other or with the
fragment's user data.
- The columns have an executable schema: three non-nullable `uint64`
fields, each holding exactly `physical_rows` values in physical row
order. A null or a length mismatch is corruption and is rejected, never
defaulted.
- The marker is valid only in a fragment whose data files are Lance v2
files: a legacy v1 `DataFile` has no `column_indices` to locate the
column, and a fragment cannot mix v1 and v2 files. A writer on a v1
dataset leaves every sequence inline.
- New feature flag `FLAG_UNSTABLE_SPILLED_ROW_LINEAGE = 1 << 11` (value
2048), the next free bit after the two fragment-reuse flags;
`FLAG_UNKNOWN` moves to `1 << 12`. Every earlier build has its unknown
boundary at or below bit 11 (v11.0.0 at 256, the v12/v13 pre-releases at
512, main before this PR at 2048), so each already refuses such a
dataset. Like data overlay files, release builds understand the bit only
with `LANCE_ENABLE_UNSTABLE_SPILLED_ROW_LINEAGE=1`; debug builds always
do.
- The `external_*` arms (field numbers 6, 8, 10) are retired: no writer
ever emitted them and the column arms replace that design. Their numbers
and names are reserved in the proto, and the `External` variants of
`RowIdMeta` and `RowDatasetVersionMeta` go away with them.
`ExternalFile` itself stays; the fragment reuse index uses it.

Placement rule, which the writers in later PRs follow: a value the
commit assigns -- an appended fragment's row ids, an inserted row's
created-at, every row's last-updated-at -- can change when a commit
conflict is retried, so it stays inline where the retry can rewrite it.
A value carried over from existing rows is fixed before the commit and
may go to a data file.

## Code

- `RowIdMeta::Column` and `RowDatasetVersionMeta::Column` unit variants,
with proto and JSON round trips (`{"column": {}}`, mirroring the empty
proto message) and manifest interning.
`Fragment::row_lineage_file(field_id)` finds the carrier among `files`
and reports more than one as corruption. Because the carrier is an
ordinary entry of `files`, cleanup, shallow-clone `base_id` rewriting,
file listing and validation already cover it; validation skips the
negative ids in a file's `fields`.
- `apply_feature_flags` sets the flag when any fragment uses a column
arm.
- Nothing writes the columns yet. Reading one -- in the row id loader,
the fragment reader, and the commit-time paths that resolve an update's
lineage from existing fragments -- returns `NotSupported` instead of
falling back to defaults.
- `assign_row_ids` treats a spilled sequence as covering every physical
row, as it always does.

## Validation

- `cargo test -p lance-table`
- `cargo test -p lance --lib -- rowid row_version stable_row
optimize::tests dataset_transactions fragment::tests feature_flag
update::tests merge_insert::tests dataset_io`
- `cargo clippy --all --tests --benches -- -D warnings`, `cargo fmt
--all`
- `cargo check --manifest-path python/Cargo.toml`, `cargo check
--manifest-path java/lance-jni/Cargo.toml`

Refs #8931, #9250

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Will Jones <willjones127@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch 2 times, most recently from 4c075ac to 7442882 Compare September 23, 2026 17:05
BubbleCal added a commit that referenced this pull request Sep 24, 2026
Second of the §5.3 stack (#8931, #9250), on top of #9253. The previous
PR defined the hidden row lineage columns; this one reads them and adds
the primitives that write them. Nothing calls the writer yet; compaction
does in the next PR.

## Stack

1. #9253 feat(format): define hidden row lineage columns
2. this PR feat(dataset): read and write spilled row lineage columns
3. #9337 feat(dataset): spill row lineage at compaction
4. #9338 feat(table): load spilled row lineage ahead of a commit
5. #9339 feat(dataset): spill row lineage when updating rows
6. #9347 feat(dataset): write compaction's spilled lineage into the
fragment's data file
7. #9340 test(bench): row lineage spill benchmark

## Read

- `load_row_id_sequence` gains a `Column` arm that reads the `_rowid`
column back through the ordinary file reader, projected by field id,
into the same per-fragment cache as inline sequences.
- New `load_row_version_sequence(dataset, fragment, RowVersionKind)`
loads either version sequence wherever it is stored; a spilled one is
cached per fragment and file, an inline one decodes from the manifest
bytes as before.
- `FileFragment::open` loads the version sequences asynchronously
alongside the row ids and hands them to the reader, instead of the
reader builder decoding them synchronously and silently falling back to
version 1 on any failure. The `NotSupported` guard from the previous PR
goes away with it.
- `Dataset::validate` checks the length of spilled version sequences
too.

## Write primitives

`place_row_lineage(dataset, &RowLineage)` encodes a fragment's three
sequences and, for each one whose encoding exceeds the table's inline
budget, writes it as a column of one new lineage file per fragment. It
returns a `PlacedRowLineage`: the three arms to put on the fragment plus
the lineage file, which `apply` adds to the fragment's `files` so the
marker can be resolved. It is only correct for lineage a commit conflict
cannot change, which is what its callers carry over from existing rows.

Spilling is opt-in per table: `lance.row_lineage.spill=true`, with
`lance.row_lineage.inline_max_bytes` overriding the 200 KiB default. A
table that never sets it is unchanged, and a build that does not
understand the feature flag never spills either, so it cannot write a
dataset it then refuses to open. A legacy v1 dataset never spills
regardless of its config, since the format only allows the columns in v2
files.

The primitives are public so a writer outside this crate that assembles
its own transactions can spill at write time.

## Validation

- `cargo test -p lance --lib rowids::` (round trip of all three columns
through one shared file, found by field id among the fragment's files;
only the sequences over the budget spill; a table that has not opted in
never spills; a v1 table never spills)
- `cargo test -p lance --lib -- rowid stable_row fragment::tests
dataset_transactions`
- `cargo clippy --all --tests --benches -- -D warnings`, `cargo fmt
--all`
- `cargo check --manifest-path python/Cargo.toml`, `cargo check
--manifest-path java/lance-jni/Cargo.toml`

Refs #8931, #9250

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Will Jones <willjones127@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from 7442882 to 4d8ca88 Compare September 24, 2026 09:18
BubbleCal added a commit that referenced this pull request Sep 25, 2026
Third of the §5.3 stack (#8931, #9250), on top of #9336. Compaction
becomes the first writer of the hidden row lineage columns.

## Stack

1. #9253 feat(format): define hidden row lineage columns
2. #9336 feat(dataset): read and write spilled row lineage columns
3. this PR feat(dataset): spill row lineage at compaction
4. #9338 feat(table): load spilled row lineage ahead of a commit
5. #9339 feat(dataset): spill row lineage when updating rows
6. #9347 feat(dataset): write compaction's spilled lineage into the
fragment's data file
7. #9340 test(bench): row lineage spill benchmark

## Change

Compaction used to rechunk the row ids and the two version sequences in
two separate passes and write each back inline. It now computes all
three together, then places each one through `place_row_lineage`: inline
when its encoding fits the table's budget, otherwise as a column of one
lineage file per output fragment, listed among the fragment's files
after its data file. The sixth PR of the stack moves those columns into
the data file itself. Compaction's output is retry-stable -- every value
is carried over from the input fragments -- so spilling it needs nothing
from the commit.

Only tables that set `lance.row_lineage.spill=true` are affected; a
table that never opts in compacts exactly as before.

## Not in this PR

Updating rows of a table whose compaction has spilled still returns
`NotSupported` from the commit, which cannot read a data file; the next
PR lifts that. With this PR alone, a user who opts in must not update
the table until then, which is why the flag stays unstable and the
config defaults off.

## Benchmarks

`LANCE_ENABLE_UNSTABLE_SPILLED_ROW_LINEAGE=1 cargo bench --bench
rowid_spill` (the benchmark lands as the last PR of the stack), 8
fragments x 1,000,000 rows, local NVMe (macOS), one run per arm,
measured on the pre-split branch whose compaction and read code is what
this stack carries. The two arms differ only in the table config: the
inline arm never opts in (today's behavior), the spilled arm sets
`lance.row_lineage.spill=true` with the default 200 KiB budget. Byte
counts are deterministic; latencies are single samples (cold open
averaged over 10, sequence loads over 3), so treat differences under
about 2x as noise. Lower is better for every row.

**deleted**: 30% of rows deleted, then compacted. The row id sequences
still run-encode as range plus bitmap, so this is the workload where
spilling is a bad trade on bytes.

| Scenario / metric | Baseline (inline) | This PR (spilled) | Benefit |
| --- | ---: | ---: | ---: |
| manifest size | 5.73 MiB | < 0.01 MiB | >500x smaller |
| compaction transaction file | 2.86 MiB | < 0.01 MiB | >500x smaller |
| cold dataset open | 1.45 ms | 0.55 ms | 2.6x speedup |
| append commit (mean of 5) | 3.65 ms | 1.46 ms | 2.5x speedup |
| load one fragment's sequence (cold) | 0.15 ms | 4.88 ms | 33x slowdown
|
| row id index build (cold) | 0.11 ms | 16.34 ms | 150x slowdown |
| take by row id, index built | 0.54 ms | 4.55 ms | 8x slowdown (first
touch, see below) |
| compaction | 221 ms | 258 ms | 1.2x slowdown |
| data files on disk | 74.4 MiB | 89.2 MiB | 1.2x larger |

**shuffled**: every row rewritten in random order, then compacted. No
run structure survives, which is the workload the design is for.

| Scenario / metric | Baseline (inline) | This PR (spilled) | Benefit |
| --- | ---: | ---: | ---: |
| manifest size | 30.52 MiB | < 0.01 MiB | >3000x smaller |
| compaction transaction file | 61.04 MiB | 30.52 MiB | 2x smaller |
| cold dataset open | 5.54 ms | 0.20 ms | 28x speedup |
| append commit (mean of 5) | 15.73 ms | 1.34 ms | 12x speedup |
| load one fragment's sequence (cold) | 2.35 ms | 2.30 ms | 1.0x |
| row id index build (cold) | 142 ms | 165 ms | 1.2x slowdown |
| take by row id, index built | 2.65 ms | 2.75 ms | 1.0x |
| compaction | 305 ms | 296 ms | 1.0x |
| data files on disk | 136.0 MiB | 158.1 MiB | 1.2x larger |

The `deleted` take row: `FileFragment::open` loads the fragment's row id
sequence whenever the table uses stable row ids, whether or not `_rowid`
is projected, and the index build reads sequences uncached, so the first
take after it pays one spilled-file read per fragment. That eager load
predates this stack; making it conditional on the projection is a
follow-up.


## Validation

- `cargo test -p lance --lib rowids::` (compaction spills all three
sequences into one lineage file among the fragment's files and reads
them back through the scan and the loaders; cleanup keeps the live
lineage file; a cold reopen serves the columns)
- `cargo test -p lance --lib -- rowid stable_row optimize::tests
cleanup::tests dataset_transactions`
- `cargo clippy --all --tests --benches -- -D warnings`, `cargo fmt
--all`

Refs #8931, #9250

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Will Jones <willjones127@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from 4d8ca88 to 281ddfb Compare September 25, 2026 15:22
Base automatically changed from yang/oss-2269-4-row-lineage-commit to main September 28, 2026 15:59
UpdateBuilder reads the rewritten rows' created-at versions alongside
their row ids while it scans them, and places each new fragment's lineage
through place_row_lineage: inline on a table that has not opted in, as
before, and spilled on one that has, so an update-heavy table no longer
regrows its manifest between compactions. The values it places are the
ones a commit conflict cannot change; the last-updated-at version is the
commit's, and the commit now stamps it on any new fragment whose
created-at versions the writer placed, spilled or inline, instead of
resolving them again from the existing fragments.

Legacy V1 datasets, whose reader does not serve the version columns, keep
leaving the created-at lookup to the commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
BubbleCal added a commit that referenced this pull request Sep 28, 2026
Fourth of the §5.3 stack (#8931, #9250), on top of #9337. Updates on a
table with spilled lineage keep every row's lineage, the way they do on
an inline table.

## Stack

1. #9253 feat(format): define hidden row lineage columns
2. #9336 feat(dataset): read and write spilled row lineage columns
3. #9337 feat(dataset): spill row lineage at compaction
4. this PR feat(table): load spilled row lineage ahead of a commit
5. #9339 feat(dataset): spill row lineage when updating rows
6. #9347 feat(dataset): write compaction's spilled lineage into the
fragment's data file
7. #9340 test(bench): row lineage spill benchmark

## Problem

Building a manifest is synchronous and has no object store. Two
commit-time paths need existing lineage: resolving which rows an update
rewrote, to carry each row's created-at version over
(`resolve_update_version_metadata`), and overlaying a partial column
rewrite's patched offsets onto the fragment's last-updated-at sequence
(`refresh_row_latest_update_meta_for_partial_frag_rewrite_cols`). Once a
fragment's sequences are spilled, neither can read them, and the
previous PRs made them refuse with `NotSupported`.

## Change

The commit path in `lance` reads every spilled sequence of the current
manifest ahead of each build attempt (`load_spilled_row_lineage`, served
from the same caches the readers use) and hands them over in
`ManifestBuildConfig::spilled_row_lineage`. The two paths consult that
map for a spilled fragment and still refuse if a sequence they need is
missing, so a caller of `build_manifest` that skips the read-ahead gets
an error rather than defaulted lineage. Loaded per attempt, so a rebase
onto a newer manifest sees that manifest's fragments; only `Update` and
`DataOverlay` operations pay for it.

`UpdateBuilder`, `merge_insert` in both write modes, and externally
assembled `Operation::Update`s therefore work on a spilled table
unchanged. The lineage they produce for the rewritten rows is inline; a
refreshed last-updated-at sequence goes back inline as well. An
update-heavy table thus regrows its manifest between compactions and the
next compaction spills it again.

## Validation

- `cargo test -p lance-table row_version` (spilled source lineage
resolves from the config; missing lineage still refuses)
- `cargo test -p lance --lib rowids::` (update keeps the rewritten row's
id and created-at; merge_insert keeps matched rows' lineage and stamps
inserted ones; a partial column rewrite on a spilled fragment stamps
only the patched rows)
- `cargo test -p lance --lib -- rowid row_version stable_row
update::tests merge_insert::tests dataset_transactions`
- `cargo clippy --all --tests --benches -- -D warnings`, `cargo fmt
--all`

Refs #8931, #9250

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
@BubbleCal
BubbleCal force-pushed the yang/oss-2269-5-row-lineage-update branch from 281ddfb to be440e2 Compare September 28, 2026 15:59
BubbleCal and others added 4 commits September 29, 2026 13:36
A spilled row lineage sequence lives only in the data file whose fields
carry its reserved id (-3, -4 or -5), and those ids are never in the
dataset schema. Three commit paths decide whether a data file is still
live by asking whether any of its fields is in the schema: Project
(drop_columns, and alter_columns renames and nullability changes),
DataReplacement, and the alter_columns cast. Each of them dropped the
carrier while the fragment kept its Column marker. After an update on an
opted-in table, a single rename or drop left a manifest whose next
row-id read failed, and cleanup would then delete the only copy of the
rows' ids and created-at versions.

Add Fragment::spilled_row_lineage_field_ids and have the three retain
sites also keep a file that carries one of the fragment's spilled ids.
Add Fragment::validate_row_lineage_carriers and call it on every fragment
in build_manifest before the manifest is assembled. It requires each
Column arm to have exactly one v2 carrier, so any future path that loses
a carrier fails the commit instead of publishing lineage no reader can
load. The check reads metadata only and skips fragments that spill
nothing. It covers every fragment, not only the ones the operation
touched, so it also refuses to build on a manifest that already lost a
carrier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The commit resolved every Update new fragment's created-at versions from
the existing fragments before noticing that a writer had already placed
them, and then discarded the result: it collected the rewritten row ids,
walked the row ids of every overlapping existing fragment and decoded
the source fragments' created-at sequences. It also read every spilled
lineage sequence of the table ahead of each attempt, whether or not the
build would look at it.

Writer-placed lineage is now defined narrowly: spilled row ids or
spilled created-at versions. Those are the only values the commit can
neither resolve nor rewrite, and only a writer carrying rows' lineage
over produces them. Such fragments skip the lookup and only have their
last-updated-at versions stamped; inline created-at versions next to
spilled row ids must cover every physical row. Inline created-at
versions next to inline row ids are resolved again, as before this
stack, so a hand-built transaction cannot commit stale ones and the
documented Python contract still holds.

The commit skips the spilled-lineage read-ahead for an update whose new
fragments all carry writer-placed lineage and that patched no offsets
in place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UpdateBuilder read every rewritten row's created-at version on any v2
table with stable row ids, held them as a Vec<u64> for the whole write,
and placed them on every new fragment, inline or spilled, next to a
placeholder last-updated-at sequence that could itself spill into the
lineage file. Tables that never opted in paid for a created-at column
on every scanned row and a larger transaction, and a malformed
lance.row_lineage.inline_max_bytes was only reported after the whole
rewrite had been written.

The update now resolves the inline budget once, before any IO, and
reads created-at versions only when the table may spill. A crate-private
place_carried_row_lineage places what a rewritten fragment carries
over: when neither the row ids nor the created-at versions exceed the
budget, only the row ids are placed, inline, and the commit resolves
created-at as it always has; once either spills, both are placed and
only the over-budget ones go to the lineage file. Last-updated-at is
never placed, since the commit stamps it. The public place_row_lineage
is unchanged.

Created-at versions are captured run-length encoded, extending a run
across batch boundaries, and split per output fragment with
rechunk_version_sequences. A capture that does not match the rows
written is now an internal error instead of a silent fallback to the
commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The two update tests were copies that differed only in the opt-in and a
few layout checks, and neither could tell carried-over lineage from a
spilled last-updated-at placeholder: the opted-in one only checked that
the lineage file's fields were negative. They are now one rstest that
pins the lineage file to exactly the row id and created-at columns,
checks the feature flags this update is the first to raise, and reads
the lineage back from a cold open.

The spilled-source update test drew its one rewritten row from the
middle of a created-at run and still described commit-time resolution,
which UpdateBuilder no longer relies on once its lineage spills. It now
deletes rows first and rewrites a selection across the compacted
fragment's run boundaries, checking every row, once with the writer
placing the lineage and once under a budget that leaves the created-at
lookup to the commit and its read-ahead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@BubbleCal
BubbleCal marked this pull request as ready for review September 29, 2026 06:04

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.

Tip: disable this comment in your organization's Code Review settings.

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

Keeping retry-stable row IDs and created-at versions with rewritten rows, while assigning last-updated-at at commit, addresses manifest growth without changing the non-spilling path. The focused spill, update, and schema-change tests support that boundary; I found no blocking issue in the current writer path.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 29, 2026
@BubbleCal
BubbleCal merged commit eedd50c into main Sep 29, 2026
41 checks passed
@BubbleCal
BubbleCal deleted the yang/oss-2269-5-row-lineage-update branch September 29, 2026 07:14
BubbleCal added a commit that referenced this pull request Sep 29, 2026
… data file (#9347)

Part of the §5.3 stack (#8931, #9250); based on `main` now that #9339
has merged. Compaction writes the lineage sequences it spills into each
output fragment's own data file, as hidden columns next to the user
columns, instead of a separate lineage file. It also fixes three bugs in
reading, placing and reclaiming spilled lineage that the review of this
stack found.

## Stack

1. #9253, #9336, #9337, #9338, #9339 (merged)
2. this PR: feat(dataset): write compaction's spilled lineage into the
fragment's data file
3. #9340 test(bench): row lineage spill benchmark (lands after this PR)

## Change

**Lineage in the data file.**
- On a table that can spill, compaction computes its output fragments'
lineage before writing. The row ids and versions carry over from the
inputs.
- A sequence type that is over the inline budget in every output
fragment rides along as an extra non-nullable `uint64` column of the
batches being written. The fragment's metadata then marks it as spilled
into that same data file.
- A type that is over budget in some outputs and not others
(`RowLineagePlan::PerFragment`) writes no hidden columns. After the
write, each fragment is placed on its own, so a cheap `Range` sequence
never turns into a column read.
- Tables that never opted in compact exactly as before: lineage is
rechunked after the write from the written sizes.
- `versions::write_fragments` sets the hidden fields aside by their
reserved ids before the schema check and puts them back on the written
schema. `Schema::validate` admits the reserved ids only on top-level
fields with those names.
- Binary-copy compaction cannot add columns and still writes a separate
lineage file.

**Fixes.**
- **Written sizes, not planned sizes.** A byte limit
(`max_bytes_per_file`) can close a file early and re-split the rows, so
the plan's file sizes are not the written ones. Placement used to zip
the plan with the written fragments. That made a stable-row-id
compaction fail outright, or shift row ids by one when the fragment
count happened to match. It now rechunks the planned lineage to each
fragment's `physical_rows` and checks every inline sequence's length.
- **Reading a spilled column.** `read_spilled_column` took the projected
field from the file schema by column index. Once lineage columns follow
nested or list user columns, the column index is no longer the top-level
position, so the read errored or took the wrong field. The field is now
built from the reserved id. Only the needed column's metadata is opened.
- **Carriers whose user columns are gone.** Now that #9339 keeps
carriers through drop, cast and replace, a data file can hold lineage
and only dead user columns. `FileFragment::validate` no longer tries to
open it as user data. The compaction planner picks such a fragment up,
so the lineage moves into a fresh file and the dead bytes are reclaimed.
Lineage-only files never match, so this cannot loop.

**Memory.** Spilled sequences are moved into their column values rather
than copied. Each batch gets its own buffer, so the encoder no longer
cuts a page per batch from a shared slice. Pre-write planning only
happens on tables that can spill.

## Validation

Run on an x86 build host (r8i.8xlarge):
- `cargo fmt --all -- --check`, `cargo clippy --all --tests --benches --
-D warnings`
- `cargo test -p lance-table`: 573 passed.
- `cargo test -p lance --lib -- rowids:: row_version stable_row rowid
optimize::tests dataset_transactions fragment::tests cleanup::tests
feature_flag update::tests merge_insert::tests io::commit dataset_io
schema_evolution versions:: write::tests utils::tests`: 1576 passed.
- New tests:
- `compaction_spills_and_reads_back_row_lineage` cases: `flat`,
`nested_v2_2`, `list_v2_0`, `multi_output`, `with_deletions`.
  - `compact_twice_reads_back_in_file_lineage`.
- `place_row_lineage_in_files_follows_written_sizes`: extra file,
shifted split, mismatched totals.
- `plan_row_lineage_spill_decides_per_kind` and
`per_fragment_plan_spills_only_the_fragments_over_budget`.
  - `append_row_lineage_columns_gives_each_batch_its_own_buffer`.
-
`dropping_every_column_of_a_lineage_carrier_keeps_lineage_until_compaction`:
drop, cast, replace.
  - `compaction_reclaims_lineage_carriers_of_dead_user_columns`.
- `cleanup_keeps_a_live_spilled_file`: in-file and lineage-file layouts.
  - `split_row_lineage_fields_keys_on_reserved_ids`.
- `cargo check --manifest-path python/Cargo.toml`, `cargo check
--manifest-path java/lance-jni/Cargo.toml`

Refs #8931, #9250

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
BubbleCal added a commit that referenced this pull request Sep 29, 2026
…#9340)

Last of the §5.3 stack (#8931, #9250). The benchmark for row lineage
placement, inline in the manifest versus spilled to data file columns.

Everything it measures is on `main`: this PR only adds
`rust/lance/benches/rowid_spill.rs` and its `[[bench]]` entry in
`rust/lance/Cargo.toml`.

## Stack

1. #9253, #9336, #9337, #9338, #9339, #9347 (merged)
2. this PR: test(bench): measure row lineage placement against a checked
workload

## What it measures

```bash
LANCE_ENABLE_UNSTABLE_SPILLED_ROW_LINEAGE=1 cargo bench -p lance --profile release-with-debug --bench rowid_spill
```

It is sized with `BENCH_FRAGMENTS`, `BENCH_ROWS_PER_FRAGMENT`,
`BENCH_DELETE_PERCENT`, `BENCH_APPENDS`, `BENCH_SCENARIOS` and
`BENCH_INLINE_MAX_BYTES`. Bad values are rejected up front.

There are two workloads:
- `deleted`: a share of rows is deleted, then compacted. The sequences
still run-encode.
- `shuffled`: every row is rewritten in random order, with its real
created-at and last-updated-at versions, then compacted. No run
structure survives.

Within each workload, the inline arm never opts in, and the spilled arm
sets `lance.row_lineage.spill=true` with the given budget.

Per arm it reports:
- steady-state manifest bytes, split by sequence;
- the compaction manifest and transaction file;
- cold open;
- the commit of a small append, timed apart from the data write;
- one cold sequence load;
- the row id index build;
- a take by row id with the fragment's sequence preloaded;
- compaction time and the bytes compaction wrote.

Timed rows report the median and the minimum of their samples.

The bench checks itself outside the timers:
- the spilled arm must actually have spilled, and the inline arm must
not;
- a full scan asserts every row's `_rowid`, created-at and
last-updated-at against the workload;
- every take probe must return the row it asked for.

## Smoke run

This is one small run, not a performance claim. It ran on an x86
r8i.8xlarge (local NVMe) with 2 fragments x 200,000 rows,
`BENCH_INLINE_MAX_BYTES=0` so everything spills, the
`release-with-debug` profile, and each arm once. Lower is better for
every row, and the ratio is inline / spilled.

| Scenario / metric | Inline | Spilled | Ratio |
| --- | ---: | ---: | ---: |
| shuffled: steady-state manifest | 7.49 MB | < 0.01 MB | 10367x smaller
|
| shuffled: cold dataset open (median) | 4.32 ms | 0.12 ms | 36.5x
faster |
| shuffled: append commit (median) | 6.27 ms | 0.78 ms | 8.1x faster |
| shuffled: row id index build, cold (median) | 36.11 ms | 35.53 ms |
1.0x |
| shuffled: one sequence load, cold (median) | 0.18 ms | 1.97 ms | 0.09x
(spilled is slower) |
| shuffled: compaction output (data files) | 1.93 MB | 2.96 MB | 0.65x
(spilled is larger) |
| deleted: steady-state manifest | 0.14 MB | < 0.01 MB | 237x smaller |
| deleted: row id index build, cold (median) | 2.53 ms | 5.48 ms | 0.46x
(spilled is slower) |

The trade-off is the one the design expects. Spilling takes the per-row
lineage out of every manifest read and write, so open and commit stop
scaling with the table. The cost moves to a column read the first time a
sequence is needed, and 2 to 3 bytes per row in the compacted data
files. On the `deleted` workload the sequences are small either way,
which is why the default 200 KiB budget keeps them inline.

## Validation

- `cargo fmt --all -- --check`, `cargo clippy --all --tests --benches --
-D warnings` on the head of this branch
- The smoke run above completed with every self-check passing.

Refs #8931, #9250

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Will Jones <willjones127@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request K-approved Latest Gatekeeper recommendation permits acceptance.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants