Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions docs/src/format/table/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,6 +119,17 @@ Field ids might be replaced with `-2`, a tombstone value.
In this case that column should be ignored. This used, for example, when rewriting a column:
The old data file replaces the field id with `-2` to ignore the old data, and a new data file is appended to the fragment.

Every negative field id is reserved for system use and never names a field of the
dataset schema. A reader MUST skip any negative id when it projects the dataset schema
onto a data file, rather than treat it as a schema field or reject the file. Besides
`-1` (not yet assigned; only ever exists in memory and must not be written) and the
`-2` tombstone above, `-3`, `-4` and `-5` are the hidden `_rowid`,
`_row_created_at_version` and `_row_last_updated_at_version` columns that hold a
fragment's row lineage sequences when they are not stored in the manifest; see
[Row ID and Lineage](row_id_lineage.md). Such a column always lives in one of the
fragment's `files`, next to user columns or in a file holding nothing else, and at
most one file of a fragment may carry each of these ids.

## Data Files

Data files store column data for a fragment using the Lance file format.
Expand Down
85 changes: 76 additions & 9 deletions docs/src/format/table/row_id_lineage.md
Original file line number Diff line number Diff line change
Expand Up @@ -185,14 +185,81 @@ The implementation selects the most compact encoding based on the value range, c

</details>

#### Inline and External Storage
#### Inline and Spilled Storage

`DataFragment` defines inline and column alternatives for row ID sequences and row
version sequences. This allows small sequences to stay in the manifest (fewer
IOPS) while larger sequences (resulting from frequent updates) move outside the
manifest.

Sequences small enough (~200KB encoded and under) are stored inline in the fragment
metadata to avoid additional I/O. An inline sequence is rewritten into every manifest
version.

A larger sequence is **spilled to a hidden column of one of the fragment's data
files**. The column arm of the oneof (`column_row_ids`, `column_created_at_versions`
or `column_last_updated_at_versions`) is an empty `RowLineageColumn` marker: it carries
no file reference, because the file is one of the fragment's `files` and is found by a
reserved negative field id in that entry's `fields`. Exactly one file of the fragment
carries each reserved field; zero or more than one is corruption. The `column_indices`
entry paired with the field id locates the column the way it does for a user column, so
a lineage column
may share a file with the user columns or with the other lineage columns, or sit in a
file that holds nothing else. Reading it uses the ordinary data file reader and its
encodings.

The marker is valid only in a fragment whose data files are Lance v2 files. A legacy
v1 data file has no `column_indices` to locate the column by, and a fragment cannot
mix v1 and v2 files, so a writer on a dataset that stores v1 files MUST leave every
sequence inline; a marker in a fragment with a v1 data file is corruption.
Comment on lines +211 to +214

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We could maybe put a more general blanket statement just say somewhere that the stable row id feature is only enabled for v2 files.


The columns have this schema, with the field ids `-3`, `-4` and `-5` respectively.
The three names are reserved: a writer MUST reject a user column with any of them
(every Lance write path does), so a hidden column never collides with a field of the
dataset schema.

```python
import pyarrow as pa

row_lineage_columns = pa.schema([
pa.field("_rowid", pa.uint64(), nullable=False),
pa.field("_row_created_at_version", pa.uint64(), nullable=False),
pa.field("_row_last_updated_at_version", pa.uint64(), nullable=False),
Comment on lines +225 to +227

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does that mean these are reserved field names? Do we prevent users from writing them?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Created: #9455

])
```

`DataFragment` defines inline and external metadata fields as valid wire alternatives for row ID sequences and row version sequences.
These fields do not currently imply a size-based switching threshold.
Current Lance writers store all three sequence types inline in the fragment metadata regardless of their encoded size and do not emit the external alternatives.
Each column present holds exactly `physical_rows` values, one per physical row in
physical row order, deleted rows included; the value at offset `i` is the row id or
version of the row at offset `i`, the same thing the inline encoding's `i`-th entry
would be. A null value, or a column whose length differs from `physical_rows`, is
corruption and MUST be rejected rather than read as a default. A file may carry any
subset of the three columns; a sequence whose arm is not the column marker is not read
from any file, whatever the file holds.
Comment on lines +231 to +237

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can keep this paragraph if you want but it seems kind of redundant.


Which sequences may leave the manifest follows from when their values are known.
A value the commit assigns -- an appended fragment's row ids, an inserted row's
created-at version, every row's last-updated-at version -- can change when a commit
conflict is retried, so it stays inline where the retry can rewrite it; those
sequences are single runs and cost a few bytes. A value carried over from existing
rows -- the row ids and created-at versions that compaction or a row rewrite
preserves -- is fixed before the commit and may be written to a data file.
A writer spills only on a table that opts in through the `lance.row_lineage.spill`
config key; a table that never sets it is unchanged.

A writer that emits any column arm MUST set the spilled row lineage feature flag
(bit 11, value 2048) in both the reader and writer flag words. A reader without that
bit sees an unset oneof and would take the fragment to have no row IDs at all, on a
table whose manifest says every fragment has them.

Field numbers 6, 8 and 10 of `DataFragment` (`external_row_ids`,
`external_last_updated_at_versions`, `external_created_at_versions`) once named an
opaque byte range in a file holding the same encoding as the inline arm. No Lance
writer ever emitted them; the column arms replace that design, and the numbers and
names are reserved.

Current Lance readers can load externally stored row ID sequences.
The format also permits external created-at and last-updated-at version sequences, but current Lance readers cannot load them; this is an implementation limitation, not an invalid encoding.
!!! note
Spilled row lineage sequences are not yet a released feature. A released build
treats bit 11 as an unknown feature flag and refuses the dataset.

<details>
<summary>DataFragment row_id_sequence field</summary>
Expand All @@ -201,7 +268,7 @@ The format also permits external created-at and last-updated-at version sequence
message DataFragment {
oneof row_id_sequence {
bytes inline_row_ids = 5;
ExternalFile external_row_ids = 6;
RowLineageColumn column_row_ids = 12;
}
}
```
Expand Down Expand Up @@ -288,7 +355,7 @@ RowDatasetVersionSequence {
message DataFragment {
oneof created_at_version_sequence {
bytes inline_created_at_versions = 9;
ExternalFile external_created_at_versions = 10;
RowLineageColumn column_created_at_versions = 14;
}
}
```
Expand Down Expand Up @@ -331,7 +398,7 @@ New physical row (current):
message DataFragment {
oneof last_updated_at_version_sequence {
bytes inline_last_updated_at_versions = 7;
ExternalFile external_last_updated_at_versions = 8;
RowLineageColumn column_last_updated_at_versions = 13;
}
}
```
Expand Down
3 changes: 2 additions & 1 deletion docs/src/format/table/versioning.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,8 @@ they should return an "unsupported" error on any read or write operation.
| 256 | `FLAG_MIXED_DATA_FILE_VERSIONS` | Yes | Yes | The snapshot may reference recognized V2 data files with different exact versions. Both bits must be set and remain set on later versions. |
| 512 | `FLAG_FRAG_REUSE_WITH_STABLE_ROW_IDS` | Yes | Yes | The table uses stable row IDs and carries a [Fragment Reuse Index](../index/system/frag_reuse.md). |
| 1024 | `FLAG_FRAGMENT_REUSE_INDEX` | Yes | Yes | The fragment reuse index records tagged transitions (`IndexMetadata.index_version >= 1`). Readers must translate row addresses through them; writers must preserve them. An implementation without this flag would decode the details as the legacy format and silently drop the transitions when it next rewrites the fragment reuse index. See [FRI index versions](../index/system/frag_reuse.md#fri-index-versions). |
| 2048 | `FLAG_UNSTABLE_SPILLED_ROW_LINEAGE` | Yes | Yes | Some fragment stores its row ids or row version sequences as hidden columns of a data file rather than inline. A reader without this flag would see the fragment as having no row ids. Unstable: release builds reject it unless explicitly opted in. |

</div>

Flags with bit values 2048 and above are unknown; unknown flags cause implementations to reject the dataset with an "unsupported" error. The paired mixed-version reader and writer bits must either both be set or both be clear; a half-set manifest is invalid.
Flags with bit values 4096 and above are unknown; unknown flags cause implementations to reject the dataset with an "unsupported" error. The paired mixed-version reader and writer bits must either both be set or both be clear; a half-set manifest is invalid.
59 changes: 51 additions & 8 deletions protos/table.proto
Original file line number Diff line number Diff line change
Expand Up @@ -412,28 +412,46 @@ message DataFragment {
oneof row_id_sequence {
// Current Lance writers store row ids inline regardless of encoded size.
bytes inline_row_ids = 5;
// Supported by current Lance readers, but not emitted by current Lance writers.
ExternalFile external_row_ids = 6;
/* The row ids are a hidden column of one of this fragment's `files`: the
* entry whose `fields` carries the reserved id -3. See RowLineageColumn.
*
* A writer MUST set the spilled row lineage feature flag (2048) when any
* fragment uses this arm or either of the other column arms below. A
* reader that does not understand the flag would see an unset oneof and
* treat the fragment as having no row ids at all.
*/
RowLineageColumn column_row_ids = 12;
} // row_id_sequence

oneof last_updated_at_version_sequence {
// Current Lance writers store last-updated versions inline regardless of encoded size.
bytes inline_last_updated_at_versions = 7;
/* Valid external alternative. Current Lance writers do not emit this field,
* and current Lance readers cannot load it.
/* The versions are a hidden column of one of this fragment's `files`: the
* entry whose `fields` carries the reserved id -5. Gated like
* `column_row_ids`.
*/
ExternalFile external_last_updated_at_versions = 8;
RowLineageColumn column_last_updated_at_versions = 13;
} // last_updated_at_version_sequence

oneof created_at_version_sequence {
// Current Lance writers store created-at versions inline regardless of encoded size.
bytes inline_created_at_versions = 9;
/* Valid external alternative. Current Lance writers do not emit this field,
* and current Lance readers cannot load it.
/* The versions are a hidden column of one of this fragment's `files`: the
* entry whose `fields` carries the reserved id -4. Gated like
* `column_row_ids`.
*/
ExternalFile external_created_at_versions = 10;
RowLineageColumn column_created_at_versions = 14;
} // created_at_version_sequence

/* The `external_*` arms of the three sequence oneofs above: an opaque byte
* range in a file, holding the same encoding as the inline arm. No Lance
* writer ever emitted them; the column arms replace that design. Reserved so
* a reader never has to interpret them.
*/
reserved 6, 8, 10;
reserved "external_row_ids", "external_last_updated_at_versions",
"external_created_at_versions";

/* Number of original rows in the fragment, this includes rows that are now marked with
* deletion tombstones. To compute the current number of rows, subtract
* `deletion_file.num_deleted_rows` from this value.
Expand All @@ -451,6 +469,11 @@ message DataFile {
* used for "tombstoned", meaning a field that is no longer in use. This is often
* because the original field id was reassigned to a different data file.
*
* Every negative value is reserved for system use and never names a field of the
* dataset schema. -3, -4 and -5 are the hidden row lineage columns (see
* RowLineageColumn); a reader projecting the dataset schema onto a file skips
* them, and any other negative value, rather than reject the file.
*
* In Lance v1 IDs are assigned based on position in the file, offset by the max
* existing field id in the table (if any already). So when a fragment is first created
* with one file of N columns, the field ids will be 1, 2, ..., N. If a second fragment
Expand Down Expand Up @@ -634,6 +657,26 @@ message DeletionFile {
optional uint32 base_id = 7;
} // DeletionFile

/* Marks a row lineage sequence that is stored as a hidden column of one of the
* fragment's `files` rather than inline: `_rowid` under the reserved field id
* -3, `_row_created_at_version` under -4, `_row_last_updated_at_version` under
* -5. Exactly one entry of `DataFragment.files` carries that id in its
* `fields`, and its `column_indices` locates the column the way it does for a
* user column, so the column may share a file with the user columns or with the
* other lineage columns, or sit in a file that holds nothing else. The column is
* a non-nullable uint64 with exactly `physical_rows` values, in physical row
* order; see the table format documentation for the schema.
*
* Carries no fields: the file is found by its field id, so the file's metadata
* lives in `files` alone.
*
* Valid only in a fragment whose data files are Lance v2 files. A legacy v1
* data file has no `column_indices` to locate the column by, so a writer on a
* dataset that stores v1 files leaves every sequence inline.
*/
message RowLineageColumn {}
Comment thread
BubbleCal marked this conversation as resolved.

// A byte range of a file. Used by the fragment reuse index details.
message ExternalFile {
// Path to the file, relative to the root of the table.
string path = 1;
Expand Down
4 changes: 2 additions & 2 deletions python/src/rowids.rs
Original file line number Diff line number Diff line change
Expand Up @@ -48,8 +48,8 @@ impl PyRowIdSequence {
fn from_inline_metadata(metadata: PyRef<'_, PyRowIdMeta>) -> PyResult<Self> {
match &metadata.0 {
RowIdMeta::Inline(data) => read_row_ids(data).infer_error().map(Self),
RowIdMeta::External(_) => Err(PyNotImplementedError::new_err(
"Row ids stored in an external file cannot be read into a RowIdSequence",
RowIdMeta::Column => Err(PyNotImplementedError::new_err(
"Row ids stored outside the manifest cannot be read into a RowIdSequence",
)),
}
}
Expand Down
Loading
Loading