Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
08bebbe
feat: stabilize field IDs across schema evolution
Xuanwo Aug 20, 2026
2964690
fix: canonicalize stable field ids at commit boundaries
Xuanwo Aug 20, 2026
7e46c86
fix: enforce stable field ids in namespace rewrites
Xuanwo Aug 20, 2026
7d59c81
fix: preserve stable identities across schema operations
Xuanwo Aug 20, 2026
2b45dbe
fix: reject ambiguous stable field mappings
Xuanwo Aug 20, 2026
8851484
fix: centralize binding field id remapping
Xuanwo Aug 20, 2026
bb42780
fix(java): preserve legacy field id remapping
Xuanwo Aug 20, 2026
d199a9a
fix(java): validate project field identities
Xuanwo Aug 20, 2026
3b041d1
test(java): verify legacy project values
Xuanwo Aug 20, 2026
981ff8e
fix: address stable field ID review feedback
Xuanwo Aug 31, 2026
41867ac
Merge origin/main into xuanwo/stable-field-ids
Xuanwo Aug 31, 2026
4a17fc2
fix: complete stable field ID main integration
Xuanwo Aug 31, 2026
d4e8ab6
Merge origin/main into xuanwo/stable-field-ids
Xuanwo Aug 31, 2026
aff26e7
fix: make stable field IDs writer-only
Xuanwo Aug 31, 2026
3a560ff
fix: preserve raw field bindings across commit retries
Xuanwo Aug 31, 2026
d952779
fix: fence automatic stable field ID activation
Xuanwo Aug 31, 2026
4ced65b
Merge remote-tracking branch 'origin/main' into xuanwo/stable-field-ids
Xuanwo Aug 31, 2026
399260c
fix: align stable field IDs in overwrite fragments
Xuanwo Aug 31, 2026
6fde30f
fix: enforce stable field ID commit boundaries
Xuanwo Sep 1, 2026
678e855
fix: align stable field ID activation contract
Xuanwo Sep 2, 2026
7927343
fix: clarify stable field IDs and make schema input explicit
Xuanwo Sep 21, 2026
eed6d73
refactor: resolve Arrow field IDs at conversion boundaries
Xuanwo Sep 21, 2026
32f617c
Merge origin/main into xuanwo/stable-field-ids
Xuanwo Sep 21, 2026
f73688a
fix: reconcile Arrow schema conversion after merging main
Xuanwo Sep 21, 2026
f4afe81
refactor: simplify stable field ID allocation and tests
Xuanwo Sep 21, 2026
122f7f7
Merge branch 'main' into xuanwo/stable-field-ids
Xuanwo Sep 21, 2026
477cea2
docs: clarify stable field ID migration requirements
Xuanwo Oct 2, 2026
8af6bf5
Merge origin/main into xuanwo/stable-field-ids
Xuanwo Oct 2, 2026
e35633e
fix: normalize new dataset IDs before blob promotion
Xuanwo Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions docs/src/format/table/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,19 @@ A manifest describes a single version of the dataset.
It contains the complete schema definition including nested fields, the list of data fragments comprising this version,
a monotonically increasing version number, and an optional reference to the index section that describes a list of index metadata.

`max_allocated_field_id` is optional. If a manifest sets it, the dataset uses stable field IDs from
that version onward. See [Field IDs](schema.md#field-ids).

The field is a high-water mark. It starts at the largest field ID that the activation manifest
references. It then records the largest ID assigned after activation. When a manifest sets it:

- Every field ID of 0 or greater in the manifest schema, a `DataFile.fields` mapping, or an overlay
mapping must be less than or equal to `max_allocated_field_id`.
- A writer must not lower `max_allocated_field_id` from the value in the previous manifest.
- A writer that adds fields must assign IDs greater than the previous `max_allocated_field_id` and
set the new high-water mark to at least the largest ID it assigned. The writer must fail if an
assigned ID does not fit in an `int32`.

<details>
<summary>Manifest protobuf message</summary>

Expand Down
118 changes: 110 additions & 8 deletions docs/src/format/table/schema.md
Comment thread
Xuanwo marked this conversation as resolved.
Original file line number Diff line number Diff line change
Expand Up @@ -221,15 +221,60 @@ Assigned IDs with parent relationships:
Note: A `parent_id` of -1 indicates a top-level field. For nested fields, `parent_id` references the ID of the parent field. Child fields reference their parent via `parent_id` rather than being stored as separate "children" arrays in the protobuf message (though the Rust in-memory representation maintains a children vector for convenience).

**New field assignment (incremental):**
When fields are added later (e.g., through schema evolution), they receive the next available ID
incrementally. This preserves the history of field additions.

`Manifest.max_allocated_field_id` selects between two behaviors:

- If the manifest does not set the field, the dataset uses the legacy behavior. A writer may choose
the next ID from fields the current version still references. It may therefore reuse the ID of a
dropped field.
- If the manifest sets the field, the dataset uses stable field IDs. A writer assigns each new field
an ID greater than `max_allocated_field_id`. It does not reuse an ID dropped or replaced after
activation.

For stable field IDs, a caller cannot choose the ID of a new field. An Arrow schema may carry
field-ID metadata, but the writer discards that metadata for new fields and assigns the IDs. The IDs
do not have to be consecutive, which leaves room for a future reservation mechanism.

The first manifest that sets `max_allocated_field_id` initializes it to the largest field ID of 0
or greater in the manifest schema, base data files, and overlay files. Earlier versions keep the
legacy behavior. Activation cannot recover an ID that an earlier version dropped or reused.

`max_allocated_field_id` stores the allocator state. `FLAG_STABLE_FIELD_IDS` tells writers that they
must honor that state. A legacy manifest sets neither value. A stable manifest sets both. A manifest
that sets only one is invalid. The reader flag for stable field IDs must remain unset because the
feature does not change read behavior.

A dataset changes to stable field IDs only through an explicit migration commit. Before activation,
operators must ensure that all clients that can write to the dataset enforce writer feature flags,
rejecting writes when they do not support a required flag. Clients that ignore these flags must no
longer write to the dataset: they may discard the high-water mark and allow field IDs to be reused.

A dataset cannot return to the legacy behavior. After activation, a restore must fail if it targets
a version that does not set `max_allocated_field_id`. Before activation, different fields may have
used the same ID in different versions. For example, an old version may assign ID 1 to an integer
field `x`, while the activation version assigns it to a string field `y`. Restoring the old version
would make ID 1 refer to `x` again. Keeping the current high-water mark prevents future allocation
from reusing IDs, but does not resolve this existing conflict. Reading old versions remains supported.

### Field ID Properties

- **Immutable**: Once assigned, a field's ID never changes
- **Unique**: Each field within a table has a unique ID
- **Stable**: IDs are preserved across schema evolution operations
- **Sparse**: Field IDs may not form a contiguous sequence after schema evolution
- **Stable**: A field keeps the same ID for as long as the field exists.
- **Unique**: No two fields in one dataset version have the same ID.
- **Sparse**: The field IDs in one version do not have to be consecutive.

When `max_allocated_field_id` is set, two more properties apply:

- **Not reused**: After activation, no later version uses the ID of a dropped or replaced field.
- **Increasing**: Every new ID is greater than the activation high-water mark and every ID assigned
after activation.

A field ID is unique within one branch of one dataset. It is not unique across datasets or across
branches that changed independently. A reference stored outside the dataset must name the dataset
and branch as well as the field ID.

Two branches can assign the same field ID after they diverge. Lance does not yet merge branches. A
future merge operation must fail if the branches assigned the same ID to different fields. It must
not pick one field, change an ID stored by an existing version, or match the fields by name.

### Using Field IDs

Expand Down Expand Up @@ -296,13 +341,70 @@ The complete schema is represented as a collection of top-level fields plus meta
Field IDs enable efficient schema evolution:

- **Add Column**: Assign a new field ID and add to schema
- **Drop Column**: Remove field from schema; its ID may be reused in some systems
- **Drop Column**: Remove the field from the schema; when `max_allocated_field_id` is set, later
versions must not reuse its ID
- **Rename Column**: Change field name; ID remains the same
- **Reorder Columns**: Change field order in schema; IDs remain the same
- **Type Evolution**: Data type can be changed. This might require rewriting the column in the data, depending on how the type was changed.
- **Metadata or Nullability Change**: Preserve the field ID
- **Type Replacement**: A cast creates a replacement field with a new ID and retires the old
identity. This keeps one logical type bound to an ID in every version that references it
- **Overwrite**: Preserve compatible logical identities; allocate new IDs for added fields and type
replacements

The use of field IDs ensures that data files can be correctly interpreted even as the schema changes over time.

### Blob Identity Namespace

The rules above apply to a Blob field in the manifest schema and to its logical children. For
example, assume `image` has field ID 0, `data` has ID 1, and `uri` has ID 2. Writer input and the
manifest schema have this logical shape:

```python
pa.schema([
pa.field(
"image",
pa.struct([
pa.field("data", pa.large_binary()),
pa.field("uri", pa.string()),
]),
metadata={b"ARROW:extension:name": b"lance.blob.v2"},
),
])
```

The writer may temporarily add `kind`, `blob_id`, `blob_size`, and `position`. A Lance data file
stores this descriptor shape:

```python
pa.schema([
pa.field(
"image",
pa.struct([
pa.field("kind", pa.uint8(), nullable=False),
pa.field("position", pa.uint64(), nullable=False),
pa.field("size", pa.uint64(), nullable=False),
pa.field("blob_id", pa.uint32(), nullable=False),
pa.field("blob_uri", pa.string(), nullable=False),
]),
),
])
```
Comment on lines +375 to +391

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question(non-blocking): when these fields are stored, how is that represented in both DataFile.fields and DataFile.column_indices? Below it says DataFile.fields = [0]. Are they really missing from the DataFile metadata? If so, how does a reader know where in the Lance file to find those columns?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The descriptor occupies one physical column: fields = [0] maps to column_indices = [k]. The reader decodes its children from column k using the Blob page layout, so they need no separate mapping entries.


The entire descriptor is encoded in one physical column. The data file maps it with
`DataFile.fields = [0]` and `DataFile.column_indices = [k]`, where `k` is that column's index in the
Lance file. The reader locates column `k` and decodes the descriptor using its Blob page layout;
the descriptor children do not have separate entries in either mapping. Their IDs may be `-1` or
file-local, and they do not change `max_allocated_field_id`.

A descriptor scan returns the stored struct. A materialized scan returns this public shape:

```python
pa.schema([pa.field("image", pa.large_binary())])
```

Both scans refer to the top-level field ID 0. The synthetic descriptor children do not become
dataset fields. The `blob_id` value identifies a stored Blob object; it is not a field ID.

## Example Schemas

The examples below use a simplified representation of the field structure. In the actual protobuf format, `type` refers to the field type enum (PARENT/REPEATED/LEAF) and `logical_type` contains the data type string representation.
Expand Down
3 changes: 2 additions & 1 deletion docs/src/format/table/versioning.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,8 @@ they should return an "unsupported" error on any read or write operation.
| 2048 | `FLAG_UNSTABLE_SPILLED_ROW_LINEAGE` | Yes | Yes | Some fragment stores its row ids or row version sequences as hidden columns of a data file rather than inline. A reader without this flag would see the fragment as having no row ids. Unstable: release builds reject it unless explicitly opted in. |
| 4096 | `FLAG_FRAGMENT_TREE` | Yes | Yes | Fragment records live in a [fragment tree](fragment_metadata.md). `Manifest.fragments` is empty. |
| 8192 | `FLAG_INDEPENDENT_COVERING_FIELDS` | Yes | Yes | Requires `FLAG_COVERED_INDEX_METADATA` and is retained together with it. `IndexMetadata.fields` contains only key fields, while `covering_fields` independently declares carried fields and may overlap `fields`. Implementations that only support the legacy suffix contract must reject the dataset. |
| 16384 | `FLAG_STABLE_FIELD_IDS` | No | Yes | The manifest sets `max_allocated_field_id`, and a writer must assign new field IDs above it. Before activation, all clients that can write to the dataset must enforce writer feature flags. See [Field IDs](schema.md#field-ids). |

</div>

Flags with bit values 16384 and above are unknown; unknown flags cause implementations to reject the dataset with an "unsupported" error. The paired mixed-version reader and writer bits must either both be set or both be clear; a half-set manifest is invalid.
Flags with bit values 32768 and above are unknown; unknown flags cause implementations to reject the dataset with an "unsupported" error. The paired mixed-version reader and writer bits must either both be set or both be clear; a half-set manifest is invalid.
Loading
Loading