Skip to content

feat(compaction): repack a fragment's columns as a compact_files task - #9604

Open
LuciferYang wants to merge 27 commits into
lance-format:mainfrom
LuciferYang:feat/horizontal-compaction
Open

LuciferYang wants to merge 27 commits into
lance-format:mainfrom
LuciferYang:feat/horizontal-compaction

Conversation

@LuciferYang

@LuciferYang LuciferYang commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

What's wrong today

Each add_columns backfill gives every fragment one more data file. Scans and random reads slow down as the count grows, and compact_files never fixes it: its candidates are fragments with too many deletions or overlays, or too few rows, so a large, clean fragment split across many files is never picked. This is the "horizontal compaction" gap from #4286 and #3995.

What this adds

compact_files gets a second kind of task. Besides rewriting fragments into new fragments, a task can repack the columns of one fragment into fewer data files without moving its rows, so the fragment id, row addresses, deletions, overlays and index coverage don't change.

The default planner plans fragment rewrites exactly as before, then looks at every fragment no rewrite took and plans a repack for it when either:

  • it holds columns in more than max_data_files_per_fragment files, as Dataset::column_layout_stats counts them (default None, so nothing changes unless it is set), or
  • its files don't match column_groups.

With column_groups, a group that no single file holds exactly is moved into a file of its own. When the fragment is also over the limit, the files holding only columns no group names are merged into one file.

A merge only takes files it empties, two or more at a time, so it always lowers the file count. The planner leaves a file alone if it holds a blob column, a column only partly present in the fragment, or the fragment's spilled row lineage, and also if it shares a column with the lineage file. A V2.0 file can keep a struct's header after the struct's last child in that file was dropped, and such a file is merged only together with the struct. Without groups, when every file could go, one of them stays (the largest, if some merge leaves it in place) unless the limit is 1.

scope limits a run to one kind of task: RewriteFragments, RepackColumns, or All (the default). max_data_files_per_fragment and column_groups can also be set as table config (lance.compaction.max_data_files_per_fragment, lance.compaction.column_groups). scope is a per-run option only. Repacks aren't planned under ForceBinaryCopy, since they reencode.

column_layout_stats reports, for each fragment, how many data files hold a schema column, each file's recorded size and number of schema fields, the share of field slots that hold no live data (tombstoned or left by a dropped column), and how many overlays it has. It reads only manifest metadata. The planner's file-count trigger uses it, and callers can read it to see why a fragment was or wasn't picked: dataset.stats.column_layout_stats() in Python, Dataset.getColumnLayoutStatistics() in Java.

The compaction section of the user guide documents the new options and config keys.

How the tasks run and commit

Tasks run through a new object-safe CompactionExecutor trait. DefaultCompactionExecutor runs both kinds. compact_files_with_executor(dataset, remap_options, planner, executor) lets a caller pass its own executor (for example, one that runs tasks on a cluster) alongside the existing CompactionPlanner. compact_files_with_planner and CompactionTask::execute use the default executor.

A repack task reads only the base values of the columns it moves. Overlays are left off the read on purpose, so they keep shadowing the new files the same way they shadowed the old ones. It writes one new file per planned file, lined up with the fragment's physical rows, and deletes what it wrote if a later file fails.

commit_compaction commits fragment rewrites as one Operation::Rewrite, as before, then commits repacks as one Operation::DataReplacement { data_change: false }. A run with both kinds therefore makes two versions. The default planner never puts one fragment in both kinds of task, and commit_compaction rejects results that do before either commit, so the second commit doesn't conflict with the first. If the second commit fails, the first stays committed and the repacks can be planned again.

Why DataReplacement with data_change: false

Repacking moves values without changing them. DataReplacement already means "these fragments get new files for these fields", and the composite transactions draft (#8644, AddDataFile) uses a data_change flag for the same idea. With data_change: false:

  • index coverage of the moved fields is kept, and a concurrent index build doesn't conflict;
  • overlays aren't tombstoned, since they still hold the newer values and keep shadowing the base;
  • no row's last-updated version moves, so an incremental reader doesn't see the moved rows as changed;
  • the commit doesn't count as supplying values to a column whose nullability a concurrent change is tightening.

DataReplacement is applied to the fragment as it stands at commit, not to a post-image built at read time. If a concurrent commit changes the fragment's other columns, deletes some of its rows, or drops a column the repack leaves alone, that change is kept and the repack still commits. A concurrent update or replacement of a moved column, a rewrite of the fragment, a delete of the whole fragment, or a drop_columns of a moved column makes the repack retry. With data_change: false, the last two are retryable, while a data_change: true replacement treats them as incompatible.

To carry a repack, DataReplacement now accepts several groups for one fragment (one per new file), applied in order. The check that every new file in one operation has the same fields is gone. data_change is an optional bool in the protobuf. A transaction written before it existed reads as true, which is what it meant.

What changes for existing callers

  • TaskData gains kind and RewriteResult gains repacked_files, both #[serde(default)]. A plan or result serialized by an older release reads as a fragment rewrite.
  • An older worker given a repack task ignores the unknown field and rewrites the fragment instead. That's more work, but it's a valid compaction and the coordinator commits it as a rewrite.
  • Java TaskData and RewriteResult now declare the serialVersionUID their current shape had, so streams from older workers still deserialize. Java and Python expose the three options, the task kind (TaskData.getRepackFiles(), CompactionTask.kind), and DataReplacement.dataChange.
  • try_harvest_seeds opens the data file holding the indexed field instead of the fragment's first file. Under column_groups an indexed column can sit in a later file, and after a repack the old file only holds a tombstone for it. Single-file fragments read the same file as before.
  • A column_groups name that isn't a top-level column is skipped with a warning instead of failing the run. A persisted group can outlive the column it names, since drop_columns and renames leave the config as it is.
  • max_source_bytes counts only the files a repack task reads, which are the files holding the columns it moves, with no overlays.

Known limits

  • A run that both rewrites and repacks makes two versions until composite transactions can commit them as one.
  • A repack keeps overlays as overlays. It doesn't fold them into the new base; a fragment rewrite still does that.
  • On a table with spilled row lineage, the file carrying the lineage keeps its columns, and a repack that would move them all out isn't planned. The new files carry no lineage, so that file would stay referenced for the lineage alone.
  • Fragments with legacy (V1) files aren't repacked, and blob columns stay where they are, so such a fragment can stay over the limit.
  • A V2.0 file holding only a struct header counts toward the limit. A repack removes it only when it moves that struct; otherwise a fragment rewrite does.
  • Under column_groups, max_bytes_per_file is ignored so every group's files split at the same rows.

Tests

Test What it pins
repack_collapses_backfilled_files two backfills give three files per fragment; one run brings them to one, in one commit, with values and fragment ids unchanged; a second run commits nothing
repack_keeps_the_biggest_file, repack_follows_column_groups the layout chosen with and without groups, and that a matching layout plans nothing
repack_keeps_deletions, repack_drops_file_left_with_only_dropped_columns deleted rows stay deleted; a file left with only dropped columns is removed
repack_keeps_scalar_index, repack_columns_preserves_vector_index, repack_columns_preserves_data_overlay BTree and IVF_PQ indices keep uuid and coverage and return the same rows; an overlay still shadows the repacked base
compaction_rewrites_and_repacks_in_one_run, compaction_scope_selects_task_kinds a mixed run commits Rewrite then DataReplacement and keeps the two task kinds on separate fragments; scope selects one kind
repack_tasks_run_distributed, task_data_without_kind_rewrites_fragments repack tasks and results survive serde and commit together; a pre-kind task reads as a rewrite
repack_against_concurrent_delete, repack_against_concurrent_drop, repack_against_concurrent_update_of_moved_column which concurrent changes the repack commits over and which make it retry
repack_merges_unclaimed_columns_over_the_limit with groups set, the unclaimed columns merge only over the limit
repack_keeps_the_spilled_lineage_file, repack_merges_around_a_blob_column, repack_merges_around_a_blob_outside_the_kept_file, repack_never_adds_a_file, repack_skips_legacy_files the lineage, blob and V1 rules; no plan adds a file
repack_ignores_struct_header_without_children, repack_moves_a_struct_with_its_header, repack_merges_only_files_it_empties, repack_keeps_the_largest_file_that_shares_a_struct, repack_keeps_the_largest_file_next_to_a_header_only_file, repack_prefers_leaving_the_largest_file_in_place, repack_plans_nothing_for_header_only_files, repack_with_a_dropped_struct_child_converges V2.0 struct headers: a header doesn't make a file hold the struct, a merge counts the files a header keeps alive, which file stays, and that a second run plans nothing
repack_is_not_planned_under_force_binary_copy, repack_counts_only_the_files_it_reads_against_the_byte_budget, commit_rejects_a_fragment_in_both_kinds_of_task the ForceBinaryCopy, max_source_bytes and mixed-result guards
test_conflicts_data_replacement (new cases), test_data_replacement_data_change_roundtrips, test_data_replacement_without_data_change_keeps_overlays, test_data_replacement_applies_several_groups_to_one_fragment, prepare_indices_keeps_coverage_for_moved_values the DataReplacement changes on their own
column_layout_stats_counts_files_per_fragment, column_layout_stats_counts_dropped_columns_as_dead_slots, column_layout_stats_skips_files_without_schema_columns which files count as live, the per-file field counts and sizes, and that tombstones and dropped columns count as dead slots while spilled lineage counts on neither side
test_compact_files_repacks_columns, test_distributed_repack (Python), testRepackColumns (Java) the options, the task kind and the layout stats across the bindings, including pickling and Java serialization

Repack a fragment's per-column data files into fewer files without moving rows. Adds CompactionExecutor/CompactionCommitter traits (vertical reuses the existing rewrite path), a horizontal RewriteColumns path committing Operation::Update with an empty fields_modified (indices and overlays preserved), column_layout_stats, and a compaction column_groups option.
Add a generic run_compaction_pipeline<E,C> that both vertical and horizontal compaction run through, so the traits are a real plug point. Route Dataset::rewrite_columns through it (bounded concurrency, no bespoke cleanup). Surface a missing fragment and an uncovered indexed column as errors instead of silently miscounting or dropping a seed. Add tests for column_groups seed routing (ZoneMap), split-preserved-across-compaction, config parsing, and index/overlay preservation.
@github-actions

Copy link
Copy Markdown
Contributor

ACTION NEEDED
Lance follows the Conventional Commits specification for release automation.

The PR title and description are used as the merge commit message. Please update your PR title and description to match the specification.

For details on the error please inspect the "PR Title Check" action.

@LuciferYang LuciferYang changed the title Horizontal compaction via a pluggable compaction pipeline (overview) horizontal compaction via a pluggable compaction pipeline (overview) Sep 29, 2026
@LuciferYang

Copy link
Copy Markdown
Contributor Author

@wjones127 @Xuanwo, picking up the horizontal-compaction thread (#4286, discussion #3995, and the stalled #9291). I built the whole thing out on a branch and split it into reviewable PRs; here's where it stands.

The gap, restated: repeated add_columns backfills leave a fragment split across many per-column files, which slows scans and random reads, and compact_files can't help because it's vertical and its candidate check never looks at per-fragment file count. Horizontal compaction repacks a fragment's columns into fewer files without moving rows, so indices and row addresses stay valid.

Two things you each asked for shaped the design:

Correctness rests on committing horizontal compaction as Operation::Update{RewriteColumns} with an empty fields_modified, reusing the existing transaction path instead of adding a new Operation. Empty means index-maintenance prunes nothing (coverage kept) and overlay tombstoning is a no-op (overlays carried forward, still shadowing). I tested the parts that were only argued before: a vector index keeps its uuid and coverage across a rewrite of the indexed column and returns the same KNN; a data overlay survives a rewrite of an unrelated column; the file count drops from three to one after two backfills.

Layout of the stack:

A couple of things I'd rather settle with you than alone:

  • Dataset::rewrite_columns reuses the RewriteColumns name that merge-insert already uses for a different operation. Keep it for continuity with feat: keep column groups through compaction and add rewrite_columns #9291, or rename (say repack_columns) to avoid the collision?
  • The framework keeps vertical and horizontal as separate entries committing their own operation, rather than one planner emitting both task kinds (which would need a CompactionPlan wire-format change). Does that match what you had in mind, or did you want the unified planner?

Review of #9606 and #9605 would be much appreciated; #9604 has the full context for the direction question.

@LuciferYang LuciferYang changed the title horizontal compaction via a pluggable compaction pipeline (overview) feat: horizontal compaction via a pluggable compaction pipeline (overview) Sep 29, 2026
@github-actions github-actions Bot added the enhancement New feature or request label Sep 29, 2026
@LuciferYang
LuciferYang marked this pull request as ready for review September 29, 2026 14:31

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-changes Latest Gatekeeper recommendation requests changes. label Sep 29, 2026
RewriteColumnsCommitter committed against the current manifest version, so in the distributed path a fragment changed between its read and the commit would slip past the conflict check. Carry each rewrite's read version in RewriteColumnsResult and commit against the minimum, mirroring vertical compaction's RewriteResult::read_version. Single-machine rewrite_columns is unaffected (read version == commit version).
@lance-gatekeeper lance-gatekeeper Bot removed the K-changes Latest Gatekeeper recommendation requests changes. label Sep 29, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

❌ Gate recommendation: request changes.

The read-version change closes the stale distributed-commit data-loss path from the previous review. The remaining acceptance step is a checked-in horizontal rewrite regression test: execute at an older snapshot, delete before commit, and verify the commit conflicts while the delete remains. That protects the new result-to-transaction contract in distributed use.

Please mark this PR with the breaking-change label.

// Commit against the earliest version any rewritten fragment was read
// at, not the current one, so the conflict check catches a fragment
// changed between the read and this commit (as vertical compaction does).
let read_version = results

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new read-version path has no horizontal regression test, leaving this data-integrity fix unguarded. A stale rewrite committed after a delete previously restored the deleted row, and the repository requires a corresponding test for bugfixes. I ran the scenario at this head: execute RewriteColumnsExecutor at version V, delete a = 1 at V+1, commit its RewriteColumnsResult, then assert a conflict, seven remaining rows, and no a = 1; all assertions pass. Please add this scenario to this module's tests so the version binding cannot silently regress.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in b15e1e9: the checked-in test executes a fragment rewrite before a delete, verifies the stale commit fails, and confirms the deleted row stays absent. I ran the focused test at this head; it passed.

@lance-gatekeeper lance-gatekeeper Bot added the K-changes Latest Gatekeeper recommendation requests changes. label Sep 29, 2026
Regression test for the read-version binding: a rewrite read at V, committed after a delete lands on the same fragment at V+1, must conflict rather than resurrect the deleted row. Guards the fix so the distributed commit path cannot silently regress.
@lance-gatekeeper lance-gatekeeper Bot removed the K-changes Latest Gatekeeper recommendation requests changes. label Sep 29, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

The read-version binding and new regression test close the stale distributed-rewrite path: a delete between execution and commit now causes the rewrite to fail, and the deleted row remains absent. Horizontal rewriting reduces per-fragment file fan-out while preserving row addresses and index coverage.

Please mark this PR with the breaking-change label.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 29, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 2, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

Horizontal rewriting still reduces column-file fan-out while preserving row addresses, indexes, overlays, and row lineage. This revision protects distributed commits against deletes and column drops, and keeps grouped compaction's file alignment and index seed lookup consistent.

Please mark this PR with the breaking-change label.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 2, 2026
@github-actions github-actions Bot added A-python Python bindings A-java Java bindings + JNI A-format On-disk format: protos and format spec docs format-change A change to the format spec, which requires a vote. Remove if minor (e.g. fixing typo). labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Important

Format specification vote

This PR modifies the Lance format specification, so it requires 3 binding +1 votes from PMC members (excluding the proposer), at least one of them on the latest commit, and a minimum 72-hour voting period, weekends excluded, before it can merge. Vote by approving this PR (+1) or requesting changes (−1, a veto). See the voting process.

Approvals carry over across pushes, so a rebase or a typo fix does not send everyone back to re-vote. Whoever approves the latest commit is vouching that nothing substantive has changed since the earlier approvals; if something has, ask for fresh votes.

Status: ❌ Blocked — 0 of 3 required approvals

Approvals none (0/3)
Latest commit approved by none — one PMC member must approve the latest commit
Vetoes none
Voting period ends Wed 2026-10-07 19:37 UTC (12:37 PDT)

Updated automatically by the format-spec vote gate, which re-checks every 15 minutes — just voted? Re-check now (press Run workflow; leave the input blank to re-check every open format PR). A PMC member may apply the format-waived label to waive the vote for a trivial edit (typo, wording, formatting).

@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 2, 2026
@LuciferYang LuciferYang changed the title feat: horizontal compaction via a pluggable compaction pipeline (overview) feat(compaction): repack a fragment's columns as a compact_files task Oct 2, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

Integrating column repacks into compact_files follows the pluggable planner/executor direction requested in #9291. The replacement commit preserves row addresses, deletion vectors, overlays, index coverage, and row lineage while detecting overlapping base changes.

Mixed rewrite/repack runs commit two versions. If the repack phase fails, refresh and replan; the completed rewrite remains committed.

Please mark this PR with the breaking-change label.

Comment thread python/python/tests/test_optimize.py Outdated
expected = dataset.to_table()

plan = Compaction.plan(
dataset, options=dict(column_groups=[["d"]], scope="repack_columns")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This test fails before exercising pickling or commit: _backfilled already stores d alone, so this group matches the existing layout and the planner correctly emits no tasks. Use column_groups=[["c", "d"]] to force a repack. I ran that variant: two tasks survive both pickle round trips, commit to two files per fragment, and preserve the data and fragment IDs.

Reproduced from python/ with uv run pytest python/tests/test_optimize.py::test_distributed_repack: the following assertion expects two task kinds but receives [].

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in bf6f0a1: the fixture now groups c and d, which start in separate files. The checked-in distributed test passes and verifies both pickle round trips, the commit, and unchanged values.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 2, 2026
@LuciferYang

Copy link
Copy Markdown
Contributor Author

@wjones127 @Xuanwo I reworked this PR against your review comments on #9291 and #8614. Here is how each one is handled now; the description has the details.

Two decisions you may want changed:

  1. A run with both kinds commits two versions, a Rewrite and then a DataReplacement. I didn't want to block this on feat(transaction): implement the Transaction V2 action vocabulary #8644. Once composite transactions land, the two can become one commit.
  2. TaskData.kind is a serde-default field, so an older worker given a repack task ignores it and rewrites the fragment. That's more work but still a valid compaction. The alternative is a breaking task format where old workers fail.

A column_groups name that isn't a top-level column now logs a warning instead of failing the run. Otherwise a group naming a dropped or renamed column would stop every later compaction.

#9606 now carries only the CompactionExecutor part, matching this shape. An earlier version of this branch also patched the conflict resolver so that an in-place Update couldn't bring back a file a concurrent drop_columns had removed. #9219 has since fixed that on main in build_manifest, and repacks no longer commit an Update, so that patch is gone from here.

@lance-gatekeeper lance-gatekeeper Bot added K-approved Latest Gatekeeper recommendation permits acceptance. and removed K-approved Latest Gatekeeper recommendation permits acceptance. labels Oct 2, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 2, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

The distributed repack test now requests a layout change and exercises the task/result round trip. Integrating repacks into compact_files follows the pluggable planner/executor direction requested in #9291. The replacement commit preserves row addresses, deletion vectors, overlays, index coverage, and row lineage while detecting overlapping base changes.

Mixed rewrite/repack runs commit two versions. If the repack phase fails, refresh and replan; the completed rewrite remains committed.

Please mark this PR with the breaking-change label.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 2, 2026
Comment thread protos/transaction.proto
* moved to new files: index coverage, overlays, and row update versions are
* left as they are.
*/
optional bool data_change = 2;

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@wjones127 This is the data_change field you suggested in #8644 r3833888750, and it puts this PR under the format-spec vote. Two questions before I ask for votes:

  1. Do you still want it on the legacy DataReplacement, now that feat(transaction): implement the Transaction V2 action vocabulary #8644 models the same thing as TombstoneFieldData plus AddDataFile with data_change: false? I kept it because, after feat(transaction): implement the Transaction V2 action vocabulary #8644 merges, release builds still need LANCE_ENABLE_UNSTABLE_TRANSACTION_V2 to write V2, so repacks would keep committing as DataReplacement there. If the field isn't in the transaction file, an index build racing a repack of its column has to retry. Once V2 is generally available, repacks can move to a CompositeOperation, and a run that both rewrites and repacks can then commit once.
  2. If you do want it, should this field, the spec text in transaction.md and the DataReplacement changes go to a separate PR for the vote? That PR would be about 600 lines instead of 4,300, and later changes to the compaction code here wouldn't fall under the spec vote.

cc @Xuanwo

@github-actions github-actions Bot added the A-docs Documentation label Oct 3, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 3, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

Repacking through compact_files reduces column-file fan-out while preserving row addresses, deletion vectors, overlays, index coverage, and row lineage. The new Python and Java statistics expose the same manifest counts used by the planner.

For the transaction-format question, retaining DataReplacement with data_change: false follows the maintainer's suggested mechanism and supports value-preserving commits in the current format. Composite transactions remain unmerged.

Mixed rewrite/repack runs commit two versions. If repacking fails, refresh and replan; the completed rewrite remains committed.

Please mark this PR with the breaking-change label.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 3, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 3, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Gate recommendation: approve.

Repacking through compact_files reduces column-file fan-out while preserving row addresses, deletion vectors, overlays, index coverage, and row lineage. The expanded layout statistics expose recorded file sizes, fields per file, and dead-slot ratios in Rust, Python, and Java using manifest metadata only.

For the transaction-format question, retaining DataReplacement with data_change: false follows the maintainer's suggested mechanism and supports value-preserving commits in the current format. Composite transactions remain unmerged.

Mixed rewrite/repack runs commit two versions. If repacking fails, refresh and replan; the completed rewrite remains committed.

Please mark this PR with the breaking-change label.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Oct 3, 2026
@wjones127
wjones127 self-requested a review October 4, 2026 04:31

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

A-docs Documentation A-format On-disk format: protos and format spec docs A-java Java bindings + JNI A-python Python bindings enhancement New feature or request format-change A change to the format spec, which requires a vote. Remove if minor (e.g. fixing typo). K-approved Latest Gatekeeper recommendation permits acceptance.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant