Skip to content

fix: a mid-stream backfill overwrites newer changes - #89

Merged
hasyimibhar merged 1 commit into
mainfrom
fix/backfill-order
Oct 5, 2026
Merged

hasyimibhar merged 1 commit into
mainfrom
fix/backfill-order

Conversation

@hasyimibhar

@hasyimibhar hasyimibhar commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

This PR fixes backfill_does_not_overwrite_newer_changes.

A table added mid-stream has its backfill staged into its log alongside its CDC, in no useful order. A change streamed before the backfill reached its row was staged ahead of the row's older snapshot copy, and applied in log order the old copy won. A row deleted during the backfill came back. An update leaving a TOASTed column unchanged had no row to resolve that column from.

Such a table's snapshot rows are now applied first, then its changes:

  • Marker. register_table_pending creates a second mat_cursor row for the table, <group>#snapshot. It marks, durably, that the table's snapshot rows go first.
  • Snapshot rows first. Once the backfill is complete, the materializer folds the log's snapshot rows in bounded steps and commits them together. The table goes from empty to its snapshot at once, never half loaded. Only then does it mark them applied. A crash before that re-applies them, which changes nothing.
  • Changes after. The table's changes then apply in log order. Snapshot rows are skipped, and so are changes committed before the first snapshot's LSN, which every snapshot row already reflects. A backfill resumed after a restart takes a newer LSN, which is why the first one counts.
  • Completeness check. Snapshot rows wait for the coordinator to report the backfill complete, even if the table is ungated early. Otherwise rows staged later would be skipped as already applied.

A table added mid-stream has its backfill staged into its log alongside
its CDC, in no useful order: a change streamed before the backfill
reached its row was staged ahead of the row's older snapshot copy, and
applied in log order the old copy won. A row deleted during the backfill
came back, and an update leaving a TOASTed column unchanged had no row to
resolve it from.

Such a table's snapshot rows are now applied first, then its changes:

- `register_table_pending` creates a second cursor for the table
  (`<group>#snapshot`), marking that its snapshot rows go first.
- Once the backfill is complete, the materializer folds the log's
  snapshot rows in bounded steps and commits them together — the table
  goes from empty to its snapshot at once — then marks them applied.
- Its changes then apply in log order, skipping the snapshot rows and
  any change committed before the first snapshot's LSN (the snapshot
  already reflects those).

`SNAPSHOT_XID_BASE` moves to pg2iceberg-core (re-exported from
pg2iceberg-snapshot) so the materializer can tell snapshot rows apart.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@hasyimibhar
hasyimibhar merged commit 5ac3dfa into main Oct 5, 2026
2 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant