Skip to content

fix: columns tracked by name instead of Iceberg field id - #87

Merged
hasyimibhar merged 1 commit into
mainfrom
fix/field-ids
Oct 5, 2026
Merged

hasyimibhar merged 1 commit into
mainfrom
fix/field-ids

Conversation

@hasyimibhar

@hasyimibhar hasyimibhar commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Fixes restart_after_dropping_a_column_keeps_field_ids, re_added_column_does_not_bring_back_dropped_values and production_keeps_large_oids.

  • Restart after a drop. A restart registered each table with its discovered schema, and discovery numbers field ids by column position. After a column was dropped, the columns behind it took other columns' ids and were written under the wrong Iceberg column. An existing table now keeps its Iceberg schema, and the stream's Relation messages evolve it, in order. Discovery's view of the source's current columns doesn't apply yet at startup, because the stream may still replay transactions from before a change.
  • Re-added column. A column dropped and re-added under the same name kept its old field id, so the dropped values came back. A re-add now renames the old column to <name>__dropped_<field id>, with its values intact, and adds a new column. A re-add is recognized in two ways: the column returns after it was seen dropped, or it arrives out of order. Postgres appends a re-added column after every surviving column and never otherwise reorders columns. One case can't be detected: a column re-added while it was already last, with no change to the table in between, looks exactly like a column that was never dropped.
  • Reader by field id. pg2iceberg's own Parquet reader matched columns by name. That reader serves compaction, TOAST resolution, FileIndex rebuilds and verify. After a rename, it would have copied dropped values into the re-added column. It now matches by field id, as Iceberg does. It also reads files written before an int→long or float→double promotion as the new type; before, it failed on them.
  • oid → long. oid is unsigned 32-bit, so values above 2³¹ wrapped negative. Existing int oid columns are promoted by the next Relation message.

- A restart registered each table with its discovered schema, whose
  field ids are column positions: after a column was dropped, the ones
  behind it took others' ids and were written under the wrong Iceberg
  column. An existing table now keeps its Iceberg schema; the stream's
  Relation messages evolve it, in order.
- A column dropped and re-added under the same name kept the old field
  id, so the dropped values came back. A re-add — seen as the column
  returning after it was seen dropped, or out of the order Postgres
  appends re-added columns in — now renames the old column to
  `<name>__dropped_<field id>`, values intact, and adds a new column.
- pg2iceberg's own Parquet reader (compaction, TOAST resolution,
  FileIndex rebuild, verify) matched columns by name. It now matches by
  field id, as Iceberg does, and reads files written before an
  `int`→`long` / `float`→`double` promotion as the new type.
- `oid` maps to `long` (it's unsigned 32-bit; values past 2^31 wrapped
  negative). Existing `int` columns are promoted by the next Relation.

Rows staged before a re-add but materialized after it still land in the
new column (schema changes apply ahead of rows already staged); pinned
as a failing DST test for a follow-up that applies schema changes in log
order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@hasyimibhar
hasyimibhar merged commit a7cff63 into main Oct 5, 2026
2 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant