fix: columns tracked by name instead of Iceberg field id - #87
Merged
Merged
Conversation
- A restart registered each table with its discovered schema, whose field ids are column positions: after a column was dropped, the ones behind it took others' ids and were written under the wrong Iceberg column. An existing table now keeps its Iceberg schema; the stream's Relation messages evolve it, in order. - A column dropped and re-added under the same name kept the old field id, so the dropped values came back. A re-add — seen as the column returning after it was seen dropped, or out of the order Postgres appends re-added columns in — now renames the old column to `<name>__dropped_<field id>`, values intact, and adds a new column. - pg2iceberg's own Parquet reader (compaction, TOAST resolution, FileIndex rebuild, verify) matched columns by name. It now matches by field id, as Iceberg does, and reads files written before an `int`→`long` / `float`→`double` promotion as the new type. - `oid` maps to `long` (it's unsigned 32-bit; values past 2^31 wrapped negative). Existing `int` columns are promoted by the next Relation. Rows staged before a re-add but materialized after it still land in the new column (schema changes apply ahead of rows already staged); pinned as a failing DST test for a follow-up that applies schema changes in log order. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes
restart_after_dropping_a_column_keeps_field_ids,re_added_column_does_not_bring_back_dropped_valuesandproduction_keeps_large_oids.<name>__dropped_<field id>, with its values intact, and adds a new column. A re-add is recognized in two ways: the column returns after it was seen dropped, or it arrives out of order. Postgres appends a re-added column after every surviving column and never otherwise reorders columns. One case can't be detected: a column re-added while it was already last, with no change to the table in between, looks exactly like a column that was never dropped.verify. After a rename, it would have copied dropped values into the re-added column. It now matches by field id, as Iceberg does. It also reads files written before anint→longorfloat→doublepromotion as the new type; before, it failed on them.oid→long.oidis unsigned 32-bit, so values above 2³¹ wrapped negative. Existingintoid columns are promoted by the next Relation message.