fix: schema changes applied ahead of rows staged before them - #88
Merged
Merged
Conversation
Relation messages evolved the Iceberg schema the moment they arrived, ahead of rows already staged under the old schema and keyed by column name. A column dropped and re-added took those rows' values for the dropped one. Schema changes now go through the change log, in order with the rows: - The pipeline stages a Relation message that changes a table's columns as a schema event (staged op `R`) where the stream put it, starting a new log entry. - The materializer, reaching an entry that starts with one, commits the rows before it, then applies the change against the catalog's current schema. A crash in between re-reads that entry, and applying the change again is a no-op. - The lifecycle no longer applies Relation messages itself (`Materializer::apply_relation` is gone). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #87. Fixes
rows_staged_before_a_re_add_keep_their_values_out_of_the_new_column.Relation messages evolved the Iceberg schema the moment they arrived, ahead of rows already staged under the old schema. Staged rows are keyed by column name, so a column dropped and re-added took those rows' values for the dropped one.
Schema changes now go through the change log, in order with the rows:
R, at the position where the stream put it. The event always starts a new log entry. Unchanged Relations aren't staged; pgoutput resends them at every new session and after every cache invalidation.Materializer::apply_relationis gone. So is the materializer's PG→Iceberg name translation: the log is already keyed by the Iceberg name.Compatibility. The staged format gains op
R, carrying the column list in_data. Binaries older than this one can't decode logs that contain it, so a rollback across this version would fail to materialize them.