Skip to content

fix: schema changes applied ahead of rows staged before them - #88

Merged
hasyimibhar merged 1 commit into
mainfrom
fix/log-ordered-schema
Oct 5, 2026
Merged

hasyimibhar merged 1 commit into
mainfrom
fix/log-ordered-schema

Conversation

@hasyimibhar

@hasyimibhar hasyimibhar commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Follow-up to #87. Fixes rows_staged_before_a_re_add_keep_their_values_out_of_the_new_column.

Relation messages evolved the Iceberg schema the moment they arrived, ahead of rows already staged under the old schema. Staged rows are keyed by column name, so a column dropped and re-added took those rows' values for the dropped one.

Schema changes now go through the change log, in order with the rows:

  • Pipeline. When a Relation message changes a table's columns, the pipeline stages it as a schema event, staged op R, at the position where the stream put it. The event always starts a new log entry. Unchanged Relations aren't staged; pgoutput resends them at every new session and after every cache invalidation.
  • Materializer. When it reaches an entry that starts with a schema event, the materializer first commits the rows before it, then applies the change. It reconciles against the catalog's current schema. A crash in between re-reads that entry, and applying the change again is a no-op. Durable, replay-safe, and correct under distributed workers.
  • Lifecycle. It no longer applies Relation messages itself, and Materializer::apply_relation is gone. So is the materializer's PG→Iceberg name translation: the log is already keyed by the Iceberg name.

Compatibility. The staged format gains op R, carrying the column list in _data. Binaries older than this one can't decode logs that contain it, so a rollback across this version would fail to materialize them.

Relation messages evolved the Iceberg schema the moment they arrived,
ahead of rows already staged under the old schema and keyed by column
name. A column dropped and re-added took those rows' values for the
dropped one.

Schema changes now go through the change log, in order with the rows:

- The pipeline stages a Relation message that changes a table's columns
  as a schema event (staged op `R`) where the stream put it, starting a
  new log entry.
- The materializer, reaching an entry that starts with one, commits the
  rows before it, then applies the change against the catalog's current
  schema. A crash in between re-reads that entry, and applying the
  change again is a no-op.
- The lifecycle no longer applies Relation messages itself
  (`Materializer::apply_relation` is gone).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@hasyimibhar
hasyimibhar merged commit 1376cfe into main Oct 5, 2026
2 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant