Skip to content

Fix source shutdown hang and reconnect at the switchpoint - #197

Merged
serprex merged 3 commits into
ClickHouse:mainfrom
timmclaughlin:fix/source-shutdown-and-switchpoint-crossing
Oct 5, 2026
Merged

serprex merged 3 commits into
ClickHouse:mainfrom
timmclaughlin:fix/source-shutdown-and-switchpoint-crossing

Conversation

@timmclaughlin

@timmclaughlin timmclaughlin commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #194, fixes #195.

#194, source fast shutdown waits on walshadow. send_status repeated the segment-aligned resume floor as flush. WalSndDone only exits once the reported position equals sentPtr, and a shutdown checkpoint ends mid-segment, so the walsender never exits. The postmaster waits until walshadow stops answering: 4m24s in a CloudNativePG in-place restart, and a CNPG switchover runs out switchoverDelay and then immediate-stops the old primary. It's also why the switchover tests here tear walshadow down first.

flush now goes out only when it advances past what the connection last confirmed. Otherwise it is InvalidXLogRecPtr, the way pg_receivewal reports it. PostgreSQL moves a physical slot only on a valid flush (ProcessStandbyReplyMessage), so restart_lsn still never passes the floor. WalSndDone falls back to write, which already tracks the source's end of WAL. Between floor advances, pg_stat_replication.flush_lsn reads NULL for walshadow.

#195, reconnect refuses a stream sitting exactly at the switchpoint. After a clean switchover, or a promotion with nothing in flight, the consumed frontier equals the descendant's switchpoint. The reconnect learns that switchpoint from the live chain before the pump's history does. prove_branch rightly won't resume a branch there, since StartReplication answers a historic start at its switchpoint with the next timeline and no COPY. So the pump fell back to redial forever with consumed frontier X sits past the fork X, and only a restart recovered it.

resume_point now reports that case as AtSwitchpoint, and the pump hands it to the existing crossing, the same way it already does when its own chain shows the branch exhausted. prove_branch and its tests are unchanged, and a frontier past the switchpoint still refuses.

Records past a segment a push seals waited for the next push. WalStream::push drained completed records once per chunk, then sealed segments. The walker stops at a segment end, so a chunk crossing a boundary left the next segment's records buffered. With archive_mode on, shutdown switches segments before its checkpoint, so the source's final chunk carries the XLOG_SWITCH plus the checkpoint that opens the fork segment. The checkpoint never dispatched, the idle ack stayed at the switch, and the fork barrier waited on emitter_ack until a restart. Once the shutdown fix above lets a switchover end cleanly, this is what a CloudNativePG switchover hits. On a quiet source the same gap holds back a commit at a segment's head. push now drains again after each seal.

Tests:

  • wire_flush_reports_only_advances, resume_point_hands_the_switchpoint_to_the_crossing, push_dispatches_records_past_a_segment_it_seals (unit)
  • unpaused_switchover_stops_promptly_and_crosses_at_the_switchpoint (control_plane_e2e): no pause, default wal_sender_timeout, fast-stop the primary, promote, then repoint. On main it fails at the stop (pg_ctl: server does not shut down). Here it passes.
  • archiving_switchover_crosses_a_fork_segment_with_no_transactions (control_plane_e2e): the same switchover with archive_mode = on. Without the third commit it stalls at the fork barrier.

Ran locally on PG 18.6 / ClickHouse 26.9 (arm64): fmt, clippy -D warnings, cargo doc -D warnings, and the control_plane_e2e, source_on_promoted_timeline_e2e, bin_stream_e2e, backfill_gap_across_promotion and source_reconnect suites. Also ran on a CloudNativePG 1.28 cluster under write load (archive_mode on). An in-place primary restart stops in under a second and walshadow reconnects. A planned switchover and a primary kill both cross live with no daemon restart. ClickHouse matches the source after each.

🤖 Generated with Claude Code

@CLAassistant

CLAassistant commented Oct 4, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@serprex
serprex self-requested a review October 5, 2026 12:04
@serprex
serprex force-pushed the fix/source-shutdown-and-switchpoint-crossing branch from bebb448 to 91ebd9a Compare October 5, 2026 12:45
@serprex
serprex marked this pull request as ready for review October 5, 2026 12:45
timmclaughlin and others added 3 commits October 5, 2026 12:57
WalSndDone exits once the reported position equals sentPtr, reading
write when flush is invalid. Repeating the segment-aligned floor never
reaches a mid-segment shutdown checkpoint, so the source's fast shutdown
waited on walshadow until it stopped answering. Report flush only when
it moves past what the connection confirmed, otherwise InvalidXLogRecPtr
as pg_receivewal does. PostgreSQL moves a physical slot only on a valid
flush, so restart_lsn still never passes the floor.

Fixes ClickHouse#194

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
push() drained completed records once per chunk, then sealed finished
segments. The walker stops at a segment end, so records a chunk carried
past the boundary waited for the next push. When that chunk is the
source's last, they never dispatch: with archive_mode on, shutdown
switches segments before its checkpoint, so the checkpoint opens the
fork segment, the idle ack stays at the XLOG_SWITCH, and the fork
barrier waits on emitter_ack until a restart re-reads the segment. On a
quiet source the same gap holds back a commit at a segment's head until
more WAL arrives.

Drain again after each seal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A clean switchover leaves the consumed frontier exactly at the
descendant's switchpoint. Reconnect read that off the live chain before
the pump's history did and refused it as past the fork, so the pump
redialed forever. prove_branch now reports it as AtSwitchpoint and
reconnect returns its connection in simple-query mode, so the crossing
fetches history and proves the slot over it instead of dialing again.
A frontier past the switchpoint still refuses. Refused reconnect proofs
count in the switchover failure metrics alongside refused crossings.

Adds e2e drills of an unpaused switchover under the default
wal_sender_timeout, via repoint, with archiving on, and with the
promoted standby taking over the source address. They cover this, the
flush fix and the segment drain fix.

Fixes ClickHouse#195

Co-Authored-By: serprex <159546+serprex@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@serprex
serprex force-pushed the fix/source-shutdown-and-switchpoint-crossing branch from 91ebd9a to 4057975 Compare October 5, 2026 12:58
@serprex

serprex commented Oct 5, 2026

Copy link
Copy Markdown
Member

@timmclaughlin please sign CLA

@serprex
serprex merged commit efc7f16 into ClickHouse:main Oct 5, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

4 participants