Upgrade check: ignore injected query memory faults in the post-upgrade log scan - #123581
Conversation
…e log scan Stress worker 1 (`--database=test_1`) runs with `memory_tracker_fault_probability=0.05`. Work that the stress phase leaves behind keeps the settings of the query that created it and is executed again by the upgraded server after the restart: - a distributed DDL entry stores the initiator's changed settings (`DDLLogEntry`) and `DDLTaskBase::makeQueryContext` applies them, so a queued `BACKUP ... ON CLUSTER` of `test_1` replays with fault injection and `BackupsWorker` logs `Failed to make internal backup ... fault injected` at `<Error>`; - a pending `Distributed` batch is sent with its stored settings, and the receiving async-insert flush logs `AsynchronousInsertQueue: Failed insertion ... fault injected` at `<Error>` with an empty query id, so the existing `} <Error> executeQuery: Code:` entry does not cover it. Seen on three unrelated PRs in 90 days (ClickHouse#123464, ClickHouse#117943, ClickHouse#120725); in each job these lines were the only output of the scan. Nothing enables fault injection on the upgraded server itself, so the message can only come from carried-over stress work. Ignore it as a fixed string, the same string the stress smoke check already tolerates (`ci/jobs/scripts/stress/stress.py`). It is MemoryTracker's message for both injection sites and does not match a real `memory limit exceeded` error. The large DDL backlog in the ClickHouse#123464 job comes from the previous release (26.9) wedging its DDLWorker on `KILL PART_MOVE_TO_SHARD ... ON CLUSTER`, which ClickHouse#122132 fixed on master only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Internal second-model review: adjudication log (click to expand)Pre-publication review: my own cold review plus an independent model (engine: codex; 0 findings). 1 finding total.
Severity: ❌ blocker / Session id: cron:clickhouse-review-slot-48:20261002-153100 |
|
Workflow [PR], commit [cdc39b0] Summary: ✅ AI ReviewSummaryThis PR teaches the upgrade-check log scrubber to ignore Final Verdict
|
CI finish ledger - ea5ff28Every failure below has an owner: a fixing PR (mine or external), or a full-effort fix task
All 149 checks completed, Session id: cron:our-pr-ci-monitor:20261002-190341 |
Changelog category (leave one):
Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md):
Stop the upgrade check from failing when work left behind by the stress phase hits that phase's own memory fault injection after the upgrade restart.
Description
Upgrade check (amd_release)reds onError message in clickhouse-server.logwhen the only line the post-upgrade scan finds is an injectedCode: 241 ... Query memory tracker: fault injected. Three runs in 90 days, all on unrelated PRs, 0 on master; in each job these lines were the whole scan output.Root cause: stress workers 1 and 6 run with
memory_tracker_fault_probability=0.05(worker 1 in the shared databasetest_1). Work they leave behind keeps the settings of the query that created it and runs again on the upgraded server:BACKUP ... ON CLUSTERoftest_1replays with fault injection (BackupsWorker: Failed to make internal backup);Distributedbatch is sent with its stored settings, and the receiving async-insert flush fails (AsynchronousInsertQueue: Failed insertion, empty query id, so the existingexecuteQueryentry misses it).Nothing enables fault injection on the upgraded server itself, so this message can only come from carried-over stress work. The change adds it as one fixed-string scan entry, the same string the stress smoke check already tolerates (
ci/jobs/scripts/stress/stress.py). It is MemoryTracker's message for both injection sites and does not match a realmemory limit exceedederror.Validation on the three jobs' real
clickhouse-server.upgrade.log, with the scan cut from the script itself: base reports 1/1/2 lines, the fix 0, and removing the new entry restores the base output. A real backup failure, memory-limit error and DDL error appended to the log are still reported. 19 green upgrade logs give identical (empty) output.The large DDL backlog in the #123464 job comes from 26.9 wedging its
DDLWorkeronKILL PART_MOVE_TO_SHARD, fixed on master by #122132.The three failing runs
BackupsWorkerAsynchronousInsertQueueAsynchronousInsertQueuex2Related: #122132
Related: #120990
Workflow [PR]
Sync PR [sync-upstream/pr/123581]
Version info
26.10.1.1454-master(included in26.10and later)