Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions content/docs/arc-enterprise/configuration/clustering.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,8 @@ A node is considered ready when **all** of the following are true:
3. No catch-up-batch pulls failed after retries (`catchup_failed == 0`).
4. No catch-up-batch pulls were dropped due to queue saturation (`catchup_dropped == 0`).

Since 26.09.2 the walker also checks the files this node originated instead of assuming it still holds them: a node that comes back with an empty data disk under the same `cluster.node_id` pulls its own files back from a peer's replica as part of the same batch, so the gate covers them too. On such a cluster the reconciler's orphan-manifest sweep waits for this convergence before it proposes anything and reports a held sweep as `manifest_sweep_held` in the run ([#959](https://github.com/Basekick-Labs/arc/issues/959)).

Failures and drops outside the catch-up window do **not** keep the gate red. They're operational concerns surfaced via puller stats but not correctness blockers — by the time the catch-up batch has settled, the reader has reconciled its view of the manifest as of walker start. Steady-state failures are handled by reactive FSM callbacks (which re-enqueue), the Phase 5 reconciler, and operator alerting via the cumulative `failed` / `dropped` counters.

**Self-heal**: catch-up failures and drops both clear without a process restart, in two ways. When a later pull succeeds for a previously-affected path (a reactive FSM callback re-enqueueing after the underlying issue resolves, or a subsequent catch-up scan), the corresponding scoped counter decrements and the gate re-opens automatically. And when the entry is removed from the cluster manifest, whether by retention, compaction, the reconciliation sweep, or an operator, the reader forgets it: a recorded failure or drop is cleared, a pull still queued or in flight for it is dropped from the catch-up batch, and the gate re-opens at once (since 26.09.2, [#759](https://github.com/Basekick-Labs/arc/issues/759)). That second path is the remedy for an entry no peer can serve, for example a file that was lost on every node, which no later pull will ever fix: the reconciliation sweep on the origin writer (`reconciliation.enabled`, or `POST /api/v1/reconciliation/trigger?dry_run=false&act=true`) removes such entries. A key no backend can address has no automated remover yet (retention cannot read it and the sweep only reports it); that class waits on the operator delete endpoint tracked in [#794](https://github.com/Basekick-Labs/arc/issues/794). The puller tracks affected paths in dedicated sets so it can attribute either event back to the original failure or drop. Pulls abandoned because their entry left the manifest are counted in `skipped_gone`, in the 503 body and in `replication_catchup_status` on `/api/v1/cluster`.
Expand All @@ -94,6 +96,8 @@ Both `catchup_failed` and `catchup_dropped` are surfaced in the 503 body so oper

<Callout type="warn" title="Combining with `replication_catchup_enabled=false`">
If you set `cluster.replication_catchup_enabled=false` (the emergency off-switch for pathologically large manifests), the catch-up walker never runs and the gate would never clear. Arc detects this combination at startup, logs a `WARN`, and **auto-disables the gate** so the node isn't permanently 503'd. Operators see a clear log line and can fix the configuration at their leisure. Don't enable the gate if you've also disabled the walker.

The same switch turns off the own-file re-pull and the reconciler's manifest-sweep hold: a node restored with an empty data disk does not pull back the files it originated, and its reconciliation runs are not held while it is missing them. Keep reconciliation in dry run after such a restore, or re-enable the walker first.
</Callout>

### Endpoints affected
Expand Down