-
Notifications
You must be signed in to change notification settings - Fork 34
fix(topology): report quorum health and failover readiness #681
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Draft
Verolop
wants to merge
2
commits into
multigres:main
Choose a base branch
from
Verolop:fix/topology-health
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Draft
Changes from all commits
Commits
Show all changes
2 commits
Select commit
Hold shift + click to select a range
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.
Oops, something went wrong.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -22,6 +22,9 @@ spec: | |
| Controller {{ $labels.controller }} has had a non-zero reconcile | ||
| error rate for {{ $labels.namespace }}/{{ $labels.name }} over the | ||
| last 5 minutes (current: {{ $value | humanize }}/s). | ||
| Inspect TopologyReady for topology errors and FailoverReady for their | ||
| impact on automatic failover; MultigresFailoverUnavailable reports | ||
| these failures separately at critical severity. | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/MultigresClusterReconcileErrors.md" | ||
|
|
||
| - alert: MultigresClusterDegraded | ||
|
|
@@ -78,6 +81,126 @@ spec: | |
| (current: {{ $value | humanize }}/s). | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/MultigresWebhookErrors.md" | ||
|
|
||
| # ── Topology and failover ────────────────────────────────── | ||
|
|
||
| - alert: MultigresTopologyQuorumUnavailable | ||
| expr: multigres_operator_toposerver_quorum_available == 0 | ||
| for: 1m | ||
| labels: | ||
| severity: critical | ||
| annotations: | ||
| summary: "Topology unavailable for failover in {{ $labels.target_namespace }}/{{ $labels.name }}" | ||
| description: >- | ||
| {{ $labels.reason }}: the operator detected an etcd status error or cannot confirm quorum for one consistent etcd cluster. | ||
| Cluster {{ $labels.cluster }} has lost verified failover protection; | ||
| existing SQL connections may still serve traffic. | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| - alert: MultigresTopologyHealthUnknown | ||
| expr: >- | ||
| (multigres_operator_toposerver_quorum_available < 0) | ||
| or (time() - multigres_operator_toposerver_health_checked_timestamp_seconds > 120) | ||
| for: 1m | ||
| labels: | ||
| severity: warning | ||
| annotations: | ||
| summary: "Topology health observation missing for {{ $labels.target_namespace }}/{{ $labels.name }}" | ||
| description: "Etcd quorum checks are failing or stale. Failover protection cannot be verified for cluster {{ $labels.cluster }}." | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| - alert: MultigresTopologyMemberUnavailable | ||
| expr: multigres_operator_toposerver_member_up == 0 | ||
| for: 2m | ||
| labels: | ||
| severity: warning | ||
| annotations: | ||
| summary: "Etcd member {{ $labels.member }} is unavailable" | ||
| description: >- | ||
| Member {{ $labels.target_namespace }}/{{ $labels.member }} cannot complete | ||
| a linearizable read. Check the quorum alert before restarting other members. | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| - alert: MultigresTopologyBackendNearQuota | ||
| expr: >- | ||
| multigres_operator_toposerver_backend_bytes | ||
| / multigres_operator_toposerver_backend_quota_bytes > 0.8 | ||
| for: 5m | ||
| labels: | ||
| severity: warning | ||
| annotations: | ||
| summary: "Etcd backend nearing quota on {{ $labels.member }}" | ||
| description: >- | ||
| Backend usage on {{ $labels.target_namespace }}/{{ $labels.member }} | ||
| is {{ $value | humanizePercentage }} of its quota. Exhausting the quota | ||
| blocks topology writes and puts failover protection at risk. | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| - alert: MultigresTopologyMemoryPressure | ||
| expr: >- | ||
| label_replace(label_replace( | ||
| max by (namespace, pod) (container_memory_working_set_bytes{container="etcd", pod!=""}), | ||
| "target_namespace", "$1", "namespace", "(.+)"), | ||
| "member", "$1", "pod", "(.+)") | ||
| / on (target_namespace, member) group_left (cluster, name) | ||
| (max by (cluster, name, target_namespace, member) | ||
| (multigres_operator_toposerver_memory_limit_bytes) > 0) | ||
| > 0.8 | ||
| for: 1m | ||
| labels: | ||
| severity: warning | ||
| annotations: | ||
| summary: "Etcd memory pressure on {{ $labels.member }}" | ||
| description: >- | ||
| Member {{ $labels.target_namespace }}/{{ $labels.member }} is using | ||
| {{ $value | humanizePercentage }} of its memory limit. An OOM kill | ||
| can remove a quorum member and interrupt failover protection for {{ $labels.cluster }}. | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| - alert: MultigresTopologyMemberOOMKilled | ||
| expr: >- | ||
| (time() - multigres_operator_toposerver_member_oom_timestamp_seconds < 300) | ||
| and (multigres_operator_toposerver_member_oom_timestamp_seconds > 0) | ||
| labels: | ||
| severity: critical | ||
| annotations: | ||
| summary: "Etcd member {{ $labels.member }} was OOM-killed" | ||
| description: "An OOM kill removed member {{ $labels.target_namespace }}/{{ $labels.member }}. Check quorum and FailoverReady for cluster {{ $labels.cluster }} before restarting another member." | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| - alert: MultigresTopologyMemberRestarting | ||
| expr: increase(multigres_operator_toposerver_member_restarts_total[10m]) >= 3 | ||
| labels: | ||
| severity: warning | ||
| annotations: | ||
| summary: "Etcd member {{ $labels.member }} is repeatedly restarting" | ||
| description: "Member {{ $labels.target_namespace }}/{{ $labels.member }} restarted at least three times in ten minutes. Check termination reasons, memory pressure, and quorum health." | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| - alert: MultigresFailoverUnavailable | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I think this one, and others will trigger for every new cluster before they become ready? |
||
| expr: multigres_operator_cluster_failover_ready == 0 | ||
| for: 1m | ||
| labels: | ||
| severity: critical | ||
| annotations: | ||
| summary: "Failover protection unavailable for {{ $labels.target_namespace }}/{{ $labels.cluster }}" | ||
| description: >- | ||
| FailoverReady is false: {{ $labels.reason }}. Topology access or shard | ||
| orchestrators are unavailable. SQL availability does not imply that | ||
| automatic failover is working. | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| - alert: MultigresFailoverHealthUnknown | ||
| expr: >- | ||
| (multigres_operator_cluster_failover_ready < 0) | ||
| or (time() - multigres_operator_cluster_failover_checked_timestamp_seconds > 120) | ||
| for: 2m | ||
| labels: | ||
| severity: warning | ||
| annotations: | ||
| summary: "Failover protection unverified for {{ $labels.target_namespace }}/{{ $labels.cluster }}" | ||
| description: "Failover observations are unknown or stale. Inspect TopologyQuorumAvailable, TopologyReady, and shard OrchReady conditions." | ||
| runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md" | ||
|
|
||
| # ── Backup & Drain ──────────────────────────────────────── | ||
|
|
||
| - alert: MultigresBackupStale | ||
|
|
@@ -138,7 +261,7 @@ spec: | |
| # ── Saturation ──────────────────────────────────────────── | ||
|
|
||
| - alert: MultigresControllerSaturated | ||
| expr: workqueue_depth{name=~"multigrescluster|tablegroup|shard|cell|toposerver"} > 50 | ||
| expr: workqueue_depth{name=~"multigrescluster|tablegroup|shard|cell|toposerver|toposerver-health"} > 50 | ||
| for: 10m | ||
| labels: | ||
| severity: warning | ||
|
|
||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,142 @@ | ||
| # Topology health and failover protection | ||
|
|
||
| These alerts distinguish topology and failover failures from SQL availability. | ||
| An existing primary may continue serving queries while topology access or | ||
| multiorch is unavailable. `Available=True` alone does not confirm that a new | ||
| primary can be elected. | ||
|
|
||
| ## Conditions and probes | ||
|
|
||
| The operator checks each managed etcd member every 30 seconds with a status RPC | ||
| and a read-only, linearizable Get. The probes run independently of resource | ||
| reconciliation and defragmentation, with a five-second deadline. A successful | ||
| linearizable read confirms quorum even when the operator cannot reach every | ||
| member. Status RPCs alone do not confirm quorum. A status error from any member | ||
| sets `QuorumAvailable=False` with reason `EtcdStatusError`, even when reads succeed. | ||
| For example, a NOSPACE alarm allows reads but blocks the topology writes required | ||
| during failover. The condition message identifies each member reporting an error. | ||
|
|
||
| `TopoServer.status.conditions[QuorumAvailable]` reports the result. | ||
| `TopologyUnreachable` means no member answered the operator; it does not prove | ||
| that the members have lost quorum among themselves. `QuorumUnavailable` means | ||
| members answered status requests but none completed the read. `ClusterMismatch` | ||
| means the endpoints reported different etcd cluster IDs. Client setup | ||
| failures, including invalid TLS credentials, produce `Unknown` with reason | ||
| `ProbeFailed`. `status.healthCheckedAt` records the latest probe completion. | ||
| The TopoServer `Ready` condition continues to report StatefulSet readiness. | ||
|
|
||
| On a MultigresCluster: | ||
|
|
||
| - `TopologyReady` reports the result of topology registration, including failures | ||
| during the startup grace period. | ||
| - `TopologyQuorumAvailable` aggregates managed global and cell-local topology | ||
| servers. Missing servers report false; observations older than two minutes or | ||
| from an earlier generation report unknown. | ||
| - `FailoverReady` requires successful topology access, current managed quorum | ||
| observations, and `OrchReady` for every desired shard at its current generation. | ||
| A missing or unready shard reports false. Unknown observations prevent a true | ||
| result. A previously initialized cluster with lost failover readiness is degraded. | ||
|
|
||
| External topology is not probed directly. Its quorum condition is unknown with | ||
| reason `ExternalTopology`; failover readiness uses topology registration and shard | ||
| orchestrator readiness. Monitor external etcd through its owner. | ||
|
|
||
| ## Alert thresholds | ||
|
|
||
| | Alert | Threshold | Severity | | ||
| | --- | --- | --- | | ||
| | `MultigresTopologyQuorumUnavailable` | Failed quorum check, inconsistent cluster IDs, or etcd status errors for one minute | critical | | ||
| | `MultigresTopologyMemberUnavailable` | A member cannot complete reads for two minutes | warning | | ||
| | `MultigresTopologyHealthUnknown` | Unknown probe or observation over two minutes old, for one minute | warning | | ||
| | `MultigresTopologyBackendNearQuota` | Backend over 80% of quota for five minutes | warning | | ||
| | `MultigresTopologyMemoryPressure` | Working set over 80% of memory limit for one minute | warning | | ||
| | `MultigresTopologyMemberOOMKilled` | An observed OOM termination within the last five minutes | critical | | ||
| | `MultigresTopologyMemberRestarting` | At least three restarts in ten minutes | warning | | ||
| | `MultigresFailoverUnavailable` | FailoverReady false for one minute | critical | | ||
| | `MultigresFailoverHealthUnknown` | Unknown failover readiness or an observation over two minutes old, for two minutes | warning | | ||
|
|
||
| Alert delays are in addition to the probe, scrape, and evaluation intervals. | ||
| The pressure alerts can warn before resource exhaustion; a rapid memory spike | ||
| can reach the limit between scrapes. | ||
|
|
||
| ## Investigate | ||
|
|
||
| Read the conditions and their messages, then inspect the affected members: | ||
|
|
||
| ```bash | ||
| kubectl -n <namespace> get multigrescluster <cluster> -o yaml | ||
| kubectl -n <namespace> get toposervers -l multigres.com/cluster=<cluster> -o yaml | ||
| kubectl -n <namespace> get shards -l multigres.com/cluster=<cluster> -o yaml | ||
| kubectl -n <namespace> describe pod <member> | ||
| kubectl -n <namespace> logs <member> -c etcd --previous | ||
| ``` | ||
|
|
||
| For quorum or connectivity failures, check member logs, headless Service and DNS, | ||
| network reachability from the operator, and TLS Secrets. Check all members before | ||
| restarting one: taking down another voter can turn a single-member failure into | ||
| quorum loss. When topology is healthy but `FailoverReady` names an orchestrator, | ||
| inspect the affected shard's multiorch Deployments and pod readiness. | ||
|
|
||
| For backend pressure, compare total backend bytes with bytes in use. Review | ||
| compaction and defragmentation settings in [Topology maintenance](../../topology-maintenance.md). | ||
| If the condition reports NOSPACE, reclaim backend space before disarming the alarm | ||
| with `etcdctl alarm disarm`. Successful reads alone do not confirm recovery. | ||
| For memory pressure or OOMs, inspect working-set history and the running pod's | ||
| memory limit. Raising the backend quota does not increase available memory. | ||
|
|
||
| ## Metrics and queries | ||
|
|
||
| Topology metrics use `cluster`, `name`, and `target_namespace`. Member metrics also | ||
| include `member` (the pod name). These labels avoid collisions with the operator | ||
| scrape target's own `namespace` and `pod` labels. | ||
|
|
||
| All names below begin with `multigres_operator_`: | ||
|
|
||
| | Suffix | Observation | | ||
| | --- | --- | | ||
| | `toposerver_quorum_available` | 1 true, 0 false, -1 unknown; includes `reason` | | ||
| | `toposerver_health_checked_timestamp_seconds` | Last completed probe | | ||
| | `toposerver_member_up` | Member completed a linearizable read; may remain 1 during a NOSPACE alarm | | ||
| | `toposerver_backend_bytes` | Total backend size | | ||
| | `toposerver_backend_in_use_bytes` | Backend bytes in use | | ||
| | `toposerver_backend_quota_bytes` | Quota configured on the running pod | | ||
| | `toposerver_revision` | Current MVCC revision | | ||
| | `toposerver_memory_limit_bytes` | Running pod's memory limit; zero means unlimited | | ||
| | `toposerver_member_restarts_total` | Observed container restart count, reset on pod replacement | | ||
| | `toposerver_member_oom_timestamp_seconds` | Latest OOM termination still recorded in pod status, or zero | | ||
| | `cluster_failover_ready` | 1 true, 0 false, -1 unknown; labels `cluster`, `target_namespace`, `reason` | | ||
| | `cluster_failover_checked_timestamp_seconds` | Last failover readiness observation | | ||
|
|
||
| Backend and revision series are removed when a member cannot be observed. They | ||
| are not carried forward as current measurements. Kubernetes retains only the | ||
| current and last container termination, so the OOM timestamp is not a complete | ||
| history of OOM events. Restart counts are exported as gauges because they come | ||
| from pod status; `increase` handles their resets. | ||
|
|
||
| Revision growth per second: | ||
|
|
||
| ```promql | ||
| deriv(multigres_operator_toposerver_revision[5m]) | ||
| ``` | ||
|
|
||
| Backend quota utilization: | ||
|
|
||
| ```promql | ||
| multigres_operator_toposerver_backend_bytes | ||
| / multigres_operator_toposerver_backend_quota_bytes | ||
| ``` | ||
|
|
||
| The memory alert requires kubelet/cAdvisor | ||
| `container_memory_working_set_bytes{container="etcd"}` with `namespace` and `pod` | ||
| labels. The local `deploy-observability` overlay configures this scrape using the | ||
| Prometheus service account. Other installations must enable their kubelet scrape. | ||
| A missing scrape or an unlimited container produces no memory-pressure alert; | ||
| check scrape health and configure limits when enabling this alert. The local | ||
| scrape accepts the kubelet's self-signed serving certificate; production scrape | ||
| configuration should use the cluster's trusted kubelet certificates. | ||
|
|
||
| To run the alert evaluation fixtures: | ||
|
|
||
| ```bash | ||
| PROMTOOL=/path/to/promtool go test ./pkg/monitoring -run TestTopologyAlertRules | ||
| ``` |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
What's going to be reading this? We'd be doing a PUT to the api server every 30s for every multigrescluster, so that could become problematic? If we're also publishing metrics, we'd be alerting from the metrics instead of by reading k8s .statuses?
I'd suggest only updating on transitions instead of a timestamp?
If we want a timestamp to make a decision in another controller, we could store an xsync.Map or something like that to store the timestamp per project in memory so other controllers would still be able to use the timestamp without impacting the API server?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
If we need .status.healthCheckAt I'd suggest to keep polling at 30s, then immediately updating on transitions and otherwise only write it to the API server every 5m or something like that?
Also- we'd want to add a predicate on the MultigresCluster's Owns(&TopoServer{}) to ignore updates to the TopoServer when the only thing that changes is the healthCheckedAt field.