Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions api/v1alpha1/toposerver_types.go
Original file line number Diff line number Diff line change
Expand Up @@ -71,6 +71,10 @@ type TopoServerSpec struct {

// TopoServerStatus defines the observed state of TopoServer.
type TopoServerStatus struct {
// HealthCheckedAt is the completion time of the latest direct etcd health probe.
// +optional
HealthCheckedAt *metav1.Time `json:"healthCheckedAt,omitempty"`

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's going to be reading this? We'd be doing a PUT to the api server every 30s for every multigrescluster, so that could become problematic? If we're also publishing metrics, we'd be alerting from the metrics instead of by reading k8s .statuses?

I'd suggest only updating on transitions instead of a timestamp?

If we want a timestamp to make a decision in another controller, we could store an xsync.Map or something like that to store the timestamp per project in memory so other controllers would still be able to use the timestamp without impacting the API server?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we need .status.healthCheckAt I'd suggest to keep polling at 30s, then immediately updating on transitions and otherwise only write it to the API server every 5m or something like that?

Also- we'd want to add a predicate on the MultigresCluster's Owns(&TopoServer{}) to ignore updates to the TopoServer when the only thing that changes is the healthCheckedAt field.


// ObservedGeneration is the most recent generation observed.
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
Expand Down
4 changes: 4 additions & 0 deletions api/v1alpha1/zz_generated.deepcopy.go

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

5 changes: 5 additions & 0 deletions config/crd/bases/multigres.com_toposervers.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1520,6 +1520,11 @@ spec:
- inProgress
- lastAttemptTime
type: object
healthCheckedAt:
description: HealthCheckedAt is the completion time of the latest
direct etcd health probe.
format: date-time
type: string
message:
description: Message provides details about the current phase.
type: string
Expand Down
26 changes: 26 additions & 0 deletions config/deploy-observability/prometheus.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,9 @@ spec:
- exemplar-storage
- native-histograms
enableOTLPReceiver: true
additionalScrapeConfigs:
name: kubelet-cadvisor-scrape
key: scrape-configs.yaml
serviceMonitorSelector:
matchLabels:
app.kubernetes.io/name: multigres-operator
Expand Down Expand Up @@ -125,3 +128,26 @@ spec:
targetPort: 9090
name: http
type: ClusterIP

---
# cAdvisor supplies per-container memory usage for the topology pressure alert.
apiVersion: v1
kind: Secret
metadata:
name: kubelet-cadvisor-scrape
stringData:
scrape-configs.yaml: |
- job_name: kubelet-cadvisor
scheme: https
metrics_path: /metrics/cadvisor
scrape_interval: 30s
kubernetes_sd_configs:
- role: node
authorization:
credentials_file: /var/run/secrets/kubernetes.io/serviceaccount/token
tls_config:
insecure_skip_verify: true
metric_relabel_configs:
- source_labels: [__name__, container]
regex: 'container_memory_working_set_bytes;etcd'
action: keep
125 changes: 124 additions & 1 deletion config/monitoring/prometheus-rules.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,9 @@ spec:
Controller {{ $labels.controller }} has had a non-zero reconcile
error rate for {{ $labels.namespace }}/{{ $labels.name }} over the
last 5 minutes (current: {{ $value | humanize }}/s).
Inspect TopologyReady for topology errors and FailoverReady for their
impact on automatic failover; MultigresFailoverUnavailable reports
these failures separately at critical severity.
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/MultigresClusterReconcileErrors.md"

- alert: MultigresClusterDegraded
Expand Down Expand Up @@ -78,6 +81,126 @@ spec:
(current: {{ $value | humanize }}/s).
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/MultigresWebhookErrors.md"

# ── Topology and failover ──────────────────────────────────

- alert: MultigresTopologyQuorumUnavailable
expr: multigres_operator_toposerver_quorum_available == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Topology unavailable for failover in {{ $labels.target_namespace }}/{{ $labels.name }}"
description: >-
{{ $labels.reason }}: the operator detected an etcd status error or cannot confirm quorum for one consistent etcd cluster.
Cluster {{ $labels.cluster }} has lost verified failover protection;
existing SQL connections may still serve traffic.
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

- alert: MultigresTopologyHealthUnknown
expr: >-
(multigres_operator_toposerver_quorum_available < 0)
or (time() - multigres_operator_toposerver_health_checked_timestamp_seconds > 120)
for: 1m
labels:
severity: warning
annotations:
summary: "Topology health observation missing for {{ $labels.target_namespace }}/{{ $labels.name }}"
description: "Etcd quorum checks are failing or stale. Failover protection cannot be verified for cluster {{ $labels.cluster }}."
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

- alert: MultigresTopologyMemberUnavailable
expr: multigres_operator_toposerver_member_up == 0
for: 2m
labels:
severity: warning
annotations:
summary: "Etcd member {{ $labels.member }} is unavailable"
description: >-
Member {{ $labels.target_namespace }}/{{ $labels.member }} cannot complete
a linearizable read. Check the quorum alert before restarting other members.
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

- alert: MultigresTopologyBackendNearQuota
expr: >-
multigres_operator_toposerver_backend_bytes
/ multigres_operator_toposerver_backend_quota_bytes > 0.8
for: 5m
labels:
severity: warning
annotations:
summary: "Etcd backend nearing quota on {{ $labels.member }}"
description: >-
Backend usage on {{ $labels.target_namespace }}/{{ $labels.member }}
is {{ $value | humanizePercentage }} of its quota. Exhausting the quota
blocks topology writes and puts failover protection at risk.
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

- alert: MultigresTopologyMemoryPressure
expr: >-
label_replace(label_replace(
max by (namespace, pod) (container_memory_working_set_bytes{container="etcd", pod!=""}),
"target_namespace", "$1", "namespace", "(.+)"),
"member", "$1", "pod", "(.+)")
/ on (target_namespace, member) group_left (cluster, name)
(max by (cluster, name, target_namespace, member)
(multigres_operator_toposerver_memory_limit_bytes) > 0)
> 0.8
for: 1m
labels:
severity: warning
annotations:
summary: "Etcd memory pressure on {{ $labels.member }}"
description: >-
Member {{ $labels.target_namespace }}/{{ $labels.member }} is using
{{ $value | humanizePercentage }} of its memory limit. An OOM kill
can remove a quorum member and interrupt failover protection for {{ $labels.cluster }}.
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

- alert: MultigresTopologyMemberOOMKilled
expr: >-
(time() - multigres_operator_toposerver_member_oom_timestamp_seconds < 300)
and (multigres_operator_toposerver_member_oom_timestamp_seconds > 0)
labels:
severity: critical
annotations:
summary: "Etcd member {{ $labels.member }} was OOM-killed"
description: "An OOM kill removed member {{ $labels.target_namespace }}/{{ $labels.member }}. Check quorum and FailoverReady for cluster {{ $labels.cluster }} before restarting another member."
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

- alert: MultigresTopologyMemberRestarting
expr: increase(multigres_operator_toposerver_member_restarts_total[10m]) >= 3
labels:
severity: warning
annotations:
summary: "Etcd member {{ $labels.member }} is repeatedly restarting"
description: "Member {{ $labels.target_namespace }}/{{ $labels.member }} restarted at least three times in ten minutes. Check termination reasons, memory pressure, and quorum health."
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

- alert: MultigresFailoverUnavailable

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this one, and others will trigger for every new cluster before they become ready?

expr: multigres_operator_cluster_failover_ready == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Failover protection unavailable for {{ $labels.target_namespace }}/{{ $labels.cluster }}"
description: >-
FailoverReady is false: {{ $labels.reason }}. Topology access or shard
orchestrators are unavailable. SQL availability does not imply that
automatic failover is working.
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

- alert: MultigresFailoverHealthUnknown
expr: >-
(multigres_operator_cluster_failover_ready < 0)
or (time() - multigres_operator_cluster_failover_checked_timestamp_seconds > 120)
for: 2m
labels:
severity: warning
annotations:
summary: "Failover protection unverified for {{ $labels.target_namespace }}/{{ $labels.cluster }}"
description: "Failover observations are unknown or stale. Inspect TopologyQuorumAvailable, TopologyReady, and shard OrchReady conditions."
runbook_url: "https://github.com/multigres/multigres-operator/blob/main/docs/monitoring/runbooks/TopologyHealth.md"

# ── Backup & Drain ────────────────────────────────────────

- alert: MultigresBackupStale
Expand Down Expand Up @@ -138,7 +261,7 @@ spec:
# ── Saturation ────────────────────────────────────────────

- alert: MultigresControllerSaturated
expr: workqueue_depth{name=~"multigrescluster|tablegroup|shard|cell|toposerver"} > 50
expr: workqueue_depth{name=~"multigrescluster|tablegroup|shard|cell|toposerver|toposerver-health"} > 50
for: 10m
labels:
severity: warning
Expand Down
142 changes: 142 additions & 0 deletions docs/monitoring/runbooks/TopologyHealth.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,142 @@
# Topology health and failover protection

These alerts distinguish topology and failover failures from SQL availability.
An existing primary may continue serving queries while topology access or
multiorch is unavailable. `Available=True` alone does not confirm that a new
primary can be elected.

## Conditions and probes

The operator checks each managed etcd member every 30 seconds with a status RPC
and a read-only, linearizable Get. The probes run independently of resource
reconciliation and defragmentation, with a five-second deadline. A successful
linearizable read confirms quorum even when the operator cannot reach every
member. Status RPCs alone do not confirm quorum. A status error from any member
sets `QuorumAvailable=False` with reason `EtcdStatusError`, even when reads succeed.
For example, a NOSPACE alarm allows reads but blocks the topology writes required
during failover. The condition message identifies each member reporting an error.

`TopoServer.status.conditions[QuorumAvailable]` reports the result.
`TopologyUnreachable` means no member answered the operator; it does not prove
that the members have lost quorum among themselves. `QuorumUnavailable` means
members answered status requests but none completed the read. `ClusterMismatch`
means the endpoints reported different etcd cluster IDs. Client setup
failures, including invalid TLS credentials, produce `Unknown` with reason
`ProbeFailed`. `status.healthCheckedAt` records the latest probe completion.
The TopoServer `Ready` condition continues to report StatefulSet readiness.

On a MultigresCluster:

- `TopologyReady` reports the result of topology registration, including failures
during the startup grace period.
- `TopologyQuorumAvailable` aggregates managed global and cell-local topology
servers. Missing servers report false; observations older than two minutes or
from an earlier generation report unknown.
- `FailoverReady` requires successful topology access, current managed quorum
observations, and `OrchReady` for every desired shard at its current generation.
A missing or unready shard reports false. Unknown observations prevent a true
result. A previously initialized cluster with lost failover readiness is degraded.

External topology is not probed directly. Its quorum condition is unknown with
reason `ExternalTopology`; failover readiness uses topology registration and shard
orchestrator readiness. Monitor external etcd through its owner.

## Alert thresholds

| Alert | Threshold | Severity |
| --- | --- | --- |
| `MultigresTopologyQuorumUnavailable` | Failed quorum check, inconsistent cluster IDs, or etcd status errors for one minute | critical |
| `MultigresTopologyMemberUnavailable` | A member cannot complete reads for two minutes | warning |
| `MultigresTopologyHealthUnknown` | Unknown probe or observation over two minutes old, for one minute | warning |
| `MultigresTopologyBackendNearQuota` | Backend over 80% of quota for five minutes | warning |
| `MultigresTopologyMemoryPressure` | Working set over 80% of memory limit for one minute | warning |
| `MultigresTopologyMemberOOMKilled` | An observed OOM termination within the last five minutes | critical |
| `MultigresTopologyMemberRestarting` | At least three restarts in ten minutes | warning |
| `MultigresFailoverUnavailable` | FailoverReady false for one minute | critical |
| `MultigresFailoverHealthUnknown` | Unknown failover readiness or an observation over two minutes old, for two minutes | warning |

Alert delays are in addition to the probe, scrape, and evaluation intervals.
The pressure alerts can warn before resource exhaustion; a rapid memory spike
can reach the limit between scrapes.

## Investigate

Read the conditions and their messages, then inspect the affected members:

```bash
kubectl -n <namespace> get multigrescluster <cluster> -o yaml
kubectl -n <namespace> get toposervers -l multigres.com/cluster=<cluster> -o yaml
kubectl -n <namespace> get shards -l multigres.com/cluster=<cluster> -o yaml
kubectl -n <namespace> describe pod <member>
kubectl -n <namespace> logs <member> -c etcd --previous
```

For quorum or connectivity failures, check member logs, headless Service and DNS,
network reachability from the operator, and TLS Secrets. Check all members before
restarting one: taking down another voter can turn a single-member failure into
quorum loss. When topology is healthy but `FailoverReady` names an orchestrator,
inspect the affected shard's multiorch Deployments and pod readiness.

For backend pressure, compare total backend bytes with bytes in use. Review
compaction and defragmentation settings in [Topology maintenance](../../topology-maintenance.md).
If the condition reports NOSPACE, reclaim backend space before disarming the alarm
with `etcdctl alarm disarm`. Successful reads alone do not confirm recovery.
For memory pressure or OOMs, inspect working-set history and the running pod's
memory limit. Raising the backend quota does not increase available memory.

## Metrics and queries

Topology metrics use `cluster`, `name`, and `target_namespace`. Member metrics also
include `member` (the pod name). These labels avoid collisions with the operator
scrape target's own `namespace` and `pod` labels.

All names below begin with `multigres_operator_`:

| Suffix | Observation |
| --- | --- |
| `toposerver_quorum_available` | 1 true, 0 false, -1 unknown; includes `reason` |
| `toposerver_health_checked_timestamp_seconds` | Last completed probe |
| `toposerver_member_up` | Member completed a linearizable read; may remain 1 during a NOSPACE alarm |
| `toposerver_backend_bytes` | Total backend size |
| `toposerver_backend_in_use_bytes` | Backend bytes in use |
| `toposerver_backend_quota_bytes` | Quota configured on the running pod |
| `toposerver_revision` | Current MVCC revision |
| `toposerver_memory_limit_bytes` | Running pod's memory limit; zero means unlimited |
| `toposerver_member_restarts_total` | Observed container restart count, reset on pod replacement |
| `toposerver_member_oom_timestamp_seconds` | Latest OOM termination still recorded in pod status, or zero |
| `cluster_failover_ready` | 1 true, 0 false, -1 unknown; labels `cluster`, `target_namespace`, `reason` |
| `cluster_failover_checked_timestamp_seconds` | Last failover readiness observation |

Backend and revision series are removed when a member cannot be observed. They
are not carried forward as current measurements. Kubernetes retains only the
current and last container termination, so the OOM timestamp is not a complete
history of OOM events. Restart counts are exported as gauges because they come
from pod status; `increase` handles their resets.

Revision growth per second:

```promql
deriv(multigres_operator_toposerver_revision[5m])
```

Backend quota utilization:

```promql
multigres_operator_toposerver_backend_bytes
/ multigres_operator_toposerver_backend_quota_bytes
```

The memory alert requires kubelet/cAdvisor
`container_memory_working_set_bytes{container="etcd"}` with `namespace` and `pod`
labels. The local `deploy-observability` overlay configures this scrape using the
Prometheus service account. Other installations must enable their kubelet scrape.
A missing scrape or an unlimited container produces no memory-pressure alert;
check scrape health and configure limits when enabling this alert. The local
scrape accepts the kubelet's self-signed serving certificate; production scrape
configuration should use the cluster's trusted kubelet certificates.

To run the alert evaluation fixtures:

```bash
PROMTOOL=/path/to/promtool go test ./pkg/monitoring -run TestTopologyAlertRules
```
13 changes: 13 additions & 0 deletions docs/observability.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,19 @@ kubectl apply -f config/monitoring/prometheus-rules.yaml

Each alert links to a dedicated runbook with investigation steps, PromQL queries, and remediation actions.

## Topology health

Managed etcd members are probed every 30 seconds. The operator exports quorum,
backend size and quota, MVCC revision, memory limits, restarts, and observed OOM
terminations. `TopologyQuorumAvailable` and `FailoverReady` on MultigresCluster
separate topology and orchestrator health from SQL availability. Quorum loss and
lost failover readiness raise critical alerts after one minute.

The [topology health runbook](monitoring/runbooks/TopologyHealth.md) lists the
metrics, thresholds, condition meanings, and investigation steps. Memory-pressure
alerts require kubelet/cAdvisor metrics; the local observability overlay includes
that scrape.

## Grafana Dashboards

Three Grafana dashboards are included in `config/monitoring/`:
Expand Down
Loading
Loading