Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion api/v1alpha1/shard_types.go
Original file line number Diff line number Diff line change
Expand Up @@ -326,7 +326,7 @@ type ShardStatus struct {
// +kubebuilder:validation:Enum=full;diff;incr
LastBackupType string `json:"lastBackupType,omitempty"`

// PodRoles maps pod names to their database roles (e.g. PRIMARY, REPLICA, DRAINED).
// PodRoles maps pod names to their database roles (e.g. PRIMARY, REPLICA, QUARANTINED).
// +optional
PodRoles map[string]string `json:"podRoles,omitempty"`

Expand Down
2 changes: 1 addition & 1 deletion config/crd/bases/multigres.com_shards.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -3042,7 +3042,7 @@ spec:
additionalProperties:
type: string
description: PodRoles maps pod names to their database roles (e.g.
PRIMARY, REPLICA, DRAINED).
PRIMARY, REPLICA, QUARANTINED).
type: object
poolsReady:
description: PoolsReady indicates if all data pools are ready.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -1540,6 +1540,7 @@ spec:
* **2026-03-20:** Added `PostgresConfigRef` type (ConfigMap reference) at the shard level for custom `postgresql.conf` parameters. Operator mounts the user-provided ConfigMap and sets `POSTGRES_INITDB_EXTRA_CONF` on pgctld so the user-supplied lines are appended to pgctld's auto-tuned config. See design rationale below.
* **2026-03-20:** Added `ExternalGatewayConfig` (`spec.externalGateway`) for external exposure of the global multigateway Service via `externalIPs` and user-provided annotations. Added `GatewayStatus` (`status.gateway.externalEndpoint`) and `GatewayExternalReady` condition. Added `InitdbArgs` field to `ShardTemplateSpec`, `ShardInlineSpec`, and `ShardOverrides` for passing extra arguments to `initdb` during PostgreSQL data directory initialization. Added `IPAddress` and `InitdbArgs` validated types to `common_types.go`.
* **2026-03-20:** Added `ExternalAdminWebConfig` (`spec.externalAdminWeb`) for external exposure of the global multiadmin-web Service via `externalIPs` and user-provided annotations. Mirrors the `ExternalGatewayConfig` pattern. Added `AdminWebStatus` (`status.adminWeb.externalEndpoint`) and `AdminWebExternalReady` condition. Readiness is driven by the admin-web Deployment's `ReadyReplicas` (unlike the gateway which aggregates across Cell CRs).
* **2026-09-25:** Removed the dormant DRAINED stand-in machinery (superseding the 2026-03-17 entry). Multigres no longer emits the `DRAINED` topology role — poolers whose postgres cannot start are surfaced as `QUARANTINED` and remediated in place (delete pod + wipe data PVC + re-bootstrap from backup), so the stand-in-replica model and the `multigres.com/role=DRAINED` PVC handling were dead code. `PodRoles` values are now `PRIMARY`, `REPLICA`, or `QUARANTINED`.

## Design Rationale: PostgreSQL Configuration (`postgresConfigRef`)

Expand Down
8 changes: 4 additions & 4 deletions docs/development/pod-management-architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -136,7 +136,7 @@ PVC lifecycle is managed through **conditional owner references** based on the `
- When `WhenDeleted` is `Delete`: PVCs are created with an ownerRef pointing to the Shard CR, enabling Kubernetes garbage collection to cascade-delete them when the Shard is removed.
- When `WhenDeleted` is `Retain`: PVCs are created without ownerRefs, ensuring they persist after Shard deletion.
- The shard controller's `reconcilePVCOwnerRefs` function ensures existing PVCs stay in sync with the current policy — adding or removing ownerRefs as the policy changes mid-lifecycle.
- During scale-down, `cleanupDrainedPod` checks `WhenScaled` and deletes data PVCs directly if the policy is `Delete` (the default). For DRAINED pods (identified by the `multigres.com/role=DRAINED` label), PVCs are always deleted regardless of the `WhenScaled` policy because DRAINED pod data is known-bad.
- During scale-down, `cleanupDrainedPod` checks `WhenScaled`: for a scaled-down pod (index >= replicas) under the `Delete` policy (the default) it orphans the data PVC for garbage collection, while a rolling-update pod (index < replicas) keeps its PVC.

---

Expand Down Expand Up @@ -408,7 +408,7 @@ Metrics are emitted per pool via `monitoring.SetShardPoolReplicas()`. A `PoolEmp
| **Etcd topology cleanup** | `UnregisterMultiPooler` called during drain flow; stale entries removed on pod termination |
| **Topology registration & pruning** | Cell and database registration centralized in MultigresCluster controller; stale entries pruned when `topologyPruning.enabled` (default) |
| **Backup health reporting** | Shard controller calls `GetBackups` RPC, sets `BackupHealthy` condition and `LastBackupTime` status |
| **DRAINED pod handling** | DRAINED pods (diverged data, pg_rewind failure) are kept alive for admin investigation. Stand-in replicas created at next index for availability. Admin discards via `kubectl delete pod`, triggering drain + PVC deletion |
| **QUARANTINED pod handling** | QUARANTINED pods (postgres cannot start) are remediated in place: the operator deletes the pod, wipes its data PVC, and re-bootstraps from backup at the same index |
| **Shard-wide drain serialization** | Planned drains check all persisted drain/termination state across reconciles and wait for data-plane recovery before the next removal |
| **Two-cell maintenance surge** | Creates and verifies temporary same-cell capacity before disrupting the final ready member of a two-cell cross-cell shard |
| **Scale-down health gate** | Drains deferred when either the current pool or another shard pool/cell has non-ready pods |
Expand Down Expand Up @@ -523,7 +523,7 @@ Pods reference this via `spec.subdomain`, combined with `spec.hostname` (set to
### Etcd Topology

The operator reads from etcd topology (via the shard controller) to:
- Determine pod roles (`PRIMARY`, `REPLICA`, `DRAINED`) for scale-down and rolling-update decisions.
- Determine pod roles (`PRIMARY`, `REPLICA`, `QUARANTINED`) for scale-down and rolling-update decisions.
- Clean up stale topology entries on permanent pod removal (`UnregisterMultiPooler`).

The operator writes to etcd topology (via the MultigresCluster controller) to:
Expand Down Expand Up @@ -559,7 +559,7 @@ Through the shard controller:
| Scale-up (new replicas) | Create PVC + Pod, multigres handles the rest |
| Scale-down | Drain state machine with standby removal + etcd unregistration |
| Rolling update | Spec-hash detection, ordered recreation |
| DRAINED pod handling | Detects DRAINED role from etcd, keeps pod alive for investigation, creates stand-in replica |
| QUARANTINED pod handling | Detects QUARANTINED role from etcd, remediates in place (delete pod + wipe data PVC + re-bootstrap from backup) |
| Backup health reporting | Calls `GetBackups` RPC, sets `BackupHealthy` condition |
| PVC lifecycle | Direct creation/deletion per policy |
| Certificate provisioning | `pkg/cert` for pgBackRest TLS |
Expand Down
3 changes: 3 additions & 0 deletions docs/development/pod-management-design.md
Original file line number Diff line number Diff line change
Expand Up @@ -529,6 +529,9 @@ When scaling down, the operator must choose which pod to delete. The selection a

### DRAINED Pooler Replacement

> [!WARNING]
> **Obsolete — needs rewrite.** This section (and the other `DRAINED` references in this design doc) describes the retired "stand-in replica" model. Multigres no longer emits the `DRAINED` topology role (`PoolerType.DRAINED` is deprecated and never produced); poolers whose postgres cannot start are surfaced as `QUARANTINED` and remediated in place (delete pod + wipe data PVC + re-bootstrap from backup). The stand-in machinery has been removed from the operator. This narrative should be rewritten around quarantine remediation; it is left in place here pending that rewrite.

Multiorch may mark a pooler as `DRAINED` in etcd independently of the operator (e.g., during internal recovery or rebalancing). When this happens, the operator must:

1. **Detect the DRAINED pooler** by reading etcd topology during reconciliation and comparing `PoolerType` for each pod's `service-id`
Expand Down
2 changes: 1 addition & 1 deletion docs/development/pvc-lifecycle.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,7 @@ sts.Spec.PersistentVolumeClaimRetentionPolicy = pvc.BuildRetentionPolicy(
```

**Shard Pool Pods** (`pkg/resource-handler/controller/shard/reconcile_pool_pods.go`):
Pool pod PVCs are managed directly by the operator. During scale-down, the `cleanupDrainedPod` function checks `shard.Spec.PVCDeletionPolicy.WhenScaled` and deletes the data PVC if the policy is `Delete` (the default). For DRAINED pods (detected via the `multigres.com/role=DRAINED` label), the PVC is always deleted regardless of the `WhenScaled` policy because DRAINED pod data is known-bad. During shard/cluster deletion, PVCs are garbage-collected by Kubernetes via conditional owner references: when `WhenDeleted` is `Delete`, PVCs are created with an ownerRef to the Shard CR, enabling cascade deletion. When `WhenDeleted` is `Retain`, PVCs have no ownerRef and persist after deletion. The `reconcilePVCOwnerRefs` function ensures existing PVCs stay in sync with the current policy during mid-lifecycle changes.
Pool pod PVCs are managed directly by the operator. During scale-down, the `cleanupDrainedPod` function checks `shard.Spec.PVCDeletionPolicy.WhenScaled` and, for a scaled-down pod (index >= replicas) under the `Delete` policy (the default), orphans the data PVC for garbage collection; a rolling-update pod (index < replicas) keeps its PVC. During shard/cluster deletion, PVCs are garbage-collected by Kubernetes via conditional owner references: when `WhenDeleted` is `Delete`, PVCs are created with an ownerRef to the Shard CR, enabling cascade deletion. When `WhenDeleted` is `Retain`, PVCs have no ownerRef and persist after deletion. The `reconcilePVCOwnerRefs` function ensures existing PVCs stay in sync with the current policy during mid-lifecycle changes.

**Utility Function** (`pkg/util/pvc/retention.go:BuildRetentionPolicy`):
- Used by TopoServer StatefulSets to convert operator's `PVCDeletionPolicy` to Kubernetes `StatefulSetPersistentVolumeClaimRetentionPolicy`
Expand Down
9 changes: 5 additions & 4 deletions docs/operator-capability-levels.md
Original file line number Diff line number Diff line change
Expand Up @@ -142,7 +142,7 @@ The operator continuously updates status on all CRs:
- **MultigresCluster**: Phase (Healthy/Progressing/Error/Deleting), conditions
(Available, Progressing), per-cell and per-database status summaries
- **Shard**: Phase, conditions, per-cell pool status, PodRoles map
(PRIMARY/REPLICA/DRAINED), LastBackupTime, LastBackupType
(PRIMARY/REPLICA/QUARANTINED), LastBackupTime, LastBackupType
- **Cell/TableGroup/TopoServer**: Phase, conditions, ObservedGeneration
- All CRs expose Phase, Available, and Age via `kubectl get` print columns

Expand Down Expand Up @@ -293,7 +293,7 @@ Two-dimensional PVC deletion policy:
| Policy | Retain (default) | Delete |
|:-------|:-----------------|:-------|
| **WhenDeleted** | PVCs kept after Shard deletion | PVCs removed with Shard |
| **WhenScaled** | PVCs kept after scale-down (explicit `Retain`) | PVCs removed on scale-down (default); always deleted for DRAINED pods |
| **WhenScaled** | PVCs kept after scale-down (explicit `Retain`) | PVCs orphaned for garbage collection on scale-down (default) |

---

Expand Down Expand Up @@ -411,8 +411,9 @@ The operator continuously reconciles desired state and recovers from failures:
entries are missing
- **PodRoles refresh**: Continuously updated from topology server to reflect
actual database roles
- **DRAINED pod handling**: Detects DRAINED role from etcd, keeps pod alive
for admin investigation, creates stand-in replica for availability
- **QUARANTINED pod handling**: Detects QUARANTINED role from etcd (a pooler
whose postgres cannot start) and remediates in place — deletes the pod, wipes
its data PVC, and re-bootstraps from backup at the same index
- **Scale-down safety**: Blocks scale-down when pool is already degraded

### Auto-healing (upstream Multigres)
Expand Down
3 changes: 1 addition & 2 deletions pkg/data-handler/topo/pooler.go
Original file line number Diff line number Diff line change
Expand Up @@ -57,8 +57,7 @@ func GetPoolerStatus(
}
// Quarantined poolers are unrecoverable (postgres cannot start).
// They are surfaced with a distinct QUARANTINED role so they are
// visible in Shard.Status.PodRoles, but they are not routed and do
// not drive the stand-in-replica path (which keyed on DRAINED).
// visible in Shard.Status.PodRoles, but they are not routed.
// The operator replaces them via quarantine remediation (delete pod
// + wipe data PVC + re-bootstrap from backup); GetQuarantinedPods
// carries the reason for that.
Expand Down
2 changes: 1 addition & 1 deletion pkg/resource-handler/controller/shard/disruption.go
Original file line number Diff line number Diff line change
Expand Up @@ -144,7 +144,7 @@ func (r *ShardReconciler) selectShardScaleDownPod(
for poolName, pool := range shard.Spec.Pools {
for _, cell := range pool.Cells {
group := groups[string(poolName)+"/"+string(cell)]
replicas := poolReplicas(pool) + countDrainedPods(shard, group)
replicas := poolReplicas(pool)
for _, pod := range group {
index, ok := resolvePodIndex(pod.Name)
if ok && index >= int(replicas) && !isMaintenanceSurge(pod) {
Expand Down
15 changes: 2 additions & 13 deletions pkg/resource-handler/controller/shard/drain_helpers.go
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,8 @@ import (
"github.com/multigres/multigres-operator/pkg/util/metadata"
)

// resolvePodRole returns the role (e.g. "PRIMARY", "REPLICA", "DRAINED") for a
// pod by checking shard.Status.PodRoles. It checks both the exact pod name and
// resolvePodRole returns the role (e.g. "PRIMARY", "REPLICA", "QUARANTINED") for
// a pod by checking shard.Status.PodRoles. It checks both the exact pod name and
// FQDN prefix (podName.subdomain...) since the data-handler may store either.
func resolvePodRole(shard *multigresv1alpha1.Shard, podName string) string {
if shard.Status.PodRoles == nil {
Expand All @@ -31,17 +31,6 @@ func resolvePodRole(shard *multigresv1alpha1.Shard, podName string) string {
return ""
}

// countDrainedPods returns the number of pods whose topology role is DRAINED.
func countDrainedPods(shard *multigresv1alpha1.Shard, existingPods map[string]*corev1.Pod) int32 {
var count int32
for _, pod := range existingPods {
if resolvePodRole(shard, pod.Name) == "DRAINED" {
count++
}
}
return count
}

// clearDrainAnnotations removes all drain annotations from a pod via merge patch,
// cancelling a drain that is no longer needed (e.g. scale-down reversed).
func clearDrainAnnotations(ctx context.Context, k8sClient client.Client, pod *corev1.Pod) error {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -99,9 +99,6 @@ func (r *ShardReconciler) reconcileCellMaintenanceSurge(
baseUnsettled = true
continue
}
if resolvePodRole(shard, pod.Name) == "DRAINED" {
continue
}
stable := isAvailablePooler(pod) &&
pod.Annotations[metadata.AnnotationDrainState] == ""
if !stable {
Expand Down Expand Up @@ -220,7 +217,7 @@ func (r *ShardReconciler) createOrAdoptMaintenanceSurge(
replicas int32,
) error {
logger := log.FromContext(ctx)
index := replicas + countDrainedPods(shard, existingPods)
index := replicas
podName := BuildPoolPodName(shard, poolName, cellName, int(index))
pvcName := BuildPoolDataPVCName(shard, poolName, cellName, int(index))

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -586,14 +586,6 @@ func (r *ShardReconciler) isDrainStale(
return false
}

// A drain on a DRAINED pod comes from external deletion (kubectl delete),
// which also sets DeletionTimestamp (handled above). If we somehow reach
// here with a DRAINED pod in requested state without a DeletionTimestamp,
// the drain should still complete — it should never be cancelled.
if resolvePodRole(shard, pod.Name) == "DRAINED" {
return false
}

poolName := pod.Labels[metadata.LabelMultigresPool]
cellName := pod.Labels[metadata.LabelMultigresCell]
if poolName == "" || cellName == "" {
Expand Down
Loading
Loading