Repository navigation
feat: explain Jobs, CronJobs, PersistentVolumeClaims, and Nodes - #789
Conversation
X had dedicated analysis only for workloads and pods. Every other kind fell back to a Ready condition, so a failed Job read as "exposes no Ready condition to assess", and a Pending claim or a broken CronJob got no explanation at all. Jobs report their Failed or Complete reason, progress, and failures against backoffLimit. CronJobs report suspension, schedule, last runs, runs blocked by concurrencyPolicy Forbid, and the Jobs they own. Claims report why they are not bound, pending resizes, ReadWriteOnce claims used on several nodes, and the pods that mount them. Nodes report readiness, cordon, taints, pod capacity, and their unhealthy pods. Explain and the diagnostic bundle now share one evidence gatherer. Discussions: #677, #678, #690
|
|
Review found several findings that claimed more than the evidence showed. A failed Job or pod list read as an empty one, so a CronJob "had no runs" and a claim "had no consumer". Storage sizes were compared as strings, so 1024Mi against 1Gi, or a volume larger than requested, read as a pending resize. Finished pods counted toward the ReadWriteOnce multi-node check, and a WaitForFirstConsumer claim blamed scheduling after its pod was scheduled. Unlisted evidence is now reported as unknown, sizes compare numerically, finished pods are ignored for attachment, and a scheduled consumer moves the blame to provisioning. Job pods must be owned by the Job, so a manual selector cannot pull in other pods. Container restarts under OnFailure count against backoffLimit. Events without a UID match on namespace as well as name. Ready nodes show since when, and Enter on a CronJob's Job opens it.
Failed pods and container restarts were added together and compared with backoffLimit, and the restarts of pods that had already failed were counted again. The Job controller checks the two separately: failed pods against the limit, and the restarts of running or pending pods, init containers included, against the same limit. The finding now uses the larger count, so a Job is no longer shown as out of retries before it is.
Without a pods kind nothing was listed, but the empty pod set still counted as listed, so a Pending WaitForFirstConsumer claim said no pod used it. StorageClasses were also read only when the pods kind existed. The pod set is now unknown in that case, and StorageClasses are read for every claim.
Summary
Xonly had dedicated analysis for workloads and pods. Every other kind fell back to its Ready condition. A failed Job read as "exposes no Ready condition to assess", and a Pending claim or a broken CronJob got no explanation at all. Three separate discussions asked for this: #677, #678 and #690.Job: the Failed or Complete condition and its reason, progress, failures used against
backoffLimitwhile it runs, and failed pods with their exit reasons.CronJob: suspension, schedule and time zone, last scheduled and successful runs, runs skipped by
concurrencyPolicy: Forbid, its five newest Jobs (each a target forE), and the pods of the latest Job.PersistentVolumeClaim: why a claim is not bound:
WaitForFirstConsumerwith no pod using the claimAlso covered: a resize that has not finished, a ReadWriteOnce claim used on several nodes, and the pods that mount the claim. Listing StorageClasses needs cluster-scope RBAC. Without it, a class is reported as unknown, never as missing.
Node: Ready state and since when, cordon, taints, pod capacity, and up to ten unhealthy pods on the node. Pressure-condition handling is unchanged.
Evidence gathering for these kinds lives in one
gather_evidencethat both explain and:bundleuse, so diagnostic bundles get the same analysis.Discussions: #677, #678, #690
Test plan
just check: fmt, clippy-D warnings, 2002 testssrc/explain.rsXkey tests for a Pending claim (class lookup, consumer pods filtered by claim), a CronJob (only its own Jobs), and a Node (cordon, unhealthy pods)Xon a failed Job, a Pending PVC, and a cordoned node on a real cluster