Skip to content

feat: explain Jobs, CronJobs, PersistentVolumeClaims, and Nodes - #789

Merged
nklmilojevic merged 4 commits into
mainfrom
feat/explain-pvc-node-job
Oct 6, 2026
Merged

nklmilojevic merged 4 commits into
mainfrom
feat/explain-pvc-node-job

Conversation

@nklmilojevic

Copy link
Copy Markdown
Owner

Summary

X only had dedicated analysis for workloads and pods. Every other kind fell back to its Ready condition. A failed Job read as "exposes no Ready condition to assess", and a Pending claim or a broken CronJob got no explanation at all. Three separate discussions asked for this: #677, #678 and #690.

  • Job: the Failed or Complete condition and its reason, progress, failures used against backoffLimit while it runs, and failed pods with their exit reasons.

  • CronJob: suspension, schedule and time zone, last scheduled and successful runs, runs skipped by concurrencyPolicy: Forbid, its five newest Jobs (each a target for E), and the pods of the latest Job.

  • PersistentVolumeClaim: why a claim is not bound:

    • a missing StorageClass, or no default class
    • WaitForFirstConsumer with no pod using the claim
    • an unprovisioned volume
    • a pending bind to a named volume

    Also covered: a resize that has not finished, a ReadWriteOnce claim used on several nodes, and the pods that mount the claim. Listing StorageClasses needs cluster-scope RBAC. Without it, a class is reported as unknown, never as missing.

  • Node: Ready state and since when, cordon, taints, pod capacity, and up to ten unhealthy pods on the node. Pressure-condition handling is unchanged.

Evidence gathering for these kinds lives in one gather_evidence that both explain and :bundle use, so diagnostic bundles get the same analysis.

Discussions: #677, #678, #690

Test plan

  • just check: fmt, clippy -D warnings, 2002 tests
  • Unit tests for each new analysis in src/explain.rs
  • X key tests for a Pending claim (class lookup, consumer pods filtered by claim), a CronJob (only its own Jobs), and a Node (cordon, unhealthy pods)
  • Manual: X on a failed Job, a Pending PVC, and a cordoned node on a real cluster

X had dedicated analysis only for workloads and pods. Every other kind fell
back to a Ready condition, so a failed Job read as "exposes no Ready
condition to assess", and a Pending claim or a broken CronJob got no
explanation at all.

Jobs report their Failed or Complete reason, progress, and failures against
backoffLimit. CronJobs report suspension, schedule, last runs, runs blocked
by concurrencyPolicy Forbid, and the Jobs they own. Claims report why they
are not bound, pending resizes, ReadWriteOnce claims used on several nodes,
and the pods that mount them. Nodes report readiness, cordon, taints, pod
capacity, and their unhealthy pods. Explain and the diagnostic bundle now
share one evidence gatherer.

Discussions: #677, #678, #690
@kritikal-github

kritikal-github Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Re-runKritika Review

Explains Jobs, CronJobs, claims, and Nodes with shared evidence gathering.

No findings

Confidence 4/5 · medium risk: src/app/explain.rs skips Job pods when a valid selector uses only matchExpressions, leaving failures and restarts out of the explanation; this case is not covered by the reported review. Risk is medium because this affects application diagnostics rather than cluster state.

Summary

The change adds dedicated diagnostics for four resource kinds and shares evidence gathering between Explain and bundles. The latest adjustment reads StorageClasses independently of Pod discovery and marks unavailable Pod evidence as unknown, preserving the distinction between missing and unlisted resources.

What's good

  • Shared evidence gathering keeps interactive explanations and bundles consistent.
  • Tests exercise missing Pod discovery and the distinction between unreadable and absent evidence.

58 context chunk(s) left out of the prompt to fit its budget.

Reviews (4) · Last reviewed commit: "fix: treat pods as unknown when discover..." · kritika with chatgpt/gpt-6-sol

Comment thread src/app/explain.rs Outdated
Comment thread src/app/explain.rs Outdated
Comment thread src/app/tests.rs Outdated
Comment thread src/explain.rs
Comment thread src/explain.rs
Comment thread src/explain.rs Outdated
Comment thread src/explain.rs
@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Adds diagnostic analysis for four Kubernetes resource types.

The PR appears safe to merge; no new actionable issue was identified.

Summary

The PR adds dedicated explanations for Jobs, CronJobs, PersistentVolumeClaims, and Nodes, and shares evidence gathering between Explain and diagnostic bundles.

  • Related pods, Jobs, StorageClasses, and events inform the new findings.
  • The latest changes distinguish unavailable pod evidence from an empty pod list.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Selected resource] --> B[gather_evidence]
  B --> C[Related resources and events]
  C --> D[Kind-specific explanation]
  D --> E[Explain view]
  D --> F[Diagnostic bundle]
Loading

Reviews (4) · Last reviewed commit: "fix: treat pods as unknown when discover..."

Comment thread src/app/explain.rs Outdated
Comment thread src/app/explain.rs Outdated
Comment thread src/explain.rs
Comment thread src/explain.rs
Comment thread src/explain.rs
Comment thread src/explain.rs Outdated
Comment thread src/app/explain.rs
Comment thread src/explain.rs
Review found several findings that claimed more than the evidence showed.
A failed Job or pod list read as an empty one, so a CronJob "had no runs"
and a claim "had no consumer". Storage sizes were compared as strings, so
1024Mi against 1Gi, or a volume larger than requested, read as a pending
resize. Finished pods counted toward the ReadWriteOnce multi-node check, and
a WaitForFirstConsumer claim blamed scheduling after its pod was scheduled.

Unlisted evidence is now reported as unknown, sizes compare numerically,
finished pods are ignored for attachment, and a scheduled consumer moves the
blame to provisioning. Job pods must be owned by the Job, so a manual
selector cannot pull in other pods. Container restarts under OnFailure count
against backoffLimit. Events without a UID match on namespace as well as
name. Ready nodes show since when, and Enter on a CronJob's Job opens it.
Comment thread src/explain.rs Outdated
Comment thread src/explain.rs
Failed pods and container restarts were added together and compared with
backoffLimit, and the restarts of pods that had already failed were counted
again. The Job controller checks the two separately: failed pods against
the limit, and the restarts of running or pending pods, init containers
included, against the same limit. The finding now uses the larger count, so
a Job is no longer shown as out of retries before it is.
Without a pods kind nothing was listed, but the empty pod set still counted
as listed, so a Pending WaitForFirstConsumer claim said no pod used it.
StorageClasses were also read only when the pods kind existed. The pod set
is now unknown in that case, and StorageClasses are read for every claim.
@nklmilojevic
nklmilojevic merged commit 3c20d79 into main Oct 6, 2026
5 checks passed
@nklmilojevic
nklmilojevic deleted the feat/explain-pvc-node-job branch October 6, 2026 14:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant