Skip to content

feat: show committed node capacity as request and limit percentages - #788

Merged
nklmilojevic merged 2 commits into
mainfrom
feat/node-committed-capacity
Oct 6, 2026
Merged

nklmilojevic merged 2 commits into
mainfrom
feat/node-committed-capacity

Conversation

@nklmilojevic

Copy link
Copy Markdown
Owner

Why

The nodes view showed live usage against allocatable, but the scheduler places pods by requests. "Why won't my pod schedule while the node is 30% busy?" had no answer in sofka. Users needed kubectl describe node or kube-capacity instead (k9s #764, #2723, #3529, #2846).

What

  • %CPU/R and %MEM/R are default nodes-view columns: the summed requests of the pods on each node as a share of allocatable.
  • %CPU/L and %MEM/L show in wide mode. Containers without a limit add nothing, as in kubectl describe node, so limits can exceed 100%.
  • New metric sources: node-cpu-request, node-memory-request, node-cpu-limit, node-memory-limit, and node-request:<resource> / node-limit:<resource> for extended resources such as nvidia.com/gpu. A node that does not advertise the resource shows -.
  • Pods are counted the way the scheduler counts them: app containers and native sidecars, raised to the largest init container step, plus overhead. Pod-level declarations take priority.
  • The existing nodes pods watch keeps the per-node totals up to date as pods change. It adds no extra API calls and works without Metrics Server.
  • The new columns sort, filter (%cpu/r>=80) and color like the other node percentages.

Tests

  • Unit tests for the scheduler-style pod totals and for tracking pods that move, resize or are deleted.
  • An app test that drives render, wide mode, sort and filter, including a GPU column, through handle_key.

Nodes only showed live usage against allocatable, so "why won't my pod
schedule while the node is 30% busy?" had no answer without kubectl
describe node or kube-capacity. The scheduler places pods by requests,
not usage.

The nodes pods watch now keeps per-node request and limit totals next to
the pod count, counting each pod the way the scheduler does: app
containers and native sidecars, raised to the largest init container
step, plus overhead, with pod-level declarations taking priority.
%CPU/R and %MEM/R are default node columns, %CPU/L and %MEM/L show in
wide mode, and node-request:<resource> / node-limit:<resource> cover
extended resources such as nvidia.com/gpu. They sort, filter, and color
like the other node percentages and work without Metrics Server.
@kritikal-github

kritikal-github Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Re-runKritika Review

Shows node requests and limits as allocatable percentages.

No findings

Confidence 5/5 · medium risk: The risk rests on application logic that aggregates pod resources and displays node capacity percentages; it does not alter cluster state or configuration.

Incremental review of the changes since 25ea107.

Findings

Earlier findings (1 resolved)

Summary

The nodes view now derives committed resource percentages from its pod watch and exposes request and limit columns, including configurable extended resources. The latest change uses i128 totals and tests petabyte-scale memory, addressing the earlier large-quantity overflow finding.

Flow
flowchart LR
  A["spawn_node_pods_poll"] --> B["NodeLoads.apply"]
  B --> C["NodeLoads.snapshot"]
  C --> D["Msg::NodePods"]
  D --> E["App.node_loads"]
  E --> F["App.metric_value"]
  F --> G["MetricColumn.value"]
  G --> H["committed_pct"]
Loading

What's good

  • Reuses the existing node pod watch rather than adding another API request path.
  • Covers large memory totals and percentage calculation with a focused regression test.

59 context chunk(s) left out of the prompt to fit its budget.

Reviewed 75435e4 by kritika with chatgpt/gpt-6-sol.

Comment thread src/columns_metrics.rs Outdated
@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Adds request and limit percentage columns to the nodes view.

The PR appears safe to merge; no outstanding finding or actionable new issue was identified.

Summary

The PR adds scheduler-style per-node request and limit percentages, including configurable extended-resource columns, using the existing pods watch. The latest changes widen quantity and aggregate storage to i128, use saturating sums, and add a petabyte-scale regression test. The previously reported memory-total overflow is fixed.

  • Adds default request columns and wide-mode limit columns to the nodes view.
  • Keeps derived totals current as watched pods change.
  • Covers rendering, sorting, filtering, and pod-load updates in tests.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
    A[Pods watch] --> B[Per-pod effective resources]
    B --> C[Per-node request and limit totals]
    C --> D[Node percentage columns]
    D --> E[Display, sort, and filter]
Loading

Reviews (2) · Last reviewed commit: "fix: keep petabyte-scale node totals fro..."

Comment thread src/columns_metrics.rs Outdated
Quantities were stored in thousandths as i64, which saturates near 8Pi
and overflows when two such pods are summed on one node, panicking in
debug builds and showing wrong %MEM in release. Store quantities and
totals as i128 with saturating sums and divide in f64.
@nklmilojevic
nklmilojevic merged commit c8e9518 into main Oct 6, 2026
5 checks passed
@nklmilojevic
nklmilojevic deleted the feat/node-committed-capacity branch October 6, 2026 12:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant