Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions mkdocs/blog/posts/agentic-orchestration.md
Original file line number Diff line number Diff line change
Expand Up @@ -250,12 +250,11 @@ $ dstack event --within-run train-qwen
```shell
$ dstack metrics train-qwen

UTILIZATION MEMORY
cpu ▅▄▄▆▆▆▆▆▆▆▆▆▆▆▆▆▆▅▆▆▆▆▆▆▆▆▆ 91% of 32 ▃▃▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄ 116GB/200GB
UTILIZATION MEMORY
job=0 cpu ▂▂▁▁▁▁▁▁▁▁▁▁▂▂▂▂▂▂▂▂▂▂▂▂▁▁ 15% ▃▃▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄ 116GB/200GB
gpu=0 ▃▃▃▃▃▃▃▃▃▃▃▃▃▄▄▄▄▅▅▄▄▄▃▃▃▃ 43% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 70GB/80GB

gpu=0 ▁▂▃▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 92% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 70GB/80GB

4 Aug 14:10 ┄┄┄┄┄┄┄┄┄┄┄ now 4 Aug 14:10 ┄┄┄┄┄┄┄┄┄┄┄ now
6 Aug 13:14 ┄┄┄┄┄┄┄┄┄┄ now 6 Aug 13:14 ┄┄┄┄┄┄┄┄┄┄ now
```

</div>
Expand Down
20 changes: 8 additions & 12 deletions mkdocs/blog/posts/dstack-metrics.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,15 +21,14 @@ for monitoring container metrics, including GPU usage for `NVIDIA`, `AMD`, and o
```shell
$ dstack metrics llama-70b-sft

UTILIZATION MEMORY
cpu ▅▄▄▄▃▃▃▃▃▃▃▃▃▃▃▃▅▅▄▂▃▃▃▃▃▃▃ 39% of 64 ▃▃▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄ 297GB/480GB

gpu=0 ▁▂▃▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 89% ▄▅▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 67GB/80GB
gpu=1 ▁▂▆▆▆▅▆▆▆▆▆▆▆▆▆▆▁▁▅▆▆▅▆▆▆▆▆ 84% ▄▅▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 67GB/80GB
gpu=2 ▁▂▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆ 87% ▄▅▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 67GB/80GB
gpu=3 ▂▃▆▅▅▅▅▅▅▆▅▅▆▆▆▆▁▅▅▅▅▅▆▅▅▅▅ 82% ▄▅▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 67GB/80GB

4 Aug 14:10 ┄┄┄┄┄┄┄┄┄┄┄ now 4 Aug 14:10 ┄┄┄┄┄┄┄┄┄┄┄ now
UTILIZATION MEMORY
job=0 cpu ▁▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▁▂▂▂▂▂▂▂▂▂ 33% ▃▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄ 296GB/480GB
gpu=0 ▁▄▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 93% ▃▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 67GB/80GB
gpu=1 ▂▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 93% ▃▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 67GB/80GB
gpu=2 ▃▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆ 93% ▃▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 67GB/80GB
gpu=3 ▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆ 93% ▃▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 67GB/80GB

6 Aug 13:14 ┄┄┄┄┄┄┄┄┄┄ now 6 Aug 13:14 ┄┄┄┄┄┄┄┄┄┄ now
```

</div>
Expand All @@ -47,9 +46,6 @@ difference is that `dstack stats` includes GPU VRAM usage and GPU utilization pe
Similar to `kubectl top`, if a run consists of multiple jobs (such as distributed training or an auto-scalable service),
`dstack stats` will display metrics per job.

> Note, `dstack metrics` now shows one job at a time, like `dstack logs`. Use `--replica` and `--job` to
> choose it; both default to `0`.

!!! info "HTTP API"
In addition to the `dstack stats` CLI commands, metrics can be obtained via the
[`/api/project/{project_name}/metrics/job/{run_name}`](../../docs/reference/http/metrics.md) HTTP endpoint.
Expand Down
33 changes: 15 additions & 18 deletions mkdocs/docs/concepts/metrics.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,34 +19,31 @@ This tab displays key CPU, memory, and GPU metrics collected during the last hou
## CLI

As an alternative to the UI, you can track essential metrics via the CLI.
The `dstack metrics` command charts CPU, memory, and GPU utilization over the last hour of the
job, with the latest value beside each chart.
The `dstack metrics` command charts CPU, memory, and GPU utilization over the last hour.

<div class="termy">

```shell
dstack metrics gentle-mayfly-1

UTILIZATION MEMORY
cpu ▅▄▄▄▃▃▃▃▃▃▃▃▃▃▃▃▅▅▄▃▃▃▃▃▃▃▃ 41% of 128 ▃▃▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄ 581GB/960GB

gpu=0 ▁▂▃▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 89% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=1 ▁▂▆▆▆▅▆▆▆▆▆▆▆▆▆▆▁▁▅▆▆▅▆▆▆▆▆ 84% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=2 ▁▂▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆ 87% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=3 ▂▃▆▅▅▅▅▅▅▆▅▅▆▆▆▆▁▅▅▅▅▅▆▅▅▅▅ 82% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=4 ▂▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▁▆▆▆▆▆▆▆▆▆▆ 90% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=5 ▂▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆▆ 85% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=6 ▃▆▆▅▅▆▆▆▆▆▆▆▆▆▆▁▅▅▅▅▆▆▆▅▅▅▅ 83% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=7 ▃▆▆▆▆▆▆▆▆▆▆▆▆▆▁▁▆▆▆▆▆▆▆▆▆▆▆ 88% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB

4 Aug 14:10 ┄┄┄┄┄┄┄┄┄┄┄ now 4 Aug 14:10 ┄┄┄┄┄┄┄┄┄┄┄ now
UTILIZATION MEMORY
job=0 cpu ▁▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▁▂▂▂▂▂▂▂▂▂ 33% ▃▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄ 579GB/960GB
gpu=0 ▁▄▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=1 ▂▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=2 ▃▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=3 ▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=4 ▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=5 ▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=6 ▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB
gpu=7 ▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB

6 Aug 13:14 ┄┄┄┄┄┄┄┄┄┄ now 6 Aug 13:14 ┄┄┄┄┄┄┄┄┄┄ now
```

</div>

Like [`dstack logs`](../reference/cli/dstack/logs.md), the command shows a single job.
Use `--replica` and `--job` to select one; both default to `0`.
Pass `-w` to keep the charts updating.
By default, metrics are shown for all jobs and replicas. Use `--replica` or `--job` to
show a single one, and `-w` to keep the charts updating.

## Prometheus

Expand Down
9 changes: 4 additions & 5 deletions mkdocs/docs/guides/migration/slurm.md
Original file line number Diff line number Diff line change
Expand Up @@ -1476,12 +1476,11 @@ Check real-time metrics:
```shell
$ dstack metrics training-job

UTILIZATION MEMORY
cpu ▅▄▄▄▃▃▃▃▃▃▃▃▃▃▃▃▅▅▄▃▃▃▃▃▃▃▃ 45% of 32 ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ 16GB/200GB
UTILIZATION MEMORY
job=0 cpu ▁▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▁▂▂▂▂▂▂▂▂▂ 33% ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ 16GB/200GB
gpu=0 ▁▄▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 93% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB

gpu=0 ▁▂▃▆▆▆▆▆▆▆▆▆▆▆▆▆▆▁▆▆▆▆▆▆▆▆▆ 90% ▄▅▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆ 71GB/80GB

4 Aug 14:10 ┄┄┄┄┄┄┄┄┄┄┄ now 4 Aug 14:10 ┄┄┄┄┄┄┄┄┄┄┄ now
6 Aug 13:14 ┄┄┄┄┄┄┄┄┄┄ now 6 Aug 13:14 ┄┄┄┄┄┄┄┄┄┄ now
```

</div>
Expand Down
47 changes: 26 additions & 21 deletions src/dstack/_internal/cli/commands/metrics.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
import argparse
import time
from typing import Optional

from rich.live import Live

Expand Down Expand Up @@ -36,50 +37,45 @@ def _register(self):
)
self._parser.add_argument(
"--replica",
help="The replica number. Defaults to 0.",
help="Show only this replica. By default, all jobs are shown.",
type=int,
default=0,
)
self._parser.add_argument(
"--job",
help="The job number inside the replica. Defaults to 0.",
help="Show only this job number. By default, all jobs are shown.",
type=int,
default=0,
)

def _command(self, args: argparse.Namespace):
super()._command(args)
job, metrics = self._fetch(args)
jobs, metrics = self._fetch(args)

if not args.watch:
console.print(get_metrics_table(job, metrics))
console.print(get_metrics_table(jobs, metrics))
return

try:
with Live(console=console, refresh_per_second=LIVE_TABLE_REFRESH_RATE_PER_SEC) as live:
while True:
live.update(get_metrics_table(job, metrics))
live.update(get_metrics_table(jobs, metrics))
time.sleep(WATCH_INTERVAL_SECONDS)
job, metrics = self._fetch(args)
jobs, metrics = self._fetch(args)
except KeyboardInterrupt:
pass

def _fetch(self, args: argparse.Namespace) -> tuple[Job, JobMetrics]:
def _fetch(self, args: argparse.Namespace) -> tuple[list[Job], list[JobMetrics]]:
run = self.api.runs.get(run_name=args.run_name)
if run is None:
raise CLIError(f"Run {args.run_name} not found")
job = _get_job(run, args.replica, args.job)
return job, _get_job_metrics(self.api, run, job)


def _get_job(run: Run, replica_num: int, job_num: int) -> Job:
for job in run._run.jobs:
if job.job_spec.replica_num == replica_num and job.job_spec.job_num == job_num:
return job
raise CLIError(
f"Run {run.name} has no replica={replica_num} job={job_num}."
" Use --replica and --job to select one."
)
jobs = select_jobs(run._run.jobs, args.replica, args.job)
if not jobs:
wanted = " ".join(
f"{name}={value}"
for name, value in (("replica", args.replica), ("job", args.job))
if value is not None
)
raise CLIError(f"Run {args.run_name} has no job matching {wanted}")
return jobs, [_get_job_metrics(self.api, run, job) for job in jobs]


def _get_job_metrics(api: Client, run: Run, job: Job) -> JobMetrics:
Expand All @@ -92,3 +88,12 @@ def _get_job_metrics(api: Client, run: Run, job: Job) -> JobMetrics:
job_num=job.job_spec.job_num,
limit=MAX_SAMPLES,
)


def select_jobs(jobs: list[Job], replica: Optional[int], job_num: Optional[int]) -> list[Job]:
return [
job
for job in jobs
if (replica is None or job.job_spec.replica_num == replica)
and (job_num is None or job.job_spec.job_num == job_num)
]
Loading