Skip to content

[BUG] Metric aggregation: dead code path makes average ~10⁶ too small; all non-GpuTime metrics read as fp32 regardless of type; EU percentages advertise 0.1 % precision but carry 1 % #167

Description

@AlexanderScha

Repo: intel/xpumanager · Affected: v2.1.0 and current main (hal/core/metric.cpp
is unchanged between the tag and origin/main as of 06.09.2026)
Package as installed: xpu-smi 2.1.0+26.33.6468cec-1~26.04.
Hardware: 8 × Intel Arc Pro B60 (8086:e211, BMG G21), Linux 7.0.0-31, xe, L0 1.32.0.

Summary

Three independent defects in hal/core/metric.cpp corrupt numbers that the user reads as
measurements. All three are visible by inspection; no special hardware is required.

1. The time-weighting branch is dead — average is off by the report interval

// hal/core/metric.cpp:350-356
for (auto &metricEntry : aggregatedGroupData.data) {
	if (metricEntry.type == "time") {
		metricEntry.total += static_cast<double>(reportElapsedTime) * metricEntry.current;
	} else {
		metricEntry.total += metricEntry.current;
	}
}
cumulativeTime += reportElapsedTime;

type is assigned in exactly one place, and it can only ever take two values:

// hal/core/metric.cpp:113
metricData->type = (metricProps.resultType == ZET_VALUE_TYPE_UINT64) ? "uint64_t" : "double";

Nothing ever assigns the string "time", so the weighting branch never executes and
total accumulates raw instantaneous values, while the divisor keeps accumulating elapsed
time:

// hal/core/metric.cpp:361-365
metricEntry.average = metricEntry.total / static_cast<double>(cumulativeTime);

Arithmetic. With 100 reports at 1 ms each, cumulativeTime ≈ 1e8 ns. For a metric
sitting at 50 %, total is 100 × 50 = 5000, so average = 5000 / 1e8 = 5e-5 instead of 50.
Had the branch worked: 1e6 × 50 × 100 / 1e8 = 50 ✓.

⇒ The average column is understated by roughly the report interval — about 10⁶ at typical
settings. current is computed separately and is unaffected.

Suggested fix: either set type = "time" where appropriate, or drop the branch and
divide by the sample count.

2. Every metric except GpuTime is read as fp32, whatever its declared type

// hal/core/metric.cpp:318-323 (and again at :336-341 for the new-entry path)
if (name == PERF_GPU_TIME_METRIC) {
	reportElapsedTime = value.value.ui64;
	metricEntry.current = static_cast<double>(value.value.ui64);
} else {
	metricEntry.current = value.value.fp32;
}

The union member is chosen by metric name, not by the declared resultType — even
though that type is computed a few hundred lines earlier (:113) and stored on the same
object.

⇒ Any metric with resultType == ZET_VALUE_TYPE_UINT64 other than GpuTime is reinterpreted
from the union as float32, producing garbage. Event counters — the majority of EU metrics
— fall exactly into this case.

Suggested fix: switch on metricProps.resultType (or the stored type), not on the
name.

3. EU percentages: scaleFactor = 1000 promises per-mille, the data carries whole percent

// hal/core/metric.cpp — accumulation in integers
totalEuStall  += static_cast<uint64_t>(static_cast<double>(currentGPUElapsedTime) * currentEuStall);
totalEuActive += static_cast<uint64_t>(static_cast<double>(currentGPUElapsedTime) * currentEuActive);
...
// integer division, then scaling
uint64_t euActive = totalEuActive / totalGPUElapsedTime;
uint64_t euStall  = totalEuStall  / totalGPUElapsedTime;
...
data.scaleFactor = 1000;
data.euActive = euActive * data.scaleFactor;

The division is uint64_t / uint64_t, so the fraction is discarded before scaling. A
true 37.9 % becomes 37, then 37000.

⇒ scaleFactor = 1000 advertises thousandths of a percent, but every emitted value is a
multiple of 1000 — the real resolution is 1 %. The truncation is one-sided (always down)
and propagates into euIdle = 100 − euActive − euStall.

Practical consequence: comparing two runs whose EU utilisation differs by less than 1 %
is meaningless — the difference does not exist in the data.

⚠️ Note that current main adds

// hal/core/metric.h:32-34
/// Divisor to convert EuMetricsData fields to percent: value / EU_PERMILLE_SCALE → 0–100.
inline constexpr double EU_PERMILLE_SCALE = 1000.0;

but a repository-wide search finds no use of it — the constant was introduced while the
computation above stayed unchanged.

Suggested fix: divide in floating point (or scale the numerator by 1000 before the
integer division) so the per-mille resolution the scale factor promises actually exists.

Possibly the cause of an existing report

Issue #147 ("The data coming out of
xpu-smi are incoherent when the load is high") shows, among other things,
utilization.gpu = 23104788116198.93. A value of that shape is what defect 2 above
produces: a uint64_t counter reinterpreted from the union as float32. The report is
still open and the cause was not identified there — this may be it.

Why these matter together

current is trustworthy; average is not (1). Counter-typed metrics are garbage (2). EU
percentages have 100× less resolution than their own scale factor claims (3). A user reading
xpu-smi output has no way to tell which of the printed numbers is which.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions