Repo: intel/xpumanager · Affected: v2.1.0 and current main (hal/core/metric.cpp
is unchanged between the tag and origin/main as of 06.09.2026)
Package as installed: xpu-smi 2.1.0+26.33.6468cec-1~26.04.
Hardware: 8 × Intel Arc Pro B60 (8086:e211, BMG G21), Linux 7.0.0-31, xe, L0 1.32.0.
Summary
Three independent defects in hal/core/metric.cpp corrupt numbers that the user reads as
measurements. All three are visible by inspection; no special hardware is required.
1. The time-weighting branch is dead — average is off by the report interval
// hal/core/metric.cpp:350-356
for (auto &metricEntry : aggregatedGroupData.data) {
if (metricEntry.type == "time") {
metricEntry.total += static_cast<double>(reportElapsedTime) * metricEntry.current;
} else {
metricEntry.total += metricEntry.current;
}
}
cumulativeTime += reportElapsedTime;
type is assigned in exactly one place, and it can only ever take two values:
// hal/core/metric.cpp:113
metricData->type = (metricProps.resultType == ZET_VALUE_TYPE_UINT64) ? "uint64_t" : "double";
Nothing ever assigns the string "time", so the weighting branch never executes and
total accumulates raw instantaneous values, while the divisor keeps accumulating elapsed
time:
// hal/core/metric.cpp:361-365
metricEntry.average = metricEntry.total / static_cast<double>(cumulativeTime);
Arithmetic. With 100 reports at 1 ms each, cumulativeTime ≈ 1e8 ns. For a metric
sitting at 50 %, total is 100 × 50 = 5000, so average = 5000 / 1e8 = 5e-5 instead of 50.
Had the branch worked: 1e6 × 50 × 100 / 1e8 = 50 ✓.
⇒ The average column is understated by roughly the report interval — about 10⁶ at typical
settings. current is computed separately and is unaffected.
Suggested fix: either set type = "time" where appropriate, or drop the branch and
divide by the sample count.
2. Every metric except GpuTime is read as fp32, whatever its declared type
// hal/core/metric.cpp:318-323 (and again at :336-341 for the new-entry path)
if (name == PERF_GPU_TIME_METRIC) {
reportElapsedTime = value.value.ui64;
metricEntry.current = static_cast<double>(value.value.ui64);
} else {
metricEntry.current = value.value.fp32;
}
The union member is chosen by metric name, not by the declared resultType — even
though that type is computed a few hundred lines earlier (:113) and stored on the same
object.
⇒ Any metric with resultType == ZET_VALUE_TYPE_UINT64 other than GpuTime is reinterpreted
from the union as float32, producing garbage. Event counters — the majority of EU metrics
— fall exactly into this case.
Suggested fix: switch on metricProps.resultType (or the stored type), not on the
name.
3. EU percentages: scaleFactor = 1000 promises per-mille, the data carries whole percent
// hal/core/metric.cpp — accumulation in integers
totalEuStall += static_cast<uint64_t>(static_cast<double>(currentGPUElapsedTime) * currentEuStall);
totalEuActive += static_cast<uint64_t>(static_cast<double>(currentGPUElapsedTime) * currentEuActive);
...
// integer division, then scaling
uint64_t euActive = totalEuActive / totalGPUElapsedTime;
uint64_t euStall = totalEuStall / totalGPUElapsedTime;
...
data.scaleFactor = 1000;
data.euActive = euActive * data.scaleFactor;
The division is uint64_t / uint64_t, so the fraction is discarded before scaling. A
true 37.9 % becomes 37, then 37000.
⇒ scaleFactor = 1000 advertises thousandths of a percent, but every emitted value is a
multiple of 1000 — the real resolution is 1 %. The truncation is one-sided (always down)
and propagates into euIdle = 100 − euActive − euStall.
Practical consequence: comparing two runs whose EU utilisation differs by less than 1 %
is meaningless — the difference does not exist in the data.
⚠️ Note that current main adds
// hal/core/metric.h:32-34
/// Divisor to convert EuMetricsData fields to percent: value / EU_PERMILLE_SCALE → 0–100.
inline constexpr double EU_PERMILLE_SCALE = 1000.0;
but a repository-wide search finds no use of it — the constant was introduced while the
computation above stayed unchanged.
Suggested fix: divide in floating point (or scale the numerator by 1000 before the
integer division) so the per-mille resolution the scale factor promises actually exists.
Possibly the cause of an existing report
Issue #147 ("The data coming out of
xpu-smi are incoherent when the load is high") shows, among other things,
utilization.gpu = 23104788116198.93. A value of that shape is what defect 2 above
produces: a uint64_t counter reinterpreted from the union as float32. The report is
still open and the cause was not identified there — this may be it.
Why these matter together
current is trustworthy; average is not (1). Counter-typed metrics are garbage (2). EU
percentages have 100× less resolution than their own scale factor claims (3). A user reading
xpu-smi output has no way to tell which of the printed numbers is which.
Repo: intel/xpumanager · Affected: v2.1.0 and current
main(hal/core/metric.cppis unchanged between the tag and
origin/mainas of 06.09.2026)Package as installed:
xpu-smi 2.1.0+26.33.6468cec-1~26.04.Hardware: 8 × Intel Arc Pro B60 (
8086:e211, BMG G21), Linux 7.0.0-31,xe, L0 1.32.0.Summary
Three independent defects in
hal/core/metric.cppcorrupt numbers that the user reads asmeasurements. All three are visible by inspection; no special hardware is required.
1. The time-weighting branch is dead —
averageis off by the report intervaltypeis assigned in exactly one place, and it can only ever take two values:Nothing ever assigns the string
"time", so the weighting branch never executes andtotalaccumulates raw instantaneous values, while the divisor keeps accumulating elapsedtime:
Arithmetic. With 100 reports at 1 ms each,
cumulativeTime ≈ 1e8ns. For a metricsitting at 50 %,
totalis 100 × 50 = 5000, soaverage = 5000 / 1e8 = 5e-5instead of 50.Had the branch worked:
1e6 × 50 × 100 / 1e8 = 50✓.⇒ The
averagecolumn is understated by roughly the report interval — about 10⁶ at typicalsettings.
currentis computed separately and is unaffected.Suggested fix: either set
type = "time"where appropriate, or drop the branch anddivide by the sample count.
2. Every metric except GpuTime is read as
fp32, whatever its declared typeThe union member is chosen by metric name, not by the declared
resultType— eventhough that type is computed a few hundred lines earlier (
:113) and stored on the sameobject.
⇒ Any metric with
resultType == ZET_VALUE_TYPE_UINT64other than GpuTime is reinterpretedfrom the union as
float32, producing garbage. Event counters — the majority of EU metrics— fall exactly into this case.
Suggested fix: switch on
metricProps.resultType(or the storedtype), not on thename.
3. EU percentages:
scaleFactor = 1000promises per-mille, the data carries whole percentThe division is
uint64_t / uint64_t, so the fraction is discarded before scaling. Atrue 37.9 % becomes 37, then 37000.
⇒
scaleFactor = 1000advertises thousandths of a percent, but every emitted value is amultiple of 1000 — the real resolution is 1 %. The truncation is one-sided (always down)
and propagates into
euIdle = 100 − euActive − euStall.Practical consequence: comparing two runs whose EU utilisation differs by less than 1 %
is meaningless — the difference does not exist in the data.
mainaddsbut a repository-wide search finds no use of it — the constant was introduced while the
computation above stayed unchanged.
Suggested fix: divide in floating point (or scale the numerator by 1000 before the
integer division) so the per-mille resolution the scale factor promises actually exists.
Possibly the cause of an existing report
Issue #147 ("The data coming out of
xpu-smi are incoherent when the load is high") shows, among other things,
utilization.gpu = 23104788116198.93. A value of that shape is what defect 2 aboveproduces: a
uint64_tcounter reinterpreted from the union asfloat32. The report isstill open and the cause was not identified there — this may be it.
Why these matter together
currentis trustworthy;averageis not (1). Counter-typed metrics are garbage (2). EUpercentages have 100× less resolution than their own scale factor claims (3). A user reading
xpu-smioutput has no way to tell which of the printed numbers is which.