Repo: intel/xpumanager · Affected: v2.1.0 and current main
Package as installed: xpu-smi 2.1.0+26.33.6468cec-1~26.04.
Hardware: 8 × Intel Arc Pro B60 (8086:e211), Linux 7.0.0-31, xe, Level Zero 1.32.0.
1. Three code paths report ECC state; one of them is wrong
Reproduced on our system, three commands in a row, same device, same moment:
$ xpu-smi dump -d 0 --metrics ecc.mode.current --number 1
... Enabled
$ xpu-smi discovery -d 0 | grep "ECC State"
ECC State: enabled
$ xpu-smi # summary table, "Volatile Uncorr. ECC" column
| 0 Intel(R) Arc(TM) Pro B Off | 0000:18:00.0 Off | Disabled |
... Disabled for all eight chips
ECC is genuinely enabled on these cards — confirmed independently: turning it off on one
chip raised memory.total from 21414 MiB to 24480 MiB (exactly 1/8 of the memory), and
the change persisted across a reboot.
The summary table is the odd one out because it never asks Sysman:
// ial/cmn/cmd_smi.cpp:99-104
ze_device_properties_t zeDevProp = {};
if (dev->getDevProps(di->deviceHdl, &zeDevProp) == ZE_RESULT_SUCCESS) {
stats.name = zeDevProp.name;
stats.eccEnabled = (zeDevProp.flags & ZE_DEVICE_PROPERTY_FLAG_ECC) != 0;
}
ZE_DEVICE_PROPERTY_FLAG_ECC is a capability bit — on our hardware it is not set — whereas
dump and discovery call zesDeviceGetEccState and get the actual state.
⇒ One binary, one field, three paths, contradictory answers. This is reproducible on any
machine where the capability flag and the Sysman state disagree, with no special setup.
Suggested fix: have the summary path query Sysman like the other two, and fall back to
the flag only if that fails (as discovery already does).
2. utilization.memory returns occupancy, not utilisation
Description and implementation sit in the same initialiser:
// ial/cmn/metrics/utilization.cpp:160-171
.description = "GPU memory utilization as a fraction of elapsed time memory was being read "
"or written; per tile or device, ...",
...
.getter = [](devInfo &d, MetricValue &out, const MetricCache &) -> ze_result_t {
auto *mem = d.dev->getMemory();
...
const auto r = mem->getMemoryUsed(nullptr, &val);
There is no time window and no pair of samples — it simply returns used memory. Measured on
our hardware, the arithmetic confirms it exactly:
DeviceId utilization.memory memory.used memory.total
0 0.16 % 34.234375 21414
4 0.16 % 34.109375 21414
check: 34.234375 / 21414 × 100 = 0.1599 % → 0.16
⇒ utilization.memory is a duplicate of memory.used expressed as a percentage, while
sitting in the UTILIZATION group next to metrics that do measure time-based activity.
Suggested fix: either implement the documented semantics, or rename the field and correct
the description.
3. Five discovery fields are placeholders printed like measurements
| field |
what is actually returned |
location |
| SKU Type |
the string literal "Production ES", always |
cmd_discovery.cpp:2019 |
| Number of Tiles |
numSlices > 0 ? 1 : 0 — always 1 |
:1663 |
| Max Command Queue Priority |
both branches return "0" |
:1638-1650 |
| Stepping |
parsed out of the model name; on a parse failure substitutes "A0" |
:1076-1083 |
| Memory ECC State (fallback) |
the "device supports ECC" flag presented as "ECC is on" |
:1782-1786 |
// cmd_discovery.cpp:2016-2020
}
// SKU type is typically derived from model name or board number
// For now, return a generic identifier based on flags
*outputLine = "Production ES";
return ZE_RESULT_SUCCESS;
Note the value is also self-contradictory: "Production" and "ES" (Engineering Sample) at once.
We initially took this as a statement about our boards and were preparing to raise it with the
board vendor — it is a constant in the tool.
🔑 This is a regression, not an unimplemented field. The 1.x line had a real
implementation — core/src/device/skuType.cpp (still visible at tag v1.3.8), with an enum
including PCH_PRODUCTION_STATE_PRQ and a "Production PRQ" return path. A user in
#122 running xpumcli on the same
0xe211 device reports SKU Type: Production PRQ — a genuinely read value. In the 2.x line
that module is gone and the field became a literal (verified in v2.0.1, v2.1.0 and
current main).
Stepping deserves a separate note:
// cmd_discovery.cpp:1073-1083
if (lastSpace != std::string::npos && lastSpace + 1 < modelName.length()) {
std::string stepping = modelName.substr(lastSpace + 1);
if (stepping.length() >= 2 && isalpha(stepping[0]) && isdigit(stepping[1])) {
*outputLine = stepping;
} else {
*outputLine = "A0";
}
} else {
*outputLine = "A0"; // Default fallback
}
Our Device Name is Intel(R) Arc(TM) Pro B60 Graphics, so the last word is Graphics, the
check fails, and A0 is printed. All eight of our chips report Stepping = A0, which is
almost certainly this substitution rather than a reading.
Suggested fix: print N/A when a value is not actually available. A placeholder that
looks like a measurement is worse than an honest gap — it cannot be told apart from real data.
4. UUID is documented as globally unique and immutable, but is derived from the PCI address
ial/cmn/metrics/identity.h:80-82 describes the UUID as "globally unique immutable
identifier". On our eight cards the values differ only in the PCI bus number — so moving a
card to another slot changes the "immutable" identifier, and two identical systems can
produce the same UUIDs.
Repo: intel/xpumanager · Affected: v2.1.0 and current
mainPackage as installed:
xpu-smi 2.1.0+26.33.6468cec-1~26.04.Hardware: 8 × Intel Arc Pro B60 (
8086:e211), Linux 7.0.0-31,xe, Level Zero 1.32.0.1. Three code paths report ECC state; one of them is wrong
Reproduced on our system, three commands in a row, same device, same moment:
ECC is genuinely enabled on these cards — confirmed independently: turning it off on one
chip raised
memory.totalfrom 21414 MiB to 24480 MiB (exactly 1/8 of the memory), andthe change persisted across a reboot.
The summary table is the odd one out because it never asks Sysman:
ZE_DEVICE_PROPERTY_FLAG_ECCis a capability bit — on our hardware it is not set — whereasdumpanddiscoverycallzesDeviceGetEccStateand get the actual state.⇒ One binary, one field, three paths, contradictory answers. This is reproducible on any
machine where the capability flag and the Sysman state disagree, with no special setup.
Suggested fix: have the summary path query Sysman like the other two, and fall back to
the flag only if that fails (as
discoveryalready does).2.
utilization.memoryreturns occupancy, not utilisationDescription and implementation sit in the same initialiser:
There is no time window and no pair of samples — it simply returns used memory. Measured on
our hardware, the arithmetic confirms it exactly:
⇒
utilization.memoryis a duplicate ofmemory.usedexpressed as a percentage, whilesitting in the
UTILIZATIONgroup next to metrics that do measure time-based activity.Suggested fix: either implement the documented semantics, or rename the field and correct
the description.
3. Five
discoveryfields are placeholders printed like measurements"Production ES", alwayscmd_discovery.cpp:2019numSlices > 0 ? 1 : 0— always 1:1663"0":1638-1650"A0":1076-1083:1782-1786Note the value is also self-contradictory: "Production" and "ES" (Engineering Sample) at once.
We initially took this as a statement about our boards and were preparing to raise it with the
board vendor — it is a constant in the tool.
🔑 This is a regression, not an unimplemented field. The 1.x line had a real
implementation —
core/src/device/skuType.cpp(still visible at tagv1.3.8), with an enumincluding
PCH_PRODUCTION_STATE_PRQand a"Production PRQ"return path. A user in#122 running
xpumclion the same0xe211device reportsSKU Type: Production PRQ— a genuinely read value. In the 2.x linethat module is gone and the field became a literal (verified in
v2.0.1,v2.1.0andcurrent
main).Stepping deserves a separate note:
Our
Device NameisIntel(R) Arc(TM) Pro B60 Graphics, so the last word isGraphics, thecheck fails, and
A0is printed. All eight of our chips reportStepping = A0, which isalmost certainly this substitution rather than a reading.
Suggested fix: print
N/Awhen a value is not actually available. A placeholder thatlooks like a measurement is worse than an honest gap — it cannot be told apart from real data.
4. UUID is documented as globally unique and immutable, but is derived from the PCI address
ial/cmn/metrics/identity.h:80-82describes the UUID as "globally unique immutableidentifier". On our eight cards the values differ only in the PCI bus number — so moving a
card to another slot changes the "immutable" identifier, and two identical systems can
produce the same UUIDs.