Skip to content

Commit 9ebd2c5

Browse files
committed
amx: the full mnemonic surface (INT8/BF16/FP16/COMPLEX/FP8/TF32/MOVRS/AVX512) with const tile operands
src/hpc/amx_ops.rs exposes every tile op LLVM's X86InstrAMX.td defines as `asm!` mnemonics on stable 1.98.1 (LLVM 22.1.8 assembles all of them with no target feature), the tile index as an `asm_const` generic, all eight tiles. Three-tile ops assert distinct operands in a `const` block, so the same-tile #UD (Gotcha 11) is a compile error. `amx_features()` returns the per-tier CPUID bits at LLVM Host.cpp's positions (7.0:EDX 22/24/25, 7.1:EAX[21], 7.1:EDX[8], 1E.1:EAX 4/6/7/8); `amx_report()` prints them. The AMX-AVX512 row ops (tcvtrow*, tilemovrow) need a zmm operand and therefore exist under the avx512f cfg only — compile-time selection, no target_feature. Encoding falsifiers read each monomorphized op's bytes back out of the text segment (bounded at the wrapper's own `ret`, since the linker packs the wrappers back to back) and pin them to the EMR-validated `.byte` table and to the LLVM encodings of the never-executed tiers. In doing so the "mirrored operand convention" (Gotcha 12) resolved as a misread of the byte table: Intel order is tdpbusd D, S1(rm, M×K), S2(vvvv, VNNI K×N), and the validated C4 E2 71 5E C2 is `tdpbusd tmm0, tmm2, tmm1` — the kernel's placement, plain SDM semantics. Recorded as Gotcha 15 and knowledge doc §5b. Executed tier is still only TILE/INT8/BF16 (Emerald Rapids); every other tier is assembler-verified only and the docs say so. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X6y3drwKSE2zSgoexheLFX
1 parent a55d841 commit 9ebd2c5

6 files changed

Lines changed: 599 additions & 2 deletions

File tree

‎.claude/AMX_GOTCHAS.md‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -267,6 +267,19 @@ in production under load, AVX-512 siblings unaffected.
267267

268268
---
269269

270+
## Gotcha 15: the operand "mirror" was a misread of the byte table — use mnemonics
271+
272+
`src/hpc/amx_ops.rs` (2026-09-14) assembles every AMX mnemonic on stable
273+
1.98.1 with `const` tile operands. Intel order is `tdpbusd tmmD, tmmS1, tmmS2`
274+
= `D += S1·S2`, S1 = ModRM.rm (plain M×K), S2 = VEX.vvvv (VNNI K×N). The
275+
validated `C4 E2 71 5E C2` is `tdpbusd tmm0, tmm2, tmm1`, i.e. the kernel's
276+
"A in tmm2, B in tmm1" placement is the plain SDM semantics, not a mirror.
277+
Gotcha 12's *placement* stays correct; its *explanation* is superseded.
278+
Aliased tile operands are now a compile error (`const` assert), so the SIGILL
279+
of Gotcha 11 cannot be written. Encodings of every tier are pinned by tests
280+
that read the emitted bytes back; the extended tiers (FP16/COMPLEX/FP8/TF32/
281+
MOVRS/AVX512) are assembler-verified ONLY — no host here has executed them.
282+
270283
## Hardware tiers
271284

272285
```

‎.claude/blackboard.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -150,6 +150,15 @@ simd_{avx512,avx2,neon,wasm,scalar}.rs each owns its realization, as a PEER
150150

151151
**What landed:**
152152

153+
> **AMX fill (2026-09-14):** `src/hpc/amx_ops.rs` — the whole `X86InstrAMX.td`
154+
> surface as mnemonics with `const` tile operands (INT8×4, BF16, FP16,
155+
> COMPLEX×2, FP8×4, TF32, MOVRS×2, AVX512 row ops ×6 under `avx512f` cfg,
156+
> STTILECFG/TILELOADDT1, all 8 tiles), `amx_features()` per LLVM `Host.cpp`
157+
> bits, `amx_report()` prints the tiers. Encoding falsifiers read the emitted
158+
> bytes back on any x86 host and pin them to the EMR-validated table; the
159+
> "mirrored operand convention" turned out to be a misread of that table
160+
> (Gotcha 15). Extended tiers are assembler-verified only.
161+
153162
> **CI finding (2026-09-14, e730109 red on `tier4-avx512-check`):** `ci.yaml`'s
154163
> workflow-global `RUSTFLAGS: "-D warnings"` REPLACES every `.cargo/config*`
155164
> rustflags entry (cargo precedence: RUSTFLAGS > target.<triple> > target.<cfg>

‎.claude/knowledge/amx-enablement-and-kernel.md‎

Lines changed: 46 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -169,6 +169,52 @@ which is bug #2. For the 16×16 int8/bf16 tile, all three tiles are 16 rows ×
169169

170170
---
171171

172+
## 5b. The mnemonic surface — `src/hpc/amx_ops.rs` (2026-09-14, LLVM 22.1.8)
173+
174+
The `.byte` tables above were forced by 1.94. Measured on **1.98.1 (LLVM
175+
22.1.8)**: the integrated assembler accepts EVERY AMX mnemonic inside `asm!`
176+
with no target feature, and `asm_const` makes the tile index a generic
177+
parameter (`tilezero tmm{t}`, `t = const T`). `amx_ops.rs` exposes the whole
178+
`X86InstrAMX.td` surface that way, and its tests read the emitted bytes back
179+
out of the text segment and pin them to this table — on any x86_64 host, no
180+
EMR needed. Tile-operand aliasing (Gotcha 11) is a `const` assert, so `#UD`
181+
is now a compile error.
182+
183+
**The "mirror" (Gotcha 12, §4) is a reading of the byte table, not a hardware
184+
quirk.** `MRMSrcReg4VOp3` puts `dst` in ModRM.reg, **S1 in ModRM.rm, S2 in
185+
VEX.vvvv**, and Intel syntax names them in that order: `tdpbusd tmmD, tmmS1,
186+
tmmS2` = `D += S1(M×K, plain) · S2(K×N, VNNI)`, with the `U`/`S` letters
187+
naming S1 then S2. The table row `C4 E2 71 5E C2` (rm = tmm2, vvvv = tmm1)
188+
is the mnemonic `tdpbusd tmm0, tmm2, tmm1` — `amx_ops::tdpbusd::<0, 2, 1>` —
189+
which is exactly the kernel's placement (A u8 → tmm2, B VNNI i8 → tmm1). The
190+
row's comment "dst=tmm0,vvvv=tmm1,rm=tmm2" had been read as the operand
191+
list `(tmm0, tmm1, tmm2)`; the assembler reads `(tmm0, tmm1, tmm2)` as
192+
rm = tmm1, vvvv = tmm2 = `C4 E2 69 5E C1`. Pinned two-sided in
193+
`operand_order_is_intel_order_rm_then_vvvv`.
194+
195+
Assembler-verified encodings for the tiers beyond GEMM (LLVM `Host.cpp`
196+
CPUID bits; NONE of these has executed in this workspace — no GNR/DMR host):
197+
198+
```
199+
CPUID bytes (tmm0,tmm1,tmm2 / tmm3,[rdi+rsi])
200+
TDPBSSD/TDPBSUD/TDPBUSD/TDPBUUD INT8 7.0:EDX[25] C4 E2 {6B,6A,69,68} 5E C1
201+
TDPBF16PS BF16 7.0:EDX[22] C4 E2 6A 5C C1
202+
TDPFP16PS FP16 7.1:EAX[21] C4 E2 6B 5C C1
203+
TCMMIMFP16PS / TCMMRLFP16PS COMPLEX 7.1:EDX[8] C4 E2 {69,68} 6C C1
204+
TDPBF8PS/TDPBHF8PS/TDPHBF8PS/ FP8 1E.1:EAX[4] C4 E5 {68,6B,6A,69} FD C1 (map5)
205+
TDPHF8PS
206+
TMMULTF32PS TF32 1E.1:EAX[6] C4 E2 69 48 C1 (dropped from LLVM main; 22.1.8 assembles it)
207+
TILELOADDRS / TILELOADDRST1 MOVRS 1E.1:EAX[8] C4 E2 {7B,79} 4A 1C 37
208+
TCVTROWD2PS zmm0,tmm1,edi / ,3 AVX512 1E.1:EAX[7] 62 F2 46 48 4A C1 / 62 F3 7E 48 07 C1 03 (EVEX; needs avx512f cfg)
209+
TCVTROWPS2{PHH,PHL,BF16H,BF16L} AVX512 62 F2 {44,46,47,45} 48 6D C1
210+
TILEMOVROW zmm0,tmm1,edi / ,5 AVX512 62 F2 45 48 4A C1 / 62 F3 7D 48 07 C1 05
211+
STTILECFG [rdi] / TILELOADDT1 TILE 7.0:EDX[24] C4 E2 79 49 07 / C4 E2 79 4B 14 16
212+
```
213+
214+
Detection: `amx_ops::amx_features()` (cached) returns the per-tier bits;
215+
`amx_available()` remains the execute gate (XCR0 + arch_prctl). `amx_report()`
216+
prints both. Gotcha 14 (VM tile-state corruption) applies to every tier.
217+
172218
## 6. Detection API (cached, CPU-aware)
173219

174220
```rust

0 commit comments

Comments
 (0)