Repository navigation
Conversation
|
THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS Validation record for head Re-ported 2026-09-16, and the figures below replace every earlier one on this pull request. The previous head
Encoding-space A/BTwo builds of the same PyVEX worktree differing only in
Every change is in one direction: a lift that returned an empty IRSB pointing at the start of the block now ends at the offending instruction with the prefix intact. No encoding moves any other way, and no encoding this guest decodes today stops decoding -- the guest being the one pyvex builds, Cavium with MIPS32R1 and R2. On both 64-bit guests not one of the 114,719 results differs, which is the check that the guard is inert when The predicate is narrower than the mnemonics, and that is load-bearingAn earlier draft of this re-port tested only the opcode and function, as the pre-import patch did, and the sweep above caught it taking away 111 translations on MIPS32 BE that the baseline decodes, plus 36 encodings whose clean
Reading the file turned up a second class the sweep could not see. Three whole opcodes change meaning on a MIPSR6 guest, where the same bytes are a 32-bit instruction this file decodes: The test is deliberately confined to those three. An earlier draft stood down for every slot holding a MIPSR6 alternative anywhere in it, which is the wrong criterion: in the SPECIAL doubleword multiply group, in COP1 and in SPECIAL3 The MIPSR6 arm is unmeasured either way: One group added since the previous head
Caveats: this repository has no test suite, so the executable evidence is the consumer's. The sweep covers the encoding space, not real binaries; the corpus and CFG measurements from the pre-import head are withdrawn and have not been retaken. 455 of the 114,719 sampled encodings still abort on a 32-bit guest in the SPECIAL2, SPECIAL3 and COP1 slots; those are mixed slots rather than a clean MIPS64-only group and this change does not address them. |
|
THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS The block reported in pyvex issue 76 -- six MIPS32 instructions then Before -- the whole superblock is discarded and vex master (875f7c9), via pyvex bdd5441After -- the reserved encoding takes the ordinary with this change (vex 561795c, via pyvex dca2871) |
A MIPS64-only instruction is a Reserved Instruction on a 32-bit MIPS, so the decoder has to report a decode failure for it. This file instead decodes ld, daddiu, the doubleword shifts and multiplies, ldl/ldr, sdl/sdr, dmfc1/dmtc1, dext/dins, dsbh/dshd, dclz/dclo, the OCTEON bbit032/bbit132 branches and their siblings unconditionally, assigning an I64 to an I32 guest register. That trips either the vassert in putIReg or the IR sanity check at the end of the superblock, and both unwind out of the whole translation: the caller gets an empty IRSB whose Ijk_NoDecode points at the start of the block instead of at the instruction that failed, so the instructions decoded ahead of it are thrown away too. Reject those encodings before the per-opcode dispatch, which reaches the same decode_failure every other unsupported encoding uses and ends the block cleanly at the offending instruction. Three of the tests are narrower than the mnemonic, because the slot is shared with an encoding a 32-bit guest does decode and the narrowing is on the field that selects between them: the SPECIAL doubleword multiply and divide group is taken at sa == 0, where the MIPSR6 DMUL/DMUH/DDIV/ DMOD group sits at sa 2 and 3; COP1 rs 1 and 5 are DMFC1 and DMTC1 only with the low 11 bits clear; and SPECIAL3 DBSHFL selects DSBH and DSHD through sa 2 and 5, leaving DBITSWAP and DALIGN alone. Three whole opcodes change meaning on a MIPSR6 guest rather than merely sharing a field, so the predicate takes hwcaps and rejects them only when the guest is not MIPSR6: DADDI is BOVC/BEQC/BEQZALC there, and BBIT032 and BBIT132 are JIC/BEQZC and JIALC/BNEZC. The narrowed cases need no such test, because at the field values covered here the file builds 64-bit IR whatever hwcaps says. LWU, SD, BBIT0 and BBIT1 decode in 32-bit mode and are left alone, and LLD and SCD test mode64 themselves. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
561795c to
bc077e2
Compare
THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS
Problem
A MIPS64-only encoding is a Reserved Instruction on a 32-bit MIPS, but
priv/guest_mips_toIR.cdecodes a large group of them without consultingmode64first. The block reported in pyvex issue 76 is six ordinary MIPS32 instructions followed bydaddu $a0, $v0, $zero(0x0040202d) at0x3c5f0:The four loads, the store and the
addiuahead of thedadduare lost with it, andnextpoints at the start of the block rather than at the encoding that failed, so the caller is handed no bytes at all and no address to resume from.Root cause
These cases build 64-bit IR and hand it to a 32-bit guest register. libVEX then
longjmps out of the whole translation, fromputIReg'svassert(typeOfIRExpr(irsb->tyenv, e) == ty)forld, from the IR sanity check (Iex.Binop: arg tys don't match op tys) fordsll32, and fromvassert(mode64)fordmtc1. The unwind happens afterdisInstr_MIPS_WRKhas already emitted IR for the preceding instructions, and past the code that would have kept it.Fix
is_MIPS64_only_insnrecognises the encodings that fail this way, anddisInstr_MIPS_WRKconsults it before the per-opcode dispatch:decode_failureis the path every other unsupported encoding already takes, so the block ends at the offending address with everything before it intact:Only encodings that fail today are listed, so the predicate cannot take away a translation that currently succeeds. In three places it is narrower than the mnemonic, because this file shares those opcode slots with instructions a 32-bit guest does decode: SPECIAL function
0x1C–0x1Falso hold the MIPSR6DMUL/DMUH/DDIV/DMODgroup, selected bysa, so onlysa == 0is rejected; COP1rs1 and 5 areDMFC1andDMTC1only when the low 11 bits are clear; and SPECIAL3DBSHFLselectsDSBHandDSHDthroughsa. It also takeshwcaps, because three whole opcodes change meaning on a MIPSR6 guest rather than merely sharing a field, and there the same bytes are a 32-bit instruction the decoder handles:DADDIisBOVC/BEQC/BEQZALC, andBBIT032andBBIT132areJIC/BEQZCandJIALC/BNEZC. Those three are rejected only when the guest is not MIPSR6. The three narrowed cases above need no such test, because at the field values they cover this file builds 64-bit IR whateverhwcapssays.LWU,SD,BBIT0andBBIT1decode in 32-bit mode and are left alone, andLLDandSCDtestmode64themselves.BBIT032andBBIT132, the Cavium OCTEON "plus 32" branches, test a bit above 31 and abort on a non-MIPSR6 32-bit guest, so there they are covered.Testing
tests/test_mips32_reserved.py, which lands with the consumer, names one encoding per affected group and asserts the two-instruction prefix survives each; on the merge base it fails at its first entry withassert 0 == 8, the block being empty where the prefix should be. Its companion tests pin that MIPS64 still decodes the same encodings, and that seven encodings a coarser guard would have swallowed still keep their prefix on MIPS32. Four of them decode:SD,LWU,BBIT0andBBIT1, whose opcodes the predicate never covers. Three sit inside a covered slot at a field value it excludes — the MIPSR6DMULform, a COP1 encoding with the low bits set, and aDBSHFLthat is neitherDSBHnorDSHD. The whole PyVEX suite is 93 passed at this head.A two-build A/B over 114,719 encodings per guest, each lifted behind two decodable instructions, moves 16,786 results on each 32-bit guest, every one of them an abort becoming a clean failure that keeps the prefix, with no encoding losing a translation and no result changing on either 64-bit guest. That covers the guest pyvex builds, which pins Cavium with MIPS32R1 and R2; the MIPSR6 case is read from the decoder rather than lifted, because these bindings cannot construct a MIPSR6 guest. Validation: #91 (comment)
Consumed by angr/pyvex#567, which carries the regression tests and the submodule bump. Merge order: this lands first, then angr/pyvex#567, whose
vexsubmodule is pinned to this head. Merging the pyvex half on its own leaves pyvex master's submodule pointing at a commit that is on no vex branch. Nothing in this repository reads async:line —.github/workflows/build.ymlis the only workflow and it is three build jobs — so the line below is a note to the reader.sync: angr/pyvex#567
session: sharpen