Repository navigation
Conversation
|
THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS Validation record for head Re-ported 2026-09-16, and this record replaces every figure that stood here before. The previous head
Regression and suite
All of the above run through the store environment the workspace builds, never a virtual environment; the package under test is the store build of these worktrees. The block issue #76 reports
Encoding-space A/BTwo builds of this worktree differing only in
Every change is in one direction, and no encoding this configuration decodes today stops decoding. Not one of the 114,719 results differs on either 64-bit guest, which is the check that the guard is inert when What the sweep caught before publicationA first re-port tested only the opcode and function, as the pre-import patch did, and this A/B showed it taking away 111 translations on MIPS32 BE, plus 36 encodings whose clean A second finding, from reading rather than from the A/B: three whole opcodes change meaning on a MIPSR6 guest, where the same bytes are a 32-bit instruction this file decodes -- Two groups also moved relative to the previous head. Caveats: the sweep covers the encoding space, not real binaries. This removes the lost superblock; it does not make MIPS n32 objects decode past the reserved instruction, which is what the architecture, lifter, loader and consumer changes do. 455 of the 114,719 sampled encodings still abort on a 32-bit guest in SPECIAL2, SPECIAL3 and COP1 slots that are not a clean MIPS64-only group; this change does not address them. |
|
Corpus decompilation diffs can be found at angr/dec-snapshots@master...angr/pyvex_567 |
28dd87a to
dca2871
Compare
|
THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS The block reported in #76 -- six MIPS32 instructions then Before -- pyvex master (bdd5441, vex 875f7c9)After -- the reserved encoding takes vex's ordinary with this change (dca2871, vex 561795c) |
libVEX decoded about thirty MIPS64-only encodings on a 32-bit MIPS guest and assigned 64-bit values to 32-bit guest registers, which failed either vassert(mode64) or the IR sanity check and unwound the whole translation. LibVEXLifter._lift then saw size == 0 and raised, and lift() fell through to an empty IRSB whose Ijk_NoDecode pointed at the start of the block rather than at the instruction that could not be decoded -- so the instructions decoded ahead of it were lost, and the "decode more bytes and extend" path never ran, because it is guarded on final_irsb.size > 0. Carry the vex fix and add the regression tests: the prefix survives for one encoding from each affected group and for the two OCTEON branches, MIPS64 still decodes all of them, and the encodings that share those opcode slots with a 32-bit instruction -- SD, LWU, BBIT0, BBIT1, the MIPSR6 DMUL group, a COP1 format with the low bits set, and the rest of DBSHFL -- still decode, so a guard that was too coarse would fail the suite. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dca2871 to
8737c95
Compare
THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS
Problem
A MIPS64-only encoding met by the 32-bit MIPS guest costs the caller the whole superblock, not just that instruction. The block reported in #76 is six ordinary MIPS32 instructions followed by
daddu $a0, $v0, $zero(0x0040202d) at0x3c5f0:nextis the address the block started at, so a caller that follows it re-lifts the same bytes and gets the same empty block. Every one of the fifteen encodings the regression covers behaves this way.Root cause
guest_mips_toIR.cdecodes MIPS64-only encodings without consultingmode64, assigns the 64-bit result to a 32-bit guest register, and libVEXlongjmps out of the translation --putIReg'svassert(typeOfIRExpr(irsb->tyenv, e) == ty)forld, the IR sanity check (Iex.Binop: arg tys don't match op tys) fordsll32,vassert(mode64)fordmtc1. Sweeping 114,719 encodings against a 32-bit guest finds 16,786 of them losing the whole superblock this way.Nothing above it recovers the decoded prefix, because there is nothing left to recover: the C side returns
size == 0, soLibVEXLifter._liftraisesand
lift()falls through itsfor ... elsetoIRSB.empty_block(arch, addr, size=0, nxt=Const(addr), jumpkind="Ijk_NoDecode"). The path that would extend and re-lift is guarded onfinal_irsb.size > 0 and final_irsb.jumpkind == "Ijk_NoDecode", which a zero-size block never satisfies.Fix
Bump the
vexsubmodule to angr/vex#91, which rejects those encodings through the samedecode_failurepath every other unsupported encoding uses. The block then ends at the offending address with the prefix intact:No Python changes; the submodule bump and
tests/test_mips32_reserved.pyare the whole diff.Testing
tests/test_mips32_reserved.pynames one encoding per affected group --ld,daddiu,daddu,dsll,dsll32,dsrl32,dmult,ldl,ldr,sdl,sdr,dmtc1,dext,dsbh,dclz, and the two OCTEONbbit032/bbit132branches, seventeen in all -- and asserts for each that the block keeps its two-instruction prefix, thatnextis the reserved address, and thatIjk_NoDecodeis still the jumpkind. On the merge base the file is 2 failed, 4 passed -- the non-branch test stopping at its first entry withassert 0 == 8, the branch test failing the same way onbbit032. At this head the file is 6 passed and the whole suite is 93 passed.The remaining tests pin that MIPS64 still decodes every one of those encodings, and that seven encodings a coarser guard would have swallowed still lift with their prefix intact on MIPS32.
sd $ra, 0x5c8($sp),lwu $a0, 0x10($v1)and the 32-bitbbit0/bbit1decode, their opcodes being ones the guard never covers; the MIPSR6dmulform that sharesdmult's function code, a COP1 encoding that sharesdmfc1'srsbut sets the low bits, and adbshflthat is neitherdsbhnordshdraiseILLEGAL_INSTRUCTON. A two-build A/B over 114,719 encodings per guest is what caught a coarser first draft of the guard taking 111 of them away, so the narrowing is tested rather than asserted.This removes the crash #76 reports; it does not make those objects decode past the reserved instruction, which is what the n32 architecture, lifter, loader and consumer changes do. Validation: #567 (comment)
Merge order: angr/vex#91 first. This head's
vexsubmodule is that pull request's head, one commit ahead ofvexmaster, so merging this half on its own leaves pyvex master's submodule pointing at a commit that is on no vex branch.angr/vexis not in the CI repository list, so thesync:line below is a note to the reader rather than an instruction to CI.sync: angr/vex#91
session: sharpen