Skip to content

Support MIPSN32 in the lifter - #576

Merged
ltfish merged 2 commits into
masterfrom
feature/mips-n32-lifter
Aug 29, 2026
Merged

ltfish merged 2 commits into
masterfrom
feature/mips-n32-lifter

Conversation

@zardus

@zardus zardus commented Aug 27, 2026 •

Copy link
Copy Markdown
Member

THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS

Problem

MIPS n32 lifts nothing. main at 0x10000110 in the chain's n32 fixture n32_be_static is 18 instructions of ordinary MIPS64 code, and every block of it comes back undecodable:

$ pyvex.lift(<72 bytes of main>, 0x10000110, archinfo.ArchMIPSN32(BE), opt_level=0)
  arch.bits=32  arch.vex_arch=VexArchMIPS64
  size=0  instructions=0  jumpkind=Ijk_NoDecode  Ity_I64 present=False
  instruction_addresses=()

The same bytes with ArchMIPS64 give size=44 instructions=11 jumpkind=Ijk_Call. Nothing is recovered for an n32 or O64 object: no blocks, so no functions, no calls and no data references.

Root cause

LibVEXLifter is selected by architecture name, not by vex_arch, and MIPSN32 is not in LIBVEX_SUPPORTED_ARCHES in pyvex/lifting/libvex.py. The name misses, no libVEX lifter is registered for the architecture, and the block falls through to Ijk_NoDecode with size 0 even though arch.vex_arch already reads VexArchMIPS64.

Falling back to MIPS32 is not a repair. It lifts the same 44 bytes, but sd $gp, 8($sp) — the 64-bit spill in a non-leaf n32 prologue, at 0x10000114 — silently becomes a 4-byte store:

   05 | ------ IMark(0x10000114, 4, 0) ------      <- ArchMIPS32
   06 | t8 = GET:I32(sp)
   09 | t9 = GET:I32(gp)
   10 | STbe(t0) = t9
   07 | ------ IMark(0x10000114, 4, 0) ------      <- ArchMIPS64
   08 | t15 = GET:I64(sp)
   11 | t16 = GET:I64(gp)
   12 | STbe(t0) = t16

Ity_I64 appears nowhere in the MIPS32 type environment, so the 32-bit guest reports a block that decodes cleanly and stores the wrong width.

Fix

Add MIPSN32 to LIBVEX_SUPPORTED_ARCHES and to the PyvexArch guest and instruction-pointer tables, mapped to VexArchMIPS64, and export ARCH_MIPSN32_BE / ARCH_MIPSN32_LE. n32 and O64 differ from MIPS64 in pointer width, not in the instruction set, so the correct guest is the one that already exists; no new VEX guest is added and libpyvex.so is byte-identical either side of this change (sha256 9e5b9532fde21223… on both, in the before/after capture). The architecture's word size stays 32 while vex_arch is VexArchMIPS64, which is the whole of what n32 means here. After:

$ pyvex.lift(<72 bytes of main>, 0x10000110, archinfo.ArchMIPSN32(BE), opt_level=0)
  arch.bits=32  arch.vex_arch=VexArchMIPS64
  size=44  instructions=11  jumpkind=Ijk_Call  Ity_I64 present=True

statement for statement identical to the ArchMIPS64 lift of the same bytes.

Testing

tests/test_mipsn32.py lifts that prologue and pins that n32 decodes exactly as ARCH_MIPS64_BE does — same statement list — that Ity_I64 is present for n32 and absent for ARCH_MIPS32_BE, and that ARCH_MIPSN32_BE.bits is 32 while its vex_arch equals the 64-bit architecture's. On the merge base it is 1 failed, AttributeError: module 'pyvex' has no attribute 'ARCH_MIPSN32_BE'; at this head, 1 passed.

These four must land together: with the loader and the architecture definition ahead of the lifter, MIPSN32 resolves and reaches a lifter with no entry for it, and recovery on the affected objects drops to zero blocks. Merge order: archinfo 375, then pyvex 576, then cle 795, then angr 6982.

Validation: #576 (comment)

sync: angr/archinfo#375

session: mega-corpus

lifters are registered by architecture name, so an architecture whose name libvex.py does not
list gets no lifter at all and every block comes back Ijk_NoDecode. MIPSN32 -- the n32 and O64
ABIs, a 64-bit MIPS instruction stream in an ELFCLASS32 container -- needs the same LibVEXLifter
as MIPS64, which it then dispatches through its own vex_arch of VexArchMIPS64.

PyvexArch gains the matching entry so the standalone arch objects cover it too, which is what
lets this be tested without archinfo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@angr-bot

Copy link
Copy Markdown
Member

Corpus decompilation diffs can be found at angr/dec-snapshots@master...angr/pyvex_576

@zardus

zardus commented Aug 27, 2026 •

Copy link
Copy Markdown
Member Author

THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS

Validation record for head 2bf2cff41aa1beb9a502a311a5154411f8d41158 against baseline bdd5441035e02920eaa72c1c3cf9a4f0d572104d, the merge base with master. CPython 3.12.13, pytest 9.1.1, Linux x86-64, against the libpyvex.so already built for this worktree — the branch changes no C, so no rebuild was needed.

  • Regression: python -m pytest tests/test_mipsn32.py — 1 passed on head. With pyvex/__init__.py, pyvex/arches.py and pyvex/lifting/libvex.py restored to the baseline and the test left in place, it fails with AttributeError: module 'pyvex' has no attribute 'ARCH_MIPSN32_BE'
  • Focused: python -m pytest tests/test_mipsn32.py — 1 passed
  • Full suite: python -m pytest tests — 65 passed, no skips, no xfails
  • Mutation: deleting "MIPSN32" from LIBVEX_SUPPORTED_ARCHES on head, leaving everything else, makes pyvex.lift(bytes.fromhex("27bdffe0ffbc00083c1c0002"), 0x10000110, pyvex.ARCH_MIPSN32_BE, opt_level=0) return size=0 insns=0 jumpkind=Ijk_NoDecode; with the entry present the same call returns size=12 insns=3 jumpkind=Ijk_Boring. That set is read at import time, when the lifter registers itself, so the entry is what does the work
  • Lint: ruff check on the three changed modules and the new test — all checks passed, ruff 0.16.4, the revision pinned in .pre-commit-config.yaml
  • Hosted CI: all 22 checks pass on this head (run 33038820956), pre-commit.ci included

Reproducer, no fixture and no angr needed. The three words are the start of a non-leaf n32 prologue — addiu $sp, $sp, -0x20, sd $gp, 8($sp), lui $gp, 2:

import pyvex
data = bytes.fromhex("27bdffe0" "ffbc0008" "3c1c0002")
kw = {"data": data, "mem_addr": 0x10000110, "num_inst": 3, "opt_level": 0}
sorted(set(pyvex.IRSB(arch=pyvex.ARCH_MIPS32_BE, **kw).tyenv.types))   # ['Ity_I32']
sorted(set(pyvex.IRSB(arch=pyvex.ARCH_MIPSN32_BE, **kw).tyenv.types))  # ['Ity_I32', 'Ity_I64']

The 32-bit guest decodes the sd as a 4-byte store rather than refusing it, which is why no 64-bit type appears; tests/test_mipsn32.py asserts both sides of that, and that the n32 statement list is identical to ARCH_MIPS64_BE's.

Caveats: this is a focused pyvex record, not the cross-repository workspace gate — no angr, cle or archinfo suite was run against it. ArchMIPSN32 itself comes from angr/archinfo#375, so until that merges nothing in the ecosystem selects this architecture and the entry added here is unreachable in practice; the test drives pyvex.ARCH_MIPSN32_BE directly and needs no archinfo change. No MIPS n32 object exists in angr/binaries, so the regression lifts three instruction words rather than loading a container. Only the big-endian variant is exercised; ARCH_MIPSN32_LE shares the same guest selection and is untested here.

Merge order, and what an incomplete chain does

The four pull requests in this chain are angr/archinfo#375 (defines ArchMIPSN32),
angr/pyvex#576 (teaches the lifter that name), angr/cle#795 (resolves the ELF variant)
and angr/angr#6982 (calling conventions, plus a CFGFast assertion). They must land
together, in that order: archinfo#375, then pyvex#576, then cle#795, then angr#6982.

Merging the loader and the architecture definition ahead of the lifter is the dangerous
partial state, and it is also the natural reading order. In it ArchMIPSN32 resolves, the
name reaches a lifter with no entry for it, every block comes back Ijk_NoDecode, and angr
recovers 0 blocks and 0 functions on every object in this population — worse than the
unfixed baseline, which at least recovers the 32-bit-decodable parts.

Measured 2026-08-28, at the heads under review, one process per object:

arm archinfo pyvex cle angr
master bf85c7e47bb469878c564480e28c677abbe35acc bdd5441035e02920eaa72c1c3cf9a4f0d572104d d2ecea068794d20b1f14d90eecc1bc4bc4cfa431 0e18fa25ef83c3db9813fe5a156da0013052e0e5
loader + arch only #375 9e9d9e26498bb5553ff5450a6764be161138f51d master #795 4aaf40ece48f07e863507e80769984ae138d11c2 master
all four #375 9e9d9e26… #576 2bf2cff41aa1beb9a502a311a5154411f8d41158 #795 4aaf40ec… #6982 3ac333750ad5b2ce76a62e6f3818c4000dbcf1d8

Population: 16 ELF objects, every one ELFCLASS32 EM_MIPS. Fourteen declare a MIPS III
ISA in e_flags — ten big-endian O64 (0x20002001) and four big-endian n32
(0x20000027) — and two are ordinary little-endian o32 (0x1001), held as controls. The
corpus is not public, so they are cited by digest below. Metric: DWARF line-table addresses
covered by no recovered block, under CFGFast(normalize=True, resolve_indirect_jumps=True).

Over the fourteen MIPS III objects, of 261,039 line addresses:

arm missed blocks functions missed function starts
master 59,268 151,093 41,104 684
loader + arch only 261,039 0 0 9,131
all four 2,366 167,600 15,989 16
  • The partial state misses 100% of every object's own line table, on all fourteen. The
    two o32 controls are untouched by it, so the damage is confined to the objects the chain
    is about.
  • The complete chain removes 96.0% of the master survivors, with block counts up on
    every one of the fourteen and no object regressing. Example: a9271ebdaaef5111… goes
    6,969 → 219 missed and 14,851 → 16,628 blocks; 761f33eb0e35fca5… goes 46 → 1 missed and
    360 → 373 blocks.
  • The function count falls because master's is inflated: each Ijk_NoDecode boundary starts
    a fresh pseudo-function. Two independent counters move the right way — declared function
    entry points with no covering block go 684 → 16, and FDE starts 60 → 0.
  • Controls: the two o32 objects report identical block counts, function counts and
    missed-address sets under every arm measured here, so nothing in the chain reaches
    ordinary MIPS32.
  • Determinism: the three arms were run in both orders (master → partial → chain, then chain
    → partial → master). All sixteen objects agree on every field in both passes.

angr/pyvex#567 overlaps this chain; do not add the two figures. Measured alone against
the same master (pyvex dca28718d29976fa0a775199076c8197bd177f89, vex
561795c17f97a008108d8642d4854c56e79513a2, libpyvex.so rebuilt for it), it takes the
fourteen from 59,268 to 47,526 — 11,742 addresses, 19.8%, per object 13.8%–32.6%. Added
on top of the complete four-PR chain it changes nothing at all: 2,366 both ways, object
for object, byte for byte. It repairs the discarded prefix, which the chain already avoids
by resolving the right architecture. Its two o32 controls are identical under every arm, so
it has no collateral effect on ordinary MIPS32 either.

Method notes: each arm asserts the resolved __file__ of angr, cle, pyvex and
archinfo lies inside the intended worktree before measuring, because a PYTHONPATH entry
that does not exist is skipped silently and falls through to the editable install. The
pyvex#567 native build was made in its own worktree and gated on a behavioural probe
(ld $a0, -0x3098($v1) at 0x80010938 lifts to size=8 under #567 and size=0 under
master), not on a resolved path; the shared checkout's prebuilt libpyvex.so was
hash-checked before and after that build and is unchanged.

Objects, sha256, all ELF:

O64  a9ab23096d37b67368d0ff00e6c37059e631f8f04c5b337a98176ac4f70ff377
O64  a9271ebdaaef51110e0a6fc23b1f3949e447ec825ae47fb4c9dc2b9c0d1b2634
O64  cbd909ae5fe39e44506a4975d106e044289706afdd659d37463a247211261dbf
O64  3b08741fdbc63589870323fe5f40373d543f34fdacb487218aad3ffcb65381dc
O64  7e65689497320f65d5951375c27116f79e5cec6cb20970169e5dc6028c242a78
O64  b860cfbe35c32968c8194bb29d223076964b452cd38ebb16e0bcf3f00328af19
O64  f2afd5b17235d3c9a10c993c03fad3d6883d89d2b861022f317d270e181f4c5e
O64  ebea75270e2c9d00170704b675d2b18c45b1b2aebbe34cdb13703bcec54b2f4c
O64  b634a05e39a6000fa52d276ba5518ffb67b03b3ddc31f79b4378bb319a41c89b
O64  a7d9003ec00bd8bb499573617fb51b3ce5b2d9a18c746f404606cce203c182e8
n32  acbc7e75858097e906084988261e7cd646cb4ee507cd7573d8719617bd45b196
n32  d254234ebbc95e26bba689de241193439828e975bd0945170d4df664c75588df
n32  ca2055345f9d85b4c66209e67a59facbc68cb5820c230c6f154123c47eef5b76
n32  761f33eb0e35fca5d108e729bd8066909163b3212b6238834a3cb922224b6739
o32  728368792e9a34a3a7a78ed9f0ccdfaee3f96ffd15ef09c8e6678ca6d098e882
o32  04b8ade9cac4172a71d90fdedf7bc7545fa0e32bd69b7d0c77fd8f835f8e65f6

Public reproducer for the ordering, on the fixtures in angr/binaries#206 at
f54a80068c7e67e1a48c44eab772a03ec1c74476, no corpus needed. For each of
tests/mipsn32/{n32_be_static,o64_el_static,n32_el_dynamic}, run
angr.Project(path, auto_load_libs=False).analyses.CFGFast(normalize=True) and count
cfg.model.nodes() and cfg.kb.functions:

arm n32_be_static o64_el_static n32_el_dynamic
master MIPS32, 13 blocks / 8 functions MIPS32, 13 / 8 MIPS32, 35 / 20
archinfo#375 + cle#795, pyvex at master MIPSN32, 0 / 0 MIPSN32, 0 / 0 MIPSN32, 0 / 0
all four MIPSN32, 13 / 6 MIPSN32, 13 / 6 MIPSN32, 35 / 15

@zardus

zardus commented Aug 28, 2026

Copy link
Copy Markdown
Member Author

THIS MESSAGE WAS GENERATED BY AN AUTOMATED PROCESS

Full lifted IR for main at 0x10000110 in the chain's n32 fixture n32_be_static, as ArchMIPSN32, ArchMIPS64 and ArchMIPS32, before and after this change. The archinfo checkout is the same on both sides, so the only thing that moves is pyvex's lifter selection; libpyvex.so is byte-identical too, sha256 9e5b9532fde21223… in both captures.

$ PYTHONPATH=<pyvex>:<archinfo> python n32_576.py <pyvex> <archinfo>
  # <archinfo> is angr/archinfo#375 at 39bab86b70890c41bd57a6fff55e467c73986007 on both sides

Before — MIPSN32 is not in LIBVEX_SUPPORTED_ARCHES, so the n32 lift is Ijk_NoDecode with size 0 and no instruction addresses at all; ArchMIPS32 decodes the same 44 bytes but its type environment holds no Ity_I64, which is the sd at 0x10000114 narrowed to a 4-byte store:

pyvex master (bdd5441, the merge base)
pyvex package:  .../scratch/sh20-07/wt/mips-base/pyvex/__init__.py
libpyvex.so:    .../scratch/sh20-07/wt/mips-base/pyvex/lib/libpyvex.so
  sha256        9e5b9532fde2122351d37d89881db0673b4bb79f8da04ce78db466cdc5fd3535
archinfo:       .../scratch/sh20-07/wt/archinfo-375/archinfo/__init__.py

MIPSN32 in pyvex.lifting.libvex.LIBVEX_SUPPORTED_ARCHES: False

$ pyvex.lift(<72 bytes of main>, 0x10000110, archinfo.ArchMIPSN32(BE), opt_level=0)
  arch.bits=32  arch.vex_arch=VexArchMIPS64
  size=0  instructions=0  jumpkind=Ijk_NoDecode  Ity_I64 present=False
  instruction_addresses=()
IRSB {
   

   NEXT: PUT(pc) = 0x10000110; Ijk_NoDecode
}

$ pyvex.lift(<72 bytes of main>, 0x10000110, archinfo.ArchMIPS64(BE), opt_level=0)
  arch.bits=64  arch.vex_arch=VexArchMIPS64
  size=44  instructions=11  jumpkind=Ijk_Call  Ity_I64 present=True
  instruction_addresses=('0x10000110', '0x10000114', '0x10000118', '0x1000011c', '0x10000120', '0x10000124', '0x10000128', '0x1000012c', '0x10000130', '0x10000134', '0x10000138')
IRSB {
   t0:Ity_I64 t1:Ity_I64 t2:Ity_I64 t3:Ity_I64 t4:Ity_I64 t5:Ity_I64 t6:Ity_I64 t7:Ity_I64 t8:Ity_I64 t9:Ity_I1 t10:Ity_I64 t11:Ity_I32 t12:Ity_I32 t13:Ity_I64 t14:Ity_I64 t15:Ity_I64 t16:Ity_I64 t17:Ity_I64 t18:Ity_I32 t19:Ity_I32 t20:Ity_I64 t21:Ity_I32 t22:Ity_I64 t23:Ity_I64 t24:Ity_I32 t25:Ity_I32 t26:Ity_I64 t27:Ity_I64 t28:Ity_I64 t29:Ity_I64 t30:Ity_I32 t31:Ity_I64 t32:Ity_I64 t33:Ity_I64 t34:Ity_I64 t35:Ity_I32 t36:Ity_I32 t37:Ity_I64 t38:Ity_I64 t39:Ity_I64 t40:Ity_I64 t41:Ity_I64 t42:Ity_I64 t43:Ity_I1 t44:Ity_I64

   00 | ------ IMark(0x10000110, 4, 0) ------
   01 | t13 = GET:I64(sp)
   02 | t12 = 64to32(t13)
   03 | t11 = Add32(t12,0xffffffe0)
   04 | t10 = 32Sto64(t11)
   05 | PUT(sp) = t10
   06 | PUT(pc) = 0x0000000010000114
   07 | ------ IMark(0x10000114, 4, 0) ------
   08 | t15 = GET:I64(sp)
   09 | t14 = Add64(t15,0x0000000000000008)
   10 | t0 = t14
   11 | t16 = GET:I64(gp)
   12 | STbe(t0) = t16
   13 | PUT(pc) = 0x0000000010000118
   14 | ------ IMark(0x10000118, 4, 0) ------
   15 | PUT(gp) = 0x0000000000020000
   16 | PUT(pc) = 0x000000001000011c
   17 | ------ IMark(0x1000011c, 4, 0) ------
   18 | t20 = GET:I64(t9)
   19 | t19 = 64to32(t20)
   20 | t22 = GET:I64(gp)
   21 | t21 = 64to32(t22)
   22 | t18 = Add32(t21,t19)
   23 | t17 = 32Sto64(t18)
   24 | PUT(gp) = t17
   25 | PUT(pc) = 0x0000000010000120
   26 | ------ IMark(0x10000120, 4, 0) ------
   27 | t26 = GET:I64(gp)
   28 | t25 = 64to32(t26)
   29 | t24 = Add32(t25,0xffff8170)
   30 | t23 = 32Sto64(t24)
   31 | PUT(gp) = t23
   32 | PUT(pc) = 0x0000000010000124
   33 | ------ IMark(0x10000124, 4, 0) ------
   34 | t28 = GET:I64(gp)
   35 | t27 = Add64(t28,0xffffffffffff8020)
   36 | t1 = t27
   37 | t30 = LDbe:I32(t1)
   38 | t29 = 32Sto64(t30)
   39 | PUT(t9) = t29
   40 | PUT(pc) = 0x0000000010000128
   41 | ------ IMark(0x10000128, 4, 0) ------
   42 | t32 = GET:I64(sp)
   43 | t31 = Add64(t32,0x0000000000000010)
   44 | t2 = t31
   45 | t33 = GET:I64(s8)
   46 | STbe(t2) = t33
   47 | PUT(pc) = 0x000000001000012c
   48 | ------ IMark(0x1000012c, 4, 0) ------
   49 | t36 = 64to32(0x0000000000000000)
   50 | t35 = Add32(t36,0x00000005)
   51 | t34 = 32Sto64(t35)
   52 | PUT(a0) = t34
   53 | PUT(pc) = 0x0000000010000130
   54 | ------ IMark(0x10000130, 4, 0) ------
   55 | t38 = GET:I64(sp)
   56 | t37 = Add64(t38,0x0000000000000018)
   57 | t3 = t37
   58 | t39 = GET:I64(ra)
   59 | STbe(t3) = t39
   60 | PUT(pc) = 0x0000000010000134
   61 | ------ IMark(0x10000134, 4, 0) ------
   62 | t5 = 0x0000000000000000
   63 | t6 = GET:I64(s1)
   64 | t8 = 0x0000000000000000
   65 | PUT(ra) = 0x000000001000013c
   66 | t9 = CmpLT64S(t5,0x0000000000000000)
   67 | t40 = 1Uto64(t9)
   68 | t7 = t40
   69 | PUT(pc) = 0x0000000010000138
   70 | ------ IMark(0x10000138, 4, 0) ------
   71 | t42 = GET:I64(sp)
   72 | t41 = Or64(t42,0x0000000000000000)
   73 | PUT(s8) = t41
   74 | t43 = CmpEQ64(t7,t8)
   75 | if (t43) { PUT(pc) = 0x10000180; Ijk_Call }
   76 | PUT(pc) = 0x000000001000013c
   77 | t44 = GET:I64(pc)
   NEXT: PUT(pc) = t44; Ijk_Call
}

$ pyvex.lift(<72 bytes of main>, 0x10000110, archinfo.ArchMIPS32(BE), opt_level=0)
  arch.bits=32  arch.vex_arch=VexArchMIPS32
  size=44  instructions=11  jumpkind=Ijk_Call  Ity_I64 present=False
  instruction_addresses=('0x10000110', '0x10000114', '0x10000118', '0x1000011c', '0x10000120', '0x10000124', '0x10000128', '0x1000012c', '0x10000130', '0x10000134', '0x10000138')
IRSB {
   t0:Ity_I32 t1:Ity_I32 t2:Ity_I32 t3:Ity_I32 t4:Ity_I1 t5:Ity_I32 t6:Ity_I32 t7:Ity_I32 t8:Ity_I32 t9:Ity_I32 t10:Ity_I32 t11:Ity_I32 t12:Ity_I32 t13:Ity_I32 t14:Ity_I32 t15:Ity_I32 t16:Ity_I32 t17:Ity_I32 t18:Ity_I32 t19:Ity_I32 t20:Ity_I32 t21:Ity_I32 t22:Ity_I32 t23:Ity_I32 t24:Ity_I32 t25:Ity_I1 t26:Ity_I32 t27:Ity_I32 t28:Ity_I32 t29:Ity_I32

   00 | ------ IMark(0x10000110, 4, 0) ------
   01 | t6 = GET:I32(sp)
   02 | t5 = Add32(t6,0xffffffe0)
   03 | PUT(sp) = t5
   04 | PUT(pc) = 0x10000114
   05 | ------ IMark(0x10000114, 4, 0) ------
   06 | t8 = GET:I32(sp)
   07 | t7 = Add32(t8,0x00000008)
   08 | t0 = t7
   09 | t9 = GET:I32(gp)
   10 | STbe(t0) = t9
   11 | PUT(pc) = 0x10000118
   12 | ------ IMark(0x10000118, 4, 0) ------
   13 | PUT(gp) = 0x00020000
   14 | PUT(pc) = 0x1000011c
   15 | ------ IMark(0x1000011c, 4, 0) ------
   16 | t11 = GET:I32(t9)
   17 | t12 = GET:I32(gp)
   18 | t10 = Add32(t12,t11)
   19 | PUT(gp) = t10
   20 | PUT(pc) = 0x10000120
   21 | ------ IMark(0x10000120, 4, 0) ------
   22 | t14 = GET:I32(gp)
   23 | t13 = Add32(t14,0xffff8170)
   24 | PUT(gp) = t13
   25 | PUT(pc) = 0x10000124
   26 | ------ IMark(0x10000124, 4, 0) ------
   27 | t16 = GET:I32(gp)
   28 | t15 = Add32(t16,0xffff8020)
   29 | t1 = t15
   30 | t17 = LDbe:I32(t1)
   31 | PUT(t9) = t17
   32 | PUT(pc) = 0x10000128
   33 | ------ IMark(0x10000128, 4, 0) ------
   34 | t19 = GET:I32(sp)
   35 | t18 = Add32(t19,0x00000010)
   36 | t2 = t18
   37 | t20 = GET:I32(s8)
   38 | STbe(t2) = t20
   39 | PUT(pc) = 0x1000012c
   40 | ------ IMark(0x1000012c, 4, 0) ------
   41 | t21 = Add32(0x00000000,0x00000005)
   42 | PUT(a0) = t21
   43 | PUT(pc) = 0x10000130
   44 | ------ IMark(0x10000130, 4, 0) ------
   45 | t23 = GET:I32(sp)
   46 | t22 = Add32(t23,0x00000018)
   47 | t3 = t22
   48 | t24 = GET:I32(ra)
   49 | STbe(t3) = t24
   50 | PUT(pc) = 0x10000134
   51 | ------ IMark(0x10000134, 4, 0) ------
   52 | PUT(ra) = 0x1000013c
   53 | t26 = And32(0x00000000,0x80000000)
   54 | t25 = CmpEQ32(t26,0x00000000)
   55 | t4 = t25
   56 | PUT(pc) = 0x10000138
   57 | ------ IMark(0x10000138, 4, 0) ------
   58 | t28 = GET:I32(sp)
   59 | t27 = Or32(t28,0x00000000)
   60 | PUT(s8) = t27
   61 | if (t4) { PUT(pc) = 0x10000180; Ijk_Call }
   62 | PUT(pc) = 0x1000013c
   63 | t29 = GET:I32(pc)
   NEXT: PUT(pc) = t29; Ijk_Call
}

After — n32 lifts 11 instructions over 44 bytes, statement for statement identical to the ArchMIPS64 lift beside it, with Ity_I64 present and arch.bits still 32:

with this change (2bf2cff)
pyvex package:  .../scratch/sh20-07/wt/mips-576/pyvex/__init__.py
libpyvex.so:    .../scratch/sh20-07/wt/mips-576/pyvex/lib/libpyvex.so
  sha256        9e5b9532fde2122351d37d89881db0673b4bb79f8da04ce78db466cdc5fd3535
archinfo:       .../scratch/sh20-07/wt/archinfo-375/archinfo/__init__.py

MIPSN32 in pyvex.lifting.libvex.LIBVEX_SUPPORTED_ARCHES: True

$ pyvex.lift(<72 bytes of main>, 0x10000110, archinfo.ArchMIPSN32(BE), opt_level=0)
  arch.bits=32  arch.vex_arch=VexArchMIPS64
  size=44  instructions=11  jumpkind=Ijk_Call  Ity_I64 present=True
  instruction_addresses=('0x10000110', '0x10000114', '0x10000118', '0x1000011c', '0x10000120', '0x10000124', '0x10000128', '0x1000012c', '0x10000130', '0x10000134', '0x10000138')
IRSB {
   t0:Ity_I64 t1:Ity_I64 t2:Ity_I64 t3:Ity_I64 t4:Ity_I64 t5:Ity_I64 t6:Ity_I64 t7:Ity_I64 t8:Ity_I64 t9:Ity_I1 t10:Ity_I64 t11:Ity_I32 t12:Ity_I32 t13:Ity_I64 t14:Ity_I64 t15:Ity_I64 t16:Ity_I64 t17:Ity_I64 t18:Ity_I32 t19:Ity_I32 t20:Ity_I64 t21:Ity_I32 t22:Ity_I64 t23:Ity_I64 t24:Ity_I32 t25:Ity_I32 t26:Ity_I64 t27:Ity_I64 t28:Ity_I64 t29:Ity_I64 t30:Ity_I32 t31:Ity_I64 t32:Ity_I64 t33:Ity_I64 t34:Ity_I64 t35:Ity_I32 t36:Ity_I32 t37:Ity_I64 t38:Ity_I64 t39:Ity_I64 t40:Ity_I64 t41:Ity_I64 t42:Ity_I64 t43:Ity_I1 t44:Ity_I64

   00 | ------ IMark(0x10000110, 4, 0) ------
   01 | t13 = GET:I64(sp)
   02 | t12 = 64to32(t13)
   03 | t11 = Add32(t12,0xffffffe0)
   04 | t10 = 32Sto64(t11)
   05 | PUT(sp) = t10
   06 | PUT(pc) = 0x0000000010000114
   07 | ------ IMark(0x10000114, 4, 0) ------
   08 | t15 = GET:I64(sp)
   09 | t14 = Add64(t15,0x0000000000000008)
   10 | t0 = t14
   11 | t16 = GET:I64(gp)
   12 | STbe(t0) = t16
   13 | PUT(pc) = 0x0000000010000118
   14 | ------ IMark(0x10000118, 4, 0) ------
   15 | PUT(gp) = 0x0000000000020000
   16 | PUT(pc) = 0x000000001000011c
   17 | ------ IMark(0x1000011c, 4, 0) ------
   18 | t20 = GET:I64(t9)
   19 | t19 = 64to32(t20)
   20 | t22 = GET:I64(gp)
   21 | t21 = 64to32(t22)
   22 | t18 = Add32(t21,t19)
   23 | t17 = 32Sto64(t18)
   24 | PUT(gp) = t17
   25 | PUT(pc) = 0x0000000010000120
   26 | ------ IMark(0x10000120, 4, 0) ------
   27 | t26 = GET:I64(gp)
   28 | t25 = 64to32(t26)
   29 | t24 = Add32(t25,0xffff8170)
   30 | t23 = 32Sto64(t24)
   31 | PUT(gp) = t23
   32 | PUT(pc) = 0x0000000010000124
   33 | ------ IMark(0x10000124, 4, 0) ------
   34 | t28 = GET:I64(gp)
   35 | t27 = Add64(t28,0xffffffffffff8020)
   36 | t1 = t27
   37 | t30 = LDbe:I32(t1)
   38 | t29 = 32Sto64(t30)
   39 | PUT(t9) = t29
   40 | PUT(pc) = 0x0000000010000128
   41 | ------ IMark(0x10000128, 4, 0) ------
   42 | t32 = GET:I64(sp)
   43 | t31 = Add64(t32,0x0000000000000010)
   44 | t2 = t31
   45 | t33 = GET:I64(s8)
   46 | STbe(t2) = t33
   47 | PUT(pc) = 0x000000001000012c
   48 | ------ IMark(0x1000012c, 4, 0) ------
   49 | t36 = 64to32(0x0000000000000000)
   50 | t35 = Add32(t36,0x00000005)
   51 | t34 = 32Sto64(t35)
   52 | PUT(a0) = t34
   53 | PUT(pc) = 0x0000000010000130
   54 | ------ IMark(0x10000130, 4, 0) ------
   55 | t38 = GET:I64(sp)
   56 | t37 = Add64(t38,0x0000000000000018)
   57 | t3 = t37
   58 | t39 = GET:I64(ra)
   59 | STbe(t3) = t39
   60 | PUT(pc) = 0x0000000010000134
   61 | ------ IMark(0x10000134, 4, 0) ------
   62 | t5 = 0x0000000000000000
   63 | t6 = GET:I64(s1)
   64 | t8 = 0x0000000000000000
   65 | PUT(ra) = 0x000000001000013c
   66 | t9 = CmpLT64S(t5,0x0000000000000000)
   67 | t40 = 1Uto64(t9)
   68 | t7 = t40
   69 | PUT(pc) = 0x0000000010000138
   70 | ------ IMark(0x10000138, 4, 0) ------
   71 | t42 = GET:I64(sp)
   72 | t41 = Or64(t42,0x0000000000000000)
   73 | PUT(s8) = t41
   74 | t43 = CmpEQ64(t7,t8)
   75 | if (t43) { PUT(pc) = 0x10000180; Ijk_Call }
   76 | PUT(pc) = 0x000000001000013c
   77 | t44 = GET:I64(pc)
   NEXT: PUT(pc) = t44; Ijk_Call
}

$ pyvex.lift(<72 bytes of main>, 0x10000110, archinfo.ArchMIPS64(BE), opt_level=0)
  arch.bits=64  arch.vex_arch=VexArchMIPS64
  size=44  instructions=11  jumpkind=Ijk_Call  Ity_I64 present=True
  instruction_addresses=('0x10000110', '0x10000114', '0x10000118', '0x1000011c', '0x10000120', '0x10000124', '0x10000128', '0x1000012c', '0x10000130', '0x10000134', '0x10000138')
IRSB {
   t0:Ity_I64 t1:Ity_I64 t2:Ity_I64 t3:Ity_I64 t4:Ity_I64 t5:Ity_I64 t6:Ity_I64 t7:Ity_I64 t8:Ity_I64 t9:Ity_I1 t10:Ity_I64 t11:Ity_I32 t12:Ity_I32 t13:Ity_I64 t14:Ity_I64 t15:Ity_I64 t16:Ity_I64 t17:Ity_I64 t18:Ity_I32 t19:Ity_I32 t20:Ity_I64 t21:Ity_I32 t22:Ity_I64 t23:Ity_I64 t24:Ity_I32 t25:Ity_I32 t26:Ity_I64 t27:Ity_I64 t28:Ity_I64 t29:Ity_I64 t30:Ity_I32 t31:Ity_I64 t32:Ity_I64 t33:Ity_I64 t34:Ity_I64 t35:Ity_I32 t36:Ity_I32 t37:Ity_I64 t38:Ity_I64 t39:Ity_I64 t40:Ity_I64 t41:Ity_I64 t42:Ity_I64 t43:Ity_I1 t44:Ity_I64

   00 | ------ IMark(0x10000110, 4, 0) ------
   01 | t13 = GET:I64(sp)
   02 | t12 = 64to32(t13)
   03 | t11 = Add32(t12,0xffffffe0)
   04 | t10 = 32Sto64(t11)
   05 | PUT(sp) = t10
   06 | PUT(pc) = 0x0000000010000114
   07 | ------ IMark(0x10000114, 4, 0) ------
   08 | t15 = GET:I64(sp)
   09 | t14 = Add64(t15,0x0000000000000008)
   10 | t0 = t14
   11 | t16 = GET:I64(gp)
   12 | STbe(t0) = t16
   13 | PUT(pc) = 0x0000000010000118
   14 | ------ IMark(0x10000118, 4, 0) ------
   15 | PUT(gp) = 0x0000000000020000
   16 | PUT(pc) = 0x000000001000011c
   17 | ------ IMark(0x1000011c, 4, 0) ------
   18 | t20 = GET:I64(t9)
   19 | t19 = 64to32(t20)
   20 | t22 = GET:I64(gp)
   21 | t21 = 64to32(t22)
   22 | t18 = Add32(t21,t19)
   23 | t17 = 32Sto64(t18)
   24 | PUT(gp) = t17
   25 | PUT(pc) = 0x0000000010000120
   26 | ------ IMark(0x10000120, 4, 0) ------
   27 | t26 = GET:I64(gp)
   28 | t25 = 64to32(t26)
   29 | t24 = Add32(t25,0xffff8170)
   30 | t23 = 32Sto64(t24)
   31 | PUT(gp) = t23
   32 | PUT(pc) = 0x0000000010000124
   33 | ------ IMark(0x10000124, 4, 0) ------
   34 | t28 = GET:I64(gp)
   35 | t27 = Add64(t28,0xffffffffffff8020)
   36 | t1 = t27
   37 | t30 = LDbe:I32(t1)
   38 | t29 = 32Sto64(t30)
   39 | PUT(t9) = t29
   40 | PUT(pc) = 0x0000000010000128
   41 | ------ IMark(0x10000128, 4, 0) ------
   42 | t32 = GET:I64(sp)
   43 | t31 = Add64(t32,0x0000000000000010)
   44 | t2 = t31
   45 | t33 = GET:I64(s8)
   46 | STbe(t2) = t33
   47 | PUT(pc) = 0x000000001000012c
   48 | ------ IMark(0x1000012c, 4, 0) ------
   49 | t36 = 64to32(0x0000000000000000)
   50 | t35 = Add32(t36,0x00000005)
   51 | t34 = 32Sto64(t35)
   52 | PUT(a0) = t34
   53 | PUT(pc) = 0x0000000010000130
   54 | ------ IMark(0x10000130, 4, 0) ------
   55 | t38 = GET:I64(sp)
   56 | t37 = Add64(t38,0x0000000000000018)
   57 | t3 = t37
   58 | t39 = GET:I64(ra)
   59 | STbe(t3) = t39
   60 | PUT(pc) = 0x0000000010000134
   61 | ------ IMark(0x10000134, 4, 0) ------
   62 | t5 = 0x0000000000000000
   63 | t6 = GET:I64(s1)
   64 | t8 = 0x0000000000000000
   65 | PUT(ra) = 0x000000001000013c
   66 | t9 = CmpLT64S(t5,0x0000000000000000)
   67 | t40 = 1Uto64(t9)
   68 | t7 = t40
   69 | PUT(pc) = 0x0000000010000138
   70 | ------ IMark(0x10000138, 4, 0) ------
   71 | t42 = GET:I64(sp)
   72 | t41 = Or64(t42,0x0000000000000000)
   73 | PUT(s8) = t41
   74 | t43 = CmpEQ64(t7,t8)
   75 | if (t43) { PUT(pc) = 0x10000180; Ijk_Call }
   76 | PUT(pc) = 0x000000001000013c
   77 | t44 = GET:I64(pc)
   NEXT: PUT(pc) = t44; Ijk_Call
}

$ pyvex.lift(<72 bytes of main>, 0x10000110, archinfo.ArchMIPS32(BE), opt_level=0)
  arch.bits=32  arch.vex_arch=VexArchMIPS32
  size=44  instructions=11  jumpkind=Ijk_Call  Ity_I64 present=False
  instruction_addresses=('0x10000110', '0x10000114', '0x10000118', '0x1000011c', '0x10000120', '0x10000124', '0x10000128', '0x1000012c', '0x10000130', '0x10000134', '0x10000138')
IRSB {
   t0:Ity_I32 t1:Ity_I32 t2:Ity_I32 t3:Ity_I32 t4:Ity_I1 t5:Ity_I32 t6:Ity_I32 t7:Ity_I32 t8:Ity_I32 t9:Ity_I32 t10:Ity_I32 t11:Ity_I32 t12:Ity_I32 t13:Ity_I32 t14:Ity_I32 t15:Ity_I32 t16:Ity_I32 t17:Ity_I32 t18:Ity_I32 t19:Ity_I32 t20:Ity_I32 t21:Ity_I32 t22:Ity_I32 t23:Ity_I32 t24:Ity_I32 t25:Ity_I1 t26:Ity_I32 t27:Ity_I32 t28:Ity_I32 t29:Ity_I32

   00 | ------ IMark(0x10000110, 4, 0) ------
   01 | t6 = GET:I32(sp)
   02 | t5 = Add32(t6,0xffffffe0)
   03 | PUT(sp) = t5
   04 | PUT(pc) = 0x10000114
   05 | ------ IMark(0x10000114, 4, 0) ------
   06 | t8 = GET:I32(sp)
   07 | t7 = Add32(t8,0x00000008)
   08 | t0 = t7
   09 | t9 = GET:I32(gp)
   10 | STbe(t0) = t9
   11 | PUT(pc) = 0x10000118
   12 | ------ IMark(0x10000118, 4, 0) ------
   13 | PUT(gp) = 0x00020000
   14 | PUT(pc) = 0x1000011c
   15 | ------ IMark(0x1000011c, 4, 0) ------
   16 | t11 = GET:I32(t9)
   17 | t12 = GET:I32(gp)
   18 | t10 = Add32(t12,t11)
   19 | PUT(gp) = t10
   20 | PUT(pc) = 0x10000120
   21 | ------ IMark(0x10000120, 4, 0) ------
   22 | t14 = GET:I32(gp)
   23 | t13 = Add32(t14,0xffff8170)
   24 | PUT(gp) = t13
   25 | PUT(pc) = 0x10000124
   26 | ------ IMark(0x10000124, 4, 0) ------
   27 | t16 = GET:I32(gp)
   28 | t15 = Add32(t16,0xffff8020)
   29 | t1 = t15
   30 | t17 = LDbe:I32(t1)
   31 | PUT(t9) = t17
   32 | PUT(pc) = 0x10000128
   33 | ------ IMark(0x10000128, 4, 0) ------
   34 | t19 = GET:I32(sp)
   35 | t18 = Add32(t19,0x00000010)
   36 | t2 = t18
   37 | t20 = GET:I32(s8)
   38 | STbe(t2) = t20
   39 | PUT(pc) = 0x1000012c
   40 | ------ IMark(0x1000012c, 4, 0) ------
   41 | t21 = Add32(0x00000000,0x00000005)
   42 | PUT(a0) = t21
   43 | PUT(pc) = 0x10000130
   44 | ------ IMark(0x10000130, 4, 0) ------
   45 | t23 = GET:I32(sp)
   46 | t22 = Add32(t23,0x00000018)
   47 | t3 = t22
   48 | t24 = GET:I32(ra)
   49 | STbe(t3) = t24
   50 | PUT(pc) = 0x10000134
   51 | ------ IMark(0x10000134, 4, 0) ------
   52 | PUT(ra) = 0x1000013c
   53 | t26 = And32(0x00000000,0x80000000)
   54 | t25 = CmpEQ32(t26,0x00000000)
   55 | t4 = t25
   56 | PUT(pc) = 0x10000138
   57 | ------ IMark(0x10000138, 4, 0) ------
   58 | t28 = GET:I32(sp)
   59 | t27 = Or32(t28,0x00000000)
   60 | PUT(s8) = t27
   61 | if (t4) { PUT(pc) = 0x10000180; Ijk_Call }
   62 | PUT(pc) = 0x1000013c
   63 | t29 = GET:I32(pc)
   NEXT: PUT(pc) = t29; Ijk_Call
}

Comment thread pyvex/arches.py Outdated
Comment thread pyvex/lifting/libvex.py Outdated
@ltfish ltfish self-assigned this Aug 29, 2026
@ltfish
ltfish merged commit 2654946 into master Aug 29, 2026
7 of 8 checks passed
@ltfish
ltfish deleted the feature/mips-n32-lifter branch August 29, 2026 14:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants