Skip to content

dat: show inferred symbols in objdiff, add directories - #3597

Merged
ribbanya merged 2 commits into
doldecomp:masterfrom
ribbanya:pr/dat
Oct 1, 2026
Merged

ribbanya merged 2 commits into
doldecomp:masterfrom
ribbanya:pr/dat

Conversation

@ribbanya

@ribbanya ribbanya commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Units are grouped by module, the first two letters of the archive's name:
every build directory holds / (Pl/PlMr), and objdiff
names them dat/Pl/PlMr, like the code's main/.

The target objects held only the samples, so units showed every symbol
matching while not complete. Now they hold the whole archive in two
sections, which objdiff lists by name:

  • .0.sampled: the sampled objects, with pointers as relocations
  • .1.inferred: the archive's other global symbols (its publics, and the
    data the samples point to), uninitialized, so objdiff never diffs
    their bytes

The base defines the globals the walk explains in its own .1.inferred
(target/.rest.o, linked with the C), so objdiff pairs them by name
and matches them by size; the rest show as missing. A global is
explained when it and everything it reaches before another global or a
sample is typed data with no unexplained relocation. Each target
symbol's offset in the archive is its virtual address in a .note.split,
as decomp-toolkit writes for split code, which objdiff shows. objdiff
takes section kinds from their ELF type, so the names are free; the
numbers set the order.

Raw u8 data says nothing of its format, so it isn't explained. DAT_BLOB
on a u8 typedef names one: HSD_FObjData, the keyframe streams of
HSD_FObjDesc and FigaTrack, sized by DAT_COUNT(length). DAT_TERMINATED
compares elements smaller than a word whole (FigaTree.nodes, up to -1).
The walk records how far each object's typed data reaches, so counted
and DAT_EXTENT arrays count whole. Generated C declares untyped data as
UNK_T instead of a DatBlob array.

@ribbanya ribbanya added tooling ai-assisted Utilizes a LLM to do the heavy lifting portability Improves non-matching builds labels Oct 1, 2026
@ribbanya
ribbanya marked this pull request as ready for review October 1, 2026 03:53
Units are grouped by module, the first two letters of the archive's name:
every build directory holds <module>/<archive> (Pl/PlMr), and objdiff
names them dat/Pl/PlMr, like the code's main/.

The target objects held only the samples, so units showed every symbol
matching while not complete. Now they hold the whole archive in two
sections, which objdiff lists by name:

- .0.sampled: the sampled objects, with pointers as relocations
- .1.inferred: the archive's other global symbols (its publics, and the
  data the samples point to), uninitialized, so objdiff never diffs
  their bytes

The base defines the globals the walk explains in its own .1.inferred
(target/<unit>.rest.o, linked with the C), so objdiff pairs them by name
and matches them by size; the rest show as missing. A global is
explained when it and everything it reaches before another global or a
sample is typed data with no unexplained relocation. Each target
symbol's offset in the archive is its virtual address in a .note.split,
as decomp-toolkit writes for split code, which objdiff shows. objdiff
takes section kinds from their ELF type, so the names are free; the
numbers set the order.

Raw u8 data says nothing of its format, so it isn't explained. DAT_BLOB
on a u8 typedef names one: HSD_FObjData, the keyframe streams of
HSD_FObjDesc and FigaTrack, sized by DAT_COUNT(length). DAT_TERMINATED
compares elements smaller than a word whole (FigaTree.nodes, up to -1).
The walk records how far each object's typed data reaches, so counted
and DAT_EXTENT arrays count whole. Generated C declares untyped data as
UNK_T instead of a DatBlob array.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ribbanya
ribbanya enabled auto-merge (squash) October 1, 2026 04:36
@ribbanya
ribbanya merged commit 082daf7 into doldecomp:master Oct 1, 2026
10 checks passed
@ribbanya
ribbanya deleted the pr/dat branch October 1, 2026 04:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai-assisted Utilizes a LLM to do the heavy lifting portability Improves non-matching builds tooling

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant