Skip to content

32-bit truncation/overflow in WAL length arithmetic wipes or panics on WALs near 4 GiB #560

Description

@yacovm

Details

The record framing code performs all length arithmetic in uint32 while file sizes are int64 and payload lengths are Go ints. Four wraparounds result, all in the crash-recovery path that runs at every node startup:

  1. wal.go ReadAll: readRecord(w.file, uint32(bytesToRead)) truncates the remaining file size to 32 bits. For ANY single WAL file >= 2^32 bytes, uint32(remaining) understates the true remaining size as reads approach a 4 GiB multiple: with near-certainty (probability ~1 - 12/recordSize) the next genuine record's payloadLen exceeds the truncated maxSize, readRecord errors, and ReadAll treats the perfectly valid record as a torn tail, calling truncateAt(fileSize - remaining). Everything past offset ~(fileSize - 2^32) — about 4 GiB of valid, fsynced consensus records — is silently destroyed and ReadAll returns nil error. When fileSize mod 2^32 is smaller than the first record's payload length, the failure occurs on the very first read and the ENTIRE log is truncated to zero. The consensus layer (simplex/epoch.go restoreFromWal) then starts as if the lost records never existed, so the node can re-vote in rounds it already voted in (equivocation risk).

  2. record.go readRecord: make([]byte, payloadLen+recordChecksumLen) wraps for payloadLen in [2^32-8, 2^32-1]; the buffer becomes 0..7 bytes, io.ReadFull succeeds, and payloadAndChecksum[:payloadLen] panics with slice bounds out of range. The payloadLen > maxSize guard does not prevent this because maxSize is itself the truncated uint32 cast: when the remaining size's low 32 bits are >= 2^32-8 (file within 8 bytes below a 4 GiB multiple), a corrupted or maliciously written length field converts the intended graceful truncate-and-recover behavior into a deterministic panic on every startup.

  3. record.go readRecord return value: recordSizeLen + payloadLen + recordChecksumLen wraps in uint32 for payloadLen >= 2^32-12, so ReadAll's bytesToRead accounting drifts and a later truncateAt offset falls inside valid records (requires a crafted ~4 GiB record with valid CRC, i.e. file tampering).

  4. record.go writeRecord: uint32(len(payload)) silently truncates the length prefix for payloads >= 4 GiB, writing a self-inconsistent record that fails CRC on replay and poisons the log from that point.

Reachability: in this repository's wiring (instance.go:429) the file WAL is always wrapped by GarbageCollectedWAL, which rotates once accumulated payload bytes exceed maxWalSize — default 100 MB (wal/gc.go) when ParameterConfig.WALMaxEntryCount is 0 — so a single file reaches 4 GiB only if the embedder passes a cap above 2^32 bytes or uses the exported wal.WriteAheadLog directly (it enforces no size bound). The wipe variant (1) needs no file tampering — only a >= 4 GiB file and a restart; variants 2 and 3 additionally require a corrupted/crafted length field on disk. Dominant impacts: silent wholesale destruction of durable consensus state (divergent replay/equivocation risk) and a persistent startup crash loop.

Evidence

  1. wal/wal.go:82–94
    bytesToRead is the int64 file size; it is truncated to uint32 when passed to readRecord as maxSize. For ANY file >= 2^32 bytes the truncated maxSize understates the remaining size: just before the remaining size crosses a 4 GiB multiple, the next valid record's payloadLen near-certainly exceeds uint32(remaining), the record is rejected as corrupt, and truncateAt(fileInfo.Size() - bytesToRead) silently destroys everything from that offset onward (~4 GiB of fsynced records; the ENTIRE log when fileSize mod 2^32 is below the first record's payload length) while returning a nil error.
  2. wal/record.go:59–63
    payloadLen is a uint32 read from disk. payloadLen + recordChecksumLen (8) is computed in uint32 arithmetic: for payloadLen >= 2^32-8 it wraps, so make() allocates a tiny buffer (e.g. 7 bytes for payloadLen 0xFFFFFFFF), ReadFull succeeds, and payloadAndChecksum[:payloadLen] panics with slice bounds out of range. The line 55 payloadLen > maxSize check does not prevent this because maxSize is itself the truncated cast: reachable when the remaining size's low 32 bits are >= 2^32-8 (file within 8 bytes below a 4 GiB multiple) and the on-disk length field is corrupted/crafted to a huge value. The node then crash-loops on every startup instead of performing its intended graceful truncation recovery.
  3. wal/record.go:73
    The returned bytesRead = recordSizeLen + payloadLen + recordChecksumLen also wraps in uint32 for payloadLen >= 2^32-12, causing ReadAll's bytesToRead accounting to drift and the eventual truncateAt offset to fall inside valid records (mistruncation of durable data). Requires a crafted ~4 GiB record whose CRC verifies, i.e. local file tampering.
  4. wal/record.go:28
    Write side: uint32(len(payload)) silently truncates the length prefix for payloads >= 4 GiB, producing a record whose declared length mismatches its data; on replay it fails CRC and everything from that record onward is truncated away. Unreachable with realistic consensus record sizes but the same root-cause family in the exported API.

Impact

Integrity: an entirely valid durable WAL is silently truncated to zero (or mistruncated mid-log) with a success return, rolling back consensus safety state and enabling divergent replay/equivocation. Availability: the payloadLen+8 wrap turns recovery into a deterministic slice-bounds panic on every startup, keeping the node down until manual file surgery.

Reproduction steps

  1. Triggered when the node restarts and replays its WAL. Requires a single WAL file near/above 4 GiB, which the wired GarbageCollectedWAL prevents by default (100 MB payload rotation) — needs an embedder-supplied cap above 2^32 bytes or direct use of the exported WriteAheadLog, hence AT PRESENT. The whole-log wipe variant needs no corruption: only file size >= 2^32 and a restart. The panic and mistruncation variants additionally need a corrupted/crafted length field on disk. No privileges or user interaction; the trigger is local on-disk state, not a network-deliverable input (AV LOCAL).

Recommended fix

  1. Remaining-file-size and record-length arithmetic is performed in uint32 while actual sizes are 64-bit, so casts and additions wrap: uint32(bytesToRead) in ReadAll, payloadLen+recordChecksumLen and the bytesRead sum in readRecord, and uint32(len(payload)) in writeRecord. Fix criteria: All framing arithmetic must be carried out in a width that cannot wrap for any representable file/payload size (e.g. int64 with explicit bounds checks), and declared payload lengths must be validated against the true remaining size minus header/checksum overhead before allocation or slicing. Verified by replay tests with file sizes straddling 4 GiB and with length fields of 0xFFFFFFF0-0xFFFFFFFF: no panic, no truncation of valid records, oversize declarations rejected as corruption at the correct offset.
  2. writeRecord silently truncates the length prefix for payloads >= 4 GiB, producing an unreadable record. Fix criteria: Append must reject payloads whose length cannot be represented in the record length field, returning an error instead of writing a malformed record. Verified by unit test that an oversized payload is refused and the log remains readable.

Severity: MEDIUM
Status: Open
Category: Integer overflow
CWE: CWE-197
Repository: ava-labs/Simplex
Branch: main
Date created: 2026-08-21


Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions