Skip to content

hl-node RSS grows ~10-12 GiB/h on binary index 111; index 110 was stable for 6 days. child_low_memory restart now fires every 1.5-9 h #173

Description

@susuper

Summary

After hl-visor upgraded our Mainnet non-validator from binary index 110 to index 111 on
2026-09-19T08:34:19Z, hl-node RSS began growing without bound at roughly 10-12 GiB/h until
hl-visor's own child_low_memory guard restarts it at memory_usage ≈ 0.95. The node self-heals
and follows consensus fine, so this is not an outage — but the host is now cycling to 95% memory
several times a day where it previously ran for weeks.

The growth is 99.9% private dirty anonymous memory and is fully reclaimed by the restart, so it
looks like retained allocations rather than a larger steady-state working set.

Related: #150 and #147 (memory growth / allocator) and #136 (the child_low_memory policy itself,
whose questions are still unanswered). This report adds the thing those don't have: a version
boundary
.

The version boundary

Complete child history from hl-visor's own stdout log, which covers 2026-09-12T12:08 onward:

child start (UTC) binary index why
2026-09-12T12:08:30 109 visor start
2026-09-13T08:37:05 110 binary upgrade
2026-09-19T08:34:19 111 binary upgrade
2026-09-20T14:20:16 111 child_low_memory: true, memory_usage: 0.9502
2026-09-20T15:32:11 111 host reboot (our OS upgrade, see Confound)
2026-09-21T00:15:52 111 child_low_memory: true, memory_usage: 0.9565
2026-09-21T07:11:25 111 child_low_memory: true, memory_usage: 0.9500
2026-09-21T10:19:46 111 child_low_memory: true, memory_usage: 0.9568
2026-09-21T13:31:33 111 child_low_memory: true, memory_usage: 0.9504
2026-09-21T15:02:26 111 child_low_memory: true, memory_usage: 0.9502
  • index 110: 5 d 23 h 57 m, zero low-memory restarts.
  • index 111: 2 d 9 h, six of them. child_stuck: false on every one; child_stuck: true appears zero times anywhere in the log window.

Restart trigger line (one line in the log, reflowed here):

2026-09-21T15:02:20.441Z ERROR >>> hl-visor @@ ... visor child in bad state, restarting
  @@ [child_stuck: false] @ [child_low_memory: true]
  @ [memory_usage: 0.9502360915579388] @ [child_running: true]
  @ [visor_abci_state: Ok(VisorAbciState { initial_height: 1153141000, height: 1155885302, ... })]

Measured growth

48 samples at 60 s, 2026-09-21 16:16:35Z -> 17:03:36Z, taken from /proc/<pid>/status and the
service cgroup. The child had started at 15:02:26.

net:  RSS 35.50 -> 44.63 GiB  = +9.13 GiB in 47 min = +11.7 GiB/h
      MemAvailable 52.02 -> 42.9 GiB (host total 93.4 GiB)
  • Inter-spike baseline: +250, +248 MiB/min early in the window, decaying to +136..+209 MiB/min
    by 17:00 — decaying, but never flat.
  • On top of that, a sawtooth from abci state serialization: single-minute jumps of
    +1972, +1443, +4335, +1977, +4280 MiB, each unwinding the next minute
    (-1556, -997, -3762, -1595, -3904 MiB) and leaving the trend line intact — e.g. across
    16:18:35 -> 16:20:35 the net is +420 MiB, at or below the two-minute baseline. These line up with
    serialized abci state @@ [elapsed: Duration(3.x-4.x)] and the 1.60 GiB .rmp written every
    ~12 min, so the sawtooth is expected — it just rides on top of the trend and adds ~4 GiB to the
    peak that trips the 0.95 guard.
  • Whole cycle: restart at 15:02:26 -> VmRSS 50.6 GiB, VmHWM 53.6 GiB at 18:03:23 (3 h 01 m).
  • Service cgroup memory.peak since the last visor start: 87.6 GiB of 93.4 GiB (93.7%),
    memory.swap peak 2.0 GiB.

Memory profile

smaps_rollup, single snapshot at 2026-09-21T18:11:21Z, RSS 49.98 GiB, child started 15:02:26Z:

Rss:             49.98 GiB
Private_Dirty:   49.93 GiB
Pss_File:         0.04 GiB
AnonHugePages:    0.00 GiB
Swap:             0.00 GiB

99.9% private dirty anon: no file-backed growth, no THP, nothing swapped.

366 rw-anon mappings, of which the 18 largest hold 41.4 GiB of the 50.0 GiB. So this is a
modest number of retained multi-GiB allocations, not a proliferation of small glibc arena heaps —
the remaining ~348 mappings average 25 MiB.

What actually changed as RSS went 37.6 -> 50.0 GiB (same process, 16:14Z -> 18:11Z):

  • The largest mapping is flat. 7a7356a00000-7a75d0000000: 7950 -> 7932 MiB.
  • The second largest oscillates rather than grows. 7a7131200000-7a7353c00000 measured
    5729 / 7364 / 5724 MiB at 16:14 / 18:03 / 18:11. The ~1.6 GiB swing tracks abci state
    serialization, so this one behaves like a working buffer.
  • The growth is in nine mappings in an address range where nothing exceeded 195 MiB at 16:14
    (0x7a62…-0x7a67…). They now hold 4016, 2441, 2174, 2075, 1207, 1194, 1002, 949 and 934 MiB —
    ~15.6 GiB combined, partly offset by decreases elsewhere (7a6b55a80000 2585 -> 1600 MiB),
    against +12.4 GiB of net RSS growth over the window.
  • Mappings that already existed barely move: +29, +25, +4 MiB.

New large mappings keep appearing while existing ones stay put. That is the shape of allocating a
fresh large structure per unit of work and retaining the previous one, rather than fragmentation
inside existing arenas.

Allocator, since #147 and #150 both asked and neither got an answer:

linked libs: ld-linux-x86-64.so.2, libc.so.6, libcrypto.so.3, libgcc_s.so.1,
             libm.so.6, libssl.so.3, libstdc++.so.6.0.33

Default glibc malloc, no jemalloc/mimalloc, MALLOC_ARENA_MAX unset, 62 threads on 24 cores.

Ruled out

  • Not a kernel OOM kill. dmesg and the journal have no OOM record; the service cgroup reports
    oom 0 / oom_kill 0. Every visor_child_stderr/* file is 0 bytes — the child exits silently
    because hl-visor terminates it, nothing panics. Note that hl-visor's own kill reports failure
    first and restarts anyway, on every occurrence:
    ERROR >>> hl-visor @@ could not kill child process: Shell::wait_inner failed, cmd: (pkill run-client).
  • No cgroup limit involved. MemoryMax=infinity, MemoryHigh=infinity.
  • Not THP. /sys/kernel/mm/transparent_hugepage/enabled = [madvise], AnonHugePages: 2048 kB
    host-wide, 0 in the process.
  • Not swap pressure. vm.swappiness=1, 64 GiB swap, the cgroup has never used more than 2.0 GiB.
  • Not io_uring. Zero io_uring fds.
  • Not gossip egress shaping. We rate-limit our own gossip serving with nftables; its counters
    were byte-identical across two reads ~2 min apart, during which RSS climbed ~300 MiB. Daily
    distinct peers did not fall either (243-508 per day across the week, spanning both sides of the
    change).

Confound I cannot separate

On 2026-09-20T15:23Z — after the first low-memory restart but before the other five — the host
took an apt upgrade (kernel 6.8.0-137 -> 6.8.0-139, glibc 2.39-0ubuntu8.8 -> 8.9) and rebooted at
15:30. So I can attribute the onset to index 111 (the 110 -> 111 boundary predates the upgrade
by a day), but not the acceleration from one restart in 30 h to five per day. If the underlying
mechanism is allocator behaviour, a glibc patch landing at exactly that point is at least worth
noting.

Environment

  • Ubuntu 24.04.5 LTS, kernel 6.8.0-139-generic, 24 vCPU, 93.4 GiB RAM, NVMe
  • hl-visor run-non-validator --serve-eth-rpc --replica-cmds-style recent-actions --disable-output-file-buffering
  • hl-node 8b149bdcc879c1e7ea83b29204af240a883c85c5 | 2026-09-19 15:31:57 +0800 (index 111)
  • hl-visor a1524269af1832d2c4fbdb5259779b889c444bfa | 2026-06-20 16:02:44 +0800
  • override_gossip_config.json: 12 root IPs, try_new_peers: false
  • Steady co-tenant load on --serve-eth-rpc (:3001): ~28 eth_call/s, measured over the last
    24 h. I have no measurement of this rate from before 2026-09-20, so I cannot myself rule out a
    load change as a contributor.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions