Skip to content

GC nursery pacing (#11645): accepted regressions to recover #11699

Description

@proggeramlug

The owner accepted these regressions on 2026-09-30 when landing #11645 ("land 11645 now, but let's have an issue for the small regression with as many details as we can right now"). This issue tracks getting them back. Part of #11549.

#11645 (survival-aware nursery pacing) starts the scavenge nursery ladder at 4 MB instead of 16 MB. That wins big on package workloads: dotenv −30% instructions, and validator, date-fns, jwt and uuid −20–30% peak RSS. It costs small or fully-live programs one or more extra minors. The rows below are the ones that got worse.

Numbers (main vs PR)

row Δ instructions Δ peak RSS minors/fulls, main → PR
dotenv_parse −30.5% −20.3% 0/7 → 61/0
moment_parse_format −20.6% +3.1% (21 reps; +5.4% at 11 reps) 1/33 → 43/0
validator_batch −3.7% −26.0% 46/0 → 211/0
date-fns_format_add −1.0% −30.0% 17/0 → 71/0
qs_parse_nested +0.36% +1.2% (ranges overlap) 11/0 → 14/0
qs_stringify_nested −0.3% +0.6% 24/10 → 24/10
jsonwebtoken_decode −3.6% −22.0% 2/0 → 17/0
uuid_v4 −0.1% −21.0% 1/0 → 6/0
alloc loop +0.57% −20.0% 67/0 → 203/0
binary-trees n=3 +746% (24.95 M → 211.09 M) +63.8% (15.5 → 25.3 MB) 0/0 → 1/0
binary-trees n=6 +649% (28.69 M → 214.83 M) +59.0% 0/0 → 1/0
binary-trees n=10 +553% (33.68 M → 219.81 M) +54.7% 0/0 → 1/0
binary-trees n=20 −4.0% −8.9% 1/0 → 2/0
gc_ratchet 01_nursery_churn −4.5% −25.5% 1/1 → 4/1
gc_ratchet 02_survivor_promotion +17.3% (276.8 M → 324.6 M) −11.6% 1/1 → 2/1
gc_ratchet 12_large_live_set −5.7% +2.6% (105.6 → 108.4 MB) 5/1 → 6/1

The rows in bold are the accepted regressions.

Method

Per-row mechanisms

  • binary-trees n=3/6/10. The program builds a ~5 MB tree that stays fully live and exits after 7–9 MB of allocation. Main never collects it. The PR runs one minor at the 4 MB floor, and that minor promotes the whole tree in place: about +186 M instructions. Earlier profiling put this minor's cost at ~120 M of trace, which is the collector's general per-object visit/classify cost (visit_gc_layout_slot_descriptors, scan_object_fields, classify_arena, visit_slot_with_weak_fact), plus ~33 M of non-trace fixed setup. Main pays the same per-object cost from n≥20 on, where n=20 is −4.0% here. The extra RSS comes from the minor's worklist and moved_headers doubling from a zero estimate. A never-reallocating header list was tried on wip/11549-header-list-experiment: it cut about 3 MB, but traded RSS elsewhere under THP (binary-trees 10→40 +7%, gc_ratchet 01 +3.6%), so it was not landed.
  • gc_ratchet 02 (+17.3%). The PR runs 2 minors where main runs 1. Minor 2 re-copies minor 1's surviving cohort, and the final gc() full then sweeps a young generation still full of garbage. The old perf(gc): survival-aware nursery pacing, 4 MB floor on the existing influx ladder (includes #11612) #11645 table showed +4.3%, but that compared main without perf(gc): the first collection's barrier-arming walk skips wholly-nursery blocks #11668 against the PR with it. With perf(gc): the first collection's barrier-arming walk skips wholly-nursery blocks #11668 on main (d7df6e7), this is the second minor's true cost.
  • gc_ratchet 12 (+2.6% RSS). One extra early minor (6 vs 5) promotes part of the large live set earlier than main does.
  • alloc loop (+0.57%). 203 minors against 67. Each minor has a fixed cost of ~100 k instructions (perf(gc): cut a copying minor's fixed cost (intern young log, skip empty array-tail tables, young-only prunes) #11634 line of work). The 136 extra minors × ~106 k account for the whole delta. To break even, the fixed cost would have to fall to ~35 k per minor.
  • qs_parse_nested (+0.36%). 14 minors against 11. See the qs prunes below.

Unexplained: moment's RSS shift

On 10ece9958 (the previous #11645 measurement), moment's THP-off RSS was main 40.8 MB against PR 35.3 MB (−13.5%). On 7fa094cb4, main is 39.6 MB (range 37.7–45.2) and the PR is 40.8 MB (range 40.0–41.2). Main barely moved; the PR's median rose ~5.5 MB. Minor and full counts are unchanged (1/33 against 43/0), and the −20.6% instruction win still holds.

Nothing has been bisected. The candidates are the GC- or rooting-touching commits in 10ece9958..7fa094cb4:

The first step is to bisect the PR arm over this range with MIMALLOC_ALLOW_THP=0, looking at moment's peak RSS.

Expected help from #11676 (per-object trace cost)

#11676 cuts the per-object trace ~35%. With #11645 stacked on it earlier, binary-trees came out at +474/+412/+351% (n=3/6/10) instead of +746/+649/+553%, and 02 at −2.8% instead of +17.3%. So once #11676 lands, 02 should drop off this list, and binary-trees should roughly halve.

Candidate next steps

  1. Prune the fixed per-minor work that qs's minors spend time in. Shares of qs minor time: closure box-capture prune 16%, layout-owner prune/sort 12%, remembered-set rebuild 12%. Making these incremental or skipping them when nothing changed helps qs, the alloc loop and every high-minor-count row.

  2. Cut the first minor's fixed setup (the ~33 M of non-trace work in binary-trees' single minor).

  3. A survival-aware first-minor policy. At the first minor, survival alone does not separate binary-trees' fully-live tree from a package's startup cohort, which is 10–31% alive (0.4–1.3 MB). Measured alternatives:

    • Powering on at the 16 MB base fixes binary-trees and 02 but costs dotenv +22.6% and moment +25.1% RSS.
    • An 8 MB start is a size-specific fit that costs 02 +20.6%.

    A policy would need a signal beyond first-minor survival, e.g. live bytes relative to allocated bytes, or deferring the promotion decision.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugConfirmed defect or regressionperformanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions