Skip to content

perf: OpenCode native binary startup must beat the bun binary — attribute and remove the module-init cost (measured 13× bun CPU on the Aug 21 build) #10106

Description

@proggeramlug

Parent: #10107. Related: PERRY_STARTUP_PLAN.md (PR #10066 tiny-path work), #6532/#6533/#6534 (cc --version 227 ms), #9441 (idle tail), scavenge/tenuring notes.

Measured 2026-09-12 (dev Mac, load avg ~100 — wall inflated; CPU is the fair number)

binary --version wall user CPU max RSS
official bun-compiled opencode 1.17.7 1.0–1.5 s 0.54 s 185 MB
bun run src/index.ts --version (transpiles on the fly) 2.8 s 0.85 s —
perry Aug-21 binary (minimal profile, 2,454 modules, 213 MB, perry 0.5.1512) 15–29 s 6.8–7.0 s 563 MB
same, --help 14.8 s 7.0 s —
A sample of the stripped binary shows the main thread 100 % in generated code (no wait states), i.e. eager module initialization of the graph (Effect's 225 modules build Schema/Layer/Context objects at import; drizzle tables; yargs command tree). No startup number exists for the full graph (4,063 modules on v1.18.30; 7,136 on the Windows full build #9133), so expect worse before optimization.

Target

On the quiet Mac mini (perry@perry-macos.local), same host, five warm runs each, paired A/B per PERRY_STARTUP_PLAN S1 discipline:

  • opencode --version: user CPU ≤ 0.5 s, wall ≤ 1.0 s, RSS ≤ 185 MB (beat the official v1.18.30 binary from gh release download v1.18.30 -R anomalyco/opencode -p 'opencode-darwin-arm64.zip').
  • opencode --help and TUI time-to-first-frame: below bun.

Scope

  1. Attribute: symbols build (strip=none, PERRY_MAIN_STACK_MB if needed), sample/Instruments over --version, plus PERRY_DEBUG_INIT module-init trace; produce a per-module init-cost table (top 50) and a per-runtime-helper table (class registration, closure allocation, object-literal construction, string interning, regex compile — regex construction is eager and SipHash-keyed per pattern in perry, see the regex analysis note in secret-tests memory, GC minors during init).
  2. Lazy work: anything perry does per module that bun does not (e.g. eager class/shape registration, module-global slot promotion, side-table setup) becomes lazy or batched; check the young-gen / tenuring configuration during init (563 MB RSS vs 185 MB suggests allocation volume, not just code size).
  3. Graph-level: confirm the compiled graph does not initialize modules bun would never evaluate (dynamic-import-only subgraphs such as Code Mode/typescript, prettier, @babel/* via the runtime Solid plugin — must not be in the eager init order; fix(compile): support full OpenCode source builds #9133 kept dynamic imports out of eager init, verify on this graph).
  4. Binary size / page-in: 213 MB → the first-run variance (16→29 s) hints at page-in; measure cold vs warm and consider PERRY_LL_SIZE_OPT=1 + optnone threshold (perf(compile): reduce generated bundle bloat #8418) trade-offs.
  5. Re-measure after each lever with the A/B script; keep raw samples in secret-tests/opencode-1.18.30-inventory/perf/.

Acceptance

  • Attribution report with the top-50 module-init costs and the runtime-helper breakdown, on the full v1.18.30 graph.
  • --version user CPU and RSS below the official bun binary on the Mac mini, two independent batches.
  • TUI first frame not slower than bun.
  • No correctness regressions: the 22 node startup oracles (PERRY_STARTUP_PLAN) and the cc --help parity gate stay green.

Activity

  1. added
    enhancementNew capability or improvement
    performanceRuntime, compile-time, build-size, or memory performance
    on Sep 12, 2026
  2. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Startup attribution, first hard number (2026-09-13, before the symbols profile)

    The --help family is not init cost: it is string-width. yargs 18 formats help through cliui → string-width@7.2.0, and a 9-module perry probe calling stringWidth() on three typical help lines 3,000 times each (9,000 calls) takes

    perry (linux-x64, --no-auto-optimize) bun 1.3.14
    9,000 stringWidth() calls 340,204 ms (≈38 ms per call) 72 ms

    GC is not the issue there (share_permille=13, arena 0.4 MB) — it is pure CPU inside the width computation: string-width 7 iterates every grapheme with Intl.Segmenter and, per grapheme, tests emojiRegex() (constructs the ~14 KB emoji regex inside the loop, once per character) plus /^\p{Default_Ignorable_Code_Point}$/u. A finer probe separating regex construction, regex .test, the segmenter and strip-ansi is running; whichever it is, this single path explains opencode --help at 31 s / run --help at 20 s (the time scales with help text) and is the first lever for this ticket. The 1.6 GB live arena seen under PERRY_GC_DIAG on --help is a separate observation to attribute with the symbols build.

  3. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Finer attribution of the string-width cost (perry vs bun per operation): emojiRegex() construction 631 µs vs 2.3 µs (274×), reused .test 11 µs vs 0.4 µs (28×), strip-ansi replace 14 µs vs 0.1 µs (143×), Intl.Segmenter line 68 µs vs 12 µs, eastAsianWidth 10×. Construction per grapheme dominates. Filed as #10179.

  4. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    perf profiles on the debug-symbols build (linux-x64, perf record -F 499 -g)

    --help (27 s user): regex construction and matching plus the GC they cause.

    --version (2.0 s user): flat, no single hot symbol. Top entries are generic by-name property access during module initialization — get_field_by_name_object_tail 2.3 %, keys_find_slot_by_bytes 1.9 %, shape_descriptor_ensure_with_holes 1.7 %, js_object_get_field_by_name 1.6 %, get_accessor_descriptor 1.2 % — plus core::str::from_utf8 1.7 %, GC layout/trace bookkeeping ≈ 6 %, js_array_get_f64/js_array_length 1.8 %. In other words: 7,897 module initializers (#10180) each doing dictionary-style property reads and writes on freshly built objects, not one pathological site. The levers are the module count (#10180) and cheaper object construction/property access in init code (shape-known stores instead of by-name lookups; from_utf8 on every string constant materialization is worth a look).

  5. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Status 2026-09-13 evening: --help fix PR #10193 (regex construction cache, stringWidth 41.8 ms → 0.39 ms) is rebased onto 0.5.1557 and waiting for the merge train; the module-pruning ask #10180 (7,897 → toward bun's 4,063 modules, the main lever on the 1.9 s --version init cost) has a codex lane on it. Current numbers on perrymaster with the pinned compiler: --version 1.9 s user (bun 0.35), --help 28 s (bun 0.37), models --help 6.7 s.

  6. proggeramlug commented on Sep 14, 2026

    @proggeramlug
    ContributorAuthor

    Measured 2026-09-15 on the full v1.18.30 graph (6,926 modules, not the minimal profile)

    Same Mac, five warm runs each, --version. The binary is the lineage-8 build (perry 0.5.1568-era, PERRY_LL_SIZE_OPT=1, 855 MB darwin-arm64) and the oracle is the official opencode-darwin-arm64 release binary.

    perry official bun binary ratio
    wall (warm median) 2.32 s 0.32 s 7.3×
    user CPU 2.10 s 0.41 s 5.1×
    sys CPU 0.21 s 0.02 s —
    instructions retired 25.26 G 3.77 G 6.7×
    cycles elapsed 7.27 G 1.37 G 5.3×
    max RSS 548 MB 182 MB 3.0×
    page faults 1,460 70 —

    The instruction count is the number that matters: this is not I/O, page-in of an 855 MB image, or dyld. It is 21.5 G extra instructions of real work before --version prints.

    Where it comes from

    crates/perry-codegen/src/codegen/entry.rs (the entry main emitter) ends its init prelude with

    for (index, prefix) in non_entry_module_prefixes.iter().enumerate() {
        if cross_module.deferred_module_prefixes.contains(prefix) { continue; }
        blk.call_void(&format!("{}__init", prefix), &[]);
    }

    so every non-entry module's __init runs eagerly at startup, and the only exemption is deferred_module_prefixes. That set is populated in crates/perry/src/commands/compile/run_pipeline.rs from ModuleInitKind::Deferred, which per the #753 doc comment means reachable from the entry only through dynamic import() edges. A module reached by even one static edge anywhere in the graph is eager, even when the program never calls into it.

    OpenCode is exactly the adversarial shape for that rule. packages/opencode/src/index.ts statically imports all 25 command modules to build the yargs tree, and those command modules then defer their heavy work behind await import(...) inside the handler (src/cli/cmd/tui.ts dynamically imports effect, ../tui/layer and the plugin host). Under bun the handler's imports never run for --version; under perry any of those subgraphs that is also statically reachable from some other module is initialized before argv is parsed.

    Suggested next step for whoever takes this

    Attribution before fixes, per the scope already in this issue: a PERRY_DEBUG_INIT build that emits the eager init order, turned into a per-module init-cost table, so we can say how many of the 6,926 modules bun never evaluates for --version. The graph-level lever (point 3 in the scope) now looks like the dominant one rather than a side item, and the general form of it — init-on-first-cross-module-use rather than init-everything-at-entry — is what the target in this issue requires.

    Tracker: #10107.

  7. proggeramlug commented on Sep 15, 2026

    @proggeramlug
    ContributorAuthor

    Correction to my comment above

    A code read from the Proxy/CJS lane work, which I have since confirmed against init_order.rs, shows I framed the mechanism too strongly. Retracting the part that matters and keeping the part that is measured.

    What I got wrong. I implied perry eagerly initializes subgraphs bun would never evaluate, and that this is the dominant lever. classify_eager_modules already exempts more than dynamic-import-only modules:

    .filter(|i| !i.is_dynamic && !i.type_only && !i.runtime_erased && !i.is_deferred_require)

    so function-local and conditional require() edges do not chain into a module's init either. The 2,239 eager modules are therefore mostly genuine static ESM imports — and bun evaluates a static ESM graph at startup too. Its bundle with splitting: true hoists the static graph exactly the same way; dynamic imports become separate lazily-loaded chunks, which is the analogue of perry's Deferred set. So "perry initializes modules bun skips" is not established by the eager/deferred split alone.

    What stands, because it is measured rather than inferred. On the full v1.18.30 graph, --version costs 25.26 G instructions against the official binary's 3.77 G, with 548 MB peak RSS against 182 MB, and 2,239 of 6,949 modules initialize before argv is parsed. The gap is real and it is CPU, not I/O.

    Where the difference is more likely to live, and what I would attribute before changing anything: bun tree-shakes and minifies, so the module count that survives into its bundle is much smaller than 6,949 to begin with; its __commonJS wrappers make CJS module bodies lazy until first require; and its functions are compiled lazily, where an AOT binary has already paid that cost differently. A per-module init-cost table from a symbolized build is still the right first deliverable — it just needs to be read against bun's surviving module set, not against perry's total.

    A colleague is taking this issue with a symbolized build (PERRY_KEEP_SYMBOLS=1 is not an object-cache key, so it relinks an existing compile rather than rebuilding it) and will post the attribution. Their sample of a --version run so far: ~19% GC (four minors at 40-125 ms each at ~100% survival, plus a 307 ms full incremental cycle that frees 12 MB), ~7% realpath syscalls under js_register_path_init → canonicalize_module_path, ~6% stack-map index build, the rest module-body execution.

    Recipe for anyone reproducing the split: PERRY_COLLECT_ONLY=1 writes module-graph.json into the cache dir with a per-module eager/deferred field, without running codegen.

  8. proggeramlug commented on Sep 15, 2026

    @proggeramlug
    ContributorAuthor

    Attribution of the --version gap: module count is not the lever

    Linux, perf stat -r5, opencode --version:

    official Bun Perry (opencode-v4, 6,949 modules)
    user instructions 3.53 B 23.52 B
    wall 0.32 s 2.01 s
    max RSS 196 MB 605 MB

    Bun runs the same modules. A Bun.build with the release options (splitting, bun/node conditions, Solid plugin) plus a per-module evaluation counter shows Bun evaluates 1,645 modules on --version. Static import x from "cjs" lowers to a top-level __toESM(require_x()), so semver, ajv, fastify and js-yaml run at startup under Bun too.

    Perry initializes 2,238 modules. Some of the difference is intended granularity: Perry compiles ai / @ai-sdk/* from src/*.ts (225 modules where Bun loads one dist file). That is ~1.36× the modules for 6.7× the instructions. Function-local and conditional requires are already deferred; #10285 fixes the conditional ones that were still hoisted, a correctness fix with a small startup effect.

    Where the 23.5 B go (DWARF-unwound perf on a PERRY_KEEP_SYMBOLS=1 relink, which is not an object-cache key: 8.5 min for the full graph):

    by code by outermost module init by runtime path
    zod 40% @agentclientprotocol/sdk 21% GC 34% inclusive (incremental full cycle started during startup 14.5%, copying minors 13%)
    effect 33% schema package 13% generic [[Set]] (ordinary_set_with_receiver) 24%
    @modelcontextprotocol/sdk 11% native method dispatch 12%
    core 7%, ai 6% defineProperty/descriptors/assign ~8%
    zstd decode of the 34 MB embedded web UI in the pre-main constructor 2.8%
    path-registry realpath ~1.7%

    Startup is dominated by building zod and Effect schemas at module scope. Perry is ~29× slower than Bun at it: 300 zod z.object schemas take 23.2 B instructions, against ~0.8 B for Bun. The main pathology is filed as #10287: one Object.defineProperty (zod's _zod) sends every later store on that object to the slow path, 5–17× the instructions and ~6× the memory. The GC share follows from the same allocation volume. Env knobs (nursery, pacing, incremental) move it at most ~6%.

    Per-module fixed costs are small for ESM (~3.3 K instructions per trivial module). A CJS wrapper costs ~0.7 M instructions and ~120 KB per module, which is real but ~1–2% here.

    Next levers, in order: #10287 (descriptor-bearing objects keep store fast paths and shared shapes), lazy decode of zstd-embedded assets (implementation in progress), then startup GC (no full cycle before init completes; minor root-scan growth).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew capability or improvementperformanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions