Skip to content

ParseHard allocates ~12 KB/op more for every function token added to the grammar, whatever the input #1441

Description

@Rafael-SOWNet

ParseHard (MathS.FromString of a fixed 80-character expression, no cache) allocates about 12 KB/op more for every 'name(' token added to the grammar, whatever the token — the input never mentions the new names.

Measured by ablation on the image-and-preimage branch, steady state (200 warm-up parses, then 2000 measured, GC.GetAllocatedBytesForCurrentThread):

grammar ParseHard B/op
baseline 86af5774 (before powerset(, union(, intersection(, complement(, subset, superset, …, image(, preimage() 3,611,690
with subset/superset/powerset( (#1432) 3,652,370
with union(/intersection(/complement( (#1435) 3,688,778
with the pattern operator's alternatives (#1439) 3,702,858
with image(/preimage( 3,727,130
the same build with the two 'image('/'preimage(' lines removed from the grammar and nothing else changed 3,703,015

So the growth is in the ANTLR runtime's per-parse work, not in any action: the lexer's and parser's static DFA caches should make the number of token types irrelevant in steady state, and they do not. Candidates I have not settled: the lexer's edge cache covering only characters 0..127 per state, so that a decision reached through a wider alphabet is recomputed with an ATNConfigSet each time; a parser decision with a semantic context that is not cached; or the Vocabulary lookups in Parser.cs's implicit-operator pass. Each is a few lines to test with the harness above (scratchpad/pa, a console project referencing the library).

Why it matters: the gate holds allocation within 3% of a baseline, and every function name the docket adds spends 0.3% of that; the pattern operator's PR reached 2.5% and image/preimage 3.2%, so the baseline moved with the explanation (this issue). A parser that pays per token type it does not see is a v3 review point beside the syntax review on #1019; a fix in the runtime's caching would give the allocation back to 2.x.

Activity

  1. Happypig375 commented on Sep 21, 2026

    @Happypig375
    Member

    Why aren't you using issue type and assigning a milestone for this? You should update agentic instructions to do this instead of using labels (#1383/#1384)

  2. added theissue type on Sep 22, 2026
  3. added this to the 2.7.0 milestone on Sep 27, 2026
  4. Happypig375 commented on Sep 28, 2026

    @Happypig375
    Member

    Solved by #1528?

  5. Rafael-SOWNet commented on Sep 28, 2026

    @Rafael-SOWNet
    MemberAuthor

    Yes. The 12 KB per token was ANTLR's full-context LL prediction, which #1528 stopped paying on every parse. The growth is gone, not just the level: ParseHard was 59,871 B/op in #1528's gate run, and it is 59,871 B/op again in today's gate run for #1537. That build has the four tokens #1529 added since (Si, Ci, Shi and Chi), which under LL would have added 4 × 12,136 = 48,544 B/op.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions