ParseHard (MathS.FromString of a fixed 80-character expression, no cache) allocates about 12 KB/op more for every 'name(' token added to the grammar, whatever the token — the input never mentions the new names.
Measured by ablation on the image-and-preimage branch, steady state (200 warm-up parses, then 2000 measured, GC.GetAllocatedBytesForCurrentThread):
| grammar |
ParseHard B/op |
baseline 86af5774 (before powerset(, union(, intersection(, complement(, subset, superset, …, image(, preimage() |
3,611,690 |
with subset/superset/powerset( (#1432) |
3,652,370 |
with union(/intersection(/complement( (#1435) |
3,688,778 |
| with the pattern operator's alternatives (#1439) |
3,702,858 |
with image(/preimage( |
3,727,130 |
the same build with the two 'image('/'preimage(' lines removed from the grammar and nothing else changed |
3,703,015 |
So the growth is in the ANTLR runtime's per-parse work, not in any action: the lexer's and parser's static DFA caches should make the number of token types irrelevant in steady state, and they do not. Candidates I have not settled: the lexer's edge cache covering only characters 0..127 per state, so that a decision reached through a wider alphabet is recomputed with an ATNConfigSet each time; a parser decision with a semantic context that is not cached; or the Vocabulary lookups in Parser.cs's implicit-operator pass. Each is a few lines to test with the harness above (scratchpad/pa, a console project referencing the library).
Why it matters: the gate holds allocation within 3% of a baseline, and every function name the docket adds spends 0.3% of that; the pattern operator's PR reached 2.5% and image/preimage 3.2%, so the baseline moved with the explanation (this issue). A parser that pays per token type it does not see is a v3 review point beside the syntax review on #1019; a fix in the runtime's caching would give the allocation back to 2.x.
ParseHard(MathS.FromStringof a fixed 80-character expression, no cache) allocates about 12 KB/op more for every'name('token added to the grammar, whatever the token — the input never mentions the new names.Measured by ablation on the
image-and-preimagebranch, steady state (200 warm-up parses, then 2000 measured,GC.GetAllocatedBytesForCurrentThread):86af5774(beforepowerset(,union(,intersection(,complement(,subset,superset,…,image(,preimage()subset/superset/powerset((#1432)union(/intersection(/complement((#1435)image(/preimage('image('/'preimage('lines removed from the grammar and nothing else changedSo the growth is in the ANTLR runtime's per-parse work, not in any action: the lexer's and parser's static DFA caches should make the number of token types irrelevant in steady state, and they do not. Candidates I have not settled: the lexer's edge cache covering only characters
0..127per state, so that a decision reached through a wider alphabet is recomputed with anATNConfigSeteach time; a parser decision with a semantic context that is not cached; or theVocabularylookups inParser.cs's implicit-operator pass. Each is a few lines to test with the harness above (scratchpad/pa, a console project referencing the library).Why it matters: the gate holds allocation within 3% of a baseline, and every function name the docket adds spends 0.3% of that; the pattern operator's PR reached 2.5% and
image/preimage3.2%, so the baseline moved with the explanation (this issue). A parser that pays per token type it does not see is a v3 review point beside the syntax review on #1019; a fix in the runtime's caching would give the allocation back to 2.x.