Repository navigation
feat!: bound untrusted-input cost and parse inlines on a delimiter stack - #4
Merged
Merged
Conversation
Adversarial Markdown could make parsing quadratic, cubic, or exponential (unclosed brackets, nested links, unclosed MDX JSX tags and expressions) and could abort the process by overflowing the stack on deep nesting. - Memoize forward closer scans per input (src/memo.rs): position tables, path memos, and bracket memos share step functions with the unmemoized walks, so answers are identical. - Index MDX JSX tags once per text and match closing tags over same-name chains through a persistent per-name map (n log n); block JSX and expression blocks reuse the same walks over joined lines. - Cap nesting at 32 block / 32 inline / 16 emphasis levels; content past a limit stays literal text. Make block-quote paragraph continuation iterative. - Budget the start-dependent `++` / `==` / underline closer search; past the budget further openers in the span stay literal. - Make serializer escaping and line-length tracking linear, and check `html_inline` before copying raw HTML. Output is byte-identical to the previous parser on all non-adversarial inputs checked (fixture corpus plus seeded random inputs, all dialects). README documents the cost contract and limits; decision 006 records the rationale. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Add docs/overview.md and behavior specs under docs/specs/ (public API, block and inline syntax, serialization, validation, HTML rendering, untrusted-input cost), moving the behavior contract out of README.md. - Renumber decisions to 0001-0006 in the spec-flow record format and rewrite 0001 for the published package boundary. - Archive the two completed plans and drop docs/index.md. - Replace the AGENTS.md docs section with the spec-flow conventions. - Add the inline-delimiter-stack plan: CommonMark BOM/NUL handling and moving emphasis-like marks and brackets onto a delimiter stack. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Implement docs/plans/inline-delimiter-stack: - Ignore one leading U+FEFF and read every U+0000 as U+FFFD when recognizing structure and in node values; spans stay in original input coordinates. - Pair `*`, `_`, `~~`, `~`, `^`, `++`, `==`, `||`, and underline `__` on one delimiter stack in closing order. Properly nested marks parse as before; crossing marks now resolve in closing order. - Resolve links, images, inline footnotes, and footnote references with a bracket stack when `]` arrives; a formed link deactivates earlier `[` openers. Remove the link-label verdict memo, link recursion, and the `++` / `==` / underline closer-search budget. - Share the 32-level inline nesting limit across every inline container (emphasis keeps its 16-level cap), so the deepest inline tree fits a 2 MiB stack. - Rewrite decision 0006 for the resulting design. Conformance: 2233/2236 (was 2229). The three remaining failures are comrak's "no literal autolinks inside brackets" cases, which conflict with five micromark (GitHub) cases; GitHub's behavior is kept. BREAKING CHANGE: parse output changes for crossing emphasis-like marks, for input with a leading BOM or embedded NUL, for bracket precedence where the previous forward scan chose an outer bracket, and for nesting past the limits. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…litting - Cut every emphasis-like span from its consumed delimiters' text nodes, so Emphasis and Strong spans cover exactly their own delimiters and content (`***a***` was giving a Strong span outside its Emphasis). - Fall back to a shortcut reference when the `[` after a link label opens no reference label (`[foo][bar` with `[foo]: /u`), as CommonMark specifies. - Accept a hard line break at the end of link text, image alt text, inline footnotes, and text directive labels in validation; a `]` closes after a line break, and the parser already produced that shape. - Replace the three diverging table-row spoiler predictions (row splitting, the block-quote row check, and the serializer's cell check) with one linear scanner that pairs `||` runs as the inline parser pairs them in a cell, reads code spans and escaped backticks as the inline parser does, and never lets an escaped pipe merge two cells. - State the `^[` caret rule as the lexical decision it is. Merge the inline-delimiter-stack spec changes into docs/specs and archive the plan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merged
plimeor
pushed a commit
that referenced
this pull request
Oct 7, 2026
…nto the branch Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011ShSm97dW5me9qB345EiTs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bounds the cost of parsing untrusted Markdown and moves the inline parser onto CommonMark's delimiter-stack algorithm. SemVer-breaking: parse output changes for crossing marks, BOM/NUL input, bracket precedence, and input nested past the limits.
Commits
fix: bound parse and serialize cost on untrusted input: removes the quadratic, cubic, and exponential paths. Raw-content scans (code spans, autolinks, raw HTML, MDX, wikilinks, directive labels) go through per-input memo tables. MDX JSX tags are matched through a tag index. Nesting limits keep stack use bounded.docs: adopt spec-flow layout and plan inline delimiter stack: moves the docs to theoverview/specs/plans/archive/decisionslayout and adds the plan.feat!: parse inline marks and brackets on a delimiter stack:*,_,~~,~,^,++,==,||, underline__) now pair on one delimiter stack. Links, images, and footnotes resolve on a bracket stack when]arrives. The closer-search budget, the link-label verdict memo, and link recursion are gone.fix: correct emphasis spans, reference fallback, and table spoiler splitting:[foo][barfalls back to a shortcut reference.docs/specs/.Behavior changes reviewers should know
**a ==b** c==now gives Strong(a ==b) +c==. Marks that don't cross parse as before."\u{feff}# t"is now a heading.[https://…].docs/specs/block-syntax.md).Design and trade-offs:
docs/decisions/0006-bounded-cost-on-untrusted-input.mddocs/archive/2026-10-05-inline-delimiter-stack/plan.mdValidation
cargo testpasses with 225 tests, and with--features html244 pass.cargo fmt --check,RUSTDOCFLAGS='-D warnings' cargo doc --no-deps, the wasm32 build, andcargo +1.82 buildall pass. Clippy has no new warning kinds.Known gaps (not addressed here)
docs/overview.md.🤖 Generated with Claude Code