From 5fcf5b03a7b6d5d37d32bda56359708b02763221 Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 5 Oct 2026 11:47:54 +0000 Subject: [PATCH 01/40] docs: propose source span mapping and CommonMark fixes Plan to map every parsed node's span back to its source bytes across container prefixes, CRLF, table cells, and split tabs, and to fix the rule of three, image reference fallback, and serializer escaping of backticks and text after shortcut references. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3 --- .../plan.md | 188 ++++++++++++++++++ .../specs/inline-syntax.md | 34 ++++ .../specs/public-api.md | 98 +++++++++ .../specs/serialization.md | 36 ++++ 4 files changed, 356 insertions(+) create mode 100644 docs/plans/source-span-mapping-and-commonmark-fixes/plan.md create mode 100644 docs/plans/source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md create mode 100644 docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md create mode 100644 docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md b/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md new file mode 100644 index 0000000..f3c1956 --- /dev/null +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md @@ -0,0 +1,188 @@ +# Source span mapping and CommonMark fixes + +## Why + +Spans are absolute only up to the first stripped byte. Container content (block +quotes, list items, and the other containers) is joined into one string and +re-read from a single base offset, and paragraph, heading, and table-cell inline +input is joined the same way. So every line after a stripped prefix, a CRLF, or +a cell pipe is shifted: block and inline spans slice the wrong text, and code +blocks inside containers get spans past the end of the input. Separately, three +CommonMark and round-trip defects predate the delimiter-stack change: the rule +of three uses the remaining run length, an image with an invalid `(…)` does not +fall back to a reference, and the serializer writes text that reparses as a code +span, an inline link, or a link reference definition. + +## What changes + +- Every parsed node's span maps to the source bytes it was read from, on every + line of every container, across CRLF line endings, inside table cells, and + across split tabs. +- Table cells carry spans; today they carry none. +- Every span lies within its parent's span, in source order among its siblings, + and this holds for the whole tree. +- The leading-whitespace qualifier on "Emphasis-like spans cover their + delimiters" is removed. +- The rule of three and its opener floor use each delimiter run's original + length, as CommonMark specifies. +- An image whose `(…)` is not a valid resource falls back to the full, + collapsed, and shortcut reference forms, as a link does. +- The serializer writes every backtick in text as `` \` ``. This changes + canonical output for text that holds a backtick it left bare before. +- The serializer escapes text after a shortcut link or image reference so that + reparsing keeps the reference: a `(` directly after it, and a `:` directly + after one that starts a paragraph. + +Specs: +- `public-api` (modified) +- `inline-syntax` (modified) +- `serialization` (modified) + +## Out of scope + +- The final line of a paragraph keeps its trailing whitespace (`a \n` gives + `Text("a ")`), which CommonMark strips. This is block-level whitespace + handling that touches hard breaks and spans, so it goes to its own task. +- Renaming `src/parse/nul.rs`, a reserved Windows filename that `cargo package` + warns about. It is a packaging fix with no behavior change and lands on its own + before the crate is published. + +## Design + +### Decisions + +- One source-map type owns the translation from a derived string to the + original input. Each derived string (container content, the inline input of a + paragraph, heading, table cell, directive label, description term, or HTML + container) is built together with its map. A `Line` takes its positions + through the map when it is created, and an inline parse translates its spans + once, after the pass. This is because the span construction sites in the + parser (about 96 `base_offset +` expressions) then stay as they are. + - Turned down: translating at each construction site, for the reason the BOM + and NUL change turned down a second coordinate system. + - Turned down: a dense per-byte offset table, which costs a machine word per + content byte at each of up to 32 container levels. + - Turned down: mapping only leading whitespace, the scope first proposed. It + leaves container continuation lines, CRLF, table cells, and out-of-bounds + code-block spans wrong. +- A map is always in original-input coordinates. A container composes its + parent's map as it copies the parent's lines. This is because a lookup then + costs the same at any depth. + - Turned down: chaining a child map to its parent's, which makes each lookup + walk the nesting. +- A map is a sequence of segments, each pairing a content range with the source + range it came from. + - Verbatim segments map byte for byte. + - A segment whose content replaces source maps every content byte to the + whole source range it replaces. This covers the spaces split from a tab, the + `\n` that joins lines ending in `\r\n`, and the `|` read from `\|`. + - A start position at a segment boundary maps through the segment that begins + there. An end position maps through the segment that ends there. + - This is because these rules give each scenario its span directly: a soft + break covers its whole line ending, and a node ending before a line break + ends on its own line. +- Inline spans are computed in inline-input coordinates and translated by one + walk over the finished subtree. + - In preorder, starts never decrease, so a forward cursor maps them. + - Each end is found by searching forward from its start's segment. + - The walk is linear because inline nesting is capped at 32 levels. + - Turned down: binary search per position, which is `n log n` against the + linear-time requirement. +- A table cell's span covers its trimmed content, the text its inline parse + reads. + - An empty cell gets the empty range just before the pipe that closes it. + - A cell missing from a short row gets the empty range at the row's end. + - This is because every parsed node must carry a span, and children then start + where their cell starts. + - Turned down: leaving cells without spans, which breaks "Source spans". + - Turned down: including padding and pipes, which makes adjacent cells share + a delimiter byte. +- "Spans nest" is checked over the whole tree for the fixture corpus and + generated inputs, because a missed construction site at any depth then fails a + test instead of waiting for a scenario. +- Each delimiter run keeps its original length beside its remaining length. + The rule of three and the `openers_bottom` key both use the original length, + because cmark and commonmark.js do both. The opener floor is sound only when + keyed on what the predicate reads. + - Turned down: changing only the predicate, which leaves the floor keyed on a + length the predicate no longer reads. + - During planning, this change reduced mismatches against commonmark.js on + 20,476 generated `*`/text inputs from 304 to 0, with `cargo test` and + conformance unchanged. +- The image-only early return in `match_link_target` is deleted, so `![` and `[` + resolve what follows `]` with one sequence. The experiment during planning + left tests and conformance unchanged. + - Turned down: an image-specific fallback path. +- Every backtick in text is escaped, and the code-span prediction + `text_code_span_can_start` is removed. A backslash stops a backtick from + opening a code span but not from closing one, so a prediction must reason + about escaped backticks as closers. Dropping the prediction removes the + serializer's copy of the code-span rules. During planning, 5 fixture inputs + changed canonical output. + - Turned down: predicting per run, which keeps output byte-stable but keeps a + mirrored copy of parser rules. +- After a shortcut reference, the serializer escapes the character that would + re-read its brackets, because the AST must round-trip. + - Turned down: writing the reference in collapsed form, which changes the + reparsed `ReferenceKind`. + +### Risks + +- [A derived string built without its map leaves a shifted span] → The "Spans + nest" check runs over the corpus and over generated inputs that mix block + quotes, list items, tables, CRLF, and tabs. Each container kind has a + scenario or test. +- [Spaces split from a tab are assumed to occur only at a line's front, so a + nested container that strips more could meet them elsewhere] → The generated + inputs mix tabs with nested container markers. A counterexample changes the + segment representation, not the requirements. +- [Map lookups break linear time] → A growth check over long, nested block + quotes and list items and over long tables joins `tests/pathological_inputs.rs` + and the growth sweep. +- [The rule-of-three change moves emphasis output beyond the targeted inputs] → + Parse output is compared with the plan's starting commit. Every diff must + involve a `*`, `_`, or underline `__` run that pairs more than once. +- [Downstream users rely on bare backticks in canonical output] → The commit + states the output change, and the "Literal backticks are always escaped" + scenarios pin it. + +## Tasks + +### 1. Rule of three +- [ ] 1.1 Add tests for "Rule of three counts whole delimiter runs" and for `***a*a*a`; verified by both failing on the current code. +- [ ] 1.2 Keep each delimiter run's original length, and use it in `emphasis_delimiters_match` and `openers_bottom_key`; verified by: + - the 1.1 tests and `cargo test` passing + - conformance not below 2233/2236 + - a parse-output comparison with the starting commit, over the fixture corpus and seeded generated inputs, differing only in the class listed under Risks + +### 2. Image fallback and text after shortcut references +- [ ] 2.1 Add tests for "Image whose resource is invalid" and the three shortcut-reference scenarios of "Escaping keeps text literal"; verified by each failing on the current code. +- [ ] 2.2 Delete the image-only early return in `match_link_target` and correct its doc comment; verified by the image test passing and conformance unchanged. +- [ ] 2.3 Escape a `(` directly after a shortcut `LinkReference` or `ImageReference`, and a `:` directly after one that starts a paragraph; verified by the shortcut-reference tests and the round-trip fixtures passing. + +### 3. Backticks +- [ ] 3.1 Add tests for both scenarios of "Literal backticks are always escaped"; verified by the first failing on the current code. +- [ ] 3.2 Escape every backtick in text, and remove `text_code_span_can_start` and any scan state only it used; verified by: + - the 3.1 tests and `cargo test` passing + - every moved `.canonical.md` golden regenerated and read for correctness + - every canonical diff from the starting commit being a backtick in text + +### 4. Source spans +- [ ] 4.1 Add tests: + - every scenario of "Spans map stripped lines back to the source", the new "Source spans" scenarios, and "Emphasis on a block quote continuation line" + - a "Spans nest" check over the fixture corpus and over seeded generated inputs mixing containers, tables, CRLF, and tabs + + Verified by the changed-behavior tests failing on the current code. +- [ ] 4.2 Add the source-map type and build every container's content with its map, composed in original coordinates. The containers are block quotes and alerts, list items, footnote definitions, container directives, description details, and HTML containers. `Line` positions come from the map. Verified by: + - the "Later block inside a block quote" and "Split tab" tests passing + - `tests/parse_span_contract.rs` passing + - the nest check finding no block span outside its parent +- [ ] 4.3 Build every inline input with its map: paragraphs, ATX and setext headings, table cells, directive labels, description terms, and HTML containers. Translate spans in one walk per block-level inline parse. Verified by the inline scenarios and the nest check passing at every depth. +- [ ] 4.4 Give table cells their spans, as described under Decisions; verified by the table-cell scenarios passing. +- [ ] 4.5 Extend `inline_container_spans_cover_their_delimiters_and_content` to inputs with leading whitespace, block quotes, list items, tables, and CRLF; verified by it passing. +- [ ] 4.6 Add a growth check for long, nested block quotes and list items and for long tables to `tests/pathological_inputs.rs`; verified by linear growth in debug and release builds. + +### 5. Integration checks +- [ ] 5.1 `cargo fmt --check`, `cargo test` with and without `html`, `RUSTDOCFLAGS='-D warnings' cargo doc --no-deps`, `cargo build --target wasm32-unknown-unknown`, and `cargo +1.82 build` all pass. +- [ ] 5.2 `tests/pathological_inputs.rs`, the growth sweep, and the 2 MiB stack check pass in debug and release builds. +- [ ] 5.3 The conformance bench result is measured and reported with the change, along with a benign-document benchmark against the starting commit. diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md new file mode 100644 index 0000000..7b91cdb --- /dev/null +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md @@ -0,0 +1,34 @@ +# Inline syntax — spec changes + +## MODIFIED Requirements + +### Requirement: CommonMark inlines +The parser SHALL recognize CommonMark inline constructs (backslash escapes, +entity and numeric character references, code spans, emphasis and strong +emphasis, links, images, autolinks, raw HTML, and hard and soft line breaks) as +the CommonMark specification defines them, including its precedence of code +spans, links, and emphasis. + +#### Scenario: Emphasis +- **WHEN** `"Hello *world*."` is parsed +- **THEN** the paragraph holds `Text("Hello ")`, an `Emphasis` containing `world`, and `Text(".")` + +#### Scenario: Link inside a link label +- **WHEN** `"[foo [bar](/u)](/v)"` is parsed +- **THEN** only `[bar](/u)` becomes a link and the surrounding brackets and `(/v)` stay text + +#### Scenario: Shortcut reference before an unclosed label +- **WHEN** `"[foo][bar\n\n[foo]: /u"` is parsed with the CommonMark preset +- **THEN** the paragraph holds a shortcut `LinkReference` to `foo` followed by `Text("[bar")` + +#### Scenario: Rule of three counts whole delimiter runs +- **WHEN** `"*a***b*"` is parsed with the CommonMark preset +- **THEN** the paragraph holds an `Emphasis` containing `a`, `Text("*")`, and an `Emphasis` containing `b` + +#### Scenario: Image whose resource is invalid +- **WHEN** `"![foo](a b)\n\n[foo]: /u"` is parsed with the CommonMark preset +- **THEN** the paragraph holds a shortcut `ImageReference` to `foo` followed by `Text("(a b)")` + +#### Scenario: CommonMark oracle cases +- **WHEN** the inline cases under `tests/fixtures/conformance/commonmark/` are parsed and rendered with the `html` feature +- **THEN** the output matches the expected HTML diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md new file mode 100644 index 0000000..39afd13 --- /dev/null +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md @@ -0,0 +1,98 @@ +# Public API — spec changes + +## ADDED Requirements + +### Requirement: Spans map stripped lines back to the source +A parsed node's span SHALL start at the source byte where its first character +was read and end after the source byte where its last character was read, +wherever the parser removes indentation, container markers, or table-cell +padding before reading a line, joins lines whose source line ending is `\r\n`, +or reads `\|` in a table cell as `|`; a space the parser produces by splitting a +tab SHALL map to that tab. + +#### Scenario: Leading whitespace on a paragraph line +- **WHEN** `parse(" a *b*")` runs +- **THEN** the paragraph holds `Text("a ")` spanning bytes 2..4 and an `Emphasis` spanning bytes 4..7 + +#### Scenario: Block quote continuation line +- **WHEN** `parse("> a\n> b *c*")` runs +- **THEN** the block quote's paragraph spans bytes 2..11 and holds `Text("b ")` spanning bytes 6..8 and an `Emphasis` spanning bytes 8..11 + +#### Scenario: List item continuation line +- **WHEN** `parse("- a\n b *c*")` runs +- **THEN** the item's paragraph spans bytes 2..11 and holds an `Emphasis` spanning bytes 8..11 + +#### Scenario: Later block inside a block quote +- **WHEN** `parse("> a\n>\n> b")` runs +- **THEN** the block quote's second paragraph spans bytes 8..9 + +#### Scenario: CRLF soft break +- **WHEN** `parse("a\r\nb")` runs +- **THEN** the paragraph holds a `SoftBreak` spanning bytes 1..3 and `Text("b")` spanning bytes 3..4 + +#### Scenario: Table cell content +- **WHEN** `parse("| a *b* |\n|-|")` runs +- **THEN** the header cell spans bytes 2..7 and holds `Text("a ")` spanning bytes 2..4 and an `Emphasis` spanning bytes 4..7 + +#### Scenario: Escaped pipe in a table cell +- **WHEN** `parse("| a\\|b |\n|-|")` runs +- **THEN** the header cell holds `Text("a|b")` spanning bytes 2..6 + +#### Scenario: Split tab +- **WHEN** `parse(">\t\tfoo")` runs +- **THEN** the block quote holds an indented code block with value ` foo` spanning bytes 1..6 + +### Requirement: Spans nest +Every parsed node's span SHALL lie on UTF-8 character boundaries within the +input and within the span of the node that contains it, and the spans of a +node's children SHALL be in source order and SHALL NOT overlap. + +#### Scenario: Fixture corpus and generated inputs +- **WHEN** every fixture input and every seeded generated input is parsed in each dialect +- **THEN** every node, at every depth, satisfies these conditions + +## MODIFIED Requirements + +### Requirement: Source spans +Every parsed node SHALL carry an absolute, half-open UTF-8 byte range into the +original input; hand-built nodes SHALL carry `None`; `LineIndex` SHALL convert a +span to 1-based line and column positions. + +#### Scenario: First block position +- **WHEN** `parse("# Title\n\nHello.")` runs and the first block's span is passed to `LineIndex::new(source).span(span)` +- **THEN** the start position is line 1, column 1 + +#### Scenario: Hand-built node +- **WHEN** a `Heading` is built with `Heading::new(1, [Text::from("Title")])` +- **THEN** its `span()` is `None` + +#### Scenario: Empty table cell +- **WHEN** `parse("| a | |\n|-|-|")` runs +- **THEN** the second header cell's span is the empty range at byte 6, just before the pipe that closes it + +#### Scenario: Missing table cell +- **WHEN** `parse("| a | b |\n|-|-|\n| c")` runs +- **THEN** the body row's second cell has no children and its span is the empty range at the end of the row, byte 19 + +### Requirement: Emphasis-like spans cover their delimiters +The span of a parsed emphasis-like container (`Emphasis`, `Strong`, +`Underline`, `Delete`, `Insert`, `Mark`, `Spoiler`, `Subscript`, or +`Superscript`) SHALL run from the first character of the delimiters that open +it to the last character of the delimiters that close it, and SHALL lie within +the span of the node that contains it. + +#### Scenario: Strong inside emphasis +- **WHEN** `parse("***a***")` runs +- **THEN** the paragraph holds an `Emphasis` spanning bytes 0..7 that holds a `Strong` spanning bytes 1..6 + +#### Scenario: Emphasis inside strong +- **WHEN** `parse("x ***a* b**")` runs +- **THEN** the paragraph holds a `Strong` spanning bytes 2..11 that holds an `Emphasis` spanning bytes 4..7 + +#### Scenario: Leftover opening delimiter +- **WHEN** `parse("**a*")` runs +- **THEN** the paragraph holds `Text("*")` spanning bytes 0..1 and an `Emphasis` spanning bytes 1..4 + +#### Scenario: Emphasis on a block quote continuation line +- **WHEN** `parse("> a\n> *b*")` runs +- **THEN** the block quote's paragraph holds an `Emphasis` spanning bytes 6..9 diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md new file mode 100644 index 0000000..0273863 --- /dev/null +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md @@ -0,0 +1,36 @@ +# Serialization — spec changes + +## ADDED Requirements + +### Requirement: Literal backticks are always escaped +The serializer SHALL write every backtick in a text value as `` \` ``. + +#### Scenario: Backtick run before a lone backtick +- **WHEN** a hand-built paragraph holding ```Text("b ``a`")``` is serialized +- **THEN** `to_markdown()` returns ``"b \`\`a\`\n"`` and reparsing it yields the same text and no code span + +#### Scenario: Paired backticks +- **WHEN** ``parse("Test \\`hello world` here.").document.to_markdown()`` runs +- **THEN** it returns ``"Test \`hello world\` here.\n"`` + +## MODIFIED Requirements + +### Requirement: Escaping keeps text literal +The serializer SHALL escape text so that reparsing the output yields the same +text, and leaves the nodes beside it as they were, rather than new constructs. + +#### Scenario: Literal asterisks in text +- **WHEN** a hand-built paragraph holding `Text("*not emphasis*")` is serialized and reparsed +- **THEN** the reparsed paragraph holds the same text and no `Emphasis` + +#### Scenario: Parenthesis after a shortcut reference +- **WHEN** the document parsed from `"[foo]\\(a)\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed paragraph holds a shortcut `LinkReference` to `foo` followed by `Text("(a)")` + +#### Scenario: Parenthesis after a shortcut image reference +- **WHEN** the document parsed from `"![foo]\\(a)\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed paragraph holds a shortcut `ImageReference` to `foo` followed by `Text("(a)")` + +#### Scenario: Colon after a shortcut reference that starts a paragraph +- **WHEN** the document parsed from `"[foo]\\: /x\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed document still holds the paragraph, with a shortcut `LinkReference` to `foo` followed by `Text(": /x")` From 264c1b9839f15d491e6b5fbd69a1c251a6c893cd Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 5 Oct 2026 11:57:16 +0000 Subject: [PATCH 02/40] docs: fold issue #6 container span cases into the span plan Add the issue's 31-case regression test to the plan's tasks and representative scenarios for nested lists, nested block quotes, alerts, footnote continuations, HTML containers, container directives, tabs, and CRLF to the public-api spec change. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3 --- .../plan.md | 14 +++++--- .../specs/public-api.md | 36 +++++++++++++++++++ 2 files changed, 46 insertions(+), 4 deletions(-) diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md b/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md index f3c1956..61c7893 100644 --- a/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md @@ -7,7 +7,11 @@ quotes, list items, and the other containers) is joined into one string and re-read from a single base offset, and paragraph, heading, and table-cell inline input is joined the same way. So every line after a stripped prefix, a CRLF, or a cell pipe is shifted: block and inline spans slice the wrong text, and code -blocks inside containers get spans past the end of the input. Separately, three +blocks inside containers get spans past the end of the input. Issue +plimeor/markdown-syntax#6 measures this on 31 container cases, 26 of which fail +on 0.3.0. Editors that rewrite source in place by span then panic when a span +lands inside a multi-byte character, or silently replace the wrong bytes. +Separately, three CommonMark and round-trip defects predate the delimiter-stack change: the rule of three uses the remaining run length, an image with an invalid `(…)` does not fall back to a reference, and the serializer writes text that reparses as a code @@ -17,7 +21,8 @@ span, an inline link, or a link reference definition. - Every parsed node's span maps to the source bytes it was read from, on every line of every container, across CRLF line endings, inside table cells, and - across split tabs. + across split tabs. This resolves issue plimeor/markdown-syntax#6, whose + 31-case regression test joins the suite. - Table cells carry spans; today they carry none. - Every span lies within its parent's span, in source order among its siblings, and this holds for the whole tree. @@ -170,14 +175,15 @@ Specs: ### 4. Source spans - [ ] 4.1 Add tests: - every scenario of "Spans map stripped lines back to the source", the new "Source spans" scenarios, and "Emphasis on a block quote continuation line" + - issue plimeor/markdown-syntax#6's regression test, as `inline_spans_address_source_inside_containers` in `tests/parse_span_contract.rs`, with its 31 cases unchanged - a "Spans nest" check over the fixture corpus and over seeded generated inputs mixing containers, tables, CRLF, and tabs - Verified by the changed-behavior tests failing on the current code. + Verified by the changed-behavior tests failing on the current code, the issue's test failing 26 of 31 cases. - [ ] 4.2 Add the source-map type and build every container's content with its map, composed in original coordinates. The containers are block quotes and alerts, list items, footnote definitions, container directives, description details, and HTML containers. `Line` positions come from the map. Verified by: - the "Later block inside a block quote" and "Split tab" tests passing - `tests/parse_span_contract.rs` passing - the nest check finding no block span outside its parent -- [ ] 4.3 Build every inline input with its map: paragraphs, ATX and setext headings, table cells, directive labels, description terms, and HTML containers. Translate spans in one walk per block-level inline parse. Verified by the inline scenarios and the nest check passing at every depth. +- [ ] 4.3 Build every inline input with its map: paragraphs, ATX and setext headings, table cells, directive labels, description terms, and HTML containers. Translate spans in one walk per block-level inline parse. Verified by the inline scenarios, all 31 cases of `inline_spans_address_source_inside_containers`, and the nest check passing at every depth. - [ ] 4.4 Give table cells their spans, as described under Decisions; verified by the table-cell scenarios passing. - [ ] 4.5 Extend `inline_container_spans_cover_their_delimiters_and_content` to inputs with leading whitespace, block quotes, list items, tables, and CRLF; verified by it passing. - [ ] 4.6 Add a growth check for long, nested block quotes and list items and for long tables to `tests/pathological_inputs.rs`; verified by linear growth in debug and release builds. diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md index 39afd13..39b9ee4 100644 --- a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md @@ -26,6 +26,38 @@ tab SHALL map to that tab. - **WHEN** `parse("> a\n>\n> b")` runs - **THEN** the block quote's second paragraph spans bytes 8..9 +#### Scenario: Nested list item after multi-byte text +- **WHEN** `parse("- 项目\n - 嵌套 [[library/工作/买菜]]\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 20..45 + +#### Scenario: Nested block quote +- **WHEN** `parse("> 外层\n> > 内层 [[A]]\n")` runs +- **THEN** the inner block quote's `WikiLink` spans bytes 20..25 + +#### Scenario: Alert body +- **WHEN** `parse("> [!NOTE]\n> 见 [[A]]\n")` runs +- **THEN** the alert's `WikiLink` spans bytes 16..21 + +#### Scenario: Footnote definition continuation line +- **WHEN** `parse("正文[^1]\n\n[^1]: 见 [[A]]\n 续 [[B]]\n")` runs +- **THEN** the footnote definition's second `WikiLink` spans bytes 36..41 + +#### Scenario: Inside an HTML container +- **WHEN** `parse("
\n更多\n\n- 项目\n - [[A]]\n\n
\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 50..55 + +#### Scenario: Inside a container directive +- **WHEN** `parse(":::note\n- 项目\n - [[A]]\n:::\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 21..26 + +#### Scenario: Tab-indented nested list +- **WHEN** `parse("- 项目\n\t- 嵌套 [[A]]\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 19..24 + +#### Scenario: CRLF nested list +- **WHEN** `parse("- 项目\r\n - 嵌套 [[A]]\r\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 21..26 + #### Scenario: CRLF soft break - **WHEN** `parse("a\r\nb")` runs - **THEN** the paragraph holds a `SoftBreak` spanning bytes 1..3 and `Text("b")` spanning bytes 3..4 @@ -42,6 +74,10 @@ tab SHALL map to that tab. - **WHEN** `parse(">\t\tfoo")` runs - **THEN** the block quote holds an indented code block with value ` foo` spanning bytes 1..6 +#### Scenario: Container span regression cases +- **WHEN** `inline_spans_address_source_inside_containers` in `tests/parse_span_contract.rs` parses each of its 31 cases (lists, task lists, block quotes, alerts, tables, footnote definitions, HTML containers, container directives, frontmatter, CRLF, and tabs) +- **THEN** every `WikiLink`, `Link`, `Image`, and `#`-holding `Text` it collects spans exactly the literal it occupies in the input + ### Requirement: Spans nest Every parsed node's span SHALL lie on UTF-8 character boundaries within the input and within the span of the node that contains it, and the spans of a From 0603ec915b9ec30dbde12769fd503e43b7992bcb Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 5 Oct 2026 11:59:02 +0000 Subject: [PATCH 03/40] fix: rename the NUL replacement module off a reserved Windows filename `src/parse/nul.rs` cannot be created on Windows, so the packaged crate would fail to unpack there; `cargo package` warns about it. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3 --- src/parse.rs | 6 +++--- src/parse/{nul.rs => nul_replacement.rs} | 0 2 files changed, 3 insertions(+), 3 deletions(-) rename src/parse/{nul.rs => nul_replacement.rs} (100%) diff --git a/src/parse.rs b/src/parse.rs index c2cadab..fa087db 100644 --- a/src/parse.rs +++ b/src/parse.rs @@ -18,7 +18,7 @@ use crate::{ validate::is_directive_name, }; -mod nul; +mod nul_replacement; #[cfg(test)] mod scan_tests; @@ -167,7 +167,7 @@ fn parse_checked(input: &str, options: &SyntaxOptions) -> Result Result char { if char == '\0' { '\u{FFFD}' diff --git a/src/parse/nul.rs b/src/parse/nul_replacement.rs similarity index 100% rename from src/parse/nul.rs rename to src/parse/nul_replacement.rs From a73cee53ad9aed4f88c20d428b99fdbc88eb584e Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 5 Oct 2026 12:29:53 +0000 Subject: [PATCH 04/40] fix: CommonMark rule of three, image reference fallback, and literal text escaping MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - The rule of three reads each delimiter run's original length, as CommonMark specifies, instead of what earlier pairings left. - An image whose `(…)` is not a valid resource falls back to the full, collapsed, and shortcut reference forms, as a link does. - The serializer escapes every backtick in text: a backslash keeps a backtick from opening a code span but not from closing one. Canonical output changes for text with a backtick it left bare. - Text after a shortcut reference escapes a leading `(`, and a leading `:` when the reference opens the line, so reparsing keeps the reference. - Inside `_`-delimited emphasis, a `_` in text that can close is escaped. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3 --- .../plan.md | 25 +++-- .../specs/serialization.md | 4 + src/parse.rs | 43 +++++--- src/serialize.rs | 94 ++++++++++++---- src/serialize/escape_scan_tests.rs | 13 --- tests/delimiter_stack.rs | 29 +++++ .../commonmark_character_escapes.canonical.md | 2 +- .../spec/commonmark_html_blocks.canonical.md | 2 +- .../spec/commonmark_html_inlines.canonical.md | 2 +- tests/serialize_regressions.rs | 104 +++++++++++++++++- 10 files changed, 255 insertions(+), 63 deletions(-) diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md b/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md index 61c7893..660d56b 100644 --- a/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md @@ -108,7 +108,9 @@ Specs: - Each delimiter run keeps its original length beside its remaining length. The rule of three and the `openers_bottom` key both use the original length, because cmark and commonmark.js do both. The opener floor is sound only when - keyed on what the predicate reads. + keyed on what the predicate reads, so the key uses the original length for `*` + and `_`, whose predicate is the rule of three, and the remaining length for + the other marks, whose predicates read it. - Turned down: changing only the predicate, which leaves the floor keyed on a length the predicate no longer reads. - During planning, this change reduced mismatches against commonmark.js on @@ -126,6 +128,12 @@ Specs: changed canonical output. - Turned down: predicting per run, which keeps output byte-stable but keeps a mirrored copy of parser rules. +- Inside an `_`-delimited emphasis, the serializer escapes a `_` in text that + can close, as text inside `*`-delimited emphasis already encodes its `*`, + because the reparse would close the emphasis there. Turned down: escaping + every `*` / `_` that can close, or every character of an escaped run, which + moved canonical output for about 1,000 corpus inputs while the defect needs a + delimiter outside the text node. - After a shortcut reference, the serializer escapes the character that would re-read its brackets, because the AST must round-trip. - Turned down: writing the reference in collapsed form, which changes the @@ -154,20 +162,21 @@ Specs: ## Tasks ### 1. Rule of three -- [ ] 1.1 Add tests for "Rule of three counts whole delimiter runs" and for `***a*a*a`; verified by both failing on the current code. -- [ ] 1.2 Keep each delimiter run's original length, and use it in `emphasis_delimiters_match` and `openers_bottom_key`; verified by: +- [x] 1.1 Add tests for "Rule of three counts whole delimiter runs" and for `***a*a*a`; verified by both failing on the current code. +- [x] 1.2 Keep each delimiter run's original length, and use it in `emphasis_delimiters_match` and in `openers_bottom_key` for `*` / `_`; verified by: - the 1.1 tests and `cargo test` passing - conformance not below 2233/2236 - a parse-output comparison with the starting commit, over the fixture corpus and seeded generated inputs, differing only in the class listed under Risks +- [x] 1.3 Escape a `_` in text that can close when the text sits inside an `_`-delimited emphasis, because the new parse produces that shape for inputs that round-tripped before; verified by `an_underscore_that_can_close_stays_inside_underscore_emphasis` (failing before) and the parse-output comparison showing no corpus input whose canonical output moves ### 2. Image fallback and text after shortcut references -- [ ] 2.1 Add tests for "Image whose resource is invalid" and the three shortcut-reference scenarios of "Escaping keeps text literal"; verified by each failing on the current code. -- [ ] 2.2 Delete the image-only early return in `match_link_target` and correct its doc comment; verified by the image test passing and conformance unchanged. -- [ ] 2.3 Escape a `(` directly after a shortcut `LinkReference` or `ImageReference`, and a `:` directly after one that starts a paragraph; verified by the shortcut-reference tests and the round-trip fixtures passing. +- [x] 2.1 Add tests for "Image whose resource is invalid" and the three shortcut-reference scenarios of "Escaping keeps text literal"; verified by each failing on the current code. +- [x] 2.2 Delete the image-only early return in `match_link_target` and correct its doc comment; verified by the image test passing and conformance unchanged. +- [x] 2.3 Escape a `(` directly after a shortcut `LinkReference` or `ImageReference`, and a `:` directly after one that starts a paragraph; verified by the shortcut-reference tests and the round-trip fixtures passing. ### 3. Backticks -- [ ] 3.1 Add tests for both scenarios of "Literal backticks are always escaped"; verified by the first failing on the current code. -- [ ] 3.2 Escape every backtick in text, and remove `text_code_span_can_start` and any scan state only it used; verified by: +- [x] 3.1 Add tests for both scenarios of "Literal backticks are always escaped"; verified by the first failing on the current code. +- [x] 3.2 Escape every backtick in text, and remove `text_code_span_can_start` and any scan state only it used; `ordinary_punctuation_text_does_not_reparse_as_character_escapes` drops the backtick from its sample of punctuation left unescaped; verified by: - the 3.1 tests and `cargo test` passing - every moved `.canonical.md` golden regenerated and read for correctness - every canonical diff from the starting commit being a backtick in text diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md index 0273863..5b885a5 100644 --- a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md @@ -23,6 +23,10 @@ text, and leaves the nodes beside it as they were, rather than new constructs. - **WHEN** a hand-built paragraph holding `Text("*not emphasis*")` is serialized and reparsed - **THEN** the reparsed paragraph holds the same text and no `Emphasis` +#### Scenario: Underscore that can close inside underscore emphasis +- **WHEN** a hand-built paragraph holding an `Emphasis` around a `Strong` around `Text("(a b)_.")`, followed by `Text("*#")`, is serialized and reparsed with the CommonMark preset +- **THEN** `to_markdown()` returns `"_**(a b)\\_.**_\\*#\n"` and the reparsed paragraph equals the original apart from spans + #### Scenario: Parenthesis after a shortcut reference - **WHEN** the document parsed from `"[foo]\\(a)\n\n[foo]: /u"` is serialized and reparsed - **THEN** the reparsed paragraph holds a shortcut `LinkReference` to `foo` followed by `Text("(a)")` diff --git a/src/parse.rs b/src/parse.rs index fa087db..29fdd94 100644 --- a/src/parse.rs +++ b/src/parse.rs @@ -3691,6 +3691,10 @@ struct DelimMarker { marker: u8, /// Remaining unmatched delimiter characters in this run. length: usize, + /// The run's length as scanned. CommonMark's rule of three, and the + /// `openers_bottom` key that caches its verdicts, read this rather than + /// what is left after earlier pairings. + run_length: usize, can_open: bool, can_close: bool, /// The `~` subscript / `^` superscript roles of a run. @@ -3972,7 +3976,6 @@ fn close_bracket( input, opener.position + 2, close, - true, definitions, ); if let Some((end, target)) = target { @@ -4002,7 +4005,6 @@ fn close_bracket( input, opener.position + 1, close, - false, definitions, ); if let Some((end, target)) = target { @@ -4086,6 +4088,7 @@ fn push_delimiter( node_index, marker, length, + run_length: length, can_open: roles.can_open, can_close: roles.can_close, single_open: roles.single_open, @@ -4687,7 +4690,7 @@ fn merge_adjacent_text(nodes: &mut Vec) { } /// Index into `openers_bottom` for a nested closer's (marker, both-flags, -/// length%3) key. +/// length % 3) key. fn openers_bottom_key(closer: &DelimMarker) -> usize { let marker = match closer.marker { b'_' => 1, @@ -4697,7 +4700,14 @@ fn openers_bottom_key(closer: &DelimMarker) -> usize { _ => 0, }; let both = usize::from(closer.can_open && closer.can_close); - let modulo = closer.length % 3; + // The key holds what `emphasis_delimiters_match` reads of the closer: the + // whole run's length for the rule of three on `*` / `_`, and what is left + // of the run for the other marks. + let length = match closer.marker { + b'*' | b'_' => closer.run_length, + _ => closer.length, + }; + let modulo = length % 3; ((marker * 2) + both) * 3 + modulo } @@ -4711,13 +4721,15 @@ fn emphasis_delimiters_match(opener: &DelimMarker, closer: &DelimMarker) -> bool b'+' | b'=' => opener.length >= 2 && closer.length >= 2, _ => { // Rule of three: if either delimiter can both open and close, the - // sum of the two run lengths must not be a multiple of three, unless - // both lengths are themselves multiples of three. + // sum of the lengths of the runs containing them must not be a + // multiple of three, unless both are themselves multiples of three. + // CommonMark counts whole runs, not what earlier pairings left. let opener_both = opener.can_open && opener.can_close; let closer_both = closer.can_open && closer.can_close; if opener_both || closer_both { - let sum = opener.length + closer.length; - if sum % 3 == 0 && !(opener.length % 3 == 0 && closer.length % 3 == 0) { + let (opener_run, closer_run) = (opener.run_length, closer.run_length); + let sum = opener_run + closer_run; + if sum % 3 == 0 && !(opener_run % 3 == 0 && closer_run % 3 == 0) { return false; } } @@ -4852,15 +4864,14 @@ enum LinkTarget { /// Matches what follows the `]` at `close` of a link or image whose label is /// `input[label_start..close]`: an inline `(…)` resource, a full or collapsed -/// reference, or a shortcut reference to a defined label. An image whose `(…)` -/// is not a valid resource is no image; a link falls back to the reference -/// forms. Returns the end of the construct and its target. +/// reference, or a shortcut reference to a defined label. A `(…)` that is not a +/// valid resource leaves the reference forms to try, for images and links +/// alike. Returns the end of the construct and its target. fn match_link_target( scan: &mut InlineScan, input: &str, label_start: usize, close: usize, - image: bool, definitions: &[String], ) -> Option<(usize, LinkTarget)> { let label = &input[label_start..close]; @@ -4868,10 +4879,10 @@ fn match_link_target( if input.as_bytes().get(after) == Some(&b'(') { match parse_link_resource(&mut scan.lookups, input, after) { Some((end, resource)) => return Some((end, LinkTarget::Resource(resource))), - // A present-but-invalid `(...)` resource is not an inline link, - // but CommonMark still resolves `[label]` as a shortcut reference - // and leaves the invalid `(...)` as literal text (links 568). - None if image => return None, + // A present-but-invalid `(...)` resource is not an inline link or + // image, but CommonMark still resolves `[label]` as a shortcut + // reference and leaves the invalid `(...)` as literal text (links + // 568). None => {} } } diff --git a/src/serialize.rs b/src/serialize.rs index cdc76ea..c5d0b4f 100644 --- a/src/serialize.rs +++ b/src/serialize.rs @@ -664,6 +664,9 @@ fn serialize_table_row( struct InlineSerializeContext { table_cell: bool, avoid_star_edges: bool, + /// Inside an `_`-delimited emphasis, where a `_` in text that can close + /// would close it on reparse. + in_underscore_emphasis: bool, } impl InlineSerializeContext { @@ -671,13 +674,21 @@ impl InlineSerializeContext { Self { table_cell: true, avoid_star_edges: false, + in_underscore_emphasis: false, } } const fn avoiding_star_edges(self) -> Self { Self { - table_cell: self.table_cell, avoid_star_edges: true, + ..self + } + } + + const fn inside_underscore_emphasis(self) -> Self { + Self { + in_underscore_emphasis: true, + ..self } } } @@ -742,6 +753,33 @@ fn is_gfm_literal_autolink(inline: &Inline) -> bool { ) } +// True when `inline` is a shortcut link or image reference, whose `[label]` +// a following `(` would turn into an inline link, and a following `:` at the +// start of a line into a link reference definition. +fn is_shortcut_reference(inline: &Inline) -> bool { + matches!( + inline, + Inline::LinkReference(LinkReference { + kind: ReferenceKind::Shortcut, + .. + }) | Inline::ImageReference(ImageReference { + kind: ReferenceKind::Shortcut, + .. + }) + ) +} + +// The escaped first character of text after a shortcut reference, when that +// character would re-read the reference's brackets: a `(` always, and a `:` +// when the reference opens the inline sequence, where a line can start. +fn escape_leading_char_after_shortcut(value: &str, reference_index: usize) -> Option<&str> { + match value.as_bytes().first() { + Some(b'(') => Some("\\("), + Some(b':') if reference_index == 0 => Some("\\:"), + _ => None, + } +} + fn is_gfm_literal_email(inline: &Inline) -> bool { matches!( inline, @@ -787,12 +825,20 @@ fn serialize_inlines_with_context( // Leading guard: a non-ASCII char abutting the END of a literal // autolink would merge into its URL on reparse — encode it. + let after_shortcut = index + .checked_sub(1) + .filter(|&prev| is_shortcut_reference(&inlines[prev])); let (lead, body) = match after_literal_autolink .then(|| encode_leading_char_after_autolink(&node.value)) .flatten() { Some((encoded, rest)) => (encoded, rest), - None => (String::new(), node.value.as_str()), + None => match after_shortcut.and_then(|reference| { + escape_leading_char_after_shortcut(&node.value, reference) + }) { + Some(escaped) => (escaped.into(), &node.value[1..]), + None => (String::new(), node.value.as_str()), + }, }; // Trailing guard: when this text is immediately followed by a @@ -825,7 +871,15 @@ fn serialize_inlines_with_context( } Inline::CharacterReference(node) => output.push_str(&node.reference), Inline::Emphasis(node) => { - let children = serialize_inlines_with_context(&node.children, options, context)?; + // Rendered as `_` content first: the choice below reads only + // the children's edges and `*`s, which escaping a `_` that can + // close does not change, so only the `*` choice renders them + // again and nesting never multiplies the work. + let children = serialize_inlines_with_context( + &node.children, + options, + context.inside_underscore_emphasis(), + )?; let touches_underscore = children.starts_with('_') || children.ends_with('_') || children.starts_with("\\_") @@ -1225,7 +1279,6 @@ struct TextScan<'a> { marker_starts: [Positions; ATTENTION_MARKERS.len()], marker_closers: [PathMemo; ATTENTION_MARKERS.len()], last_occurrences: Vec<(&'static str, Option)>, - backtick_runs: Option, dollar_runs: Option, /// The byte, start, and end of the last run `run_len_from` measured. current_run: Option<(u8, usize, usize)>, @@ -1238,7 +1291,6 @@ impl<'a> TextScan<'a> { marker_starts: Default::default(), marker_closers: Default::default(), last_occurrences: Vec::new(), - backtick_runs: None, dollar_runs: None, current_run: None, } @@ -1307,14 +1359,12 @@ impl<'a> TextScan<'a> { } /// Whether some position at or after `from` begins exactly `run_len` - /// trailing bytes of a run of `needle` (an ASCII byte). - fn exact_run_follows(&mut self, needle: u8, from: usize, run_len: usize) -> bool { + /// trailing bytes of a run of `$`. + fn exact_dollar_run_follows(&mut self, from: usize, run_len: usize) -> bool { let input = self.input; - let runs = match needle { - b'`' => &mut self.backtick_runs, - _ => &mut self.dollar_runs, - } - .get_or_insert_with(|| SameCharRuns::new(input, needle)); + let runs = self + .dollar_runs + .get_or_insert_with(|| SameCharRuns::new(input, b'$')); runs.has_run_ending_at_or_after(from + run_len, run_len) } } @@ -1367,6 +1417,7 @@ fn escape_text_with_context( context: InlineSerializeContext, ) -> String { let avoid_star_edges = context.avoid_star_edges; + let in_underscore_emphasis = context.in_underscore_emphasis; let mut output = String::new(); let mut output_line = OutputLine::default(); let mut scan = TextScan::new(input); @@ -1475,7 +1526,10 @@ fn escape_text_with_context( output.push('\\'); output.push(char); } - '`' if text_code_span_can_start(input, offset, &mut scan) => { + // A backslash keeps a backtick from opening a code span but not + // from closing one, so every backtick is escaped: a bare one could + // open a span that an escaped one closes. + '`' => { output.push('\\'); output.push(char); } @@ -1483,6 +1537,10 @@ fn escape_text_with_context( output.push('\\'); output.push(char); } + '_' if in_underscore_emphasis && text_delimiter_can_close(input, offset, 1, true) => { + output.push('\\'); + output.push(char); + } '_' if text_attention_delimiter_can_start(input, offset, "_", true, &mut scan) => { output.push('\\'); output.push(char); @@ -1542,14 +1600,6 @@ fn escape_text_with_context( output } -fn text_code_span_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool { - let marker_len = scan.run_len_from(b'`', offset); - if marker_len == 0 || text_char_at_edge(input, offset, marker_len) { - return true; - } - scan.exact_run_follows(b'`', offset + marker_len, marker_len) -} - fn text_attention_delimiter_can_start( input: &str, offset: usize, @@ -1681,7 +1731,7 @@ fn text_math_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool if marker_len == 0 || text_char_at_edge(input, offset, marker_len) { return true; } - scan.exact_run_follows(b'$', offset + marker_len, marker_len) + scan.exact_dollar_run_follows(offset + marker_len, marker_len) } fn text_tilde_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool { diff --git a/src/serialize/escape_scan_tests.rs b/src/serialize/escape_scan_tests.rs index fe423a9..61f0277 100644 --- a/src/serialize/escape_scan_tests.rs +++ b/src/serialize/escape_scan_tests.rs @@ -12,14 +12,6 @@ use crate::test_support::{boundaries, generated_inputs, query_orders, Rng}; mod reference { use super::super::*; - pub(super) fn text_code_span_can_start(input: &str, offset: usize) -> bool { - let marker_len = same_char_run_len(input, offset, '`'); - if marker_len == 0 || text_char_at_edge(input, offset, marker_len) { - return true; - } - find_same_char_run(input, offset + marker_len, '`', marker_len).is_some() - } - pub(super) fn text_attention_delimiter_can_start( input: &str, offset: usize, @@ -181,11 +173,6 @@ fn run_and_lookahead_checks_match_the_reference_scan() { for_each_scan(22, |input, scan, offset| { let char = input[offset..].chars().next().expect("offset below len"); match char { - '`' => assert_eq!( - text_code_span_can_start(input, offset, scan), - reference::text_code_span_can_start(input, offset), - "{input:?} at {offset}" - ), '$' => assert_eq!( text_math_can_start(input, offset, scan), reference::text_math_can_start(input, offset), diff --git a/tests/delimiter_stack.rs b/tests/delimiter_stack.rs index e3d7d8d..f8b1f91 100644 --- a/tests/delimiter_stack.rs +++ b/tests/delimiter_stack.rs @@ -359,3 +359,32 @@ fn an_image_whose_label_cannot_close_yields_to_a_wikilink() { paragraph.children ); } + +#[test] +fn the_rule_of_three_counts_whole_delimiter_runs() { + let commonmark = SyntaxOptions::commonmark(); + assert_eq!( + parsed(&commonmark, "*a***b*"), + r#"Emphasis["a"]"*"Emphasis["b"]"# + ); + assert_eq!( + parsed(&commonmark, "***a*a*a"), + r#""*"Emphasis[Emphasis["a"]"a"]"a""# + ); +} + +#[test] +fn an_image_whose_resource_is_invalid_falls_back_to_a_shortcut_reference() { + let document = SyntaxOptions::commonmark() + .parse("![foo](a b)\n\n[foo]: /u") + .document; + let Some(Block::Paragraph(paragraph)) = document.children.first() else { + panic!("expected a paragraph"); + }; + assert!( + matches!(paragraph.children.as_slice(), [Inline::ImageReference(image), Inline::Text(rest)] + if image.kind == ReferenceKind::Shortcut && image.identifier == "foo" && rest.value == "(a b)"), + "{:?}", + paragraph.children + ); +} diff --git a/tests/fixtures/roundtrip/spec/commonmark_character_escapes.canonical.md b/tests/fixtures/roundtrip/spec/commonmark_character_escapes.canonical.md index 553452d..f4e08c5 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_character_escapes.canonical.md +++ b/tests/fixtures/roundtrip/spec/commonmark_character_escapes.canonical.md @@ -1,4 +1,4 @@ -!"#$%&'()*+,-./:;\<=>?@\[\]^_`\{|}\~ +!"#$%&'()*+,-./:;\<=>?@\[\]^_\`\{|}\~ \\→\\A\\a\\ \\3\\φ\\« diff --git a/tests/fixtures/roundtrip/spec/commonmark_html_blocks.canonical.md b/tests/fixtures/roundtrip/spec/commonmark_html_blocks.canonical.md index 6fc7c6a..42c07f4 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_html_blocks.canonical.md +++ b/tests/fixtures/roundtrip/spec/commonmark_html_blocks.canonical.md @@ -16,6 +16,6 @@
ok and \bad. +Text ok and \bad. ").document.to_markdown()` runs +- **THEN** the output contains `` diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/untrusted-input-cost.md b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/untrusted-input-cost.md new file mode 100644 index 0000000..1e469de --- /dev/null +++ b/docs/plans/source-span-mapping-and-commonmark-fixes/specs/untrusted-input-cost.md @@ -0,0 +1,27 @@ +# Untrusted input cost — spec changes + +## MODIFIED Requirements + +### Requirement: Linear time +Parsing, `to_markdown`, `to_html`, and `validate` SHALL take time linear in the +input size, except MDX JSX tag matching, which SHALL take at most `n log n`. + +#### Scenario: Unclosed openers +- **WHEN** an input of tens of thousands of unclosed openers of one construct (for example `[`, `++`, ` Date: Tue, 6 Oct 2026 02:28:15 +0000 Subject: [PATCH 39/40] fix: read the spans around a text instead of what follows it The re-review found these regressions from group 19: - Adjacent strongs merged into one wherever no read-back runs (table cells, description lists, directive labels). A closing `**` and an opening `**` form a run of four that by the rule of three closes neither, so the plain rendering writes the second strong with `__` exactly there. - Three escapes guessed the span a text sits in from the delimiter chars written anywhere after or before it, and escaped ordinary prose (`**Note:** use snake_case: here`, `a +\nb`). The context now carries the spans around the inlines and the delimiter beside a text's edges, and the escapes read those. - A reference's raw label backtick was detected in the whole output of a span, code spans included; spans are now judged by the labels of the references inside them. - The render memo now holds only content with a nested span, where nesting can double the work, which brings flat emphasis back to the base's cost. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3 --- src/serialize.rs | 162 ++++++++++++++++++++++++--------- tests/serialize_regressions.rs | 27 ++++++ 2 files changed, 147 insertions(+), 42 deletions(-) diff --git a/src/serialize.rs b/src/serialize.rs index 07776ae..a019ac2 100644 --- a/src/serialize.rs +++ b/src/serialize.rs @@ -1387,6 +1387,13 @@ struct InlineSerializeContext { /// raw label, which an escaped backtick after it could close as a code /// span. raw_backtick_before: bool, + /// The delimiter chars of the spans around the inlines, at every level. + inside: DelimiterChars, + /// The delimiter char of the span whose content the inlines are. + enclosed: Option, + /// For a text, the `+` or `=` of a `++` or `==` delimiter written right + /// before and right after it. + text_edges: (Option, Option), } /// A set of the chars `*`, `_`, `~`, `+`, `=`, `^`, `|`, `$`, and `:`, which @@ -1516,6 +1523,9 @@ impl InlineSerializeContext { raw_edge: None, autolink_edges: AutolinkEdges::Plain, raw_backtick_before: false, + inside: DelimiterChars(0), + enclosed: None, + text_edges: (None, None), } } @@ -1551,6 +1561,9 @@ impl InlineSerializeContext { raw_edge: None, autolink_edges: AutolinkEdges::Plain, raw_backtick_before: false, + inside: DelimiterChars(0), + enclosed: None, + text_edges: (None, None), } } @@ -1566,6 +1579,15 @@ impl InlineSerializeContext { } } + /// For the content of a span delimited by `delimiter`. + fn enclosed_by(self, delimiter: char) -> Self { + Self { + inside: self.inside.union(DelimiterChars::of_char(delimiter)), + enclosed: Some(delimiter), + ..self + } + } + const fn inside_underscore_emphasis(self) -> Self { Self { in_underscore_emphasis: true, @@ -1765,6 +1787,18 @@ impl RenderMemo { options: &SerializeOptions, context: InlineSerializeContext, ) -> Result { + // Only content holding a span can double its work by nesting; + // rendering other content twice is cheaper than keeping it. + let nests = inlines.iter().any(|inline| { + span_children(inline).is_some() + || matches!( + inline, + Inline::Image(_) | Inline::ImageReference(_) | Inline::TextDirective(_) + ) + }); + if !nests { + return render_inlines(self, inlines, options, context); + } let key = (inlines.as_ptr() as usize, inlines.len(), context); if let Some(rendered) = self.0.get(&key) { return Ok(rendered.clone()); @@ -1783,11 +1817,14 @@ fn render_inlines( ) -> Result { let opens_line = context.opens_line; let opens_span = context.opens_span; + let enclosed = context.enclosed; // Nested inlines follow their parent's opening delimiter. let base_context = InlineSerializeContext { opens_line: false, text_opens_line: false, opens_span: false, + enclosed: None, + text_edges: (None, None), ..context }; let mut output = String::new(); @@ -1867,8 +1904,22 @@ fn render_inlines( let at_line_start = output_line.len(&output) == 0; let opens_block_line = breaks_line_start(&output, &mut output_line, opens_line); let at_line_end = text_is_at_line_end(inlines, index); + let doubled_delimiter = |inline: &Inline| match inline { + Inline::Insert(_) => Some('+'), + Inline::Mark(_) => Some('='), + _ => None, + }; + let edge_before = match index.checked_sub(1) { + Some(previous) => doubled_delimiter(&inlines[previous]), + None => enclosed.filter(|_| opens_span), + }; + let edge_after = match inlines.get(index + 1) { + Some(next) => doubled_delimiter(next), + None => enclosed, + }; let text_context = InlineSerializeContext { text_opens_line: opens_block_line, + text_edges: (edge_before, edge_after), ..context }; @@ -2023,7 +2074,7 @@ fn render_inlines( | Inline::Emphasis(_), ) => true, Some(_) => false, - None => written_later.contains('+') || written_later.contains('_'), + None => matches!(enclosed, Some('+' | '_')), }; if (inlines .get(index + 1) @@ -2049,7 +2100,10 @@ fn render_inlines( let children = memo.render( &node.children, options, - context.inside_underscore_emphasis().opening_span(), + context + .inside_underscore_emphasis() + .opening_span() + .enclosed_by('_'), )?; // An escaped `_` at an edge joins no run. let touches_underscore = @@ -2114,24 +2168,28 @@ fn render_inlines( options, context.avoiding_star_edges().opening_span(), )?; - // A `**` right after a `*` that closes a span would join its - // run, so the read-back choice writes the strong with `__` - // there when `_` can flank and its content does not touch - // `_`. The plain rendering keeps `**`: `__` reparses as - // `Underline` when that construct is enabled, which the - // serializer has no signal for. - // So is a strong opening or closing its parent's run, which - // its `**` would lengthen. + // A `**` right after a closing `**` joins it into a run of + // four, which by the rule of three closes neither strong, so + // the strong is written with `__` there when `_` can flank + // and its content does not touch `_`. After a lone closing + // `*` the run of three splits as written, and `__` would read + // back as `Underline` where that construct is enabled, so + // only the read-back choices write `__` there, or at the edge + // of the run around the strong. + let raw_star_edge = context.raw_edge == Some('*'); + let after_strong = ends_with_unescaped(&output, '*') + && ends_with_unescaped(&output[..output.len() - 1], '*') + && !raw_star_edge; let edge_of_run = context.inside_run() && (index == 0 || index + 1 == inlines.len()); - let after_star = matches!( - context.run_style, - RunStyle::StrongUnderscore | RunStyle::EdgeStrongUnderscore - ) && ((ends_with_unescaped(&output, '*') - && context.raw_edge != Some('*')) - || (edge_of_run && context.run_style == RunStyle::EdgeStrongUnderscore) - || children.starts_with('*') - || children.ends_with('*')); + let after_star = after_strong + || (matches!( + context.run_style, + RunStyle::StrongUnderscore | RunStyle::EdgeStrongUnderscore + ) && ((ends_with_unescaped(&output, '*') && !raw_star_edge) + || (edge_of_run && context.run_style == RunStyle::EdgeStrongUnderscore) + || children.starts_with('*') + || children.ends_with('*'))); // A `_` opening the next text is escaped beside the run. let underscore_fits = !children.starts_with('_') && !ends_with_unescaped(&children, '_') @@ -2161,7 +2219,7 @@ fn render_inlines( memo, &node.children, options, - context.opening_span(), + context.opening_span().enclosed_by('_'), )?); output.push_str("__"); } @@ -2186,7 +2244,7 @@ fn render_inlines( memo, &node.children, options, - context.opening_span().delimited_by('+'), + context.opening_span().delimited_by('+').enclosed_by('+'), )?); output.push_str("++"); } @@ -2196,7 +2254,7 @@ fn render_inlines( memo, &node.children, options, - context.opening_span().delimited_by('='), + context.opening_span().delimited_by('=').enclosed_by('='), )?); output.push_str("=="); } @@ -2461,9 +2519,15 @@ fn render_inlines( } } } - if holds_unescaped_backtick(&output[segment_start..]) && holds_reference(inline) { - raw_backtick_before = true; - } + // A reference writes its raw label, whose backtick is unescaped; a + // span is judged by the labels of the references inside it, since + // its output also holds its code spans' backticks. + raw_backtick_before |= match inline { + Inline::FootnoteReference(_) | Inline::LinkReference(_) | Inline::ImageReference(_) => { + holds_unescaped_backtick(&output[segment_start..]) + } + other => holds_raw_label_backtick(other), + }; } Ok(output) } @@ -2476,13 +2540,21 @@ fn holds_unescaped_backtick(written: &str) -> bool { .any(|(index, _)| !ends_with_unescaped(&written[..index], '\\')) } -/// Whether `inline` is a reference or a span holding one, whose raw label -/// the serializer writes as its source. -fn holds_reference(inline: &Inline) -> bool { +/// Whether `inline` is a reference whose raw label holds an unescaped +/// backtick, or a span holding one. +fn holds_raw_label_backtick(inline: &Inline) -> bool { match inline { - Inline::FootnoteReference(_) | Inline::LinkReference(_) | Inline::ImageReference(_) => true, - Inline::Image(node) => node.alt.iter().any(holds_reference), - other => span_children(other).is_some_and(|children| children.iter().any(holds_reference)), + Inline::FootnoteReference(node) => holds_unescaped_backtick(&node.label), + Inline::LinkReference(node) => { + holds_unescaped_backtick(&node.label) + || node.children.iter().any(holds_raw_label_backtick) + } + Inline::ImageReference(node) => { + holds_unescaped_backtick(&node.label) || node.alt.iter().any(holds_raw_label_backtick) + } + Inline::Image(node) => node.alt.iter().any(holds_raw_label_backtick), + other => span_children(other) + .is_some_and(|children| children.iter().any(holds_raw_label_backtick)), } } @@ -2990,16 +3062,16 @@ fn escape_text_with_context( // autolink's local part, so a run is escaped whole. '+' if run_escaped(view, offset, b'+', &mut scan, &mut plus_run, |scan, at| { text_attention_delimiter_can_start(view, at, "++", false, scan) - || text_doubled_delimiter_can_close(view, at, "++", scan) - || text_edge_joins_delimiter(view, at, '+', scan) + || text_doubled_delimiter_can_close(view, at, "++", context.inside) + || text_edge_joins_delimiter(view, at, '+', context.text_edges) }) => { output.push('\\'); output.push(char); } '=' if text_attention_delimiter_can_start(view, offset, "==", false, &mut scan) - || text_doubled_delimiter_can_close(view, offset, "==", &scan) - || text_edge_joins_delimiter(view, offset, '=', &scan) => + || text_doubled_delimiter_can_close(view, offset, "==", context.inside) + || text_edge_joins_delimiter(view, offset, '=', context.text_edges) => { output.push('\\'); output.push(char); @@ -3042,25 +3114,31 @@ fn text_attention_delimiter_can_start( } /// Whether the `+` or `=` at `offset` opens or ends the text beside a `++` or -/// `==` delimiter written before or after it, whose run it would lengthen. -fn text_edge_joins_delimiter(input: &str, offset: usize, char: char, scan: &TextScan) -> bool { - (offset == 0 && scan.written_before.contains(char)) - || (offset + char.len_utf8() == input.len() && scan.written_later.contains(char)) +/// `==` delimiter written right before or after it (`edges`), whose run it +/// would lengthen. +fn text_edge_joins_delimiter( + input: &str, + offset: usize, + char: char, + edges: (Option, Option), +) -> bool { + (offset == 0 && edges.0 == Some(char)) + || (offset + char.len_utf8() == input.len() && edges.1 == Some(char)) } -/// Whether the `++` or `==` at `offset` could close one written before the -/// text, such as the delimiter of the insert or mark the text sits in. +/// Whether the `++` or `==` at `offset` could close an insert or mark the +/// text sits in (`inside`). fn text_doubled_delimiter_can_close( input: &str, offset: usize, marker: &str, - scan: &TextScan, + inside: DelimiterChars, ) -> bool { input[offset..].starts_with(marker) && marker .chars() .next() - .is_some_and(|char| scan.written_before.contains(char)) + .is_some_and(|char| inside.contains(char)) && text_delimiter_can_close(input, offset, marker.len(), false) } diff --git a/tests/serialize_regressions.rs b/tests/serialize_regressions.rs index 7521a96..1a14133 100644 --- a/tests/serialize_regressions.rs +++ b/tests/serialize_regressions.rs @@ -2092,4 +2092,31 @@ mod round_trip_edges { assert_round_trips(source); } } + + #[test] + fn adjacent_strongs_split_wherever_their_inlines_are_written() { + // Without a read-back, the plain rendering itself keeps them apart. + for source in [ + "| h |\n| - |\n| **a**__b__ |", + "::d[**a**__b__]", + "Term **a**__b__\n: def", + "Term\n: **a**__b__", + ] { + assert_round_trips(source); + } + } + + #[test] + fn prose_beside_spans_keeps_its_plain_spelling() { + for source in [ + "**Note:** use snake_case: here", + "*Warning:* set MY_VAR: 1", + "a +\nb + c", + "a =\nb = c", + "if a == b\nthen c== d", + "**See [docs] and `cfg`** then use \\` quote\n\n[docs]: /u", + ] { + assert_eq!(assert_round_trips(source), format!("{source}\n")); + } + } } From 937e35ff4e1d94f23cf261e057e812ec4723529a Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 6 Oct 2026 02:36:38 +0000 Subject: [PATCH 40/40] docs: archive source span mapping and CommonMark fixes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Merges the plan's spec changes into docs/specs (block-syntax, inline-syntax, public-api, serialization, untrusted-input-cost, validation), moves the plan to docs/archive/2026-10-06-…, and points the overview's Next at the panic and the root-cause rewrite in plimeor/markdown-syntax#11. The findings the re-verify left open were accepted for archive and are tracked there. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3 --- .../plan.md | 9 + .../specs/block-syntax.md | 0 .../specs/inline-syntax.md | 0 .../specs/public-api.md | 0 .../specs/serialization.md | 0 .../specs/untrusted-input-cost.md | 0 .../specs/validation.md | 0 docs/overview.md | 13 +- docs/specs/block-syntax.md | 182 +++++++++++- docs/specs/inline-syntax.md | 64 +++++ docs/specs/public-api.md | 111 +++++++- docs/specs/serialization.md | 260 +++++++++++++++++- docs/specs/untrusted-input-cost.md | 12 + docs/specs/validation.md | 17 +- 14 files changed, 641 insertions(+), 27 deletions(-) rename docs/{plans/source-span-mapping-and-commonmark-fixes => archive/2026-10-06-source-span-mapping-and-commonmark-fixes}/plan.md (98%) rename docs/{plans/source-span-mapping-and-commonmark-fixes => archive/2026-10-06-source-span-mapping-and-commonmark-fixes}/specs/block-syntax.md (100%) rename docs/{plans/source-span-mapping-and-commonmark-fixes => archive/2026-10-06-source-span-mapping-and-commonmark-fixes}/specs/inline-syntax.md (100%) rename docs/{plans/source-span-mapping-and-commonmark-fixes => archive/2026-10-06-source-span-mapping-and-commonmark-fixes}/specs/public-api.md (100%) rename docs/{plans/source-span-mapping-and-commonmark-fixes => archive/2026-10-06-source-span-mapping-and-commonmark-fixes}/specs/serialization.md (100%) rename docs/{plans/source-span-mapping-and-commonmark-fixes => archive/2026-10-06-source-span-mapping-and-commonmark-fixes}/specs/untrusted-input-cost.md (100%) rename docs/{plans/source-span-mapping-and-commonmark-fixes => archive/2026-10-06-source-span-mapping-and-commonmark-fixes}/specs/validation.md (100%) diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/plan.md similarity index 98% rename from docs/plans/source-span-mapping-and-commonmark-fixes/plan.md rename to docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/plan.md index 4da200f..10c2cc3 100644 --- a/docs/plans/source-span-mapping-and-commonmark-fixes/plan.md +++ b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/plan.md @@ -302,6 +302,15 @@ Specs: - a container directive opener inside an HTML block, which the directive counts as nested (`:::e\n\n:::e`). +- [Findings accepted at archive] → The re-verify of group 19 found that its + quote and list column tracking (`OpenParagraphIn`, `OpenBlock.item_column`) + fixes 378 generated nested-container inputs and newly breaks 13, a panic in + `prefix_ends_with_gfm_email` on `"\u{a0}e+@"` that predates this plan, and + gaps in the growth and corpus tests. They were accepted for archive and are + tracked, with the residual classes above, in plimeor/markdown-syntax#11, + which replaces the container re-prediction and the serializer's copies of + parser rules at their root. + - [Tabs after a split tab] → Once a container marker splits a tab, the later tabs that open the line are expanded to spaces, inside a fence or HTML block too: `>\t\t\tfoo` gives a code block of six spaces and `foo` where diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/block-syntax.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/block-syntax.md similarity index 100% rename from docs/plans/source-span-mapping-and-commonmark-fixes/specs/block-syntax.md rename to docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/block-syntax.md diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md similarity index 100% rename from docs/plans/source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md rename to docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/public-api.md similarity index 100% rename from docs/plans/source-span-mapping-and-commonmark-fixes/specs/public-api.md rename to docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/public-api.md diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/serialization.md similarity index 100% rename from docs/plans/source-span-mapping-and-commonmark-fixes/specs/serialization.md rename to docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/serialization.md diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/untrusted-input-cost.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/untrusted-input-cost.md similarity index 100% rename from docs/plans/source-span-mapping-and-commonmark-fixes/specs/untrusted-input-cost.md rename to docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/untrusted-input-cost.md diff --git a/docs/plans/source-span-mapping-and-commonmark-fixes/specs/validation.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/validation.md similarity index 100% rename from docs/plans/source-span-mapping-and-commonmark-fixes/specs/validation.md rename to docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/validation.md diff --git a/docs/overview.md b/docs/overview.md index 128d490..378efc7 100644 --- a/docs/overview.md +++ b/docs/overview.md @@ -1,5 +1,5 @@ # Project overview -Updated 2026-10-05 +Updated 2026-10-06 ## What this is @@ -26,14 +26,15 @@ Shipped features are described in `docs/specs/`. ## Current focus -Releasing the delimiter-stack inline parser and CommonMark input handling -(leading BOM, NUL) as a SemVer-breaking version. +Releasing the delimiter-stack inline parser, CommonMark input handling (leading +BOM, NUL), and source-mapped spans as a SemVer-breaking version. ## Next -- Inline spans on lines that carry leading whitespace inside their block are - offset by the stripped whitespace; spans need mapping back to source - coordinates. +- `parse` panics on `"\u{a0}e+@"`; fix it before the release. +- Rebuild block parsing on one stack of open blocks, take the serializer's + escapes and delimiter choices from the parser, and fold autolinks into + `Link` (plimeor/markdown-syntax#11). ## Non-goals diff --git a/docs/specs/block-syntax.md b/docs/specs/block-syntax.md index e9f8871..9ad995c 100644 --- a/docs/specs/block-syntax.md +++ b/docs/specs/block-syntax.md @@ -17,6 +17,98 @@ specification defines them. - **WHEN** `"# Title\n\nHello."` is parsed - **THEN** the document holds a `Heading` of depth 1 followed by a `Paragraph` +#### Scenario: Lazy list marker ends a block quote +- **WHEN** `"> a\n- "` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` with a paragraph `a`, followed by a `List` holding one empty item + +#### Scenario: Lazy ordered item not starting at 1 +- **WHEN** `"> > a\n2. b"` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` followed by an ordered `List` starting at 2 + +#### Scenario: Final whitespace of a paragraph +- **WHEN** `"aaa \nbbb "` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("aaa")`, a `LineBreak`, and `Text("bbb")` + +#### Scenario: Final whitespace of a setext heading +- **WHEN** `"Foo \n-----"` is parsed with the CommonMark preset +- **THEN** the document holds a level-2 setext `Heading` holding `Text("Foo")` + +#### Scenario: Lazy line in an item that started blank +- **WHEN** `"- \n a\nb"` is parsed with the CommonMark preset +- **THEN** the list's one item holds a paragraph `a`, a soft break, and `b` + +#### Scenario: Blank line indented four columns ends a block quote +- **WHEN** `"> a\n \n> b"` is parsed with the CommonMark preset +- **THEN** the document holds two `BlockQuote`s + +#### Scenario: Blank line between an empty item and the next +- **WHEN** `"* \n\n * b c"` is parsed with the CommonMark preset +- **THEN** the document holds one loose `List` of two items + +#### Scenario: Complete HTML tag on a lazy line ends a list item +- **WHEN** `"- a\n"` is parsed with the CommonMark preset +- **THEN** the document holds a `List` followed by an `HtmlBlock` + +#### Scenario: ATX-like line that is no heading +- **WHEN** `"a\n#)"` is parsed with the CommonMark preset +- **THEN** the document holds one `Paragraph` holding `Text("a")`, a `SoftBreak`, and `Text("#)")` + +#### Scenario: Backtick run with a backtick in its info +- **WHEN** ````"a\n``` `` ```"```` is parsed with the CommonMark preset +- **THEN** the document holds one `Paragraph` whose second line is a code span + +#### Scenario: Unclosed fence in a block quote +- **WHEN** ``"> ```\n> x\na"`` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` whose fenced `CodeBlock` holds `"x\n"`, followed by a `Paragraph` holding `Text("a")` + +#### Scenario: Last line ending of indented code +- **WHEN** `"\ta\r\tb"` is parsed with the CommonMark preset +- **THEN** the document holds an indented `CodeBlock` whose value is `"a\rb\r"` + +#### Scenario: Container content ending in a carriage return +- **WHEN** ``"- ```\n ~\r"`` is parsed with the CommonMark preset +- **THEN** the list item's fenced `CodeBlock` holds `"~\n"` + +#### Scenario: Complete tag after a definition +- **WHEN** `"[o]: u\n"` is parsed with the CommonMark preset +- **THEN** the document holds a `Definition` followed by a `Paragraph` holding an `Html` inline + +#### Scenario: Indented table header row +- **WHEN** `"a\n |b\n----"` is parsed with `parse` +- **THEN** the document holds one setext `Heading` + +#### Scenario: Tab after a top-level block quote marker +- **WHEN** `"> \tcode"` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` holding a `Paragraph`, since the tab spans columns 2 to 4 + +#### Scenario: Quoted paragraph opening with backticks +- **WHEN** `"> ``a\nb"` is parsed with the CommonMark preset +- **THEN** the document holds one `BlockQuote` whose paragraph ends with `Text("b")` + +#### Scenario: Blank line inside a nested item's open fence +- **WHEN** ``"2. a\n 1. ```\n\n2. b"`` is parsed with the CommonMark preset +- **THEN** the document holds one tight `List` + +#### Scenario: Tab inside nested containers +- **WHEN** `"* - \tb c"` is parsed with the CommonMark preset +- **THEN** the inner list item holds an indented `CodeBlock`, since the tab spans columns 4 to 8 + +#### Scenario: Lazy list marker inside a quoted item +- **WHEN** `"> - a\n> 2.\nz"` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` holding two lists, followed by a `Paragraph` holding `Text("z")` + +#### Scenario: Setext-like line short of a quoted item +- **WHEN** `"> 1. a\n> ===\nb"` is parsed with the CommonMark preset +- **THEN** the item holds one `Paragraph` holding `Text("a")`, `Text("===")`, and `Text("b")` with soft breaks between + +#### Scenario: Sibling item ending a nested fence +- **WHEN** ``"- - ```\n - a\n\n- b"`` is parsed with the CommonMark preset +- **THEN** the outer `List` is loose + +#### Scenario: Lazy fence-like line in an item +- **WHEN** ``"1. a\n ```\n\nb"`` is parsed with the CommonMark preset +- **THEN** the document holds a `List` followed by a `Paragraph` holding `Text("b")` + #### Scenario: CommonMark oracle cases - **WHEN** the block cases under `tests/fixtures/conformance/commonmark/` are parsed and rendered with the `html` feature - **THEN** the output matches the expected HTML @@ -50,7 +142,8 @@ containers (`details` / `summary`). ### Requirement: Block directives When directives are enabled, the parser SHALL recognize `::name[label]{attrs}` leaf directives and `:::name` container directives, whose container closes at a -fence of at least the opening fence's length. +fence of at least the opening fence's length, and a line with a malformed +opener SHALL NOT end a paragraph. #### Scenario: Container directive - **WHEN** `":::note\nbody\n:::"` is parsed with `parse` @@ -64,6 +157,14 @@ fence of at least the opening fence's length. - **WHEN** a leaf directive opener has a malformed name - **THEN** an error-severity `InvalidDirectiveName` diagnostic is reported +#### Scenario: Directive attribute without a valid name +- **WHEN** `":b{<} :c{a <=1 d}"` is parsed with `parse` and serialized +- **THEN** `to_markdown()` returns `":b :c{a d}\n"` + +#### Scenario: Malformed directive line inside a paragraph +- **WHEN** `"a\n::1bad"` or `"a\n:::"` is parsed with `parse` +- **THEN** the document holds one `Paragraph` holding both lines + ### Requirement: Directives are not MDX The parser SHALL treat `:name`, `::name`, and `:::name` as the directive family under every option set and SHALL never parse them as MDX. @@ -113,3 +214,82 @@ pipes it held as text in the cell. #### Scenario: Empty cell between bars - **WHEN** `"| x | y | z |\n|---|---|---|\n|a||b|"` is parsed with `parse` - **THEN** the body row's cells are `a`, empty, and `b` + +### Requirement: Spaces and tabs are block whitespace +The parser SHALL read only spaces and tabs as whitespace in block structure: a +blank line holds only spaces and tabs; indentation, the space after a block +marker, the trailing whitespace a thematic break, setext underline, closing +fence, ATX closing sequence, HTML block start line, definition, or table row +allows, the indentation before flow MDX JSX, and the final whitespace of a +paragraph are spaces and tabs. Any +other whitespace char, such as a no-break space or a form feed, is content. + +#### Scenario: Other whitespace after a thematic break +- **WHEN** `"***\u{a0}"` is parsed with the CommonMark preset +- **THEN** the document holds a `Paragraph`, not a `ThematicBreak` + +#### Scenario: Line holding only a no-break space +- **WHEN** `"a\n\u{a0}\nb"` is parsed with the CommonMark preset +- **THEN** the document holds one `Paragraph` + +#### Scenario: No-break space before MDX JSX +- **WHEN** `"\u{a0}

"` is parsed with the MDX preset +- **THEN** the document holds a `Paragraph`, not a flow `MdxJsx` block + +#### Scenario: Form feed ending a paragraph +- **WHEN** `"a\u{c}"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("a\u{c}")` + +### Requirement: Fenced code inside a container directive +A fenced code block inside a container directive SHALL hold its lines as code: +a line in it that looks like a directive opener opens no nested directive, +while a closing fence of the directive still closes it, and a fence that a +nested directive leaves open ends with that directive. + +#### Scenario: Directive opener inside fenced code +- **WHEN** `":::t\n```\n:::e\n```\n:::"` is parsed +- **THEN** the document holds one `ContainerDirective` named `t` holding a `CodeBlock` whose value is `":::e\n"` + +#### Scenario: Fence left open in a nested directive +- **WHEN** ``":::outer\n:::inner\n```\n:::\n:::inner2\nx\n:::\n:::\nafter"`` is parsed +- **THEN** the `ContainerDirective` named `outer` holds the directives `inner` and `inner2`, a `Paragraph` holding `Text("after")` follows it, and no diagnostic is reported + +### Requirement: Footnote definition content +A footnote definition's content SHALL start after the spaces and tabs that +follow its `]:` and keep the trailing spaces of its first line, which may +make a hard break. + +#### Scenario: Hard break on a footnote definition's first line +- **WHEN** `"[^1]: a \nb"` is parsed +- **THEN** the definition's paragraph holds `Text("a")`, a `LineBreak`, and `Text("b")` + +### Requirement: List after a definition +A list SHALL start on the line right after a definition only when its first +item could interrupt a paragraph: a bullet or an ordered item starting at 1, +with content. Otherwise the line continues the paragraph the definition was +read from. + +#### Scenario: Ordered item not starting at 1 after a definition +- **WHEN** `"[foo]: /url\n2) a"` is parsed with the CommonMark preset +- **THEN** the document holds the `Definition` and a `Paragraph` + +### Requirement: GFM table start +A GFM table SHALL start only where the line after its header row is a +delimiter row that is neither a lazy continuation line nor a setext +underline; such a line keeps its other reading. + +#### Scenario: Delimiter row without pipes +- **WHEN** `"| --- |\n-- "` is parsed with the GFM preset +- **THEN** the document holds a level-2 setext `Heading`, not a `Table` + +#### Scenario: Setext underline below a header row +- **WHEN** `"a\n|b\n---"` is parsed with the GFM preset +- **THEN** the document holds one level-2 setext `Heading` holding both lines + +#### Scenario: Lazy delimiter row +- **WHEN** `"1. ---(\n:-:"` is parsed with the GFM preset +- **THEN** the list item holds a `Paragraph`, not a `Table` + +#### Scenario: Table header row that looks like an empty list item +- **WHEN** `"a\n+\n|-"` is parsed with the GFM preset +- **THEN** the document holds a `Paragraph` and a `Table` whose header cell holds `+` diff --git a/docs/specs/inline-syntax.md b/docs/specs/inline-syntax.md index 155828d..af31aa8 100644 --- a/docs/specs/inline-syntax.md +++ b/docs/specs/inline-syntax.md @@ -28,6 +28,26 @@ spans, links, and emphasis. - **WHEN** `"[foo][bar\n\n[foo]: /u"` is parsed with the CommonMark preset - **THEN** the paragraph holds a shortcut `LinkReference` to `foo` followed by `Text("[bar")` +#### Scenario: Rule of three counts whole delimiter runs +- **WHEN** `"*a***b*"` is parsed with the CommonMark preset +- **THEN** the paragraph holds an `Emphasis` containing `a`, `Text("*")`, and an `Emphasis` containing `b` + +#### Scenario: Image whose resource is invalid +- **WHEN** `"![foo](a b)\n\n[foo]: /u"` is parsed with the CommonMark preset +- **THEN** the paragraph holds a shortcut `ImageReference` to `foo` followed by `Text("(a b)")` + +#### Scenario: Underscore after Unicode punctuation +- **WHEN** `"«_**]**_"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("«")` and an `Emphasis` containing a `Strong` containing `]` + +#### Scenario: Escaped backslash before a line ending +- **WHEN** `"a\\\\\nb"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("a\\")`, a `SoftBreak`, and `Text("b")` + +#### Scenario: Space inside a bare destination's parentheses +- **WHEN** `"[a](( ))"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("[a](( ))")` and no `Link` + #### Scenario: CommonMark oracle cases - **WHEN** the inline cases under `tests/fixtures/conformance/commonmark/` are parsed and rendered with the `html` feature - **THEN** the output matches the expected HTML @@ -195,3 +215,47 @@ with a closer inside the same label. #### Scenario: Emphasis opener before a link - **WHEN** `"*[foo*](/u)"` is parsed with `parse` - **THEN** the paragraph holds `Text("*")` followed by a `Link` whose text is `foo*` + +### Requirement: Footnote labels +The parser SHALL read `[^label]` as a footnote reference, and `[^label]:` as a +footnote definition, only when the label is non-empty, holds no space, tab, or +line ending, and, as a link label, holds no unescaped `[` or `]`. + +#### Scenario: Bracket inside a footnote label +- **WHEN** `"^*[^[^]]"` and `"[^a[b]"` are parsed with `parse` +- **THEN** neither paragraph holds a `FootnoteReference`, while `"[^a\\[b]"` holds one + +### Requirement: Reference label matching +Two link labels SHALL match when they agree after Unicode case folding, +trimming, and collapsing each run of spaces, tabs, and line endings to one +space; any other whitespace char is matched as written. + +#### Scenario: No-break space in a label +- **WHEN** `"[a\u{a0}b]\n\n[a b]: /u"` is parsed +- **THEN** the paragraph holds no `LinkReference` + +### Requirement: Angle-bracket autolink URI +The parser SHALL read `` as an autolink when the scheme is valid +and the rest holds no space, ASCII control char, `<`, or `>`; any other +whitespace char is part of the URI. + +#### Scenario: No-break space in an angle-bracket autolink +- **WHEN** `""` is parsed with the CommonMark preset +- **THEN** the paragraph holds an `Autolink` to `http://a\u{a0}b` + +### Requirement: Hard line breaks from spaces +A line ending SHALL be a hard break when two or more spaces the source holds, +and no tab, end the line; spaces or tabs a character reference writes are +text, and only the source's spaces and tabs before a soft break are removed. + +#### Scenario: Referenced space before a line ending +- **WHEN** `"a \nb"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("a ")`, a `SoftBreak`, and `Text("b")` + +### Requirement: Processing instructions +Raw inline HTML SHALL read `` +after the `` is text +- **WHEN** `"a b"` is parsed with the CommonMark preset +- **THEN** the paragraph holds no `Html` inline diff --git a/docs/specs/public-api.md b/docs/specs/public-api.md index e7f71f1..5a33656 100644 --- a/docs/specs/public-api.md +++ b/docs/specs/public-api.md @@ -137,6 +137,14 @@ span to 1-based line and column positions. - **WHEN** a `Heading` is built with `Heading::new(1, [Text::from("Title")])` - **THEN** its `span()` is `None` +#### Scenario: Empty table cell +- **WHEN** `parse("| a | |\n|-|-|")` runs +- **THEN** the second header cell's span is the empty range at byte 6, just before the pipe that closes it + +#### Scenario: Missing table cell +- **WHEN** `parse("| a | b |\n|-|-|\n| c")` runs +- **THEN** the body row's second cell has no children and its span is the empty range at the end of the row, byte 19 + ### Requirement: Top-level spans tile the source The spans of a parsed document's top-level blocks SHALL be in source order, non-overlapping, on UTF-8 character boundaries, within the input, and separated @@ -192,11 +200,9 @@ stay in the coordinates of the original input. ### Requirement: Emphasis-like spans cover their delimiters The span of a parsed emphasis-like container (`Emphasis`, `Strong`, `Underline`, `Delete`, `Insert`, `Mark`, `Spoiler`, `Subscript`, or -`Superscript`) in a block whose lines carry no leading whitespace after their container -markers SHALL -run from the first character of the delimiters that open it to the last -character of the delimiters that close it, and SHALL lie within the span of the -node that contains it. +`Superscript`) SHALL run from the first character of the delimiters that open +it to the last character of the delimiters that close it, and SHALL lie within +the span of the node that contains it. #### Scenario: Strong inside emphasis - **WHEN** `parse("***a***")` runs @@ -209,3 +215,98 @@ node that contains it. #### Scenario: Leftover opening delimiter - **WHEN** `parse("**a*")` runs - **THEN** the paragraph holds `Text("*")` spanning bytes 0..1 and an `Emphasis` spanning bytes 1..4 + +#### Scenario: Emphasis on a block quote continuation line +- **WHEN** `parse("> a\n> *b*")` runs +- **THEN** the block quote's paragraph holds an `Emphasis` spanning bytes 6..9 + +### Requirement: Spans map stripped lines back to the source +A parsed node's span SHALL end after the source byte where its last character +was read, and SHALL start at the source byte where its first character was +read, or for a block, where its first line starts after the markers and +indentation of the containers around it. This SHALL hold wherever the parser +removes indentation, container markers, or table-cell padding before reading a +line, joins lines whose source line ending is `\r\n`, or reads `\|` in a table +cell as `|`, which maps to both of its bytes; a space the parser produces by +splitting a tab SHALL map to that tab. + +#### Scenario: Leading whitespace on a paragraph line +- **WHEN** `parse(" a *b*")` runs +- **THEN** the paragraph holds `Text("a ")` spanning bytes 2..4 and an `Emphasis` spanning bytes 4..7 + +#### Scenario: Block quote continuation line +- **WHEN** `parse("> a\n> b *c*")` runs +- **THEN** the block quote's paragraph spans bytes 2..11 and holds `Text("b ")` spanning bytes 6..8 and an `Emphasis` spanning bytes 8..11 + +#### Scenario: List item continuation line +- **WHEN** `parse("- a\n b *c*")` runs +- **THEN** the item's paragraph spans bytes 2..11 and holds an `Emphasis` spanning bytes 8..11 + +#### Scenario: Later block inside a block quote +- **WHEN** `parse("> a\n>\n> b")` runs +- **THEN** the block quote's second paragraph spans bytes 8..9 + +#### Scenario: Nested list item after multi-byte text +- **WHEN** `parse("- 项目\n - 嵌套 [[library/工作/买菜]]\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 20..45 + +#### Scenario: Nested block quote +- **WHEN** `parse("> 外层\n> > 内层 [[A]]\n")` runs +- **THEN** the inner block quote's `WikiLink` spans bytes 20..25 + +#### Scenario: Alert body +- **WHEN** `parse("> [!NOTE]\n> 见 [[A]]\n")` runs +- **THEN** the alert's `WikiLink` spans bytes 16..21 + +#### Scenario: Footnote definition continuation line +- **WHEN** `parse("正文[^1]\n\n[^1]: 见 [[A]]\n 续 [[B]]\n")` runs +- **THEN** the footnote definition's second `WikiLink` spans bytes 36..41 + +#### Scenario: Inside an HTML container +- **WHEN** `parse("

\n更多\n\n- 项目\n - [[A]]\n\n
\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 50..55 + +#### Scenario: Inside a container directive +- **WHEN** `parse(":::note\n- 项目\n - [[A]]\n:::\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 21..26 + +#### Scenario: Tab-indented nested list +- **WHEN** `parse("- 项目\n\t- 嵌套 [[A]]\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 19..24 + +#### Scenario: CRLF nested list +- **WHEN** `parse("- 项目\r\n - 嵌套 [[A]]\r\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 21..26 + +#### Scenario: CRLF soft break +- **WHEN** `parse("a\r\nb")` runs +- **THEN** the paragraph holds a `SoftBreak` spanning bytes 1..3 and `Text("b")` spanning bytes 3..4 + +#### Scenario: Table cell content +- **WHEN** `parse("| a *b* |\n|-|")` runs +- **THEN** the header cell spans bytes 2..7 and holds `Text("a ")` spanning bytes 2..4 and an `Emphasis` spanning bytes 4..7 + +#### Scenario: Escaped pipe in a table cell +- **WHEN** `parse("| a\\|b |\n|-|")` runs +- **THEN** the header cell holds `Text("a|b")` spanning bytes 2..6 + +#### Scenario: Escaped pipe opening a table cell +- **WHEN** `parse("| \\|a |\n|-|")` runs +- **THEN** the header cell holds `Text("|a")` spanning bytes 2..5 + +#### Scenario: Split tab +- **WHEN** `parse(">\t\tfoo")` runs +- **THEN** the block quote holds an indented code block with value `" foo\n"` spanning bytes 1..6 + +#### Scenario: Container span regression cases +- **WHEN** `inline_spans_address_source_inside_containers` in `tests/parse_span_contract.rs` parses each of its 31 cases (lists, task lists, block quotes, alerts, tables, footnote definitions, HTML containers, container directives, frontmatter, CRLF, and tabs) +- **THEN** every `WikiLink`, `Link`, `Image`, and `#`-holding `Text` it collects spans exactly the literal it occupies in the input + +### Requirement: Spans nest +Every parsed node's span SHALL lie on UTF-8 character boundaries within the +input and within the span of the node that contains it, and the spans of a +node's children SHALL be in source order and SHALL NOT overlap. + +#### Scenario: Fixture corpus and generated inputs +- **WHEN** every fixture input and every seeded generated input is parsed in each dialect +- **THEN** every node, at every depth, satisfies these conditions diff --git a/docs/specs/serialization.md b/docs/specs/serialization.md index 444aaba..aed2b8a 100644 --- a/docs/specs/serialization.md +++ b/docs/specs/serialization.md @@ -8,14 +8,19 @@ by `src/serialize.rs`. ## Requirements ### Requirement: Canonical output -`Document::to_markdown` SHALL emit canonical Markdown: one fixed spelling per -construct, chosen by `SerializeOptions`, independent of how the source spelled -it. +`Document::to_markdown` SHALL emit canonical Markdown: for each construct, the +spelling the AST records for it, such as a list marker, a fence's char and +length, a heading's style, or a reference's kind, or else one fixed spelling, +independent of source details the AST does not record. #### Scenario: Paragraph and heading - **WHEN** `parse("# Title\n\nHello *world*.").document.to_markdown()` runs - **THEN** it returns `"# Title\n\nHello *world*.\n"` +#### Scenario: Marker the AST records +- **WHEN** `parse("+ a").document.to_markdown()` runs +- **THEN** it returns `"+ a\n"` + ### Requirement: Round-trip stability For a parsed document, parsing the serialized Markdown SHALL yield the same AST apart from spans, and serializing that reparsed document SHALL yield the same @@ -26,14 +31,19 @@ text. - **THEN** the reparsed AST matches the first and the two serialized texts are identical ### Requirement: Serialize options -`SerializeOptions` SHALL control the line ending, the trailing newline, the -bullet marker, the ordered-list delimiter, and the code fence character, and -SHALL be constructed by mutating `SerializeOptions::default()`. +`SerializeOptions` SHALL control the line ending and the trailing newline; a +bullet marker, ordered-list delimiter, or code fence character other than its +default SHALL replace the one the AST records, while the default keeps it. +Options SHALL be constructed by mutating `SerializeOptions::default()`. #### Scenario: CRLF without final newline - **WHEN** `parse("# Title").document.to_markdown_with(&options)` runs with `line_ending = LineEnding::CrLf` and `final_newline = false` - **THEN** it returns `"# Title"` +#### Scenario: Bullet override +- **WHEN** `parse("- a\n\n+ b").document.to_markdown_with(&options)` runs with `bullet = ListDelimiter::Plus` +- **THEN** both lists are written with `+` + ### Requirement: Invalid documents are rejected Serialization SHALL validate the document first and return `SerializeError::InvalidDocument` with the validation diagnostics when it is @@ -45,17 +55,249 @@ invalid, and `SerializeError::UnsupportedNode` for a node kind it cannot write. ### Requirement: Escaping keeps text literal The serializer SHALL escape text so that reparsing the output yields the same -text rather than new constructs. +text and leaves the nodes beside it unchanged, rather than forming new +constructs. #### Scenario: Literal asterisks in text - **WHEN** a hand-built paragraph holding `Text("*not emphasis*")` is serialized and reparsed - **THEN** the reparsed paragraph holds the same text and no `Emphasis` +#### Scenario: Underscore that can close inside underscore emphasis +- **WHEN** a hand-built paragraph holding an `Emphasis` around a `Strong` around `Text("(a b)_.")`, followed by `Text("*#")`, is serialized and reparsed with the CommonMark preset +- **THEN** `to_markdown()` returns `"_**(a b)\\_.**_\\*#\n"` and the reparsed paragraph equals the original apart from spans + +#### Scenario: Parenthesis after a shortcut reference +- **WHEN** the document parsed from `"[foo]\\(a)\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed paragraph holds a shortcut `LinkReference` to `foo` followed by `Text("(a)")` + +#### Scenario: Parenthesis after a shortcut image reference +- **WHEN** the document parsed from `"![foo]\\(a)\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed paragraph holds a shortcut `ImageReference` to `foo` followed by `Text("(a)")` + +#### Scenario: Colon after a shortcut reference that starts a paragraph +- **WHEN** the document parsed from `"[foo]\\: /x\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed document still holds the paragraph, with a shortcut `LinkReference` to `foo` followed by `Text(": /x")` + +#### Scenario: Pipe ending a level-two setext heading +- **WHEN** the document parsed from `"a |\n-"` is serialized and reparsed +- **THEN** `to_markdown()` returns `"a \\|\n---\n"` and the reparsed document holds the same setext `Heading` and no `Table` + +#### Scenario: Empty fenced code block +- **WHEN** ``parse("```\n```").document.to_markdown()`` runs +- **THEN** it returns ``"```\n```\n"`` + +#### Scenario: Whitespace at the ends of an info string +- **WHEN** the document parsed from ``"``` a \nb\n```"`` is serialized +- **THEN** `to_markdown()` returns ``"``` a \nb\n```\n"`` and reparsing it yields the info string `" a\t"` + +#### Scenario: Text right after a literal autolink +- **WHEN** the documents parsed from `"://&"` and `"www.}"` are serialized and reparsed +- **THEN** each reparsed paragraph holds the same `Autolink` and `Text` as the parsed one + +#### Scenario: Paragraph that opens with a soft break +- **WHEN** the document parsed from `" \na"` is serialized +- **THEN** `to_markdown()` returns `" \na\n"` + +#### Scenario: Text line that would open a block +- **WHEN** the documents parsed from `"a\n\\
"` and `"a\n\\::b"` are serialized and reparsed +- **THEN** each reparsed document holds the same single `Paragraph` + +#### Scenario: HTML block value +- **WHEN** the document parsed from `"