diff --git a/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/plan.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/plan.md new file mode 100644 index 0000000..10c2cc3 --- /dev/null +++ b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/plan.md @@ -0,0 +1,464 @@ +# Source span mapping and CommonMark fixes + +## Why + +Spans are absolute only up to the first stripped byte. Container content (block +quotes, list items, and the other containers) is joined into one string and +re-read from a single base offset, and paragraph, heading, and table-cell inline +input is joined the same way. So every line after a stripped prefix, a CRLF, or +a cell pipe is shifted: block and inline spans slice the wrong text, and code +blocks inside containers get spans past the end of the input. Issue +plimeor/markdown-syntax#6 measures this on 31 container cases, 26 of which fail +on 0.3.0. Editors that rewrite source in place by span then panic when a span +lands inside a multi-byte character, or silently replace the wrong bytes. +Separately, three +CommonMark and round-trip defects predate the delimiter-stack change: the rule +of three uses the remaining run length, an image with an invalid `(…)` does not +fall back to a reference, and the serializer writes text that reparses as a code +span, an inline link, or a link reference definition. + +## What changes + +- Every parsed node's span maps to the source bytes it was read from, on every + line of every container, across CRLF line endings, inside table cells, and + across split tabs. This resolves issue plimeor/markdown-syntax#6, whose + 31-case regression test joins the suite. +- Table cells carry spans; today they carry none. +- Every span lies within its parent's span, in source order among its siblings, + and this holds for the whole tree. +- The leading-whitespace qualifier on "Emphasis-like spans cover their + delimiters" is removed. +- The rule of three and its opener floor use each delimiter run's original + length, as CommonMark specifies. +- An image whose `(…)` is not a valid resource falls back to the full, + collapsed, and shortcut reference forms, as a link does. +- The serializer writes every backtick in text as `` \` ``. This changes + canonical output for text that holds a backtick it left bare before. +- The serializer escapes text after a shortcut link or image reference so that + reparsing keeps the reference: a `(` directly after it, and a `:` directly + after one that starts a paragraph. +- A lazy line that opens a list ends the block quote it would otherwise + continue, including an empty item and an ordered item not starting at 1. +- A paragraph and a setext heading drop the final whitespace of their content. + A level-two setext heading whose text ends in `|` is written with that pipe + escaped, so its `---` underline is not read as a table delimiter row. +- Lazy lines continue a paragraph in list items and block quotes as cmark, + commonmark.js, and micromark read them: after an item that started blank, + after a single-line block, across nested containers, and at the quote level + the paragraph sits in. A blank line ends a block quote however far it is + indented, and a blank line between two items loosens the list. +- A complete HTML tag (a type-7 HTML block start) on a lazy line ends a list + item, as it already ends a block quote. +- Only a line that opens an ATX heading or a code fence interrupts a + paragraph: `#)` and a backtick run whose info string holds a backtick + continue it. +- A line without a `>` after an unclosed fence, math block, or HTML block in a + block quote ends the quote rather than continuing that block, and a fence + that closes leaves the quote's next paragraph open to lazy lines again. +- A code or math block that the input ends ends its last line with the value's + first line ending, `\n` when it has none. +- The serializer writes these so that they reparse to the same tree: an empty + fenced code block, spaces and tabs at the ends of an info string, text right + after a literal autolink, a paragraph that opens with a soft break, an HTML + block value, a text line that would open an HTML block or a leaf directive, + and an indented code block ending in `\r` or `\r\n`. The CRLF line-ending + option leaves a value's `\r\n` as it is. +- A `_` run preceded or followed by Unicode punctuation opens and closes as + one beside ASCII punctuation does; an escaped backslash before a line ending + is text, not a hard break; a bare link destination ends at a space inside + parentheses; a line with a malformed directive opener does not end a + paragraph; and a container's last content line ends in `\n` like the others. +- The serializer escapes a run of `*`, `_`, or `$` in text whole or not at + all, reads the text's neighbours as the reparse sees them (delimiters the + later inlines write, tabs, and chars written as references), writes `_` + emphasis only where `_` can open and close, indents a continuation line + inside an inline that would start a block, writes a break that opens a line + or a delimited span after a ` `, writes a space in a bare destination + as a reference, starts a list item whose first block opens with whitespace + on the line after its marker, and lengthens a code fence only past lines + that would close it. +- A complete-tag line right after a definition continues the paragraph the + definition was read from, a table header row indented four columns or more + starts no table, and a directive attribute without a valid name is + dropped, so every parsed document serializes. +- The serializer also writes these so that they reparse to the same tree: + text that would open a text directive, shortcode, raw HTML, or math with + an inline after it; a pipe that raw HTML, an autolink, or math writes in a + table cell, escaped with the cell; a list before a block indented one to + three columns, whose markers are indented past it; a paragraph after a + block quote in a tight item; a list item whose first line would make the + marker's line a thematic break; a code fence that a content line would + close, indented when that keeps the fence's length; raw HTML that opens a + paragraph or setext heading after a definition; a line that would open + description details; frontmatter ending in an empty line; and a literal `~` + beside an emphasis run, which stays literal unless it could pair. +- Every line carries the source column it starts at, so a tab spans the + columns it spans in the source line at any container depth: in the + indentation after a block quote's marker or a list item's content indent + (up to the four columns that decide indentation), after a list marker, and + in a fence's indentation. A line inside a fence or HTML block keeps its + tabs. +- A blank line inside a fence that a nested item left open loosens no list. +- A quoted paragraph that opens with backticks takes lazy lines; the GH-19 + rule applies to the lazy line, not to the quoted line before it. +- `src/parse/nul.rs` is named `nul_replacement.rs`, since `nul` is a reserved + Windows filename that `cargo package` warns about. +- Block structure reads only spaces and tabs as whitespace: a line holding a + no-break space, form feed, or other whitespace char is neither blank nor + indented, and such a char ends no thematic break, setext underline, fence, + ATX heading, HTML block start line, table row, flow MDX JSX indentation, or + paragraph. Label matching collapses only spaces, tabs, and line endings; a + footnote label holds no space, tab, line ending, or unescaped bracket; and an + angle-bracket autolink's URI may hold any other whitespace. +- A hard break comes only from spaces the source holds, a `` is text, a + list that could not interrupt a paragraph continues the one a definition + was read from, a directive opener inside fenced code in a container + directive is code, a footnote definition's first line keeps its trailing + spaces, and a GFM table starts below a delimiter row that is neither lazy + nor a setext underline, before any block its header row would otherwise + open. +- A paragraph or heading whose content the plain rendering does not read + back, and whose runs abut, meet a `~` or `*`, or meet a literal autolink's + edge, takes the first other delimiter choice that reads back under the + default preset, else under GFM or MDX: `__` strong, `_` inner or outer runs, + all-`*` runs, raw edge chars, and spaces around a literal autolink written + raw or as references. The check gives each reference label a stand-in + definition. +- The serializer also writes these so that they reparse to the same tree: + text before a construct whose delimiters it could pair with (a footnote's + `^`, a math fence, a literal autolink's or wiki link's chars, a `$` or `}` + in code, raw HTML, or MDX); a reference label's raw backtick, a bang + before a wiki link, whitespace control chars, labels, and alert titles as + their source; a bare text directive followed by what could go on with it; + email-local chars before an email; a paragraph that MDX would read as ESM + or flow; an empty container directive; and task item text that opens with + whitespace. +- Collecting definitions reads the block structure alone, once the input + holds `]:`, over the lines the main pass reads; inline content is scanned + for literal autolinks only when it holds `://`, `www.`, or `@`; and a table + delimiter row is checked by its chars before its cells are split. +- Serialization stays linear on nested emphasis, including when every + delimiter choice of a paragraph is tried, and on long `~` runs; parsing + stays linear on inline diagnostics that run to a paragraph's end. +- A quoted line short of the list item an open paragraph sits in, or of a + block quote inside that item, continues it only as a lazy line; a block + open in a list item ends at a line indented less than the item's content; + a lazy line in an item opens no fence; a fence a nested container + directive leaves open ends with it; a setext underline is no delimiter row + wherever a table start is checked; and the `|` read from a cell's `\|` + spans both of its bytes. +- The plain rendering keeps `**` after a `*`, and the serializer also writes + these so that they reparse to the same tree: runs inside links, images, + and marks; a strong at the edge of the run around it; a cell pipe after an + escaped backslash; text after a reference's raw label backtick, inside and + after spans; an `==` or `++` that could close its span, or a `+` or `=` + beside one; math opening a definition's paragraph; a `:` that a span's + delimiters would make a shortcode; whitespace opening a cell, or before a + text directive; a list in a list item before an indented block; and a + quote whose first line would read as an alert marker. + +Specs: +- `public-api` (modified) +- `inline-syntax` (modified) +- `serialization` (modified) +- `block-syntax` (modified) +- `validation` (modified) +- `untrusted-input-cost` (modified) + +## Design + +### Decisions + +- One source-map type owns the translation from a derived string to the + original input. Each derived string (container content, the inline input of a + paragraph, heading, table cell, directive label, description term, or HTML + container) is built together with its map. A `Line` takes its positions + through the map when it is created, and an inline parse translates its spans + once, after the pass. This is because the span construction sites in the + parser (about 96 `base_offset +` expressions) then stay as they are. + - Turned down: translating at each construction site, for the reason the BOM + and NUL change turned down a second coordinate system. + - Turned down: a dense per-byte offset table, which costs a machine word per + content byte at each of up to 32 container levels. + - Turned down: mapping only leading whitespace, the scope first proposed. It + leaves container continuation lines, CRLF, table cells, and out-of-bounds + code-block spans wrong. +- A map is always in original-input coordinates. A container composes its + parent's map as it copies the parent's lines. This is because a lookup then + costs the same at any depth. + - Turned down: chaining a child map to its parent's, which makes each lookup + walk the nesting. +- A map is a sequence of segments, each pairing a content range with the source + range it came from. + - Verbatim segments map byte for byte. + - A segment whose content replaces source maps every content byte to the + whole source range it replaces. This covers the spaces split from a tab, the + `\n` that joins lines ending in `\r\n`, and the `|` read from `\|`. + - A start position at a segment boundary maps through the segment that begins + there. An end position maps through the segment that ends there. + - This is because these rules give each scenario its span directly: a soft + break covers its whole line ending, and a node ending before a line break + ends on its own line. +- Inline spans are computed in inline-input coordinates and translated by one + walk over the finished subtree. + - In preorder, starts never decrease, so a forward cursor maps them. + - Each end is found by searching forward from its start's segment. + - The walk is linear because inline nesting is capped at 32 levels. + - Turned down: binary search per position, which is `n log n` against the + linear-time requirement. +- A table cell's span covers its trimmed content, the text its inline parse + reads. + - An empty cell gets the empty range just before the pipe that closes it. + - A cell missing from a short row gets the empty range at the row's end. + - This is because every parsed node must carry a span, and children then start + where their cell starts. + - Turned down: leaving cells without spans, which breaks "Source spans". + - Turned down: including padding and pipes, which makes adjacent cells share + a delimiter byte. +- "Spans nest" is checked over the whole tree for the fixture corpus and + generated inputs, because a missed construction site at any depth then fails a + test instead of waiting for a scenario. +- Each delimiter run keeps its original length beside its remaining length. + The rule of three and the `openers_bottom` key both use the original length, + because cmark and commonmark.js do both. The opener floor is sound only when + keyed on what the predicate reads, so the key uses the original length for `*` + and `_`, whose predicate is the rule of three, and the remaining length for + the other marks, whose predicates read it. + - Turned down: changing only the predicate, which leaves the floor keyed on a + length the predicate no longer reads. + - During planning, this change reduced mismatches against commonmark.js on + 20,476 generated `*`/text inputs from 304 to 0, with `cargo test` and + conformance unchanged. +- The image-only early return in `match_link_target` is deleted, so `![` and `[` + resolve what follows `]` with one sequence. The experiment during planning + left tests and conformance unchanged. + - Turned down: an image-specific fallback path. +- Every backtick in text is escaped, and the code-span prediction + `text_code_span_can_start` is removed. A backslash stops a backtick from + opening a code span but not from closing one, so a prediction must reason + about escaped backticks as closers. Dropping the prediction removes the + serializer's copy of the code-span rules. During planning, 5 fixture inputs + changed canonical output. + - Turned down: predicting per run, which keeps output byte-stable but keeps a + mirrored copy of parser rules. +- Inside an `_`-delimited emphasis, the serializer escapes a `_` in text that + can close, as text inside `*`-delimited emphasis already encodes its `*`, + because the reparse would close the emphasis there. Turned down: escaping + every `*` / `_` that can close, or every character of an escaped run, which + moved canonical output for about 1,000 corpus inputs while the defect needs a + delimiter outside the text node. +- A list marker on a lazy line ends the block quote whatever its content or + start number — because cmark, commonmark.js, markdown-it, and micromark all + read `> a\n- ` and `> a\n2. b` as a quote followed by a list: the empty-item + and start-number restrictions apply only when the paragraph is in the + container the list would open in. Turned down: applying those restrictions to + lazy lines too, which the spec's "would otherwise count as paragraph + continuation text" could be read to ask but no reference implementation does. +- Final whitespace is dropped from the derived inline input, with its map, after + the last line is read — because the inline parser then never sees it, and + hard line breaks on earlier lines keep their trailing spaces. Turned down: + trimming the last `Text` node after the inline parse, which leaves the spaces + to interact with the inline scan first. +- The serializer escapes a `|` that ends a level-two setext heading only when a + text node supplies it — because a pipe ending a literal autolink belongs to + its URL, and escaping it there changes the autolink. +- A complete HTML tag on a lazy line ends the container, in list items as in + block quotes — because cmark-gfm (GitHub) and micromark start a type-7 HTML + block there, and the micromark case `html_flow` 147 pins it. Turned down: + upstream cmark's and commonmark.js's reading, which keeps the tag in the + paragraph, against the GitHub rendering. +- Whether a paragraph is open is tracked with the number of nested block quotes + it sits in — because a line that reaches that level continues the paragraph + by the paragraph-interruption rules, while a line that stops short of it can + only continue it as a lazy line. Turned down: a yes/no flag, which reads + `> > a\n> 1. ` as continuation instead of a list in the outer quote. +- The GH-19 rule that a fence-like lazy line ends a block quote's paragraph is + kept for block quotes and not extended to list items — because extending it + moved 22 generated inputs away from cmark, commonmark.js, and micromark. +- After a shortcut reference, the serializer escapes the character that would + re-read its brackets, because the AST must round-trip. + - Turned down: writing the reference in collapsed form, which changes the + reparsed `ReferenceKind`. + +### Risks +- [Round-trip classes the serializer cannot reach without the parse's + context] → A paragraph whose `Strong` and `Emphasis` runs abut is + reparsed once under the default dialect, and other delimiter choices are + tried only when the first does not read back. Generated inputs still fail + to round-trip in five classes, and this plan adds a fix for none: + - a relaxed `://` literal autolink followed by a span whose content opens + with a char that only a named character reference stops the URL scan + before (`://~∓~`), which needs a reverse named reference table; + - a `www` literal autolink whose domain scan an escape ends early + (`www.\[]_(`), where escaping the `]` after it lets the scan reach a `_` + that rejects the domain; + - attention runs that abut inside another run (`***b_*_b_*`, + `_^*^*_c__`), whose delimiters must be chosen together, which no choice + the read-back tries does: an emphasis-heavy fuzz over 100,000 inputs + fails on 59; + - an `_` run after a letter that the default preset's strikethrough bonus + opens beside a `~` (`t_~>___~`), a reading CommonMark and GFM do not + share; + - a container directive opener inside an HTML block, which the directive + counts as nested (`:::e\n\n:::e`). + +- [Findings accepted at archive] → The re-verify of group 19 found that its + quote and list column tracking (`OpenParagraphIn`, `OpenBlock.item_column`) + fixes 378 generated nested-container inputs and newly breaks 13, a panic in + `prefix_ends_with_gfm_email` on `"\u{a0}e+@"` that predates this plan, and + gaps in the growth and corpus tests. They were accepted for archive and are + tracked, with the residual classes above, in plimeor/markdown-syntax#11, + which replaces the container re-prediction and the serializer's copies of + parser rules at their root. + +- [Tabs after a split tab] → Once a container marker splits a tab, the later + tabs that open the line are expanded to spaces, inside a fence or HTML + block too: `>\t\t\tfoo` gives a code block of six spaces and `foo` where + cmark, commonmark.js, and micromark keep ` \tfoo`. The behavior predates + this plan, whose comparisons generated no line with three tabs after a + split, and is left to a later one. + +- [A derived string built without its map leaves a shifted span] → The "Spans + nest" check runs over the corpus and over generated inputs that mix block + quotes, list items, tables, CRLF, and tabs. Each container kind has a + scenario or test. +- [Spaces split from a tab are assumed to occur only at a line's front, so a + nested container that strips more could meet them elsewhere] → The generated + inputs mix tabs with nested container markers. A counterexample changes the + segment representation, not the requirements. +- [Map lookups break linear time] → A growth check over long, nested block + quotes and list items and over long tables joins `tests/pathological_inputs.rs` + and the growth sweep. +- [The rule-of-three change moves emphasis output beyond the targeted inputs] → + Parse output is compared with the plan's starting commit. Every diff must + involve a `*`, `_`, or underline `__` run that pairs more than once. +- [Downstream users rely on bare backticks in canonical output] → The commit + states the output change, and the "Literal backticks are always escaped" + scenarios pin it. + +## Tasks + +### 1. Rule of three +- [x] 1.1 Add tests for "Rule of three counts whole delimiter runs" and for `***a*a*a`; verified by both failing on the current code. +- [x] 1.2 Keep each delimiter run's original length, and use it in `emphasis_delimiters_match` and in `openers_bottom_key` for `*` / `_`; verified by: + - the 1.1 tests and `cargo test` passing + - conformance not below 2233/2236 + - a parse-output comparison with the starting commit, over the fixture corpus and seeded generated inputs, differing only in the class listed under Risks +- [x] 1.3 Escape a `_` in text that can close when the text sits inside an `_`-delimited emphasis, because the new parse produces that shape for inputs that round-tripped before; verified by `an_underscore_that_can_close_stays_inside_underscore_emphasis` (failing before) and the parse-output comparison showing no corpus input whose canonical output moves + +### 2. Image fallback and text after shortcut references +- [x] 2.1 Add tests for "Image whose resource is invalid" and the three shortcut-reference scenarios of "Escaping keeps text literal"; verified by each failing on the current code. +- [x] 2.2 Delete the image-only early return in `match_link_target` and correct its doc comment; verified by the image test passing and conformance unchanged. +- [x] 2.3 Escape a `(` directly after a shortcut `LinkReference` or `ImageReference`, and a `:` directly after one that starts a paragraph; verified by the shortcut-reference tests and the round-trip fixtures passing. + +### 3. Backticks +- [x] 3.1 Add tests for both scenarios of "Literal backticks are always escaped"; verified by the first failing on the current code. +- [x] 3.2 Escape every backtick in text, and remove `text_code_span_can_start` and any scan state only it used; `ordinary_punctuation_text_does_not_reparse_as_character_escapes` drops the backtick from its sample of punctuation left unescaped; verified by: + - the 3.1 tests and `cargo test` passing + - every moved `.canonical.md` golden regenerated and read for correctness + - every canonical diff from the starting commit being a backtick in text + +### 4. Source spans +- [x] 4.1 Add tests: + - every scenario of "Spans map stripped lines back to the source", the new "Source spans" scenarios, and "Emphasis on a block quote continuation line" + - issue plimeor/markdown-syntax#6's regression test, as `inline_spans_address_source_inside_containers` in `tests/parse_span_contract.rs`, with its 31 cases unchanged + - a "Spans nest" check over the fixture corpus and over seeded generated inputs mixing containers, tables, CRLF, and tabs + + Verified by the changed-behavior tests failing on the current code, the issue's test failing 26 of 31 cases. +- [x] 4.2 Add the source-map type and build every container's content with its map, composed in original coordinates. The containers are block quotes and alerts, list items, footnote definitions, container directives, description details, and HTML containers. `Line` positions come from the map. A container whose first content line is empty keeps that line, which the old string join dropped; this realigns lazy-line flags, so `> \n> a\n- ` parses as `> a\n- ` does. Verified by: + - the "Later block inside a block quote" and "Split tab" tests passing + - `tests/parse_span_contract.rs` passing + - the nest check finding no block span outside its parent +- [x] 4.3 Build every inline input with its map: paragraphs, ATX and setext headings, table cells, directive labels, description terms, and HTML containers. Translate spans in one walk per block-level inline parse. Verified by the inline scenarios, all 31 cases of `inline_spans_address_source_inside_containers`, and the nest check passing at every depth. +- [x] 4.4 Give table cells their spans, as described under Decisions; verified by the table-cell scenarios passing. +- [x] 4.5 Extend `inline_container_spans_cover_their_delimiters_and_content` to inputs with leading whitespace, block quotes, list items, tables, and CRLF; verified by it passing. +- [x] 4.6 Add a growth check for long, nested block quotes and list items and for long tables to `tests/pathological_inputs.rs`; verified by linear growth in debug and release builds. + +### 5. Lazy list markers and final whitespace +- [x] 5.1 Add tests for "Lazy list marker ends a block quote", "Lazy ordered item not starting at 1", "Final whitespace of a paragraph", and "Final whitespace of a setext heading", and for `> - a\n- `; verified by each failing on the group-4 code. +- [x] 5.2 Let any list marker end a lazy line in `lazy_line_starts_block`; verified by the lazy-list tests and by a comparison with commonmark.js on 30,000 generated block-structure inputs in which no input that matched before stops matching. +- [x] 5.3 Trim the final whitespace of paragraph and setext heading content before the inline parse, regenerating and reading the moved goldens; verified by the final-whitespace tests and `tests/fixtures.rs` passing. +- [x] 5.4 Escape a text `|` that ends a level-two setext heading; verified by `a_pipe_ending_a_level_two_setext_heading_does_not_start_a_table` (failing before) and a parse-output comparison with the group-4 end in which no input newly fails to round-trip and every tree diff comes from 5.2 or 5.3. + +### 6. Container laziness +- [x] 6.1 Add tests for "Lazy line in an item that started blank", "Blank line indented four columns ends a block quote", "Blank line between an empty item and the next", and "Complete HTML tag on a lazy line ends a list item", and for lazy lines after a thematic break, through nested containers, and short of the paragraph's quote level; verified by each failing on the group-5 code, apart from the quote-level guard. +- [x] 6.2 End a block quote at any blank line; track an item's open paragraph afresh after a blank line or a single-line block; track the open paragraph's quote level; carry a lazy flag into nested lists; loosen a list at a blank line between items; and read a list item's lazy lines with the block quote's rule, GH-19 aside; verified by a comparison with cmark/commonmark.js and micromark, where they agree, on 43,805 generated block-structure inputs without tabs or fences, with no mismatch left and no input that matched before stopping matching, and by a parse-output comparison with the group-5 end in which no input newly fails to round-trip. + +### 7. Integration checks +- [x] 7.1 `cargo fmt --check`, `cargo test` with and without `html`, `RUSTDOCFLAGS='-D warnings' cargo doc --no-deps`, `cargo build --target wasm32-unknown-unknown`, and `cargo +1.82 build` all pass. +- [x] 7.2 `tests/pathological_inputs.rs`, the growth sweep, and the 2 MiB stack check pass in debug and release builds. +- [x] 7.3 The conformance bench result is measured and reported with the change, along with a benign-document benchmark against the starting commit. + +### 8. Paragraph interruption, verbatim blocks in block quotes, and round-trips +- [x] 8.1 Add tests for "ATX-like line that is no heading", "Backtick run with a backtick in its info", "Unclosed fence in a block quote", and "Last line ending of indented code", and for each serializer case under "Escaping keeps text literal" that this group adds, with guards that a closed fence and a fence inside a quoted list item leave lazy lines as they were; verified by each failing on the group-6 code, apart from the two guards. +- [x] 8.2 Read `#` and fence lines in `likely_block_start` with the ATX and fence opener rules; track a block quote's open fence, math block, and HTML block in `content_line_state`, which list items share; end a code or math block's last line with the value's first line ending; regenerate and read the `commonmark_code_spans`, `commonmark_blockquotes`, and `commonmark_tabs` goldens; verified by the comparison with cmark/commonmark.js and micromark on the 95,474 inputs where they agree (792 mismatches before, 751 after, none newly mismatching; of the rest, 744 involve tabs, 6 the GH-19 rule, and 1 a blank line inside an unclosed fence of a nested item) and by conformance staying at 2233 of 2236. +- [x] 8.3 Write the serializer cases above; verified by the round-trip fuzz over 30,000 generated inputs dropping from 812 failures at the group-6 end to 372, and by `tests/fixtures.rs` passing. + +### 9. Inline flanking, delimiter runs, and further round-trips +- [x] 9.1 Add tests for "Underscore after Unicode punctuation", "Escaped backslash before a line ending", "Space inside a bare destination's parentheses", "Container content ending in a carriage return", "Malformed directive line inside a paragraph", and each serializer scenario this group adds; verified by each failing on the group-8 code. +- [x] 9.2 Read the `_` rules' punctuation as Unicode punctuation; take a hard break only after an unescaped backslash; end a bare destination at any space; require a valid opener for a directive line to interrupt a paragraph; end a container's last line with `\n`; verified by the tests, by conformance staying at 2233 of 2236, and by the comparison with cmark/commonmark.js and micromark on the 95,474 inputs where they agree with no input newly mismatching. +- [x] 9.3 Write the serializer cases above, regenerating the `commonmark_attention` canonical output, whose `\__foo\__bar` becomes `\_\_foo\_\_bar`; verified by the round-trip fuzz over 30,000 generated inputs dropping from 372 failures at the group-8 end to none. + +### 10. Extension constructs, cells, list edges, and definitions +- [x] 10.1 Add tests for "Complete tag after a definition", "Indented table header row", and "Directive attribute without a valid name", and for each serializer scenario this group adds; verified by each failing on the group-9 code. +- [x] 10.2 Skip type-7 HTML blocks right after a definition; reject table header rows indented four columns or more; drop directive attributes whose names do not validate; and skip HTML blocks of types 1–5 when looking for an unclosed fence in a closed container; verified by the tests, by conformance staying at 2233 of 2236, and by the comparison with cmark/commonmark.js and micromark on the 95,474 inputs where they agree with no input newly mismatching. +- [x] 10.3 Write the serializer cases above, regenerating the `commonmark_character_escapes` and `math_edges` canonical outputs (a final `~` stays literal; a `$` that a later line's `$` could close is escaped) and updating two serializer regression tests whose expected output changed for the better (a thematic-break item keeps its marker; dollar math in a table cell keeps its kind); verified by the round-trip fuzz over 200,000 generated inputs on four seeds dropping from 125–168 failures to 20–27, all in the classes listed under Risks. +- [x] 10.4 Keep the added checks off the common path: a byte prefilter before the continuation-line and referenced-char scans, a lookup for delimiter chars, sibling char sets built only for runs that can need them, a cell pipe prefilter, and preallocated text and fence buffers; verified by the serializer's instruction count on a 400 KB document of the fixtures falling from 179M to 138M (114M at 0.3.0), parsing at 10.7% more instructions than 0.3.0 (the source maps), and linear growth of serialization over emphasis-, line-, quote-, and cell-heavy inputs up to 80,000 repetitions. + +### 11. Top-level tab columns +- [x] 11.1 Add a test for "Tab after a top-level block quote marker", with guards for a list item's continuation, a tab past an indented code block's indentation, and a tab inside a fence; verified by the quote case failing on the group-10 code. +- [x] 11.2 Expand the leading tabs of a top-level block quote's or list item's content line at their source columns, up to four columns and not inside an open fence or HTML block; verified by conformance staying at 2233 of 2236, by the comparison with cmark/commonmark.js and micromark on the 95,474 inputs where they agree going from 751 mismatches to 289 with none newly mismatching, and by fourteen hand-written tab layouts (tab-indented code in lists and quotes, tab-separated markers, nested tab lists) all matching commonmark.js. +- [x] 11.3 Give each line the source column it starts at, recorded by `DerivedText` for derived lines, and read tabs from it in list markers, list continuations, block quote markers, fence indentation, and the container line classification; verified by `a_tab_inside_nested_containers_spans_its_source_columns` (failing before), by the block comparison going from 286 mismatches to 8 (all under the GH-19 rule) with none newly mismatching, by the hand-written tab layouts all matching, and by conformance staying at 2233 of 2236. + +### 12. Inline comparison +- [x] 12.1 Add a test for "Quoted paragraph opening with backticks"; verified by it failing on the group-11 code. +- [x] 12.2 Read a quoted content line's block without the lazy-line GH-19 rule; verified by a comparison with commonmark.js and micromark, with raw HTML allowed, on 30,000 generated inline inputs (links, references, images, code spans, entities, escapes, autolinks, emphasis, raw HTML) where they agree on 29,971: 5 mismatches before and 3 after, one the GH-19 rule the conformance oracle (markdown-rs) keeps and two the link-text autolink demotion the 0.3.0 delimiter stack chose; and by the block comparison staying with no input newly mismatching. +- [x] 12.3 Leave a list tight across a blank line that a fence opened in a nested item holds; verified by `a_blank_line_inside_a_nested_items_open_fence_leaves_the_list_tight` (failing before) and by the block comparison's last non-tab, non-GH-19 mismatch matching, with none newly mismatching (286 left: 280 tab cases in nested containers and 6 under the GH-19 rule). + +### 13. Abutting attention runs and literal autolink neighbours +- [x] 13.1 Add tests for "Abutting attention runs" and "Text beside a literal autolink"; verified by `nested_attention_runs_pick_delimiters_that_reparse_to_them` and `text_around_a_literal_autolink_does_not_move_its_end` failing on the group-12 code. +- [x] 13.2 Reparse a paragraph whose `Strong` and `Emphasis` runs abut and, when it does not read back, take the first of `__` strong, `_` inner emphasis, and all-`*` delimiters that does; carry a literal autolink's end through the span delimiters written after it; and write a scheme char ending a text before a `://` literal autolink as an escape or a character reference; verified by the tests, by the round-trip fuzz over 200,000 generated inputs on four seeds dropping from 20–27 failures to 0–15, all in the classes listed under Risks, and by serialization of the 4 MB fixture benchmark taking 5% longer (74 ms to 78 ms median). +- [x] 13.3 Add tests for "Bracket inside a footnote label", "Backtick after a reference's raw label", and "Bang before a wiki link"; verified by `a_footnote_label_holds_no_unescaped_bracket` and `text_after_a_reference_or_before_a_wikilink_keeps_its_parse` failing on the 13.2 code. +- [x] 13.4 Reject a footnote label with an unescaped bracket; write a text backtick after a reference whose raw label holds a backtick as ```; escape a `!` before a wiki link; verified by the tests and by the round-trip fuzz over 200,000 generated inputs on four seeds dropping from 0–15 failures to 0–10. +- [x] 13.5 Add a test for "Tilde beside an attention run"; verified by `a_tilde_beside_an_attention_run_keeps_the_runs_bonus` failing on the 13.4 code. +- [x] 13.6 Read a text `*` or `_` touching a `~` as able to open or close, as the parser's strikethrough bonus does, and add raw edge tildes to the delimiter choices a paragraph that does not read back tries; verified by the test and by the round-trip fuzz over 200,000 generated inputs on four seeds dropping from 0–10 failures to 0–9, with no `~` case left outside the literal autolink class. +- [x] 13.7 Extend `nested_attention_runs_pick_delimiters_that_reparse_to_them` with `__***/***__`, `**#****]***_**`, and `***_\**#*`; verified by each failing on the 13.6 code. +- [x] 13.8 Add `_` for the outermost run only, and raw edge `*` text joining a run, to the delimiter choices a paragraph that does not read back tries; verified by the test and by the round-trip fuzz over 200,000 generated inputs on four seeds dropping from 0–9 failures to 0–7, none of them a strong or emphasis run without a literal autolink or a NUL char. +- [x] 13.9 Add tests for "Delimiters a following construct writes" and "Run of bars before a spoiler", and for a strong inside a strong and an emphasis touching a `~` inside an emphasis; verified by `text_before_a_construct_escapes_the_delimiters_it_writes` failing on the 13.8 code. +- [x] 13.10 Count a math span's `$`, a footnote's `^`, and a literal autolink's URL chars among the delimiters an inline writes; guard any relaxed scheme, not only `://`, against a scheme char before it; write every bar but the last of a run that can open a spoiler as a reference; and try every delimiter choice with each raw edge char, including for a strong inside a strong; verified by the tests and by the round-trip fuzz over 200,000 generated inputs on eight seeds failing on 0–7 inputs, all a relaxed `://` literal autolink beside a span or a table. +- [x] 13.11 Read what sits beside a paragraph's runs in one walk, find a literal autolink's end in the output of the inline before a text instead of searching the output, and scan inline content for literal autolinks only when it holds `://`, `www.`, or `@`; verified by the AST of 155,000 corpus inputs under the default and GFM presets staying byte-identical, and by instruction counts on the 400 KB fixture document: parsing at 229M against 261M before, and serialization at 48M against 46M at the group-12 end. + +### 14. Definition pass, table rows, and block whitespace +- [x] 14.1 Collect definitions from the block structure alone, leaving inline content unparsed, and only when the input holds `]:`; reject a table delimiter row with a char other than `|`, `-`, `:`, space, or tab before splitting it; verified by the AST of the 155,000 corpus inputs under the default and GFM presets staying byte-identical and by parsing the 400 KB fixture document at 154M instructions against 229M after 13.11. +- [x] 14.2 Add tests for "Other whitespace after a thematic break", "Line holding only a no-break space", "Form feed ending a paragraph", "No-break space in a label", "Whitespace control chars in text", and "Hard break opening a span"; verified by `a_line_with_other_whitespace_is_text`, `a_form_feed_ending_a_paragraph_stays_its_text`, and `labels_collapse_only_spaces_tabs_and_line_endings` failing on the group-13 code, and by the round-trip fuzz with no-break space, form feed, and ideographic space added to its pieces failing on over 2,000 of 200,000 inputs per seed before 14.4. +- [x] 14.3 Read only spaces and tabs as whitespace for blank lines, indentation, thematic breaks, setext underlines, fence info and closing lines, ATX content, HTML block start lines, footnote and alert lines, definition titles, table rows and cells, and the final whitespace of a paragraph; collapse only spaces, tabs, and line endings in label matching; and allow any other whitespace in a footnote label; verified by the tests, by 388 CommonMark and 54 GFM probes of up to six whitespace chars across block constructs, where all 356 structures commonmark.js, markdown-it, and micromark agree on match and the 32 disputed ones match micromark (54 of 54 GFM match micromark with its GFM extensions), by the 155,000 corpus inputs staying byte-identical, and by conformance staying at 2233 of 2236. +- [x] 14.4 Write a line tabulation, form feed, or next-line char in text as itself, a reference or footnote label's control chars as its source, and a hard break opening a span as ` `; validate a definition's identifier as blank only when it holds spaces, tabs, and line endings alone; count a wiki link's text and a `$` in code or raw HTML among the delimiters an inline writes; escape a `:` ending a text before a span whose delimiters could name a shortcode; and let a strong's content end with an escaped `_` under `__`; verified by the tests and by the round-trip fuzz over 200,000 generated inputs on eight seeds failing on 0–7 inputs, all a relaxed `://` literal autolink beside a span. +- [x] 14.5 Accept any char but a space, an ASCII control char, `<`, and `>` in an angle-bracket autolink's URI, as the parser and as validation; verified by `an_angle_autolink_holds_whitespace_other_than_a_space` (failing before) and by 108 inline probes of four whitespace chars, where all 84 results commonmark.js and micromark agree on match and the 24 disputed ones match micromark. + +### 15. Literal autolink neighbours and the reparse check +- [x] 15.1 Add tests for "Space between a literal autolink and a span delimiter" and "Text between a literal autolink and a span"; verified by `spaces_and_text_beside_a_literal_autolink_keep_its_end` failing on the group-14 code. +- [x] 15.2 Add raw line-edge spaces, a referenced space before a literal autolink, and an escaped first char after one to the choices a paragraph that does not read back tries, gated on a literal autolink meeting a space at a span edge or followed by a whitespace-free text and another inline; give the reparse check a stand-in definition for each reference label; and compare span-free trees instead of their debug text; verified by the tests and by the round-trip fuzz over 200,000 generated inputs on eight seeds failing on 0–1 inputs, each a relaxed `://` literal autolink followed by a span whose content opens with a char only a named reference such as `∓` writes, and by serialization of the 400 KB fixture document at 50M instructions against 46M at the group-12 end. + +### 16. Block-oriented round trips across presets +- [x] 16.1 Add a round-trip fuzz over block-oriented pieces (list, quote, heading, fence, directive, footnote, table, alert, HTML, MDX, and autolink pieces) under the CommonMark, default, GFM, and MDX presets, and add tests for "Directive opener inside fenced code", "Hard break on a footnote definition's first line", "Text after a bare text directive", "Email-local char before an email", "Alert title and empty container directive", and "Paragraph opening with an ESM keyword"; verified by `a_directive_opener_inside_fenced_code_is_code`, `a_footnote_definitions_first_line_keeps_a_hard_break`, and `constructs_after_a_directive_email_or_alert_keep_their_parse` failing on the group-15 code, and by the fuzz failing on 282–325 of 100,000 inputs per seed before 16.2 and 16.3. +- [x] 16.2 Skip directive openers inside a fenced code block in a container directive, and keep the trailing spaces of a footnote definition's first line; verified by the tests, by the 155,000 corpus inputs under the default and GFM presets staying byte-identical, and by conformance staying at 2233 of 2236. +- [x] 16.3 End a text directive with no attributes by an empty label, or an empty attribute list before a `{`, when what follows could go on with it; escape an email-local char or `@` before an email or a literal autolink and a `+` run whole; write an alert title as its source, an empty container directive without a blank line, and a paragraph's leading `import ` or `export ` with a referenced first char; read a relaxed scheme only from a run holding a letter; and try the literal autolink choices for a text at a span's end; verified by the tests, by the autolinks canonical output regenerated for `\`, which reparses to the same text under both presets, and by the block-oriented fuzz failing on 0–4 of 100,000 inputs per seed. + +### 17. Read-back across presets, headings, and MDX flow +- [x] 17.1 Add tests for "No-break space before MDX JSX", "Content that reads back only under its preset", "Heading content that does not read back", "Flow-like first line under MDX", and "Task item text opening with whitespace"; verified by `a_paragraph_opening_with_an_esm_keyword_stays_a_paragraph_under_mdx`, `mdx_and_gfm_content_reads_back_under_its_preset`, and `delimiters_around_autolinks_tasks_and_alerts_keep_their_parse` failing on the group-16 code. +- [x] 17.2 Read a flow MDX JSX line's indentation as spaces and tabs; give headings the paragraph's delimiter choices; take a choice that reads back under the default preset first, then one under GFM or MDX; count a `}` among the delimiters an inline writes; keep a paragraph's or setext heading's first line off MDX ESM and flow, and an angle-bracket autolink at its start unescaped; separate a paragraph after an alert in a tight item; write whitespace opening a task item's text raw; and try run choices for a `www` autolink after a text `*`, `_`, or `~` and the literal autolink choices for a link's or inline footnote's text; verified by the tests and by the block-oriented fuzz over 100,000 inputs on ten seeds and the inline fuzz over 200,000 inputs on eight seeds failing on 0–2 inputs per seed, each `www.\[]_(` or a relaxed `://` literal autolink before a span opening with a char only a named reference stops the URL scan before; serialization of the 400 KB fixture document stays at 50M instructions. +- [x] 17.3 Split the input into lines once for the definition pass and the main pass; verified by the 155,000 corpus inputs under the default and GFM presets staying byte-identical and by parsing the 400 KB fixture document at 137M instructions against 145M. + +### 18. Mixed-construct comparison +- [x] 18.1 Add tests for "`` is text", "Referenced space before a line ending", and "Ordered item not starting at 1 after a definition"; verified by `a_processing_instruction_closes_after_its_opener`, `a_referenced_space_makes_no_hard_break`, and `a_list_that_cannot_interrupt_a_paragraph_continues_a_definitions_paragraph` failing on the group-17 code. +- [x] 18.2 Close a processing instruction only at a `?>` after its ` a\n- "` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` with a paragraph `a`, followed by a `List` holding one empty item + +#### Scenario: Lazy ordered item not starting at 1 +- **WHEN** `"> > a\n2. b"` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` followed by an ordered `List` starting at 2 + +#### Scenario: Final whitespace of a paragraph +- **WHEN** `"aaa \nbbb "` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("aaa")`, a `LineBreak`, and `Text("bbb")` + +#### Scenario: Final whitespace of a setext heading +- **WHEN** `"Foo \n-----"` is parsed with the CommonMark preset +- **THEN** the document holds a level-2 setext `Heading` holding `Text("Foo")` + +#### Scenario: Lazy line in an item that started blank +- **WHEN** `"- \n a\nb"` is parsed with the CommonMark preset +- **THEN** the list's one item holds a paragraph `a`, a soft break, and `b` + +#### Scenario: Blank line indented four columns ends a block quote +- **WHEN** `"> a\n \n> b"` is parsed with the CommonMark preset +- **THEN** the document holds two `BlockQuote`s + +#### Scenario: Blank line between an empty item and the next +- **WHEN** `"* \n\n * b c"` is parsed with the CommonMark preset +- **THEN** the document holds one loose `List` of two items + +#### Scenario: Complete HTML tag on a lazy line ends a list item +- **WHEN** `"- a\n"` is parsed with the CommonMark preset +- **THEN** the document holds a `List` followed by an `HtmlBlock` + +#### Scenario: ATX-like line that is no heading +- **WHEN** `"a\n#)"` is parsed with the CommonMark preset +- **THEN** the document holds one `Paragraph` holding `Text("a")`, a `SoftBreak`, and `Text("#)")` + +#### Scenario: Backtick run with a backtick in its info +- **WHEN** ````"a\n``` `` ```"```` is parsed with the CommonMark preset +- **THEN** the document holds one `Paragraph` whose second line is a code span + +#### Scenario: Unclosed fence in a block quote +- **WHEN** ``"> ```\n> x\na"`` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` whose fenced `CodeBlock` holds `"x\n"`, followed by a `Paragraph` holding `Text("a")` + +#### Scenario: Last line ending of indented code +- **WHEN** `"\ta\r\tb"` is parsed with the CommonMark preset +- **THEN** the document holds an indented `CodeBlock` whose value is `"a\rb\r"` + +#### Scenario: Container content ending in a carriage return +- **WHEN** ``"- ```\n ~\r"`` is parsed with the CommonMark preset +- **THEN** the list item's fenced `CodeBlock` holds `"~\n"` + +#### Scenario: Complete tag after a definition +- **WHEN** `"[o]: u\n"` is parsed with the CommonMark preset +- **THEN** the document holds a `Definition` followed by a `Paragraph` holding an `Html` inline + +#### Scenario: Indented table header row +- **WHEN** `"a\n |b\n----"` is parsed with `parse` +- **THEN** the document holds one setext `Heading` + +#### Scenario: Tab after a top-level block quote marker +- **WHEN** `"> \tcode"` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` holding a `Paragraph`, since the tab spans columns 2 to 4 + +#### Scenario: Quoted paragraph opening with backticks +- **WHEN** `"> ``a\nb"` is parsed with the CommonMark preset +- **THEN** the document holds one `BlockQuote` whose paragraph ends with `Text("b")` + +#### Scenario: Blank line inside a nested item's open fence +- **WHEN** ``"2. a\n 1. ```\n\n2. b"`` is parsed with the CommonMark preset +- **THEN** the document holds one tight `List` + +#### Scenario: Tab inside nested containers +- **WHEN** `"* - \tb c"` is parsed with the CommonMark preset +- **THEN** the inner list item holds an indented `CodeBlock`, since the tab spans columns 4 to 8 + +#### Scenario: Lazy list marker inside a quoted item +- **WHEN** `"> - a\n> 2.\nz"` is parsed with the CommonMark preset +- **THEN** the document holds a `BlockQuote` holding two lists, followed by a `Paragraph` holding `Text("z")` + +#### Scenario: Setext-like line short of a quoted item +- **WHEN** `"> 1. a\n> ===\nb"` is parsed with the CommonMark preset +- **THEN** the item holds one `Paragraph` holding `Text("a")`, `Text("===")`, and `Text("b")` with soft breaks between + +#### Scenario: Sibling item ending a nested fence +- **WHEN** ``"- - ```\n - a\n\n- b"`` is parsed with the CommonMark preset +- **THEN** the outer `List` is loose + +#### Scenario: Lazy fence-like line in an item +- **WHEN** ``"1. a\n ```\n\nb"`` is parsed with the CommonMark preset +- **THEN** the document holds a `List` followed by a `Paragraph` holding `Text("b")` + +#### Scenario: CommonMark oracle cases +- **WHEN** the block cases under `tests/fixtures/conformance/commonmark/` are parsed and rendered with the `html` feature +- **THEN** the output matches the expected HTML + +### Requirement: Block directives +When directives are enabled, the parser SHALL recognize `::name[label]{attrs}` +leaf directives and `:::name` container directives, whose container closes at a +fence of at least the opening fence's length, and a line with a malformed +opener SHALL NOT end a paragraph. + +#### Scenario: Container directive +- **WHEN** `":::note\nbody\n:::"` is parsed with `parse` +- **THEN** the document holds a `ContainerDirective` named `note` whose children hold a paragraph `body` + +#### Scenario: Unclosed container +- **WHEN** `":::note\nunclosed container"` is parsed +- **THEN** a `ContainerDirective` holds the remaining content and an error-severity `UnclosedDirectiveContainer` diagnostic is reported + +#### Scenario: Invalid name +- **WHEN** a leaf directive opener has a malformed name +- **THEN** an error-severity `InvalidDirectiveName` diagnostic is reported + +#### Scenario: Directive attribute without a valid name +- **WHEN** `":b{<} :c{a <=1 d}"` is parsed with `parse` and serialized +- **THEN** `to_markdown()` returns `":b :c{a d}\n"` + +#### Scenario: Malformed directive line inside a paragraph +- **WHEN** `"a\n::1bad"` or `"a\n:::"` is parsed with `parse` +- **THEN** the document holds one `Paragraph` holding both lines + +## ADDED Requirements + +### Requirement: Spaces and tabs are block whitespace +The parser SHALL read only spaces and tabs as whitespace in block structure: a +blank line holds only spaces and tabs; indentation, the space after a block +marker, the trailing whitespace a thematic break, setext underline, closing +fence, ATX closing sequence, HTML block start line, definition, or table row +allows, the indentation before flow MDX JSX, and the final whitespace of a +paragraph are spaces and tabs. Any +other whitespace char, such as a no-break space or a form feed, is content. + +#### Scenario: Other whitespace after a thematic break +- **WHEN** `"***\u{a0}"` is parsed with the CommonMark preset +- **THEN** the document holds a `Paragraph`, not a `ThematicBreak` + +#### Scenario: Line holding only a no-break space +- **WHEN** `"a\n\u{a0}\nb"` is parsed with the CommonMark preset +- **THEN** the document holds one `Paragraph` + +#### Scenario: No-break space before MDX JSX +- **WHEN** `"\u{a0}

"` is parsed with the MDX preset +- **THEN** the document holds a `Paragraph`, not a flow `MdxJsx` block + +#### Scenario: Form feed ending a paragraph +- **WHEN** `"a\u{c}"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("a\u{c}")` + +### Requirement: Fenced code inside a container directive +A fenced code block inside a container directive SHALL hold its lines as code: +a line in it that looks like a directive opener opens no nested directive, +while a closing fence of the directive still closes it, and a fence that a +nested directive leaves open ends with that directive. + +#### Scenario: Directive opener inside fenced code +- **WHEN** `":::t\n```\n:::e\n```\n:::"` is parsed +- **THEN** the document holds one `ContainerDirective` named `t` holding a `CodeBlock` whose value is `":::e\n"` + +#### Scenario: Fence left open in a nested directive +- **WHEN** ``":::outer\n:::inner\n```\n:::\n:::inner2\nx\n:::\n:::\nafter"`` is parsed +- **THEN** the `ContainerDirective` named `outer` holds the directives `inner` and `inner2`, a `Paragraph` holding `Text("after")` follows it, and no diagnostic is reported + +### Requirement: Footnote definition content +A footnote definition's content SHALL start after the spaces and tabs that +follow its `]:` and keep the trailing spaces of its first line, which may +make a hard break. + +#### Scenario: Hard break on a footnote definition's first line +- **WHEN** `"[^1]: a \nb"` is parsed +- **THEN** the definition's paragraph holds `Text("a")`, a `LineBreak`, and `Text("b")` + +### Requirement: List after a definition +A list SHALL start on the line right after a definition only when its first +item could interrupt a paragraph: a bullet or an ordered item starting at 1, +with content. Otherwise the line continues the paragraph the definition was +read from. + +#### Scenario: Ordered item not starting at 1 after a definition +- **WHEN** `"[foo]: /url\n2) a"` is parsed with the CommonMark preset +- **THEN** the document holds the `Definition` and a `Paragraph` + +### Requirement: GFM table start +A GFM table SHALL start only where the line after its header row is a +delimiter row that is neither a lazy continuation line nor a setext +underline; such a line keeps its other reading. + +#### Scenario: Delimiter row without pipes +- **WHEN** `"| --- |\n-- "` is parsed with the GFM preset +- **THEN** the document holds a level-2 setext `Heading`, not a `Table` + +#### Scenario: Setext underline below a header row +- **WHEN** `"a\n|b\n---"` is parsed with the GFM preset +- **THEN** the document holds one level-2 setext `Heading` holding both lines + +#### Scenario: Lazy delimiter row +- **WHEN** `"1. ---(\n:-:"` is parsed with the GFM preset +- **THEN** the list item holds a `Paragraph`, not a `Table` + +#### Scenario: Table header row that looks like an empty list item +- **WHEN** `"a\n+\n|-"` is parsed with the GFM preset +- **THEN** the document holds a `Paragraph` and a `Table` whose header cell holds `+` diff --git a/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md new file mode 100644 index 0000000..e7af140 --- /dev/null +++ b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/inline-syntax.md @@ -0,0 +1,92 @@ +# Inline syntax — spec changes + +## MODIFIED Requirements + +### Requirement: CommonMark inlines +The parser SHALL recognize CommonMark inline constructs (backslash escapes, +entity and numeric character references, code spans, emphasis and strong +emphasis, links, images, autolinks, raw HTML, and hard and soft line breaks) as +the CommonMark specification defines them, including its precedence of code +spans, links, and emphasis. + +#### Scenario: Emphasis +- **WHEN** `"Hello *world*."` is parsed +- **THEN** the paragraph holds `Text("Hello ")`, an `Emphasis` containing `world`, and `Text(".")` + +#### Scenario: Link inside a link label +- **WHEN** `"[foo [bar](/u)](/v)"` is parsed +- **THEN** only `[bar](/u)` becomes a link and the surrounding brackets and `(/v)` stay text + +#### Scenario: Shortcut reference before an unclosed label +- **WHEN** `"[foo][bar\n\n[foo]: /u"` is parsed with the CommonMark preset +- **THEN** the paragraph holds a shortcut `LinkReference` to `foo` followed by `Text("[bar")` + +#### Scenario: Rule of three counts whole delimiter runs +- **WHEN** `"*a***b*"` is parsed with the CommonMark preset +- **THEN** the paragraph holds an `Emphasis` containing `a`, `Text("*")`, and an `Emphasis` containing `b` + +#### Scenario: Image whose resource is invalid +- **WHEN** `"![foo](a b)\n\n[foo]: /u"` is parsed with the CommonMark preset +- **THEN** the paragraph holds a shortcut `ImageReference` to `foo` followed by `Text("(a b)")` + +#### Scenario: Underscore after Unicode punctuation +- **WHEN** `"«_**]**_"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("«")` and an `Emphasis` containing a `Strong` containing `]` + +#### Scenario: Escaped backslash before a line ending +- **WHEN** `"a\\\\\nb"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("a\\")`, a `SoftBreak`, and `Text("b")` + +#### Scenario: Space inside a bare destination's parentheses +- **WHEN** `"[a](( ))"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("[a](( ))")` and no `Link` + +#### Scenario: CommonMark oracle cases +- **WHEN** the inline cases under `tests/fixtures/conformance/commonmark/` are parsed and rendered with the `html` feature +- **THEN** the output matches the expected HTML + +## ADDED Requirements + +### Requirement: Footnote labels +The parser SHALL read `[^label]` as a footnote reference, and `[^label]:` as a +footnote definition, only when the label is non-empty, holds no space, tab, or +line ending, and, as a link label, holds no unescaped `[` or `]`. + +#### Scenario: Bracket inside a footnote label +- **WHEN** `"^*[^[^]]"` and `"[^a[b]"` are parsed with `parse` +- **THEN** neither paragraph holds a `FootnoteReference`, while `"[^a\\[b]"` holds one + +### Requirement: Reference label matching +Two link labels SHALL match when they agree after Unicode case folding, +trimming, and collapsing each run of spaces, tabs, and line endings to one +space; any other whitespace char is matched as written. + +#### Scenario: No-break space in a label +- **WHEN** `"[a\u{a0}b]\n\n[a b]: /u"` is parsed +- **THEN** the paragraph holds no `LinkReference` + +### Requirement: Angle-bracket autolink URI +The parser SHALL read `` as an autolink when the scheme is valid +and the rest holds no space, ASCII control char, `<`, or `>`; any other +whitespace char is part of the URI. + +#### Scenario: No-break space in an angle-bracket autolink +- **WHEN** `""` is parsed with the CommonMark preset +- **THEN** the paragraph holds an `Autolink` to `http://a\u{a0}b` + +### Requirement: Hard line breaks from spaces +A line ending SHALL be a hard break when two or more spaces the source holds, +and no tab, end the line; spaces or tabs a character reference writes are +text, and only the source's spaces and tabs before a soft break are removed. + +#### Scenario: Referenced space before a line ending +- **WHEN** `"a \nb"` is parsed with the CommonMark preset +- **THEN** the paragraph holds `Text("a ")`, a `SoftBreak`, and `Text("b")` + +### Requirement: Processing instructions +Raw inline HTML SHALL read `` +after the `` is text +- **WHEN** `"a b"` is parsed with the CommonMark preset +- **THEN** the paragraph holds no `Html` inline diff --git a/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/public-api.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/public-api.md new file mode 100644 index 0000000..e86d4c6 --- /dev/null +++ b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/public-api.md @@ -0,0 +1,140 @@ +# Public API — spec changes + +## ADDED Requirements + +### Requirement: Spans map stripped lines back to the source +A parsed node's span SHALL end after the source byte where its last character +was read, and SHALL start at the source byte where its first character was +read, or for a block, where its first line starts after the markers and +indentation of the containers around it. This SHALL hold wherever the parser +removes indentation, container markers, or table-cell padding before reading a +line, joins lines whose source line ending is `\r\n`, or reads `\|` in a table +cell as `|`, which maps to both of its bytes; a space the parser produces by +splitting a tab SHALL map to that tab. + +#### Scenario: Leading whitespace on a paragraph line +- **WHEN** `parse(" a *b*")` runs +- **THEN** the paragraph holds `Text("a ")` spanning bytes 2..4 and an `Emphasis` spanning bytes 4..7 + +#### Scenario: Block quote continuation line +- **WHEN** `parse("> a\n> b *c*")` runs +- **THEN** the block quote's paragraph spans bytes 2..11 and holds `Text("b ")` spanning bytes 6..8 and an `Emphasis` spanning bytes 8..11 + +#### Scenario: List item continuation line +- **WHEN** `parse("- a\n b *c*")` runs +- **THEN** the item's paragraph spans bytes 2..11 and holds an `Emphasis` spanning bytes 8..11 + +#### Scenario: Later block inside a block quote +- **WHEN** `parse("> a\n>\n> b")` runs +- **THEN** the block quote's second paragraph spans bytes 8..9 + +#### Scenario: Nested list item after multi-byte text +- **WHEN** `parse("- 项目\n - 嵌套 [[library/工作/买菜]]\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 20..45 + +#### Scenario: Nested block quote +- **WHEN** `parse("> 外层\n> > 内层 [[A]]\n")` runs +- **THEN** the inner block quote's `WikiLink` spans bytes 20..25 + +#### Scenario: Alert body +- **WHEN** `parse("> [!NOTE]\n> 见 [[A]]\n")` runs +- **THEN** the alert's `WikiLink` spans bytes 16..21 + +#### Scenario: Footnote definition continuation line +- **WHEN** `parse("正文[^1]\n\n[^1]: 见 [[A]]\n 续 [[B]]\n")` runs +- **THEN** the footnote definition's second `WikiLink` spans bytes 36..41 + +#### Scenario: Inside an HTML container +- **WHEN** `parse("

\n更多\n\n- 项目\n - [[A]]\n\n
\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 50..55 + +#### Scenario: Inside a container directive +- **WHEN** `parse(":::note\n- 项目\n - [[A]]\n:::\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 21..26 + +#### Scenario: Tab-indented nested list +- **WHEN** `parse("- 项目\n\t- 嵌套 [[A]]\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 19..24 + +#### Scenario: CRLF nested list +- **WHEN** `parse("- 项目\r\n - 嵌套 [[A]]\r\n")` runs +- **THEN** the nested item's `WikiLink` spans bytes 21..26 + +#### Scenario: CRLF soft break +- **WHEN** `parse("a\r\nb")` runs +- **THEN** the paragraph holds a `SoftBreak` spanning bytes 1..3 and `Text("b")` spanning bytes 3..4 + +#### Scenario: Table cell content +- **WHEN** `parse("| a *b* |\n|-|")` runs +- **THEN** the header cell spans bytes 2..7 and holds `Text("a ")` spanning bytes 2..4 and an `Emphasis` spanning bytes 4..7 + +#### Scenario: Escaped pipe in a table cell +- **WHEN** `parse("| a\\|b |\n|-|")` runs +- **THEN** the header cell holds `Text("a|b")` spanning bytes 2..6 + +#### Scenario: Escaped pipe opening a table cell +- **WHEN** `parse("| \\|a |\n|-|")` runs +- **THEN** the header cell holds `Text("|a")` spanning bytes 2..5 + +#### Scenario: Split tab +- **WHEN** `parse(">\t\tfoo")` runs +- **THEN** the block quote holds an indented code block with value `" foo\n"` spanning bytes 1..6 + +#### Scenario: Container span regression cases +- **WHEN** `inline_spans_address_source_inside_containers` in `tests/parse_span_contract.rs` parses each of its 31 cases (lists, task lists, block quotes, alerts, tables, footnote definitions, HTML containers, container directives, frontmatter, CRLF, and tabs) +- **THEN** every `WikiLink`, `Link`, `Image`, and `#`-holding `Text` it collects spans exactly the literal it occupies in the input + +### Requirement: Spans nest +Every parsed node's span SHALL lie on UTF-8 character boundaries within the +input and within the span of the node that contains it, and the spans of a +node's children SHALL be in source order and SHALL NOT overlap. + +#### Scenario: Fixture corpus and generated inputs +- **WHEN** every fixture input and every seeded generated input is parsed in each dialect +- **THEN** every node, at every depth, satisfies these conditions + +## MODIFIED Requirements + +### Requirement: Source spans +Every parsed node SHALL carry an absolute, half-open UTF-8 byte range into the +original input; hand-built nodes SHALL carry `None`; `LineIndex` SHALL convert a +span to 1-based line and column positions. + +#### Scenario: First block position +- **WHEN** `parse("# Title\n\nHello.")` runs and the first block's span is passed to `LineIndex::new(source).span(span)` +- **THEN** the start position is line 1, column 1 + +#### Scenario: Hand-built node +- **WHEN** a `Heading` is built with `Heading::new(1, [Text::from("Title")])` +- **THEN** its `span()` is `None` + +#### Scenario: Empty table cell +- **WHEN** `parse("| a | |\n|-|-|")` runs +- **THEN** the second header cell's span is the empty range at byte 6, just before the pipe that closes it + +#### Scenario: Missing table cell +- **WHEN** `parse("| a | b |\n|-|-|\n| c")` runs +- **THEN** the body row's second cell has no children and its span is the empty range at the end of the row, byte 19 + +### Requirement: Emphasis-like spans cover their delimiters +The span of a parsed emphasis-like container (`Emphasis`, `Strong`, +`Underline`, `Delete`, `Insert`, `Mark`, `Spoiler`, `Subscript`, or +`Superscript`) SHALL run from the first character of the delimiters that open +it to the last character of the delimiters that close it, and SHALL lie within +the span of the node that contains it. + +#### Scenario: Strong inside emphasis +- **WHEN** `parse("***a***")` runs +- **THEN** the paragraph holds an `Emphasis` spanning bytes 0..7 that holds a `Strong` spanning bytes 1..6 + +#### Scenario: Emphasis inside strong +- **WHEN** `parse("x ***a* b**")` runs +- **THEN** the paragraph holds a `Strong` spanning bytes 2..11 that holds an `Emphasis` spanning bytes 4..7 + +#### Scenario: Leftover opening delimiter +- **WHEN** `parse("**a*")` runs +- **THEN** the paragraph holds `Text("*")` spanning bytes 0..1 and an `Emphasis` spanning bytes 1..4 + +#### Scenario: Emphasis on a block quote continuation line +- **WHEN** `parse("> a\n> *b*")` runs +- **THEN** the block quote's paragraph holds an `Emphasis` spanning bytes 6..9 diff --git a/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/serialization.md b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/serialization.md new file mode 100644 index 0000000..7b5565a --- /dev/null +++ b/docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/specs/serialization.md @@ -0,0 +1,282 @@ +# Serialization — spec changes + +## ADDED Requirements + +### Requirement: Literal backticks are always escaped +The serializer SHALL write every backtick in a text value as `` \` ``. + +#### Scenario: Backtick run before a lone backtick +- **WHEN** a hand-built paragraph holding ```Text("b ``a`")``` is serialized +- **THEN** `to_markdown()` returns ``"b \\`\\`a\\`\n"`` and reparsing it yields the same text and no code span + +#### Scenario: Paired backticks +- **WHEN** ``parse("Test \\`hello world` here.").document.to_markdown()`` runs +- **THEN** it returns ``"Test \\`hello world\\` here.\n"`` + +## MODIFIED Requirements + +### Requirement: Escaping keeps text literal +The serializer SHALL escape text so that reparsing the output yields the same +text and leaves the nodes beside it unchanged, rather than forming new +constructs. + +#### Scenario: Literal asterisks in text +- **WHEN** a hand-built paragraph holding `Text("*not emphasis*")` is serialized and reparsed +- **THEN** the reparsed paragraph holds the same text and no `Emphasis` + +#### Scenario: Underscore that can close inside underscore emphasis +- **WHEN** a hand-built paragraph holding an `Emphasis` around a `Strong` around `Text("(a b)_.")`, followed by `Text("*#")`, is serialized and reparsed with the CommonMark preset +- **THEN** `to_markdown()` returns `"_**(a b)\\_.**_\\*#\n"` and the reparsed paragraph equals the original apart from spans + +#### Scenario: Parenthesis after a shortcut reference +- **WHEN** the document parsed from `"[foo]\\(a)\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed paragraph holds a shortcut `LinkReference` to `foo` followed by `Text("(a)")` + +#### Scenario: Parenthesis after a shortcut image reference +- **WHEN** the document parsed from `"![foo]\\(a)\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed paragraph holds a shortcut `ImageReference` to `foo` followed by `Text("(a)")` + +#### Scenario: Colon after a shortcut reference that starts a paragraph +- **WHEN** the document parsed from `"[foo]\\: /x\n\n[foo]: /u"` is serialized and reparsed +- **THEN** the reparsed document still holds the paragraph, with a shortcut `LinkReference` to `foo` followed by `Text(": /x")` + +#### Scenario: Pipe ending a level-two setext heading +- **WHEN** the document parsed from `"a |\n-"` is serialized and reparsed +- **THEN** `to_markdown()` returns `"a \\|\n---\n"` and the reparsed document holds the same setext `Heading` and no `Table` + +#### Scenario: Empty fenced code block +- **WHEN** ``parse("```\n```").document.to_markdown()`` runs +- **THEN** it returns ``"```\n```\n"`` + +#### Scenario: Whitespace at the ends of an info string +- **WHEN** the document parsed from ``"``` a \nb\n```"`` is serialized +- **THEN** `to_markdown()` returns ``"``` a \nb\n```\n"`` and reparsing it yields the info string `" a\t"` + +#### Scenario: Text right after a literal autolink +- **WHEN** the documents parsed from `"://&"` and `"www.}"` are serialized and reparsed +- **THEN** each reparsed paragraph holds the same `Autolink` and `Text` as the parsed one + +#### Scenario: Paragraph that opens with a soft break +- **WHEN** the document parsed from `" \na"` is serialized +- **THEN** `to_markdown()` returns `" \na\n"` + +#### Scenario: Text line that would open a block +- **WHEN** the documents parsed from `"a\n\\
"` and `"a\n\\::b"` are serialized and reparsed +- **THEN** each reparsed document holds the same single `Paragraph` + +#### Scenario: HTML block value +- **WHEN** the document parsed from `"", index)? + 3); } + // The `?>` closing a processing instruction follows its `", index)? + 2); + return Some(lookups.find("?>", index + 2)? + 2); } if rest.starts_with("", index)? + 3); @@ -8526,7 +9286,7 @@ fn is_uri_autolink(input: &str) -> bool { input[colon + 1..] .chars() .map(source_char) - .all(|char| !matches!(char, '<' | '>') && !char.is_control() && !char.is_whitespace()) + .all(|char| !matches!(char, '<' | '>' | ' ') && !char.is_ascii_control()) } fn is_email_autolink(input: &str) -> bool { @@ -8556,6 +9316,16 @@ fn is_email_autolink(input: &str) -> bool { // returned destination is the synthesized href (a `http://`/`mailto:` prefix // may be prepended); the caller keeps `input[index..end]` as the visible // original. +/// The lengths of the literal autolink that starts `input` under GFM autolinks +/// and under GFM plus relaxed autolinks, for the serializer to check that text +/// it writes after an autolink leaves the URL as it was. +pub(crate) fn literal_autolink_extents(input: &str) -> [Option; 2] { + [false, true].map(|relaxed| { + parse_literal_autolink(input, 0, true, relaxed, &mut LiteralAutolinkScan::default()) + .map(|(end, _)| end) + }) +} + fn parse_literal_autolink( input: &str, index: usize, @@ -9398,10 +10168,15 @@ fn is_email_domain(input: &str, min_labels: usize) -> bool { label_count >= min_labels } +/// A footnote label: no space, tab, or line ending and, as in a link label, +/// no unescaped bracket. fn is_footnote_label(label: &str) -> bool { !label.is_empty() && reference_label_is_within_limit(label) - && !label.chars().any(char::is_whitespace) + && !label.contains([' ', '\t', '\n', '\r']) + && !label + .match_indices(['[', ']']) + .any(|(index, _)| !is_escaped_at(label, index)) } fn find_footnote_definition_label_end(input: &str) -> Option { diff --git a/src/parse/nul.rs b/src/parse/nul_replacement.rs similarity index 100% rename from src/parse/nul.rs rename to src/parse/nul_replacement.rs diff --git a/src/parse/scan_tests.rs b/src/parse/scan_tests.rs index afc6b28..942e5be 100644 --- a/src/parse/scan_tests.rs +++ b/src/parse/scan_tests.rs @@ -505,7 +505,7 @@ mod reference { // character; Unicode whitespace (e.g. U+00A0) is ordinary. A backslash // before a space is NOT an escape (only ASCII punctuation is escapable), // so `\ ` still terminates the destination → `[a](\ b)` is not a link. - if (char == ' ' || char.is_ascii_control()) && depth == 0 { + if char == ' ' || char.is_ascii_control() { break; } if char == '(' && !is_escaped_at(input, cursor) { @@ -1308,7 +1308,8 @@ fn literal_autolink_scans_match_the_reference_scan() { fn html_container_closes_match_the_reference_scan() { let mut rng = Rng(17); for input in generated_inputs(700, 40, 17) { - let lines = collect_lines(&input, 0); + let map = SourceMap::verbatim(input.len(), 0); + let lines = collect_lines(&input, &map); let starts: Vec = (0..=lines.len()).collect(); for order in query_orders(&starts, &mut rng) { let mut closes = BracketMemo::default(); @@ -1367,7 +1368,8 @@ fn flow_inputs() -> Vec { fn flow_jsx_and_expression_closes_match_the_reference_scan() { let mut rng = Rng(19); for input in flow_inputs() { - let lines = collect_lines(&input, 0); + let map = SourceMap::verbatim(input.len(), 0); + let lines = collect_lines(&input, &map); let starts: Vec = (0..lines.len()).collect(); for order in query_orders(&starts, &mut rng) { let mut flow = MdxFlowScan::default(); @@ -1427,7 +1429,9 @@ fn table_row_spoilers_form_where_the_row_scan_predicts() { .count(); let text = table_cell_text(&row[start..end]); let mut diagnostics = Vec::new(); - let formed = parse_inlines(text.trim(), 0, &options, &[], &mut diagnostics) + let content = text.trim(); + let map = SourceMap::verbatim(content.len(), 0); + let formed = parse_inlines(content, &map, &options, Some(&[]), &mut diagnostics) .iter() .filter(|inline| matches!(inline, Inline::Spoiler(_))) .count(); diff --git a/src/parse/source_map.rs b/src/parse/source_map.rs new file mode 100644 index 0000000..58a0a1e --- /dev/null +++ b/src/parse/source_map.rs @@ -0,0 +1,476 @@ +//! Where the text the parser reads came from in the original input. +//! +//! Containers and inline parsing read derived strings: lines with container +//! markers and indentation removed, tabs split into spaces, lines joined with +//! `\n`, and table cells with `\|` read as `|`. A [`SourceMap`] pairs runs of a +//! derived string with the source bytes they were read from, always in +//! original-input coordinates, so a position in any derived string translates +//! to a span of the input without walking the nesting that produced it. + +use alloc::{string::String, vec::Vec}; + +use super::Line; +use crate::{ + ast::{Inline, NodeMeta}, + diagnostic::Diagnostic, + span::Span, +}; + +/// A run of derived text and the source bytes it was read from. When the two +/// lengths are equal the run maps byte for byte; otherwise every text byte of +/// the run stands for the whole source range (the spaces split from a tab, the +/// `\n` that joins lines ending in `\r\n`, the `|` read from `\|`, or a byte the +/// parser inserted, whose source range is empty). +#[derive(Clone, Copy, Debug)] +pub(super) struct Segment { + text_start: usize, + text_end: usize, + source_start: usize, + source_end: usize, +} + +impl Segment { + pub(super) fn text_start(&self) -> usize { + self.text_start + } + + pub(super) fn text_end(&self) -> usize { + self.text_end + } + + fn verbatim(&self) -> bool { + self.text_end - self.text_start == self.source_end - self.source_start + } + + /// The source position where a node starting at `position` (within this + /// segment's text) starts. + fn start_at(&self, position: usize) -> usize { + if self.verbatim() { + self.source_start + (position - self.text_start) + } else { + self.source_start + } + } + + /// The source position where a node ending at `position` (within this + /// segment's text) ends. + fn end_at(&self, position: usize) -> usize { + if self.verbatim() { + self.source_start + (position - self.text_start) + } else { + self.source_end + } + } +} + +/// The source start of a node starting at `position`: a position on a segment +/// boundary maps through the segment that begins there. +pub(super) fn start_of(segments: &[Segment], position: usize) -> usize { + match segments.iter().find(|segment| position < segment.text_end) { + Some(segment) => segment.start_at(position.max(segment.text_start)), + None => segments + .last() + .map_or(0, |segment| segment.end_at(segment.text_end)), + } +} + +/// The source end of a node ending at `position`: a position on a segment +/// boundary maps through the segment that ends there. +pub(super) fn end_of(segments: &[Segment], position: usize) -> usize { + match segments + .iter() + .rev() + .find(|segment| position > segment.text_start) + { + Some(segment) => segment.end_at(position.min(segment.text_end)), + None => segments + .first() + .map_or(0, |segment| segment.start_at(segment.text_start)), + } +} + +/// The segments of a derived string, in text order, covering it without gaps. +#[derive(Clone, Debug, Default)] +pub(super) struct SourceMap { + segments: Vec, +} + +impl SourceMap { + /// A string read verbatim from `source_start` on. + pub(super) fn verbatim(len: usize, source_start: usize) -> Self { + // An empty string still keeps where it sits in the input. + Self { + segments: alloc::vec![Segment { + text_start: 0, + text_end: len, + source_start, + source_end: source_start + len, + }], + } + } + + pub(super) fn segments(&self) -> &[Segment] { + &self.segments + } + + /// Records that text `text_start..text_start + text_len` was read from + /// `source_start..source_end`. Runs are pushed in text order. + fn push(&mut self, text_start: usize, text_len: usize, source_start: usize, source_end: usize) { + if text_len == 0 { + return; + } + let segment = Segment { + text_start, + text_end: text_start + text_len, + source_start, + source_end, + }; + if let Some(last) = self.segments.last_mut() { + if last.verbatim() + && segment.verbatim() + && last.text_end == segment.text_start + && last.source_end == segment.source_start + { + last.text_end = segment.text_end; + last.source_end = segment.source_end; + return; + } + } + self.segments.push(segment); + } + + /// Forgets what the map holds for text from `len` on. + fn truncate(&mut self, len: usize) { + while self + .segments + .last() + .is_some_and(|last| last.text_start >= len) + { + self.segments.pop(); + } + if let Some(last) = self.segments.last_mut() { + if last.text_end > len { + if last.verbatim() { + last.source_end -= last.text_end - len; + } + last.text_end = len; + } + } + } + + /// Records that text from `text_start` on repeats what `segments` map for + /// their text `from..to`. + pub(super) fn copy(&mut self, text_start: usize, segments: &[Segment], from: usize, to: usize) { + let mut covered = from; + for segment in segments { + if segment.text_end <= from || segment.text_start >= to { + continue; + } + let start = segment.text_start.max(from); + let end = segment.text_end.min(to); + let (source_start, source_end) = if segment.verbatim() { + (segment.start_at(start), segment.end_at(end)) + } else { + (segment.source_start, segment.source_end) + }; + self.push( + text_start + (start - from), + end - start, + source_start, + source_end, + ); + covered = end; + } + if covered < to { + // Text the segments do not cover (none is expected) stays anchored + // at the end of what they do. + let at = end_of(segments, covered); + self.push(text_start + (covered - from), to - covered, at, at); + } + } +} + +/// A derived string built line by line, with its source map. Lines are joined +/// with `\n`; a joiner stands for the line ending of the line before it. +#[derive(Clone, Debug, Default)] +pub(super) struct DerivedText { + pub(super) text: String, + map: SourceMap, + /// The source range of the line ending the next joiner stands for. + pending_eol: Option<(usize, usize)>, + /// The source column each pushed line starts at. + columns: Vec, +} + +impl DerivedText { + pub(super) fn map(&self) -> &SourceMap { + &self.map + } + + pub(super) fn into_map(self) -> SourceMap { + self.map + } + + fn join(&mut self) { + if !self.text.is_empty() || self.pending_eol.is_some() { + let (start, end) = self.pending_eol.unwrap_or_else(|| { + let at = end_of(self.map.segments(), self.text.len()); + (at, at) + }); + self.map.push(self.text.len(), 1, start, end); + self.text.push('\n'); + } + } + + /// Appends `derived`, which `line` read from its text: either a slice of + /// `line.text`, or the text from byte `from` of `line.text` on with its + /// leading whitespace expanded (a tab split into spaces), joined to the + /// previous line. + pub(super) fn push_line(&mut self, line: &Line<'_>, derived: &str, from: usize) { + self.join(); + self.columns.push(derived_column(line, derived, from)); + self.append(line, derived, from); + self.pending_eol = Some(line.eol_source()); + } + + /// The lines of the text, each starting at the source column it was read + /// from. + pub(super) fn lines(&self) -> Vec> { + let mut lines = super::collect_lines(&self.text, &self.map); + for (line, column) in lines.iter_mut().zip(&self.columns) { + line.column = *column; + } + lines + } + + /// Appends `derived` to the current line without a joiner. + pub(super) fn append(&mut self, line: &Line<'_>, derived: &str, from: usize) { + let at = self.text.len(); + match slice_offset(line.text, derived) { + Some(offset) => line.copy_into(&mut self.map, at, offset, offset + derived.len()), + None => { + // `derived` expands the whitespace at the start of + // `line.text[from..]`; the rest is verbatim. + let raw = &line.text[from..]; + let suffix = common_suffix_len(raw, derived); + let head = derived.len() - suffix; + let raw_head_end = from + raw.len() - suffix; + let source_start = line.source_start(from); + let source_end = line.source_end(raw_head_end); + self.map.push(at, head, source_start, source_end); + line.copy_into(&mut self.map, at + head, raw_head_end, line.text.len()); + } + } + self.text.push_str(derived); + } + + /// Appends `line.text` with `inserted` placed before byte `offset`, joined + /// to the previous line. The inserted bytes have no source of their own. + pub(super) fn push_line_with_insertion( + &mut self, + line: &Line<'_>, + offset: usize, + inserted: &str, + ) { + self.join(); + self.columns.push(line.column); + let at = self.text.len(); + line.copy_into(&mut self.map, at, 0, offset); + let source = line.source_start(offset); + self.map.push(at + offset, inserted.len(), source, source); + line.copy_into( + &mut self.map, + at + offset + inserted.len(), + offset, + line.text.len(), + ); + self.text.push_str(&line.text[..offset]); + self.text.push_str(inserted); + self.text.push_str(&line.text[offset..]); + self.pending_eol = Some(line.eol_source()); + } + + /// Ends the last pushed line with `\n`, mapped to the line ending it was + /// read with, as the lines before it are joined. + pub(super) fn push_pending_eol(&mut self) { + self.join(); + self.pending_eol = None; + } + + /// Drops the spaces and tabs that end the text, with the map runs they + /// covered. + pub(super) fn trim_final_whitespace(&mut self) { + let len = self.text.trim_end_matches([' ', '\t']).len(); + self.text.truncate(len); + self.map.truncate(len); + } + + /// Appends `text`, read from the whole source range `source_start.. + /// source_end`, which it replaces. + pub(super) fn append_replacing(&mut self, text: &str, source_start: usize, source_end: usize) { + self.map + .push(self.text.len(), text.len(), source_start, source_end); + self.text.push_str(text); + } + + /// Appends text the parser adds that the source does not hold. + pub(super) fn push_synthetic(&mut self, text: &str) { + let at = end_of(self.map.segments(), self.text.len()); + self.map.push(self.text.len(), text.len(), at, at); + self.text.push_str(text); + } +} + +/// The source column `derived` starts at, which `line` read from byte `from` +/// of its text (see [`DerivedText::push_line`]): the column of its offset +/// when it is a slice of the text, or else the column its verbatim tail starts +/// at less the spaces its expanded head writes. +pub(super) fn derived_column(line: &Line<'_>, derived: &str, from: usize) -> usize { + if let Some(offset) = slice_offset(line.text, derived) { + return line.column_at(offset); + } + let raw = &line.text[from..]; + let suffix = common_suffix_len(raw, derived); + let head = derived.len() - suffix; + line.column_at(from + raw.len() - suffix) + .saturating_sub(head) +} + +/// The offset of `slice` inside `text` when it is a borrowed sub-slice of it. +fn slice_offset(text: &str, slice: &str) -> Option { + let start = text.as_ptr() as usize; + let offset = (slice.as_ptr() as usize).checked_sub(start)?; + (offset + slice.len() <= text.len()).then_some(offset) +} + +fn common_suffix_len(a: &str, b: &str) -> usize { + a.bytes() + .rev() + .zip(b.bytes().rev()) + .take_while(|(x, y)| x == y) + .count() +} + +/// Translates spans that an inline parse produced in its input's coordinates +/// into original-input spans, with one walk over the nodes it produced and the +/// diagnostics it pushed. +pub(super) fn translate_inlines( + map: &SourceMap, + inlines: &mut [Inline], + diagnostics: &mut [Diagnostic], +) { + let mut translator = Translator { + segments: map.segments(), + cursor: 0, + }; + translator.inlines(inlines); + // Diagnostics arrive in scan order, so their starts mostly increase too. + translator.cursor = 0; + for diagnostic in diagnostics { + if let Some(span) = diagnostic.span { + diagnostic.span = Some(translator.span(span)); + } + } +} + +/// Maps spans with a cursor that only moves forward: in preorder, node starts +/// never decrease, and each end is found by searching on from its start. +struct Translator<'a> { + segments: &'a [Segment], + cursor: usize, +} + +impl Translator<'_> { + fn span(&mut self, span: Span) -> Span { + if self.segments.is_empty() { + return span; + } + if span.start < self.segments[self.cursor].text_start { + // Out of order: find the segment again. + self.cursor = self + .segments + .partition_point(|segment| segment.text_end <= span.start) + .min(self.segments.len() - 1); + } + while self.cursor + 1 < self.segments.len() + && span.start >= self.segments[self.cursor].text_end + { + self.cursor += 1; + } + let start = start_of(&self.segments[self.cursor..], span.start); + let last = self.end_segment(span.end); + let end = end_of(&self.segments[self.cursor..=last], span.end); + Span::new(start, end.max(start)) + } + + /// The first segment at or after the cursor that `end` falls within, found + /// by galloping: diagnostics do not nest, and many run to the input's end, + /// so walking every segment they cross would make the pass quadratic. + fn end_segment(&self, end: usize) -> usize { + let last = self.segments.len() - 1; + if end > self.segments[last].text_start { + return last; + } + let covers = |index: usize| end <= self.segments[index].text_end; + let mut low = self.cursor; + let mut step = 1; + while !covers(low) { + let next = (low + step).min(last); + if covers(next) { + let found = + self.segments[low + 1..=next].partition_point(|segment| end > segment.text_end); + return low + 1 + found; + } + low = next; + step *= 2; + } + low + } + + fn meta(&mut self, meta: &mut NodeMeta) { + if let Some(span) = meta.span { + meta.span = Some(self.span(span)); + } + } + + fn inlines(&mut self, inlines: &mut [Inline]) { + for inline in inlines { + match inline { + Inline::Text(node) => self.meta(&mut node.meta), + Inline::Escape(node) => self.meta(&mut node.meta), + Inline::CharacterReference(node) => self.meta(&mut node.meta), + Inline::SoftBreak(node) => self.meta(&mut node.meta), + Inline::LineBreak(node) => self.meta(&mut node.meta), + Inline::Shortcode(node) => self.meta(&mut node.meta), + Inline::Code(node) => self.meta(&mut node.meta), + Inline::Autolink(node) => self.meta(&mut node.meta), + Inline::Html(node) => self.meta(&mut node.meta), + Inline::Math(node) => self.meta(&mut node.meta), + Inline::FootnoteReference(node) => self.meta(&mut node.meta), + Inline::WikiLink(node) => self.meta(&mut node.meta), + Inline::MdxExpression(node) => self.meta(&mut node.meta), + Inline::MdxJsx(node) => self.meta(&mut node.meta), + Inline::Emphasis(node) => self.container(&mut node.meta, &mut node.children), + Inline::Strong(node) => self.container(&mut node.meta, &mut node.children), + Inline::Underline(node) => self.container(&mut node.meta, &mut node.children), + Inline::Delete(node) => self.container(&mut node.meta, &mut node.children), + Inline::Insert(node) => self.container(&mut node.meta, &mut node.children), + Inline::Mark(node) => self.container(&mut node.meta, &mut node.children), + Inline::Subscript(node) => self.container(&mut node.meta, &mut node.children), + Inline::Superscript(node) => self.container(&mut node.meta, &mut node.children), + Inline::Spoiler(node) => self.container(&mut node.meta, &mut node.children), + Inline::InlineFootnote(node) => self.container(&mut node.meta, &mut node.children), + Inline::Link(node) => self.container(&mut node.meta, &mut node.children), + Inline::Image(node) => self.container(&mut node.meta, &mut node.alt), + Inline::LinkReference(node) => self.container(&mut node.meta, &mut node.children), + Inline::ImageReference(node) => self.container(&mut node.meta, &mut node.alt), + Inline::TextDirective(node) => self.container(&mut node.meta, &mut node.label), + } + } + } + + fn container(&mut self, meta: &mut NodeMeta, children: &mut [Inline]) { + self.meta(meta); + self.inlines(children); + } +} diff --git a/src/serialize.rs b/src/serialize.rs index cdc76ea..a019ac2 100644 --- a/src/serialize.rs +++ b/src/serialize.rs @@ -5,8 +5,11 @@ //! with a [`SerializeError`]. use alloc::{ + borrow::Cow, + collections::BTreeMap, format, string::{String, ToString}, + vec, vec::Vec, }; @@ -14,8 +17,12 @@ use crate::{ ast::*, diagnostic::Diagnostic, memo::{pattern_starts, PathMemo, Positions, Step}, - parse::{gfm_table_can_start_source, line_starts_html_block}, - validate::validate_document, + parse::{ + continuation_line_breaks_paragraph, gfm_table_can_start_source, is_flanking_punctuation, + line_opens_alert, line_starts_html_block, line_starts_interrupting_html_block, + line_starts_math_block, literal_autolink_extents, + }, + validate::{is_directive_name, validate_document}, }; /// The newline style emitted by the serializer. @@ -97,17 +104,49 @@ fn serialize_document_body( ) -> Result { let mut output = serialize_blocks_at_start(&document.children, options, true)?; if options.line_ending == LineEnding::CrLf { - output = output.replace('\n', "\r\n"); + output = lf_to_crlf(&output); } - if options.final_newline - && !output.is_empty() - && !output.ends_with(options.line_ending.as_str()) - { + // Blocks end without a line ending, except an HTML block whose last line + // is empty (the final newline ends that line too) and an indented code + // block that keeps its value's `\r` or `\r\n` ending. + if options.final_newline && !output.is_empty() && !ends_with_carriage_return_ending(&output) { output.push_str(options.line_ending.as_str()); } Ok(output) } +/// Rewrites each bare `\n` as `\r\n`; existing `\r\n` and `\r` endings that +/// verbatim values carry stay as they are. +fn lf_to_crlf(input: &str) -> String { + let mut output = String::with_capacity(input.len()); + let mut previous = '\0'; + for ch in input.chars() { + if ch == '\n' && previous != '\r' { + output.push('\r'); + } + output.push(ch); + previous = ch; + } + output +} + +fn ends_with_carriage_return_ending(input: &str) -> bool { + input.ends_with('\r') || input.ends_with("\r\n") +} + +/// Separates two blocks with one blank line. A block that already ends its +/// last line with `\r` or `\r\n` gets only the blank line, written so that it +/// cannot merge into that ending. +fn push_block_gap(output: &mut String) { + if output.ends_with('\r') { + output.push('\r'); + } else if output.ends_with("\r\n") { + output.push('\n'); + } else { + output.push_str("\n\n"); + } +} + /// Serialize a block sequence. `document_start` is true only for the top-level /// document body, where the first block sits at byte 0 and a contiguous `---` /// would open frontmatter; that one position emits a spaced dash thematic break @@ -117,26 +156,60 @@ fn serialize_blocks_at_start( options: &SerializeOptions, document_start: bool, ) -> Result { + // Written last to first: a list reads the indentation of the block after + // it, which would join its last item unless the items' content starts + // further in. + let mut outputs: Vec = Vec::with_capacity(blocks.len()); + for (index, block) in blocks.iter().enumerate().rev() { + let at_document_start = document_start && index == 0; + let next_indent = outputs + .last() + .map(|next: &String| next.len() - next.trim_start_matches(' ').len()); + let written = match (block, next_indent) { + (Block::List(list), Some(indent @ 1..=3)) => { + serialize_list_with_marker_spacing(list, options, &" ".repeat(indent), " ")? + } + (Block::List(list), Some(4..)) => { + serialize_list_with_marker_spacing(list, options, " ", " ")? + } + _ => serialize_block(block, options, at_document_start)?, + }; + outputs.push(written); + } let mut output = String::new(); - for (index, block) in blocks.iter().enumerate() { + for (index, written) in outputs.iter().rev().enumerate() { if index > 0 { - output.push_str("\n\n"); - } - let at_document_start = document_start && index == 0; - if let ( - Block::List(list), - Some(Block::CodeBlock(CodeBlock { - kind: CodeBlockKind::Indented, - .. - })), - ) = (block, blocks.get(index + 1)) - { - output.push_str(&serialize_list_with_marker_spacing( - list, options, " ", " ", - )?); - } else { - output.push_str(&serialize_block(block, options, at_document_start)?); + let first_line = written.split('\n').next().unwrap_or(""); + let first_inline = match &blocks[index] { + Block::Paragraph(paragraph) => paragraph.children.first(), + Block::Heading(heading) if heading.kind == HeadingKind::Setext => { + heading.children.first() + } + _ => None, + }; + // Raw HTML that would start an HTML block, or math that would + // start a math block, opens such content only as the continuation + // of the paragraph a definition was read from; a line that would + // interrupt it is indented as well. + let continues_definition = matches!(blocks[index - 1], Block::Definition(_)) + .then(|| match first_inline { + Some(Inline::Html(_)) if line_starts_html_block(first_line) => { + Some(line_starts_interrupting_html_block(first_line)) + } + Some(Inline::Math(_)) if line_starts_math_block(first_line) => Some(true), + _ => None, + }) + .flatten(); + if let Some(indented) = continues_definition { + output.push('\n'); + if indented { + output.push_str(" "); + } + } else { + push_block_gap(&mut output); + } } + output.push_str(written); } Ok(output) } @@ -148,21 +221,37 @@ fn serialize_block( ) -> Result { match block { Block::Paragraph(node) => serialize_paragraph(node, options), - Block::Heading(node) => { - let content = serialize_inlines(&node.children, options)?; + Block::Heading(node) => serialize_reading_back(&node.children, options, |content| { // A setext underline can only express depth 1 (`=`) or 2 (`-`); any // other depth must fall back to ATX, otherwise the depth is lost. // Multi-line content stays setext because ATX is single-line and // would split a heading the parser legitimately produces. let setext_representable = matches!(node.depth, 1 | 2); - Ok(match node.kind { + match node.kind { HeadingKind::Setext if setext_representable => { let marker = if node.depth == 1 { '=' } else { '-' }; - format!( - "{}\n{}", - content, - marker.to_string().repeat(content.len().max(3)) - ) + let underline = marker.to_string().repeat(content.len().max(3)); + let ends_with_text_pipe = matches!( + node.children.last(), + Some(Inline::Text(text)) if text.value.trim_end().ends_with('|') + ); + let content = if node.depth == 2 && ends_with_text_pipe { + escape_pipe_ending_table_header(content) + } else { + content + }; + let mut content = content; + keep_first_line_off_html_block(&node.children, &mut content); + keep_first_line_off_esm(&mut content); + let mut content = indent_block_starting_continuations(content); + if let Some(last_line_start) = content.rfind('\n').map(|end| end + 1) { + // A continuation line would read as a table header + // over the underline; indented, it reads as before. + if gfm_table_can_start_source(&content[last_line_start..], &underline) { + content.insert_str(last_line_start, " "); + } + } + format!("{content}\n{underline}") } _ if content.is_empty() => "#".repeat(node.depth as usize), _ => format!( @@ -170,8 +259,8 @@ fn serialize_block( "#".repeat(node.depth as usize), escape_atx_heading_content(&content) ), - }) - } + } + }), Block::ThematicBreak(node) => Ok(match node.marker { // A Dash break is normally written contiguous (`---`) — the form // that survives after a `-` bullet list, where the spaced `- - -` @@ -187,6 +276,11 @@ fn serialize_block( let inner = serialize_blocks_at_start(&node.children, options, false)?; if inner.is_empty() { Ok(">".into()) + } else if line_opens_alert(inner.split('\n').next().unwrap_or("")) { + // A raw label such as a definition's `[!NOTE]` on the quote's + // first line would make it an alert; an empty first line + // keeps it a quote. + Ok(alloc::format!(">\n{}", prefix_lines(&inner, "> "))) } else { Ok(prefix_lines(&inner, "> ")) } @@ -195,7 +289,10 @@ fn serialize_block( Block::List(node) => serialize_list(node, options), Block::DescriptionList(node) => serialize_description_list(node, options), Block::CodeBlock(node) => serialize_code_block(node, options), - Block::HtmlBlock(node) => Ok(trim_trailing_newline(&node.value).into()), + // An HTML block's lines are joined with `\n`, so a value ending in one + // ends with an empty line that belongs to the block (an unclosed + // comment, say); it is written as it is. + Block::HtmlBlock(node) => Ok(node.value.clone()), Block::HtmlContainer(node) => serialize_html_container(node, options), Block::Definition(node) => { let destination = serialize_destination_kind( @@ -234,20 +331,16 @@ fn serialize_block( Block::Table(node) => serialize_table(node, options), Block::MathBlock(node) => { let fence = block_math_fence(&node.value); - Ok(format!( - "{fence}\n{}\n{fence}", - trim_trailing_newline(&node.value) - )) + Ok(fenced_body(&fence, &node.value, &fence)) } Block::Frontmatter(node) => { let fence = match node.kind { FrontmatterKind::Yaml => "---", FrontmatterKind::Toml => "+++", }; - Ok(format!( - "{fence}\n{}\n{fence}", - trim_trailing_newline(&node.value) - )) + // The value holds the lines between the fences joined with `\n`, + // so a final `\n` is an empty last line. + Ok(format!("{fence}\n{}\n{fence}", node.value)) } Block::MdxEsm(node) => Ok(node.value.clone()), Block::MdxExpression(node) => Ok(format!("{{{}}}", node.value)), @@ -259,14 +352,19 @@ fn serialize_block( serialize_attributes(&node.attributes) )), Block::ContainerDirective(node) => { - let inner = serialize_blocks_at_start(&node.children, options, false)?; + let mut inner = serialize_blocks_at_start(&node.children, options, false)?; let fence = directive_fence(&inner); + // Content ends with a line ending before the closing fence; an + // empty directive takes no blank line, which would loosen a list + // holding it. + if !inner.is_empty() { + inner.push('\n'); + } Ok(format!( - "{fence}{}{}{}\n{}\n{fence}", + "{fence}{}{}{}\n{inner}{fence}", node.name, serialize_directive_label(&node.label, options)?, serialize_attributes(&node.attributes), - inner )) } } @@ -301,6 +399,29 @@ fn serialize_html_container( /// closing hash sequence. CommonMark treats a final run of `#` preceded by /// whitespace (after trailing whitespace is trimmed) as the optional closing /// sequence; escaping the first `#` of that run keeps it as literal text. +/// `content`, ending in text, with an unescaped `|` that ends its last line +/// escaped: above a `---` underline, a line ending in a bare pipe is a one-cell +/// table header and the underline its delimiter row. +fn escape_pipe_ending_table_header(content: String) -> String { + let last_line = content.rsplit('\n').next().unwrap_or_default(); + let trimmed = last_line.trim_end(); + let Some(before_pipe) = trimmed.strip_suffix('|') else { + return content; + }; + let backslashes = before_pipe + .bytes() + .rev() + .take_while(|byte| *byte == b'\\') + .count(); + if backslashes % 2 == 1 { + return content; + } + let pipe = content.len() - (last_line.len() - trimmed.len()) - 1; + let mut escaped = content; + escaped.insert(pipe, '\\'); + escaped +} + fn escape_atx_heading_content(content: &str) -> String { let trimmed_len = content.trim_end_matches([' ', '\t']).len(); let trimmed = &content[..trimmed_len]; @@ -323,14 +444,504 @@ fn serialize_paragraph( node: &Paragraph, options: &SerializeOptions, ) -> Result { - let mut output = serialize_inlines(&node.children, options)?; - if let Some(offset) = paragraph_html_block_escape_offset(&output) { - output.insert(offset, '\\'); + serialize_reading_back(&node.children, options, |mut output| { + keep_first_line_off_html_block(&node.children, &mut output); + keep_first_line_off_esm(&mut output); + keep_first_line_off_mdx_flow(&node.children, &mut output); + indent_block_starting_continuations(output) + }) +} + +/// The block that `finish` writes around the rendering of `children`, a +/// paragraph's or heading's content. +fn serialize_reading_back( + children: &[Inline], + options: &SerializeOptions, + finish: impl Fn(String) -> String, +) -> Result { + let render = |run_style: RunStyle, + raw_edge: Option, + autolink_edges: AutolinkEdges| + -> Result { + let context = InlineSerializeContext { + run_style, + raw_edge, + autolink_edges, + ..InlineSerializeContext::block_content() + }; + Ok(finish(serialize_inlines_with_context( + children, options, context, + )?)) + }; + let output = render(RunStyle::Plain, None, AutolinkEdges::Plain)?; + let mut expected = None; + // The dialect the content came from is not known here: a rendering that + // reads back under the default preset is taken first, and only when none + // does, one that reads back under GFM or MDX. + let default_preset = [crate::options::SyntaxOptions::default()]; + let other_presets = [ + crate::options::SyntaxOptions::gfm(), + crate::options::SyntaxOptions::mdx(), + ]; + let mut reads_back = |markdown: &str, presets: &[crate::options::SyntaxOptions]| { + let expected = expected.get_or_insert_with(|| without_spans(children)); + reparses_to(markdown, children, expected, presets) + }; + // A strong or emphasis run abutting another splits on reparse only as its + // flanking allows, which the rest of the paragraph decides; one beside a + // `~` opens or closes only as the GFM bonus for a raw `~` allows, and one + // beside a text `*` may take that `*` into its run. A literal autolink's + // URL scan runs on through a space written as a reference, which a span + // delimiter beside it may need. When the plain rendering does not read + // back, the first other style that does is taken. + let mut runs = RunNeighbours::default(); + runs.read(children, 0); + let RunNeighbours { + abut_runs, + edge_tildes, + edge_stars, + } = runs; + let autolink_spaces = autolink_meets_space_in_span(children, false); + let autolink_leads = autolink_text_runs_on(children, false); + if (abut_runs || edge_tildes || edge_stars || autolink_spaces || autolink_leads) + && !reads_back(&output, &default_preset) + { + let styles = [ + RunStyle::Plain, + RunStyle::StrongUnderscore, + RunStyle::EdgeStrongUnderscore, + RunStyle::InnerUnderscore, + RunStyle::AllStar, + RunStyle::OuterUnderscore, + ]; + let raw_edges = [None, edge_tildes.then_some('~'), edge_stars.then_some('*')]; + let run_alternates = raw_edges + .into_iter() + .enumerate() + .filter(|&(at, raw_edge)| at == 0 || raw_edge.is_some()) + .flat_map(|(_, raw_edge)| styles.map(|style| (style, raw_edge, AutolinkEdges::Plain))) + .skip(1) + .filter(|_| abut_runs || edge_tildes || edge_stars); + let autolink_alternates = [ + AutolinkEdges::RawEdges, + AutolinkEdges::EncodedBefore, + AutolinkEdges::EncodedLead, + ] + .into_iter() + .filter(|_| autolink_spaces || autolink_leads) + .map(|edges| (RunStyle::Plain, None, edges)); + let mut alternates = Vec::new(); + for (style, raw_edge, spaces) in run_alternates.chain(autolink_alternates) { + let alternate = render(style, raw_edge, spaces)?; + if reads_back(&alternate, &default_preset) { + return Ok(alternate); + } + alternates.push(alternate); + } + if !reads_back(&output, &other_presets) { + if let Some(alternate) = alternates + .into_iter() + .find(|alternate| reads_back(alternate, &other_presets)) + { + return Ok(alternate); + } + } } - if let Some(offset) = paragraph_table_escape_offset(&output) { + Ok(output) +} + +/// How a paragraph writes the text around a literal autolink (see +/// `serialize_paragraph`). +#[derive(Clone, Copy, Debug, Default, Eq, Ord, PartialEq, PartialOrd)] +enum AutolinkEdges { + /// A space or tab at a line's edge is a reference, one before a literal + /// autolink is raw, and a text after one opens as written unless the URL + /// scan would read on into it. + #[default] + Plain, + /// Every space or tab at a line's edge is raw. + RawEdges, + /// A space or tab before a literal autolink is written as at any edge. + EncodedBefore, + /// A text after a literal autolink and before another inline opens with + /// its first char escaped or written as a reference. + EncodedLead, +} + +/// Whether a literal autolink among `inlines`, within spans, is followed by a +/// text without whitespace short of its end, which leaves the URL scan to read +/// on into what follows: another inline, a span's closing delimiter, or a +/// space written as a reference; `in_span` when `inlines` are a span's content. +fn autolink_text_runs_on(inlines: &[Inline], in_span: bool) -> bool { + inlines.windows(2).enumerate().any(|(index, pair)| { + is_gfm_literal_autolink(&pair[0]) + && matches!(&pair[1], Inline::Text(text) + if !text.value.trim_end_matches([' ', '\t']).contains(char::is_whitespace) + && (in_span || index + 2 < inlines.len())) + }) || inlines.iter().any(|inline| { + span_children(inline).is_some_and(|children| autolink_text_runs_on(children, true)) + }) +} + +/// The content of `inline` when it is a span such as an emphasis, or a +/// link's or an inline footnote's text. +fn span_children(inline: &Inline) -> Option<&[Inline]> { + Some(match inline { + Inline::InlineFootnote(node) => &node.children, + Inline::Link(node) => &node.children, + Inline::LinkReference(node) => &node.children, + Inline::Emphasis(node) => &node.children, + Inline::Strong(node) => &node.children, + Inline::Underline(node) => &node.children, + Inline::Delete(node) => &node.children, + Inline::Insert(node) => &node.children, + Inline::Mark(node) => &node.children, + Inline::Subscript(node) => &node.children, + Inline::Superscript(node) => &node.children, + Inline::Spoiler(node) => &node.children, + _ => return None, + }) +} + +/// Whether `inline`, or the first (`at_start`) or last of its span content, +/// is a text meeting that side with a space or tab; `delimited` when a span +/// delimiter stands between. +fn meets_space(inline: Option<&Inline>, at_start: bool, delimited: bool) -> bool { + match inline { + Some(Inline::Text(text)) => { + delimited + && if at_start { + text.value.starts_with([' ', '\t']) + } else { + text.value.ends_with([' ', '\t']) + } + } + Some(inline) => span_children(inline).is_some_and(|children| { + meets_space( + if at_start { + children.first() + } else { + children.last() + }, + at_start, + true, + ) + }), + None => false, + } +} + +/// Whether a literal autolink among `inlines`, within spans, meets a space or +/// tab across a span delimiter or inside a span, which the delimiter may need +/// written as a reference; `in_span` when `inlines` are a span's content. +fn autolink_meets_space_in_span(inlines: &[Inline], in_span: bool) -> bool { + inlines.iter().enumerate().any(|(index, inline)| { + if is_gfm_literal_autolink(inline) { + return meets_space(inlines.get(index + 1), true, in_span) + || meets_space( + index.checked_sub(1).map(|previous| &inlines[previous]), + false, + in_span, + ); + } + span_children(inline).is_some_and(|children| autolink_meets_space_in_span(children, true)) + }) +} + +/// What sits beside the strong and emphasis runs of a paragraph, at any +/// depth within runs. +#[derive(Default)] +struct RunNeighbours { + /// Two runs sit side by side, a run opens or closes right beside one + /// inside it, or a strong holds a strong or an emphasis an emphasis. + abut_runs: bool, + /// A text opens or closes with a `~` right beside a run's delimiter. + edge_tildes: bool, + /// The same with a `*`. + edge_stars: bool, +} + +impl RunNeighbours { + /// Reads `inlines`; `inside` holds a bit for each run kind around them, + /// `1` for strong and `2` for emphasis. + fn read(&mut self, inlines: &[Inline], inside: u8) { + for (index, inline) in inlines.iter().enumerate() { + let (children, kind): (&[Inline], u8) = match inline { + Inline::Text(node) => { + let after_run = (index == 0 && inside != 0) + || index + .checked_sub(1) + .is_some_and(|previous| is_attention_run(&inlines[previous])); + let before_run = (index + 1 == inlines.len() && inside != 0) + || inlines.get(index + 1).is_some_and(is_attention_run); + let value = node.value.as_bytes(); + for (edge, found) in + [(b'~', &mut self.edge_tildes), (b'*', &mut self.edge_stars)] + { + *found |= (after_run && value.first() == Some(&edge)) + || (before_run && value.last() == Some(&edge)); + } + // A `www` literal autolink needs the `*`, `_`, or `~` + // before it raw, which a run's delimiter choice decides. + self.abut_runs |= inside != 0 + && matches!(value.last(), Some(b'*' | b'_' | b'~')) + && inlines + .get(index + 1) + .and_then(literal_autolink_original) + .is_some_and(|original| { + original.len() >= 3 && original[..3].eq_ignore_ascii_case("www") + }); + continue; + } + Inline::Strong(node) => (&node.children, 1), + Inline::Emphasis(node) => (&node.children, 2), + // A link, image, or mark opens no run, but the runs inside it + // choose their delimiters too. + Inline::Image(node) => (&node.alt, 0), + Inline::ImageReference(node) => (&node.alt, 0), + Inline::TextDirective(node) => (&node.label, 0), + other => match span_children(other) { + Some(children) => (children, 0), + None => continue, + }, + }; + if kind == 0 { + self.read(children, 0); + continue; + } + self.abut_runs |= inside & kind != 0 + || inlines.get(index + 1).is_some_and(is_attention_run) + || children.first().is_some_and(is_attention_run) + || children.last().is_some_and(is_attention_run); + self.read(children, inside | kind); + } + } +} + +fn is_attention_run(inline: &Inline) -> bool { + matches!(inline, Inline::Strong(_) | Inline::Emphasis(_)) +} + +/// `rendered` text with the run of `edge` chars at its start, or at its end, +/// written raw where it was escaped or written as a character reference. +fn unescape_edge(rendered: &str, edge: char, at_start: bool, at_end: bool) -> String { + let escaped = alloc::format!("\\{edge}"); + let reference = alloc::format!("&#x{:X};", edge as u32); + let mut text = rendered; + let mut head = 0; + if at_start { + while let Some(rest) = text + .strip_prefix(escaped.as_str()) + .or_else(|| text.strip_prefix(reference.as_str())) + { + head += 1; + text = rest; + } + } + let mut tail = 0; + if at_end { + loop { + if let Some(before) = text.strip_suffix(reference.as_str()) { + text = before; + } else if let Some(before) = text.strip_suffix(escaped.as_str()) { + let backslashes = before.len() - before.trim_end_matches('\\').len(); + if backslashes % 2 == 1 { + break; + } + text = before; + } else { + break; + } + tail += 1; + } + } + let mut output = String::with_capacity(rendered.len()); + output.extend(core::iter::repeat_n(edge, head)); + output.push_str(text); + output.extend(core::iter::repeat_n(edge, tail)); + output +} + +/// Whether `markdown` parses, under one of `presets`, to one paragraph or +/// heading holding `inlines`, which `expected` holds without spans. +fn reparses_to( + markdown: &str, + inlines: &[Inline], + expected: &[Inline], + presets: &[crate::options::SyntaxOptions], +) -> bool { + // The references in the paragraph resolve against definitions elsewhere + // in the document, which a definition per label stands in for. + let mut labels = Vec::new(); + reference_labels(inlines, &mut labels); + let mut source = String::from(markdown); + for label in labels { + source.push_str("\n\n["); + source.push_str(label); + source.push_str("]: u"); + } + presets.iter().any(|options| { + let document = options.parse(&source).document; + match document.children.as_slice() { + [Block::Paragraph(Paragraph { children, .. }) + | Block::Heading(Heading { children, .. }), definitions @ ..] + if definitions + .iter() + .all(|block| matches!(block, Block::Definition(_))) => + { + let mut reparsed = children.clone(); + clear_spans(&mut reparsed); + reparsed == expected + } + _ => false, + } + }) +} + +/// The labels of the link and image references in `inlines`, at any depth. +fn reference_labels<'a>(inlines: &'a [Inline], labels: &mut Vec<&'a str>) { + for inline in inlines { + let children = match inline { + Inline::LinkReference(node) => { + labels.push(&node.label); + &node.children + } + Inline::ImageReference(node) => { + labels.push(&node.label); + &node.alt + } + Inline::Link(node) => &node.children, + Inline::Image(node) => &node.alt, + Inline::InlineFootnote(node) => &node.children, + inline => match span_children(inline) { + Some(children) => children, + None => continue, + }, + }; + reference_labels(children, labels); + } +} + +/// `inlines` with every span cleared, which compare by their content. +fn without_spans(inlines: &[Inline]) -> Vec { + let mut inlines = inlines.to_vec(); + clear_spans(&mut inlines); + inlines +} + +fn clear_spans(inlines: &mut [Inline]) { + for inline in inlines { + let (meta, children) = match inline { + Inline::Text(node) => (&mut node.meta, None), + Inline::Escape(node) => (&mut node.meta, None), + Inline::SoftBreak(node) => (&mut node.meta, None), + Inline::LineBreak(node) => (&mut node.meta, None), + Inline::CharacterReference(node) => (&mut node.meta, None), + Inline::Emphasis(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Strong(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Underline(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Delete(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Insert(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Mark(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Subscript(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Superscript(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Spoiler(node) => (&mut node.meta, Some(&mut node.children)), + Inline::InlineFootnote(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Shortcode(node) => (&mut node.meta, None), + Inline::Code(node) => (&mut node.meta, None), + Inline::Link(node) => (&mut node.meta, Some(&mut node.children)), + Inline::Image(node) => (&mut node.meta, Some(&mut node.alt)), + Inline::LinkReference(node) => (&mut node.meta, Some(&mut node.children)), + Inline::ImageReference(node) => (&mut node.meta, Some(&mut node.alt)), + Inline::Autolink(node) => (&mut node.meta, None), + Inline::Html(node) => (&mut node.meta, None), + Inline::Math(node) => (&mut node.meta, None), + Inline::FootnoteReference(node) => (&mut node.meta, None), + Inline::WikiLink(node) => (&mut node.meta, None), + Inline::MdxExpression(node) => (&mut node.meta, None), + Inline::MdxJsx(node) => (&mut node.meta, None), + Inline::TextDirective(node) => (&mut node.meta, Some(&mut node.label)), + }; + meta.span = None; + if let Some(children) = children { + clear_spans(children); + } + } +} + +/// Keeps inline content opening with `import ` or `export `, which MDX reads +/// as ESM, a paragraph, by writing its first char as a reference. +fn keep_first_line_off_esm(output: &mut String) { + if output.starts_with("import ") || output.starts_with("export ") { + let reference = if output.starts_with('i') { + "i" + } else { + "e" + }; + output.replace_range(..1, reference); + } +} + +/// Keeps a first line holding only an MDX expression or JSX, which MDX reads +/// as a flow block, the paragraph's, by ending it with a referenced space. +fn keep_first_line_off_mdx_flow(children: &[Inline], output: &mut String) { + let value = match children.first() { + Some(Inline::MdxExpression(node)) => alloc::format!("{{{}}}", node.value), + Some(Inline::MdxJsx(node)) => node.value.clone(), + _ => return, + }; + if !output.starts_with(&value) { + return; + } + if output[value.len()..].starts_with('\n') { + output.insert_str(value.len(), " "); + } +} + +/// Keeps the first line of inline content that would start an HTML block +/// from starting one. Text gets an escape. Raw HTML takes none: raw HTML +/// opening such a line opens a paragraph only as the continuation of the one +/// a definition was read from, which `serialize_blocks_at_start` writes it +/// right after. +fn keep_first_line_off_html_block(children: &[Inline], output: &mut String) { + // An angle-bracket autolink that looks like an HTML block start comes + // from a dialect without raw HTML, which reads it back as written; raw + // HTML is written as it is, as above. + if matches!( + children.first(), + Some(Inline::Autolink(_) | Inline::Html(_)) + ) { + return; + } + if let Some(offset) = paragraph_html_block_escape_offset(output) { output.insert(offset, '\\'); } - Ok(output) +} + +/// Indents each continuation line of inline content that would start a block +/// (a line inside a code span, raw HTML, or a link title, which text escaping +/// does not reach) past a block start; the paragraph drops that indentation. +fn indent_block_starting_continuations(output: String) -> String { + let mut lines = output.split('\n'); + let Some(first) = lines.next() else { + return output; + }; + let mut previous = first; + let mut indented = None::; + let mut written = first.len(); + for line in lines { + if continuation_line_breaks_paragraph(previous, line) { + let result = indented.get_or_insert_with(|| String::from(&output[..written])); + result.push_str("\n "); + result.push_str(line); + } else if let Some(result) = indented.as_mut() { + result.push('\n'); + result.push_str(line); + } + written += 1 + line.len(); + previous = line; + } + indented.unwrap_or(output) } fn paragraph_html_block_escape_offset(input: &str) -> Option { @@ -348,25 +959,6 @@ fn paragraph_html_block_escape_offset(input: &str) -> Option { ) } -fn paragraph_table_escape_offset(input: &str) -> Option { - let first_line_end = input.find('\n')?; - let first_line = &input[..first_line_end]; - let second_line_start = first_line_end + 1; - let second_line_end = input[second_line_start..] - .find('\n') - .map(|offset| second_line_start + offset) - .unwrap_or(input.len()); - let second_line = &input[second_line_start..second_line_end]; - - if !gfm_table_can_start_source(first_line, second_line) { - return None; - } - - second_line - .find('-') - .map(|offset| second_line_start + offset) -} - fn serialize_alert(node: &Alert, options: &SerializeOptions) -> Result { let mut output = String::from("> [!"); output.push_str(alert_kind_name(node.kind)); @@ -395,16 +987,10 @@ fn alert_kind_name(kind: AlertKind) -> &'static str { } } +/// An alert title is kept as written, so only a line ending, which would end +/// its line, is written as a space. fn escape_alert_title(input: &str) -> String { - let mut output = String::new(); - for char in input.chars() { - match char { - '\n' | '\r' => output.push(' '), - char if char.is_control() => output.push_str(&format!("&#x{:X};", char as u32)), - _ => output.push(char), - } - } - output + input.replace(['\n', '\r'], " ") } fn serialize_list(node: &List, options: &SerializeOptions) -> Result { @@ -448,13 +1034,29 @@ fn serialize_list_with_marker_spacing( ) }; let mut inner = serialize_item_blocks(&item.children, options, node.tight)?; - if !node.ordered && unordered_list_marker(list_delimiter) == '*' { - inner = disambiguate_asterisk_list_item(inner); - } if let Some(checked) = item.checked { if let Some(rest) = inner.strip_prefix("- ") { inner = rest.into(); } + // The checkbox keeps the whitespace after it as text, so text + // opening with a space or tab, which a line's start would need as + // a reference, is written raw there. + if matches!(item.children.first(), Some(Block::Paragraph(paragraph)) + if matches!(paragraph.children.first(), Some(Inline::Text(text)) + if text.value.starts_with([' ', '\t']))) + { + for (reference, raw) in [(" ", " "), (" ", "\t")] { + // Content must follow on the line, or the paragraph's end + // would drop the whitespace. + if inner.starts_with(reference) + && !inner[reference.len()..].starts_with(['\n', ' ', '\t']) + && inner.len() > reference.len() + { + inner.replace_range(..reference.len(), raw); + break; + } + } + } let checkbox = if checked { "[x] " } else { "[ ] " }; inner = format!("{checkbox}{inner}"); } @@ -468,31 +1070,47 @@ fn serialize_list_with_marker_spacing( output.push_str(&prefix_lines(&inner, &" ".repeat(marker.len()))); continue; } + let first_line = inner.split('\n').next().unwrap_or(""); + let may_break = first_line.starts_with(['-', '*', '_']); + if inner.starts_with([' ', '\t']) + || (may_break && is_thematic_break_line(&format!("{marker}{first_line}"))) + { + // Content after the marker's padding would move the item's content + // column, so whitespace that opens the item's first block (an HTML + // block's indentation) starts on the line after the marker; so does + // a first line that would make the marker's line a thematic break. + // An item that starts blank has its content one column past the + // marker, whatever padding the other items use. + let content_indent = marker.trim_end().len() + 1; + output.push_str(marker.trim_end()); + output.push('\n'); + output.push_str(&prefix_lines(&inner, &" ".repeat(content_indent))); + continue; + } output.push_str(&marker); output.push_str(&indent_after_first_line(&inner, marker.len())); } Ok(output) } -fn disambiguate_asterisk_list_item(inner: String) -> String { - let first_line_end = inner.find('\n').unwrap_or(inner.len()); - let first_line = &inner[..first_line_end]; - if !asterisk_bullet_first_line_is_thematic_break(first_line) { - return inner; +/// Whether `line` is a thematic break: up to three spaces, then three or more +/// of one of `-`, `*`, `_`, with only spaces and tabs between and after them. +fn is_thematic_break_line(line: &str) -> bool { + let trimmed = line.trim_start_matches(' '); + if line.len() - trimmed.len() > 3 { + return false; } - let mut output = String::from("---"); - output.push_str(&inner[first_line_end..]); - output -} - -/// Whether a `*`-bullet item's first content line, once prefixed by the `* ` -/// marker, would escape the list as an asterisk thematic break. This is the -/// rendering of a `ThematicBreak` child: a contiguous run of asterisks (`***`, -/// rendered with no internal whitespace). A line with interior spaces such as -/// `* *` is a genuine nested bullet and must be left alone, since `* * *` -/// re-parses back into the nested list it came from. -fn asterisk_bullet_first_line_is_thematic_break(first_line: &str) -> bool { - first_line.len() >= 2 && first_line.bytes().all(|byte| byte == b'*') + let Some(marker) = trimmed + .chars() + .next() + .filter(|char| matches!(char, '-' | '*' | '_')) + else { + return false; + }; + trimmed + .chars() + .all(|char| matches!(char, ' ' | '\t') || char == marker) + && trimmed.chars().filter(|char| *char == marker).count() >= 3 } fn serialize_item_blocks( @@ -500,8 +1118,25 @@ fn serialize_item_blocks( options: &SerializeOptions, tight: bool, ) -> Result { + // Written last to first, as at the top level: a list reads the + // indentation of the block after it. + let mut written: Vec = Vec::with_capacity(blocks.len()); + for block in blocks.iter().rev() { + let next_indent = written + .last() + .map(|next: &String| next.len() - next.trim_start_matches(' ').len()); + written.push(match (block, next_indent) { + (Block::List(list), Some(indent @ 1..=3)) => { + serialize_list_with_marker_spacing(list, options, &" ".repeat(indent), " ")? + } + (Block::List(list), Some(4..)) => { + serialize_list_with_marker_spacing(list, options, " ", " ")? + } + _ => serialize_block(block, options, false)?, + }); + } let mut output = String::new(); - for (index, block) in blocks.iter().enumerate() { + for (index, (block, written)) in blocks.iter().zip(written.iter().rev()).enumerate() { if index > 0 { if tight { output.push('\n'); @@ -509,7 +1144,16 @@ fn serialize_item_blocks( output.push_str("\n\n"); } } - output.push_str(&serialize_block(block, options, false)?); + output.push_str(written); + if tight + && matches!(block, Block::BlockQuote(_) | Block::Alert(_)) + && matches!(blocks.get(index + 1), Some(Block::Paragraph(_))) + { + // The paragraph's first line would continue the quote's or + // alert's paragraph lazily; an empty quote line ends that + // paragraph first. + output.push_str("\n>"); + } } Ok(output) } @@ -542,7 +1186,11 @@ fn serialize_description_list( output.push('\n'); } output.push_str("\n:"); - let inner = serialize_blocks_at_start(&detail.children, options, false)?; + let inner = if node.tight { + serialize_item_blocks(&detail.children, options, true)? + } else { + serialize_blocks_at_start(&detail.children, options, false)? + }; if !inner.is_empty() { output.push('\n'); output.push_str(&indent_lines(&inner, 4)); @@ -557,27 +1205,50 @@ fn serialize_code_block( options: &SerializeOptions, ) -> Result { match node.kind { - CodeBlockKind::Indented => Ok(prefix_lines(trim_trailing_newline(&node.value), " ")), + CodeBlockKind::Indented => { + // Each value line ends with a line ending; the block gap or the + // final newline writes a last `\n`, so only `\r` and `\r\n` stay. + let body = trim_trailing_newline(&node.value); + let mut output = prefix_lines(body, " "); + let ending = &node.value[body.len()..]; + if matches!(ending, "\r" | "\r\n") { + output.push_str(ending); + } + Ok(output) + } CodeBlockKind::Fenced { marker, length } => { let marker = code_block_fence_marker(node, marker, options); - let fence = fence_for(&node.value, marker, length.max(3)); + let (fence, indent) = code_block_fence(&node.value, marker, length.max(3)); let mut opener = fence.clone(); if let Some(info) = &node.info { opener.push(' '); opener.push_str(&escape_code_info(info)); } - let mut output = opener; - output.push('\n'); - output.push_str(&node.value); - if !ends_with_line_ending(&node.value) { - output.push('\n'); - } - output.push_str(&fence); - Ok(output) + let body = fenced_body(&opener, &node.value, &fence); + Ok(if indent == 0 { + body + } else { + prefix_lines(&body, &" ".repeat(indent)) + }) } } } +/// A fenced block: `opener`, then `value` (whose lines each keep their line +/// ending, the last one optionally), then `closer`. An empty value writes no +/// line between the fences. +fn fenced_body(opener: &str, value: &str, closer: &str) -> String { + let mut output = String::with_capacity(opener.len() + value.len() + closer.len() + 2); + output.push_str(opener); + output.push('\n'); + output.push_str(value); + if !value.is_empty() && !ends_with_line_ending(value) { + output.push('\n'); + } + output.push_str(closer); + output +} + fn code_block_fence_marker( node: &CodeBlock, marker: FenceMarker, @@ -595,8 +1266,15 @@ fn code_block_fence_marker( fn escape_code_info(input: &str) -> String { let mut output = String::new(); - for char in input.chars() { + // The parser trims the info string, so whitespace at either end is + // written as a character reference. + let inner_start = input.len() - input.trim_start_matches([' ', '\t']).len(); + let inner_end = input.trim_end_matches([' ', '\t']).len().max(inner_start); + for (offset, char) in input.char_indices() { match char { + ' ' | '\t' if offset < inner_start || offset >= inner_end => { + output.push_str(&format!("&#x{:X};", char as u32)); + } '\n' => output.push_str(" "), '\r' => output.push_str(" "), '\t' => output.push(char), @@ -650,20 +1328,184 @@ fn serialize_table_row( options, InlineSerializeContext::table_cell(), )?; - if table_cell_has_unescaped_pipe(&cell) { - return Err(SerializeError::UnsupportedNode( - "table cell inline contains a pipe that cannot be escaped without changing source", - )); - } - cells.push(cell); + cells.push(escape_cell_delimiter_pipes(cell)); } Ok(format!("| {} |", cells.join(" | "))) } -#[derive(Clone, Copy, Debug, Default, Eq, PartialEq)] +/// How strong and emphasis runs choose between `*` and `_`. +#[derive(Clone, Copy, Debug, Default, Eq, Ord, PartialEq, PartialOrd)] +enum RunStyle { + /// Each run reads its neighbours: `_` where a `*` would join a run beside + /// it, `*` otherwise. + #[default] + Plain, + /// As `Plain`, and a strong whose content opens or closes with a `*` run + /// is written `__`. + StrongUnderscore, + /// As `StrongUnderscore`, and so is a strong that opens or closes the run + /// around it. + EdgeStrongUnderscore, + /// Every run is written with `*` where nothing before it would join it. + AllStar, + /// An emphasis with no strong or emphasis inside is written with `_`, and + /// every other run with `*`. + InnerUnderscore, + /// A run inside no other is written with `_`, and every other with `*`. + OuterUnderscore, +} + +#[derive(Clone, Copy, Debug, Default, Eq, Ord, PartialEq, PartialOrd)] struct InlineSerializeContext { table_cell: bool, avoid_star_edges: bool, + /// Inside an `_`-delimited emphasis, where a `_` in text that can close + /// would close it on reparse. + in_underscore_emphasis: bool, + /// The inlines open a line of the block, rather than following a + /// delimiter such as a link's `[` on it. + opens_line: bool, + /// For a text, the delimiter chars that the inlines after it may write; + /// for nested inlines, those after their parent, at every level. + written_later: DelimiterChars, + /// The same for the inlines before. + written_before: DelimiterChars, + /// For a text, whether it opens a line of the block. + text_opens_line: bool, + /// How strong and emphasis runs choose their char (see + /// `serialize_paragraph`). + run_style: RunStyle, + /// The inlines open a span delimited by a run such as `*` or `++`, which + /// a line ending right after it could not open. + opens_span: bool, + /// The char whose run opening or closing a text beside a strong or + /// emphasis delimiter is written raw (see `serialize_paragraph`). + raw_edge: Option, + /// How spaces and tabs around a literal autolink are written. + autolink_edges: AutolinkEdges, + /// A reference before the inlines, at any level, wrote a backtick in its + /// raw label, which an escaped backtick after it could close as a code + /// span. + raw_backtick_before: bool, + /// The delimiter chars of the spans around the inlines, at every level. + inside: DelimiterChars, + /// The delimiter char of the span whose content the inlines are. + enclosed: Option, + /// For a text, the `+` or `=` of a `++` or `==` delimiter written right + /// before and right after it. + text_edges: (Option, Option), +} + +/// A set of the chars `*`, `_`, `~`, `+`, `=`, `^`, `|`, `$`, and `:`, which +/// delimit inline spans or shortcodes that pair across sibling inlines, and +/// `>`, which ends raw HTML or an autolink that a `<` before it may open. +#[derive(Clone, Copy, Debug, Default, Eq, Ord, PartialEq, PartialOrd)] +struct DelimiterChars(u16); + +impl DelimiterChars { + const CHARS: [char; 11] = ['*', '_', '~', '+', '=', '^', '|', '$', ':', '>', '}']; + + /// The set holding `char` when it is one of [`Self::CHARS`], in order. + const fn of_char(char: char) -> Self { + Self(match char { + '*' => 1, + '_' => 1 << 1, + '~' => 1 << 2, + '+' => 1 << 3, + '=' => 1 << 4, + '^' => 1 << 5, + '|' => 1 << 6, + '$' => 1 << 7, + ':' => 1 << 8, + '>' => 1 << 9, + '}' => 1 << 10, + _ => 0, + }) + } + + fn of_str(input: &str) -> Self { + // Every char of the set is ASCII, so the bytes suffice. + input.bytes().fold(Self(0), |set, byte| { + set.union(Self::of_char(char::from(byte))) + }) + } + + const fn union(self, other: Self) -> Self { + Self(self.0 | other.0) + } + + fn contains(self, char: char) -> bool { + let bit = Self::of_char(char).0; + bit != 0 && self.0 & bit == bit + } + + /// The `>` and `$` in raw text, which can close what a char before it + /// opens. + fn raw_closers(raw: &str) -> Self { + let mut set = Self(0); + for char in ['>', '$'] { + if raw.contains(char) { + set = set.union(Self::of_char(char)); + } + } + set + } + + /// The delimiter chars that `inline` may write that can pair with a run + /// before it: a link's or image's text pairs only within its brackets. + fn written_by(inline: &Inline) -> Self { + let within = |own: &str, children: &[Inline]| { + children.iter().fold(Self::of_str(own), |set, child| { + set.union(Self::written_by(child)) + }) + }; + match inline { + Inline::Text(node) => Self::of_str(&node.value), + Inline::Emphasis(node) => within("*_", &node.children), + Inline::Strong(node) => within("*_", &node.children), + Inline::Underline(node) => within("_", &node.children), + Inline::Delete(node) => within("~", &node.children), + Inline::Insert(node) => within("+", &node.children), + Inline::Mark(node) => within("=", &node.children), + Inline::Subscript(node) => within("~", &node.children), + Inline::Superscript(node) => within("^", &node.children), + Inline::Spoiler(node) => within("|", &node.children), + // Their contents pair with nothing outside them; only a `>` in + // them can end raw HTML that a `<` before them opens, and a `$` + // close math that a `$` before them opens, as a math span's own + // `$` fence can. + Inline::Math(MathInline { value, .. }) => { + Self::of_char('$').union(Self::raw_closers(value)) + } + Inline::Html(HtmlInline { value: raw, .. }) | Inline::Code(CodeInline { raw, .. }) => { + Self::raw_closers(raw) + } + // A footnote's `^` can close a superscript before it; its label + // or content pairs with nothing outside it but a `<`. + Inline::FootnoteReference(node) => { + let caret = Self::of_char('^'); + if node.label.contains('>') { + caret.union(Self::of_char('>')) + } else { + caret + } + } + Inline::InlineFootnote(node) => { + let caret = Self::of_char('^'); + if within("", &node.children).contains('>') { + caret.union(Self::of_char('>')) + } else { + caret + } + } + // A URL's chars are scanned after the delimiters before it pair. + Inline::Autolink(node) => Self::of_str(&node.destination).union(Self::of_char('>')), + // A wiki link's text can close what a char before it opens. + Inline::WikiLink(node) => Self::of_str(&node.target).union(Self::of_str(&node.label)), + Inline::Shortcode(_) | Inline::TextDirective(_) => Self::of_char(':'), + _ => Self(0), + } + } } impl InlineSerializeContext { @@ -671,13 +1513,85 @@ impl InlineSerializeContext { Self { table_cell: true, avoid_star_edges: false, + in_underscore_emphasis: false, + opens_line: false, + written_later: DelimiterChars(0), + written_before: DelimiterChars(0), + text_opens_line: false, + run_style: RunStyle::Plain, + opens_span: false, + raw_edge: None, + autolink_edges: AutolinkEdges::Plain, + raw_backtick_before: false, + inside: DelimiterChars(0), + enclosed: None, + text_edges: (None, None), + } + } + + /// The context of a span's content, delimited by `delimiter` on both + /// sides. + fn delimited_by(self, delimiter: char) -> Self { + let delimiter = DelimiterChars::of_char(delimiter); + Self { + written_later: self.written_later.union(delimiter), + written_before: self.written_before.union(delimiter), + ..self + } + } + + const fn opening_span(self) -> Self { + Self { + opens_span: true, + ..self + } + } + + const fn block_content() -> Self { + Self { + table_cell: false, + avoid_star_edges: false, + in_underscore_emphasis: false, + opens_line: true, + written_later: DelimiterChars(0), + written_before: DelimiterChars(0), + text_opens_line: false, + run_style: RunStyle::Plain, + opens_span: false, + raw_edge: None, + autolink_edges: AutolinkEdges::Plain, + raw_backtick_before: false, + inside: DelimiterChars(0), + enclosed: None, + text_edges: (None, None), } } + /// Whether the inlines sit inside a strong or emphasis. + const fn inside_run(self) -> bool { + self.avoid_star_edges || self.in_underscore_emphasis + } + const fn avoiding_star_edges(self) -> Self { Self { - table_cell: self.table_cell, avoid_star_edges: true, + ..self + } + } + + /// For the content of a span delimited by `delimiter`. + fn enclosed_by(self, delimiter: char) -> Self { + Self { + inside: self.inside.union(DelimiterChars::of_char(delimiter)), + enclosed: Some(delimiter), + ..self + } + } + + const fn inside_underscore_emphasis(self) -> Self { + Self { + in_underscore_emphasis: true, + ..self } } } @@ -686,11 +1600,11 @@ fn serialize_inlines( inlines: &[Inline], options: &SerializeOptions, ) -> Result { - serialize_inlines_with_context(inlines, options, InlineSerializeContext::default()) + serialize_inlines_with_context(inlines, options, InlineSerializeContext::block_content()) } /// Escape a trailing unescaped `!` already in `output` before emitting a -/// following `[`-starting node (link / reference / footnote), so the pair does +/// following `[`-starting node (link / reference / footnote / wikilink), so the pair does /// not reparse as an image (`![…]`). The within-text `!`-before-`[` escaper /// only sees a single text node, so this handles the cross-node boundary. fn escape_trailing_bang(output: &mut String) { @@ -722,15 +1636,16 @@ fn escape_trailing_email_local(output: &mut String) { let Some(last) = output.chars().next_back() else { return; }; - // Only an ASCII alphanumeric immediately before the email forces a leftward + // An email-local char immediately before the email forces a leftward // re-anchor on reparse (the local part starts at the leftmost local-char - // run, and the dispatch reaches that alnum first). The punctuation - // local-chars (`.+-_`) are handled by the text serializer's own escaping. - if !last.is_ascii_alphanumeric() { - return; + // run): an ASCII alphanumeric is written as a reference, and an unescaped + // `.`, `+`, `-`, or `_` takes a backslash. + if last.is_ascii_alphanumeric() { + output.pop(); + output.push_str(&alloc::format!("&#{};", last as u32)); + } else if matches!(last, '.' | '+' | '-' | '_') && ends_with_unescaped(output, last) { + output.insert(output.len() - 1, '\\'); } - output.pop(); - output.push_str(&alloc::format!("&#{};", last as u32)); } // True when `inline` is a GFM literal autolink (its raw URL serialization can @@ -742,6 +1657,33 @@ fn is_gfm_literal_autolink(inline: &Inline) -> bool { ) } +// True when `inline` is a shortcut link or image reference or a footnote +// reference, whose `[label]` a following `(` would turn into an inline link, +// and a following `:` at the start of a line into a definition. +fn is_shortcut_reference(inline: &Inline) -> bool { + matches!( + inline, + Inline::LinkReference(LinkReference { + kind: ReferenceKind::Shortcut, + .. + }) | Inline::ImageReference(ImageReference { + kind: ReferenceKind::Shortcut, + .. + }) | Inline::FootnoteReference(_) + ) +} + +// The escaped first character of text after a shortcut reference, when that +// character would re-read the reference's brackets: a `(` always, and a `:` +// when the reference opens the inline sequence, where a line can start. +fn escape_leading_char_after_shortcut(value: &str, reference_index: usize) -> Option<&str> { + match value.as_bytes().first() { + Some(b'(') => Some("\\("), + Some(b':') if reference_index == 0 => Some("\\:"), + _ => None, + } +} + fn is_gfm_literal_email(inline: &Inline) -> bool { matches!( inline, @@ -751,22 +1693,75 @@ fn is_gfm_literal_email(inline: &Inline) -> bool { ) } -// A GFM literal autolink's URL scan stops at whitespace, `<`, `]`, and a -// backslash-escaped punctuation char, and trims trailing punctuation/entities. -// A following text node whose first char is none of those — in particular a -// non-ASCII char such as `©` decoded from `©` — would otherwise be pulled -// into the URL on reparse. Re-emit that leading char as a hex numeric character -// reference (`&#xNN;`), which `autolink_delim` trims back off the URL and which -// decodes to the same text, keeping the boundary stable. -fn encode_leading_char_after_autolink(value: &str) -> Option<(String, &str)> { - let first = value.chars().next()?; - if first.is_ascii() { - // ASCII merge chars are handled by the text serializer's own backslash - // escaping (`\[`, `\&`, …) and the parser's matching `\` stop. - return None; +/// The source spelling of a GFM literal autolink. +fn literal_autolink_original(inline: &Inline) -> Option<&str> { + match inline { + Inline::Autolink(Autolink { + kind: AutolinkKind::GfmLiteral { original }, + .. + }) => Some(original), + _ => None, + } +} + +/// The spelling of the literal autolink that `inline` is, or that its last +/// child is, through spans. +fn last_literal_autolink(inline: &Inline) -> Option<&str> { + if let Some(original) = literal_autolink_original(inline) { + return Some(original); + } + let children = match inline { + Inline::Emphasis(node) => &node.children, + Inline::Strong(node) => &node.children, + Inline::Underline(node) => &node.children, + Inline::Delete(node) => &node.children, + Inline::Insert(node) => &node.children, + Inline::Mark(node) => &node.children, + Inline::Subscript(node) => &node.children, + Inline::Superscript(node) => &node.children, + Inline::Spoiler(node) => &node.children, + _ => return None, + }; + last_literal_autolink(children.last()?) +} + +/// How `inline` is written when that takes no escaping: a literal autolink or +/// a shortcode, which a literal autolink's URL scan may run on into; empty +/// for any other inline. +fn plain_spelling(inline: &Inline) -> String { + match inline { + Inline::Shortcode(node) => alloc::format!(":{}:", node.name), + _ => literal_autolink_original(inline).map_or(String::new(), String::from), + } +} + +/// Whether `rendered` text written right after the literal autolink +/// `original`, and before `following` (the [`plain_spelling`] of the next +/// inline), leaves the autolink as it is under at least one +/// autolink dialect (GFM, or GFM plus relaxed): the parser's own scan reads +/// exactly `original`. The dialect the document came from is not known here; +/// the source spelling keeps it under that one. +fn text_keeps_literal_autolink(original: &str, rendered: &str, following: &str) -> bool { + let mut joined = String::from(original); + joined.push_str(rendered); + joined.push_str(following); + literal_autolink_extents(&joined).contains(&Some(original.len())) +} + +/// Spellings of `value`'s first char that a literal autolink's URL scan may +/// stop at, most readable first: a backslash escape for ASCII punctuation, then +/// a character reference, which the scan trims back off the URL's end. +fn leading_char_encodings(value: &str) -> Vec<(String, &str)> { + let Some(first) = value.chars().next() else { + return Vec::new(); + }; + let rest = &value[first.len_utf8()..]; + let mut encodings = Vec::new(); + if first.is_ascii_punctuation() { + encodings.push((alloc::format!("\\{first}"), rest)); } - let encoded = alloc::format!("&#x{:X};", first as u32); - Some((encoded, &value[first.len_utf8()..])) + encodings.push((alloc::format!("&#x{:X};", first as u32), rest)); + encodings } fn serialize_inlines_with_context( @@ -774,25 +1769,158 @@ fn serialize_inlines_with_context( options: &SerializeOptions, context: InlineSerializeContext, ) -> Result { + render_inlines(&mut RenderMemo::default(), inlines, options, context) +} + +/// Renders an emphasis has already made of its content, by the content's +/// address and the context it was rendered in. An emphasis may render its +/// content in two contexts to choose its delimiter; reusing the renders keeps +/// nested emphases from doubling the work at every level, since the contexts +/// a run can be rendered in are few. +#[derive(Default)] +struct RenderMemo(BTreeMap<(usize, usize, InlineSerializeContext), String>); + +impl RenderMemo { + fn render( + &mut self, + inlines: &[Inline], + options: &SerializeOptions, + context: InlineSerializeContext, + ) -> Result { + // Only content holding a span can double its work by nesting; + // rendering other content twice is cheaper than keeping it. + let nests = inlines.iter().any(|inline| { + span_children(inline).is_some() + || matches!( + inline, + Inline::Image(_) | Inline::ImageReference(_) | Inline::TextDirective(_) + ) + }); + if !nests { + return render_inlines(self, inlines, options, context); + } + let key = (inlines.as_ptr() as usize, inlines.len(), context); + if let Some(rendered) = self.0.get(&key) { + return Ok(rendered.clone()); + } + let rendered = render_inlines(self, inlines, options, context)?; + self.0.insert(key, rendered.clone()); + Ok(rendered) + } +} + +fn render_inlines( + memo: &mut RenderMemo, + inlines: &[Inline], + options: &SerializeOptions, + context: InlineSerializeContext, +) -> Result { + let opens_line = context.opens_line; + let opens_span = context.opens_span; + let enclosed = context.enclosed; + // Nested inlines follow their parent's opening delimiter. + let base_context = InlineSerializeContext { + opens_line: false, + text_opens_line: false, + opens_span: false, + enclosed: None, + text_edges: (None, None), + ..context + }; let mut output = String::new(); let mut output_line = OutputLine::default(); + // The delimiter chars the inlines from each index on may write, and those + // before each index, within this run and around it. A run of plain text + // and breaks needs no more than the chars around it, so its inlines are + // read only when a text holds such a char or an inline holds others. + let needs_written = inlines.len() > 1 + && inlines.iter().any(|inline| match inline { + Inline::Text(node) => node.value.contains(DelimiterChars::CHARS), + Inline::SoftBreak(_) | Inline::LineBreak(_) => false, + _ => true, + }); + // Each inline's own chars, and the chars from each index on. + let written = needs_written.then(|| { + let own = inlines + .iter() + .map(DelimiterChars::written_by) + .collect::>(); + let mut from = vec![context.written_later; inlines.len() + 1]; + for index in (0..inlines.len()).rev() { + from[index] = from[index + 1].union(own[index]); + } + (own, from) + }); + let mut written_until = context.written_before; + let mut raw_backtick_before = context.raw_backtick_before; + // Where the output of the inline before the current one starts. + let mut segment_start = 0; for (index, inline) in inlines.iter().enumerate() { + let written_before = written_until; + let previous_start = segment_start; + segment_start = output.len(); + let written_later = match &written { + Some((own, from)) => { + written_until = written_until.union(own[index]); + from[index + 1] + } + None => context.written_later, + }; + let context = InlineSerializeContext { + written_later, + written_before, + raw_backtick_before, + ..base_context + }; match inline { Inline::Text(node) => { - let after_literal_autolink = index - .checked_sub(1) - .is_some_and(|prev| is_gfm_literal_autolink(&inlines[prev])); - let before_literal_autolink = - inlines.get(index + 1).is_some_and(is_gfm_literal_autolink); - - // Leading guard: a non-ASCII char abutting the END of a literal - // autolink would merge into its URL on reparse — encode it. - let (lead, body) = match after_literal_autolink - .then(|| encode_leading_char_after_autolink(&node.value)) - .flatten() - { - Some((encoded, rest)) => (encoded, rest), - None => (String::new(), node.value.as_str()), + // A literal autolink just before the text, ending this run's + // previous inline or the last span inside it: its spelling, + // and the span delimiters written after it. + let autolink_before = index.checked_sub(1).and_then(|prev| { + let original = last_literal_autolink(&inlines[prev])?; + // The URL may itself end with a delimiter char, so each + // split of the trailing delimiters is tried, shortest tail + // first. + let segment = &output[previous_start..]; + let delimiters = segment.len() + - segment + .trim_end_matches(['*', '_', '~', '=', '+', '^', '|']) + .len(); + (0..=delimiters).find_map(|tail| { + let body = &segment[..segment.len() - tail]; + body.ends_with(original) + .then(|| (original, &segment[segment.len() - tail..])) + }) + }); + // A text directive opens only after whitespace, so the + // whitespace before one stays raw too. + let before_literal_autolink = context.autolink_edges + != AutolinkEdges::EncodedBefore + && inlines.get(index + 1).is_some_and(|next| { + is_gfm_literal_autolink(next) || matches!(next, Inline::TextDirective(_)) + }); + let raw_edges = context.autolink_edges == AutolinkEdges::RawEdges; + let at_line_start = output_line.len(&output) == 0; + let opens_block_line = breaks_line_start(&output, &mut output_line, opens_line); + let at_line_end = text_is_at_line_end(inlines, index); + let doubled_delimiter = |inline: &Inline| match inline { + Inline::Insert(_) => Some('+'), + Inline::Mark(_) => Some('='), + _ => None, + }; + let edge_before = match index.checked_sub(1) { + Some(previous) => doubled_delimiter(&inlines[previous]), + None => enclosed.filter(|_| opens_span), + }; + let edge_after = match inlines.get(index + 1) { + Some(next) => doubled_delimiter(next), + None => enclosed, + }; + let text_context = InlineSerializeContext { + text_opens_line: opens_block_line, + text_edges: (edge_before, edge_after), + ..context }; // Trailing guard: when this text is immediately followed by a @@ -801,23 +1929,163 @@ fn serialize_inlines_with_context( // otherwise re-encoded (` `/` `) at an edge or as a // control char, which would break the literal's left boundary on // reparse — emit the trailing space/tab run literally instead. - let (escape_body, trailing_ws) = if before_literal_autolink { + let render = |lead: &str, body: &str| { let head = body.trim_end_matches([' ', '\t']); - (head, &body[head.len()..]) - } else { - (body, "") + // Whitespace that opens a line or a table cell stays + // encoded: written literally, the line or cell would drop + // it. + let cell_start = context.table_cell && !opens_span && index == 0; + let whole_line_start = + (opens_block_line || cell_start) && lead.is_empty() && head.is_empty(); + let (escape_body, trailing_ws) = if before_literal_autolink && !whole_line_start + { + (head, &body[head.len()..]) + } else { + (body, "") + }; + let mut rendered = String::with_capacity(lead.len() + body.len() + 8); + rendered.push_str(lead); + rendered.push_str(&escape_text_with_context( + escape_body, + !raw_edges + && lead.is_empty() + && trailing_ws.len() != body.len() + && at_line_start, + !raw_edges && trailing_ws.is_empty() && at_line_end, + text_context, + )); + rendered.push_str(trailing_ws); + rendered }; - output.push_str(&lead); - output.push_str(&escape_text_with_context( - escape_body, - lead.is_empty() - && trailing_ws.len() != body.len() - && output_line.len(&output) == 0, - trailing_ws.is_empty() && text_is_at_line_end(inlines, index), - context, - )); - output.push_str(trailing_ws); + let after_shortcut = index + .checked_sub(1) + .filter(|&prev| is_shortcut_reference(&inlines[prev])); + let mut rendered = match after_shortcut.and_then(|reference| { + escape_leading_char_after_shortcut(&node.value, reference) + }) { + Some(escaped) => render(escaped, &node.value[1..]), + None => render("", &node.value), + }; + if let Some(edge) = context.raw_edge { + let at_start = (index == 0 && opens_span) + || index + .checked_sub(1) + .is_some_and(|previous| is_attention_run(&inlines[previous])); + let at_end = (index + 1 == inlines.len() && opens_span) + || inlines.get(index + 1).is_some_and(is_attention_run); + if at_start || at_end { + rendered = unescape_edge(&rendered, edge, at_start, at_end); + } + } + // Leading guard: text right after a literal autolink must not + // extend its URL on reparse. When it would, its first char is + // written in the first form the URL scan stops at. + if let Some((original, tail)) = autolink_before { + let following = inlines.get(index + 1).map_or(String::new(), plain_spelling); + let keeps = |rendered: &str| { + text_keeps_literal_autolink( + original, + &format!("{tail}{rendered}"), + &following, + ) + }; + let encode_lead = context.autolink_edges == AutolinkEdges::EncodedLead + && (index + 1 < inlines.len() || opens_span); + if encode_lead || !keeps(&rendered) { + for (lead, rest) in leading_char_encodings(&node.value) { + let candidate = render(&lead, rest); + if keeps(&candidate) { + rendered = candidate; + break; + } + } + } + } + // A scheme char ending the text would join the scheme of a + // literal autolink after it (`://x` or `p://x` under the + // relaxed dialect). + if inlines + .get(index + 1) + .and_then(literal_autolink_original) + .is_some_and(|original| { + let scheme = original + .bytes() + .take_while(|byte| { + byte.is_ascii_alphanumeric() || matches!(byte, b'+' | b'.' | b'-') + }) + .count(); + original[scheme..].starts_with("://") + }) + { + // A relaxed scheme opens with a letter, so a run of + // scheme chars without one joins nothing. + let scheme_run = &rendered[rendered + .trim_end_matches(|char: char| { + char.is_ascii_alphanumeric() || matches!(char, '+' | '.' | '-') + }) + .len()..]; + if let Some(last) = rendered.chars().next_back().filter(|char| { + (char.is_ascii_alphanumeric() || matches!(char, '+' | '.' | '-')) + && scheme_run.bytes().any(|byte| byte.is_ascii_alphabetic()) + }) { + if !ends_with_unescaped(&rendered, last) || last.is_ascii_alphanumeric() { + if last.is_ascii_alphanumeric() { + rendered.pop(); + rendered.push_str(&alloc::format!("&#x{:X};", last as u32)); + } + } else { + rendered.insert(rendered.len() - 1, '\\'); + } + } + } + // An escaped backtick still closes a code span that a raw + // backtick before it opens; a character reference does not. + if raw_backtick_before && rendered.contains("\\`") { + rendered = rendered.replace("\\`", "`"); + } + // A `:` opening the text would close a shortcode that a bare + // text directive before it opens. + if index.checked_sub(1).is_some_and(|prev| { + matches!(&inlines[prev], Inline::TextDirective(directive) + if directive.label.is_empty() && directive.attributes.is_empty()) + }) && rendered.starts_with(':') + { + rendered.insert(0, '\\'); + } + // An `@` ending the text would open an email whose domain the + // literal autolink after it writes. + if inlines.get(index + 1).is_some_and(is_gfm_literal_autolink) + && rendered.ends_with('@') + { + rendered.pop(); + rendered.push_str("@"); + } + // A `:` ending the text would open a shortcode that a `:` in + // the literal autolink after it closes, or that a span after it, + // or the end of the span around it, names with its `++` or `_` + // delimiters. + let names_shortcode = written_later.contains(':') + && match inlines.get(index + 1) { + Some( + Inline::Insert(_) + | Inline::Underline(_) + | Inline::Strong(_) + | Inline::Emphasis(_), + ) => true, + Some(_) => false, + None => matches!(enclosed, Some('+' | '_')), + }; + if (inlines + .get(index + 1) + .and_then(literal_autolink_original) + .is_some() + || names_shortcode) + && ends_with_unescaped(&rendered, ':') + { + rendered.insert(rendered.len() - 1, '\\'); + } + output.push_str(&rendered); } Inline::Escape(node) => { output.push('\\'); @@ -825,26 +2093,66 @@ fn serialize_inlines_with_context( } Inline::CharacterReference(node) => output.push_str(&node.reference), Inline::Emphasis(node) => { - let children = serialize_inlines_with_context(&node.children, options, context)?; - let touches_underscore = children.starts_with('_') - || children.ends_with('_') - || children.starts_with("\\_") - || children.ends_with("\\_"); + // Rendered as `_` content first: the choice below reads only + // the children's edges and `*`s, which escaping a `_` that can + // close does not change, so only the `*` choice renders them + // again and nesting never multiplies the work. + let children = memo.render( + &node.children, + options, + context + .inside_underscore_emphasis() + .opening_span() + .enclosed_by('_'), + )?; + // An escaped `_` at an edge joins no run. + let touches_underscore = + children.starts_with('_') || ends_with_unescaped(&children, '_'); // An emphasis abutting a `*` already in the output (e.g. a // preceding `*`-emphasis) would otherwise merge into one run, so // switch this run to `_` when that does not introduce a new // `_`-collision with the children. - let abuts_star = output.ends_with('*') && !touches_underscore; - let prefer_underscore = (context.avoid_star_edges && !touches_underscore) - || abuts_star - || children.starts_with('*') - || children.ends_with('*'); + // A raw edge `*` is meant to join the run. + let abuts_star = ends_with_unescaped(&output, '*') + && !touches_underscore + && context.raw_edge != Some('*'); + // `_` neither opens after nor closes before an alphanumeric. + let underscore_flanks = !output + .chars() + .next_back() + .is_some_and(char::is_alphanumeric) + && !matches!( + inlines.get(index + 1), + Some(Inline::Text(next)) if next.value.chars().next().is_some_and(char::is_alphanumeric) + ); + let innermost = !node + .children + .iter() + .any(|child| matches!(child, Inline::Strong(_) | Inline::Emphasis(_))); + let prefer_underscore = underscore_flanks + && match context.run_style { + RunStyle::Plain + | RunStyle::StrongUnderscore + | RunStyle::EdgeStrongUnderscore => { + (context.avoid_star_edges && !touches_underscore) + || abuts_star + || children.starts_with('*') + || children.ends_with('*') + } + RunStyle::AllStar => abuts_star, + RunStyle::InnerUnderscore => { + abuts_star || (innermost && !touches_underscore) + } + RunStyle::OuterUnderscore => { + abuts_star || (!context.inside_run() && !touches_underscore) + } + }; let delimiter = if prefer_underscore { '_' } else { '*' }; let children = if delimiter == '*' { - serialize_inlines_with_context( + memo.render( &node.children, options, - context.avoiding_star_edges(), + context.avoiding_star_edges().opening_span(), )? } else { children @@ -854,31 +2162,74 @@ fn serialize_inlines_with_context( output.push(delimiter); } Inline::Strong(node) => { - let children = serialize_inlines_with_context( + let children = render_inlines( + memo, &node.children, options, - context.avoiding_star_edges(), + context.avoiding_star_edges().opening_span(), )?; - // NOTE: two abutting `Strong` nodes (`**a****b**`) reparse as a - // single run. The only zero-insertion separator is flipping one - // run to `__`, but `__` reparses as `Underline` when that - // construct is enabled and the serializer has no signal for it, - // so this hand-built-AST sub-case is left as a known limitation. - output.push_str("**"); + // A `**` right after a closing `**` joins it into a run of + // four, which by the rule of three closes neither strong, so + // the strong is written with `__` there when `_` can flank + // and its content does not touch `_`. After a lone closing + // `*` the run of three splits as written, and `__` would read + // back as `Underline` where that construct is enabled, so + // only the read-back choices write `__` there, or at the edge + // of the run around the strong. + let raw_star_edge = context.raw_edge == Some('*'); + let after_strong = ends_with_unescaped(&output, '*') + && ends_with_unescaped(&output[..output.len() - 1], '*') + && !raw_star_edge; + let edge_of_run = + context.inside_run() && (index == 0 || index + 1 == inlines.len()); + let after_star = after_strong + || (matches!( + context.run_style, + RunStyle::StrongUnderscore | RunStyle::EdgeStrongUnderscore + ) && ((ends_with_unescaped(&output, '*') && !raw_star_edge) + || (edge_of_run && context.run_style == RunStyle::EdgeStrongUnderscore) + || children.starts_with('*') + || children.ends_with('*'))); + // A `_` opening the next text is escaped beside the run. + let underscore_fits = !children.starts_with('_') + && !ends_with_unescaped(&children, '_') + && !matches!( + inlines.get(index + 1), + Some(Inline::Text(next)) + if next.value.chars().next().is_some_and(char::is_alphanumeric) + ); + let outer_underscore = context.run_style == RunStyle::OuterUnderscore + && !context.inside_run() + && !output + .chars() + .next_back() + .is_some_and(char::is_alphanumeric); + let delimiter = if (after_star || outer_underscore) && underscore_fits { + "__" + } else { + "**" + }; + output.push_str(delimiter); output.push_str(&children); - output.push_str("**"); + output.push_str(delimiter); } Inline::Underline(node) => { output.push_str("__"); - output.push_str(&serialize_inlines_with_context( + output.push_str(&render_inlines( + memo, &node.children, options, - context, + context.opening_span().enclosed_by('_'), )?); output.push_str("__"); } Inline::Delete(node) => { - let children = serialize_inlines_with_context(&node.children, options, context)?; + let children = render_inlines( + memo, + &node.children, + options, + context.opening_span().delimited_by('~'), + )?; let marker = match node.marker { DeleteMarker::SingleTilde => "~", DeleteMarker::DoubleTilde => "~~", @@ -889,46 +2240,51 @@ fn serialize_inlines_with_context( } Inline::Insert(node) => { output.push_str("++"); - output.push_str(&serialize_inlines_with_context( + output.push_str(&render_inlines( + memo, &node.children, options, - context, + context.opening_span().delimited_by('+').enclosed_by('+'), )?); output.push_str("++"); } Inline::Mark(node) => { output.push_str("=="); - output.push_str(&serialize_inlines_with_context( + output.push_str(&render_inlines( + memo, &node.children, options, - context, + context.opening_span().delimited_by('=').enclosed_by('='), )?); output.push_str("=="); } Inline::Subscript(node) => { output.push('~'); - output.push_str(&serialize_inlines_with_context( + output.push_str(&render_inlines( + memo, &node.children, options, - context, + context.opening_span().delimited_by('~'), )?); output.push('~'); } Inline::Superscript(node) => { output.push('^'); - output.push_str(&serialize_inlines_with_context( + output.push_str(&render_inlines( + memo, &node.children, options, - context, + context.opening_span().delimited_by('^'), )?); output.push('^'); } Inline::Spoiler(node) => { output.push_str("||"); - output.push_str(&serialize_inlines_with_context( + output.push_str(&render_inlines( + memo, &node.children, options, - context, + context.opening_span().delimited_by('|'), )?); output.push_str("||"); } @@ -973,11 +2329,7 @@ fn serialize_inlines_with_context( Inline::Link(node) => { escape_trailing_bang(&mut output); output.push('['); - output.push_str(&serialize_inlines_with_context( - &node.children, - options, - context, - )?); + output.push_str(&render_inlines(memo, &node.children, options, context)?); output.push_str("]("); output.push_str(&serialize_destination_kind( &node.destination, @@ -992,9 +2344,7 @@ fn serialize_inlines_with_context( } Inline::Image(node) => { output.push_str("!["); - output.push_str(&serialize_inlines_with_context( - &node.alt, options, context, - )?); + output.push_str(&render_inlines(memo, &node.alt, options, context)?); output.push_str("]("); output.push_str(&serialize_destination_kind( &node.destination, @@ -1008,7 +2358,7 @@ fn serialize_inlines_with_context( output.push(')'); } Inline::LinkReference(node) => { - let children = serialize_inlines_with_context(&node.children, options, context)?; + let children = render_inlines(memo, &node.children, options, context)?; let children_identifier = normalize_reference_label(&children); escape_trailing_bang(&mut output); push_reference_body( @@ -1020,7 +2370,7 @@ fn serialize_inlines_with_context( ); } Inline::ImageReference(node) => { - let alt = serialize_inlines_with_context(&node.alt, options, context)?; + let alt = render_inlines(memo, &node.alt, options, context)?; let alt_identifier = normalize_reference_label(&alt); output.push('!'); push_reference_body( @@ -1049,7 +2399,12 @@ fn serialize_inlines_with_context( && index .checked_sub(1) .is_some_and(|prev| is_gfm_literal_email(&inlines[prev])); - if is_bare_email && !follows_literal_email_plus { + // A span's closing delimiter run is read before the + // email, so only a text's last char can join it. + let after_span = index + .checked_sub(1) + .is_some_and(|prev| span_children(&inlines[prev]).is_some()); + if is_bare_email && !follows_literal_email_plus && !after_span { escape_trailing_email_local(&mut output); } else { escape_trailing_less_than(&mut output); @@ -1058,9 +2413,27 @@ fn serialize_inlines_with_context( } }, Inline::Html(node) => output.push_str(&node.value), - Inline::SoftBreak(_) => output.push('\n'), + // A break that opens a line, as ` \n` and ` \n` parse + // (the whitespace before a line ending is dropped), is written the + // same way: a bare line ending or spaces there would end the block. + // So is one that opens a delimited span, whose opener a line ending + // right after it would keep from opening. + Inline::SoftBreak(_) => { + if breaks_line_start(&output, &mut output_line, opens_line) + || (opens_span && output.is_empty()) + { + output.push_str(" "); + } + output.push('\n'); + } Inline::LineBreak(node) => match node.kind { LineBreakKind::Backslash => output.push_str("\\\n"), + LineBreakKind::Spaces + if breaks_line_start(&output, &mut output_line, opens_line) + || (opens_span && output.is_empty()) => + { + output.push_str(" \n"); + } LineBreakKind::Spaces => output.push_str(" \n"), }, Inline::Math(node) => { @@ -1078,14 +2451,11 @@ fn serialize_inlines_with_context( } Inline::InlineFootnote(node) => { output.push_str("^["); - output.push_str(&serialize_inlines_with_context( - &node.children, - options, - context, - )?); + output.push_str(&render_inlines(memo, &node.children, options, context)?); output.push(']'); } Inline::WikiLink(node) => { + escape_trailing_bang(&mut output); output.push_str("[["); let target = escape_wikilink_part(&node.target); let label = escape_wikilink_part(&node.label); @@ -1125,12 +2495,69 @@ fn serialize_inlines_with_context( &node.attributes, context, )); + // What follows could go on with a name char, or with a `[` or + // `{` the directive could read as its label or attributes, so + // an empty label, or an empty attribute list, ends the + // directive, unless a break or a text that cannot follows. + if node.attributes.is_empty() { + let next = match inlines.get(index + 1) { + None | Some(Inline::SoftBreak(_) | Inline::LineBreak(_)) => None, + Some(Inline::Text(text)) => text.value.chars().next(), + // Another inline may open with any of them. + Some(_) => Some('a'), + }; + if node.label.is_empty() + && next.is_some_and(|char| { + char.is_ascii_alphanumeric() || matches!(char, '_' | '-' | '[' | '{') + }) + { + output.push_str("[]"); + } + if next == Some('{') { + output.push_str("{}"); + } + } } } + // A reference writes its raw label, whose backtick is unescaped; a + // span is judged by the labels of the references inside it, since + // its output also holds its code spans' backticks. + raw_backtick_before |= match inline { + Inline::FootnoteReference(_) | Inline::LinkReference(_) | Inline::ImageReference(_) => { + holds_unescaped_backtick(&output[segment_start..]) + } + other => holds_raw_label_backtick(other), + }; } Ok(output) } +/// Whether `written` holds a backtick no backslash escapes, which could open +/// a code span. +fn holds_unescaped_backtick(written: &str) -> bool { + written + .match_indices('`') + .any(|(index, _)| !ends_with_unescaped(&written[..index], '\\')) +} + +/// Whether `inline` is a reference whose raw label holds an unescaped +/// backtick, or a span holding one. +fn holds_raw_label_backtick(inline: &Inline) -> bool { + match inline { + Inline::FootnoteReference(node) => holds_unescaped_backtick(&node.label), + Inline::LinkReference(node) => { + holds_unescaped_backtick(&node.label) + || node.children.iter().any(holds_raw_label_backtick) + } + Inline::ImageReference(node) => { + holds_unescaped_backtick(&node.label) || node.alt.iter().any(holds_raw_label_backtick) + } + Inline::Image(node) => node.alt.iter().any(holds_raw_label_backtick), + other => span_children(other) + .is_some_and(|children| children.iter().any(holds_raw_label_backtick)), + } +} + fn serialize_directive_label( label: &[Inline], options: &SerializeOptions, @@ -1203,10 +2630,17 @@ fn is_directive_shorthand_value(input: &str) -> bool { .all(|char| char.is_ascii_alphanumeric() || matches!(char, '_' | '-')) } +/// Whether whitespace ending the text at `index` would be dropped: before a +/// line ending, two-space hard break, or the end of the run. A backslash hard +/// break keeps the whitespace before it. fn text_is_at_line_end(inlines: &[Inline], index: usize) -> bool { matches!( inlines.get(index + 1), - None | Some(Inline::SoftBreak(_)) | Some(Inline::LineBreak(_)) + None | Some(Inline::SoftBreak(_)) + | Some(Inline::LineBreak(LineBreak { + kind: LineBreakKind::Spaces, + .. + })) ) } @@ -1225,25 +2659,47 @@ struct TextScan<'a> { marker_starts: [Positions; ATTENTION_MARKERS.len()], marker_closers: [PathMemo; ATTENTION_MARKERS.len()], last_occurrences: Vec<(&'static str, Option)>, - backtick_runs: Option, dollar_runs: Option, /// The byte, start, and end of the last run `run_len_from` measured. current_run: Option<(u8, usize, usize)>, + /// The [`DelimiterChars`] that the inlines after the text may write. + written_later: DelimiterChars, + /// Those that the inlines before it may write. + written_before: DelimiterChars, } impl<'a> TextScan<'a> { - fn new(input: &'a str) -> Self { + #[cfg(test)] + fn new(input: &'a str, written_later: DelimiterChars) -> Self { + Self::around(input, written_later, DelimiterChars(0)) + } + + fn around( + input: &'a str, + written_later: DelimiterChars, + written_before: DelimiterChars, + ) -> Self { Self { input, marker_starts: Default::default(), marker_closers: Default::default(), last_occurrences: Vec::new(), - backtick_runs: None, dollar_runs: None, current_run: None, + written_later, + written_before, } } + /// Whether an inline after the text may write `marker`'s char, which a + /// delimiter opening in the text could pair with. + fn written_later(&self, marker: &str) -> bool { + marker + .chars() + .next() + .is_some_and(|char| self.written_later.contains(char)) + } + /// `same_char_run_len` for an ASCII `needle`, measuring each run once /// however many of its positions ask. fn run_len_from(&mut self, needle: u8, offset: usize) -> usize { @@ -1307,14 +2763,12 @@ impl<'a> TextScan<'a> { } /// Whether some position at or after `from` begins exactly `run_len` - /// trailing bytes of a run of `needle` (an ASCII byte). - fn exact_run_follows(&mut self, needle: u8, from: usize, run_len: usize) -> bool { + /// trailing bytes of a run of `$`. + fn exact_dollar_run_follows(&mut self, from: usize, run_len: usize) -> bool { let input = self.input; - let runs = match needle { - b'`' => &mut self.backtick_runs, - _ => &mut self.dollar_runs, - } - .get_or_insert_with(|| SameCharRuns::new(input, needle)); + let runs = self + .dollar_runs + .get_or_insert_with(|| SameCharRuns::new(input, b'$')); runs.has_run_ending_at_or_after(from + run_len, run_len) } } @@ -1367,17 +2821,29 @@ fn escape_text_with_context( context: InlineSerializeContext, ) -> String { let avoid_star_edges = context.avoid_star_edges; - let mut output = String::new(); + let in_underscore_emphasis = context.in_underscore_emphasis; + let mut output = String::with_capacity(input.len() + input.len() / 8); let mut output_line = OutputLine::default(); - let mut scan = TextScan::new(input); let mut line_digit_prefix = 0usize; - let trailing_start = if preserve_trailing { - input - .trim_end_matches(|char| matches!(char, ' ' | '\t')) - .len() + // Only the space or tab at a preserved edge is written as a reference: + // the ones beside it are no longer at the edge of the line or span. + let trailing_start = if preserve_trailing && input.ends_with([' ', '\t']) { + input.len() - 1 } else { input.len() }; + // Delimiter decisions read the text as the reparse sees it, where each + // char written as a character reference is punctuation. + let view = referenced_chars_as_punctuation(input, preserve_leading, trailing_start); + let view = view.as_ref(); + let mut scan = TextScan::around(view, context.written_later, context.written_before); + // The end of the current `$`, `*`, and `_` run, and whether it is escaped. + let mut dollar_run = (0usize, false); + let mut star_run = (0usize, false); + let mut underscore_run = (0usize, false); + let mut tilde_run = (0usize, false); + let mut pipe_run = (0usize, false); + let mut plus_run = (0usize, false); let mut chars = input.char_indices().peekable(); let mut at_leading_edge = preserve_leading; while let Some((offset, char)) = chars.next() { @@ -1393,13 +2859,16 @@ fn escape_text_with_context( } if (at_leading_edge || offset >= trailing_start) && char == ' ' { output.push_str(" "); + at_leading_edge = false; continue; } if (at_leading_edge || offset >= trailing_start) && char == '\t' { output.push_str(" "); + at_leading_edge = false; continue; } - if char.is_control() { + // A tab inside the text stays literal: the reparse keeps it as it is. + if written_as_reference(char) { output.push_str(&format!("&#x{:X};", char as u32)); at_leading_edge = false; continue; @@ -1410,10 +2879,9 @@ fn escape_text_with_context( line_digit_prefix += 1; continue; } - if char == ':' - && (input[..offset].ends_with("http") || input[..offset].ends_with("https")) - && input[offset + char.len_utf8()..].starts_with("//") - { + // `://` can open a literal autolink: with any scheme, or none under + // the relaxed autolink dialect. + if char == ':' && input[offset + char.len_utf8()..].starts_with("//") { output.push('\\'); output.push(char); line_digit_prefix = usize::MAX; @@ -1475,19 +2943,67 @@ fn escape_text_with_context( output.push('\\'); output.push(char); } - '`' if text_code_span_can_start(input, offset, &mut scan) => { + // A backslash keeps a backtick from opening a code span but not + // from closing one, so every backtick is escaped: a bare one could + // open a span that an escaped one closes. + '`' => { + output.push('\\'); + output.push(char); + } + // Escaping only part of a run would leave a shorter run, which + // flanks and pairs differently, so a run is escaped whole or not. + '*' if run_escaped(view, offset, b'*', &mut scan, &mut star_run, |scan, at| { + text_attention_delimiter_can_start(view, at, "*", false, scan) + }) => + { output.push('\\'); output.push(char); } - '*' if text_attention_delimiter_can_start(input, offset, "*", false, &mut scan) => { + '_' if run_escaped( + view, + offset, + b'_', + &mut scan, + &mut underscore_run, + |scan, at| { + (in_underscore_emphasis && text_delimiter_can_close(view, at, 1, true)) + || text_attention_delimiter_can_start(view, at, "_", true, scan) + }, + ) => + { output.push('\\'); output.push(char); } - '_' if text_attention_delimiter_can_start(input, offset, "_", true, &mut scan) => { + // A text that starts a line may sit on a paragraph's continuation + // line, where an HTML block start (types 1–6) or a directive + // opener would interrupt the paragraph. + '<' if output_line.len(&output) == 0 + && line_starts_interrupting_html_block(&view[offset..]) => + { output.push('\\'); output.push(char); } - '<' if text_less_than_can_start_inline(input, offset, &mut scan) => { + // A `:` or `~` and whitespace opening a line would open the + // details of a description list whose term is the line before. + ':' | '~' + if offset == 0 + && context.text_opens_line + && input[offset + 1..] + .chars() + .next() + .is_none_or(|next| matches!(next, ' ' | '\t')) => + { + output.push('\\'); + output.push(char); + } + ':' if (output_line.len(&output) == 0 && input[offset..].starts_with("::")) + || text_directive_can_start(view, offset) + || shortcode_can_form(view, offset, &mut scan) => + { + output.push('\\'); + output.push(char); + } + '<' if text_less_than_can_start_inline(view, offset, &mut scan) => { output.push('\\'); output.push(char); } @@ -1495,16 +3011,33 @@ fn escape_text_with_context( output.push('\\'); output.push(char); } - '{' if scan.occurs_from("}", offset + char.len_utf8()) => { + '{' if scan.written_later("}") || scan.occurs_from("}", offset + char.len_utf8()) => { output.push('\\'); output.push(char); } - '#' if text_atx_heading_can_start(input, offset, output_line.len(&output)) => { + '#' if text_atx_heading_can_start(view, offset, output_line.len(&output)) => { output.push('\\'); output.push(char); } - '|' if text_spoiler_can_start(input, offset, &mut scan) => output.push_str("|"), - '$' if text_math_can_start(input, offset, &mut scan) => { + // A run of two or more bars opens with its last two, so every bar + // but the last of a run that can open is written as a reference. + '|' if run_escaped(view, offset, b'|', &mut scan, &mut pipe_run, |scan, at| { + text_spoiler_can_start(view, at, scan) + }) && offset + 1 < pipe_run.0 => + { + output.push_str("|") + } + // Escaping only part of a `$` run would leave a shorter run that can + // open math, so a run is escaped whole or not at all. + '$' if run_escaped( + view, + offset, + b'$', + &mut scan, + &mut dollar_run, + |scan, at| text_math_can_start(view, at, scan), + ) => + { output.push('\\'); output.push(char); } @@ -1512,23 +3045,38 @@ fn escape_text_with_context( output.push('\\'); output.push(char); } - '~' if text_tilde_can_start(input, offset, &mut scan) => { + // A run of three opening a line would open a code fence. + '~' if run_escaped(view, offset, b'~', &mut scan, &mut tilde_run, |scan, at| { + (at == 0 && context.text_opens_line && same_byte_run_len(view, at, b'~') >= 3) + || tilde_run_can_pair(at, scan) + }) => + { output.push('\\'); output.push(char); } - '^' if text_caret_can_start(input, offset, &mut scan) => { + '^' if text_caret_can_start(view, offset, &mut scan) => { output.push('\\'); output.push(char); } - '+' if text_attention_delimiter_can_start(input, offset, "++", false, &mut scan) => { + // A lone `+` left after an escaped one could open an email + // autolink's local part, so a run is escaped whole. + '+' if run_escaped(view, offset, b'+', &mut scan, &mut plus_run, |scan, at| { + text_attention_delimiter_can_start(view, at, "++", false, scan) + || text_doubled_delimiter_can_close(view, at, "++", context.inside) + || text_edge_joins_delimiter(view, at, '+', context.text_edges) + }) => + { output.push('\\'); output.push(char); } - '=' if text_attention_delimiter_can_start(input, offset, "==", false, &mut scan) => { + '=' if text_attention_delimiter_can_start(view, offset, "==", false, &mut scan) + || text_doubled_delimiter_can_close(view, offset, "==", context.inside) + || text_edge_joins_delimiter(view, offset, '=', context.text_edges) => + { output.push('\\'); output.push(char); } - '&' if text_character_reference_can_start(input, offset) => { + '&' if text_character_reference_can_start(view, offset) => { output.push('\\'); output.push(char); } @@ -1542,14 +3090,6 @@ fn escape_text_with_context( output } -fn text_code_span_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool { - let marker_len = scan.run_len_from(b'`', offset); - if marker_len == 0 || text_char_at_edge(input, offset, marker_len) { - return true; - } - scan.exact_run_follows(b'`', offset + marker_len, marker_len) -} - fn text_attention_delimiter_can_start( input: &str, offset: usize, @@ -1569,7 +3109,37 @@ fn text_attention_delimiter_can_start( return false; } - scan.attention_closer_follows(marker, offset + marker.len(), underscore) + scan.written_later(marker) + || scan.attention_closer_follows(marker, offset + marker.len(), underscore) +} + +/// Whether the `+` or `=` at `offset` opens or ends the text beside a `++` or +/// `==` delimiter written right before or after it (`edges`), whose run it +/// would lengthen. +fn text_edge_joins_delimiter( + input: &str, + offset: usize, + char: char, + edges: (Option, Option), +) -> bool { + (offset == 0 && edges.0 == Some(char)) + || (offset + char.len_utf8() == input.len() && edges.1 == Some(char)) +} + +/// Whether the `++` or `==` at `offset` could close an insert or mark the +/// text sits in (`inside`). +fn text_doubled_delimiter_can_close( + input: &str, + offset: usize, + marker: &str, + inside: DelimiterChars, +) -> bool { + input[offset..].starts_with(marker) + && marker + .chars() + .next() + .is_some_and(|char| inside.contains(char)) + && text_delimiter_can_close(input, offset, marker.len(), false) } fn text_delimiter_can_open( @@ -1579,17 +3149,23 @@ fn text_delimiter_can_open( underscore: bool, ) -> bool { let flanking = text_delimiter_flanking(input, offset, marker_len); + if touches_tilde_bonus(input, offset, flanking.next) { + return true; + } if underscore { - flanking.left - && (!flanking.right - || flanking - .previous - .is_some_and(|char| char.is_ascii_punctuation())) + flanking.left && (!flanking.right || flanking.previous.is_some_and(is_flanking_punctuation)) } else { flanking.left } } +/// Whether the `*` or `_` run at `offset` touches a `~` on the side whose char +/// is `neighbour`: the GFM strikethrough bonus lets such a run open or close +/// whatever its flanking. +fn touches_tilde_bonus(input: &str, offset: usize, neighbour: Option) -> bool { + neighbour == Some('~') && matches!(input.as_bytes()[offset], b'*' | b'_') +} + fn text_delimiter_can_close( input: &str, offset: usize, @@ -1597,12 +3173,11 @@ fn text_delimiter_can_close( underscore: bool, ) -> bool { let flanking = text_delimiter_flanking(input, offset, marker_len); + if touches_tilde_bonus(input, offset, flanking.previous) { + return true; + } if underscore { - flanking.right - && (!flanking.left - || flanking - .next - .is_some_and(|char| char.is_ascii_punctuation())) + flanking.right && (!flanking.left || flanking.next.is_some_and(is_flanking_punctuation)) } else { flanking.right } @@ -1622,8 +3197,8 @@ fn text_delimiter_flanking(input: &str, offset: usize, marker_len: usize) -> Tex let previous_whitespace = previous.is_none_or(char::is_whitespace); let next_whitespace = next.is_none_or(char::is_whitespace); - let previous_punctuation = previous.is_some_and(|char| char.is_ascii_punctuation()); - let next_punctuation = next.is_some_and(|char| char.is_ascii_punctuation()); + let previous_punctuation = previous.is_some_and(is_flanking_punctuation); + let next_punctuation = next.is_some_and(is_flanking_punctuation); let left = next.is_some() && !next_whitespace @@ -1643,7 +3218,7 @@ fn text_delimiter_flanking(input: &str, offset: usize, marker_len: usize) -> Tex fn text_less_than_can_start_inline(input: &str, offset: usize, scan: &mut TextScan) -> bool { let after_offset = offset + '<'.len_utf8(); let after = &input[after_offset..]; - if scan.occurs_from(">", after_offset) { + if scan.written_later(">") || scan.occurs_from(">", after_offset) { let next = after.chars().next(); return next.is_some_and(|char| { char.is_ascii_alphabetic() || matches!(char, '/' | '!' | '?' | '_') @@ -1654,6 +3229,39 @@ fn text_less_than_can_start_inline(input: &str, offset: usize, scan: &mut TextSc false } +/// Whether the `:` at `offset` can open a shortcode, closed by a `:` after a +/// name in the text or written by the inlines after it. +fn shortcode_can_form(input: &str, offset: usize, scan: &mut TextScan) -> bool { + let name_end = offset + + 1 + + input[offset + 1..] + .bytes() + .take_while(|byte| byte.is_ascii_alphanumeric() || matches!(byte, b'_' | b'-' | b'+')) + .count(); + let opens = name_end > offset + 1 + && match input.as_bytes().get(name_end) { + Some(b':') => true, + Some(_) => false, + None => scan.written_later(":"), + }; + opens +} + +/// Whether the `:` at `offset` can open a text directive: after whitespace, +/// `(`, `[`, `{`, or at the text's start, and before a directive name. +fn text_directive_can_start(input: &str, offset: usize) -> bool { + let rest = &input[offset + ':'.len_utf8()..]; + let opens_after = input[..offset] + .chars() + .next_back() + .is_none_or(|char| char.is_whitespace() || matches!(char, '(' | '[' | '{')); + let name_len = rest + .bytes() + .take_while(|byte| byte.is_ascii_alphanumeric() || matches!(byte, b'_' | b'-')) + .count(); + opens_after && is_directive_name(&rest[..name_len]) +} + fn text_atx_heading_can_start(input: &str, offset: usize, output_line_len: usize) -> bool { if output_line_len != 0 { return false; @@ -1669,7 +3277,79 @@ fn text_atx_heading_can_start(input: &str, offset: usize, output_line_len: usize fn text_spoiler_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool { input[offset..].starts_with("||") && !input[offset + "||".len()..].starts_with('|') - && scan.occurs_from("||", offset + "||".len()) + && (scan.written_later("|") || scan.occurs_from("||", offset + "||".len())) +} + +/// Whether the `byte` at `offset` is escaped: its whole run is when +/// `escapes` holds at any position in it. `run` holds the end of the run read +/// last and its answer, which the run's later bytes reuse. +fn run_escaped( + input: &str, + offset: usize, + byte: u8, + scan: &mut TextScan, + run: &mut (usize, bool), + mut escapes: impl FnMut(&mut TextScan, usize) -> bool, +) -> bool { + if offset >= run.0 { + let end = offset + same_byte_run_len(input, offset, byte); + let escaped = (offset..end).any(|position| escapes(scan, position)); + *run = (end, escaped); + } + run.1 +} + +fn same_byte_run_len(input: &str, offset: usize, byte: u8) -> usize { + input.as_bytes()[offset..] + .iter() + .take_while(|item| **item == byte) + .count() +} + +/// `input` with every char that the text escaper writes as a character +/// reference replaced by as many `&` bytes: control chars other than a tab, +/// and the spaces and tabs it keeps at a preserved edge. +/// Whether text writes `char` as a character reference: a control char +/// other than the tab, line tabulation, form feed, and next line, which the +/// reparse keeps as written and reads as whitespace, as the source did. +fn written_as_reference(char: char) -> bool { + char.is_control() && !matches!(char, '\t' | '\u{b}' | '\u{c}' | '\u{85}') +} + +fn referenced_chars_as_punctuation( + input: &str, + preserve_leading: bool, + trailing_start: usize, +) -> Cow<'_, str> { + let leading_end = usize::from(preserve_leading && input.starts_with([' ', '\t'])); + let referenced = |offset: usize, char: char| { + written_as_reference(char) + || (matches!(char, ' ' | '\t') && (offset < leading_end || offset >= trailing_start)) + }; + // A control char is ASCII below a space or DEL, or a C1 char, whose UTF-8 + // form opens with 0xC2; so the bytes tell when no char is referenced. + let bytes = input.as_bytes(); + let may_reference = leading_end > 0 + || trailing_start < input.len() + || bytes + .iter() + .any(|byte| (*byte < b' ' && *byte != b'\t') || matches!(*byte, 0x7f | 0xc2)); + if !may_reference + || !input + .char_indices() + .any(|(offset, char)| referenced(offset, char)) + { + return Cow::Borrowed(input); + } + let mut view = String::with_capacity(input.len()); + for (offset, char) in input.char_indices() { + if referenced(offset, char) { + view.extend(core::iter::repeat_n('&', char.len_utf8())); + } else { + view.push(char); + } + } + Cow::Owned(view) } fn text_math_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool { @@ -1681,15 +3361,20 @@ fn text_math_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool if marker_len == 0 || text_char_at_edge(input, offset, marker_len) { return true; } - scan.exact_run_follows(b'$', offset + marker_len, marker_len) + scan.written_later("$") || scan.exact_dollar_run_follows(offset + marker_len, marker_len) } -fn text_tilde_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool { - if input[offset..].starts_with("~~") { - return text_attention_delimiter_can_start(input, offset, "~~", false, scan) - || text_simple_delimiter_can_start(input, offset, "~", scan); - } - text_simple_delimiter_can_start(input, offset, "~", scan) +/// Whether the `~` run from `offset` to its end could pair, or join a run, +/// once written literally: with a `~` later in the text or one the inlines +/// after it (or the span around it) write, or with one written right before +/// the text's start. (A run before it in the text is escaped when it could +/// pair with this one.) Escaping only when it could keeps a `*` or `_` run +/// beside a literal `~` opening or closing as it did. +fn tilde_run_can_pair(offset: usize, scan: &mut TextScan) -> bool { + let end = offset + scan.run_len_from(b'~', offset); + scan.occurs_from("~", end) + || scan.written_later("~") + || (offset == 0 && scan.written_before.contains('~')) } fn text_caret_can_start(input: &str, offset: usize, scan: &mut TextScan) -> bool { @@ -1709,7 +3394,7 @@ fn text_simple_delimiter_can_start( { return true; } - scan.occurs_from(marker, offset + marker.len()) + scan.written_later(marker) || scan.occurs_from(marker, offset + marker.len()) } fn text_character_reference_can_start(input: &str, offset: usize) -> bool { @@ -1753,10 +3438,11 @@ fn same_char_run_len(input: &str, offset: usize, needle: char) -> usize { } fn at_sign_can_start_email_autolink(input: &str, offset: usize) -> bool { + // Any email-local char before the `@` can make the local part. let before = input[..offset] .chars() .next_back() - .is_some_and(|char| char.is_ascii_alphanumeric()); + .is_some_and(|char| char.is_ascii_alphanumeric() || matches!(char, '.' | '-' | '_' | '+')); if !before { return false; } @@ -1784,6 +3470,20 @@ fn at_sign_can_start_email_autolink(input: &str, offset: usize) -> bool { saw_domain_char_after_dot } +/// Whether `output` ends with `char` that no backslash escapes. +fn ends_with_unescaped(output: &str, char: char) -> bool { + output.strip_suffix(char).is_some_and(|before| { + let backslashes = before.len() - before.trim_end_matches('\\').len(); + backslashes % 2 == 0 + }) +} + +/// Whether the next inline opens a line of the block: at the start of +/// inlines that open one, or after a line ending. +fn breaks_line_start(output: &str, output_line: &mut OutputLine, opens_line: bool) -> bool { + output_line.len(output) == 0 && (opens_line || !output.is_empty()) +} + /// The length of the last line of an append-only output, found by scanning only /// the bytes appended since the previous call, so asking once per appended /// char stays linear on a long line. @@ -1810,14 +3510,14 @@ fn escape_destination_with_pipe(input: &str, escape_pipe: bool) -> String { let mut output = String::new(); for char in input.chars() { match char { - char if char.is_control() => output.push_str(&format!("&#x{:X};", char as u32)), + // A space would end the destination, and `\ ` is no escape. + char if char.is_control() || char == ' ' => { + output.push_str(&format!("&#x{:X};", char as u32)); + } '|' if escape_pipe => { output.push('\\'); output.push(char); } - // NB: a literal space is NOT escaped here — `\ ` is not a valid - // escape, so a space-containing destination is routed to the - // angle-bracket form by `serialize_destination_kind`. '(' | ')' | '\\' | '<' | '>' | '&' => { output.push('\\'); output.push(char); @@ -1914,14 +3614,13 @@ fn escape_reference_label_source(input: &str, escape_pipe: bool) -> String { for char in input.chars() { match char { // A reference label may span several physical lines, and the parser - // matches the RAW label (whitespace collapsed, no entity decode). - // Emitting interior newlines/tabs literally (rather than ` `) - // keeps a whitespace-bearing label re-parsing as the same reference - // — crucially, a `^`-prefixed label with literal whitespace stays a - // link reference instead of becoming a footnote (which requires `^` - // followed by non-whitespace). - '\t' | '\n' | '\r' => output.push(char), - char if char.is_control() => output.push_str(&format!("&#x{:X};", char as u32)), + // matches the RAW label (whitespace collapsed, no entity decode), so + // every control char, line endings and tabs among them, is written + // as itself rather than as a reference such as ` `. This keeps + // a whitespace-bearing label re-parsing as the same reference — + // crucially, a `^`-prefixed label with a literal space stays a link + // reference instead of becoming a footnote. + char if char.is_control() => output.push(char), '|' if escape_pipe => { output.push('\\'); output.push(char); @@ -1936,15 +3635,9 @@ fn escape_reference_label_with_pipe(input: &str, escape_pipe: bool) -> String { escape_label_syntax(input, escape_pipe, false) } +/// A footnote label is matched as written, so its chars are. fn escape_footnote_label_source(input: &str) -> String { - let mut output = String::new(); - for char in input.chars() { - match char { - char if char.is_control() => output.push_str(&format!("&#x{:X};", char as u32)), - _ => output.push(char), - } - } - output + input.into() } fn escape_footnote_label_semantic(input: &str) -> String { @@ -1975,9 +3668,13 @@ fn escape_label_syntax(input: &str, escape_pipe: bool, escape_whitespace: bool) fn escape_wikilink_part(input: &str) -> String { let mut output = String::new(); - for char in input.chars() { + for (offset, char) in input.char_indices() { match char { char if char.is_control() => output.push_str(&format!("&#x{:X};", char as u32)), + '&' if text_character_reference_can_start(input, offset) => { + output.push('\\'); + output.push(char); + } '\\' | '[' | ']' | '|' => { output.push('\\'); output.push(char); @@ -2004,14 +3701,6 @@ fn serialize_destination_kind( LinkDestinationKind::Bare | LinkDestinationKind::Omitted => { if input.is_empty() { "<>".into() - } else if input.contains(' ') { - // A bare destination cannot contain a space (it would terminate - // the destination, and `\ ` is not an escape), so emit the - // angle-bracket form instead. - let mut output = String::from("<"); - output.push_str(&escape_angle_destination_with_context(input, context)); - output.push('>'); - output } else { escape_destination_with_pipe(input, context.table_cell) } @@ -2203,6 +3892,44 @@ fn fence_for(input: &str, marker: FenceMarker, min_len: usize) -> String { char.to_string().repeat(min_len.max(longest + 1)) } +/// The fence for a code block's `value` and the columns the block is indented +/// by, so that no line of `value` closes it. A closing fence is up to three +/// spaces, at least as many marker chars, and nothing else but spaces and +/// tabs. The fence keeps `min_len` chars when indenting the block (which the +/// value's lines lose again) moves every closing-like line past three spaces; +/// otherwise it grows past the longest of them. +fn code_block_fence(value: &str, marker: FenceMarker, min_len: usize) -> (String, usize) { + let char = match marker { + FenceMarker::Backtick => '`', + FenceMarker::Tilde => '~', + }; + let closing_like = |line: &str, length: usize| { + let indent = line.len() - line.trim_start_matches(' ').len(); + let rest = &line[indent..]; + let run = rest.len() - rest.trim_start_matches(char).len(); + (indent <= 3 && run >= length && rest[run..].trim_matches([' ', '\t']).is_empty()) + .then_some((indent, run)) + }; + let lines = || value.split(['\n', '\r']); + let least_indent = lines() + .filter_map(|line| closing_like(line, min_len)) + .map(|(indent, _)| indent) + .min(); + match least_indent { + None => (char.to_string().repeat(min_len), 0), + Some(indent) if indent > 0 => (char.to_string().repeat(min_len), 4 - indent), + Some(_) => { + let mut length = min_len; + for line in lines() { + if let Some((_, run)) = closing_like(line, length) { + length = run + 1; + } + } + (char.to_string().repeat(length), 0) + } + } +} + fn inline_code_fence(input: &str) -> String { fence_for(input, FenceMarker::Backtick, 1) } @@ -2239,22 +3966,10 @@ fn serialize_inline_math_with_context( node: &MathInline, context: InlineSerializeContext, ) -> Result { + // A pipe in a table cell is escaped with the cell (see + // `escape_cell_delimiter_pipes`), which the table drops again. + let _ = context; let input = node.value.as_str(); - - // A table-cell pipe cannot live inside a dollar fence (it would split the - // cell), so it is forced into the `$`…`$` code-math form regardless of the - // node's recorded kind. That form cannot represent a value that itself - // contains a `` `$ `` close. - if context.table_cell && input.contains('|') { - if input.contains("`$") { - return Err(SerializeError::UnsupportedNode( - "inline math containing a table pipe and a code-math close", - )); - } - let input = table_cell_escape_code_pipes(input); - return Ok(format!("$`{input}`$")); - } - match node.kind { MathInlineKind::Code => { if input.contains("`$") { @@ -2276,11 +3991,34 @@ fn serialize_inline_math_with_context( } } -/// Whether a serialized cell would split into more cells when parsed back, -/// judged with spoilers enabled so a spoiler's bars are never taken for -/// delimiters. -fn table_cell_has_unescaped_pipe(input: &str) -> bool { - input.contains('|') && !crate::parse::table_row_delimiters(input, true).is_empty() +/// `cell` with a backslash before each pipe that would delimit a cell: a pipe +/// that raw HTML, an autolink, or another verbatim inline writes. The table +/// drops that backslash before the cell's inline parse. +fn escape_cell_delimiter_pipes(cell: String) -> String { + if !cell.contains('|') { + return cell; + } + // A pipe after an odd run of backslashes never delimits, and the cell's + // own text writes its pipes as references, so most cells hold no other. + if !cell + .match_indices('|') + .any(|(index, _)| !ends_with_unescaped(&cell[..index], '\\')) + { + return cell; + } + let delimiters = crate::parse::table_row_delimiters(&cell, true); + if delimiters.is_empty() { + return cell; + } + let mut output = String::with_capacity(cell.len() + delimiters.len()); + let mut copied = 0; + for pipe in delimiters { + output.push_str(&cell[copied..pipe]); + output.push('\\'); + copied = pipe; + } + output.push_str(&cell[copied..]); + output } fn longest_char_streak(input: &str, needle: char) -> usize { diff --git a/src/serialize/escape_scan_tests.rs b/src/serialize/escape_scan_tests.rs index fe423a9..0a0a51a 100644 --- a/src/serialize/escape_scan_tests.rs +++ b/src/serialize/escape_scan_tests.rs @@ -12,14 +12,6 @@ use crate::test_support::{boundaries, generated_inputs, query_orders, Rng}; mod reference { use super::super::*; - pub(super) fn text_code_span_can_start(input: &str, offset: usize) -> bool { - let marker_len = same_char_run_len(input, offset, '`'); - if marker_len == 0 || text_char_at_edge(input, offset, marker_len) { - return true; - } - find_same_char_run(input, offset + marker_len, '`', marker_len).is_some() - } - pub(super) fn text_attention_delimiter_can_start( input: &str, offset: usize, @@ -82,12 +74,13 @@ mod reference { find_same_char_run(input, after_open, '$', marker_len).is_some() } - pub(super) fn text_tilde_can_start(input: &str, offset: usize) -> bool { - if input[offset..].starts_with("~~") { - return text_attention_delimiter_can_start(input, offset, "~~", false) - || text_simple_delimiter_can_start(input, offset, '~'); - } - text_simple_delimiter_can_start(input, offset, '~') + pub(super) fn tilde_run_can_pair(input: &str, offset: usize) -> bool { + let end = offset + + input[offset..] + .bytes() + .take_while(|byte| *byte == b'~') + .count(); + input[end..].contains('~') } pub(super) fn text_caret_can_start(input: &str, offset: usize) -> bool { @@ -147,7 +140,7 @@ fn for_each_scan(seed: u64, mut check: impl FnMut(&str, &mut TextScan, usize)) { for input in generated_inputs(700, 40, seed) { let positions = boundaries(&input); for order in query_orders(&positions, &mut rng) { - let mut scan = TextScan::new(&input); + let mut scan = TextScan::new(&input, DelimiterChars::default()); for offset in order { if offset < input.len() { check(&input, &mut scan, offset); @@ -181,11 +174,6 @@ fn run_and_lookahead_checks_match_the_reference_scan() { for_each_scan(22, |input, scan, offset| { let char = input[offset..].chars().next().expect("offset below len"); match char { - '`' => assert_eq!( - text_code_span_can_start(input, offset, scan), - reference::text_code_span_can_start(input, offset), - "{input:?} at {offset}" - ), '$' => assert_eq!( text_math_can_start(input, offset, scan), reference::text_math_can_start(input, offset), @@ -202,8 +190,8 @@ fn run_and_lookahead_checks_match_the_reference_scan() { "{input:?} at {offset}" ), '~' => assert_eq!( - text_tilde_can_start(input, offset, scan), - reference::text_tilde_can_start(input, offset), + tilde_run_can_pair(offset, scan), + reference::tilde_run_can_pair(input, offset), "{input:?} at {offset}" ), '^' => assert_eq!( diff --git a/src/validate.rs b/src/validate.rs index 20b4597..eee386f 100644 --- a/src/validate.rs +++ b/src/validate.rs @@ -81,7 +81,11 @@ fn validate_block(block: &Block, diagnostics: &mut Vec) { } } Block::Definition(definition) => { - if definition.identifier.trim().is_empty() { + if definition + .identifier + .trim_matches([' ', '\t', '\n', '\r']) + .is_empty() + { diagnostics.push(Diagnostic::invalid( definition.meta.span, "definition identifier cannot be empty", @@ -360,18 +364,18 @@ fn validate_escape(escape: &Escape, diagnostics: &mut Vec) { fn validate_autolink(autolink: &Autolink, diagnostics: &mut Vec) { // GFM literal autolinks carry a synthesized destination that MAY contain // `>` (the renderer percent-encodes it). Only angle-bracket autolinks - // forbid whitespace, `<`, and `>` in the destination. + // forbid a space, an ASCII control char, `<`, and `>` in the destination. if matches!(autolink.kind, AutolinkKind::GfmLiteral { .. }) { return; } if autolink .destination .chars() - .any(|char| char.is_whitespace() || char == '<' || char == '>') + .any(|char| matches!(char, ' ' | '<' | '>') || char.is_ascii_control()) { diagnostics.push(Diagnostic::invalid( autolink.meta.span, - "autolink destination cannot contain whitespace, `<`, or `>`", + "autolink destination cannot contain a space, a control char, `<`, or `>`", )); } } diff --git a/tests/delimiter_stack.rs b/tests/delimiter_stack.rs index e3d7d8d..6bb765e 100644 --- a/tests/delimiter_stack.rs +++ b/tests/delimiter_stack.rs @@ -359,3 +359,119 @@ fn an_image_whose_label_cannot_close_yields_to_a_wikilink() { paragraph.children ); } + +#[test] +fn the_rule_of_three_counts_whole_delimiter_runs() { + let commonmark = SyntaxOptions::commonmark(); + assert_eq!( + parsed(&commonmark, "*a***b*"), + r#"Emphasis["a"]"*"Emphasis["b"]"# + ); + assert_eq!( + parsed(&commonmark, "***a*a*a"), + r#""*"Emphasis[Emphasis["a"]"a"]"a""# + ); +} + +#[test] +fn an_image_whose_resource_is_invalid_falls_back_to_a_shortcut_reference() { + let document = SyntaxOptions::commonmark() + .parse("![foo](a b)\n\n[foo]: /u") + .document; + let Some(Block::Paragraph(paragraph)) = document.children.first() else { + panic!("expected a paragraph"); + }; + assert!( + matches!(paragraph.children.as_slice(), [Inline::ImageReference(image), Inline::Text(rest)] + if image.kind == ReferenceKind::Shortcut && image.identifier == "foo" && rest.value == "(a b)"), + "{:?}", + paragraph.children + ); +} + +#[test] +fn an_underscore_after_unicode_punctuation_opens_as_after_ascii_punctuation() { + let commonmark = SyntaxOptions::commonmark(); + for source in ["\u{ab}_**]**_", "\u{20ac}_**]**_", "\0_**]**_"] { + let shape = parsed(&commonmark, source); + assert!( + shape.ends_with(r#"Emphasis[Strong["]"]]"#), + "{source:?}: {shape}" + ); + } +} + +#[test] +fn an_escaped_backslash_before_a_line_ending_is_no_hard_break() { + assert_eq!( + parsed(&SyntaxOptions::commonmark(), "a\\\\\nb"), + r#""a\\"/"b""# + ); +} + +#[test] +fn a_bare_destination_ends_at_a_space_inside_parentheses() { + let commonmark = SyntaxOptions::commonmark(); + assert_eq!(parsed(&commonmark, "[a](( ))"), r#""[a](( ))""#); + let blocks = commonmark.parse("[o]:(a b)\n\n[o]").document.children; + assert!( + !blocks + .iter() + .any(|block| matches!(block, Block::Definition(_))), + "{blocks:?}" + ); +} + +#[test] +fn a_footnote_label_holds_no_unescaped_bracket() { + let options = SyntaxOptions::default(); + for source in ["^*[^[^]]", "[^a[b]", "[^a[b]: x\n\n[^a[b]"] { + let debug = format!("{:?}", options.parse(source).document.children); + assert!(!debug.contains("FootnoteReference"), "{source:?}: {debug}"); + assert!(!debug.contains("FootnoteDefinition"), "{source:?}: {debug}"); + } + let debug = format!("{:?}", options.parse("[^a\\[b]").document.children); + assert!(debug.contains("FootnoteReference"), "{debug}"); +} + +#[test] +fn an_angle_autolink_holds_whitespace_other_than_a_space() { + let document = SyntaxOptions::commonmark() + .parse("") + .document; + let debug = format!("{:?}", document.children); + assert!( + debug.contains("destination: \"http://a\\u{a0}b\""), + "{debug}" + ); + assert!(document.validate().is_empty()); + assert_eq!(document.to_markdown().unwrap(), "\n"); +} + +#[test] +fn a_referenced_space_makes_no_hard_break() { + let blocks = SyntaxOptions::commonmark() + .parse("a \nb") + .document + .children; + let [Block::Paragraph(paragraph)] = blocks.as_slice() else { + panic!("{blocks:?}"); + }; + assert!( + matches!( + paragraph.children.as_slice(), + [Inline::Text(text), Inline::SoftBreak(_), Inline::Text(_)] if text.value == "a " + ), + "{blocks:?}" + ); +} + +#[test] +fn a_processing_instruction_closes_after_its_opener() { + let blocks = SyntaxOptions::commonmark() + .parse("a b") + .document + .children; + let debug = format!("{blocks:?}"); + assert!(!debug.contains("Html"), "{debug}"); +} diff --git a/tests/fixtures/roundtrip/extensions/gfm_task_list.ast b/tests/fixtures/roundtrip/extensions/gfm_task_list.ast index 0bce6a7..db942fc 100644 --- a/tests/fixtures/roundtrip/extensions/gfm_task_list.ast +++ b/tests/fixtures/roundtrip/extensions/gfm_task_list.ast @@ -48,7 +48,7 @@ Document Text "[ ]" ListItem checked=none Paragraph - Text "[ ] " + Text "[ ]" List ordered=false tight=true ListItem checked=false Paragraph diff --git a/tests/fixtures/roundtrip/extensions/gfm_task_list.canonical.md b/tests/fixtures/roundtrip/extensions/gfm_task_list.canonical.md index 60c55d5..af1c2b4 100644 --- a/tests/fixtures/roundtrip/extensions/gfm_task_list.canonical.md +++ b/tests/fixtures/roundtrip/extensions/gfm_task_list.canonical.md @@ -22,7 +22,7 @@ + \[ \] -+ \[ \] ++ \[ \] * [ ] Text. diff --git a/tests/fixtures/roundtrip/extensions/math_edges.canonical.md b/tests/fixtures/roundtrip/extensions/math_edges.canonical.md index 683d950..d87aa3f 100644 --- a/tests/fixtures/roundtrip/extensions/math_edges.canonical.md +++ b/tests/fixtures/roundtrip/extensions/math_edges.canonical.md @@ -1,8 +1,8 @@ Let $x$ and $$y = z + 2$$ be values. -This is not math: 2000$. -And neither is this \$ 4 $. -Or this $4 +This is not math: 2000\$. +And neither is this \$ 4 \$. +Or this \$4 \$. The cost is between \$10 and 30$. diff --git a/tests/fixtures/roundtrip/spec/commonmark_attention.canonical.md b/tests/fixtures/roundtrip/spec/commonmark_attention.canonical.md index ce4f478..f3468d0 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_attention.canonical.md +++ b/tests/fixtures/roundtrip/spec/commonmark_attention.canonical.md @@ -18,7 +18,7 @@ foo-*(bar)* *(_foo_)* -\__foo\__bar +\_\_foo\_\_bar foo-**(bar)** diff --git a/tests/fixtures/roundtrip/spec/commonmark_autolinks.canonical.md b/tests/fixtures/roundtrip/spec/commonmark_autolinks.canonical.md index 284598f..e0ec439 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_autolinks.canonical.md +++ b/tests/fixtures/roundtrip/spec/commonmark_autolinks.canonical.md @@ -1,5 +1,5 @@ Autolinks and . Email . Special . -Invalid email \ \ \ \. +Invalid email \ \ \ \. Invalid \ \. diff --git a/tests/fixtures/roundtrip/spec/commonmark_blockquotes.canonical.md b/tests/fixtures/roundtrip/spec/commonmark_blockquotes.canonical.md index 93b5499..58a2914 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_blockquotes.canonical.md +++ b/tests/fixtures/roundtrip/spec/commonmark_blockquotes.canonical.md @@ -27,11 +27,9 @@ > b > ``` -> > ``` a ``` - ``` diff --git a/tests/fixtures/roundtrip/spec/commonmark_character_escapes.canonical.md b/tests/fixtures/roundtrip/spec/commonmark_character_escapes.canonical.md index 553452d..401e205 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_character_escapes.canonical.md +++ b/tests/fixtures/roundtrip/spec/commonmark_character_escapes.canonical.md @@ -1,4 +1,4 @@ -!"#$%&'()*+,-./:;\<=>?@\[\]^_`\{|}\~ +!"#$%&'()*+,-./:;\<=>?@\[\]^_\`\{|}~ \\→\\A\\a\\ \\3\\φ\\« diff --git a/tests/fixtures/roundtrip/spec/commonmark_code_spans.ast b/tests/fixtures/roundtrip/spec/commonmark_code_spans.ast index 7ca1660..987b812 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_code_spans.ast +++ b/tests/fixtures/roundtrip/spec/commonmark_code_spans.ast @@ -7,7 +7,7 @@ Document Code value=" both " raw=" both " fence=1 SoftBreak Code value="`code`" raw=" `code` " fence=2 - Paragraph + SoftBreak Code value="``" raw=" `` " fence=3 - Paragraph + SoftBreak Code value="a``b" raw="a``b" fence=3 diff --git a/tests/fixtures/roundtrip/spec/commonmark_code_spans.canonical.md b/tests/fixtures/roundtrip/spec/commonmark_code_spans.canonical.md index 2c70d6b..84b408d 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_code_spans.canonical.md +++ b/tests/fixtures/roundtrip/spec/commonmark_code_spans.canonical.md @@ -5,7 +5,5 @@ bar `` ` both ` `` `code` `` - ``` `` ``` - ```a``b``` diff --git a/tests/fixtures/roundtrip/spec/commonmark_html_blocks.canonical.md b/tests/fixtures/roundtrip/spec/commonmark_html_blocks.canonical.md index 6fc7c6a..42c07f4 100644 --- a/tests/fixtures/roundtrip/spec/commonmark_html_blocks.canonical.md +++ b/tests/fixtures/roundtrip/spec/commonmark_html_blocks.canonical.md @@ -16,6 +16,6 @@
ok and \bad. +Text ok and \bad.