Skip to content

fix: map spans to their source and fix CommonMark parse and round-trip defects - #12

Merged
plimeor merged 40 commits into
mainfrom
claude/keen-ride-o76svu
Oct 6, 2026
Merged

plimeor merged 40 commits into
mainfrom
claude/keen-ride-o76svu

Conversation

@plimeor

@plimeor plimeor commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Summary

This implements the plan source-span-mapping-and-commonmark-fixes, now archived at docs/archive/2026-10-06-source-span-mapping-and-commonmark-fixes/. The plan's spec changes are merged into docs/specs/.

Closes #6.

Source spans

Every parsed node's span maps to the source bytes it was read from:

  • on every line of every container;
  • across CRLF;
  • across split tabs;
  • inside table cells, which now carry spans.

Spans nest within their parent. The issue's 31-case regression test passes.

Parsing

The parser now matches commonmark.js and micromark where they agree on:

  • Emphasis: the rule of three and opener floor use original run lengths.
  • Image references: an invalid image resource falls back to a reference.
  • Containers: lazy lines and blank lines in block quotes and list items, including lazy list markers and an HTML tag on a lazy line.
  • Paragraph interruption: ATX-like lines and backtick runs.
  • Open blocks in quotes: unclosed fences, math, and HTML blocks.
  • Tab columns at any container depth.
  • Whitespace: only spaces and tabs count as block whitespace.
  • Labels: label matching, footnote labels, and angle-autolink URIs.
  • Breaks: a hard break comes only from source spaces.
  • Processing instructions: <?> is text.
  • Definitions: a list after a definition.
  • GFM tables: a table starts below a delimiter row that is neither lazy nor a setext underline, before any block its header row would open.

Serialization

Many more parsed documents serialize to Markdown that reads back to the same tree:

  • Text is escaped against every construct beside it.
  • A paragraph or heading that does not read back tries other delimiter choices.

Canonical output changes

  • Every backtick in text is written as \`.
  • A few constructs move to spellings that read back. The regenerated goldens show each one.

Performance

  • Parsing the 400 KB fixture document takes 137M instructions, against 261M before.
  • Serialization takes about 51M, against 46M.
  • Growth tests now check linear time for:
    • nested emphasis;
    • long ~ runs;
    • inline diagnostics that run to a paragraph's end.

Known issues

The final review left findings open. They were accepted for archive and are tracked in #11:

  • A panic that predates this branch: parse("\u{a0}e+@"). It should be fixed before the next release.
  • 13 nested quote-and-list inputs that the group-19 column tracking newly breaks. The same tracking fixes 378 others.
  • Five round-trip classes listed under the plan's Risks.

#11 replaces the container re-prediction and the serializer's copies of parser rules at their root, and folds autolinks into Link.

Verification

  • cargo fmt --check, cargo test with and without html, RUSTDOCFLAGS='-D warnings' cargo doc --no-deps, cargo build --target wasm32-unknown-unknown, cargo +1.82 build, and the release pathological tests all pass.
  • Conformance stays at 2233 of 2236.
  • The parse comparisons against commonmark.js and micromark, on the inputs where they agree, are recorded per task in the archived plan.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3


Generated by Claude Code

claude added 30 commits October 5, 2026 11:47
Plan to map every parsed node's span back to its source bytes across
container prefixes, CRLF, table cells, and split tabs, and to fix the
rule of three, image reference fallback, and serializer escaping of
backticks and text after shortcut references.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
Add the issue's 31-case regression test to the plan's tasks and
representative scenarios for nested lists, nested block quotes, alerts,
footnote continuations, HTML containers, container directives, tabs,
and CRLF to the public-api spec change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
`src/parse/nul.rs` cannot be created on Windows, so the packaged crate
would fail to unpack there; `cargo package` warns about it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…text escaping

- The rule of three reads each delimiter run's original length, as
  CommonMark specifies, instead of what earlier pairings left.
- An image whose `(…)` is not a valid resource falls back to the full,
  collapsed, and shortcut reference forms, as a link does.
- The serializer escapes every backtick in text: a backslash keeps a
  backtick from opening a code span but not from closing one.
  Canonical output changes for text with a backtick it left bare.
- Text after a shortcut reference escapes a leading `(`, and a leading
  `:` when the reference opens the line, so reparsing keeps the reference.
- Inside `_`-delimited emphasis, a `_` in text that can close is escaped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
Containers and inline parsing read derived strings: lines with container
markers and indentation removed, tabs split into spaces, lines joined
with `\n`, and table cells with `\|` read as `|`. Positions in those
strings were offset from one base, so every line after a stripped
prefix, a CRLF, or a cell pipe got a shifted span, and code blocks in
containers could span past the end of the input.

A source map now pairs runs of each derived string with the input bytes
they came from, in original coordinates. Lines take their positions
through it, and each inline parse translates its spans in one walk.
Table cells carry spans. A container whose first content line is empty
keeps that line, which realigns lazy-line flags.

Closes #6.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
`cargo fmt --check`, `cargo test` with and without `html`, `cargo doc`
with `-D warnings`, the wasm32 build, and `cargo +1.82 build` pass;
`tests/pathological_inputs.rs` passes in debug and release. Conformance
is 2233/2236, as before. On a 4 MB document of the fixture corpus,
parsing takes a median 331 ms against 312-326 ms at the starting commit,
and serialization 56 ms against 59-60 ms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…hitespace

- A lazy line that opens a list ends the block quote it would otherwise
  continue, including an empty item and an ordered item not starting at
  1, as cmark, commonmark.js, markdown-it, and micromark read it.
- Paragraphs and setext headings drop the final whitespace of their
  content, as CommonMark specifies; the last text node no longer keeps
  trailing spaces or tabs.
- A level-two setext heading whose text ends in `|` is written with the
  pipe escaped, so its `---` underline is not read as a table delimiter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
… does

- A list item tracks its open paragraph afresh after a blank line or a
  single-line block, so a lazy line continues a paragraph in an item that
  started blank or held a thematic break.
- The open paragraph is tracked with its block quote level: a line that
  reaches it continues it by the paragraph-interruption rules, a line short
  of it only as a lazy line.
- A lazy line keeps its lazy flag inside nested lists.
- A blank line ends a block quote however far it is indented.
- A blank line between two items loosens the list, and one before a
  thematic break that ends the list no longer does.
- A complete HTML tag on a lazy line ends a list item, as it ends a block
  quote (cmark-gfm and micromark behavior).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…Mark does, and round-trip their edges

- Only a real ATX heading or code fence interrupts a paragraph: `#)` and a
  backtick run whose info string holds a backtick continue it.
- A line without `>` after an unclosed fence, math block, or HTML block in a
  block quote ends the quote; block quotes and list items share
  `content_line_state` to track the open block.
- A code or math block ended by the input ends its last line with the
  value's first line ending.
- The serializer round-trips empty fenced code, info-string edge whitespace,
  text right after a literal autolink, a paragraph opening with a soft
  break, HTML block values, text lines that would open an HTML block or a
  leaf directive, and indented code ending in `\r` / `\r\n`; CRLF output
  keeps existing `\r\n`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
… CommonMark does, and round-trip delimiter runs whole

Parser:
- `_` opens and closes beside Unicode punctuation as beside ASCII punctuation.
- An escaped backslash before a line ending is text, not a hard break.
- A bare link destination ends at a space inside parentheses.
- A malformed directive line does not interrupt a paragraph.
- A container's last content line ends in `\n` like the others.

Serializer:
- Escape a `*`, `_`, or `$` run whole or not at all, reading neighbours as
  the reparse sees them (later inlines' delimiters, tabs, references).
- Choose `_` emphasis only where `_` can flank; indent continuation lines
  inside inlines that would start a block; write a break opening a line or
  a delimited span after `&#x20;`; encode spaces in bare destinations;
  start a list item whose first block opens with whitespace on the next
  line; lengthen code fences only past closing-like lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…ive attributes as CommonMark and GFM read them, and round-trip extension text edges

Parser:
- A complete-tag line right after a definition continues the paragraph the
  definition was read from.
- A table header row indented four columns or more starts no table.
- A directive attribute without a valid name is dropped, so every parsed
  document serializes.
- An unclosed HTML comment holding a fence-like line no longer counts as an
  unclosed fence when its container closes.

Serializer:
- Escape text that would open a text directive, shortcode, raw HTML, or
  math with a later sibling; escape pipes that inlines write in table cells
  with the cell instead of failing.
- Indent list markers past an indented following block; end a quote with an
  empty `>` line before a paragraph in a tight item; start a list item on
  the next line when its first line would read as a thematic break; indent
  a fenced block rather than lengthen its fence when that suffices.
- Keep raw HTML after a definition in its paragraph; escape description
  detail markers at line starts; keep frontmatter's empty last line.
- Leave a literal `~` beside an emphasis run unescaped unless it can pair.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
- Prefilter continuation lines and referenced-char scans on their first or
  any byte, look delimiter chars up by match, and build sibling char sets
  only for runs of more than one inline that can need them.
- Skip the cell pipe scan for cells without an unescaped pipe, and
  preallocate escaped text and fenced block buffers.

On a 400 KB document of the fixtures the serializer runs 138M instructions,
down from 179M (114M at 0.3.0).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…heir source columns

A tab after `> ` or in a top-level list item's continuation indentation now
spans the columns it spans in the source line (up to the four columns that
decide indentation), so `> \tcode` is a paragraph as in CommonMark. Lines
inside an open fence or HTML block keep their tabs, as does a tab past an
indented code block's indentation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
The GH-19 rule (a fence-like lazy line ends a block quote) applies to the
lazy line only; a quoted content line opening with backticks is a paragraph
and keeps lazy continuation, as cmark, commonmark.js, and micromark read it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…n fence

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
Every line now carries the source column it starts at; `DerivedText`
records it for container content lines. List markers, list continuations,
block quote markers, fence indentation, and the container line
classification read tabs from that column, so `> > \ta` and `* - \tb` hold
indented code as in CommonMark at any depth.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
… they read back

A paragraph whose Strong and Emphasis runs abut is reparsed once; when the
plain delimiter choice does not read back, `__` strong, `_` inner emphasis,
and all-`*` delimiters are tried in turn. Text after a literal autolink is
guarded through the span delimiters written after it, and a scheme char
ending a text before a `://` literal autolink is escaped or encoded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…d wiki link bangs literal

A footnote label holds no unescaped bracket, as a link label. A text
backtick after a reference whose raw label holds a backtick is written as
`&#96;`, since an escaped backtick still closes a code span, and a `!`
before a wiki link is escaped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
… the round trip

A text `*` or `_` touching a `~` is escaped as one the GFM strikethrough
bonus lets open or close, and a paragraph that does not read back also
tries writing a `~` run at a strong's or emphasis's edge raw.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…es not read back

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…//`, `www.`, or `@`

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
… a text

A text escapes a `^` a footnote's `^` could close, a `$` a math fence or a
literal autolink could close, and a scheme char before any relaxed scheme.
A run of bars that can open a spoiler is written as references up to its
last bar, and a paragraph that does not read back tries every delimiter
choice with each raw edge char, including for a strong inside a strong.
What sits beside a paragraph's runs is read in one walk.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…en a `]:` is present

The definitions pass reads block structure alone, and a table delimiter
row holding any char other than `|`, `-`, `:`, or whitespace is rejected
before its cells are split.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
A line holding a no-break space, form feed, or other non-ASCII whitespace
is neither blank nor indented, and such a char ends no thematic break,
setext underline, fence, ATX heading, HTML block start line, table row, or
paragraph. Labels collapse only spaces, tabs, and line endings, and a
footnote label may hold any other whitespace. The serializer writes those
whitespace control chars and label controls as themselves, encodes a hard
break that opens a span, and escapes delimiters that wiki links, code, and
raw HTML write after a text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
… tree

A paragraph that does not read back also tries raw line-edge spaces, a
referenced space before a literal autolink, and an escaped first char of
the text after one. The reparse check gives each reference label a
stand-in definition and compares span-free trees instead of debug text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…aragraphs

A directive opener inside fenced code in a container directive is code,
and a footnote definition's first line keeps its trailing spaces. The
serializer ends a bare text directive before what could go on with it,
escapes email-local chars and `@` before an email or literal autolink and
`+` runs whole, writes alert titles as their source and empty container
directives without a blank line, and keeps a paragraph opening with
`import ` or `export ` off MDX ESM.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…low and task text

Headings take the delimiter choices paragraphs do, and a choice that
reads back under the default preset wins before one under GFM or MDX.
Flow MDX JSX is indented by spaces and tabs only; a paragraph's first
line stays off MDX ESM and flow; a paragraph after an alert in a tight
item is separated; task item text keeps its opening whitespace; and a
`}` written later escapes a text `{`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…ly on literal spaces, and keep lists after definitions in the paragraph

A `<?>` is text; a hard break needs two spaces the source holds, so a
referenced space before a line ending stays text; and a list that could
not interrupt a paragraph continues the paragraph a definition was read
from. Raw HTML opening a paragraph is written as it is, after the
definition it continues.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
claude added 10 commits October 5, 2026 20:14
…y nor a setext underline

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
…below a lazy line

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
Fixes three cost blowups found in review, each now pinned by a growth or
bounded-time test in tests/pathological_inputs.rs:

- An emphasis rendered its content twice to choose its delimiter, so
  nested emphasis cost 2^depth; renders are reused by content and context.
- A long `~` run re-measured the run at each position; the run length is
  memoized.
- An inline diagnostic's end was found by walking every segment after its
  start; it is found by galloping, or directly at the last segment.

Also fixes round trips the review found:
- the plain rendering keeps `**` after a `*`;
- the read-back gate reads runs inside links, images, and marks;
- a strong at the edge of the run around it can be written `__`;
- an escaped `_` at an emphasis edge, or a `_` opening the next text, no
  longer blocks the `_` choice;
- a cell pipe after an escaped backslash is escaped;
- a reference's raw label backtick reaches text inside and after spans;
- an `==` or `++` that could close its mark or insert is escaped;
- math opening a definition's paragraph is written as its continuation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
- A quoted paragraph records the first list item it sits in and the block
  quotes inside that item; a line short of either continues it only as a
  lazy line, so `> - a\n> 2.\nz` ends the quote before `z`.
- A block opened in a list item records the item's content column, and a
  line indented less ends it, so a sibling item ends a nested fence.
- A lazy line in a list item opens no fence.
- A fence left open in a nested container directive ends with it.
- A setext underline is no delimiter row wherever a table start is
  checked, so `a\n|b\n---` is one heading.
- The `|` read from a cell's `\|` maps to both source bytes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
Round trips the fuzz found on further seeds:
- an escaped backtick in a reference label sets no code-span guard;
- a `:` ending the text inside an insert or underline, which the span's
  delimiters and a `:` after it would make a shortcode, is escaped;
- whitespace opening a table cell before a literal autolink is encoded;
- a list in a list item is indented past the block after it, as at the
  top level;
- whitespace before a text directive stays raw;
- a `+` or `=` beside a `++` or `==` delimiter is escaped;
- a block quote whose first line would read as an alert marker opens with
  an empty line.

Also tests: the span nesting check reads the round-trip cases, which name
their body's length, and `d_~_` pins only its round trip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
Adds group 19 (the review fixes) to the plan, its What changes, the
residual round-trip classes and the split-tab gap to Risks, and the
untrusted-input-cost spec change. Corrects spec text the review found
wrong: a broken code span, the split-tab value, the span rule for block
nodes, a sentence without SHALL, and the serialization requirements on
canonical spelling and options that disagreed with the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
The re-review found these regressions from group 19:

- Adjacent strongs merged into one wherever no read-back runs (table cells,
  description lists, directive labels). A closing `**` and an opening `**`
  form a run of four that by the rule of three closes neither, so the
  plain rendering writes the second strong with `__` exactly there.
- Three escapes guessed the span a text sits in from the delimiter chars
  written anywhere after or before it, and escaped ordinary prose
  (`**Note:** use snake_case: here`, `a +\nb`). The context now carries
  the spans around the inlines and the delimiter beside a text's edges,
  and the escapes read those.
- A reference's raw label backtick was detected in the whole output of a
  span, code spans included; spans are now judged by the labels of the
  references inside them.
- The render memo now holds only content with a nested span, where
  nesting can double the work, which brings flat emphasis back to the
  base's cost.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
Merges the plan's spec changes into docs/specs (block-syntax,
inline-syntax, public-api, serialization, untrusted-input-cost,
validation), moves the plan to docs/archive/2026-10-06-…, and points the
overview's Next at the panic and the root-cause rewrite in
#11. The findings the re-verify left open were
accepted for archive and are tracked there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hyap2mHhP25kLtUxRS8XZ3
@plimeor
plimeor merged commit f987d66 into main Oct 6, 2026
3 checks passed
@plimeor
plimeor deleted the claude/keen-ride-o76svu branch October 6, 2026 02:38
@github-actions github-actions Bot mentioned this pull request Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Inline spans are offset inside container blocks (lists, blockquotes, tables, footnotes, containers)

2 participants