Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
40 commits
Select commit Hold shift + click to select a range
5fcf5b0
docs: propose source span mapping and CommonMark fixes
claude Oct 5, 2026
264c1b9
docs: fold issue #6 container span cases into the span plan
claude Oct 5, 2026
0603ec9
fix: rename the NUL replacement module off a reserved Windows filename
claude Oct 5, 2026
a73cee5
fix: CommonMark rule of three, image reference fallback, and literal …
claude Oct 5, 2026
1d3c415
fix: map every parsed span back to its source bytes
claude Oct 5, 2026
98eafaf
docs: tick the integration checks of the span and CommonMark plan
claude Oct 5, 2026
7c1ea74
fix: end block quotes at lazy list markers and drop final paragraph w…
claude Oct 5, 2026
db37a47
fix: continue lazy paragraphs in list items and block quotes as cmark…
claude Oct 5, 2026
4d0e6f2
fix: read paragraph interruption and quoted verbatim blocks as Common…
claude Oct 5, 2026
0624021
fix: read Unicode punctuation in `_` rules and escaped hard breaks as…
claude Oct 5, 2026
56d7b2b
fix: keep definitions' paragraphs, indented table headers, and direct…
claude Oct 5, 2026
2dbcdf2
perf: keep the serializer's new round-trip checks off the common path
claude Oct 5, 2026
81fa0bf
fix: read tabs after top-level block quote and list item markers at t…
claude Oct 5, 2026
4f0f3e9
fix: let a quoted paragraph that opens with backticks take lazy lines
claude Oct 5, 2026
9d72b8f
fix: keep a list tight across a blank line inside a nested item's ope…
claude Oct 5, 2026
69bf0f5
fix: read tabs at their source columns inside nested containers
claude Oct 5, 2026
ac427ac
fix: write abutting attention runs and literal autolink neighbours so…
claude Oct 5, 2026
4bb3765
fix: reject bracketed footnote labels and keep raw label backticks an…
claude Oct 5, 2026
813909e
fix: keep the strikethrough bonus of attention runs beside a tilde on…
claude Oct 5, 2026
21b3635
fix: try outer underscore runs and raw edge stars when a paragraph do…
claude Oct 5, 2026
a9c885f
perf: scan inline content for literal autolinks only when it holds `:…
claude Oct 5, 2026
90dae8a
fix: escape the delimiters footnotes, math, and autolinks write after…
claude Oct 5, 2026
e7c29bd
perf: collect definitions without parsing inline content, and only wh…
claude Oct 5, 2026
2bee313
fix: read only spaces and tabs as block whitespace
claude Oct 5, 2026
1a787e9
fix: accept non-ASCII whitespace in an angle-bracket autolink's URI
claude Oct 5, 2026
bac286c
fix: keep a literal autolink's end beside spans and check reparses by…
claude Oct 5, 2026
1e04b25
fix: round-trip directives, emails, alerts, footnotes, and ESM-like p…
claude Oct 5, 2026
fbe01f6
fix: read back headings and paragraphs across presets, and keep MDX f…
claude Oct 5, 2026
7f825ea
perf: split the input into lines once for the definition and main passes
claude Oct 5, 2026
ec49bb8
fix: close processing instructions after their opener, break lines on…
claude Oct 5, 2026
edc36a0
fix: start a GFM table only below a delimiter row that is neither laz…
claude Oct 5, 2026
e7ecd83
fix: start a table that ends a paragraph on its header row, and none …
claude Oct 5, 2026
6848c31
test: name the table interruption test's bindings apart from its helper
claude Oct 5, 2026
f41c548
docs: summarize groups 13 to 18 in the plan's What changes
claude Oct 5, 2026
ab149b5
fix: keep serialization and span translation linear
claude Oct 6, 2026
025494d
fix: end blocks and lazy lines at the containers they leave
claude Oct 6, 2026
51d5e16
fix: keep what borders spans, cells, and nested items on reparse
claude Oct 6, 2026
c090823
docs: record group 19 and correct the plan's spec changes
claude Oct 6, 2026
3aea2f9
fix: read the spans around a text instead of what follows it
claude Oct 6, 2026
937e35f
docs: archive source span mapping and CommonMark fixes
claude Oct 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,216 @@
# Block syntax — spec changes

## MODIFIED Requirements

### Requirement: CommonMark blocks
The parser SHALL recognize CommonMark block constructs (ATX and setext headings,
thematic breaks, indented and fenced code blocks, block quotes, lists, HTML
blocks, link reference definitions, and paragraphs) as the CommonMark
specification defines them.

#### Scenario: Heading and paragraph
- **WHEN** `"# Title\n\nHello."` is parsed
- **THEN** the document holds a `Heading` of depth 1 followed by a `Paragraph`

#### Scenario: Lazy list marker ends a block quote
- **WHEN** `"> a\n- "` is parsed with the CommonMark preset
- **THEN** the document holds a `BlockQuote` with a paragraph `a`, followed by a `List` holding one empty item

#### Scenario: Lazy ordered item not starting at 1
- **WHEN** `"> > a\n2. b"` is parsed with the CommonMark preset
- **THEN** the document holds a `BlockQuote` followed by an ordered `List` starting at 2

#### Scenario: Final whitespace of a paragraph
- **WHEN** `"aaa \nbbb "` is parsed with the CommonMark preset
- **THEN** the paragraph holds `Text("aaa")`, a `LineBreak`, and `Text("bbb")`

#### Scenario: Final whitespace of a setext heading
- **WHEN** `"Foo \n-----"` is parsed with the CommonMark preset
- **THEN** the document holds a level-2 setext `Heading` holding `Text("Foo")`

#### Scenario: Lazy line in an item that started blank
- **WHEN** `"- \n a\nb"` is parsed with the CommonMark preset
- **THEN** the list's one item holds a paragraph `a`, a soft break, and `b`

#### Scenario: Blank line indented four columns ends a block quote
- **WHEN** `"> a\n \n> b"` is parsed with the CommonMark preset
- **THEN** the document holds two `BlockQuote`s

#### Scenario: Blank line between an empty item and the next
- **WHEN** `"* \n\n * b c"` is parsed with the CommonMark preset
- **THEN** the document holds one loose `List` of two items

#### Scenario: Complete HTML tag on a lazy line ends a list item
- **WHEN** `"- a\n<a>"` is parsed with the CommonMark preset
- **THEN** the document holds a `List` followed by an `HtmlBlock`

#### Scenario: ATX-like line that is no heading
- **WHEN** `"a\n#)"` is parsed with the CommonMark preset
- **THEN** the document holds one `Paragraph` holding `Text("a")`, a `SoftBreak`, and `Text("#)")`

#### Scenario: Backtick run with a backtick in its info
- **WHEN** ````"a\n``` `` ```"```` is parsed with the CommonMark preset
- **THEN** the document holds one `Paragraph` whose second line is a code span

#### Scenario: Unclosed fence in a block quote
- **WHEN** ``"> ```\n> x\na"`` is parsed with the CommonMark preset
- **THEN** the document holds a `BlockQuote` whose fenced `CodeBlock` holds `"x\n"`, followed by a `Paragraph` holding `Text("a")`

#### Scenario: Last line ending of indented code
- **WHEN** `"\ta\r\tb"` is parsed with the CommonMark preset
- **THEN** the document holds an indented `CodeBlock` whose value is `"a\rb\r"`

#### Scenario: Container content ending in a carriage return
- **WHEN** ``"- ```\n ~\r"`` is parsed with the CommonMark preset
- **THEN** the list item's fenced `CodeBlock` holds `"~\n"`

#### Scenario: Complete tag after a definition
- **WHEN** `"[o]: u\n<a>"` is parsed with the CommonMark preset
- **THEN** the document holds a `Definition` followed by a `Paragraph` holding an `Html` inline

#### Scenario: Indented table header row
- **WHEN** `"a\n |b\n----"` is parsed with `parse`
- **THEN** the document holds one setext `Heading`

#### Scenario: Tab after a top-level block quote marker
- **WHEN** `"> \tcode"` is parsed with the CommonMark preset
- **THEN** the document holds a `BlockQuote` holding a `Paragraph`, since the tab spans columns 2 to 4

#### Scenario: Quoted paragraph opening with backticks
- **WHEN** `"> ``a\nb"` is parsed with the CommonMark preset
- **THEN** the document holds one `BlockQuote` whose paragraph ends with `Text("b")`

#### Scenario: Blank line inside a nested item's open fence
- **WHEN** ``"2. a\n 1. ```\n\n2. b"`` is parsed with the CommonMark preset
- **THEN** the document holds one tight `List`

#### Scenario: Tab inside nested containers
- **WHEN** `"* - \tb c"` is parsed with the CommonMark preset
- **THEN** the inner list item holds an indented `CodeBlock`, since the tab spans columns 4 to 8

#### Scenario: Lazy list marker inside a quoted item
- **WHEN** `"> - a\n> 2.\nz"` is parsed with the CommonMark preset
- **THEN** the document holds a `BlockQuote` holding two lists, followed by a `Paragraph` holding `Text("z")`

#### Scenario: Setext-like line short of a quoted item
- **WHEN** `"> 1. a\n> ===\nb"` is parsed with the CommonMark preset
- **THEN** the item holds one `Paragraph` holding `Text("a")`, `Text("===")`, and `Text("b")` with soft breaks between

#### Scenario: Sibling item ending a nested fence
- **WHEN** ``"- - ```\n - a\n\n- b"`` is parsed with the CommonMark preset
- **THEN** the outer `List` is loose

#### Scenario: Lazy fence-like line in an item
- **WHEN** ``"1. a\n ```\n\nb"`` is parsed with the CommonMark preset
- **THEN** the document holds a `List` followed by a `Paragraph` holding `Text("b")`

#### Scenario: CommonMark oracle cases
- **WHEN** the block cases under `tests/fixtures/conformance/commonmark/` are parsed and rendered with the `html` feature
- **THEN** the output matches the expected HTML

### Requirement: Block directives
When directives are enabled, the parser SHALL recognize `::name[label]{attrs}`
leaf directives and `:::name` container directives, whose container closes at a
fence of at least the opening fence's length, and a line with a malformed
opener SHALL NOT end a paragraph.

#### Scenario: Container directive
- **WHEN** `":::note\nbody\n:::"` is parsed with `parse`
- **THEN** the document holds a `ContainerDirective` named `note` whose children hold a paragraph `body`

#### Scenario: Unclosed container
- **WHEN** `":::note\nunclosed container"` is parsed
- **THEN** a `ContainerDirective` holds the remaining content and an error-severity `UnclosedDirectiveContainer` diagnostic is reported

#### Scenario: Invalid name
- **WHEN** a leaf directive opener has a malformed name
- **THEN** an error-severity `InvalidDirectiveName` diagnostic is reported

#### Scenario: Directive attribute without a valid name
- **WHEN** `":b{<} :c{a <=1 d}"` is parsed with `parse` and serialized
- **THEN** `to_markdown()` returns `":b :c{a d}\n"`

#### Scenario: Malformed directive line inside a paragraph
- **WHEN** `"a\n::1bad"` or `"a\n:::"` is parsed with `parse`
- **THEN** the document holds one `Paragraph` holding both lines

## ADDED Requirements

### Requirement: Spaces and tabs are block whitespace
The parser SHALL read only spaces and tabs as whitespace in block structure: a
blank line holds only spaces and tabs; indentation, the space after a block
marker, the trailing whitespace a thematic break, setext underline, closing
fence, ATX closing sequence, HTML block start line, definition, or table row
allows, the indentation before flow MDX JSX, and the final whitespace of a
paragraph are spaces and tabs. Any
other whitespace char, such as a no-break space or a form feed, is content.

#### Scenario: Other whitespace after a thematic break
- **WHEN** `"***\u{a0}"` is parsed with the CommonMark preset
- **THEN** the document holds a `Paragraph`, not a `ThematicBreak`

#### Scenario: Line holding only a no-break space
- **WHEN** `"a\n\u{a0}\nb"` is parsed with the CommonMark preset
- **THEN** the document holds one `Paragraph`

#### Scenario: No-break space before MDX JSX
- **WHEN** `"\u{a0} <p/>"` is parsed with the MDX preset
- **THEN** the document holds a `Paragraph`, not a flow `MdxJsx` block

#### Scenario: Form feed ending a paragraph
- **WHEN** `"a\u{c}"` is parsed with the CommonMark preset
- **THEN** the paragraph holds `Text("a\u{c}")`

### Requirement: Fenced code inside a container directive
A fenced code block inside a container directive SHALL hold its lines as code:
a line in it that looks like a directive opener opens no nested directive,
while a closing fence of the directive still closes it, and a fence that a
nested directive leaves open ends with that directive.

#### Scenario: Directive opener inside fenced code
- **WHEN** `":::t\n```\n:::e\n```\n:::"` is parsed
- **THEN** the document holds one `ContainerDirective` named `t` holding a `CodeBlock` whose value is `":::e\n"`

#### Scenario: Fence left open in a nested directive
- **WHEN** ``":::outer\n:::inner\n```\n:::\n:::inner2\nx\n:::\n:::\nafter"`` is parsed
- **THEN** the `ContainerDirective` named `outer` holds the directives `inner` and `inner2`, a `Paragraph` holding `Text("after")` follows it, and no diagnostic is reported

### Requirement: Footnote definition content
A footnote definition's content SHALL start after the spaces and tabs that
follow its `]:` and keep the trailing spaces of its first line, which may
make a hard break.

#### Scenario: Hard break on a footnote definition's first line
- **WHEN** `"[^1]: a \nb"` is parsed
- **THEN** the definition's paragraph holds `Text("a")`, a `LineBreak`, and `Text("b")`

### Requirement: List after a definition
A list SHALL start on the line right after a definition only when its first
item could interrupt a paragraph: a bullet or an ordered item starting at 1,
with content. Otherwise the line continues the paragraph the definition was
read from.

#### Scenario: Ordered item not starting at 1 after a definition
- **WHEN** `"[foo]: /url\n2) a"` is parsed with the CommonMark preset
- **THEN** the document holds the `Definition` and a `Paragraph`

### Requirement: GFM table start
A GFM table SHALL start only where the line after its header row is a
delimiter row that is neither a lazy continuation line nor a setext
underline; such a line keeps its other reading.

#### Scenario: Delimiter row without pipes
- **WHEN** `"| --- |\n-- "` is parsed with the GFM preset
- **THEN** the document holds a level-2 setext `Heading`, not a `Table`

#### Scenario: Setext underline below a header row
- **WHEN** `"a\n|b\n---"` is parsed with the GFM preset
- **THEN** the document holds one level-2 setext `Heading` holding both lines

#### Scenario: Lazy delimiter row
- **WHEN** `"1. ---(\n:-:"` is parsed with the GFM preset
- **THEN** the list item holds a `Paragraph`, not a `Table`

#### Scenario: Table header row that looks like an empty list item
- **WHEN** `"a\n+\n|-"` is parsed with the GFM preset
- **THEN** the document holds a `Paragraph` and a `Table` whose header cell holds `+`
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
# Inline syntax — spec changes

## MODIFIED Requirements

### Requirement: CommonMark inlines
The parser SHALL recognize CommonMark inline constructs (backslash escapes,
entity and numeric character references, code spans, emphasis and strong
emphasis, links, images, autolinks, raw HTML, and hard and soft line breaks) as
the CommonMark specification defines them, including its precedence of code
spans, links, and emphasis.

#### Scenario: Emphasis
- **WHEN** `"Hello *world*."` is parsed
- **THEN** the paragraph holds `Text("Hello ")`, an `Emphasis` containing `world`, and `Text(".")`

#### Scenario: Link inside a link label
- **WHEN** `"[foo [bar](/u)](/v)"` is parsed
- **THEN** only `[bar](/u)` becomes a link and the surrounding brackets and `(/v)` stay text

#### Scenario: Shortcut reference before an unclosed label
- **WHEN** `"[foo][bar\n\n[foo]: /u"` is parsed with the CommonMark preset
- **THEN** the paragraph holds a shortcut `LinkReference` to `foo` followed by `Text("[bar")`

#### Scenario: Rule of three counts whole delimiter runs
- **WHEN** `"*a***b*"` is parsed with the CommonMark preset
- **THEN** the paragraph holds an `Emphasis` containing `a`, `Text("*")`, and an `Emphasis` containing `b`

#### Scenario: Image whose resource is invalid
- **WHEN** `"![foo](a b)\n\n[foo]: /u"` is parsed with the CommonMark preset
- **THEN** the paragraph holds a shortcut `ImageReference` to `foo` followed by `Text("(a b)")`

#### Scenario: Underscore after Unicode punctuation
- **WHEN** `"«_**]**_"` is parsed with the CommonMark preset
- **THEN** the paragraph holds `Text("«")` and an `Emphasis` containing a `Strong` containing `]`

#### Scenario: Escaped backslash before a line ending
- **WHEN** `"a\\\\\nb"` is parsed with the CommonMark preset
- **THEN** the paragraph holds `Text("a\\")`, a `SoftBreak`, and `Text("b")`

#### Scenario: Space inside a bare destination's parentheses
- **WHEN** `"[a](( ))"` is parsed with the CommonMark preset
- **THEN** the paragraph holds `Text("[a](( ))")` and no `Link`

#### Scenario: CommonMark oracle cases
- **WHEN** the inline cases under `tests/fixtures/conformance/commonmark/` are parsed and rendered with the `html` feature
- **THEN** the output matches the expected HTML

## ADDED Requirements

### Requirement: Footnote labels
The parser SHALL read `[^label]` as a footnote reference, and `[^label]:` as a
footnote definition, only when the label is non-empty, holds no space, tab, or
line ending, and, as a link label, holds no unescaped `[` or `]`.

#### Scenario: Bracket inside a footnote label
- **WHEN** `"^*[^[^]]"` and `"[^a[b]"` are parsed with `parse`
- **THEN** neither paragraph holds a `FootnoteReference`, while `"[^a\\[b]"` holds one

### Requirement: Reference label matching
Two link labels SHALL match when they agree after Unicode case folding,
trimming, and collapsing each run of spaces, tabs, and line endings to one
space; any other whitespace char is matched as written.

#### Scenario: No-break space in a label
- **WHEN** `"[a\u{a0}b]\n\n[a b]: /u"` is parsed
- **THEN** the paragraph holds no `LinkReference`

### Requirement: Angle-bracket autolink URI
The parser SHALL read `<scheme:rest>` as an autolink when the scheme is valid
and the rest holds no space, ASCII control char, `<`, or `>`; any other
whitespace char is part of the URI.

#### Scenario: No-break space in an angle-bracket autolink
- **WHEN** `"<http://a\u{a0}b>"` is parsed with the CommonMark preset
- **THEN** the paragraph holds an `Autolink` to `http://a\u{a0}b`

### Requirement: Hard line breaks from spaces
A line ending SHALL be a hard break when two or more spaces the source holds,
and no tab, end the line; spaces or tabs a character reference writes are
text, and only the source's spaces and tabs before a soft break are removed.

#### Scenario: Referenced space before a line ending
- **WHEN** `"a&#x20; \nb"` is parsed with the CommonMark preset
- **THEN** the paragraph holds `Text("a ")`, a `SoftBreak`, and `Text("b")`

### Requirement: Processing instructions
Raw inline HTML SHALL read `<?` as a processing instruction only when a `?>`
after the `<?` closes it.

#### Scenario: `<?>` is text
- **WHEN** `"a<?> b"` is parsed with the CommonMark preset
- **THEN** the paragraph holds no `Html` inline
Loading
Loading