Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .changeset/preserve-html-code-languages.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"@neuledge/context": patch
---

Preserve language fences and code indentation when HTML pre/code blocks contain surrounding formatting whitespace or comments.
50 changes: 50 additions & 0 deletions packages/context/src/html.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -34,4 +34,54 @@ describe("DocBook HTML examples", () => {
);
expect(parsed.sections[0]?.content).toContain("```sh\necho hello\n```");
});

it.each([
["spaces", " ", " "],
["newlines", "\n ", "\n"],
["tabs", "\t", "\t\n"],
["comments", "\n<!-- example -->\n ", "\n<!-- end -->\n"],
])("preserves language and indentation with %s around code", (_, before, after) => {
const parsed = parseHtml(
`<h1>Guide</h1><h2>Example</h2><pre>${before}<code class="highlight language-sh">if true; then\n echo hello\nfi\n</code>${after}</pre>`,
"guide.html",
);
expect(parsed.sections).toHaveLength(1);
expect(parsed.sections[0]?.hasCode).toBe(true);
expect(parsed.sections[0]?.content).toBe(
"```sh\nif true; then\n echo hello\nfi\n```",
);
});

it("keeps embedded backticks inside a language fence with wrapper whitespace", () => {
const parsed = parseHtml(
'<h1>Guide</h1><h2>Example</h2><pre>\n <code class="language-markdown">first\n```\nlast</code>\n</pre>',
"guide.html",
);
expect(parsed.sections).toHaveLength(1);
expect(parsed.sections[0]?.hasCode).toBe(true);
expect(parsed.sections[0]?.content).toBe(
"````markdown\nfirst\n```\nlast\n````",
);
});

it.each([
[
'prefix <code class="language-sh">echo hello</code> suffix',
"prefix echo hello suffix",
],
[
'<span>prefix </span><code class="language-sh">echo hello</code>',
"prefix echo hello",
],
[
' <code class="language-sh">echo hello</code><code> suffix</code>',
" echo hello suffix",
],
])("preserves meaningful siblings in mixed preformatted content: %s", (content, expected) => {
const parsed = parseHtml(
`<h1>Guide</h1><h2>Example</h2><pre>${content}</pre>`,
"guide.html",
);
expect(parsed.sections[0]?.content).toBe(`\`\`\`\n${expected}\n\`\`\``);
});
});
18 changes: 16 additions & 2 deletions packages/context/src/html.ts
Original file line number Diff line number Diff line change
Expand Up @@ -40,17 +40,31 @@ for (const tag of REMOVED_TAGS) {

// DocBook emits bare <pre> elements; Turndown's code rule requires <pre><code>.
// Preserve their whitespace and prevent Markdown escaping of unit-file examples.
// Also recognize a sole <code> child preceded by whitespace or comments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maintainability · comment: Add why this case exists: whitespace or comments before <code> prevent Turndown’s built-in first-child rule from recognizing the language. That rationale explains when this special case is needed.


This codebase is managed by Human0.

turndown.addRule("barePre", {
filter: (node) =>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Behavior · comment: The custom rule excludes <pre> where <code> is already the first child; Turndown’s built-in handler then emits only firstChild.textContent, so <pre><code>x</code> suffix</pre> loses its suffix. This predates the patch, but the description says mixed preformatted content retains all sibling text; narrow that claim or cover this case if intended.


This codebase is managed by Human0.

node.nodeName === "PRE" && node.firstChild?.nodeName !== "CODE",
replacement: (_content, node) => {
const code: string = node.textContent ?? "";
const codeElement = node.firstElementChild;
const isWrappedCode =
codeElement?.nodeName === "CODE" &&
[...node.childNodes].every(
(child) =>
child === codeElement ||
(child.nodeType === 3 && !/\S/.test(child.textContent ?? "")) ||

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maintainability · comment: Use Node.TEXT_NODE and Node.COMMENT_NODE (or named constants) instead of magic values 3 and 8 so the predicate reads clearly.


This codebase is managed by Human0.

child.nodeType === 8,
);
const className = isWrappedCode
? (codeElement.getAttribute("class") ?? "")
: "";
const language = className.match(/\blanguage-(\S+)/)?.[1] ?? "";
const code: string = (isWrappedCode ? codeElement : node).textContent ?? "";
const longestRun = (code.match(/`+/g) ?? []).reduce(
(longest, run) => Math.max(longest, run.length),
0,
);
const fence = "`".repeat(Math.max(3, longestRun + 1));
return `\n\n${fence}\n${code.replace(/\n$/, "")}\n${fence}\n\n`;
return `\n\n${fence}${language}\n${code.replace(/\n$/, "")}\n${fence}\n\n`;
},
});

Expand Down
Loading