Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion .github/workflows/convert.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ on:
workflow_dispatch:
inputs:
filename:
description: "Specific filename in input/ (e.g. book.pdf, notes.md, site.zip)"
description: "Specific filename in input/ (e.g. book.pdf, book.epub, notes.md, site.zip)"
required: false
type: string

Expand Down Expand Up @@ -38,6 +38,9 @@ jobs:
- name: Install dependencies
run: pip install -r requirements.txt

- name: Install pandoc
run: sudo apt-get update && sudo apt-get install -y pandoc

- name: Process content
run: python scripts/process.py
env:
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ jobs:
run: pip install -r requirements.txt

- name: Validate Python scripts
run: python3 -m py_compile scripts/build_manifest.py scripts/convert.py
run: python3 -m py_compile scripts/build_manifest.py scripts/convert.py scripts/process.py

- name: Run Python tests
run: python3 -m unittest discover -s tests/scripts -v
Expand Down
16 changes: 9 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

GitHub-hosted content shelf. Fork, upload, done.

> Fork this repo to get your own content platform on GitHub Pages. Upload PDFs (auto-converted to books), Markdown documents (rendered directly), or ZIP archives (deployed as static sites). Zero server cost.
> Fork this repo to get your own content platform on GitHub Pages. Upload PDFs or EPUBs (auto-converted to books), Markdown documents (rendered directly), or ZIP archives (deployed as static sites). Zero server cost.

## Quick Start

Expand All @@ -24,7 +24,7 @@ Your site is now live at `https://<your-username>.github.io/gitshelf/`
3. In your fork, go to **Settings > Secrets and variables > Actions**
4. Click **New repository secret**, name it `MINERU_TOKEN`, paste the token

> Only needed if you want to upload PDFs. Markdown and ZIP uploads work without this.
> Only needed if you want to upload PDFs. EPUB, Markdown, and ZIP uploads work without this.

### 3. Password Protection (Optional)

Expand All @@ -41,6 +41,7 @@ Your site is now live at `https://<your-username>.github.io/gitshelf/`
([Create one here](https://github.com/settings/tokens/new?scopes=repo&description=GitShelf))
3. Upload a file:
- **`.pdf`** — Converted to a multi-chapter book via MinerU API
- **`.epub`** — Converted to a multi-chapter book via pandoc
- **`.md`** — Rendered directly as a document
- **`.zip`** — Extracted as a static site (must contain `index.html`)
4. Wait for GitHub Actions to process (progress shown in Actions tab)
Expand All @@ -50,14 +51,14 @@ Your site is now live at `https://<your-username>.github.io/gitshelf/`

| Type | Upload | Display |
|------|--------|---------|
| **Book** | `.pdf` file | Chapter reader with TOC sidebar, keyboard navigation |
| **Book** | `.pdf` or `.epub` file | Chapter reader with TOC sidebar, keyboard navigation |
| **Document** | `.md` file | Single-page Markdown rendering with syntax highlighting |
| **Site** | `.zip` file | Static site served directly (clicks open in new tab) |

## Features

- **Reader** — Dark/light theme, chapter sidebar, keyboard navigation, code highlighting (Shiki), math rendering (KaTeX), responsive layout
- **Admin** — Upload PDFs/Markdown/ZIPs from browser, catalog management (edit, publish, hide, archive, delete), search & filter
- **Admin** — Upload PDFs/EPUBs/Markdown/ZIPs from browser, catalog management (edit, publish, hide, archive, delete), search & filter
- **Pipeline** — GitHub Actions processes uploads automatically, handles large PDFs by auto-chunking
- **Homepage** — Tab-based filtering: All / Books / Documents / Sites

Expand All @@ -66,9 +67,10 @@ Your site is now live at `https://<your-username>.github.io/gitshelf/`
```
Upload content (browser → GitHub API → input/)
→ GitHub Actions runs scripts/process.py
→ .pdf: MinerU API → Markdown → Split chapters → docs/books/{id}/
→ .md: Copy to docs/articles/{id}/content.md
→ .zip: Extract to docs/sites/{id}/
→ .pdf: MinerU API → Markdown → Split chapters → docs/books/{id}/
→ .epub: pandoc → Markdown + media → Split chapters → docs/books/{id}/
→ .md: Copy to docs/articles/{id}/content.md
→ .zip: Extract to docs/sites/{id}/
→ Build manifest → GitHub Pages deploys
```

Expand Down
16 changes: 9 additions & 7 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

基于 GitHub 的内容托管平台。Fork,上传,搞定。

> Fork 本仓库即可拥有自己的内容平台。上传 PDF(自动转换为书籍)、Markdown 文档(直接渲染)、ZIP 压缩包(部署为静态站点),全部托管在 GitHub Pages 上。零服务器成本。
> Fork 本仓库即可拥有自己的内容平台。上传 PDF 或 EPUB(自动转换为书籍)、Markdown 文档(直接渲染)、ZIP 压缩包(部署为静态站点),全部托管在 GitHub Pages 上。零服务器成本。

## 快速开始

Expand All @@ -24,7 +24,7 @@
3. 在你的 Fork 中,进入 **Settings > Secrets and variables > Actions**
4. 点击 **New repository secret**,名称填 `MINERU_TOKEN`,粘贴 Token

> 仅上传 PDF 时需要。Markdown 和 ZIP 上传无需此配置。
> 仅上传 PDF 时需要。EPUB、Markdown 和 ZIP 上传无需此配置。

### 3. 密码保护(可选)

Expand All @@ -41,6 +41,7 @@
([点此创建](https://github.com/settings/tokens/new?scopes=repo&description=GitShelf))
3. 上传文件:
- **`.pdf`** — 通过 MinerU API 转换为多章节书籍
- **`.epub`** — 通过 pandoc 转换为多章节书籍
- **`.md`** — 直接作为文档渲染展示
- **`.zip`** — 解压为静态站点(需包含 `index.html`)
4. 等待 GitHub Actions 处理完成
Expand All @@ -50,14 +51,14 @@

| 类型 | 上传格式 | 展示方式 |
|------|----------|----------|
| **书籍** | `.pdf` | 章节阅读器 + TOC 侧栏 + 键盘导航 |
| **书籍** | `.pdf` 或 `.epub` | 章节阅读器 + TOC 侧栏 + 键盘导航 |
| **文档** | `.md` | 单页 Markdown 渲染,支持代码高亮和数学公式 |
| **站点** | `.zip` | 静态站点直接托管,点击新窗口打开 |

## 功能

- **阅读器** — 明暗主题、章节侧边栏、键盘导航、代码高亮(Shiki)、数学公式(KaTeX)、响应式布局
- **管理面板** — 上传 PDF/Markdown/ZIP、目录管理(编辑、发布、隐藏、归档、删除)、搜索和筛选
- **管理面板** — 上传 PDF/EPUB/Markdown/ZIP、目录管理(编辑、发布、隐藏、归档、删除)、搜索和筛选
- **处理流水线** — GitHub Actions 自动处理上传内容,大 PDF 自动分块转换
- **首页** — 标签页筛选:全部 / 书籍 / 文档 / 站点

Expand All @@ -66,9 +67,10 @@
```
上传内容(浏览器 → GitHub API → input/)
→ GitHub Actions 运行 scripts/process.py
→ .pdf: MinerU API → Markdown → 拆分章节 → docs/books/{id}/
→ .md: 复制到 docs/articles/{id}/content.md
→ .zip: 解压到 docs/sites/{id}/
→ .pdf: MinerU API → Markdown → 拆分章节 → docs/books/{id}/
→ .epub: pandoc → Markdown + 媒体资源 → 拆分章节 → docs/books/{id}/
→ .md: 复制到 docs/articles/{id}/content.md
→ .zip: 解压到 docs/sites/{id}/
→ 构建 manifest → GitHub Pages 部署
```

Expand Down
2 changes: 1 addition & 1 deletion cli/commands/reconvert.js
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ async function run(argv) {
const { token, repo } = loadConfig(argv);
const { items } = await fetchCatalog(repo, token);
const { item } = selectCatalogItem(items, id);
if (item.type !== 'book') die('Only PDF books can be re-processed. Re-upload Markdown or ZIP sources instead.');
if (item.type !== 'book') die('Only books can be re-processed. Re-upload Markdown or ZIP sources instead.');

const clearCache = hasFlag(argv, '--clear-cache');
await triggerReconvert(item, repo, token, { clearCache });
Expand Down
27 changes: 20 additions & 7 deletions cli/github.js
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ const FAILURES_PATH = 'docs/failures.json';
const CONTENT_TYPE_DIRS = { book: 'docs/books', doc: 'docs/articles', site: 'docs/sites' };
const VISIBILITY_VALUES = ['published', 'hidden', 'archived'];
const MAX_FILE_SIZE = 100 * 1024 * 1024;
const ACCEPTED_EXTENSIONS = ['pdf', 'md', 'zip'];
const ACCEPTED_EXTENSIONS = ['pdf', 'epub', 'md', 'zip'];

let catalogSourcePath = CATALOG_DEFAULT_PATH;

Expand Down Expand Up @@ -172,17 +172,30 @@ async function listRepoTree(repo, path, token) {
}

async function getCacheDeleteOps(repo, itemId, token) {
let md5;
let pdfMd5;
let epubMd5;
try {
const f = await readJson(repo, `docs/books/${itemId}/meta.json`, token);
md5 = f.data?.pdf_md5;
pdfMd5 = f.data?.pdf_md5;
epubMd5 = f.data?.epub_md5;
} catch { /* skip */ }
if (!md5) {
return [];
const ops = [];

if (pdfMd5) {
const cacheFiles = await listRepoTree(repo, 'cache/markdown', token);
ops.push(...cacheFiles
.filter((f) => f.name.startsWith(pdfMd5))
.map((f) => ({ path: f.path, delete: true })));
}

if (epubMd5) {
const cacheFiles = await listRepoTree(repo, 'cache/epub', token);
ops.push(...cacheFiles
.filter((f) => f.name.startsWith(epubMd5))
.map((f) => ({ path: f.path, delete: true })));
}

const cacheFiles = await listRepoTree(repo, 'cache/markdown', token);
return cacheFiles.filter((f) => f.name.startsWith(md5)).map((f) => ({ path: f.path, delete: true }));
return ops;
}

function isSameCatalogItem(left, right) {
Expand Down
4 changes: 2 additions & 2 deletions cli/index.js
Original file line number Diff line number Diff line change
Expand Up @@ -23,12 +23,12 @@ Usage:
gitshelf <command> [options]

Commands:
upload <file> Upload .pdf, .md, or .zip to GitShelf
upload <file> Upload .pdf, .epub, .md, or .zip to GitShelf
list [--type TYPE] List all content items
info <id|type:id> Show details for one item
edit <id|type:id> [...] Edit item metadata
delete <id|type:id> [--yes] Delete an item permanently
reconvert <id|type:id> Trigger re-processing for a PDF book
reconvert <id|type:id> Trigger re-processing for a book source
failures List processing failures
failures dismiss <filename> Dismiss a failure
failures retry <filename> Retry a failed conversion
Expand Down
12 changes: 6 additions & 6 deletions cli/mcp-server.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -3,10 +3,10 @@
/**
* GitShelf MCP Server
*
* GitShelf is a zero-cost GitHub Pages content platform. Upload PDFs
* (auto-converted to chapter books), Markdown (rendered as documents),
* or ZIP archives (deployed as static sites). Everything is processed
* by GitHub Actions and served from GitHub Pages.
* GitShelf is a zero-cost GitHub Pages content platform. Upload PDFs or
* EPUBs (auto-converted to chapter books), Markdown (rendered as
* documents), or ZIP archives (deployed as static sites). Everything is
* processed by GitHub Actions and served from GitHub Pages.
*
* This MCP server exposes GitShelf content reading and management tools.
*
Expand Down Expand Up @@ -42,7 +42,7 @@ const server = new McpServer(
version: '0.1.3',
},
{
instructions: 'GitShelf is a zero-cost GitHub Pages content platform. Upload PDFs (auto-converted to multi-chapter books with TOC and reader UI), Markdown files (rendered as documents), or ZIP archives (deployed as static sites). Everything is processed by GitHub Actions and served from GitHub Pages. Use these tools to browse, read, and manage GitShelf content.',
instructions: 'GitShelf is a zero-cost GitHub Pages content platform. Upload PDFs or EPUBs (auto-converted to multi-chapter books with TOC and reader UI), Markdown files (rendered as documents), or ZIP archives (deployed as static sites). Everything is processed by GitHub Actions and served from GitHub Pages. Use these tools to browse, read, and manage GitShelf content.',
},
);

Expand Down Expand Up @@ -233,7 +233,7 @@ server.registerTool(
'upload',
{
title: 'Upload Content',
description: 'Upload a local file (.pdf, .md, or .zip) to GitShelf for processing.',
description: 'Upload a local file (.pdf, .epub, .md, or .zip) to GitShelf for processing.',
inputSchema: z.object({
file_path: z.string().describe('Absolute path to the file to upload'),
}),
Expand Down
13 changes: 8 additions & 5 deletions scripts/generate_structure.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,9 +30,9 @@ def _slugify_anchor(text: str) -> str:
return anchor.strip("-")


def _extract_subheadings(content: str) -> list[dict[str, str]]:
"""Find H2 headings within a chapter to produce sub-children entries with anchors."""
sub_level = 2
def _extract_subheadings(content: str, chapter_level: int = 1) -> list[dict[str, str]]:
"""Find sub-headings immediately below the chapter level for TOC children."""
sub_level = chapter_level + 1
pattern = re.compile(rf"^{'#' * sub_level}(?!#)\s+(.+)$", re.MULTILINE)
subheadings: list[dict[str, str]] = []
for match in pattern.finditer(content):
Expand All @@ -47,6 +47,7 @@ def _extract_subheadings(content: str) -> list[dict[str, str]]:
def _build_toc(
title: str,
chapters: list[Chapter],
chapter_level: int = 1,
) -> dict:
"""Build the toc.json structure from a list of chapters.

Expand All @@ -56,7 +57,7 @@ def _build_toc(
children: list[dict] = []
for chapter in chapters:
entry: dict = {"title": chapter.title, "slug": chapter.slug}
subheadings = _extract_subheadings(chapter.content)
subheadings = _extract_subheadings(chapter.content, chapter_level=chapter_level)
if subheadings:
entry["children"] = [
{
Expand Down Expand Up @@ -93,6 +94,8 @@ def generate_book_structure(
title: str,
chapters: list[Chapter],
output_dir: Path = Path("docs/books"),
*,
chapter_level: int = 1,
) -> Path:
"""Create book directory, write chapter files, generate toc.json.

Expand Down Expand Up @@ -123,7 +126,7 @@ def generate_book_structure(
chapter_path = chapters_dir / f"{chapter.slug}.md"
chapter_path.write_text(chapter.content, encoding="utf-8")

toc = _build_toc(title, chapters)
toc = _build_toc(title, chapters, chapter_level=chapter_level)
toc_path = book_dir / "toc.json"
toc_path.write_text(json.dumps(toc, indent=2, ensure_ascii=False) + "\n", encoding="utf-8")

Expand Down
Loading
Loading