Skip to content

fix: mount point-in-time snapshots into forum containers instead of the live SQLite WAL trio (macOS Docker corruption) - #18

Open
dohooo wants to merge 1 commit into
recursive-knowledge:mainfrom
dohooo:fix/forum-container-shm-corruption
Open

dohooo wants to merge 1 commit into
recursive-knowledge:mainfrom
dohooo:fix/forum-container-shm-corruption

Conversation

@dohooo

@dohooo dohooo commented Aug 14, 2026 •

Copy link
Copy Markdown

Symptom

On macOS with Docker Desktop, campaigns that enable forum phases eventually kill the authoritative knowledge store with sqlite3.DatabaseError: database disk image is malformed (host-side, on the live DB — not inside a container), or occasionally a host SIGBUS inside walIndexAppend/walIndexReadHdr. The failure probability scales with the number of concurrent forum containers and WAL churn: 15-workstream campaigns died reliably at the first forum-heavy generation boundary, while small 3-task runs usually survived. Strategies without forum phases never corrupt, because their containers never mount the DB files. Linux hosts are unaffected, which is why this never showed up in CI (ubuntu-only).

Root cause

Forum containers bind-mount the live knowledge DB together with its -wal/-shm sidecars, and the in-container MCP server opens them read_only=True. In WAL mode even a read-only SQLite connection writes to the -shm wal-index (read-marks, and full wal-index rebuilds when its view of the header looks invalid) and takes POSIX advisory locks. Docker Desktop's VirtioFS file sharing provides neither of the two things that make that safe: host and VM cannot see each other's advisory locks, and each side has its own page cache of the -shm file rather than a coherent shared mmap. A container "reader" therefore silently clobbers the host writer's wal-index mid-transaction, and the host's next WAL append/checkpoint backfill copies page images through a poisoned frame→page mapping to wrong offsets in the main DB file. Forensics on a corrupted store show exactly that signature: whole-page transplants, e.g. an index leaf sitting at a table's root page and a byte-for-byte copy of page 1 at page 13.

Reproduced synthetically with zero LLM calls: a host KnowledgeStore under forum-drain-shaped write load dies within seconds next to containers that merely open the bind-mounted trio read-only — including with a plain sqlite3 mode=ro connection and no ksi code in the container. The same load with no containers is clean.

The change

The live WAL trio never crosses the container boundary:

  • _store_common.snapshot_db_for_container_mount() cuts a point-in-time snapshot via VACUUM INTO on its own read-only connection — a consistent WAL read snapshot, safe next to the writer thread; the output is a single rollback-journal file with no sidecars.
  • container_host materializes that snapshot for per_task_forum / cross_task_forum payloads (same lifecycle as the per-task memory snapshot: created in the knowledge-DB dir, unlinked after the run) and stamps knowledge.db_is_snapshot. If the snapshot cannot be built, the container runs with no knowledge DB (logged, visible degradation) — it never falls back to mounting the live files.
  • container_mounts.ts mounts the snapshot as a single read-only file for forum containers and stops mounting the live knowledge / runtime-audit DB files there entirely.
  • The in-container MCP server opens it with the new KnowledgeStore(read_only=True, immutable=True) (KNOWLEDGE_DB_IMMUTABLE=1 from the runner): immutable=1 is truthful for a snapshot and makes SQLite skip locking and wal-index shared memory entirely.

Forum writes are unaffected: they already flow through the ForumBus JSONL mount and are drained into SQLite host-side by the single writer. Freshness is also unchanged in practice — the knowledge DB only gains rows at drain/phase boundaries, i.e. before the next container launch and snapshot.

Behavior on Linux

Semantics converge from "container reads possibly racing the live WAL" to "deterministic snapshot taken at phase start". No functional loss, and phase inputs become reproducible instead of depending on checkpoint timing.

Validation

  • 9 new regression tests: tests/memory/test_forum_container_db_snapshot.py (snapshot is standalone and complete while the writer is live, overwrites stale destinations, leaves no partial file on failure, immutable requires read_only, immutable opens create no sidecars) and tests/runtime/test_container_host.py::TestForumKnowledgeDbSnapshot (forum payloads point at a standalone snapshot and never the live path, non-forum sources keep the historical payload, snapshot failure degrades without a live fallback). The payload-contract tests fail on the previous code.
  • Full local run of the CI checks on this branch: pytest (3105 passed), ruff check / ruff format --check, mypy src/ksi, tsc --noEmit for both runners, node --test tests/js — green. (One pre-existing, unrelated flake on macOS hosts: egress_lifecycle.test.mjs "keeps a sibling lease alive…" fails identically on current main.)
  • Synthetic repro: the 15-container live-trio topology still dies in seconds; the snapshot/immutable topology is clean for the full duration.
  • Empirical: multi-generation benchmark campaigns with forum phases enabled on macOS corrupted the store 4/4 times before this change; after it, 3 full campaigns completed with zero malformed errors and a clean PRAGMA integrity_check.

Repro notes for maintainers with a Mac

Docker Desktop with the default VirtioFS file sharing. Run any forum-enabled campaign with ~15 concurrent workstreams and wait for the first generation boundary; the host store typically dies within the first forum drain. A minimal no-LLM repro: open a KnowledgeStore on the host and write in a loop (forum-drain-shaped commits), while a container bind-mounts the DB + -wal + -shm and merely opens it with sqlite3.connect("file:/mnt/db?mode=ro", uri=True) — corruption within seconds. Mount the VACUUM INTO snapshot instead (this PR's topology) and the same load runs clean.

…um containers

Forum-phase containers used to bind-mount the live knowledge DB plus its
-wal/-shm sidecars read-write, and the in-container MCP server opened them
with read_only=True. On macOS (Docker Desktop) that combination corrupts
the authoritative host store:

- In WAL mode, even a read-only SQLite connection WRITES to the -shm
  wal-index: readers update read-marks and, when their view of the header
  looks invalid, grab the recovery locks and REBUILD the wal-index.
- A VirtioFS bind mount shares neither POSIX advisory locks nor a coherent
  shared mmap between the macOS host and the Linux VM. Host fcntl locks are
  invisible in the container and vice versa, and each side has its own page
  cache of the -shm file, lazily synced underneath the other's live mapping.
- A container "reader" therefore silently clobbers the host writer's
  wal-index mid-transaction. Subsequent host WAL appends/checkpoint
  backfills use poisoned frame->page mappings and copy page images to wrong
  offsets in the main DB file: "sqlite3.DatabaseError: database disk image
  is malformed" (autopsies show cross-page transplants, e.g. an index leaf
  at the `generations` table root and a copy of page 1 at page 13), or a
  host SIGBUS inside walIndexAppend/walIndexReadHdr when the mapping is
  invalidated outright ("FS pagein error").

The failure probability scales with the number of concurrent forum
containers and WAL churn, which is why 15-workstream campaigns died at the
gen-1-to-2 boundary (first forum-heavy phase transition) while 3-task runs
usually survived, and why strategies without forum phases (whose containers
never mount the DB files) never corrupted. Reproduced synthetically with
zero LLM calls: a host KnowledgeStore under forum-drain-shaped load dies
within seconds alongside 15 (or even 1) containers that merely OPEN the
bind-mounted trio read-only - including with a plain sqlite3 mode=ro
connection and no ksi code in the container. The same load with no
containers, at the same rate and duration, is clean.

Fix - the live WAL trio never crosses the container boundary:

- _store_common.snapshot_db_for_container_mount() produces a standalone
  point-in-time copy via VACUUM INTO on its own read-only connection
  (consistent WAL read snapshot, safe next to the writer thread; output is
  a single rollback-journal file with no sidecars).
- container_host materializes that snapshot for per_task_forum /
  cross_task_forum payloads (same lifecycle as the per-task memory
  snapshot: created in the knowledge-DB dir, unlinked after the run) and
  stamps payload.knowledge.db_is_snapshot. On snapshot failure the
  container runs with NO knowledge DB (visible degradation) - it never
  falls back to mounting the live files.
- container_mounts.ts mounts the snapshot as a single read-only file for
  forum containers and stops mounting the live knowledge/runtime-audit DB
  files there entirely. Forum WRITES are unaffected: they already flow
  through the forum_bus JSONL mount and are drained into SQLite host-side
  by the single writer.
- mcp_server opens a snapshot-backed knowledge store with the new
  KnowledgeStore(read_only=True, immutable=True) (KNOWLEDGE_DB_IMMUTABLE=1
  from the runner): immutable=1 is truthful there and makes SQLite skip
  locking and wal-index shared memory entirely.

Freshness is unchanged in practice: the knowledge DB only gains rows at
drain/phase boundaries (before the next container launch and snapshot),
and mid-round forum traffic travels via the ForumBus JSONL, not the DB.

Validation: tests/memory/test_forum_container_db_snapshot.py and
tests/runtime/test_container_host.py::TestForumKnowledgeDbSnapshot fail on
the previous code and pass now; the synthetic 15-container repro is clean
for the full duration in the fixed topology (read-only immutable snapshot
mount) while the live-trio topology still dies in seconds; a previously
always-fatal 15-task ARC-1 forum campaign (3 generations) completes with
zero "malformed" errors and a clean PRAGMA integrity_check.

Claude-Session: https://claude.ai/code/session_01KNTYDZGTNzY3Yrg2wB7ixj
@dohooo
dohooo requested a review from xuefei-wang as a code owner August 14, 2026 20:40

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant