fix: mount point-in-time snapshots into forum containers instead of the live SQLite WAL trio (macOS Docker corruption) - #18
Open
dohooo wants to merge 1 commit into
Conversation
…um containers
Forum-phase containers used to bind-mount the live knowledge DB plus its
-wal/-shm sidecars read-write, and the in-container MCP server opened them
with read_only=True. On macOS (Docker Desktop) that combination corrupts
the authoritative host store:
- In WAL mode, even a read-only SQLite connection WRITES to the -shm
wal-index: readers update read-marks and, when their view of the header
looks invalid, grab the recovery locks and REBUILD the wal-index.
- A VirtioFS bind mount shares neither POSIX advisory locks nor a coherent
shared mmap between the macOS host and the Linux VM. Host fcntl locks are
invisible in the container and vice versa, and each side has its own page
cache of the -shm file, lazily synced underneath the other's live mapping.
- A container "reader" therefore silently clobbers the host writer's
wal-index mid-transaction. Subsequent host WAL appends/checkpoint
backfills use poisoned frame->page mappings and copy page images to wrong
offsets in the main DB file: "sqlite3.DatabaseError: database disk image
is malformed" (autopsies show cross-page transplants, e.g. an index leaf
at the `generations` table root and a copy of page 1 at page 13), or a
host SIGBUS inside walIndexAppend/walIndexReadHdr when the mapping is
invalidated outright ("FS pagein error").
The failure probability scales with the number of concurrent forum
containers and WAL churn, which is why 15-workstream campaigns died at the
gen-1-to-2 boundary (first forum-heavy phase transition) while 3-task runs
usually survived, and why strategies without forum phases (whose containers
never mount the DB files) never corrupted. Reproduced synthetically with
zero LLM calls: a host KnowledgeStore under forum-drain-shaped load dies
within seconds alongside 15 (or even 1) containers that merely OPEN the
bind-mounted trio read-only - including with a plain sqlite3 mode=ro
connection and no ksi code in the container. The same load with no
containers, at the same rate and duration, is clean.
Fix - the live WAL trio never crosses the container boundary:
- _store_common.snapshot_db_for_container_mount() produces a standalone
point-in-time copy via VACUUM INTO on its own read-only connection
(consistent WAL read snapshot, safe next to the writer thread; output is
a single rollback-journal file with no sidecars).
- container_host materializes that snapshot for per_task_forum /
cross_task_forum payloads (same lifecycle as the per-task memory
snapshot: created in the knowledge-DB dir, unlinked after the run) and
stamps payload.knowledge.db_is_snapshot. On snapshot failure the
container runs with NO knowledge DB (visible degradation) - it never
falls back to mounting the live files.
- container_mounts.ts mounts the snapshot as a single read-only file for
forum containers and stops mounting the live knowledge/runtime-audit DB
files there entirely. Forum WRITES are unaffected: they already flow
through the forum_bus JSONL mount and are drained into SQLite host-side
by the single writer.
- mcp_server opens a snapshot-backed knowledge store with the new
KnowledgeStore(read_only=True, immutable=True) (KNOWLEDGE_DB_IMMUTABLE=1
from the runner): immutable=1 is truthful there and makes SQLite skip
locking and wal-index shared memory entirely.
Freshness is unchanged in practice: the knowledge DB only gains rows at
drain/phase boundaries (before the next container launch and snapshot),
and mid-round forum traffic travels via the ForumBus JSONL, not the DB.
Validation: tests/memory/test_forum_container_db_snapshot.py and
tests/runtime/test_container_host.py::TestForumKnowledgeDbSnapshot fail on
the previous code and pass now; the synthetic 15-container repro is clean
for the full duration in the fixed topology (read-only immutable snapshot
mount) while the live-trio topology still dies in seconds; a previously
always-fatal 15-task ARC-1 forum campaign (3 generations) completes with
zero "malformed" errors and a clean PRAGMA integrity_check.
Claude-Session: https://claude.ai/code/session_01KNTYDZGTNzY3Yrg2wB7ixj
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Symptom
On macOS with Docker Desktop, campaigns that enable forum phases eventually kill the authoritative knowledge store with
sqlite3.DatabaseError: database disk image is malformed(host-side, on the live DB — not inside a container), or occasionally a host SIGBUS insidewalIndexAppend/walIndexReadHdr. The failure probability scales with the number of concurrent forum containers and WAL churn: 15-workstream campaigns died reliably at the first forum-heavy generation boundary, while small 3-task runs usually survived. Strategies without forum phases never corrupt, because their containers never mount the DB files. Linux hosts are unaffected, which is why this never showed up in CI (ubuntu-only).Root cause
Forum containers bind-mount the live knowledge DB together with its
-wal/-shmsidecars, and the in-container MCP server opens themread_only=True. In WAL mode even a read-only SQLite connection writes to the-shmwal-index (read-marks, and full wal-index rebuilds when its view of the header looks invalid) and takes POSIX advisory locks. Docker Desktop's VirtioFS file sharing provides neither of the two things that make that safe: host and VM cannot see each other's advisory locks, and each side has its own page cache of the-shmfile rather than a coherent shared mmap. A container "reader" therefore silently clobbers the host writer's wal-index mid-transaction, and the host's next WAL append/checkpoint backfill copies page images through a poisoned frame→page mapping to wrong offsets in the main DB file. Forensics on a corrupted store show exactly that signature: whole-page transplants, e.g. an index leaf sitting at a table's root page and a byte-for-byte copy of page 1 at page 13.Reproduced synthetically with zero LLM calls: a host
KnowledgeStoreunder forum-drain-shaped write load dies within seconds next to containers that merely open the bind-mounted trio read-only — including with a plainsqlite3mode=roconnection and no ksi code in the container. The same load with no containers is clean.The change
The live WAL trio never crosses the container boundary:
_store_common.snapshot_db_for_container_mount()cuts a point-in-time snapshot viaVACUUM INTOon its own read-only connection — a consistent WAL read snapshot, safe next to the writer thread; the output is a single rollback-journal file with no sidecars.container_hostmaterializes that snapshot forper_task_forum/cross_task_forumpayloads (same lifecycle as the per-task memory snapshot: created in the knowledge-DB dir, unlinked after the run) and stampsknowledge.db_is_snapshot. If the snapshot cannot be built, the container runs with no knowledge DB (logged, visible degradation) — it never falls back to mounting the live files.container_mounts.tsmounts the snapshot as a single read-only file for forum containers and stops mounting the live knowledge / runtime-audit DB files there entirely.KnowledgeStore(read_only=True, immutable=True)(KNOWLEDGE_DB_IMMUTABLE=1from the runner):immutable=1is truthful for a snapshot and makes SQLite skip locking and wal-index shared memory entirely.Forum writes are unaffected: they already flow through the ForumBus JSONL mount and are drained into SQLite host-side by the single writer. Freshness is also unchanged in practice — the knowledge DB only gains rows at drain/phase boundaries, i.e. before the next container launch and snapshot.
Behavior on Linux
Semantics converge from "container reads possibly racing the live WAL" to "deterministic snapshot taken at phase start". No functional loss, and phase inputs become reproducible instead of depending on checkpoint timing.
Validation
tests/memory/test_forum_container_db_snapshot.py(snapshot is standalone and complete while the writer is live, overwrites stale destinations, leaves no partial file on failure,immutablerequiresread_only, immutable opens create no sidecars) andtests/runtime/test_container_host.py::TestForumKnowledgeDbSnapshot(forum payloads point at a standalone snapshot and never the live path, non-forum sources keep the historical payload, snapshot failure degrades without a live fallback). The payload-contract tests fail on the previous code.pytest(3105 passed),ruff check/ruff format --check,mypy src/ksi,tsc --noEmitfor both runners,node --test tests/js— green. (One pre-existing, unrelated flake on macOS hosts:egress_lifecycle.test.mjs"keeps a sibling lease alive…" fails identically on currentmain.)malformederrors and a cleanPRAGMA integrity_check.Repro notes for maintainers with a Mac
Docker Desktop with the default VirtioFS file sharing. Run any forum-enabled campaign with ~15 concurrent workstreams and wait for the first generation boundary; the host store typically dies within the first forum drain. A minimal no-LLM repro: open a
KnowledgeStoreon the host and write in a loop (forum-drain-shaped commits), while a container bind-mounts the DB +-wal+-shmand merely opens it withsqlite3.connect("file:/mnt/db?mode=ro", uri=True)— corruption within seconds. Mount theVACUUM INTOsnapshot instead (this PR's topology) and the same load runs clean.