You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
mcp-server.js still saves intelligence.json non-atomically and swallows corrupt loads — #634 / #698 fixed cli.js only #995
#634 reported that Intelligence.save() overwrites .ruvector/intelligence.json with a plain fs.writeFileSync and that load() swallows a parse error into empty defaults, so a concurrent reader that lands mid-write silently restarts from an empty store. #698 fixed both in bin/cli.js (atomicWriteFileSync, readIntelStoreSafe quarantining to .corrupt-<epoch>), but bin/mcp-server.js has its own copy of the Intelligence class that was not touched:
mcp-server.jsIntelligence.save() still ends in fs.writeFileSync(this.intelPath, JSON.stringify(this.data, null, 2)) — line 396 in both 0.2.41 and 0.3.1 (line 320 in 0.2.34).
mcp-server.jsIntelligence.load() still wraps the read in try { … } catch {} and returns { patterns: {}, memories: [], … } on any parse failure, with no quarantine.
Under Claude Code the MCP server is the long-lived writer (every hooks_remember call goes through it) while each PreToolUse/PostToolUse/SessionStart hook is a separate short-lived ruvector hooks … CLI process reading the same file. The server's save therefore still opens the torn-file window #634 described; the fixed CLI side now detects the torn read and quarantines the file, but the store the user ends up with is whatever the hook wrote next.
What we observed (ruvector 0.2.34 server, 2026-09-17): a 9 MB onnx-minilm-stamped store with 757 memories was being saved by the MCP server while ~12 parallel subagents fired PostToolUse hooks. A hook-side CLI read the half-written file, moved it to intelligence.json.corrupt-1789669812640, and created a fresh store — hash-stamped, 64-dim, because the hook path resolved the hash embedder — so hooks reembed --dry-run then reported an embedder-provenance MISMATCH against a store that had been fine an hour earlier. The quarantined file was complete and parsed with all 757 memories; the server still had the good snapshot in memory. The same sequence is reachable on 0.3.1 because the server-side write is unchanged.
To Reproduce
Start the MCP server (ruvector mcp start, or via a Claude Code plugin) against a store of a few MB with vector memories.
In a loop, call hooks_remember through the server while running ~20 parallel npx ruvector@0.3.1 hooks route "x" / hooks remember "y" -t test CLI processes against the same .ruvector/.
Eventually a CLI process reads mid-write: with 0.2.41+ it prints the quarantine error and intelligence.json.corrupt-<epoch> appears; the server's next save() then overwrites intelligence.json from memory, or a CLI process creates a fresh store first, depending on timing.
Deterministic variant (no race): while the server is running, truncate intelligence.json mid-file, then trigger a server-side operation that calls load() (e.g. restart the server) — load() returns empty defaults with no error and the next save() persists the empty store.
mcp-server.jsIntelligence.load() fails loud (or quarantines, like readIntelStoreSafe) on an unparseable store instead of returning empty defaults that the next save() persists.
Reader side: rename() alone does not protect a reader that already opened the old inode before the rename (see Rename atomicity is not enough npm/write-file-atomic#64), so a parse-retry once (re-open + re-read) before quarantining would avoid quarantining a file that was simply replaced under the reader.
Environment
OS: Linux 6.18 (WSL2, Ubuntu 24.04)
Node: 24.15.0 (also reproduced reading with 22.22.0)
Ruvector version: 0.2.34 (observed); 0.2.41 and 0.3.1 verified to still have the plain writeFileSync at bin/mcp-server.js:396
Describe the bug
#634 reported that
Intelligence.save()overwrites.ruvector/intelligence.jsonwith a plainfs.writeFileSyncand thatload()swallows a parse error into empty defaults, so a concurrent reader that lands mid-write silently restarts from an empty store. #698 fixed both inbin/cli.js(atomicWriteFileSync,readIntelStoreSafequarantining to.corrupt-<epoch>), butbin/mcp-server.jshas its own copy of theIntelligenceclass that was not touched:mcp-server.jsIntelligence.save()still ends infs.writeFileSync(this.intelPath, JSON.stringify(this.data, null, 2))— line 396 in both 0.2.41 and 0.3.1 (line 320 in 0.2.34).mcp-server.jsIntelligence.load()still wraps the read intry { … } catch {}and returns{ patterns: {}, memories: [], … }on any parse failure, with no quarantine.Under Claude Code the MCP server is the long-lived writer (every
hooks_remembercall goes through it) while each PreToolUse/PostToolUse/SessionStart hook is a separate short-livedruvector hooks …CLI process reading the same file. The server's save therefore still opens the torn-file window #634 described; the fixed CLI side now detects the torn read and quarantines the file, but the store the user ends up with is whatever the hook wrote next.What we observed (ruvector 0.2.34 server, 2026-09-17): a 9 MB onnx-minilm-stamped store with 757 memories was being saved by the MCP server while ~12 parallel subagents fired PostToolUse hooks. A hook-side CLI read the half-written file, moved it to
intelligence.json.corrupt-1789669812640, and created a fresh store — hash-stamped, 64-dim, because the hook path resolved the hash embedder — sohooks reembed --dry-runthen reported an embedder-provenance MISMATCH against a store that had been fine an hour earlier. The quarantined file was complete and parsed with all 757 memories; the server still had the good snapshot in memory. The same sequence is reachable on 0.3.1 because the server-side write is unchanged.To Reproduce
ruvector mcp start, or via a Claude Code plugin) against a store of a few MB with vector memories.hooks_rememberthrough the server while running ~20 parallelnpx ruvector@0.3.1 hooks route "x"/hooks remember "y" -t testCLI processes against the same.ruvector/.intelligence.json.corrupt-<epoch>appears; the server's nextsave()then overwritesintelligence.jsonfrom memory, or a CLI process creates a fresh store first, depending on timing.Deterministic variant (no race): while the server is running, truncate
intelligence.jsonmid-file, then trigger a server-side operation that callsload()(e.g. restart the server) —load()returns empty defaults with no error and the nextsave()persists the empty store.Expected behavior
mcp-server.jsIntelligence.save()writes to a unique temp file in the same directory andrename()s over the target, ascli.jsdoes since fix(ruvector): atomic intelligence.json writes + fail-loud on corrupt load (#634) #698 (ideally by sharing that helper rather than a second copy).mcp-server.jsIntelligence.load()fails loud (or quarantines, likereadIntelStoreSafe) on an unparseable store instead of returning empty defaults that the nextsave()persists.rename()alone does not protect a reader that already opened the old inode before the rename (see Rename atomicity is not enough npm/write-file-atomic#64), so a parse-retry once (re-open + re-read) before quarantining would avoid quarantining a file that was simply replaced under the reader.Environment
writeFileSyncatbin/mcp-server.js:396Additional context
docs/solutions/integration-issues/ruvector-adr210-embedding-provenance-refusal.md, "Non-atomic store write race".