You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
graph-node: persisted writes invisible to stats()/query() across handles/processes; MATCH without label unimplemented #879
@ruvector/graph-node writes (createNode/createEdge/batchInsert) commit successfully to the redb-backed storage file (isPersistent() is true, the file grows on disk), but nothing that reads the graph afterwards — stats(), query(), kHopNeighbors(), searchHyperedges() — ever sees that data, even from a second handle in the same process, let alone a new process. This makes persistence effectively non-functional for any real (multi-process, or even multi-handle) usage pattern.
Confirmed on both the currently-installed 2.0.3 and the latest published 2.0.4.
Minimal repro
const{ GraphDatabase }=require('@ruvector/graph-node');constg1=newGraphDatabase({distanceMetric: 'Cosine',dimensions: 384,storagePath: './test.db'});consttx=awaitg1.begin();awaitg1.createNode({id: 'n1',embedding: newFloat32Array(384),labels: ['X'],properties: {}});awaitg1.commit(tx);console.log(awaitg1.stats());// { totalNodes: 1, ... } -- correct// Reopen -- same process, new handle:constg2=newGraphDatabase({distanceMetric: 'Cosine',dimensions: 384,storagePath: './test.db'});console.log(awaitg2.stats());// { totalNodes: 0, ... } -- wrong// Also wrong via the static factory, and in a brand new process:constg3=GraphDatabase.open('./test.db');console.log(awaitg3.stats());// { totalNodes: 0, ... } -- wrong
Root cause -- three stacked bugs in crates/ruvector-graph-node/src/lib.rs
Constructor never rehydrated from storage.GraphDatabase::new() and ::open() built graph_db via GraphDB::new() (empty in-memory graph) instead of GraphDB::with_storage(path) (see ruvector-graph/src/graph.rs:58, which correctly calls load_from_storage()).
Already fixed on main -- commit 31bb94401 ("fix: integrate latest PRs and issue regressions", 2026-08-12) added a hydrate_from_storage() call to both constructors.
Not yet published. Latest npm release is 2.0.4, published 2026-05-06 -- three months before that fix landed. npm/packages/graph-node/package.json on main still reads 2.0.4, so there's no released version with this fix yet.
Read paths (stats, query, kHopNeighbors, searchHyperedges) all read from self.hypergraph (a CoreHypergraphIndex), a second in-memory index entirely separate from graph_db/storage. It's populated only as a side effect of create_node/create_edge/batch_insert on that exact struct instance, with no persistence of its own. The hydrate_from_storage() fix above now populates it too, so stats() should be correct once that commit is released.
// If no labels specified, return all nodes (simplified)if node_pattern.labels.is_empty() && node_pattern.variable.is_some(){// This would need iteration over all nodes - for now just stats}
This branch does nothing. There's also no WHERE-clause evaluation anywhere in query(). Since MATCH (n) WHERE n.id = '...' RETURN n (no label filter) is the standard query shape for point lookups and most real usage, query() will keep returning an empty node list for it even after Implement Ruvector high-performance vector database #1/Set up Claude Flow swarm initialization #2 ship -- stats() would report correct counts, but query() results would not.
Summary
@ruvector/graph-nodewrites (createNode/createEdge/batchInsert) commit successfully to the redb-backed storage file (isPersistent()istrue, the file grows on disk), but nothing that reads the graph afterwards —stats(),query(),kHopNeighbors(),searchHyperedges()— ever sees that data, even from a second handle in the same process, let alone a new process. This makes persistence effectively non-functional for any real (multi-process, or even multi-handle) usage pattern.Confirmed on both the currently-installed
2.0.3and the latest published2.0.4.Minimal repro
Root cause -- three stacked bugs in
crates/ruvector-graph-node/src/lib.rsConstructor never rehydrated from storage.
GraphDatabase::new()and::open()builtgraph_dbviaGraphDB::new()(empty in-memory graph) instead ofGraphDB::with_storage(path)(seeruvector-graph/src/graph.rs:58, which correctly callsload_from_storage()).main-- commit31bb94401("fix: integrate latest PRs and issue regressions", 2026-08-12) added ahydrate_from_storage()call to both constructors.2.0.4, published 2026-05-06 -- three months before that fix landed.npm/packages/graph-node/package.jsononmainstill reads2.0.4, so there's no released version with this fix yet.Read paths (
stats,query,kHopNeighbors,searchHyperedges) all read fromself.hypergraph(aCoreHypergraphIndex), a second in-memory index entirely separate fromgraph_db/storage. It's populated only as a side effect ofcreate_node/create_edge/batch_inserton that exact struct instance, with no persistence of its own. Thehydrate_from_storage()fix above now populates it too, sostats()should be correct once that commit is released.query()'s CypherMATCHhandling is still a stub for the label-less case -- confirmed still broken onmainas of this writing, independent of Implement Ruvector high-performance vector database #1/Set up Claude Flow swarm initialization #2. In theStatement::Matcharm:This branch does nothing. There's also no
WHERE-clause evaluation anywhere inquery(). SinceMATCH (n) WHERE n.id = '...' RETURN n(no label filter) is the standard query shape for point lookups and most real usage,query()will keep returning an empty node list for it even after Implement Ruvector high-performance vector database #1/Set up Claude Flow swarm initialization #2 ship --stats()would report correct counts, butquery()results would not.Ask
@ruvector/graph-noderelease once the persistence fix (Implement Ruvector high-performance vector database #1/Set up Claude Flow swarm initialization #2, already onmain) and a real implementation of label-lessMATCH+WHEREfiltering (Reorganize repo structure and update documentation #3) are both in.main-history hygiene, but flagging: consumers of this crate that pin an older commit (we did, ~2,964 commits behind) won't get Implement Ruvector high-performance vector database #1/Set up Claude Flow swarm initialization #2 either without an explicit update.