Skip to content

graph-node: persisted writes invisible to stats()/query() across handles/processes; MATCH without label unimplemented #879

Description

@anorlandrtc

Summary

@ruvector/graph-node writes (createNode/createEdge/batchInsert) commit successfully to the redb-backed storage file (isPersistent() is true, the file grows on disk), but nothing that reads the graph afterwards — stats(), query(), kHopNeighbors(), searchHyperedges() — ever sees that data, even from a second handle in the same process, let alone a new process. This makes persistence effectively non-functional for any real (multi-process, or even multi-handle) usage pattern.

Confirmed on both the currently-installed 2.0.3 and the latest published 2.0.4.

Minimal repro

const { GraphDatabase } = require('@ruvector/graph-node');

const g1 = new GraphDatabase({ distanceMetric: 'Cosine', dimensions: 384, storagePath: './test.db' });
const tx = await g1.begin();
await g1.createNode({ id: 'n1', embedding: new Float32Array(384), labels: ['X'], properties: {} });
await g1.commit(tx);
console.log(await g1.stats()); // { totalNodes: 1, ... } -- correct

// Reopen -- same process, new handle:
const g2 = new GraphDatabase({ distanceMetric: 'Cosine', dimensions: 384, storagePath: './test.db' });
console.log(await g2.stats()); // { totalNodes: 0, ... } -- wrong

// Also wrong via the static factory, and in a brand new process:
const g3 = GraphDatabase.open('./test.db');
console.log(await g3.stats()); // { totalNodes: 0, ... } -- wrong

Root cause -- three stacked bugs in crates/ruvector-graph-node/src/lib.rs

  1. Constructor never rehydrated from storage. GraphDatabase::new() and ::open() built graph_db via GraphDB::new() (empty in-memory graph) instead of GraphDB::with_storage(path) (see ruvector-graph/src/graph.rs:58, which correctly calls load_from_storage()).

    • Already fixed on main -- commit 31bb94401 ("fix: integrate latest PRs and issue regressions", 2026-08-12) added a hydrate_from_storage() call to both constructors.
    • Not yet published. Latest npm release is 2.0.4, published 2026-05-06 -- three months before that fix landed. npm/packages/graph-node/package.json on main still reads 2.0.4, so there's no released version with this fix yet.
  2. Read paths (stats, query, kHopNeighbors, searchHyperedges) all read from self.hypergraph (a CoreHypergraphIndex), a second in-memory index entirely separate from graph_db/storage. It's populated only as a side effect of create_node/create_edge/batch_insert on that exact struct instance, with no persistence of its own. The hydrate_from_storage() fix above now populates it too, so stats() should be correct once that commit is released.

  3. query()'s Cypher MATCH handling is still a stub for the label-less case -- confirmed still broken on main as of this writing, independent of Implement Ruvector high-performance vector database #1/Set up Claude Flow swarm initialization #2. In the Statement::Match arm:

    // If no labels specified, return all nodes (simplified)
    if node_pattern.labels.is_empty() && node_pattern.variable.is_some() {
        // This would need iteration over all nodes - for now just stats
    }

    This branch does nothing. There's also no WHERE-clause evaluation anywhere in query(). Since MATCH (n) WHERE n.id = '...' RETURN n (no label filter) is the standard query shape for point lookups and most real usage, query() will keep returning an empty node list for it even after Implement Ruvector high-performance vector database #1/Set up Claude Flow swarm initialization #2 ship -- stats() would report correct counts, but query() results would not.

Ask

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions