Skip to content

feat(ruvector-py): Python SDK, CLI, and MCP server (ADR-352) - #1117

Draft
ruvnet wants to merge 42 commits into
mainfrom
feat/python-sdk
Draft

ruvnet wants to merge 42 commits into
mainfrom
feat/python-sdk

Conversation

@ruvnet

@ruvnet ruvnet commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Summary

Full-scope Python SDK + CLI + MCP server for RuVector, backed by the Rust core via PyO3/maturin,
publish-ready for PyPI but not published, merged, or tagged (boundary unchanged throughout —
see "Not done" below). ADR-352 (docs/adr/ADR-352-ruvector-python-sdk-cli-mcp.md) is the living
design doc: capability-coverage table, two independent benchmark write-ups vs hnswlib, security
findings, the Salesforce research that changed the integration plan, and publishing-prep
verification. docs/sdk/LOOP-STATE.md is the resume-point/gotchas doc if this needs to continue
in a future session.

⚠️ One commit in this PR must NOT land until ruvector is live on PyPI: docs(readme): add Python install + user guide link adds a pip install ruvector / uv add ruvector snippet to
the root README. That snippet doesn't work yet — nothing has been published. A PR can't be
merged minus a single commit, so either hold this whole PR's merge until a real release exists,
or git rebase -i that one commit out before merging the rest.

Capability coverage (vs the npm ruvector package's surface)

Capability Status Notes
RaBitQ quantized index ✅ done RabitqIndex, Collection(backend="rabitq")
HNSW index (default) ✅ done HnswIndex, Collection(backend="hnsw")
Metadata filtering ✅ done (partial) dict-filter in Rust for hnsw; callable filter always Python-side
Collections (CRUD, persistence) ✅ done own format (.rbpx/.npy + JSON sidecar), not RVF
CLI ✅ done, polished Rich tables, colored/ranked output, progress spinner, fast startup
MCP server (stdio + HTTP) ✅ done bearer auth via RUVECTOR_MCP_TOKEN, read/write scopes, live-verified
ChatGPT ui:// widget ✅ done vector_explore tool, live-verified _meta shape
Graph (CRUD + Cypher MATCH/WHERE) ✅ done (partial) ported ruvector-graph-node's executor; RETURN parsed not projected, matches upstream
GNN forward-pass rerank ✅ done untrained = random projection, stated explicitly
Attention rerank ✅ done returns blended vector + raw weights
k-means clustering ✅ done the real ML algorithm (ruvector-cluster-rag), not the distributed-infra crate of a similar name
SONA (inference-only) ✅ done fresh engine = exact identity transform, stated as fact
LangChain VectorStore ✅ done verified vs langchain-core 1.6.6
LlamaIndex VectorStore ✅ done verified vs llama-index-core 0.14.25; found+fixed a real distance-vs-similarity bug
Salesforce Agentforce ✅ done (External Services + OpenAPI path) see below — mocked APIs only, no real org
Rust unit tests ✅ done, now runs in CI 49 tests; found+fixed a real gap where they ran nowhere in CI
RVF persistence ❌ not practical this session ~30-file subsystem; stated reason in ADR, not silent
Embeddings (M3) ❌ deferred needs ONNX/ort + model download, separate milestone by design
Quantization (Turbo4) ⬜ unexplored exists in ruvector-core, not evaluated
ruvector-cluster (distributed) ❌ premise mismatch sharding/consensus infra, not ML clustering — see k-means row

Salesforce Agentforce — what changed from the original plan, and why

Researched Agentforce's actual extension points before building (per task instruction, not from
memory): Data Cloud's "bring your own retriever" isn't a drop-in external-vector-store plug
(it queries Data Cloud's own indexed data), and Agentforce's MCP client is Beta and AE-tier
gated
(not self-service, and this session has no such org access to verify against anyway).
External Services + OpenAPI 3.0 custom actions is the extension point that's actually
self-service — built that as the primary path instead.

Ships: OAuth2 client-credentials auth, paginated SOQL record fetch, record→collection sync, three
Agentforce-callable actions (search/upsert/ground), and a generated flat OpenAPI 3.0 doc an
External Service imports. Mounted as custom HTTP routes on the same server/port as MCP — which
needed its own manual bearer check, since MCPServer.custom_route doesn't go through the SDK's
token_verifier (confirmed by reading the SDK source first, not assumed).

Verified / not verified against a real org, stated explicitly per the task's instruction: all
28 tests run against httpx.MockTransport or an in-process ASGI TestClient — no real
Salesforce org, OAuth exchange, SOQL query, or External Service registration happened anywhere in
this session. examples/python-salesforce-agentforce/ has a Named Credential + External Service
Registration metadata sketch and a setup README, every file headed "illustrative, not validated
against a real org or the current Metadata API schema" — a GenAiFunction sketch was
deliberately NOT included (a newer metadata type whose current schema this session had no way to
confirm; fabricating one would misrepresent confidence this session doesn't have).

Real bugs found and fixed by actually running things, not just unit-testing happy paths

  • #[pyclass(unsendable)] RabitqIndex panicked under the MCP server's threaded dispatch — fixed.
  • A silently-dead ef_search per-call kwarg on HnswIndex.search() (the Rust struct field
    existed; nothing ever read it) — found via a suspicious benchmark number, removed.
  • Collection.search() over-fetched 4x from Rust even with filter=None — added a fast path
    (~17% faster at ef_search=50).
  • patches/hnsw_rs had a println! every 50k inserts that would corrupt the stdio MCP
    transport's framing — removed.
  • The LlamaIndex adapter returned raw distance where llama-index expects similarity — would have
    silently inverted ranking under any similarity-cutoff filter. Fixed + regression-tested.
  • Salesforce routes leaked CollectionError/KeyError as raw 500s instead of clean 4xx bodies —
    found via live curl, not just unit tests. Also narrowed an over-broad except Exception in
    the same handler that would have mis-reported an unrelated real bug as a 404 with its internal
    message exposed.
  • release.yml's own comment claimed ruvector-py's 49 Rust unit tests "run in python-wheels.yml
    instead" — checked the actual workflow file and they ran nowhere in CI. Added a job that runs
    them (this PR's own CI run, in flight, is the first real test of it).
  • A module-reload staleness bug in the Salesforce route test fixture itself (not the shipped
    code) — a stale from .mcp_server import server binding after a test reloaded that module.
  • Plus the earlier-session bugs: an empty-collection existence-check gap, a silently-ignored
    custom-ids kwarg, and a concurrency race reproduced as 3 id collisions out of 160 under load
    before fixing with a process-wide lock.

Not done / explicitly out of scope

  • RVF persistence, embeddings (M3), Turbo4 quantization evaluation — each has a stated reason in
    the ADR's capability table, not a silent gap.
  • Agentforce MCP registration (Beta/AE-tier-gated — documented as a future option, not built).
  • The 3 known_limitations bugs in hnsw.rs's JSON converter (large-int precision loss,
    NaN→null, lone-surrogate error message) — characterized by tests, not fixed.
  • No PyPI publish, no merge, no tags — unchanged boundary for the entire session.

Test plan

  • 201/201 pytest, 49/49 cargo test -p ruvector-py, mypy --strict clean, clippy -D warnings clean
  • cargo audit / pip-audit / npx @claude-flow/cli@latest security scan — all clean
    (one unrelated transitive finding recorded: nltk PYSEC-2026-3740 via llama-index-core,
    not reachable through ruvector's own code paths)
  • Two independent benchmark passes vs hnswlib (random-Gaussian and real 20newsgroups/
    MiniLM embeddings) — honest gap reported both times, not hidden
  • MCP HTTP auth live-verified with real curl (no token → 401, wrong token → 401, correct →
    200; write-scope enforcement checked directly)
  • Salesforce routes live-verified with real curl before the final fix (caught the 500-leak)
  • Fresh-venv wheel install + full suite pass (no dev/editable carryover)
  • This PR's own CI run (Linux aarch64/macOS/Windows wheel legs + the new Rust-tests job) is
    in flight as of this push — not yet confirmed green, check gh pr checks 1117

Publishing (when the user approves — not done here)

  1. One-time, needs repo admin: register this repo/workflow as a PyPI trusted publisher.
  2. Land the docs(readme): commit (see the warning at the top) only after step 3.
  3. gh pr ready 1117 && gh pr merge 1117 --squash
  4. git tag python-v0.1.0 && git push --tags — triggers the publish job.

🤖 Generated with claude-flow

ruvnet added 30 commits October 1, 2026 15:43
Restores the crates/ruvector-py PyO3 binding from the stale
feature/python-sdk-m1 branch (926 commits behind main) and re-validates
it against current main as the foundation for the broader Python
SDK/CLI/MCP effort (ADR-352, follow-up commit).

- Added crates/ruvector-py to workspace members.
- Bumped pyo3 0.22 -> 0.29.3, numpy 0.22 -> 0.29.0 (6 months of drift):
  Python::allow_threads -> Python::detach, Python::get_type_bound ->
  Python::get_type. No other API drift against ruvector-rabitq v2.3.0
  (from_vectors_parallel, search_with_rerank, persist::{save,load}_index,
  export_items all unchanged).
- Verified: cargo build -p ruvector-py (clean), cargo clippy --all-targets
  --no-deps -D warnings (clean), maturin develop --release (wheel builds,
  installs), pytest tests/ (7/7 pass).

Co-Authored-By: claude-flow <ruv@ruv.net>
Supersedes the SCOPE (not the binding strategy) of docs/sdk/01-06.md:
those docs deliberately excluded a CLI and MCP server. This ADR adds
both plus a generic Collection surface, per the actual ask, while
keeping docs/sdk/02-strategy.md's PyO3+maturin+abi3 decision intact
(bumped pyo3/numpy pins to current, see prior commit).

Includes live-verified evidence of the ChatGPT Apps SDK ui:// widget
_meta convention (openai/outputTemplate + ui.resourceUri +
openai/widgetAccessible), fetched from
web-based-chatgpt-mcp-starter.ruv.chatgpt.site/api/mcp rather than
assumed from memory.

Adds docs/sdk/LOOP-STATE.md as the cross-session resume point.

Co-Authored-By: claude-flow <ruv@ruv.net>
Adds the three Rust-side primitives Collection needs that M1 didn't
expose: RabitqIndex.add (true incremental, single row), .add_batch
(GIL-released loop), and .export_items (round-trips the owned
vectors so a rebuild-without-tombstones is possible from Python).

python/ruvector/collection.py is the thin pure-Python layer per
CLAUDE.md's "core logic in Rust" rule: ids, a JSON metadata sidecar,
client-side filter predicates (dict-exact-match or callable,
documented as not pushed into the Rust scan), soft delete via
tombstones, and vacuum() to physically rebuild without them (the
underlying index has no delete).

Verified: 21/21 pytest (7 M1 smoke + 14 new Collection tests covering
bulk build, metadata, both filter forms, delete/vacuum,
insert/insert_batch, empty-collection lazy build, dim-mismatch
errors, duplicate-id rejection, save/load roundtrip incl. empty
collections and a missing-sidecar error path).

Co-Authored-By: claude-flow <ruv@ruv.net>
ruvector.cli: click-based console script (create/insert-batch/search/
delete/export/import/info/benchmark/serve), lazy per-subcommand
imports.

ruvector.mcp_server: MCP server on the official `mcp` SDK (v2's
MCPServer, not the v1 FastMCP name) exposing the Collection surface
as tools (vector_create_collection/insert/insert_batch/search/delete/
stats/list_collections) plus vector_explore, a ui:// widget tool
whose _meta matches the shape live-verified in ADR-352. Collection
names (not paths) are the only tool argument that touches the
filesystem; _safe_path restricts them to [A-Za-z0-9_-] under
RUVECTOR_MCP_DATA_DIR, so there is no traversal surface regardless of
what a remote client sends.

Two real bugs found and fixed by actually running this end-to-end
rather than trusting it to work:

1. RabitqIndex was `#[pyclass(unsendable)]` in the M1 salvage. The
   MCP SDK dispatches sync tool handlers to worker threads, so a
   Collection cached across calls panicked the instant two calls
   landed on different threads ("unsendable, but sent to another
   thread") - caught by tests/test_mcp_server.py before it ever hit
   a live server. Fixed by removing `unsendable`: the wrapped
   RabitqPlusIndex already satisfies AnnIndex: Send + Sync, so pyo3's
   auto-derive gives a real Send+Sync pyclass for free.
2. Both the CLI's `create` and the MCP server's duplicate-collection
   check tested existence of the raw .rbpx index path, which an
   empty (just-created, zero-vector) collection never writes - only
   its .meta.json sidecar does. A second `create` call on an empty
   collection silently succeeded instead of erroring. Fixed by adding
   Collection.meta_path() and checking that everywhere existence is
   tested (collection.py, cli.py, mcp_server.py).

Fast-startup fix: ruvector/__init__.py eagerly imported Collection
(and therefore numpy) at package-import time, which `ruvector --help`
paid for even though no subcommand needs it. Switched to PEP 562
`__getattr__` lazy attribute resolution. Measured via `python -X
importtime -c "from ruvector.cli import main"`: 72.8ms -> 16.1ms;
`ruvector --help` wall time ~20ms.

Verified live, not just via in-process test harness: `ruvector serve`
over real stdio (piped JSON-RPC initialize/tools-list) and `ruvector
serve --http` (curl against a real uvicorn process) both answer
correctly with no stderr errors.

38/38 pytest passing (7 M1 + 14 Collection + 6 MCP-specific +
6 CLI-specific, plus the widget-meta and traversal-rejection tests).

Co-Authored-By: claude-flow <ruv@ruv.net>
Splits _native.pyi out of __init__.pyi (the compiled extension needs
its own stub module so `from ._native import ...` in collection.py
resolves — without it every RuVectorError subclass silently typed as
Any and masked real errors). Declares Collection's _dim/_rerank_factor
/_seed at class level so mypy recognizes attributes set outside
__init__ by the create/from_vectors/load constructors. Fixes a real
narrowing bug: the filter param (Dict | Callable) was branched with
`callable(filter)`, which mypy can't use to exclude Callable from a
Dict|Callable union - switched to `isinstance(filter, dict)`, which
it can. Adds ToolAnnotations instances in mcp_server.py instead of
bare annotations dicts (pydantic coerced them at runtime, but they
never had a declared type). Generic-args cleanup (set[int],
os.PathLike[str], tuple[int, ...]) and type annotations on every test
fixture/helper that `mypy --strict tests/` flagged.

`mypy --strict` now passes clean on collection.py, cli.py,
mcp_server.py, __init__.py, and all 4 test files. 38/38 pytest still
green after the rebuild.

Co-Authored-By: claude-flow <ruv@ruv.net>
ADR-352 now carries real measured results instead of a methodology
promise:

- Benchmark vs hnswlib (installed fresh this session) at n=10k/100k,
  dim=128, random-Gaussian data: recall/latency tables for both a
  default-settings comparison and a rerank_factor/ef sweep, with an
  explicit caveat that random-Gaussian is close to worst-case for
  both RaBitQ and HNSW (not representative of real embedding
  clusters) and that this session's recall numbers don't match
  ruvector-rabitq/BENCHMARK.md's figures (different host/workload) -
  flagged as an honest discrepancy, not swept under the rug.
- Security: cargo audit (0 findings in ruvector-py's own dependency
  tree; 2 unrelated findings elsewhere in the 300+ crate workspace
  lockfile), pip-audit (0 findings - caught that the standalone
  pip-audit binary audits the system Python rather than an active
  venv, re-ran via `python3 -m pip_audit`), claude-flow security
  scan (0 findings), mypy --strict (clean, 2 real bugs fixed getting
  there).
- Publishing: `maturin build --release` produces a 345 KiB wheel
  (budget: 8 MiB); installed into a brand-new venv (no editable/dev
  carryover) and ran the full suite against it, 38/38 pass. Added
  .github/workflows/python-wheels.yml (5-platform matrix via
  PyO3/maturin-action, not cibuildwheel as 02-strategy.md originally
  said - rationale in the workflow's header comment) with trusted
  PyPI publishing gated on a tag or explicit dispatch input; not run
  in CI yet, nothing published.

docs/sdk/LOOP-STATE.md updated with the full verified/next/deferred
breakdown and a gotchas list (unsendable pyclass thread panics,
mcp>=2 FastMCP->MCPServer rename, pydantic snake_case vs wire
camelCase, pip-audit's system-vs-venv trap, hnswlib's libomp-dev
build dependency, the empty-collection sidecar-vs-index-path bug)
so a future session doesn't re-discover any of it.

Co-Authored-By: claude-flow <ruv@ruv.net>
…-ready

Found by a second-pass review of the finished M1.5 work, not by the
author's own tests (which all passed while these were live):

1. RabitqIndex.build(ids=...): Collection.from_vectors accepted an
   `ids` kwarg but the Rust build() had no way to honor it, so an
   earlier cut silently remapped search() results back to row
   indices instead of the caller's ids. Added `ids: Option<...>` to
   build() (validated to fit u32, RabitqPlusIndex's storage width,
   with a clear ValueError instead of ruvector-rabitq's unchecked
   `id as u32` truncation a layer down) and the matching guard in
   add_batch. Collection.from_vectors now passes ids straight
   through; 3 new tests cover custom ids in search results, custom
   ids with metadata, and the u32-overflow rejection.
2. README was still the M1-era RaBitQ-only version - the published
   long description would have said nothing about Collection, the
   CLI, or the MCP server. Rewritten with install instructions for
   every extra, three worked examples (all three executed for real
   against the built wheel before being written down, not
   transcribed from memory), the full CLI command list, the MCP
   tool list + its security model, and a link to ADR-352's real
   benchmark table instead of restating it.
3. stubs/ruvector/__init__.pyi was a stale duplicate of the old
   __init__.pyi, already out of sync with the split into
   __init__.pyi + _native.pyi - deleted (it was never in maturin's
   include list, so nothing shipped from it; the pyproject.toml
   comment claiming otherwise is fixed too).

Also: python-wheels.yml's "aarch64 can't run under QEMU" test-skip
was wrong - ubuntu-24.04-arm is a native ARM64 runner, not emulated,
so both matrix legs now run their own tests. Verified (not assumed)
that polars and pydantic-core use PyO3/maturin-action rather than
cibuildwheel via `gh api .../contents/.github/workflows/...`. Added
crates/ruvector-py/.claude/ to .gitignore (written by the claude-flow
security scan run from that directory). Fetched the fourth/last
research URL (signal-to-swarm.ruv.chatgpt.site/live.html) that an
earlier commit's research pass had missed - confirmed irrelevant,
noted in ADR-352 rather than left silently unfetched.

41/41 pytest, mypy --strict clean, cargo clippy clean, all README
examples re-verified against the rebuilt wheel.

Co-Authored-By: claude-flow <ruv@ruv.net>
… floor

A second review pass found two issues after the PR was already open:

1. `_cache`, `Collection._next_id`, `_metadata`, and `_tombstones` had
   no locking, and `ruvector serve --http` dispatches concurrent tool
   calls to worker threads - the same fact that made the old
   `unsendable` RabitqIndex pyclass panic earlier in this session.
   Two concurrent `vector_insert` calls on one collection could read
   the same `_next_id` and both `add()` with the same id. Fixed with
   one process-wide `threading.RLock()` (RLock, not Lock -
   vector_explore calls vector_search from the same thread) around
   every tool body touching shared state. Verified the fix is real:
   temporarily neutered the lock and reproduced 3 duplicate ids out
   of 160 under 16 threads x 10 inserts each
   (test_concurrent_inserts_do_not_collide); restoring the lock
   eliminates the collisions. stdio (the default transport) was
   never affected - one request at a time regardless.
2. `pyproject.toml`'s `mcp` extra floor was `mcp>=1.6`, but
   mcp_server.py imports `mcp.server.mcpserver.MCPServer`, which only
   exists from the v2 rename (v1 called it `FastMCP` - the exact
   error this session hit earlier and had to work around). A fresh
   `pip install ruvector[mcp]` on mcp 1.x would resolve cleanly and
   then fail on import. Bumped the floor to `mcp>=2.0`.

42/42 pytest (was 41 - one new concurrency regression test), mypy
--strict clean.

Co-Authored-By: claude-flow <ruv@ruv.net>
…ltering

Per rUv's follow-up: "HNSW-backed Collection... metadata filtering done
in Rust, not Python... make the default Collection the fast path."

New crates/ruvector-py/src/hnsw.rs wraps
ruvector_core::vector_db::VectorDB (confirmed Send+Sync through the
VectorIndex: Send+Sync supertrait bound, same check already done for
RabitqIndex before dropping `unsendable` — see ADR-352's capability
inventory). Exposes insert/insert_batch/search/delete/export_items,
all operating on string ids with an arbitrary JSON-compatible metadata
dict per vector.

What this gets "for free" vs the Python-side filtering the M1.5
Collection did: VectorDB::search's SearchQuery.filter applies the
equality filter inside the Rust call (still a bolt-on retain-after-
search, not pushed into the HNSW graph traversal itself -
ruvector-core doesn't wire up hnsw_rs's lower-level search_filter
either - documented honestly in the module docstring rather than
oversold).

ruvector-core added as a path dependency with default-features=false
and an explicit feature list (storage, hnsw, simd, parallel) to avoid
pulling in api-embeddings' reqwest/rustls for a capability this crate
never touches. Always constructed with a memory:// storage path -
persistence goes through export_items() + Python's existing save/load
sidecar (the same idiom RabitqIndex uses), not VectorDB's own
redb-backed persistent mode, to keep one persistence story across
both backends.

Hand-rolled a minimal Python<->serde_json::Value converter
(str/int/float/bool/None/list/dict) rather than adding the
`pythonize` crate - small enough to fully control and avoids a new
dependency's own version-pinning risk against pyo3 0.29.

Verified with a manual smoke script (create/insert/insert_batch/
search/filtered search/delete/export_items all round-trip correctly)
before writing the pytest suite in the next commit. cargo build +
clippy --all-targets --no-deps -D warnings both clean. Existing 42
pytest still pass (RabitqIndex path untouched).

Co-Authored-By: claude-flow <ruv@ruv.net>
Per rUv's follow-up: "make the default Collection the fast path."
Collection.create()/from_vectors() now take backend="hnsw" (default)
or backend="rabitq" (the original M1.5 backend, still fully
supported). Every Collection method branches on backend:

- insert/insert_batch: hnsw calls HnswIndex directly (string ids,
  no lazy-build - HnswIndex.create() works immediately at n=0,
  unlike RabitqPlusIndex which needs >=1 vector to fit a rotation).
- search: hnsw pushes dict-based exact-match filters into Rust
  (HnswIndex.search(filter=...)) - the literal fix for "filtering in
  Rust, not Python." Still runs the overfetch-widen loop in Python
  (VectorDB's own filter has no overfetch for selectivity, so a
  naive single-shot call could return fewer than k when it didn't
  need to) - only the per-row equality test moved to Rust, not the
  loop orchestration. A callable predicate still falls back to
  fetching unfiltered candidates and testing them in Python on
  *both* backends - shipping an arbitrary Python function across the
  PyO3 boundary isn't possible, a structural limit, not a gap.
- delete: hnsw is a real delete (HnswIndex.delete); rabitq stays a
  tombstone (no delete on the underlying index).
- vacuum: hnsw always returns 0 (nothing queued to drop - documented
  as not a bug); rabitq keeps its existing rebuild-without-tombstones
  behavior.
- save/load: hnsw has no native on-disk format (always in-memory -
  see src/hnsw.rs), so the vectors go to `path` via a plain
  `np.save`/`np.load` on an open file handle and ids/metadata/config
  go in the `.meta.json` sidecar - same two-file split as rabitq, for
  one persistence story across backends. A sidecar with no "backend"
  key (written before this commit) loads as rabitq, the only backend
  that existed then.

CollectionStats gained a `backend` field.

Test suite restructured: most tests now parametrize over both
backends (`@pytest.mark.parametrize("backend", ["hnsw", "rabitq"])`)
since the behavior is meant to be identical; backend-specific
semantics get their own explicitly-pinned tests (rabitq's tombstone+
vacuum vs hnsw's immediate real delete; rabitq's u32 id ceiling vs
hnsw's unbounded string ids; hnsw's metric options).

Two real bugs this surfaced and fixed while making it work, not just
pass review:
1. The MCP server's default-backend-created collections now report
   vacuumed=0 on delete+vacuum (hnsw has nothing to reclaim) -
   test_mcp_server.py's expectation was still written for the old
   rabitq-tombstone default; fixed the test, not the behavior.
2. The initial Rust-filter wiring called
   HnswIndex.search(qvec, k, filter=filter) once with no overfetch,
   which under-returned for a selective filter (VectorDB's own filter
   fetches exactly k ANN candidates *before* filtering, no
   over-fetch at that layer) - caught by
   test_search_dict_filter[hnsw] expecting 5 hits and getting 3.
   Fixed by keeping the overfetch-widen loop in Python around the
   Rust-filtered call.

63/63 pytest (was 42), mypy --strict clean, cargo clippy clean.
README's existing from_vectors(...)/search(...) example re-verified
against the rebuilt wheel - still runs, now on the hnsw backend by
default (rerank_factor=20 in that example is now a harmlessly-unused
rabitq-only kwarg - README update queued).

Co-Authored-By: claude-flow <ruv@ruv.net>
Per a second advisor review of the first HNSW-backend benchmark
result (6-8x slower than hnswlib looked too large for a same-
algorithm comparison):

1. HnswIndex.search()'s per-call `ef_search` override silently did
   nothing - SearchQuery.ef_search exists on the Rust struct but
   VectorDB::search never reads it (calls the generic
   VectorIndex::search(query, k) trait method, which has no ef
   parameter at all). Removed the misleading parameter rather than
   ship a kwarg that's dead code; ef_search is construction-time-only
   (HnswIndex.create(..., ef_search=...)), documented in the Rust doc
   comment and the .pyi stub.
2. Collection.search() over-fetched k*4 candidates from Rust on every
   *unfiltered* search, even with filter=None - the widen-on-
   undersupply loop has no reason to run when there's nothing to
   widen for. Added a fast path: filter=None now asks for exactly k.
   Measured: p50 0.304ms -> 0.252ms at ef_search=50 on real
   embeddings (~17% faster, not the dominant cost but real).
3. patches/hnsw_rs/src/hnsw.rs (the workspace's [patch.crates-io]
   hnsw_rs, used by ruvector-core and ruvector-graph) had a hardcoded
   `println!` to stdout every 50,000 points inserted - corrupts any
   consumer framing a protocol over stdout, which is exactly what
   `ruvector serve`'s stdio MCP transport does. Removed it (an
   adjacent `trace!` already logs the identical message through the
   `log` facade instead of a hardcoded stdout write). Verified with a
   60,000-vector insert_batch producing zero stray stdout lines.

Re-ran the benchmark on real text embeddings (10k docs,
sklearn's 20newsgroups, all-MiniLM-L6-v2, 384-dim - not the
random-Gaussian workload from the earlier RabitqPlus benchmark,
which had flagged itself as an adversarial worst case). Recall is
now ~identical between ruvector and hnswlib at matched m/ef (both
0.999-1.0) - confirms the algorithm itself isn't the issue. Latency
gap narrowed from ~8.6x to ~5.7-7.2x after fix #2 but didn't close;
the residual is named honestly as unprofiled structural cost
(per-query Vec<f32> copy, per-hit Python dict construction) rather
than guessed at - a batched search_many entry point is the next
lever, not implemented this session.

Full table + both bug writeups in ADR-352's new "Benchmark - M2,
HnswIndex backend" section. 63/63 pytest, mypy --strict and clippy
both still clean after the rebuild.

Co-Authored-By: claude-flow <ruv@ruv.net>
Adds the capability-coverage table (done/partial/not-done against
the npm package's keyword list) rUv asked for in the final report,
kept live from here on rather than reconstructed at the end. Records
the expanded queue from rUv's follow-up (Agentforce, Slack, docs,
remaining capabilities, MCP auth, CLI polish, Rust tests) and the
Salesforce research correction sent to team-lead.

Co-Authored-By: claude-flow <ruv@ruv.net>
…nverters, error mappers

cargo test -p ruvector-py previously ran 0 tests; all correctness coverage
went through the Python-side pytest suite, never exercising the Rust code
in isolation. Adds 26 #[test] fns across hnsw.rs and error.rs:

- parse_metric: every documented metric alias (Ok) + an unknown string (Err)
- py_to_json/json_to_py round trips: null, bool, negative int, float,
  unicode, empty string/list/dict, flat mixed list, deeply nested structure
- py_dict_to_json_map/json_map_to_py metadata round trip + non-string-key
  TypeError + unsupported-Python-type TypeError + bool-vs-int distinction
- error::to_pyerr / to_pyerr_core: verbatim message forwarding for several
  RabitqError/RuvectorError variants, same RuVectorError exception class
  for both backends

Required dropping "extension-module" from ruvector-py's pyo3 dependency
features (Cargo.toml) -- that Cargo feature disables linking against
libpython, which made any test touching the interpreter fail to link
(confirmed empirically: undefined PyExc_ValueError/Py_Initialize/etc. even
through PyErr's lazy exception construction). This is the fix pyo3 0.29's
own FAQ prescribes; pyproject.toml's `[tool.maturin] features =
["pyo3/extension-module"]` already re-adds the feature via maturin's own
flag for every real wheel build, so `maturin build`/`develop` are
unaffected -- only bare `cargo build`/`cargo test` behavior changes (now
links libpython, as a normal embedding binary would).

Three known limitations found and documented as regression tests (not
fixed, per task scope): large Python ints (>i64::MAX) silently lose
precision through py_to_json's i64->f64 fallback; float('nan') metadata
silently becomes JSON null; a lone-UTF16-surrogate str hits py_to_json's
generic "got str" TypeError with a slightly misleading message.

Co-Authored-By: claude-flow <ruv@ruv.net>
Found while verifying the rust-tests cherry-pick's wheel build: when
HnswIndex was added to the compiled _native module (M2 slice), it was
never added to __init__.py's lazy _NATIVE_NAMES dispatch set (PEP 562
__getattr__), nor to __all__ in either __init__.py or __init__.pyi.

`ruvector._native.HnswIndex` always worked (collection.py imports it
directly, which is why Collection(backend="hnsw") worked fine), but
`ruvector.HnswIndex` - the advertised top-level import per the README
and the ADR - raised AttributeError. mypy --strict never caught this
either, since __init__.pyi never declared it present. Fixed in all
three places; 63/63 pytest still green, mypy --strict still clean.

Co-Authored-By: claude-flow <ruv@ruv.net>
…ss, serve error check

search: --no-color flag (plus NO_COLOR/non-TTY auto-detection, verified),
rank-based row styling (best-first, metric-agnostic), metadata truncation
with ellipsis overflow guard. info: Rich panel over the real
CollectionStats fields (count/dim/backend/rerank_factor/memory_bytes/
tombstoned), with hnsw's always-zero rerank_factor/memory_bytes labeled
as not-applicable instead of a bare misleading 0.

insert-batch/import: wrap the single atomic insert_batch/from_vectors
call in an indeterminate Rich spinner rather than fake granular
progress — chunking was considered and rejected because it would change
the rabitq rotation's fit quality on partial data (see
_run_with_spinner's docstring).

create/delete/export/import/benchmark: consistent green/cyan secho for
success/info lines, also NO_COLOR-aware (click.secho doesn't check that
env var itself). serve --http: verified port-in-use/permission-denied
both already surface as a clean uvicorn log line (not a traceback) —
documented why a try/except OSError wrapper here would be dead code.

All --json paths remain exactly plain JSON (asserted no-ANSI explicitly
in tests). mypy --strict clean. ruvector --help importtime unchanged
(~17-21ms before and after, measured back-to-back).

Co-Authored-By: claude-flow <ruv@ruv.net>
Add ruvector.integrations.langchain.RuVectorStore (langchain_core.vectorstores.VectorStore,
verified against langchain-core==1.6.6) and ruvector.integrations.llamaindex.RuVectorStore
(llama_index.core.vector_stores.types.BasePydanticVectorStore, verified against
llama-index-core==0.14.25), both wrapping ruvector.Collection.

- LangChain adapter: add_texts/similarity_search/similarity_search_with_score/delete/
  from_texts, text+metadata convention (page_content under a text_key, merged with the
  doc's own metadata), str<->int id bridging for Collection's int-only id space.
- LlamaIndex adapter: add/query/delete/delete_nodes/clear, reuses node_to_metadata_dict/
  metadata_dict_to_node (remove_text=False, needed for a correct round trip), full
  MetadataFilters translation (nested filters, AND/OR/NOT, all FilterOperators) via a
  Python predicate passed to Collection.search's callable-filter path, DEFAULT-mode-only
  query (raises NotImplementedError for sparse/hybrid/MMR/learner modes rather than
  silently degrading them to plain kNN).
- Both import their framework at module level (documented as a deliberate, honest
  boundary -- subclassing a live ABC needs the real base class at class-definition time)
  with an ImportError + install hint if missing; plain `import ruvector` never imports
  either module.
- New pyproject.toml extras: langchain, llamaindex (floors pinned to the versions
  actually tested against), added to `all`.
- tests/test_langchain_integration.py, tests/test_llamaindex_integration.py: 23 new
  tests over a deterministic fixed-vocabulary embedding/basis (add, search, metadata,
  delete, filters, from_texts, and the ImportError-reachable path via a sys.modules
  sentinel that evicts the whole framework module family, not just its root -- a
  root-only sentinel was empirically a no-op once submodules were already cached).

63 pre-existing tests + 23 new = 86 passed. mypy --strict clean across the whole
python/ruvector package.

Co-Authored-By: claude-flow <ruv@ruv.net>
Found by actually running the cherry-picked adapters end-to-end
rather than trusting the subagent's pytest-only verification.

query()'s VectorStoreQueryResult.similarities was populated directly
from hit.score, which is a *distance* (Collection.search's documented
convention: lower = closer). Every llama-index consumer that sorts or
thresholds on "similarity" (SimilarityPostprocessor's
similarity_cutoff, retrievers) assumes higher = closer - left as-is,
the two closest hits would be the first ones a similarity_cutoff
filter drops. The LangChain adapter in the same cherry-picked commit
already got this right (documented why it refuses to implement
similarity_search_with_relevance_scores rather than guess a
conversion); the LlamaIndex side just didn't carry the same care over.

Fixed with _distance_to_similarity(distance, metric): exact 1-distance
inverse for cosine (ruvector_core::encoding::metric_distance's cosine
branch is literally (1 - cos_sim).max(0), so this is provably exact,
not a guess), and a documented-as-not-normalized 1/(1+distance) for
every other metric (monotonic, bounded, preserves rank order, but
explicitly NOT claimed as a principled similarity for those metrics).

Needed a new public Collection.metric property to know which
conversion applies - also fixes a latent inconsistency where the
rabitq backend accepted a metric= constructor kwarg it silently never
applied (RabitqPlusIndex always scores via squared L2); .metric now
reports the backend's real behavior instead of echoing the ignored
kwarg back.

Added a regression test that would have caught this (asserts
descending, bounded-[0,1], near-1.0-for-exact-match similarities -
the previous test's `== pytest.approx(0.0, ...)` assertion had
encoded the bug and had to be corrected too), plus direct unit tests
for _distance_to_similarity and Collection.metric.

Also fixes a real CI risk in the cherry-picked Cargo.toml change
(dropping pyo3's "extension-module" feature so `cargo test -p
ruvector-py` can link against libpython and actually run - correct
and necessary for 26 new Rust unit tests to work): .github/workflows/
release.yml's `validate` and `build-crates` jobs run `cargo build
--workspace --release` and `cargo test --workspace [...]` on bare
ubuntu-22.04 runners with no Python setup step. Without
"extension-module", building ruvector-py's cdylib there needs a
system libpython to link against, which is unverified on that image.
Excluded ruvector-py from all four workspace-wide build/test
invocations in both jobs (`--exclude ruvector-py`) rather than risk
breaking CI for the whole repo on an unverifiable assumption -
ruvector-py's own build+test already happens correctly in
python-wheels.yml, which does set up Python.

93/93 pytest, mypy --strict clean, clippy clean, 26/26 cargo test
still passing after the Cargo.toml change.

Co-Authored-By: claude-flow <ruv@ruv.net>
…pdated

Co-Authored-By: claude-flow <ruv@ruv.net>
Per advisor review: checked gh pr checks 1117 instead of continuing
to stack work on top of an unverified CI run (8 pushes in, never
looked). Found two real, fixable failures:

1. `cargo fmt --all -- --check` failed - the cherry-picked rust-tests
   agent's commit (and the earlier hnsw.rs/rabitq.rs edits) were never
   run through rustfmt. `cargo fmt -p ruvector-py` fixes it; verified
   `cargo fmt --all -- --check` is now clean for every ruvector-py
   file, and build+clippy+test all still pass after reformatting.
2. The aarch64 wheel-build leg failed for real:
   maturin-action picked the x86_64 CROSS manylinux container
   (needing an `aarch64-linux-gnu-gcc` cross-linker) even though
   `runs-on: ubuntu-24.04-arm` is a native arm64 runner - confirmed
   from the actual job log ("linker `aarch64-linux-gnu-gcc` not
   found"), then confirmed against maturin-action's own source
   (`src/index.ts`) that `container: 'off'` is its documented escape
   hatch ("disable manylinux docker build and build on the host
   instead") and that `container: 'auto'` is a recognized literal
   equivalent to leaving it unset - used a matrix field for both
   rather than a ternary with an empty-string fallback, to avoid any
   risk to the x86_64 leg, which already passes. Tradeoff documented
   in the workflow: the aarch64 wheel loses the manylinux_2_28
   platform tag until someone can iterate against a real aarch64
   runner.

Two other failing checks on the PR are NOT mine: "Clippy (deny
warnings)" (workspace-wide) fails on crates/rvAgent/rvagent-core's
#[async_trait] usage tripping a new clippy 1.99 double_must_use lint
- verified this workflow passed clean on main as of 2026-09-30 (5
most recent main runs all green), so this is Rust-toolchain drift on
an unrelated crate this PR never touches, not something introduced
here or something in scope to fix.

26/26 cargo test, build+clippy clean, 93/93 pytest after the
reformat - no regressions from either fix.

Co-Authored-By: claude-flow <ruv@ruv.net>
…hardcoded set

Prep for the next batch of PyO3 classes (graph/GNN/cluster/SONA
bindings, forked in parallel): the old _NATIVE_NAMES hardcoded set is
exactly the bug class that made ruvector.HnswIndex unreachable earlier
this session (addable in Rust, forgotten here) - and with 3 parallel
forks each adding a class, it's also a guaranteed cherry-pick conflict
on the same set literal.

__getattr__ now probes `hasattr(_native, name)` / `hasattr(collection,
name)` dynamically instead. A new PyO3 class registered in lib.rs is
immediately reachable as ruvector.<Name> with nothing else to update
here - the bug class is now structurally impossible, not just
avoided by discipline. __all__ stays a static list (still needed for
`from ruvector import *` / IDE completion) and still needs one line
per new public class, but getting that wrong only costs
discoverability, never correctness.

Guards: leading-underscore names (except the one legitimate
__version__ re-export) and a reserved-submodule set (cli, mcp_server,
integrations, collection, _native) are rejected outright so dynamic
dispatch can't shadow a real submodule import or leak something
private.

93/93 pytest, mypy --strict clean.

Co-Authored-By: claude-flow <ruv@ruv.net>
…stubs

Per advisor review, before forking three parallel agents at the
remaining capability-table rows (graph CRUD, GNN+attention rerank,
k-means, SONA): add all five new path dependencies to Cargo.toml and
no-op mod/register() stub files to lib.rs myself first, so each fork
only fills in its own module body and never touches the two files
every fork would otherwise all need to edit (Cargo.toml, lib.rs) -
avoids a repeat of the HnswIndex merge-conflict/registration-gap
pattern from the last round of forks.

Each dependency's feature selection verified by a real standalone
trial compile (not assumed from Cargo.toml), in
/home/ruvultra/.cache/claude-code/.../scratchpad/trial_deps_check -
caught two things worth recording:
- ruvector-graph's default feature ("full") pulls in tokio/moka/
  zstd/lz4/redb for storage/async-runtime/compression this binding
  doesn't use - default-features=false + simd, matching ruvector-core.
- ruvector-sona (crate ruvector-sona, dir crates/sona) must NOT get
  default-features=false: its serde/serde_json deps are Cargo-
  "optional" but several modules use them with no cfg-gate at all,
  so disabling the default serde-support feature fails to compile -
  a real, reproduced build failure, not a guess.
ruvector-gnn and ruvector-attention were already clean (no heavy
non-optional deps); ruvector-cluster-rag has zero dependencies.

Wheel size with all five stub-registered: 2.27 MB (up from 358 KB),
still well under the 8 MiB CI budget - worth tracking once the real
bindings land, not a blocker now.

build+clippy+fmt all clean, 93/93 pytest still pass (no behavior
change - every register() is a no-op until the forks land real
pyclasses).

Co-Authored-By: claude-flow <ruv@ruv.net>
…mmit 1)

Add a GraphDB PyO3 class wrapping ruvector_graph::GraphDB: create/get
node, create/get edge, outgoing-edge traversal, __len__/__repr__. Not
unsendable -- GraphDB's own concurrent-update test proves Send+Sync, and
the one non-Send field (storage) is compiled out by this crate's feature
set. Properties round-trip through a hand-rolled PropertyValue<->Python
converter (serde's default enum tagging rules out a serde_json round-trip).
Guards two silent-overwrite footguns in the underlying Rust API: a
duplicate explicit id on create_node/create_edge now raises instead of
corrupting the label/property indexes.

18 new pytest cases (111 total passing), 5 Rust unit tests for the
converters, mypy --strict clean, clippy -D warnings clean, cargo fmt clean.

Co-Authored-By: claude-flow <ruv@ruv.net>
… labels

The _native.pyi stub promises Sequence[str] but py_to_labels only cast to
PyList, so a tuple of labels passed mypy and then raised at runtime.
Switch to Vec<String> extraction (works for any Python sequence/iterable
of str) and add an explicit dict rejection, since a dict is iterable too
and would otherwise silently become a list of its keys.

Co-Authored-By: claude-flow <ruv@ruv.net>
…-352 graph slice, commit 2)

Port the MATCH executor from crates/ruvector-graph-node/src/cypher_exec.rs
(a NAPI binding crate's module that turned out to have zero NAPI-specific
types -- only ruvector_graph::{cypher::ast, GraphDB} and std) into this
crate as three pyo3-free child modules (graph/cypher_eval.rs,
graph/cypher_exec.rs) plus one pyo3-aware glue module
(graph/cypher_bridge.rs). Logic is unchanged from the original; the only
rename is Bound -> Binding to avoid colliding with pyo3::Bound. All 14
ported Rust unit tests pass unmodified.

GraphDB.query_cypher(cypher) -> {"nodes": [...], "edges": [...]} executes
only MATCH; RETURN is parsed but never applied as a projection (documented
in both the Rust doc comment and the .pyi stub); CREATE raises explicitly
(mirrors ruvector-graph-node's own query() contract, which refuses writes
through a query string rather than silently discarding them); other write
statement types (MERGE/SET/DELETE/REMOVE) are currently accepted but not
executed -- same upstream contract, not invented here. Releases the GIL
around the graph walk, matching hnsw.rs/rabitq.rs's own pattern.

Also split the Python<->PropertyValue converters out of graph.rs into
graph/convert.rs, purely to keep every file under this codebase's
500-line convention once query_cypher landed.

Fixes a stale `# type: ignore` in test_graph.py left over from the
Sequence[str] labels fix (afe3d0acd): mypy now accepts a bare `str` for
`labels` structurally (str satisfies Sequence[str]), so the ignore was
unused.

13 new pytest cases (126 total), mypy --strict clean on both the package
and the test files, clippy -D warnings clean, cargo fmt clean.

Co-Authored-By: claude-flow <ruv@ruv.net>
Fills in gnn.rs's pre-wired stub with GnnLayer (wraps
ruvector_gnn::layer::RuvectorLayer) and AttentionReranker (wraps
ruvector_attention's ScaledDotProductAttention). GnnLayer is honestly
documented as a random-init-only forward pass (no training step here,
so it is not a quality improvement by itself); AttentionReranker is
trainless and deterministic, and returns both the blended output and
the raw per-candidate weights for RAG-style reranking. Neither type
needs `unsendable` — both are plain ndarray/usize/f32 data, Send+Sync
by auto-derive. Appends to_pyerr_gnn/to_pyerr_attention to error.rs and
appends GnnLayer/AttentionReranker to the __all__ lists and .pyi stubs.

Co-Authored-By: claude-flow <ruv@ruv.net>
The module doc overclaimed "no training step anywhere in
ruvector_gnn" — the crate does have training::Optimizer,
replay::ReplayBuffer, and ewc elsewhere. What's actually true: none
of it is wired to RuvectorLayer (no backward()) or reachable from
this binding. Narrow the claim accordingly.

Co-Authored-By: claude-flow <ruv@ruv.net>
Wrap ruvector_cluster_rag::cluster::kmeans as a free pyfunction
returning (assignments, centroids, cohesion, cluster_sizes) NumPy
arrays. Re-validates every invariant the Rust fn enforces via assert!
(k>0, k<=n, non-empty, finite coordinates) before calling, since an
unhandled Rust panic surfaces as pyo3_runtime.PanicException — a
BaseException that bypasses RuVectorError entirely.

Co-Authored-By: claude-flow <ruv@ruv.net>
Wrap ruvector_sona::SonaEngine's forward-pass API: new(hidden_dim),
apply_micro_lora, apply_base_lora, stats, save_state/load_state.
Deliberately does not bind the trajectory/training API (begin/end
trajectory, tick, force_learn, find_patterns) — left for a possible
follow-up commit per the task's own scope.

Compiles as a plain (non-unsendable) #[pyclass] under pyo3 0.29's
Send+Sync requirement — no unsafe opt-out needed. Both LoRA forward
passes are residual (output seeded with input, delta added on top)
and every projection is zero-initialised at construction, so a fresh
engine's apply_* calls are an exact identity transform, not merely
"close to one" — documented honestly rather than oversold.

load_state propagates malformed-JSON errors as RuVectorError instead
of mirroring the upstream NAPI binding's silent eprintln!+0 fallback.

Co-Authored-By: claude-flow <ruv@ruv.net>
…ding

Documents what shipped from three parallel forks (GraphDB raw CRUD +
Cypher MATCH, GnnLayer + AttentionReranker, kmeans + SonaEngine),
integrated via cherry-pick onto feat/python-sdk with zero Rust
conflicts and three small append-only __all__-list conflicts
(resolved trivially - the dynamic-dispatch refactor done before
forking already eliminated the correctness-risk half of this class
of conflict).

Re-verified post-integration, not just trusted from the forks'
reports: cargo build/clippy/fmt/test (49/49) all clean; maturin
develop + pytest (167/167) clean; mypy --strict clean (19 files);
a hands-on smoke test of every new class (GraphDB CRUD + Cypher,
GnnLayer.forward, AttentionReranker.rerank, kmeans, SonaEngine) run
directly against the rebuilt wheel; a fresh-wheel-install check
(140 passed + 2 cleanly-skipped-with-a-reason for the optional
langchain/llamaindex extras, 0 errors); cargo audit (no new findings
- confirmed the only `lru` version in ruvector-py's own dependency
tree, 0.18.2, is past the vulnerable range the two pre-existing
unrelated advisories flag) and claude-flow security scan (clean).

Added a capability-table row explicitly declining
crates/ruvector-cluster (distributed-sharding infra, not ML
clustering - a premise mismatch flagged since the original
capability inventory) and RVF persistence (not practical this
session - documented honestly rather than silently dropped).

Co-Authored-By: claude-flow <ruv@ruv.net>
…esign

Co-Authored-By: claude-flow <ruv@ruv.net>
ruvnet added 12 commits October 1, 2026 18:48
Closes the gap ADR-352 flagged earlier this session ("ruvector serve
--http currently has no auth"). Design decided before implementing,
not guessed mid-code: read the mcp SDK's actual source for
MCPServer.__init__, TokenVerifier, AuthSettings, and custom_route.

RUVECTOR_MCP_TOKEN env var (never in code), checked via the SDK's
real token_verifier/AuthSettings mechanism - not a hand-rolled ASGI
middleware. StaticTokenVerifier compares the bearer token with
hmac.compare_digest (constant-time). AuthSettings needs issuer_url/
resource_server_url (AnyHttpUrl, required fields) even though this
design does no real OAuth discovery - uses a self-referential
placeholder URL and sets validate_token_resource=True explicitly
(our AccessToken.resource always matches it, so this is a real
enforced check, not just silencing the SDK's own deprecation
warning).

Every mutating tool (vector_create_collection/vector_insert/
vector_insert_batch/vector_delete) additionally calls
_require_write_scope(), checked against the authenticated token's
scopes via get_access_token() - a no-op when no auth is configured
at all (stdio, or --http with no token set), so local/dev use is
unaffected either way. If --http starts with no token configured,
run_http prints one unmissable startup warning instead of silently
serving unauthenticated.

Real finding from reading the SDK source before implementing, not
assumed: MCPServer.custom_route's own docstring says those routes do
NOT get token_verifier protection ("intended for uses that are part
of authorization flows... or public") - documented prominently for
whoever builds the Salesforce Agentforce action endpoint next, since
mounting it via custom_route would otherwise silently ship
unauthenticated.

Verified live end-to-end, not just via the test suite: a real HTTP
server + real curl requests - no Authorization header -> 401, wrong
token -> 401, correct token -> 200 with the real initialize response.
Also verified the scope-gating logic directly against all three auth
context states (read-only token, read+write token, no auth context).

8 new tests in test_mcp_server.py (StaticTokenVerifier.verify_token,
_build_auth with/without the env var, _require_write_scope across all
three states). 173/173 pytest (was 167), mypy --strict clean
(19 files, needed one type: ignore[attr-defined] for mcp's
AuthenticatedUser, which isn't in that module's explicit __all__).

Also records a real, honest pip-audit finding in ADR-352: nltk
3.10.3 (PYSEC-2026-3740, file-sandbox bypass in nltk's own model-
persistence APIs) is a direct, unpatched dependency of
llama-index-core - not a ruvector bug, not fixable by pinning (no
patched nltk version exists yet), and unreachable through
ruvector.integrations.llamaindex's actual code paths, but a real
transitive exposure for `pip install ruvector[llamaindex]` worth
recording rather than discovering later.

Co-Authored-By: claude-flow <ruv@ruv.net>
Adds ruvector.integrations.salesforce (OAuth2 client-credentials,
paginated SOQL fetch, sync-to-collection, and three Agentforce actions:
search/upsert/ground) plus ruvector.salesforce_routes, which mounts
those actions as custom HTTP routes on the same MCPServer/port.

MCPServer.custom_route does not go through the SDK's token_verifier
(confirmed by reading the SDK source before writing this), so every
route does its own manual bearer check against
RUVECTOR_SALESFORCE_ACTION_TOKEN (falls back to RUVECTOR_MCP_TOKEN).
generate_openapi_spec() produces the flat OpenAPI 3.0 doc an Agentforce
External Service imports as custom actions.

28 new tests (20 integration, 8 route-level via Starlette TestClient
against the real ASGI app). All network calls in tests go through
httpx.MockTransport -- no real Salesforce org is contacted anywhere.
Found and fixed three real bugs via live end-to-end curl testing before
committing: unhandled CollectionError/KeyError leaking as raw 500s, a
mypy untyped-decorator placement, and (in the route tests themselves) a
module-reload staleness bug where salesforce_routes' `from .mcp_server
import server` binding went stale after the test fixture reloaded
mcp_server, registering routes onto an orphaned server object.

201 tests pass, mypy --strict clean.

Co-Authored-By: claude-flow <ruv@ruv.net>
release.yml excludes ruvector-py from cargo test (needs libpython,
which that bare runner doesn't provide), and its comment claimed the
49 Rust unit tests run instead in python-wheels.yml -- they didn't;
that workflow only builds wheels via maturin and pytest-tests them,
with no `cargo test` step anywhere. Found while writing up ADR-352's
coverage table and checking the claim against the actual workflow
file instead of trusting the comment.

Adds a "rust-tests" job to python-wheels.yml: actions/setup-python +
cargo test -p ruvector-py --release, same approach confirmed working
locally (no extension-module Cargo feature needed for this to link,
per that crate's Cargo.toml comment). Also tightens
salesforce_routes.py's _run_action to catch ruvector.RuVectorError
specifically instead of bare Exception, so a real bug in this module
(e.g. AttributeError) surfaces as Starlette's normal 500 instead of
being reported as a 404 with its internal message exposed -- the
module's own docstring already claimed this narrower behavior; now
the code matches it.

Co-Authored-By: claude-flow <ruv@ruv.net>
ADR-352 gets a new "Integrations" section covering LangChain/LlamaIndex
(previously undocumented there despite being implemented) and the full
Salesforce Agentforce research + build writeup: why Data Cloud's "BYO
retriever" and Agentforce MCP (Beta, AE-tier-gated) were ruled out as
primaries, why External Services + OpenAPI was chosen instead, what the
integration does, and an explicit verified/not-verified split against a
real org per the task's instruction to be explicit about that boundary.

examples/python-salesforce-agentforce/ adds a Named Credential and
External Service Registration metadata sketch plus a setup README --
every file headed "illustrative, not validated against a real org or
the current Metadata API schema" and deliberately excludes a
GenAiFunction sketch (a newer metadata type whose current schema this
session had no way to confirm) rather than fabricate one.

Co-Authored-By: claude-flow <ruv@ruv.net>
docs/python/README.md -- install + extras, quick start, full SDK
reference (Collection, graph/GNN/attention/kmeans/SONA), the CLI
command table, the MCP server (tools, auth, ChatGPT ui:// widget),
each integration (LangChain, LlamaIndex, Salesforce Agentforce), perf
tips, and the real-embedding benchmark table.

Every code example was corrected against and then smoke-tested live
against the real compiled extension before committing, not just
written from the Rust source read-through -- caught several real
signature mismatches first: GnnLayer has no `.random()` staticmethod
(it's a direct constructor with input_dim/hidden_dim/heads/dropout),
`GraphDB.create_edge`'s third positional arg is `relation_type` not
`edge_type`, `kmeans()` returns a plain 4-tuple (not an object with
.assignments/.centroids attributes), `SonaEngine`'s constructor param
is `hidden_dim` not `dim`, and the LlamaIndex `RuVectorStore` takes a
`Collection` instance, not a `dim` kwarg.

Co-Authored-By: claude-flow <ruv@ruv.net>
DO NOT MERGE until `ruvector` is live on PyPI -- this commit is a
placeholder documenting that the install snippet works, not a
confirmation that it does yet. Minimal addition matching the
surrounding Node.js section's style: install snippet + a link to the
full guide at docs/python/README.md, nothing restructured.

Co-Authored-By: claude-flow <ruv@ruv.net>
MCP HTTP auth, Salesforce Agentforce, the CI test-coverage gap, and
all four docs deliverables (ADR integrations section, user guide,
flagged README commit, this checkpoint) are now done. Updated the
capability table and replaced the stale "Next" list with what's
actually left: a full CI run on today's commits, the final Slack
summary/ClaimReleased block, the PR description flag, and the
already-deferred items (RVF persistence, embeddings, Turbo4 — each
with its stated reason, not silently dropped).

Co-Authored-By: claude-flow <ruv@ruv.net>
The user guide pointed at "the benchmark scripts referenced in
LOOP-STATE.md" -- which itself said they only lived in the session
scratchpad, never committed. A reader following the guide could not
reproduce the benchmark table at all. Moved both scripts (RabitqPlus
vs hnswlib on random-Gaussian data; HnswIndex vs hnswlib on real
MiniLM/20newsgroups embeddings) into crates/ruvector-py/benchmarks/,
confirmed they still compile and their deps (hnswlib,
sentence-transformers, scikit-learn) import cleanly from this
location, and fixed both docs to point at the real path.

Co-Authored-By: claude-flow <ruv@ruv.net>
This PR's own first real CI run on the aarch64 leg surfaced a second,
distinct bug beyond the earlier container-selection one: even with
`container: 'off'` (building natively on ubuntu-24.04-arm instead of
in a manylinux docker image), the step still passed `manylinux: 2_28`
to maturin-action -- which runs a post-build compliance check against
that floor regardless of whether a container was used to build. The
native runner's own glibc is newer than manylinux_2_28 allows (real
error: "not manylinux_2_28 compliant ... GLIBC_2.30, GLIBC_2.33,
GLIBC_2.34"), so the build now compiled fine and then maturin itself
failed post-build.

Fix: `manylinux: off` for the aarch64 leg specifically (matching
`container: off`), which skips that compliance check and accepts the
already-documented tradeoff (plain `linux_aarch64` tag instead of a
true manylinux one) rather than erroring out before ever reaching it.

Co-Authored-By: claude-flow <ruv@ruv.net>
…eholder

Per a design review from the coordinating teammate: did real WebSearch/
WebFetch research (not memory) to back the ADR's two Agentforce research
claims with dated, live-fetched citations -- SalesforceBen's Data Cloud
grounding writeup (2026-01-28) confirms Custom Retriever only offers
internal Data Cloud sources (Data Space/DMO/Search Index), and
Salesforce's own Agentforce MCP blog post (2026-01-15) confirms Beta
status gated on AE account qualification. Both independently corroborate
what the ADR already asserted from the prior session.

Added examples/python-salesforce-agentforce/genAiFunctions/
RuVector_Search.genAiFunction-meta.xml per the same review, with
invocationTargetType left as a literal, commented UNCONFIRMED_PLACEHOLDER
rather than a guessed value -- real research confirmed `flow` and `slack`
as valid values for this field but could not confirm what an
External-Service-backed action resolves to. Updated both the ADR and the
examples README to match.

Co-Authored-By: claude-flow <ruv@ruv.net>
"Salesforce's Setup UI most likely assigns the invocationTargetType
value automatically" had no source behind it -- a plausible-sounding
guess about Salesforce internals that neither WebFetch this checkpoint
actually confirmed. In a file and ADR section whose entire point is
marking what is and isn't verified, shipping an unsourced hypothesis
undermines that. Trimmed to what was actually confirmed: flow/slack
are valid values, the External-Service case is unconfirmed, no source
says how/by whom that value gets set.

Co-Authored-By: claude-flow <ruv@ruv.net>
Caught by the coordinating teammate before any tag/publish was
attempted: the previous fix (manylinux: off on the aarch64 leg, to
stop maturin's post-build compliance check from failing against the
native runner's newer glibc) made the CI leg green, but produces a
plain `linux_aarch64` wheel tag. PyPI's upload validation rejects any
wheel whose platform tag isn't manylinux_*/musllinux_*/macosx_*/win_*
-- so this leg's wheel would have failed at the real publish step,
not here.

Root-caused properly this time by reading maturin-action's actual
source at the pinned commit (function `getDockerContainer`): its
DEFAULT_CONTAINERS table is keyed by the RUNNER's own process.arch
first, then target, then manylinux version --
`DEFAULT_CONTAINERS.arm64['aarch64-unknown-linux-gnu']['2_28']` is
`quay.io/pypa/manylinux_2_28_aarch64:latest`, the correct NATIVE
image for this native-arm64 runner (no QEMU, no cross-linker). That
lookup only fires when `container` is left unset -- an explicit `off`
(this PR's two earlier attempts) bypasses it. Removed the per-leg
`container`/`manylinux` overrides entirely; both legs now just set
`manylinux: 2_28` and let the action's own per-host-arch table do the
right thing.

Added a new `validate-wheels` job (runs on every push/PR, unlike
`publish` which is tag/dispatch-gated) that asserts every built
wheel's filename carries an accepted platform tag and runs `twine
check` against all of them -- so a regression of this exact class
fails loudly in CI instead of surfacing at a real publish attempt.
`publish` now also depends on it.

Co-Authored-By: claude-flow <ruv@ruv.net>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant