Skip to content

# [C API] Expose stream-based svs_index_save / svs_index_load_dynamic variants #401

Description

@eshenayo

Summary

At pin 5717f6855a73a9ab7524748e2a10a57ed865737d, the SVS C API
exposes only path-based persistence:

  • bool svs_index_save(svs_index_h index, const char* directory, svs_error_h out_err);
    (bindings/c/include/svs/c/svs_c.h:1079)
  • bool svs_index_load_dynamic(svs_index_builder_t builder, const char* directory, size_t blocksize_bytes, svs_error_h out_err);
    (bindings/c/include/svs/c/svs_c.h:992)

Both variants require a filesystem directory. Consumers who persist
indexes to a non-filesystem byte stream (e.g., Valkey's RDB chunked
serialization) have to bridge the two with a temp-directory adapter:
save to a mkstemp directory, walk it in canonical order, stream
the file bytes through their transport, then unlink + rmdir;
load drains chunks back to a temp directory, hands it to
svs_index_load_dynamic, then cleans up.

This ask adds stream-based variants that call user-supplied
write/read callbacks, eliminating the temp-file round-trip for
consumers whose transport is already a byte stream.

Motivation

valkey-search integrates SVS as ALGORITHM SVS_VAMANA and persists
indexes to Valkey's RDB via RDBChunkOutputStream / RDBChunkInputStream
(valkey-io/valkey-search:src/rdb_serialization.h:289-367). Its
existing HNSW path serializes directly to the stream via
std::ostream&, so persistence has zero filesystem side effects.

Without a stream API, the SVS path must go filesystem -> stream on
save and stream -> filesystem on load. Concrete cost per RDB
save or load:

  1. Disk round-trip on every BGSAVE. For a 10 GiB index, that is
    10 GiB of temp-file writes followed by 10 GiB of reads to stream
    into RDB, on top of the RDB I/O the stream is already doing.
  2. Startup RDB load latency. Same doubling in reverse: RDB
    chunks drained to a temp directory, then re-read by
    svs_index_load_dynamic.
  3. Disk-space headroom requirement. Save now needs free space
    equal to the index size in $TMPDIR (or /tmp) even when the
    RDB destination has ample room. Deployments with small /tmp
    partitions can fail to save an index that otherwise fits.
  4. Cleanup complexity. Every error path must unlink the temp
    directory. Any missed path leaks disk.

The equivalent HNSW path has none of these because hnswlib's
saveIndex(std::ostream&) / loadIndex(std::istream&) accept a
stream directly.

Proposed change

Add stream-callback variants alongside the existing path-based
functions. Version-gate via SVS_C_API_VERSION so callers can
compile against either.

typedef bool (*svs_write_fn)(void* userdata, const void* data, size_t len);
typedef bool (*svs_read_fn)(void* userdata, void* buf, size_t len);

SVS_API bool svs_index_save_stream(
    svs_index_h index,
    svs_write_fn write_fn,     /* required */
    void* userdata,            /* opaque, passed to write_fn */
    svs_error_h out_err);

SVS_API bool svs_index_load_dynamic_stream(
    svs_index_builder_t builder,
    svs_read_fn read_fn,       /* required */
    void* userdata,            /* opaque, passed to read_fn */
    size_t blocksize_bytes,
    svs_error_h out_err);

Semantics:

  • write_fn returns true on success, false on error. SVS must
    abort the save on the first false return and propagate a
    distinguishable error code (e.g., SVS_ERROR_IO) through
    out_err.
  • read_fn fills exactly len bytes into buf and returns true,
    or returns false (partial fills are not allowed; short reads
    are errors). SVS must abort load on false and propagate
    SVS_ERROR_IO.
  • The byte stream produced by svs_index_save_stream is the
    concatenation of the files svs_index_save writes to a directory,
    in a canonical, documented order, each prefixed by a
    length-tagged header so the reader knows file boundaries. The
    stream and path forms must be interoperable: a save_stream
    output can be split back into files that load_dynamic accepts,
    and a save directory can be concatenated into a byte stream
    that load_dynamic_stream accepts. (Implementers can choose to
    implement the path variants as thin wrappers over the stream
    variants; that's the ideal but not required.)
  • Callbacks may be invoked with any len from 1 byte upward; SVS
    must not assume a minimum chunk size.
  • Callbacks may be invoked from any SVS thread. For consumers
    using SVS_THREADPOOL_KIND_SINGLE_THREAD or a
    SVS_THREADPOOL_KIND_CUSTOM with size()=1, this is the
    calling thread; for NATIVE/OMP, callbacks must be
    reentrant-safe -- document the guarantee.

Acceptance criteria

  1. svs_index_save_stream invokes write_fn repeatedly and returns
    true on success. write_fn sees the complete serialized index
    (byte-equivalent, modulo canonical-order framing, to what
    svs_index_save writes to a directory).
  2. svs_index_load_dynamic_stream invokes read_fn to drain the
    stream and constructs an index whose behavior is identical to
    one loaded via svs_index_load_dynamic from the equivalent
    directory.
  3. Round-trip interop: stream save -> stream load produces an
    index whose search results are byte-identical to the pre-save
    index for a deterministic seed. Path save -> stream load and
    stream save -> path load both work (via a documented
    directory-to-stream concatenation format).
  4. Chunk-boundary agnosticism: a unit test exercises save with
    a write_fn that fragments to 1-byte writes and a load with a
    read_fn that fragments to 1-byte reads; both succeed.
  5. Error propagation: a write_fn that returns false on the
    Nth call causes svs_index_save_stream to return false with
    svs_error_get_code(err) == SVS_ERROR_IO. Same for read_fn on
    load.
  6. Zero filesystem side effects: strace on a test binary that
    only exercises the stream variants shows no
    open/write/mkdir/unlink in the SVS code path. (Any
    filesystem access must originate from the caller's write_fn /
    read_fn, not from SVS.)
  7. Existing path-based svs_index_save / svs_index_load_dynamic
    behavior is unchanged.
  8. Unit tests in bindings/c/tests/ cover: stream save + stream
    load round-trip, path save + stream load, stream save + path
    load, chunk-boundary fragmentation, write-fn failure,
    read-fn failure/short-read.

Non-goals

  • Removing or deprecating the path-based svs_index_save /
    svs_index_load_dynamic. Consumers that persist to a filesystem
    keep the simpler API.
  • Async / non-blocking callbacks. Synchronous is sufficient; the
    consumer's stream is already synchronous.
  • Seekable-stream semantics. write_fn is append-only; read_fn
    is sequential. If SVS's on-disk format currently requires seeks,
    the stream variant should either buffer or restructure the
    format so it is streamable.
  • Stream-based variants of svs_index_build_dynamic /
    svs_index_dynamic_add_points. Only save/load is on the RDB
    hot path; build/add are not.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions