Summary
At pin 5717f6855a73a9ab7524748e2a10a57ed865737d, the SVS C API
exposes only path-based persistence:
bool svs_index_save(svs_index_h index, const char* directory, svs_error_h out_err);
(bindings/c/include/svs/c/svs_c.h:1079)
bool svs_index_load_dynamic(svs_index_builder_t builder, const char* directory, size_t blocksize_bytes, svs_error_h out_err);
(bindings/c/include/svs/c/svs_c.h:992)
Both variants require a filesystem directory. Consumers who persist
indexes to a non-filesystem byte stream (e.g., Valkey's RDB chunked
serialization) have to bridge the two with a temp-directory adapter:
save to a mkstemp directory, walk it in canonical order, stream
the file bytes through their transport, then unlink + rmdir;
load drains chunks back to a temp directory, hands it to
svs_index_load_dynamic, then cleans up.
This ask adds stream-based variants that call user-supplied
write/read callbacks, eliminating the temp-file round-trip for
consumers whose transport is already a byte stream.
Motivation
valkey-search integrates SVS as ALGORITHM SVS_VAMANA and persists
indexes to Valkey's RDB via RDBChunkOutputStream / RDBChunkInputStream
(valkey-io/valkey-search:src/rdb_serialization.h:289-367). Its
existing HNSW path serializes directly to the stream via
std::ostream&, so persistence has zero filesystem side effects.
Without a stream API, the SVS path must go filesystem -> stream on
save and stream -> filesystem on load. Concrete cost per RDB
save or load:
- Disk round-trip on every BGSAVE. For a 10 GiB index, that is
10 GiB of temp-file writes followed by 10 GiB of reads to stream
into RDB, on top of the RDB I/O the stream is already doing.
- Startup RDB load latency. Same doubling in reverse: RDB
chunks drained to a temp directory, then re-read by
svs_index_load_dynamic.
- Disk-space headroom requirement. Save now needs free space
equal to the index size in $TMPDIR (or /tmp) even when the
RDB destination has ample room. Deployments with small /tmp
partitions can fail to save an index that otherwise fits.
- Cleanup complexity. Every error path must unlink the temp
directory. Any missed path leaks disk.
The equivalent HNSW path has none of these because hnswlib's
saveIndex(std::ostream&) / loadIndex(std::istream&) accept a
stream directly.
Proposed change
Add stream-callback variants alongside the existing path-based
functions. Version-gate via SVS_C_API_VERSION so callers can
compile against either.
typedef bool (*svs_write_fn)(void* userdata, const void* data, size_t len);
typedef bool (*svs_read_fn)(void* userdata, void* buf, size_t len);
SVS_API bool svs_index_save_stream(
svs_index_h index,
svs_write_fn write_fn, /* required */
void* userdata, /* opaque, passed to write_fn */
svs_error_h out_err);
SVS_API bool svs_index_load_dynamic_stream(
svs_index_builder_t builder,
svs_read_fn read_fn, /* required */
void* userdata, /* opaque, passed to read_fn */
size_t blocksize_bytes,
svs_error_h out_err);
Semantics:
write_fn returns true on success, false on error. SVS must
abort the save on the first false return and propagate a
distinguishable error code (e.g., SVS_ERROR_IO) through
out_err.
read_fn fills exactly len bytes into buf and returns true,
or returns false (partial fills are not allowed; short reads
are errors). SVS must abort load on false and propagate
SVS_ERROR_IO.
- The byte stream produced by
svs_index_save_stream is the
concatenation of the files svs_index_save writes to a directory,
in a canonical, documented order, each prefixed by a
length-tagged header so the reader knows file boundaries. The
stream and path forms must be interoperable: a save_stream
output can be split back into files that load_dynamic accepts,
and a save directory can be concatenated into a byte stream
that load_dynamic_stream accepts. (Implementers can choose to
implement the path variants as thin wrappers over the stream
variants; that's the ideal but not required.)
- Callbacks may be invoked with any
len from 1 byte upward; SVS
must not assume a minimum chunk size.
- Callbacks may be invoked from any SVS thread. For consumers
using SVS_THREADPOOL_KIND_SINGLE_THREAD or a
SVS_THREADPOOL_KIND_CUSTOM with size()=1, this is the
calling thread; for NATIVE/OMP, callbacks must be
reentrant-safe -- document the guarantee.
Acceptance criteria
svs_index_save_stream invokes write_fn repeatedly and returns
true on success. write_fn sees the complete serialized index
(byte-equivalent, modulo canonical-order framing, to what
svs_index_save writes to a directory).
svs_index_load_dynamic_stream invokes read_fn to drain the
stream and constructs an index whose behavior is identical to
one loaded via svs_index_load_dynamic from the equivalent
directory.
- Round-trip interop: stream save -> stream load produces an
index whose search results are byte-identical to the pre-save
index for a deterministic seed. Path save -> stream load and
stream save -> path load both work (via a documented
directory-to-stream concatenation format).
- Chunk-boundary agnosticism: a unit test exercises save with
a write_fn that fragments to 1-byte writes and a load with a
read_fn that fragments to 1-byte reads; both succeed.
- Error propagation: a
write_fn that returns false on the
Nth call causes svs_index_save_stream to return false with
svs_error_get_code(err) == SVS_ERROR_IO. Same for read_fn on
load.
- Zero filesystem side effects:
strace on a test binary that
only exercises the stream variants shows no
open/write/mkdir/unlink in the SVS code path. (Any
filesystem access must originate from the caller's write_fn /
read_fn, not from SVS.)
- Existing path-based
svs_index_save / svs_index_load_dynamic
behavior is unchanged.
- Unit tests in
bindings/c/tests/ cover: stream save + stream
load round-trip, path save + stream load, stream save + path
load, chunk-boundary fragmentation, write-fn failure,
read-fn failure/short-read.
Non-goals
- Removing or deprecating the path-based
svs_index_save /
svs_index_load_dynamic. Consumers that persist to a filesystem
keep the simpler API.
- Async / non-blocking callbacks. Synchronous is sufficient; the
consumer's stream is already synchronous.
- Seekable-stream semantics.
write_fn is append-only; read_fn
is sequential. If SVS's on-disk format currently requires seeks,
the stream variant should either buffer or restructure the
format so it is streamable.
- Stream-based variants of
svs_index_build_dynamic /
svs_index_dynamic_add_points. Only save/load is on the RDB
hot path; build/add are not.
Summary
At pin
5717f6855a73a9ab7524748e2a10a57ed865737d, the SVS C APIexposes only path-based persistence:
bool svs_index_save(svs_index_h index, const char* directory, svs_error_h out_err);(
bindings/c/include/svs/c/svs_c.h:1079)bool svs_index_load_dynamic(svs_index_builder_t builder, const char* directory, size_t blocksize_bytes, svs_error_h out_err);(
bindings/c/include/svs/c/svs_c.h:992)Both variants require a filesystem directory. Consumers who persist
indexes to a non-filesystem byte stream (e.g., Valkey's RDB chunked
serialization) have to bridge the two with a temp-directory adapter:
save to a
mkstempdirectory, walk it in canonical order, streamthe file bytes through their transport, then
unlink+rmdir;load drains chunks back to a temp directory, hands it to
svs_index_load_dynamic, then cleans up.This ask adds stream-based variants that call user-supplied
write/read callbacks, eliminating the temp-file round-trip for
consumers whose transport is already a byte stream.
Motivation
valkey-search integrates SVS as
ALGORITHM SVS_VAMANAand persistsindexes to Valkey's RDB via
RDBChunkOutputStream/RDBChunkInputStream(
valkey-io/valkey-search:src/rdb_serialization.h:289-367). Itsexisting HNSW path serializes directly to the stream via
std::ostream&, so persistence has zero filesystem side effects.Without a stream API, the SVS path must go filesystem -> stream on
save and stream -> filesystem on load. Concrete cost per RDB
save or load:
10 GiB of temp-file writes followed by 10 GiB of reads to stream
into RDB, on top of the RDB I/O the stream is already doing.
chunks drained to a temp directory, then re-read by
svs_index_load_dynamic.equal to the index size in
$TMPDIR(or/tmp) even when theRDB destination has ample room. Deployments with small
/tmppartitions can fail to save an index that otherwise fits.
directory. Any missed path leaks disk.
The equivalent HNSW path has none of these because hnswlib's
saveIndex(std::ostream&)/loadIndex(std::istream&)accept astream directly.
Proposed change
Add stream-callback variants alongside the existing path-based
functions. Version-gate via
SVS_C_API_VERSIONso callers cancompile against either.
Semantics:
write_fnreturnstrueon success,falseon error. SVS mustabort the save on the first
falsereturn and propagate adistinguishable error code (e.g.,
SVS_ERROR_IO) throughout_err.read_fnfills exactlylenbytes intobufand returnstrue,or returns
false(partial fills are not allowed; short readsare errors). SVS must abort load on
falseand propagateSVS_ERROR_IO.svs_index_save_streamis theconcatenation of the files
svs_index_savewrites to a directory,in a canonical, documented order, each prefixed by a
length-tagged header so the reader knows file boundaries. The
stream and path forms must be interoperable: a
save_streamoutput can be split back into files that
load_dynamicaccepts,and a
savedirectory can be concatenated into a byte streamthat
load_dynamic_streamaccepts. (Implementers can choose toimplement the path variants as thin wrappers over the stream
variants; that's the ideal but not required.)
lenfrom 1 byte upward; SVSmust not assume a minimum chunk size.
using
SVS_THREADPOOL_KIND_SINGLE_THREADor aSVS_THREADPOOL_KIND_CUSTOMwithsize()=1, this is thecalling thread; for
NATIVE/OMP, callbacks must bereentrant-safe -- document the guarantee.
Acceptance criteria
svs_index_save_streaminvokeswrite_fnrepeatedly and returnstrueon success.write_fnsees the complete serialized index(byte-equivalent, modulo canonical-order framing, to what
svs_index_savewrites to a directory).svs_index_load_dynamic_streaminvokesread_fnto drain thestream and constructs an index whose behavior is identical to
one loaded via
svs_index_load_dynamicfrom the equivalentdirectory.
index whose search results are byte-identical to the pre-save
index for a deterministic seed. Path save -> stream load and
stream save -> path load both work (via a documented
directory-to-stream concatenation format).
a
write_fnthat fragments to 1-byte writes and a load with aread_fnthat fragments to 1-byte reads; both succeed.write_fnthat returnsfalseon theNth call causes
svs_index_save_streamto returnfalsewithsvs_error_get_code(err) == SVS_ERROR_IO. Same forread_fnonload.
straceon a test binary thatonly exercises the stream variants shows no
open/write/mkdir/unlinkin the SVS code path. (Anyfilesystem access must originate from the caller's
write_fn/read_fn, not from SVS.)svs_index_save/svs_index_load_dynamicbehavior is unchanged.
bindings/c/tests/cover: stream save + streamload round-trip, path save + stream load, stream save + path
load, chunk-boundary fragmentation, write-fn failure,
read-fn failure/short-read.
Non-goals
svs_index_save/svs_index_load_dynamic. Consumers that persist to a filesystemkeep the simpler API.
consumer's stream is already synchronous.
write_fnis append-only;read_fnis sequential. If SVS's on-disk format currently requires seeks,
the stream variant should either buffer or restructure the
format so it is streamable.
svs_index_build_dynamic/svs_index_dynamic_add_points. Only save/load is on the RDBhot path; build/add are not.