Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
57 changes: 38 additions & 19 deletions docs/src/format/index/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,10 +100,11 @@ Index segments are created and updated through a transactional process:
- `name`: The index name (must match existing segments if adding to an existing index)
- `fields`: The columns the index depends on: the keyed column(s) it is searched on, followed
by any merely-carried columns named in `covering_fields`. `fields[0]` is always a keyed column.
- `covering_fields`: The trailing subset of `fields` whose values the index carries but is not
keyed on, letting a query that only projects those columns be answered without a fragment take.
Empty for an index that carries no extra columns. Declaring a column here does not by itself
make it servable -- see [Serving carried columns](#serving-carried-columns).
- `covering_fields`: The trailing subset of `fields` whose values the index carries, letting a
query that only projects those columns be answered without a fragment take. Usually these are
columns the index is not keyed on, but a keyed column may also be carried. Empty for an index
that carries no extra columns. Declaring a column here does not by itself make it servable --
see [Serving carried columns](#serving-carried-columns).
- `fragment_bitmap`: The set of fragment IDs covered by this segment
- `index_details`: Index-specific configuration and parameters
- `version`: The format version of this index type
Expand Down Expand Up @@ -141,18 +142,36 @@ fragments that would have been covered by that segment.
carries. It does not establish that the segment's storage holds their values.

**The segment's storage schema is authoritative.** Before answering a query from a
carried column, an engine must confirm that column is present in the storage it opened,
and fall back to a take against the base table when it is not. A segment whose
declaration names a column its storage does not hold is a legal state, not corruption:
a maintenance operation that cannot carry the payload through a rebuild is permitted to
withdraw it and leave the declaration standing.

!!! note "Current state"

No index builder writes carried values yet, so today every declaration is ahead of
its storage. Engines that read `covering_fields` must therefore treat it purely as a
declaration and serve every column from the base table until they have verified the
storage themselves. This is transitional; the rule above is not.
carried column, an engine must confirm that column is present and bound to the declared
logical field in the storage it opened, and fall back to a take against the base table
when it cannot. A segment whose declaration names a column its storage does not hold is
a legal state, not corruption: a maintenance operation that cannot carry the payload
through a rebuild is permitted to withdraw it and leave the declaration standing.

!!! note "Capability varies by segment"

Whether a segment's storage holds a declared column depends on the index type, on the
writer that produced the segment, and on what later maintenance did to it, so one
logical index may hold values for some of its segments and not others. An engine
therefore verifies each selected segment rather than inferring capability from the
index type, the writer version, or the declaration alone, and serves from the base
table every column it cannot verify.

!!! note "A keyed column may also be carried"

`covering_fields` usually names columns the index is *not* keyed on, but an index is
permitted to carry a column it is also keyed on -- for instance a vector index that
keeps full-precision vectors so a refine pass can re-rank without a base-table take.
The id then appears twice in `fields`, once as `fields[0]` and again as the trailing
carried entry, and once in `covering_fields`. This is the only case in which an id
repeats in `fields`.

A reader must therefore take the carried set from `covering_fields` directly, and
never derive it by subtracting the keyed prefix from `fields`: that set difference
silently drops a column that is both. The trailing-subset rule is stated over
`covering_fields` and is unaffected: `fields[0]` remains the column the index is
searched on, and an engine serves the repeated id from storage like any other
carried column.

## Loading an index

Expand All @@ -175,9 +194,9 @@ The `IndexMetadata` message contains important information about the index segme
- `fields`: the columns the index depends on: the keyed column(s) the index is searched on, followed
by any columns it merely carries, as named in `covering_fields`. `fields[0]` is always a keyed column.
- `covering_fields`: the trailing subset of `fields` whose values the index carries alongside its own
data but is not keyed on. Empty for an index that carries no extra columns. This declaration is
not authoritative for what the segment can serve -- see
[Serving carried columns](#serving-carried-columns).
data -- usually columns it is not keyed on, though a keyed column may also be carried. Empty for an
index that carries no extra columns. This declaration is not authoritative for what the segment can
serve -- see [Serving carried columns](#serving-carried-columns).
- `fragment_bitmap`: the set of fragment IDs covered by this index segment.
- `index_details`: a protobuf `Any` message that contains index-specific details, such as index type,
parameters, and storage format. This allows different index types to store their own metadata.
Expand Down
24 changes: 24 additions & 0 deletions docs/src/format/index/vector/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -161,6 +161,19 @@ the Arrow schema of the Lance file varies depending on the quantization method u
!!! note
All partitions are stored in the same file, and partitions must be written in order.

Every quantization format below lists only its internal columns. When a V3 IVF
writer materializes carried values, it appends one trailing column per carried
field after them, named and typed exactly as in the dataset schema. This physical
payload may be a subset of the manifest's `covering_fields` declaration (see
[Index Metadata](../index.md)). A reader returns only columns whose physical
schema and source field ids it verifies across every selected segment; all other
projected columns come from a base-table take.

A reader discovers carried columns by exclusion, not by position: any column in the
auxiliary file's schema that is not one of the quantizer's internal columns is a
carried column. Writers append them in trailing order, but a reader must not depend
on that ordering to identify them.

##### FLAT

No quantization applied - stores original vectors in their full precision:
Expand Down Expand Up @@ -229,6 +242,17 @@ Contains RabitQ-specific metadata in JSON format (only present for RQ quantizati
This includes the rotation matrix position, number of bits, and packing information.
See the RQ metadata specification in the "storage_metadata" section below.

##### "covering_field_ids"

The *source dataset* field ids of the storage file's physical carried columns,
comma separated in physical schema order (only present when the storage carries
values). Arrow fields carry no Lance field id, so names and types alone cannot
prove which logical column a payload came from. Readers use these ids to bind
physical values to the segment's `covering_fields` declaration, and treat missing,
malformed, ambiguous, or mismatched metadata as no servable carried capability.
Distributed merges use the same identity to reject shards whose columns match by
name and type but come from different fields.

##### "storage_metadata"

Contains quantizer-specific metadata as a list of JSON strings.
Expand Down
14 changes: 14 additions & 0 deletions docs/src/guide/performance.md
Original file line number Diff line number Diff line change
Expand Up @@ -529,3 +529,17 @@ Set `LANCE_DISABLE_AMX=1` to take the AMX paths out of service without rebuildin
A/B measurement, or to get the previous behaviour back. Because it also moves partition
assignment back to the approximate path, an index built with it set is not equivalent to one
built without it; compare recall, not just build time.

#### Covering Columns

A vector index can carry the values of extra columns beside its vectors, so a query whose
projection they satisfy is answered from the index without a take against the base table.
That trade only pays off where the search settles into a single global top-k heap, because
the covering read is then bounded by the query's survivors. Every HNSW index, and any query
with `query_parallelism` above one, emits results per partition instead, so covering reads
scale with the number of partitions probed rather than with `k`, and are issued serially.

On those shapes covering is *slower* than the base-table take it exists to avoid — roughly
2.8x a plain index's latency warm at nprobe 32, and 1.18x cold. Prefer IVF_PQ or IVF_FLAT
when declaring covering columns; index creation logs a warning when covering is combined
with HNSW.
11 changes: 11 additions & 0 deletions java/lance-jni/src/utils.rs
Original file line number Diff line number Diff line change
Expand Up @@ -261,6 +261,11 @@ pub fn build_compaction_options(
}

// Convert from Java Optional<Query> to Rust Option<Query>
//
// This builds a `Query` directly rather than through `Scanner::nearest`, and is used
// only by `JniTestHelper.parseQuery` to check the Java-side field marshalling. Java's
// real search path is `blocking_scanner.rs`, which configures a `Scanner`, so it picks
// up every plan-derived setting -- including the covering projection below.
pub fn get_query(env: &mut JNIEnv, query_obj: JObject) -> Result<Option<Query>> {
let query = env.get_optional(&query_obj, |env, java_obj| {
let column = env.get_string_from_method(&java_obj, "getColumn")?;
Expand Down Expand Up @@ -305,6 +310,11 @@ pub fn get_query(env: &mut JNIEnv, query_obj: JObject) -> Result<Option<Query>>
dist_q_c: 0.0,
query_parallelism,
approx_mode,
// Not a user-settable search parameter: the covering projection is derived
// per plan from what the scan reads, so it stays `None` (materialize
// whatever the index declares) at the binding boundary. Java's real search
// path resolves it in `Scanner`; see the note on this function.
covering_projection: None,
})
})?;

Expand Down Expand Up @@ -515,6 +525,7 @@ pub fn get_vector_index_params(
version: IndexFileVersion::V3,
skip_transpose: false,
runtime_hints: Default::default(),
covering_columns: Default::default(),
})
},
)?;
Expand Down
9 changes: 8 additions & 1 deletion java/src/main/java/org/lance/index/IndexDescription.java
Original file line number Diff line number Diff line change
Expand Up @@ -69,7 +69,14 @@ public String getName() {
return name;
}

/** Field ids that this index is built on. */
/**
* Field ids that this index is built on -- the columns it can answer queries for.
*
* <p>This is the index's <em>keyed</em> prefix only. An index may additionally carry values for
* columns it is not keyed on; those are deliberately absent here, because the index cannot be
* searched on them. They stay reachable per segment via {@link Index#coveringFields()} on the
* entries of {@link #getMetadata()}.
*/
public List<Integer> getFieldIds() {
return fieldIds;
}
Expand Down
23 changes: 23 additions & 0 deletions protos/ann.proto
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,23 @@ enum VectorApproxMode {
Accurate = 2;
}

// The covering ("included") index columns a query needs materialized, by name.
//
// Exists as a message rather than a bare `repeated string` because proto3 has no
// presence tracking for repeated fields, and the empty list is a distinct, meaningful
// state here:
//
// * absent: no narrowing computed; materialize every covering column declared.
// * present and empty: materialize nothing, though the index does declare covering.
// * present and non-empty: materialize exactly these.
//
// Collapsing "present, empty" into "absent" costs no correctness (a covering column is
// semantically transparent -- its values otherwise arrive from a base-table read) but
// silently restores full materialization on every distributed plan.
message CoveringProjection {
repeated string columns = 1;
}

// Serialized vector query parameters.
message VectorQueryProto {
// Query vector as Arrow IPC bytes (supports Float16, Float32, Float64, UInt8, etc.)
Expand All @@ -43,6 +60,12 @@ message VectorQueryProto {
// Query-time approximation mode. Currently only affects RQ-quantized vector
// indexes, such as IVF_RQ. Other index types ignore this setting.
VectorApproxMode approx_mode = 14;
// Which covering columns the index must materialize for this query. Absent means
// "not computed" -- see CoveringProjection. Carried across the wire so a remote
// executor declares the same search output schema the planner did; without it the
// executor's node is wider than the plan it came from, and the surrounding nodes
// were built against the planner's narrower schema.
CoveringProjection covering_projection = 15;
}

// Serializable form of ANNIvfSubIndexExec — the IVF sub-index search node.
Expand Down
4 changes: 4 additions & 0 deletions python/src/dataset.rs
Original file line number Diff line number Diff line change
Expand Up @@ -5674,6 +5674,10 @@ impl PySearchFilter {
query_parallelism,
dist_q_c: 0.0,
approx_mode,
// Not a user-settable search parameter: the covering projection is derived
// from what the plan reads, so it stays `None` (materialize whatever the
// index declares) at the binding boundary.
covering_projection: None,
};

Ok(Self {
Expand Down
6 changes: 4 additions & 2 deletions python/src/indices.rs
Original file line number Diff line number Diff line change
Expand Up @@ -674,9 +674,11 @@ pub struct PyIndexDescription {
pub type_url: String,
/// The short type of the index (may not be unique)
pub index_type: String,
/// The ids of the fields that the index is built on
/// The ids of the fields the index is keyed on -- the columns it can answer queries about.
/// Covering ("included") columns are excluded; those are reported per-segment as
/// `covering_fields`.
pub fields: Vec<u32>,
/// The full paths of the fields that the index is built on
/// The full paths of the fields the index is keyed on, matching `fields`
/// (dotted, with backtick-quoted segments for non-identifier names)
pub field_names: Vec<String>,
/// The number of rows indexed by the index
Expand Down
8 changes: 7 additions & 1 deletion rust/lance-index/src/traits.rs
Original file line number Diff line number Diff line change
Expand Up @@ -344,7 +344,13 @@ pub trait IndexDescription: Send + Sync {
/// deleted.
fn rows_indexed(&self) -> u64;

/// Returns the ids of the fields that the index is built on.
/// Returns the ids of the fields that the index is built on -- the columns it can
/// answer queries for.
///
/// This is the index's *keyed* prefix only. An index may additionally carry values
/// for columns it is not keyed on (see [`IndexMetadata::covering_fields`]); those are
/// deliberately absent here, because the index cannot be searched on them. They stay
/// reachable per segment via [`Self::metadata`].
fn field_ids(&self) -> &[u32];

/// Returns a JSON string representation of the index details
Expand Down
57 changes: 57 additions & 0 deletions rust/lance-index/src/vector.rs
Original file line number Diff line number Diff line change
Expand Up @@ -166,6 +166,52 @@ pub struct Query {
/// This currently only affects RQ-quantized vector indexes, such as IVF_RQ.
/// Other index types ignore this setting.
pub approx_mode: ApproxMode,

/// The covering ("included") columns this query needs the index to materialize,
/// by name.
///
/// Unlike the other fields here this is not a user-facing search parameter. It is
/// derived per plan from what the query actually reads, so that an index covering a
/// wide payload does not pay to materialize that payload on searches which never
/// look at it.
///
/// # The three states are distinct, and must stay distinct
///
/// * `None` — no query-level narrowing or physical-capability resolution was computed.
/// A raw index search may consider every physical covering column, but an execution
/// planner must not treat this state as proof that declared values are servable. It
/// resolves storage first and converts the query to an explicit `Some` projection.
/// * `Some(&[])` — the index *does* declare covering columns, but this query needs
/// **none** of them. Consumers must do no covering work at all: not "project zero
/// columns out of a batch that was loaded anyway", but skip the covering read
/// entirely.
/// * `Some(cols)` — this query needs exactly `cols`, and the execution planner has
/// verified that every selected segment can physically serve them.
///
/// `Some(&[])` is the state this field exists for, and the one that silently
/// degrades if it is folded into `None`: a covering column is semantically
/// transparent, so conflating the two restores full materialization while every
/// result stays byte-identical (the values simply arrive from the base table
/// instead). No result-correctness test can catch that; only a cost or plan-shape
/// assertion can. Treat `Option::unwrap_or_default()` on this field, or any
/// `is_empty()` test that does not first distinguish `None`, as a bug.
///
/// # Physical capability is a ceiling
///
/// A declared column left out of `cols` is simply fetched from the base table —
/// correct, just slower. A column listed here that any selected segment cannot emit
/// leaves the covered projection short a column. Planners may therefore remove an
/// unproven field, but must never add one based on the manifest declaration alone.
///
/// # Notes
///
/// Names rather than field ids, because consumers match against a storage batch whose
/// columns carry names. The planner separately verifies source field ids, Arrow shape,
/// and compatible physical order before producing this name projection.
///
/// `Arc` rather than `Vec` because [`Query`] is cloned once per probed partition on
/// the search hot path.
pub covering_projection: Option<Arc<[String]>>,
}

impl From<pb::VectorMetricType> for DistanceType {
Expand Down Expand Up @@ -236,6 +282,17 @@ pub trait VectorIndex: Send + Sync + std::fmt::Debug + Index {
/// Get the total number of partitions in the index.
fn total_partitions(&self) -> usize;

/// Covering columns physically present and safe to serve from this index, paired with
/// their source dataset field ids and returned in storage order.
///
/// [`IndexMetadata::covering_fields`](lance_table::format::IndexMetadata::covering_fields)
/// is only the logical declaration. Query planners intersect that declaration with this
/// physical capability for every segment before omitting a base-table read. Formats that
/// do not expose a verifiable covering payload use the default empty capability.
fn physical_covering_fields(&self) -> Result<Vec<(i32, Field)>> {
Ok(Vec::new())
}

/// Search a single partition for nearest neighbors.
///
/// This method should return the same results as [`VectorIndex::search`] method except
Expand Down
Loading
Loading