Replies: 6 comments 2 replies
|
In discussing the different layers, we found it was really helpful to provide engine-agnostic implementations of operations. Good example today: |
|
To me the naming of lance/lancedb suggests that lance is the data file format and lancedb is the database implementation built on top. As noted above however there is leakage. Lance includes engine functionality that could be separated from the format, and lancedb does not include a good interface for third-party engine integration. You cannot easily integrate lancedb into a custom engine at the datafusion layer (low-level) or the SQL layer (high level), so you are forced to use lance. This leads to confusion because we end up supporting two fairly-similar looking ways to solve a given problem. Datasets and tables can be substituted in many cases, and for users who land first on datasets and later need to transition to tables/lancedb for some reason, the transition is an unnecessary inconvenience that results from this unclear separation of concerns. If with time we tend toward putting more engine functionality in lancedb instead of lance and begin to scope lance more to format essentials and immediate consequences, then direct lance consumers would begin to lose out because they would need to re-implement that functionality in their own engines. For those who are integrating with external engines, would doing so via a SQL interface in lancedb be easier? Alternatively, would it be disqualifying because you need a lower level of control? |
|
Really glad to see this discussion — at the pace things are moving, it's the right time to revisit the architecture. Let me ground this in a fairly typical model-training pipeline we run, since you asked for real use cases rather than architecture diagrams. It spans both ends of the gradient you described: distributed What we use the distributed engines for
What we fall back to
|
|
@yanghua @timsaucer @eddyxu As PMC members, would welcome your input here. Also @LuciferYang for your perspective as a Spark PMC member and someone who's worked on the MergeInsert implementation. |
|
@wjones127 Thanks for pinging me, and for framing it as core-vs-high-level. That's exactly the tension we've been living with. The team I work on runs a distributed multimodal data-preparation pipeline (image/video plus large binary payloads, tables with hundreds of millions of rows across many fragments) on Ray + Lance, built directly on Lance's primitives. I've also worked on the MergeInsert implementation and come from the Spark side, which shapes how I read the boundary. What the framework layer uses. Our framework layer (the Ray integration) rides entirely on the engine-agnostic, task-object primitives: workers call Where we do reach higher. At the application layer we use the higher-level surface where it fits: What internals we reach into. On the older Lance version we run in production we reached into a handful of internals; checking against current Lance, several are already closed: the dataset is picklable, and version enumeration and storage-option accessors are public. The one still open for us is mapping a row we've read back to its owning fragment. There's no public resolver for that, so we depend on the row-address layout, which also stops holding once row ids are stable rather than address-based. A public row-to-fragment resolver, correct under both schemes, would remove the coupling. It's partly a consequence of how we distribute the work, so it may be more our concern than a general one. Where I'd draw the line. The Spark DataSource V2 analogy is close: DSv2 separates the physical On merge_insert specifically. Since that's why you tagged me: I'd group it with the Scanner decoupling you already flagged rather than treat it as separate. The commit at the end is already the engine-agnostic task object (a |
|
I don’t think the boundary can be described simply as “format in Lance, query execution elsewhere.” In Lance-Spark’s normal local scan path, Spark owns query-level planning, partition planning, and distributed scheduling, while each Spark input partition invokes Lance’s native Scanner for a specific fragment. That Scanner builds and executes a DataFusion physical plan locally. The local plan is not merely a generic filter-and-project pipeline. It enforces Lance-specific behavior around index coverage, fallback for unindexed fragments, overlay-stale index rows, prefilter eligibility, refinement, deletion visibility, and late materialization. The code therefore suggests that asking each engine to reproduce this behavior would create duplicated correctness-sensitive implementations. What appears to be missing is a stable engine-neutral boundary between connector-level planning and the DataFusion-specific physical plan. Today Lance-Spark maintains its own representation of resolved snapshots, fragment splits, zonemap pruning, statistics, and pushed versus residual predicates, while the JNI boundary reconstructs a Rust Scanner from a large option set. A serializable Lance-owned planning/task protocol could make that boundary explicit. DataFusion could continue as the native reference executor without being exposed in the protocol’s public types. I would also distinguish scheduler-agnostic from dependency-neutral. Suppose we want to change the situation. My preference would therefore be to retain a Lance-owned native runtime, define an engine-neutral planning/task/commit boundary around it, and let host engines own global relational planning and distributed scheduling. |
Uh oh!
There was an error while loading. Please reload this page.
Lance today spans a wide range. At one end is functionality that's clearly core to a file and table format: encodings (lance-encoding), the file reader/writer (lance-file), the manifest and fragment model (lance-table), object store I/O (lance-io), and index structures.
At the other end is functionality that looks a lot like a query engine:
Scanner— filter and projection pushdown, limits, vector/FTS/hybrid search planning (dataset/scanner.rs, ~16k lines)I don't want to argue for a particular end state in this thread. What I do want to put on the table is that scope has a maintenance cost, and I think we're paying it:
scanner.rsandmerge_insert.rs, ~16k and ~15k lines with tests). That's also where a lot of the churn and breakage concentrates. I can't help but feel these need a major refactor anyways.lance-encoding,lance-file,lance-io,lance-select,lance-arrow— have no DataFusion dependency at all. The gradient already half exists.There's a "what is Lance for" question underneath this too. There are already a lot of good query engines: DuckDB, DataFusion, Spark, Trino, LanceDB. I'm skeptical that Lance should try to be another one, when the thing only Lance can provide is the format and the primitives to read it well. delta-kernel-rs is one example of the alternative scope. It contains engine-agnostic building blocks, leaving planning and execution to the engines.
But I know the high-level APIs within Lance, including pylance, get a lot of use. "Narrow the scope" might mean different things depending on which APIs are your main entrypoints to Lance. Whatever we do has to start from real use cases rather than from an architecture diagram. So I'd like to hear:
Scanner, SQL, merge insert, update/delete? What would it look like to use a downstream engine (DuckDB, Spark, etc.) for that instead?Nothing is being removed and no decision has been made. This is the start of the conversation.
All reactions