You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Chunked point-cloud payloads (COPC) as a first-class Lance workload
#9472
I have been evaluating Lance as the storage and governance layer for LiDAR
(point cloud) data. The conclusion is narrower than where I started, and I think
it is more useful to the project:
Lance already works for this, with no engine changes — and the parts that do
not work are generic, not LiDAR-specific.
Concretely: a real, unmodified COPC file stored as a blob can be queried
regionally through an unmodified COPC client, in one declarative SQL predicate
over a multi-file corpus, with only the needed byte ranges read. I would like to
propose (a) documenting that path, (b) a small correctness fix, and (c) an
optional batching improvement. I am explicitly not proposing a LiDAR type or
a point-cloud codec — reasoning in the last section.
Background
Point cloud data in geospatial and robotics workloads is overwhelmingly stored as LAZ 1.4 / COPC files: COPC is a LAZ 1.4 file with an octree hierarchy in a
VLR, and readers seek directly to the octree nodes that intersect a query
(spec). The rest of the world's index lives elsewhere —
PostGIS, a tile manifest, or a GeoParquet index — which means "select" and
"fetch" happen in two systems that must be kept consistent.
Lance's blob v2 already provides the two primitives this needs: verbatim byte
storage (the payload is not re-encoded, so there is nothing to lose) and read_blob_ranges() for byte-range access. The docs/src/guide/blob.md "decode
video frames lazily" example is the same pattern with a different media type.
So the question I tried to answer with measurements rather than argument: how far
does that get a LiDAR workload today, and what is actually missing?
What I measured
All numbers are bytes actually read (lance.bytes_read_counter()), not wall
clock. Local NVMe, pylance 12.0.0-beta.15 @ a3eeb8578, pyarrow 25.0.1.
Data: autzen-classified.copc.laz (81,123,042 B, 10,653,336 points, PDRF 7,
sha256 db2d56cd…a27fa) and a 24-file corpus generated from it (2 areas x 12
runs, 1.95 GB — see "Reproducing this").
1. Storage integrity: the payload is preserved byte-for-byte
bytes
source COPC file
81,123,042
Lance blob dataset on disk
81,125,061
overhead
2,019 B (0.0025%)
2. Regional query through an unmodified COPC client
COPC clients already compute the exact set of byte ranges a query needs. Wired
to read_blob_ranges, a 400 m x 400 m query against an 81 MB file returned 131,082 points — identical to a full-scan ground truth — reading 6.40% of the
file (15.6x fewer bytes).
3. Batching collapses the request count
laspy's CopcReader._fetch_all_chunks builds its ranges as byte_queries: [(offset, size), ...] and then issues one seek + readinto
per range. Adding a read_ranges() branch that hands the whole list to read_blob_ranges() in one call changes the request profile without changing
results:
path
bytes
requests
external index -> object store -> whole file
5,528,442
15
Lance, per-range reads
5,529,927
14
Lance, one batched call
5,529,666
1
At this query size COPC's coarse levels dominate the byte count, so the win is in request count, not bytes — which on object storage is what gets billed and
what drives tail latency.
Reads scale with query footprint, and every footprint is a single request:
footprint
points
payload bytes
% of file
requests
100 m
6,999
2,994,418
3.69%
1
200 m
34,254
2,994,418
3.69%
1
428 m
135,556
5,520,429
6.81%
1
800 m
458,811
7,682,900
9.47%
1
1600 m
2,158,111
25,749,081
31.74%
1
(The 100 m and 200 m queries read byte-identical data. That is COPC's additive
LOD: coarse nodes span the whole extent, so small queries hit a floor set by the
root node's size. It is a property of the file, not of Lance.)
4. Multi-file corpus: one predicate prunes across files
24 files, one row and one blob per file, RTREE on the file extent:
predicate
files selected
bytes
requests
region only
24
536,073,693
24
region AND area = 'areaA'
12
268,033,944
12
region AND run < 3
6
134,017,258
6
region outside the corpus
0
0
0
Spatial and attribute predicates both prune, and reads scale with the number of
surviving files. Nothing is read when nothing matches.
5. Governance on the same table
incremental registration: version 2 -> 3, rows 24 -> 27, no rewrite of prior data;
version pinning: lance.dataset(uri, version=2) returns files 24 / points
50,285,880 — identical to the pre-append answer.
That last pair is the part a pile of COPC files plus a manifest does not give
you, and it is why I think this belongs in a table format rather than in a
sidecar index.
What I would propose
1. A cookbook section for chunked/container payloads (documentation only). docs/src/guide/blob.md covers video. A sibling section for "self-indexing
container formats" — COPC is the concrete example — would make the existing
capability discoverable. Worth including: the read_blob_ranges + take_blobs patterns, and the gotcha that a client's query bounds must be the
real data bounds (passing a sentinel z-range makes some COPC readers return zero
points).
Separately, RTREE and the st_* functions currently have format-level
documentation (docs/src/format/index/scalar/rtree.md) but no user guide, so the
frame-selection half of this pattern is hard to find.
2. BlobFile.readinto should accept a memoryview (correctness fix).
python/python/lance/blob.py:463 forwards its argument straight into a binding
that requires bytearray:
io.RawIOBase.readinto is specified to accept any writable buffer, and
third-party decoders pass memoryview — laspy's COPC reader does exactly this
(ChunkIter.next() returns a memoryview), so:
>>> CopcReader.open(ds.take_blobs("copc", indices=[0])[0]) # opens fine
>>> reader.query() # TypeError
TypeError: argument 'dst': 'memoryview' object is not an instance of 'bytearray'
Reproduced directly:
blob.readinto(bytearray(16)) -> 16 OK
blob.readinto(memoryview(...)) -> TypeError
Accepting any writable buffer (as e.g. io.BytesIO does) would make take_blobs() usable as a drop-in file-like source for the whole class of
"container format with an internal index" decoders — not just COPC.
3. Consider a batched read path for file-like blob handles (optional).
A file-like interface can only express one range at a time, but many clients know
all their ranges up front. Exposing that batch (BlobFile.read_ranges already
exists) to decoders that can express it is worth a small documented helper or an
opt-in adapter, given the 14 -> 1 request result above. If maintainers prefer to
keep blob handles minimal, documenting the pattern instead is fine.
What I am not proposing
No LiDAR type, no point-cloud codec, no LAZ decoder in the engine. The
payload is already preserved losslessly as bytes and the ecosystem's readers
already handle it. In a separate evaluation of my own (a local branch, not
this repository) I measured re-encoding points as columns against a LAZ
baseline and it did not come out ahead — roughly +8% storage versus LAZ after
a producer-side reorder, and no DataLoader win when a trainer reads the full
hot attribute set. Those numbers come from that other evaluation rather than
from the scripts in "Reproducing this"; happy to share the details, but the
point here is only that I do not think that path should be pursued.
No 3D/point-level index work. Frame-level spatial selection is a 2D
problem in the workloads I looked at, and the existing 2D RTREE plus st_* was sufficient.
No new format features. Existing storage versions already do everything
measured here.
Reproducing this
Both scripts are self-contained: the sample COPC file is downloaded on first run
and verified against a known size and sha256, and the corpus is generated
locally from it.
pip install geoarrow-rust-core is needed to build the polygon column; without
it the geometry column can be replaced by plain min/max columns and a BTREE
index, at the cost of the frame-selection predicate being a range expression
rather than st_intersects.
Limitations / what I have not measured
Being explicit so the evidence is not over-read:
Everything here is local NVMe. The request-count result is the one that
should change most on object storage (S3/MinIO); I have not measured it there
yet, and that is the next step.
The corpus is generated by copying one real COPC file (autzen, 10.6 M
points) into 24 runs, so every file shares the same extent. Files, blobs and
byte ranges are all real, and the file-level predicate is real, but
partition pruning is ranking degenerate statistics. A corpus of genuinely
distinct extents would be a stronger result and I have not run one.
No concurrent write/compaction workload was tested, so the governance
claims are demonstrated but not stress-tested.
Scale is 10^1 files and 10^7 points per file. Nothing here was run at 10^5
files or as a distributed/object-store deployment.
Offer
I am happy to contribute the documentation section and the readinto fix as
PRs, and to publish the reproduction scripts. If maintainers think the batching
point (3) deserves engine work rather than a documented pattern, I would
appreciate a pointer on where you would want that to live.
Feedback on the framing would be welcome — in particular whether this belongs as
a blob/multimodal use case (how I read the project's direction) or whether you
would rather keep point clouds out of scope for the core docs.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Short version
I have been evaluating Lance as the storage and governance layer for LiDAR
(point cloud) data. The conclusion is narrower than where I started, and I think
it is more useful to the project:
Concretely: a real, unmodified COPC file stored as a blob can be queried
regionally through an unmodified COPC client, in one declarative SQL predicate
over a multi-file corpus, with only the needed byte ranges read. I would like to
propose (a) documenting that path, (b) a small correctness fix, and (c) an
optional batching improvement. I am explicitly not proposing a LiDAR type or
a point-cloud codec — reasoning in the last section.
Background
Point cloud data in geospatial and robotics workloads is overwhelmingly stored as
LAZ 1.4 / COPC files: COPC is a LAZ 1.4 file with an octree hierarchy in a
VLR, and readers seek directly to the octree nodes that intersect a query
(spec). The rest of the world's index lives elsewhere —
PostGIS, a tile manifest, or a GeoParquet index — which means "select" and
"fetch" happen in two systems that must be kept consistent.
Lance's blob v2 already provides the two primitives this needs: verbatim byte
storage (the payload is not re-encoded, so there is nothing to lose) and
read_blob_ranges()for byte-range access. Thedocs/src/guide/blob.md"decodevideo frames lazily" example is the same pattern with a different media type.
So the question I tried to answer with measurements rather than argument: how far
does that get a LiDAR workload today, and what is actually missing?
What I measured
All numbers are bytes actually read (
lance.bytes_read_counter()), not wallclock. Local NVMe,
pylance 12.0.0-beta.15@a3eeb8578, pyarrow 25.0.1.Data:
autzen-classified.copc.laz(81,123,042 B, 10,653,336 points, PDRF 7,sha256
db2d56cd…a27fa) and a 24-file corpus generated from it (2 areas x 12runs, 1.95 GB — see "Reproducing this").
1. Storage integrity: the payload is preserved byte-for-byte
2. Regional query through an unmodified COPC client
COPC clients already compute the exact set of byte ranges a query needs. Wired
to
read_blob_ranges, a 400 m x 400 m query against an 81 MB file returned131,082 points — identical to a full-scan ground truth — reading 6.40% of the
file (15.6x fewer bytes).
3. Batching collapses the request count
laspy'sCopcReader._fetch_all_chunksbuilds its ranges asbyte_queries: [(offset, size), ...]and then issues oneseek+readintoper range. Adding a
read_ranges()branch that hands the whole list toread_blob_ranges()in one call changes the request profile without changingresults:
At this query size COPC's coarse levels dominate the byte count, so the win is in
request count, not bytes — which on object storage is what gets billed and
what drives tail latency.
Reads scale with query footprint, and every footprint is a single request:
(The 100 m and 200 m queries read byte-identical data. That is COPC's additive
LOD: coarse nodes span the whole extent, so small queries hit a floor set by the
root node's size. It is a property of the file, not of Lance.)
4. Multi-file corpus: one predicate prunes across files
24 files, one row and one blob per file,
RTREEon the file extent:area = 'areaA'run < 3Spatial and attribute predicates both prune, and reads scale with the number of
surviving files. Nothing is read when nothing matches.
5. Governance on the same table
lance.dataset(uri, version=2)returns files 24 / points50,285,880 — identical to the pre-append answer.
That last pair is the part a pile of COPC files plus a manifest does not give
you, and it is why I think this belongs in a table format rather than in a
sidecar index.
What I would propose
1. A cookbook section for chunked/container payloads (documentation only).
docs/src/guide/blob.mdcovers video. A sibling section for "self-indexingcontainer formats" — COPC is the concrete example — would make the existing
capability discoverable. Worth including: the
read_blob_ranges+take_blobspatterns, and the gotcha that a client's query bounds must be thereal data bounds (passing a sentinel z-range makes some COPC readers return zero
points).
Separately,
RTREEand thest_*functions currently have format-leveldocumentation (
docs/src/format/index/scalar/rtree.md) but no user guide, so theframe-selection half of this pattern is hard to find.
2.
BlobFile.readintoshould accept amemoryview(correctness fix).python/python/lance/blob.py:463forwards its argument straight into a bindingthat requires
bytearray:io.RawIOBase.readintois specified to accept any writable buffer, andthird-party decoders pass
memoryview—laspy's COPC reader does exactly this(
ChunkIter.next()returns amemoryview), so:Reproduced directly:
Accepting any writable buffer (as e.g.
io.BytesIOdoes) would maketake_blobs()usable as a drop-in file-like source for the whole class of"container format with an internal index" decoders — not just COPC.
3. Consider a batched read path for file-like blob handles (optional).
A file-like interface can only express one range at a time, but many clients know
all their ranges up front. Exposing that batch (
BlobFile.read_rangesalreadyexists) to decoders that can express it is worth a small documented helper or an
opt-in adapter, given the 14 -> 1 request result above. If maintainers prefer to
keep blob handles minimal, documenting the pattern instead is fine.
What I am not proposing
payload is already preserved losslessly as bytes and the ecosystem's readers
already handle it. In a separate evaluation of my own (a local branch, not
this repository) I measured re-encoding points as columns against a LAZ
baseline and it did not come out ahead — roughly +8% storage versus LAZ after
a producer-side reorder, and no DataLoader win when a trainer reads the full
hot attribute set. Those numbers come from that other evaluation rather than
from the scripts in "Reproducing this"; happy to share the details, but the
point here is only that I do not think that path should be pursued.
problem in the workloads I looked at, and the existing 2D
RTREEplusst_*was sufficient.measured here.
Reproducing this
Both scripts are self-contained: the sample COPC file is downloaded on first run
and verified against a known size and sha256, and the corpus is generated
locally from it.
pip install geoarrow-rust-coreis needed to build the polygon column; withoutit the geometry column can be replaced by plain min/max columns and a
BTREEindex, at the cost of the frame-selection predicate being a range expression
rather than
st_intersects.Limitations / what I have not measured
Being explicit so the evidence is not over-read:
should change most on object storage (S3/MinIO); I have not measured it there
yet, and that is the next step.
autzen, 10.6 Mpoints) into 24 runs, so every file shares the same extent. Files, blobs and
byte ranges are all real, and the file-level predicate is real, but
partition pruning is ranking degenerate statistics. A corpus of genuinely
distinct extents would be a stronger result and I have not run one.
claims are demonstrated but not stress-tested.
files or as a distributed/object-store deployment.
Offer
I am happy to contribute the documentation section and the
readintofix asPRs, and to publish the reproduction scripts. If maintainers think the batching
point (3) deserves engine work rather than a documented pattern, I would
appreciate a pointer on where you would want that to live.
Feedback on the framing would be welcome — in particular whether this belongs as
a blob/multimodal use case (how I read the project's direction) or whether you
would rather keep point clouds out of scope for the core docs.
All reactions