Ready to be used by early adopters. Have your code depend on a particular git commit as API is not yet frozen
This is an experimental crate for processing of large scale imaging data. Modern microscopes are able to churn out TB-scale datasets, making it impossible to load them into memory at once and using old bases. The solution is (1) "out of core processing", where only a part of the data is present in memory at once. Furthermore, (2) multithreading, GPUs and computing on multiple computers in parallel is required.
Figuring out the optimal compute order is hard (likely NP-hard). The following factors need to be taken into account:
- How much memory is available?
- How many threads are available, and what CPU/how much cache memory?
- How many computers are available?
- Is a GPU available? And if so, what type, and which compute nodes have them?
- How much time does it take to read data?
- How much time does it take to write data?
- How much time does it take to compress data?
- How well does the data compress?
- What operations are performed, and in which order?
- What precision does the data need to be stored in?
This crate aims to resolve the problem using the following ingredients:
- Operations are represented as a DAG (direct acyclic graph), representing dependencies
- Borrowing from database query planners, statistics about compute times are gathered during execution
- A 6d scheduler figures out the best order and adapts in realtime based on statistics
- Designed for multiple compute nodes, GPUs and heterogenous compute environments from day one
- OME-Zarr is used to enable distributed computing on chunks of image data
In the current example benchmarks, Blockflow's planned Zarr pipelines were approximately:
- 5× faster than ImgLib2 on 10-image segmentation batches
- 4.5× faster than OpenCV on the same batches
- 9× faster than scikit-image on the same batches
- 3× faster than Dask-image on two larger images
- 5.3× faster than Python StarDist end to end on a real DAPI subset
- 1.2× faster than Python Cellpose end to end on a real DAPI subset
- 1.11× faster than original YOLOv11 on matched FP32 CUDA inference
- within 7% of Python Cellpose3D, with matching object counts on two real crops
- 1.24× faster than Python StarDist3D on the crowded crop and within 8% on the isolated crop, with exact labels
These are ratios for the specified workloads, excluding fixture preparation and compilation. See BENCHMARKS.md for each run's timed region, counts, measurements, output checks, and reproduction commands. Earlier results are archived in OLD_BENCHMARKS.md.
examples/stardist-ome-zarr shows normal
out-of-core use on an OME-Zarr channel: attach the DAPI plane, price the real
StarDist phase with the planner, execute it, and write a multiscale label layer
plus an ID-linked measurement table. The result is discovered directly by
newvolim and preserves the cell IDs needed for later per-channel measurements.
examples/cellpose-ome-zarr runs the same normal
OME-Zarr reader, planner, executor, multiscale label writer, and linked table
flow with Cellpose. It supports CPU and an optional CUDA build and writes a
second annotation layer that newvolim discovers directly.
Volumetric fluorescence uses the axis-aware variants:
cellpose-3d-ome-zarr,
stardist-3d-ome-zarr, and the separate
yolo-3d-ome-zarr distillation, training, and
inference workflow. All commands in those examples use cargo run --release.
Details about the design are located in docs/design/ ; docs need cleaning
cargo test --release
cargo test --release --features gui,distributed,zarr,model-segment
Use release builds for local tests and examples. This keeps execution behavior and performance measurements consistent with the documented benchmarks.
The suite that asserts is the suite that runs. The 39 #[ignore]d tests are
measurements — tables of nanoseconds per voxel, of resident bytes, of how
far repetitions moved — and they print rather than assert, because nothing in
this crate asserts on a duration. Run them deliberately, on a quiet machine:
cargo test --release -- --ignored --nocapture
The features CI does not cover are the ones a hosted runner cannot: fftw
wants a system libfftw3 (Linux and macOS jobs install one; there is no
Windows job), and everything *-cuda wants a device.
MIT (but note that the code is AI generated so no guarantees about provenance)