Skip to content

Latest commit

 

History

79 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

blockflow

Ready to be used by early adopters. Have your code depend on a particular git commit as API is not yet frozen

This is an experimental crate for processing of large scale imaging data. Modern microscopes are able to churn out TB-scale datasets, making it impossible to load them into memory at once and using old bases. The solution is (1) "out of core processing", where only a part of the data is present in memory at once. Furthermore, (2) multithreading, GPUs and computing on multiple computers in parallel is required.

Figuring out the optimal compute order is hard (likely NP-hard). The following factors need to be taken into account:

  • How much memory is available?
  • How many threads are available, and what CPU/how much cache memory?
  • How many computers are available?
  • Is a GPU available? And if so, what type, and which compute nodes have them?
  • How much time does it take to read data?
  • How much time does it take to write data?
  • How much time does it take to compress data?
  • How well does the data compress?
  • What operations are performed, and in which order?
  • What precision does the data need to be stored in?

This crate aims to resolve the problem using the following ingredients:

  • Operations are represented as a DAG (direct acyclic graph), representing dependencies
  • Borrowing from database query planners, statistics about compute times are gathered during execution
  • A 6d scheduler figures out the best order and adapts in realtime based on statistics
  • Designed for multiple compute nodes, GPUs and heterogenous compute environments from day one
  • OME-Zarr is used to enable distributed computing on chunks of image data

Benchmarks

In the current example benchmarks, Blockflow's planned Zarr pipelines were approximately:

  • 5× faster than ImgLib2 on 10-image segmentation batches
  • 4.5× faster than OpenCV on the same batches
  • 9× faster than scikit-image on the same batches
  • 3× faster than Dask-image on two larger images
  • 5.3× faster than Python StarDist end to end on a real DAPI subset
  • 1.2× faster than Python Cellpose end to end on a real DAPI subset
  • 1.11× faster than original YOLOv11 on matched FP32 CUDA inference
  • within 7% of Python Cellpose3D, with matching object counts on two real crops
  • 1.24× faster than Python StarDist3D on the crowded crop and within 8% on the isolated crop, with exact labels

These are ratios for the specified workloads, excluding fixture preparation and compilation. See BENCHMARKS.md for each run's timed region, counts, measurements, output checks, and reproduction commands. Earlier results are archived in OLD_BENCHMARKS.md.

Whole-slide StarDist annotation example

examples/stardist-ome-zarr shows normal out-of-core use on an OME-Zarr channel: attach the DAPI plane, price the real StarDist phase with the planner, execute it, and write a multiscale label layer plus an ID-linked measurement table. The result is discovered directly by newvolim and preserves the cell IDs needed for later per-channel measurements.

Whole-slide Cellpose annotation example

examples/cellpose-ome-zarr runs the same normal OME-Zarr reader, planner, executor, multiscale label writer, and linked table flow with Cellpose. It supports CPU and an optional CUDA build and writes a second annotation layer that newvolim discovers directly.

Volumetric fluorescence uses the axis-aware variants: cellpose-3d-ome-zarr, stardist-3d-ome-zarr, and the separate yolo-3d-ome-zarr distillation, training, and inference workflow. All commands in those examples use cargo run --release.

Design notes

Details about the design are located in docs/design/ ; docs need cleaning

Testing

cargo test --release
cargo test --release --features gui,distributed,zarr,model-segment

Use release builds for local tests and examples. This keeps execution behavior and performance measurements consistent with the documented benchmarks.

The suite that asserts is the suite that runs. The 39 #[ignore]d tests are measurements — tables of nanoseconds per voxel, of resident bytes, of how far repetitions moved — and they print rather than assert, because nothing in this crate asserts on a duration. Run them deliberately, on a quiet machine:

cargo test --release -- --ignored --nocapture

The features CI does not cover are the ones a hosted runner cannot: fftw wants a system libfftw3 (Linux and macOS jobs install one; there is no Windows job), and everything *-cuda wants a device.

License

MIT (but note that the code is AI generated so no guarantees about provenance)

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages