BlobNet is a focused U-Net pipeline for atom localization in atomic-resolution STEM images.
The maintained workflow has two explicit steps:
- Generate deterministic train, validation, and test datasets from YAML.
- Train U-Net from the saved dataset and write checkpoints, losses, and metrics.
uv syncOn Windows or Linux with an NVIDIA GPU, BlobNet uses the CUDA 12.8 PyTorch wheel index through pyproject.toml. Refresh the lockfile and environment after pulling changes:
uv lock
uv syncCheck that the uv environment can see CUDA before starting a long training run:
uv run python -c "import torch; print(torch.__version__); print(torch.version.cuda); print(torch.cuda.is_available()); print(torch.cuda.device_count()); print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'no cuda device')"If CUDA is still unavailable, confirm that the NVIDIA driver is visible outside Python:
nvidia-smiDataset parameters live in configs/dataset_configs/. The default configuration generates random atom images:
random.yamlsquare.yamlhexagonal.yaml
uv run blobnet-generate-dataset \
--config configs/dataset_configs/random.yamlUseful overrides:
uv run blobnet-generate-dataset \
--config configs/dataset_configs/random.yaml \
--output-dir /tmp/blobnet_dataset \
--train-samples 128 \
--val-samples 32 \
--test-samples 32 \
--num-workers 0The generator writes one compressed NPZ file per sample under train/, val/, and test/, plus dataset_manifest.yaml with the resolved generation parameters.
Use --num-workers 0 to use all available CPU cores, or pass a positive worker count to cap CPU use.
Model and training parameters live in configs/model_configs/:
uv run blobnet-train \
--config configs/model_configs/base_unet.yamlCommon settings can be overridden without editing YAML:
uv run blobnet-train \
--config configs/model_configs/base_unet.yaml \
--dataset-dir /tmp/blobnet_dataset \
--output-dir /tmp/blobnet_training \
--epochs 10 \
--batch-size 4 \
--learning-rate 0.001 \
--device autoTraining writes:
unet_best.pthunet_loss_history.npzloss_history.csvloss_curves.pngtraining_metrics.jsonresolved_config.yaml
The manuscript comparison uses three matched U-Net configs:
configs/model_configs/random_unet.yamlconfigs/model_configs/square_unet.yamlconfigs/model_configs/hexagonal_unet.yaml
Run the full pipeline from existing datasets to trained models and figures:
uv run blobnet-train-manuscript --device cudaThis command reuses complete random, square, and hexagonal datasets when they already match the config split counts, trains all three U-Nets from scratch, checks that each checkpoint was written, and rebuilds the manuscript figures. If a dataset is missing, the script generates it. If a dataset is partial or has mismatched split counts, the script stops instead of silently mixing old and new data.
Start from training when the datasets are already generated and should not be touched:
uv run blobnet-train-manuscript --device cuda --skip-dataset-generationRegenerate every dataset before training when you need fresh data:
uv run blobnet-train-manuscript --device cuda --regenerate-datasetsDataset generation runs in parallel by default when it is needed; use --dataset-workers N to cap it.
The manuscript and SI generation scripts load the publication-ready NSID microscopy files in experimental_data/ through pyTEMlib.file_tools.open_file.
The repository includes the five publication-ready experimental images, the exact model checkpoints used by the figure workflow, their available training records, the locked Python environment, and the manuscript source. Figure 3 uses its separately packaged figure3_random and figure3_hexagonal checkpoints so that its published detections remain unchanged. All individual binary files are below GitHub's 100 MB limit.
After cloning, install the locked environment and audit the publication inputs:
uv sync --frozen
uv run blobnet-reproduce-publication --target auditRegenerate the manuscript figures and compile the manuscript when latexmk is installed:
uv run blobnet-reproduce-publication --target manuscript --device auto --compile-latexRegenerate the full SI, including the gold-in-TiO2 figure:
uv run blobnet-reproduce-publication --target si --device auto --compile-latexUse --target all to run both workflows. Generated files are written under outputs/; manuscript figures are copied to publication/manuscript/figures/ before LaTeX compilation. The packaged checkpoints provide deterministic figure inputs. The separate blobnet-train-manuscript command independently regenerates synthetic datasets and retrains the three primary models from the tracked YAML configurations; retrained floating-point weights can vary across hardware and software backends.
For an interactive inference walkthrough, open notebooks/quickstart_inference.ipynb. It loads the three packaged models, compares them on one reproducible simulated image, and ends with an editable path that reads experimental microscopy data through pyTEMlib.file_tools.open_file before running the already-loaded models.
from blobnet import RandomAtomImageConfig, generate_atom_image
config = RandomAtomImageConfig(image_shape=(256, 256))
sample = generate_atom_image(config)
image, target = sample['image'], sample['target']blobnet/ reusable models, data generation, metrics, and plotting
configs/
dataset_configs/ saved-dataset generation parameters
model_configs/ model, loss, and training parameters
scripts/
generate_training_dataset.py
train_unet.py
notebooks/ dataset and experimental-image exploration
experimental_data/ tracked publication-ready NSID microscopy images
artifacts/manuscript_models/ exact publication checkpoints and training records
publication/manuscript/ manuscript LaTeX, bibliography, captions, and sections
See CONTRIBUTING.md for development conventions.