STORM_pipe is a Nextflow-based pipeline that generates quality control and analysis files from raw STORM-seq FASTQ
files. It is designed to work with STORM-seq data that has unique molecular indexes (UMIs) incorporated and will not
work with the original STORM-seq data that does not include UMIs.
STORM_pipe can be downloaded from GitHub via git or from the
releases page.
Via git:
git clone git@github.com:huishenlab/STORM_pipe.gitSTORM_pipe is set up such that it can be downloaded once and used across all STORM-seq experiments. This ensures
consistent processing across experiments and multiple instances of the same pipeline spread out across directories.
STORM_pipe requires three dependencies out of the box:
- Nextflow
- Minimum version: 24.10.0
- conda
- SLURM
- SLURM can easily be avoided by changing
process.executor = 'slurm'innextflow.config
- SLURM can easily be avoided by changing
All dependencies needed for running the workflows will be retrieved at runtime with conda. Conda environments files
may be found in the envs/ directory.
STORM_pipe has contains two workflows (described below in more detail):
- Downloading and creating reference files
- Config parameter:
params.mode = 'ref' - Description: Download ERCC reference FASTA, human (hg38) and mouse (mm10) reference FASTAs, human and mouse transcript GTFs, human and mouse enhancer annotation BEDs, and human and mouse repeat annotation BEDs. It then builds STAR and kallisto indexes for merged human/ERCC and mouse/ERCC references files, and extracts gene information needed for quality control.
- Other notes:
- This workflow only needs to be run one time!
- Reference files for species other than human and mouse will need to be generated by the user following the process laid out in the reference workflow.
- Config parameter:
- Process raw FASTQ files
- Config parameter:
params.mode = 'process' - Description: Trim FASTQs and process with either STARsolo or kallisto-bustools to create counts matrix files. Also produces quality control files for visualization with MultiQC and a custom visualization script.
- Other notes:
- Once the reference files (indexes and annotations) are set, this is the only mode that you will need to run.
- Config parameter:
The pipeline is governed by the nextflow.config file, which contains runtime parameters and Nextflow-specific options.
The params.mode parameter selects which workflow to run (see Workflows above for descriptions on the options). Each
mode has a few parameters that may need adjusted when running the pipeline.
params.mode = 'ref' has one parameters (dir_ref) which would need to be adjusted, though it is suggested to leave
this parameter with its default value so the index and annotation paths do not need to be updated.
params.mode = 'process' has several parameters that will need to be adjusted for each experiment:
params.experiment_nameis a unique name for each experiment that you procesparams.fastq_diris the path to your sequenced FASTQsparams.cell_metadatais a TSV file containing metadata for each samples (see below for details, including required columns)
There are a variety of other parameters that can be adjusted, but the ones described above are the primary ones you may be adjusting.
Once the configuration has been set, you can run the pipeline from the command line (or in a script submitted to a resource management system like SLURM):
NXF_SYNTAX_PARSER=v2 nextflow run \
main.nf \
-with-conda \
-resumeNote, -with-conda must be included for the environment files to be properly loaded for managing dependencies.
The only required input for generating reference and annotation files is the params.dir_ref parameter. We suggest
leaving this as the default value though. All other files will be downloaded or created from the downloaded files.
The workflow to create the reference files includes these components:
- Download ERCC reference FASTA and formatting an ERCC transcript GTF
- Download reference FASTA, transcript GTF, enhancer annotation BED, and repeat annotation BED for both mouse (hg38) and mouse (mm10)
- Merge reference FASTAs, transcript GTFs, and annotation files for human/ERCC and mouse/ERCC
- Extract gene IDs used in quality control for human and mouse
- Create human and mouse STAR index
- Create kallisto NAC index
Outputs of note are (examples shown for human, but comparable mouse files/directories are also generated):
- STAR index (
references/star_index_human) - kallisto NAC index (
references/kallisto_index_human/nac_transcriptome.idx) - kallisto NAC transcripts to gene map (
references/kallisto_index_human/nac_transcripts_to_genes.txt) - Annotated space BED (
references/GRCh38_ERCC_enhancers_rmsk.merged.sorted.bed.gz) - rRNA gene IDs (
references/gene_ids.hg38.rrna.txt) - Mitochondrial DNA gene IDs (
references/gene_ids.hg38.mito.txt) - ERCC gene IDs (
references/gene_ids.hg38.ercc.txt)
The three primary inputs you will need to set to process STORM-seq data are:
- A name for your experiment (
params.experiment_name). This must be unique to ensure outputs are not overwritten. - The path to your sequenced FASTQ data (
params.fastq_dir). Ideally, this is an absolute path. The FASTQs must be in one of these two forms:<WELL_ID>_R[1,2].fastq.gzor<WELL_ID>_R[1,2]_001.fastq.gzfor the pipeline to properly map samples to FASTQs - A metadata files describing your samples. There are four required columns. Any other columns will be ignored. The four
required columns are:
WELL_ID: Must match the first column ofresults/<experiment_name>/samplesheets/<experiment_name>.samples.tsv. This is basically everything before_R1*or_R2*in the FASTQ file name. Must be unique across samples.CELL_BC: Cell barcode for each sample.WELL: Which well the cell is in. The combination ofWELLandPLATEmust be unique to ensure each cell is correctly plotted in the QC plots.PLATE: Which plate the cell is included in. The combination ofWELLandPLATEmust be unique to ensure each cell is correctly plotted in the QC plots.
The workflow to generate quality control and counts matrices is:
- Create a samplesheet based on the FASTQs found in
params.fastq_dir. The sample ID (orWELL_ID) is everything prior to the_R[1,2]*in the file name. All instances of_001are removed from the file name, so avoid including these as best as possible. - Run FASTQC on the raw FASTQ files.
- Trim adapters from FASTQs and run FASTQC on trimmed data.
- Run STAR on the trimmed FASTQs using the STORMsolo pipeline.
- Run kb-python on the trimmed data.
- Combine individual kb-python results NAC-compliant spliced/unspliced/total setup.
- Run
samtools statandsamtools flagstaton STAR BAM. - Find fraction of reads that align to non-annotated space in STAR BAM.
- Run MultiQC.
- Collect QC data and generate QC plots.
For default parameters, the output of the data processing workflow can be found at results/<experiment_name>/analysis.
Within that directory, outputs of note are:
fastqc: FASTQC on raw FASTQ fileskallisto: kb-python outputs for individual cells and combined NAC resultsmultiqc: MultiQC resultsplots: Quality control plots and corresponding datasamplesheets: Map fromWELL_IDto file path for read 1 and read 2 FASTQsstar_align: Outputs from STAR alignmentstrimmed: Trimmed FASTQs and FASTQC on those files

