Skip to content

About

Pipeline for processing STORM-seq data

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

STORM_pipe

STORM_pipe is a Nextflow-based pipeline that generates quality control and analysis files from raw STORM-seq FASTQ files. It is designed to work with STORM-seq data that has unique molecular indexes (UMIs) incorporated and will not work with the original STORM-seq data that does not include UMIs.

Pipeline Overview

Download

STORM_pipe can be downloaded from GitHub via git or from the releases page.

Via git:

git clone git@github.com:huishenlab/STORM_pipe.git

STORM_pipe is set up such that it can be downloaded once and used across all STORM-seq experiments. This ensures consistent processing across experiments and multiple instances of the same pipeline spread out across directories.

Dependencies

STORM_pipe requires three dependencies out of the box:

  • Nextflow
    • Minimum version: 24.10.0
  • conda
  • SLURM
    • SLURM can easily be avoided by changing process.executor = 'slurm' in nextflow.config

All dependencies needed for running the workflows will be retrieved at runtime with conda. Conda environments files may be found in the envs/ directory.

Workflows

STORM_pipe has contains two workflows (described below in more detail):

  • Downloading and creating reference files
    • Config parameter: params.mode = 'ref'
    • Description: Download ERCC reference FASTA, human (hg38) and mouse (mm10) reference FASTAs, human and mouse transcript GTFs, human and mouse enhancer annotation BEDs, and human and mouse repeat annotation BEDs. It then builds STAR and kallisto indexes for merged human/ERCC and mouse/ERCC references files, and extracts gene information needed for quality control.
    • Other notes:
      • This workflow only needs to be run one time!
      • Reference files for species other than human and mouse will need to be generated by the user following the process laid out in the reference workflow.
  • Process raw FASTQ files
    • Config parameter: params.mode = 'process'
    • Description: Trim FASTQs and process with either STARsolo or kallisto-bustools to create counts matrix files. Also produces quality control files for visualization with MultiQC and a custom visualization script.
    • Other notes:
      • Once the reference files (indexes and annotations) are set, this is the only mode that you will need to run.

Running the Pipeline

The pipeline is governed by the nextflow.config file, which contains runtime parameters and Nextflow-specific options. The params.mode parameter selects which workflow to run (see Workflows above for descriptions on the options). Each mode has a few parameters that may need adjusted when running the pipeline.

params.mode = 'ref' has one parameters (dir_ref) which would need to be adjusted, though it is suggested to leave this parameter with its default value so the index and annotation paths do not need to be updated.

params.mode = 'process' has several parameters that will need to be adjusted for each experiment:

  • params.experiment_name is a unique name for each experiment that you proces
  • params.fastq_dir is the path to your sequenced FASTQs
  • params.cell_metadata is a TSV file containing metadata for each samples (see below for details, including required columns)

There are a variety of other parameters that can be adjusted, but the ones described above are the primary ones you may be adjusting.

Once the configuration has been set, you can run the pipeline from the command line (or in a script submitted to a resource management system like SLURM):

NXF_SYNTAX_PARSER=v2 nextflow run \
    main.nf \
    -with-conda \
    -resume

Note, -with-conda must be included for the environment files to be properly loaded for managing dependencies.

Reference Generation

Inputs

The only required input for generating reference and annotation files is the params.dir_ref parameter. We suggest leaving this as the default value though. All other files will be downloaded or created from the downloaded files.

Workflow Components

The workflow to create the reference files includes these components:

  • Download ERCC reference FASTA and formatting an ERCC transcript GTF
  • Download reference FASTA, transcript GTF, enhancer annotation BED, and repeat annotation BED for both mouse (hg38) and mouse (mm10)
  • Merge reference FASTAs, transcript GTFs, and annotation files for human/ERCC and mouse/ERCC
  • Extract gene IDs used in quality control for human and mouse
  • Create human and mouse STAR index
  • Create kallisto NAC index

Outputs

Outputs of note are (examples shown for human, but comparable mouse files/directories are also generated):

  • STAR index (references/star_index_human)
  • kallisto NAC index (references/kallisto_index_human/nac_transcriptome.idx)
  • kallisto NAC transcripts to gene map (references/kallisto_index_human/nac_transcripts_to_genes.txt)
  • Annotated space BED (references/GRCh38_ERCC_enhancers_rmsk.merged.sorted.bed.gz)
  • rRNA gene IDs (references/gene_ids.hg38.rrna.txt)
  • Mitochondrial DNA gene IDs (references/gene_ids.hg38.mito.txt)
  • ERCC gene IDs (references/gene_ids.hg38.ercc.txt)

Data Processing

Inputs

The three primary inputs you will need to set to process STORM-seq data are:

  • A name for your experiment (params.experiment_name). This must be unique to ensure outputs are not overwritten.
  • The path to your sequenced FASTQ data (params.fastq_dir). Ideally, this is an absolute path. The FASTQs must be in one of these two forms: <WELL_ID>_R[1,2].fastq.gz or <WELL_ID>_R[1,2]_001.fastq.gz for the pipeline to properly map samples to FASTQs
  • A metadata files describing your samples. There are four required columns. Any other columns will be ignored. The four required columns are:
    • WELL_ID: Must match the first column of results/<experiment_name>/samplesheets/<experiment_name>.samples.tsv. This is basically everything before _R1* or _R2* in the FASTQ file name. Must be unique across samples.
    • CELL_BC: Cell barcode for each sample.
    • WELL: Which well the cell is in. The combination of WELL and PLATE must be unique to ensure each cell is correctly plotted in the QC plots.
    • PLATE: Which plate the cell is included in. The combination of WELL and PLATE must be unique to ensure each cell is correctly plotted in the QC plots.

Workflow Components

The workflow to generate quality control and counts matrices is:

  • Create a samplesheet based on the FASTQs found in params.fastq_dir. The sample ID (or WELL_ID) is everything prior to the _R[1,2]* in the file name. All instances of _001 are removed from the file name, so avoid including these as best as possible.
  • Run FASTQC on the raw FASTQ files.
  • Trim adapters from FASTQs and run FASTQC on trimmed data.
  • Run STAR on the trimmed FASTQs using the STORMsolo pipeline.
  • Run kb-python on the trimmed data.
  • Combine individual kb-python results NAC-compliant spliced/unspliced/total setup.
  • Run samtools stat and samtools flagstat on STAR BAM.
  • Find fraction of reads that align to non-annotated space in STAR BAM.
  • Run MultiQC.
  • Collect QC data and generate QC plots.

Outputs

For default parameters, the output of the data processing workflow can be found at results/<experiment_name>/analysis. Within that directory, outputs of note are:

  • fastqc: FASTQC on raw FASTQ files
  • kallisto: kb-python outputs for individual cells and combined NAC results
  • multiqc: MultiQC results
  • plots: Quality control plots and corresponding data
  • samplesheets: Map from WELL_ID to file path for read 1 and read 2 FASTQs
  • star_align: Outputs from STAR alignments
  • trimmed: Trimmed FASTQs and FASTQC on those files

About

Pipeline for processing STORM-seq data

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages