Skip to content

About

Repository for development of new single cell SQANTI tool

Resources

Stars

9 stars

Watchers

1 watching

Forks

Repository files navigation

SQANTI-single-cell

⚠️⚠️ WARNING: SQANTI-single-cell is currently under development. Use at your own risk. ⚠️⚠️

SQANTI-single-cell (SQANTI-sc) is a pipeline for the structural and quality control of long-read single-cell transcriptomics datasets. It extends the capabilities of SQANTI3 and SQANTI-reads frameworks to provide cell-level structural and quality control metrics.

Table of Contents:

Prerequisites & Installation

0. Install Anaconda

Make sure you have installed Anaconda. If not, you can download the generic installer for Linux here.

1. Install SQANTI-sc

You can install SQANTI-sc by downloading the source code or cloning the repository.

Option A: Download Source Code (Recommended for General Users) We recommend this option for general users who want to use the stable version of the tool. Download the latest release from the Releases page (if available) or download the repository as a ZIP file.

wget https://github.com/ConesaLab/SQANTI-single-cell/archive/refs/heads/main.zip
unzip main.zip
mv SQANTI-single-cell-main SQANTI-single-cell
cd SQANTI-single-cell

Option B: Clone Repository (For Developers) If you intend to contribute to the development of SQANTI-sc, please clone the repository. This option sets up a git repository and is NOT recommended for general users unless you plan to submit pull requests or track the latest development changes.

git clone https://github.com/ConesaLab/SQANTI-single-cell.git
cd SQANTI-single-cell

2. Install SQANTI3

SQANTI-sc requires a functional installation of SQANTI3. Please follow the SQANTI3 Installation Instructions to install SQANTI3.

Important: SQANTI3 Location. SQANTI-sc needs to know where SQANTI3 is installed. You have two options:

  • Option A (Simpler): Place the SQANTI3 folder inside the SQANTI-single-cell directory.
  • Option B (Flexible): Install SQANTI3 anywhere and set the SQANTI3_DIR environment variable:
export SQANTI3_DIR=/path/to/your/SQANTI3/directory

3. Set up the Environment

We provide a unified Conda environment file SQANTI-sc_env.yml that includes all dependencies for both SQANTI3 and SQANTI-sc (including R packages for reporting).

conda env create -f SQANTI-sc_env.yml
conda activate SQANTI-sc_env

4. Docker Support (Coming Soon)

Future releases of SQANTI-sc will be containerized and available on DockerHub. Currently, please use the Conda installation method.

Getting Ready

Activate the SQANTI-sc conda environment:

conda activate SQANTI-sc_env

Arguments and parameters in SQANTI-sc

The SQANTI-sc quality control script accepts the following arguments:

usage: sqanti_sc.py [-h] --mode {reads,isoforms} --design DESIGN --refGTF REFGTF --refFasta REFFASTA 
                    [--out_dir OUT_DIR] [--input_dir INPUT_DIR] [--report {pdf,html,both,skip}] 
                    [-@ SAMTOOLS_CPUS] [--verbose] [--skip_hash] [--ignore_cell_summary]
                    [--multisample_report] [--multisample_report_prefix PREFIX]
                    [--write_per_cell_outputs] [--export_h5ad] [--run_clustering]
                    [--normalization {log1p,sqrt,pearson}] [--n_neighbors N] [--n_pc N]
                    [--resolution RES] [--n_top_genes N] [--clustering_method {leiden,louvain,kmeans}]
                    [--n_clusters N]
                    [--min_ref_len MIN_REF_LEN] [--genename] 
                    [--novel_gene_prefix NOVEL_GENE_PREFIX] [--ref_cov_min_pct PCT]
                    [--short_reads SHORT_READS] [--SR_bam SR_BAM] 
                    [--aligner_choice {minimap2,uLTRA,gmap,deSALT}] [-x GMAP_INDEX]
                    [--include_ORF] [--orf_input ORF_INPUT]
                    [--CAGE_peak CAGE_PEAK] [--polyA_motif_list POLYA_MOTIF_LIST] 
                    [--polyA_peak POLYA_PEAK] [--phyloP_bed PHYLOP_BED] 
                    [--isoAnnotLite] [--gff3 GFF3]
                    [--isoform_hits] [--ratio_TSS_metric {max,mean,median,3quartile}] 
                    [--min_cov MIN_COV] [--ratio_TSS RATIO_TSS] 
                    [-t CPUS] [-n CHUNKS] [-l {ERROR,WARNING,INFO,DEBUG}] 
                    [--is_fusion] [-v]
Arguments explanation
Structural and Quality Analysis of Single-Cell Isoforms

options:
  -h, --help            show this help message and exit

Required arguments:
  --refFasta REFFASTA   Reference genome (Fasta format)
  --refGTF REFGTF       Reference annotation file (GTF format)
  --design DESIGN, -de DESIGN
                        Design file (CSV) containing sampleID, file_acc, and cell/abundance metadata
  --mode {reads,isoforms}, -m {reads,isoforms}
                        Input data mode: 'reads' for long reads, 'isoforms' for collapsed transcripts

SQANTI-sc specific options:
  --out_dir OUT_DIR, -d OUT_DIR
                        Output directory (Default: current dir)
  --input_dir INPUT_DIR, -i INPUT_DIR
                        Input directory for files referenced in design (Default: current dir)
  --report {pdf,html,both,skip}
                        Report format (Default: pdf)
  --samtools_cpus CPUS, -@ CPUS
                        Threads for samtools sorting/indexing (Default: 10)
  --verbose             Print detailed logs during execution
  --skip_hash           Skip UJC hashing calculation step
  --ignore_cell_summary Don't save cell summary table in report
  --multisample_report  Generate a multisample cohort report
  --multisample_report_prefix PREFIX
                        Prefix for multisample report (Default: SQANTI_sc_multisample_report)
  --write_per_cell_outputs
                        Write per-cell gene/UJC counts and CV matrices
  --export_h5ad         Export an AnnData .h5ad file per sample with count matrix and cell QC metadata,
                        compatible with Scanpy and Seurat
  --min_cov MIN_COV     Minimum short read coverage to validate an isoform (default: 1)
  --ratio_TSS RATIO_TSS Minimum ratio_TSS to validate a TSS (default: 2.0)

SQANTI-sc Clustering and UMAP options:
  --run_clustering      Run cell clustering and UMAP analysis
  --normalization {log1p,sqrt,pearson}
                        Normalization method (Default: log1p)
  --n_neighbors N       Number of neighbors for UMAP (Default: 15)
  --n_pc N              Number of principal components (Default: 30)
  --resolution RES      Resolution for Leiden clustering (Default: 0.5)
  --n_top_genes N       Number of highly variable genes (Default: 2000)
  --clustering_method {leiden,louvain,kmeans}
                        Clustering method (Default: leiden)
  --n_clusters N        Number of clusters for K-means (Default: 10)

SQANTI3 Customization and filtering:
  --min_ref_len MIN_REF_LEN
                        Minimum reference transcript length (default: 0 bp)
  --genename            Use gene_name tag from GTF to define genes. Default: gene_id
  --novel_gene_prefix PREFIX
                        Prefix for novel isoforms
  --ref_cov_min_pct PCT
                        Minimum % of reference transcript length a read must cover (Default: 45.0)

Aligner and mapping options:
  --aligner_choice {minimap2,uLTRA,gmap,deSALT}
                        Select aligner for FASTA input (default: minimap2)
  -x GMAP_INDEX, --gmap_index GMAP_INDEX
                        Path to gmap_build index. Mandatory if using GMAP.

ORF prediction:
  --include_ORF         Run ORF prediction. Isoforms mode only: ignored in reads
                        mode, where predicting an ORF for every read is
                        prohibitively slow, so the coding, non-coding and NMD
                        columns are NA.
  --orf_input ORF_INPUT Input fasta to run ORF on.

SQANTI3 Orthogonal data inputs:
  
  --CAGE_peak CAGE_PEAK FANTOM5 Cage Peak (BED format)
  --polyA_motif_list LIST
                        Ranked list of polyA motifs (text)
  --polyA_peak POLYA_PEAK
                        PolyA Peak (BED format)
  --phyloP_bed PHYLOP_BED
                        PhyloP BED for conservation scores

Functional annotation:
  --isoAnnotLite        Run isoAnnot Lite for tappAS compatibility
  --gff3 GFF3           Precomputed tappAS species specific GFF3 file

Output & Performance options:
  --isoform_hits        Report all FSM/ISM isoform hits in a separate file
  --ratio_TSS_metric {max,mean,median,3quartile}
                        Metric for ratio_TSS column (default: max)
  -t CPUS, --mapping_cpus CPUS
                        Number of threads used during alignment (default: 10)
  -n CHUNKS, --chunks CHUNKS
                        Number of chunks to split analysis (default: 10)
  -l {ERROR,WARNING,INFO,DEBUG}, --log_level LEVEL
                        Set the logging level (Default: INFO)

Optional arguments:
  --is_fusion           Input are fusion isoforms (GTF required)
  -v, --version         Display program version number

Running SQANTI-sc

The tool operates in two main modes: reads and isoforms.

1. Reads Mode (--mode reads)

The Reads Mode is an adaptation of the SQANTI-reads tool specifically designed for single-cell data. It provides a comprehensive structural and quality control assessment of long-read single-cell RNA-seq data at the read level. This step can be useful for validating your data structure and quality before committing to the more complex steps of isoform identification and quantification.

Key Applications:

  • Pre-analysis QC: Provides actionable insights into the quality and structural properties of the library, such as splicing accuracy and the presence of technical artifacts. This information allows users to make informed decisions about downstream isoform identification strategies.
  • Multisample Benchmarking: Adapts core SQANTI-reads metrics for single-cell resolution. The full suite of quality metrics enables robust comparisons of library complexity and technical variability across different experimental conditions, sequencing platforms, preprocessing pipelines, etc.

In this mode, SQANTI-sc treats each individual read as a "transcript" instance.

Compatible Pipelines

  • PacBio Iso-Seq: The Iso-Seq single cell pipeline produces deduplicated unaligned reads in BAM format.
  • Oxford Nanopore: The EPI2ME wf-single-cell pipeline producess undeduplicated aligned reads in BAM format. To use these reads as input for SQANTI-sc, first they must be deduplicated and converted to GTF/GFF. You can find a utility script named process_ont_bam.py that performs these steps for you automatically in the scripts/ folder.
  • Any long-read pipeline producing deduplicated unaligned reads in BAM or FASTA/FASTQ format; or deduplicated aligned reads in GTF/GFF format.

Reads Input Formats

  • uBAM: Unaligned reads. Note: SQANTI-sc will automatically convert BAM files to FASTQ for mapping and processing.
  • FASTA/FASTQ: Unaligned reads. If reads are provided in FASTA or FASTQ format, SQANTI-sc will automatically add the (-fasta) option to the SQANTI3 runner and perform a mapping step. Check the SQANTI3 documentation for mapping options.
  • GTF/GFF: Transcript annotations. We recommend providing reads in GTF/GFF format, as it enables better control over mapping parameters and faster runtimes.

Minimal Input (Mandatory)

SQANTI-sc requires a specific set of input files to run. The main entry point is a Design File (CSV) that maps each sample to its corresponding reads and cell barcode information.

  • Design File (--design) A comma-separated values (CSV) file containing the metadata for your samples.
Column Description
input_file Recommended. The exact absolute or relative path to the input reads file (BAM/FASTQ/FASTA/GTF) for this sample. If provided, the pipeline reads exactly this file and ignores --input_dir.
sampleID Required. A unique descriptive identifier for the sample (e.g., pb_brain). This will be used as the prefix for output files and the display name in reports.
file_acc Required. The identifier used to name the output directory where this sample's results will be stored entirely separate from others.
Fallback Logic: If you do not provide an input_file column, SQANTI-sc will use this as a prefix to search the --input_dir for input files (e.g., SRX123456 will match --input_dir/SRX123456*.bam).
cell_association Required. Path to the file linking reads to cell barcodes for each sample.
There are 2 options for this file:
1. A TSV file mapping Read IDs to Cell Barcodes. The minimum columns required are id (identifier of the read) and cb (cell barcode of the read).
2. A uBAM/BAM file containing CB (Cell Barcode) and XM/UB (UMI) tags. Easier to give if your input reads are already in uBAM format (the uBAM file can be the same as the one used as input reads).
coverage Optional. Path to STAR splice junction output(s). Used to validate splice junctions. Can be a single file, a directory, a comma-separated list, a wildcard pattern, or a File of File Names (FOFN).
SR_bam Optional. Path to short-read BAM file(s). Used to validate TSS. Can be a single BAM file or a File of File Names (FOFN) containing paths to multiple BAMs.
color_group Optional. Group label for PCA color encoding in multisample reports (--multisample_report).
shape_group Optional. Group label for PCA shape encoding in multisample reports (--multisample_report).
shade_group Optional. Group label for PCA shade/lightness encoding in multisample reports (--multisample_report).

Each group column can be used independently or in any combination. See Regenerate multisample reports from existing outputs for details.

Example design_reads.csv:

input_file,sampleID,file_acc,cell_association,coverage,SR_bam
/data/pb/SRX123456.bam,pb_brain,SRX123456,/data/pb/SRX123456_barcodes.tsv,/data/sr/brain.SJ.out.tab,/data/sr/brain_aligned.bam
/data/pb/SRX123457.gff,pb_lung,SRX123457,/data/pb/SRX123457.bam,,

Example cell_association file (barcodes.tsv):

id	cb	umi
m64012_250421_000242/120719489/ccs/10460_11196	AAACCCAAGGTTCCTA	TATGCCCGGTAT
m64012_250421_000242/17565024/ccs/13918_14203	GGGTTTAAGGTTCCTA	GCGCGCAATTCA

Note: The umi column is optional. If included, the UMIs will be added to the reads classification under the column UMI.

  • Reference Genome (--refFasta): FASTA format (e.g., hg38.fa). The Chromosome/scaffold names must exactly match those in the reference annotation.
  • Reference Annotation (--refGTF): GTF format (e.g., gencode.v38.annotation.gtf). Used to classify transcripts and assess novelty. Make sure that it matches the reference genome's coordinate system. You can find reference transcriptomes for different species in GENCODE or CHESS.

2. Isoforms Mode (--mode isoforms)

The Isoforms Mode is designed for the in-depth characterization of unique transcript isoforms across single cells. Unlike Reads Mode, which focuses on individual reads, this mode takes collapsed or assembled consensus transcript models as input and performs quality control at the single-cell level.

Key Applications:

  • Single-Cell Isoform Characterization: Analyzes isoform expression across cell populations using user-provided abundance data, enabling the exploration of alternative splicing patterns and isoform usage heterogeneity.
  • Structural Classification: Classifies each isoform against the reference annotation following the standard SQANTI3 classification scheme, providing a detailed quality assessment of the transcriptome assembly.

Compatible Inputs (WIP)

  • Iso-Seq Collapse: Output from PacBio's Iso-Seq Collapse. To generate the isoform count matrices from the outputs provided by this tool you can use the script named make_pacbio_matrix.py, located in the scripts/ folder.
  • Spl-IsoQuant: Output from spl-IsoQuant, IsoQuant version developed for upstream analysis of single-cell and spatial long-read data analysis. To generate the isoform count matrices from the outputs provided by this tool you can use the script named make_isoquant_matrix.py, located in the scripts/ folder.
  • Any other tool producing a GTF/GFF or FASTA representing specific transcript isoforms and the corresponding quantification of said isoforms in each cell barcode.

Isoforms Input Formats

  • GTF (*.gtf): Recommended. Transcript annotations. Like SQANTI3, this is the default and recommended format that SQANTI-sc expects.
  • FASTA/FASTQ (*.fasta, *.fastq): Transcript sequences. If provided, SQANTI-sc will map them to the reference genome (like in the Reads Mode).

Minimal Input (Mandatory)

Follows similar inputs to Reads mode.

  • Design File (--design) A comma-separated values (CSV) file containing the metadata for your samples.
Column Description
input_file Recommended. The exact absolute or relative path to the input transcript models file (GTF/FASTA) for this sample. If provided, the pipeline reads exactly this file and ignores --input_dir.
sampleID Required. A unique descriptive identifier for the sample (e.g., PacBio_Brain_5k). This will be used as the prefix for output files and the display name in validation plots/reports.
file_acc Required. The identifier used to name the output directory where this sample's results will be stored entirely separate from others.
Fallback Logic: If you do not provide an input_file column, SQANTI-sc will use this as a prefix to search the --input_dir for input files.
cell_association Conditional. Path to the file linking Isoform IDs to Cell Barcodes (TSV) if no abundance matrix is present.
abundance Conditional. Path to a folder containing quantification data in Market Exchange (MEX) format. The folder MUST contain three files: matrix.mtx, features.tsv (with the transcript model IDs used in the input), and barcodes.tsv.
coverage Optional. Path to STAR splice junction output(s). Used to validate splice junctions. Can be a single file, a directory, a comma-separated list, a wildcard pattern, or a File of File Names (FOFN).
SR_bam Optional. Path to short-read BAM file(s). Used to validate TSS. Can be a single BAM file or a File of File Names (FOFN) containing paths to multiple BAMs.
color_group Optional. Group label for PCA color encoding in multisample reports (--multisample_report).
shape_group Optional. Group label for PCA shape encoding in multisample reports (--multisample_report).
shade_group Optional. Group label for PCA shade/lightness encoding in multisample reports (--multisample_report).

Each group column can be used independently or in any combination. See Regenerate multisample reports from existing outputs for details.

Note: You can provide either a cell_association file or abundance directory as your cell barcode-isoform association file. If you want to perform quality control with the quantification of the expression of the isoforms (recommended) you will need the count matrix, but is not mandatory to run the isoforms mode.

Example design_isoforms.csv:

input_file,sampleID,file_acc,abundance,coverage,SR_bam
/data/SRX987654.gtf,Iso_Brain,SRX987654,/data/SRX987654/counts/,/path/to/sr1_SJ.out.tab,/path/to/sr1_sorted.bam
/data/SRX987655.gtf,Iso_Lung,SRX987655,/data/SRX987655/counts/,/path/to/sr2_SJ.out.tab,/path/to/sr2_sorted.bam
  • Reference Files: Same as Reads Mode (--refFasta, --refGTF).

SQANTI-sc usage Example

python sqanti_sc.py \
    --mode isoforms \
    --design design_isoforms.csv \
    --refFasta /genomes/human/hg38.fa \
    --refGTF /genomes/human/hg38.gtf \
    --input_dir ./isoforms \
    --out_dir ./results \
    --CAGE_peak ./SQANTI3/data/ref_TSS_annotation/human.refTSS_v4.1.hg38.bed \
    --polyA_motif_list ./SQANTI3/data/polyA_motifs/mouse_and_human.polyA_motif.txt \
    --report both \
    --run_clustering \
    --multisample_report

Providing orthogonal data to SQANTI-sc

SQANTI-sc accepts orthogonal data to assist in the quality control and filtering of artifactual transcript models.

  • CAGE peak data (--CAGE_peak) for Transcription Start Site (TSS) validation.
  • PolyA information (--polyA_motif_list, --polyA_peak) for Transcription Termination Site (TTS) validation.
  • Short Reads data (Splice Junctions and alignment BAMs) provided via the Design File.

Short Reads Validation

SQANTI-sc supports orthogonal validation of splice junctions and TSS using bulk-like short reads as a population-level proxy. Unlike SQANTI3, it does not accept raw FASTQ inputs (--short_reads). Pre-processed files must be provided via the Design File.

Two validation types are supported, each requiring a different input:

Column Validates Required input
coverage Splice junctions (min_cov) STAR SJ.out.tab file
SR_bam TSS ratio (ratio_TSS) STAR-aligned BAM file

Both inputs require bulk short-read RNA-seq data — from your own experiment or downloaded from public repositories such as SRA or GTEx. Align the reads with STAR using the run_STAR.py script in scripts/, which applies SQANTI3-compatible parameters.

For splice junction validation only, SQANTI-sc also provides scripts/make_recount3_sj.R, a convenience script that downloads pre-computed junction counts from recount3 (GTEx for human, SRA for mouse) and converts them directly to SJ.out.tab format — no raw data download or alignment required. TSS ratio validation (SR_bam) is not available through recount3 as it requires a BAM file, which recount3 does not provide.

To learn more about the underlying metrics, visit the SQANTI3 documentation.

Why are single-cell short reads (e.g. 10x Genomics) not supported?

SQANTI-sc requires full-transcript short-read coverage to validate internal splice junctions. Single-cell short-read assays (e.g. 10x Genomics 3') generate 3'-end-biased reads that do not cover the full transcript body and are therefore unsuitable for this validation step.

Design File Columns for Short Reads:

Column Description
coverage Optional. Path to a STAR splice junction file (SJ.out.tab). Used to validate splice junctions. Multiple files accepted via: directory, comma-separated list, wildcard pattern, or FOFN (one path per line).
SR_bam Optional. Path to a sorted and indexed bulk RNA-seq BAM file. Used to calculate ratio_TSS and validate 5' ends. Multiple BAMs accepted via FOFN (one path per line).

Downloading junction data with recount3

The script scripts/make_recount3_sj.R downloads pre-computed junction counts from recount3 and converts them to STAR SJ.out.tab format. Two complementary metrics are produced:

  • Read depth (--min_cov): column 7 of the SJ.out.tab holds the total read coverage of each junction summed across all samples, so --min_cov keeps its standard SQANTI3 meaning — minimum read coverage of an isoform's weakest junction.
  • Sample support: a companion <output>.sample_support.tsv sidecar is written alongside the SJ.out.tab. SQANTI-sc uses it to populate two fields that SQANTI3 defines but which are otherwise degenerate when a single pre-aggregated file is provided:
    • *_junctions.txt: sample_with_cov is overwritten with the real number of study samples in which each junction was observed; sample_pct (new) gives that as a percentage of the total samples in the study.
    • *_classification.txt: min_sample_cov and min_sample_pct reflect the least-reproducible junction of each isoform (its minimum across all its junctions), giving an isoform-level reproducibility score independent of raw read depth.

Chromosome names keep the chr prefix by default, matching GENCODE and UCSC-style references (chr1, chr2, chrX). Pass --strip_chr if your reference uses Ensembl-style names (1, 2, X).

Human (GTEx): 54 tissue types available, up to 1,000+ samples per tissue, uniformly processed. This is the recommended path for human data.

Mouse (SRA): Individual SRA studies from the recount3 catalogue (~10,000 projects). Requires identifying the right project ID using --list_tissues and cross-referencing with SRA or recount3. Many SRA projects cover multiple tissues or mixed RNA-seq protocols — use --list_metadata and --filter_col/--filter_val to restrict to the relevant samples before summing (see below).

# List available human GTEx tissues
Rscript scripts/make_recount3_sj.R --list_tissues

# List top mouse SRA projects by sample count
Rscript scripts/make_recount3_sj.R --list_tissues --organism mouse

# Export full mouse project list to a file
Rscript scripts/make_recount3_sj.R --list_tissues --organism mouse --n_projects 0 > projects.txt

# Download human GTEx junctions
Rscript scripts/make_recount3_sj.R \
    --organism human --tissue BLOOD \
    --output blood_gtex.SJ.out.tab

# Download mouse SRA junctions (all samples in project)
Rscript scripts/make_recount3_sj.R \
    --organism mouse --project SRP150473 \
    --output mouse_brain.SJ.out.tab

Filtering by tissue or condition (multi-tissue SRA projects)

Many mouse SRA projects pool multiple tissues. For example, Tabula Muris Senis bulk (SRP199494) covers 17 organs. To build a tissue-matched junction reference, first inspect the available sample metadata, then filter to the samples of interest:

# Step 1: discover which metadata columns and values exist for the project
Rscript scripts/make_recount3_sj.R \
    --organism mouse --project SRP199494 \
    --list_metadata
# prints e.g.:
#   sra.sample_attributes   [947 unique]  ...tissue;;Heart_40... | ...tissue;;Liver_54... | ...

# Step 2: download junctions filtered to a single tissue
Rscript scripts/make_recount3_sj.R \
    --organism mouse --project SRP199494 \
    --filter_col sra.sample_attributes --filter_val Heart \
    --output heart_tms.SJ.out.tab

Matching is case-insensitive substring, so partial values work (e.g. --filter_val heart matches Heart or Heart - Left Ventricle). When a filter is active, pct_samples in the sidecar is computed relative to the filtered cohort, so min_sample_pct correctly reflects reproducibility within the chosen tissue rather than across the whole study.

The --min_samples argument (default: 1) pre-filters junctions at download time. We recommend keeping the default so that the reference file does not need to be regenerated; read-depth stringency can be controlled at runtime via --min_cov, and sample-support stringency via the min_sample_cov / min_sample_pct columns in the downstream filter step. Keep the generated <output>.sample_support.tsv sidecar next to the SJ.out.tab file — SQANTI-sc looks for it there to populate the sample-support fields (without it, the run still works but those fields are not updated).

Once generated, provide the file in the coverage column of your design CSV:

input_file,sampleID,file_acc,abundance,coverage
/data/sample.gtf,PBMC_sample,SRX987654,/data/counts/,/path/to/blood_gtex.SJ.out.tab

Cell clustering (--run_clustering)

SQANTI-sc includes an optional step to perform cell clustering based on the expression of genes. This step is useful for exploring the heterogeneity of the cell population. SQANTI-sc uses the Scanpy toolkit for cell clustering.

To enable this step, you must use the flag --run_clustering.

In isoforms mode, clustering also fills two columns of the classification, max_cluster_cells and max_cluster_pct (see the classification columns). They are recomputed every time clustering runs, in the QC run and in the filters.

Customization options

The clustering process can be customized using the following arguments:

  • --normalization: Normalization method to apply to the count matrix before clustering. Options: log1p (default, log(CPM+1)), sqrt (square root), pearson (Pearson residuals).
  • --n_neighbors: Number of neighbors for UMAP construction (Default: 15).
  • --n_pc: Number of principal components to use for clustering (Default: 30).
  • --resolution: Resolution parameter for Leiden clustering (Default: 0.5). Higher values result in more clusters.
  • --n_top_genes: Number of highly variable genes to use for dimensionality reduction (Default: 2000).
  • --clustering_method: Algorithm to use for clustering. Options: leiden (default), louvain, kmeans.
  • --n_clusters: Number of clusters to force if using K-means clustering (Default: 10).

Cell filter (sqanti_sc_filter.py cells)

Once a QC run has finished, sqanti_sc_filter.py identifies low-quality cell barcodes from the per-cell metrics in *_SQANTI_cell_summary.txt.gz. It is a separate entry point so that thresholds can be retuned without re-running SQANTI3 QC, and it needs neither the genome FASTA nor the reference GTF:

python sqanti_sc_filter.py cells \
    --design design.csv \
    --qc_dir ./qc \
    --out_dir ./filter/cells

Every run writes both the verdict and the filtered data files. --qc_dir is read and never written, so the original QC run is untouched.

There is no --mode: Reads_in_cell and Transcripts_in_cell are mutually exclusive, so the cell summary states which mode it came from.

Nothing existing is ever modified. --qc_dir is read and never written; everything the filter produces goes under --out_dir.

Output layout

The filter's output has the same shape as a QC run — <out_dir>/<file_acc>/<sampleID>_* — so every existing tool works on it unchanged, with the same design file:

qc/rep1/s1_classification.txt                  qc/rep1/clustering/umap_results.csv
filter/cells/rep1/s1_classification.txt        filter/cells/rep1/clustering/umap_results.csv

Threshold trials are therefore just different directories (--out_dir filter/cells_depth1000), and they can be compared side by side. The transcript filter writes to filter/transcripts/ the same way.

Always written, mirroring the SQANTI3 filter's contract:

File Contents
<sampleID>_CellFilter_cell_summary.txt.gz Every judged barcode, all its metrics, plus exactly one added column: filter_result (Cell/Artifact). This is the cell-summary equivalent of SQANTI3's _RulesFilter_classification.txt, which likewise adds only its verdict column.
<sampleID>_pass_cells.txt Passing barcodes, one per line — usable directly as a Scanpy/Seurat subset
<sampleID>_cell_filtering_reasons.txt Discarded barcodes and why: CB and filter_reason. Artifacts only, as in SQANTI3
<sampleID>_cell_filter_params.txt Every threshold used, plus the run statistics
<sampleID>_classification.txt Filtered: only retained cells. In isoforms mode cells_detected is recounted over them, and max_cluster_cells/max_cluster_pct are NA unless --run_clustering refits the clusters
<sampleID>_junctions.txt Filtered, with the per-row barcode list rewritten
<sampleID>_corrected.gtf Filtered: only the surviving transcript models
<sampleID>_corrected.fasta Filtered: the sequences of those same models

If the QC run produced them, the optional per-model files are filtered the same way and keep their names: <sampleID>_corrected.faa and <sampleID>_corrected.cds.gff3 (from --include_ORF), <sampleID>_corrected.sam (when SQANTI3 aligned sequence input) and <sampleID>.gff3 (from --isoAnnotLite). Nothing has to be passed for this: each is looked for in the sample's QC folder.

The cell summary is labelled, never subset — the same choice SQANTI3 makes for the classification it judges. <sampleID>_CellFilter_cell_summary.txt.gz holds every judged barcode with the verdict on it, and there is no second copy containing only the survivors. The reports drop the Artifact rows when they see the verdict column, so they show the retained cells without a separate file existing.

The filtered data files are written alongside it under their usual names: *_classification.txt, *_junctions.txt, *_corrected.gtf and *_corrected.fasta. The GTF and the FASTA are subset to exactly the models that survive in the filtered classification, so all four stay consistent with one another. This mirrors SQANTI3, whose shipped filter configuration points filter_gtf and filter_isoforms at the corrected GTF and FASTA by default.

Note that filtering cells also removes transcript models that were observed only in discarded cells. The count is reported as TranscriptModelsLostAllSupport.

Rules and defaults

Criteria live in a JSON file (--rules, default src/filter_assets/cell_filter_default.json) with the same structure and numeric vocabulary as the SQANTI3 rules filter: a list of two numbers is a [min, max] range and a bare number is a minimum. SQANTI3's string forms (all_canonical: "canonical") have no counterpart here — every cell-summary column except CB is numeric, so a string rule would match nothing and silently discard every cell. It is rejected with an error instead.

{
  "all": [
    {
      "depth": 500,
      "Annotated_genes": 200,
      "MT_perc": [0, 10]
    }
  ]
}

depth resolves to Reads_in_cell in reads mode and Transcripts_in_cell in isoforms mode, so one file works for both.

Alternatives: the value is a list

As in SQANTI3, the value is a list of rule-sets. Rules inside one set are ANDed; the sets are ORed, so each set is an alternative way for a cell to be acceptable. A single rule-set is a one-element list, as above.

{
  "all": [
    { "depth": 500, "MT_perc": [0, 10] },
    { "depth": 200, "Annotated_genes": 200, "MT_perc": [0, 10] }
  ]
}

Keep a cell with 500+ observations, or a shallower one that still detects 200+ annotated genes. SQANTI3 uses the same construction to let one kind of evidence substitute for another — its default waives the canonical-junction requirement when short reads support the junction — and the reasoning carries over to cells.

A cell is discarded only when every rule-set rejects it, so no single rule "killed" it. Its reason therefore lists the failures from all the sets, which is why a discarded cell commonly names several rules.

These are defaults, not recommendations. The right values depend on your tissue, platform and library chemistry — edit the JSON for your data.

  • depth (min 500): single-cell long-read libraries typically run in the thousands of reads or transcripts per real cell, so this mainly separates cells from empty droplets. Note this is a different question from whether a proportion is estimable: below ~10 observations a percentage moves in steps of more than 10%, so a proportion rule on a very shallow cell is noise even when the cell is real.
  • Annotated_genes (min 200): chosen because the retention curve is flat either side of it. On data that still contains empty droplets the separation between empty and real barcodes is sharp, so anything from about 50 to 300 keeps the same cells; above roughly 500 the floor starts removing real ones. 200 sits in the middle of that plateau, which also means the exact value matters little. Annotated_genes rather than Genes_in_cell, because SQANTI3 mints a fresh novelGene_* id for nearly every unassignable read, so Genes_in_cell partly counts artifacts.
  • MT_perc (max 10): a ruptured cell loses cytoplasmic mRNA while membrane-bound mitochondria stay, so high MT does indicate damage. MT transcripts are short and long-read protocols size-select, so expect a lower fraction here than the familiar short-read thresholds, not a higher one. Some cell types (cardiomyocytes, neurons) genuinely run much higher — raise it for those.

Criteria deliberately left out of the defaults

Every column of the cell summary can be used as a rule. Three families are available but not defaulted, because a sensible threshold for them is not a constant:

  • Per-cell artifact loads — Intrapriming_prop_in_cell, RTS_prop_in_cell, Non_canonical_prop_in_cell. These are the criteria a general-purpose single-cell QC tool cannot compute, and they are the reason to use this filter — but note what they measure. SQANTI3's own rules ask "is this transcript an artifact?" and answer with a definition (intra-priming is perc_A_downstream_TTS > 59). These ask "what fraction of this cell may be artifacts?", which is a tolerance, and the right tolerance depends on the chemistry. 10x 3′ protocols in particular produce a high baseline intra-priming rate, so a tolerance chosen without looking at your data can sit below the baseline and discard nearly every cell. Set them from your own distribution — a robust outlier bound such as median + 3·MAD is a reasonable starting point.
  • Structural-category shares — FSM_prop, ISM_prop and their relatives. These vary more across cells than anything else in the summary, which makes them tempting, but a fixed threshold on one does not mean the same thing twice. Two effects swamp cell quality. The transcript reconstruction tool dominates: run the same reads through different tools and the median FSM share can move from roughly a third to nearly everything, because tools differ in how readily they emit a model that is not in the annotation. Cell type comes next: large, transcriptionally active cells carry a higher share of immature and partial transcripts than small quiescent ones, so in a mixed population these columns separate cell types before they separate good cells from bad. A plausible-looking FSM_prop minimum can therefore delete an entire cell type from one sample and nothing at all from another processed differently, in both cases leaving a report that looks clean. If you use these at all, compare each cell against others of its own cluster within a single processing run, never against a global constant.
  • Columns from optional SQANTI3 inputs — CAGE_peak_support_prop, PolyA_motif_support_prop, srjunctions_support_prop, TSS_ratio_validated_prop and the ORF-derived columns. These are NA unless the matching evidence was supplied to the QC run, and a rule on an all-NA column judges nothing (the filter warns when this happens).

Values that cannot be judged

A metric can be NA because it was never measured — an undefined proportion (zero denominator), or an attribute whose SQANTI3 input was not supplied (CAGE_peak_support_prop, PolyA_motif_support_prop, the ORF-derived columns). A comparison against NA is false, so judging a cell on one would discard it for an unknown value.

Instead, a rule whose value is NA for a given cell is skipped for that cell; the cell is still judged on every other rule, and the skipped rule contributes no reason. If a rule column is NA for every cell, the filter warns that the rule is judging nothing.

This is a deliberate difference from SQANTI3, which treats NA as a failure. The two NAs do not mean the same thing. SQANTI3's is usually a missing input — supply no short reads and min_cov is NA on every transcript — and its alternative rule-sets route around it. Ours is a fact about one cell: a cell with no multi-exonic reads has no denominator for Non_canonical_prop_in_cell, so failing it would judge the cell on its own composition rather than its quality.

Rows with no cell barcode (unassigned, NA, -, *, empty) are dropped before any statistic and the count is reported — an unassigned row aggregates every read that was never assigned to a cell, so it would otherwise dominate any depth-based statistic.

Transcript filter (sqanti_sc_filter.py transcripts)

sqanti_sc_filter.py transcripts runs a SQANTI3 filter on each sample's classification, the way the QC run calls SQANTI3 QC: the rules filter by default, or the machine-learning filter with --method ml. In reads mode each row is a read, so the filter judges reads; in isoforms mode it judges transcript models.

python sqanti_sc_filter.py transcripts \
    --design design.csv \
    --qc_dir ./filter/cells \
    --out_dir ./filter/transcripts

--qc_dir can be the QC run itself or the output of the cell filter. The intended order is cells first, then transcripts, but either works: both directories have the same layout, and the transcript filter accepts the cell filter's labelled summary as its input.

Rules

The rules are SQANTI3's own, in SQANTI3's JSON format, keyed by structural category. Without --rules SQANTI3's filter_default.json applies. See the SQANTI3 documentation for the rule syntax.

Mono-exonic models are judged by the rules like any other model; SQANTI3's -e, which discards all of them regardless of the rules, is not exposed. To require multiple exons, write it as a rule on the exons column ("exons": 2), which can also be limited to the structural categories where it should apply.

Rules on cell context

In isoforms mode the classification carries three cell-context columns — cells_detected, max_cluster_cells and max_cluster_pct (see the classification columns) — and a rule can use them like any other column. Reads-mode classifications do not have them, since each row there is a single read from a single cell, so SQANTI3 stops with "column not found" if a rule names them. This is how a model that is not seen again across the cells of one cluster can be discarded. src/filter_assets/transcript_filter_cluster_example.json is SQANTI3's default with "max_cluster_cells": 3 added to both rest rule-sets: a novel model must reach at least 3 cells of one cluster, while full-splice matches are judged as before. The 3 is an example, not a calibrated value. Without --rules nothing changes.

Put the rule in every alternative rule-set of a category: a rule-set without it accepts models regardless.

The two cluster columns are measured only when the input was clustered (--run_clustering in the QC run or the cell filter). Otherwise they are NA, and SQANTI3 fails a rule on NA, so a rule on them discards every model of that category with the reason NA value in max_cluster_cells — as min_cov does without short reads. A QC run made before these columns existed has none of them; run the cell filter with --run_clustering to add them.

Machine learning (--method ml)

--method ml runs SQANTI3's ML filter. A random forest learns from a list of true isoforms and a list of artifacts, then judges every multi-exon model; an intra-priming check on the 3' end runs alongside it. See the SQANTI3 documentation for the method.

python sqanti_sc_filter.py transcripts --method ml \
    --design design.csv \
    --qc_dir ./filter/cells \
    --out_dir ./filter/transcripts_ml

The options are SQANTI3's, under SQANTI3's flags, so -j is the probability threshold here rather than a rules file. -t, -p/-n, -f, -r, -z, -i and --intermediate_files mean what they mean in SQANTI3, and SQANTI3's own defaults apply to any left unset. transcripts --method ml -h lists them. -e is not exposed, as for the rules.

SQANTI3 reads a copy of the classification made for it, deleted once it finishes:

  • In isoforms mode FL holds one count per cell, so the copy carries the model's total, the way SQANTI3 adds up the FL columns of several samples. How many cells the model is in reaches the forest separately, as cells_detected, and on a clustered run as max_cluster_cells and max_cluster_pct.
  • The cell barcodes and UMIs, and in reads mode the junction-chain columns jxn_string and jxnHash, are left out. They describe cells, not models, and the forest would otherwise use their text as a number.

SQANTI3 builds the two training lists from each sample's own classification: reference-match FSM models as true isoforms, non-canonical NNC models as artifacts. It needs at least 250 models in each. With fewer it skips the classifier and still finishes normally, so only intra-primed models are removed; the report then shows intra-priming as the only reason. --TP and --TN supply the lists yourself, for a design with a single sample: model IDs are numbered per sample, so they are refused when the design has more than one.

Output

The output has the same <out_dir>/<file_acc>/<sampleID>_* shape as a QC run. Each filter labels the table whose rows it judges and cuts the other files to match: the cell filter labels the cell summary, and the transcript filter labels the classification.

File Contents
<sampleID>_RulesFilter_classification.txt Rules: every model, plus SQANTI3's filter_result column (Isoform/Artifact)
<sampleID>_ML_classification.txt ML: every model, plus SQANTI3's POS_MLprob, NEG_MLprob, ML_classifier, intra_priming and filter_result columns
<sampleID>_filtering_reasons.txt Rules: discarded models and why (SQANTI3's)
<sampleID>_TP_list.txt, <sampleID>_TN_list.txt, <sampleID>_randomforest.RData, testSet_*, classifier_variable-importance_table.txt ML: SQANTI3's training lists (when it built them), the classifier and its test-set results
<sampleID>_pass_isoforms.txt Passing model IDs, one per line (SQANTI3's)
<sampleID>_params.txt SQANTI3's record of the run
<sampleID>_junctions.txt Only the junctions of the passing models
<sampleID>_corrected.gtf Only the passing models
<sampleID>_corrected.fasta The sequences of those same models
<sampleID>_SQANTI_cell_summary.txt.gz Recomputed from the passing models

The classification is labelled, never subset, as in SQANTI3, and there is no second copy holding only the passing models. Every reader of the classification — the cell summary, clustering, both reports and the .h5ad export — skips the Artifact rows when it sees the verdict column. The rows are exactly as the QC run wrote them. SQANTI3's rules filter writes its own copy of this file back through pandas, which rewrites NA as an empty field, TRUE as True and -1 as -1.0, and its ML filter writes the reduced copy it read; the wrapper replaces either with the original rows plus SQANTI3's added columns. A rules run and an ML run cannot share an --out_dir, since the readers take whichever labelled classification they find.

The junctions, GTF and FASTA contain only the passing models, the way SQANTI3's own filtered GTF and FASTA do. They keep their QC names and their original lines. So do the optional per-model files when the QC run produced them, as for the cell filter: the proteins and their CDS coordinates, the SAM and the tappAS GFF3.

The cell summary is recomputed, not carried over. Removing models changes every cell's depth and composition, so the per-cell metrics are rebuilt from what survives. The summary thresholds --min_cov, --ratio_TSS and --ref_cov_min_pct default as in the QC run; pass the values the QC run used if you changed them. A cell whose every model was discarded is absent from the recomputed summary, and the count is reported.

A classification that already carries filter_result is refused as input: filtering it again could keep models whose junctions, GTF and FASTA entries were already removed. The cell filter does accept a transcript filter's output, and keeps the verdict column in its own classification.

Reports

--report, --refGTF, --multisample_report and the clustering options work as for the cell filter: clustering is refitted on the filtered data, and the per-sample and multisample reports are rendered from it. Refitting recomputes the cell-context columns of the passing models; the Artifact rows keep the values they were judged on. The reports therefore describe every cell as it is once its artifactual models are removed. After the ML filter the report shows why models were removed (machine learning, intra-priming or both), per cell and per structural category, and, when the classifier was trained, its test-set performance and the importance of each variable. SQANTI3's own filter report is not produced, just as the QC run does not produce SQANTI3's QC report.

Understanding the output of SQANTI-sc

The majority of the outputs follow the same logic and structure as SQANTI3 and SQANTI-reads, with modifications to include single-cell information and new files specific to this pipeline.

1. SQANTI3-based Outputs

Standard SQANTI3 output files are generated for each sample, but they include additional columns to track cell barcode of origin.

  • *_classification.txt: The main output file containing structural classification and quality attributes.
    • New columns:
      • CB: Cell Barcode associated with the read/isoform. In isoforms mode, this column is a comma-separated list of the cell barcodes in which the isoform appears.
      • UMI: Unique Molecular Identifier (Reads Mode only, if UMI information available (through cell association file)).
      • jxn_string / jxnHash (Reads Mode only): Unique Junction Chain (UJC) representation of the transcript model structure. These columns, introduced by SQANTI-reads, allow for read grouping based on splice junctions structure. This step can be skipped using the --skip_hash option.
      • cells_detected (Isoforms Mode only): Number of cells the model is detected in: the cells in CB with a count above 0, so the fractional counts of EM-based quantifiers (Isosceles, bambu) count as detection.
      • max_cluster_cells / max_cluster_pct (Isoforms Mode only): The most cells the model reaches in any one cluster, and the highest share of a cluster's cells it reaches. Each is maximised on its own, so the two can come from different clusters — a model confined to a small cluster can have few cells but a high share. Counted the same way as cells_detected. NA unless clustering was run; a model whose cells are all outside the UMAP gets 0.
    • Changed columns:
      • FL: Full-length count. In reads mode, this column will always diplay values of 1, as each row of the classification is suppossed to represent a unique UMI. In isoforms mode it will also display values of 1 except if quantification information is provided, in which case the column will be a comma-separated list of numeric values that represent the number of counts of that transcript model in each of the cells in the CB column, following the same order.
  • *_junctions.txt: File containing all splice junctions identified.
    • New columns:
      • CB: Cell Barcode associated with the junction. In isoforms mode, this column can contain multiple comma-separated cell barcodes if the junction pertains to a transcript model present in multiple cell barcodes.

2. SQANTI-reads-based Outputs

The majority of SQANTI-reads-specific otuputs are not output by SQANTI-sc, with the exception of some tables that add the cell barcode dimension. These per-cell matrices follow the logic of SQANTI-reads and are generated optionally if the --write_per_cell_outputs flag is used. They provide detailed metrics at the single-cell level.

  • *_gene_counts.csv: Counts of reads/isoforms per gene per cell, broken down by structural category.
  • *_ujc_counts.csv: Counts of Unique Junction Chains (UJCs) per cell, including their structural classification and novelty status.
  • *_cv.csv: Coefficient of Variation (CV) metrics for splice junctions per cell, useful for identifying splicing variability.

3. SQANTI-sc Specific Outputs

  • *_report.html / *.pdf: Comprehensive quality control reports summarizing the data at the single-cell level.
  • clustering/umap_results.csv: (If --run_clustering is active) Contains the UMAP coordinates and cluster assignments for each cell barcode.
  • *_SQANTI_cell_summary.txt.gz: A GZIP-compressed tab-delimited file containing a wide array of quality control metrics aggregated per cell. This is the core file for downstream analysis of cellular transcriptome quality.
  • *.h5ad: (If --run_clustering is active) An AnnData object containing the gene × cell count matrix (.X), isoform × cell count matrix (.obsm["isoform_counts"], isoforms mode only), per-cell QC metrics (.obs), UMAP coordinates (.obsm["X_umap"]), and cluster labels (.obs["cluster"]). This file is directly compatible with Scanpy (Python) and Seurat (R, via SeuratDisk).

Glossary of Cell Summary columns

The output _SQANTI_cell_summary.txt.gz has the following possible fields:

Proportion columns can be NA. Every *_prop, *_perc and *_support column is a percentage of some denominator — the cell's reads, its junctions, or its reads of one structural category. When that denominator is 0 the column is NA, because there is no quantity to take a percentage of: a cell with no fusion reads has no fusion RT-switching rate. This is distinct from a genuine 0, which means the denominator was positive and the numerator was 0 (for example 100 canonical and 0 non-canonical junctions is 0, not NA). Count columns are never NA — a cell with no fusion reads has a Fusion count of 0. Filter or aggregate accordingly (na.rm = TRUE in R, skipna=True is the pandas default); the denominator columns (Reads_in_cell / Transcripts_in_cell, total_junctions, total_reads_no_monoexon / total_transcripts_no_monoexon, and the per-category counts) are all present in this file if you need to reconstruct which case applies.

A whole column is NA when the run never measured that attribute. The short-read (srjunctions_support_prop, TSS_ratio_validated_prop), CAGE (CAGE_peak_support_prop), polyA-motif (PolyA_motif_support_prop) and ORF (NMD_prop_in_cell, *_coding_prop, *_non_coding_prop) families are only computed when the matching input was supplied — the design file's coverage and SR_bam columns, and --CAGE_peak, --polyA_motif_list, --include_ORF. The ORF family additionally needs --mode isoforms: --include_ORF is ignored in reads mode, where predicting an ORF for every read would be prohibitively slow, so those columns are NA there whether or not the flag was given. The columns are always written so the file's shape does not depend on the flags, following SQANTI3's own convention of emitting every field and writing NA for anything it could not compute. A read whose attribute SQANTI3 left NA — a mono-exonic read has no junction for short reads to cover, and NMD is only predicted for a read with junctions — is dropped from both numerator and denominator rather than counted as unsupported, so these percentages describe the reads the attribute was actually evaluated on.

  • CB : Cell Barcode identifier.
  • Reads_in_cell / Transcripts_in_cell : Total number of reads (Reads Mode) or transcripts (Isoforms Mode) associated with the cell.
  • UMIs_in_cell : Total number of unique Molecular Identifiers (UMIs) detected (Reads Mode only) in the cell.
  • total_reads_no_monoexon / total_transcripts_no_monoexon : Count of reads/transcripts excluding mono-exons.
  • FSM, ISM, NIC, NNC, Genic_Genomic, Antisense, Fusion, Intergenic, Genic_intron : Absolute counts of reads/transcripts belonging to each SQANTI3 structural category.
  • [Category]_prop (e.g., FSM_prop, NIC_prop) : The proportion of reads/transcripts in the cell belonging to each structural category.
  • Genes_in_cell : Number of unique genes detected in the cell.
  • UJCs_in_cell : Number of Unique Junction Chains (UJCs) detected (multi-exon only).
  • MT_reads_count : Number of reads/transcripts mapping to mitochondrial genes.
  • MT_perc : Percentage of reads/transcripts mapping to mitochondrial genes.
  • Annotated_genes : Number of known (annotated) genes detected.
  • Novel_genes : Number of novel genes detected.
  • *_junctions (e.g., Known_canonical_junctions, Novel_canonical_junctions, Novel_non_canonical_junctions, Known_non_canonical_junctions) : Counts of splice junctions by type.
  • total_junctions : Total number of splice junctions identified in the cell.
  • *_junctions_prop : Proportions of each junction type relative to total junctions.
  • [Category]_[Subcategory]_prop : Proportion of reads/transcripts within a category that belong to a specific subcategory (e.g., FSM_alternative_3end_prop, ISM_intron_retention_prop, Genic_mono_exon_prop).
  • anno_bin*_perc : Proportion of annotated genes with expression levels (reads or transcript model counts) in specific bins (bin1: 1 count, bin2_4: 2-4 counts, bin5_9: 5-9 counts, bin10plus: >=10 counts).
  • novel_bin*_perc : Proportion of novel genes with expression levels in specific bins (same as above).
  • *_ujc_bin*_perc : Similar bins calculated for Unique Junction Chains (UJCs) (Reads Mode only): bin1 (1 count), bin2_3 (2-3 counts), bin4_5 (4-5 counts), bin6plus (>=6 counts).
  • [Total|Category]_*_length_[mono]_prop : Proportion of reads falling into specific length bins (Total and per-category, including mono-exonic specific). Length bins: <250bp, 250-500bp, 500-1000bp (short), 1000-2000bp (mid), >2000bp (long).
  • *_ref_coverage_prop : Proportion of reads/transcripts in each category covering at least ref_cov_min_pct of the reference transcript length.
  • ref_cov_min_pct : The minimum reference coverage percentage used for the above metric (parameter value).
  • RTS_prop_in_cell : Proportion of reads flagged as Reverse Transcriptase Switching (RTS) artifacts.
  • [Category]_RTS_prop : Proportion of RTS artifacts within each structural category.
  • Non_canonical_prop_in_cell : Proportion of reads with non-canonical splicing.
  • [Category]_noncanon_prop : Proportion of non-canonical splicing within each structural category.
  • Intrapriming_prop_in_cell : Proportion of reads flagged as intra-priming artifacts.
  • [Category]_intrapriming_prop : Proportion of intra-priming artifacts within each structural category.
  • TSSAnnotationSupport_prop : Proportion of reads where the TSS is within 50bp of an annotated TSS.
  • [Category]_TSSAnnotationSupport : Proportion of TSS support within each structural category.
  • Annotated_genes_prop_in_cell : Proportion of genes detected that are annotated.
  • Annotated_juction_strings_prop_in_cell : Proportion of UJCs that match annotated junction combinations.
  • Canonical_prop_in_cell : Proportion of reads/transcripts with canonical splicing.
  • [Category]_canon_prop : Proportion of canonical splicing within each structural category.
  • NMD_prop_in_cell : Proportion of reads predicted to be NMD candidates (if ORF prediction is on).
  • [Category]_NMD_prop : Proportion of NMD candidates within each structural category.
  • srjunctions_support_prop : Proportion of reads/transcripts supported by short read junctions (min_cov >= threshold).
  • [Category]_srjunctions_support_prop : Proportion of short read junction support within each structural category.
  • TSS_ratio_validated_prop : Proportion of reads/transcripts with validated TSS (ratio_TSS >= threshold).
  • [Category]_TSS_ratio_validated_prop : Proportion of TSS validation within each structural category.
  • [Category]_[non]_coding_prop : Proportion of coding vs non-coding transcripts within each category.
  • CAGE_peak_support_prop : Proportion of reads supported by CAGE peaks (if provided).
  • [Category]_CAGE_peak_support_prop : Proportion of CAGE support within each structural category.
  • PolyA_motif_support_prop : Proportion of reads with identified PolyA motifs (if provided).
  • [Category]_PolyA_motif_support_prop : Proportion of PolyA support within each structural category.

Exporting to Scanpy / Seurat

SQANTI-sc provides a standalone utility script (scripts/export_scanpy_seurat.py) to convert pipeline outputs into an AnnData .h5ad file for downstream analysis with Scanpy or Seurat. This script can be run independently on any existing SQANTI-sc output without re-running the pipeline.

python scripts/export_scanpy_seurat.py \
    --mode isoforms \
    --classification ./results/sample1/sample1_classification.txt \
    --cell_summary ./results/sample1/sample1_SQANTI_cell_summary.txt.gz \
    --clustering ./results/sample1/clustering/umap_results.csv \
    -o ./export \
    -p sample1

The resulting .h5ad file contains:

Slot Content
.X Gene × cell count matrix (sparse)
.obsm["isoform_counts"] Isoform × cell count matrix (isoforms mode only)
.obs Per-cell QC metrics from the cell summary
.obs["cluster"] Cluster labels (if clustering results provided)
.obsm["X_umap"] UMAP coordinates (if clustering results provided)
.var Gene/feature metadata
.uns["isoform_features"] Isoform-to-gene mapping (isoforms mode only)

Regenerate multisample reports from existing outputs

SQANTI-sc provides a standalone utility script (scripts/generate_multisample_report.py) to create multisample reports from previously generated per-sample outputs. This is useful when you want to compare only a subset of samples or conditions without rerunning the full pipeline.

The script takes a tab-separated FOFN (File of File Names) listing one sample per line. Lines starting with # are ignored.

FOFN columns (tab-separated, positional):

Col Required Description
1 Yes Path to *_SQANTI_cell_summary.txt.gz
2 No Path to *_classification.txt (enables length distribution plots when ≥ 2 provided)
3 No color_group — PCA color/hue grouping
4 No shape_group — PCA shape grouping
5 No shade_group — PCA shade/lightness grouping

Example multisample_inputs.txt:

# cell_summary	classification	color_group	shape_group	shade_group
/results/SRX123/pb_brain_5k_SQANTI_cell_summary.txt.gz	/results/SRX123/pb_brain_5k_classification.txt	PacBio	brain	5k
/results/SRX456/pb_lung_500_SQANTI_cell_summary.txt.gz	/results/SRX456/pb_lung_500_classification.txt	PacBio	lung	500
/results/SRX789/ont_brain_5k_SQANTI_cell_summary.txt.gz	/results/SRX789/ont_brain_5k_classification.txt	ONT	brain	5k
python scripts/generate_multisample_report.py \
    --mode reads \
    --fofn multisample_inputs.txt \
    --report both \
    -o ./multisample_output \
    -p my_comparison

To compare a different subset of samples, simply edit the FOFN file.

Notes:

  • At least 2 valid *_SQANTI_cell_summary.txt.gz files are required.
  • The classification column (col 2) is optional, but providing at least 2 enables multisample length distribution plots.
  • Any combination of color_group, shape_group, and shade_group triggers a grouped PCA alongside the default sampleID PCA. Each column is fully independent:
    • color_group only → color encoding, fixed circle shape;
    • shape_group only → shape encoding, all points the same color;
    • shade_group only → shade/lightness encoding, all points the same hue;
    • any two or all three → combined encoding.
  • When running via sqanti_sc.py --multisample_report, the same grouped PCA behavior is available through the optional color_group, shape_group, and shade_group columns in the design CSV.
  • In the HTML multisample report, when grouped PCA is available, the PCA Plot tab includes a dropdown with By Sample (default) and By Group options.
  • The script reuses the same R multisample report generator used by --multisample_report in sqanti_sc.py.

Citation

If you are using SQANTI-single-cell, please cite:

About

Repository for development of new single cell SQANTI tool

Resources

Stars

9 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages