scRNA-seq data sets contains systematic and random noise that obscure the biological signal.
1 Analysis Workflow
1.1 Generate count matrix
Cellranger analysis: * Trim adapters * Alignment * Barcode and UMI extraction * Deduplication UMI * Filter barcodes * Filter reads * Generate count matrix
1.1.1 Gene expression matrix
- Filtered Matrix: Contains only estimated cell-barcodes after removing empty droplets and background noise
- Unfiltered (Raw) Matrix: Contains every barcode from the known-good list that has at least one read, including background and empty droplets.
├── sample_filtered_feature_bc_matrix
│ ├── barcodes.tsv.gz
│ ├── features.tsv.gz
│ └── matrix.mtx.gz
├── sample_filtered_feature_bc_matrix.h5MEX Format (.mtx.gz, barcodes.tsv.gz, features.tsv.gz): Stored in folders like filtered_feature_bc_matrix/ or raw_feature_bc_matrix/. It includes a sparse matrix file (matrix.mtx or .mtx.gz), a barcode file (barcodes.tsv.gz), and a feature/gene annotation file (features.tsv.gz). - the rows are features (genes) - and the columns are barcodes (cells)
HDF5 Format (.h5): A single binary file (e.g., filtered_feature_bc_matrix.h5) containing the same matrix data and feature annotations, ideal for fast loading and smaller file management.
The H5 format is a single, compressed binary file based on the Hierarchical Data Format version 5 (HDF5). In single-cell genomics, Cell Ranger uses it to package the entire feature-barcode matrix into one self-contained file
1.2 Quality control
In single-cell analysis, quality control are essential for assessing the integrity and reliability of the data. These metrics help identify low-quality cells, potential artifacts, and technical variations that may affect downstream analyses.
Low quality cells can arise from various sources, including: - Doublets or multiplets - Dying cells - Empty droplets - Ambient RNA contamination
1.2.1 Doublets or multiplets
Two or more cells encapsulated within the exact same gel bead or droplet, receiving the same barcode.
- Extreme Library Size: A massive spike in total UMI counts, often doubling the average cell’s profile.
- High Gene Diversity: A significantly higher number of uniquely detected genes.
- Biological Inconsistency: Outliers when plotting UMI counts vs. gene counts. They frequently express mutually exclusive marker genes (e.g., a single barcode showing high expression for both T-cell markers and Epithelial cell markers).
1.2.2 Dying or stressed cells
Cells that began to lyse, break apart, or undergo apoptosis prior to or during encapsulation.
- Mitochondrial Dominance: High percentage of mitochondrial gene reads (typically >5% in mice, >10-20% in humans, depending on tissue).
- Low Nuclear RNA: Low total UMI and gene counts because the cytoplasmic mRNA has leaked through the compromised cell membrane and washed away.
- Metabolic Shifts: Enrichment for apoptotic pathways or heat-shock protein genes
1.2.3 Empty droplets
Droplets that contain no actual cell but capture background ambient RNA shed from broken cells in the cell suspension.
Floor-Level UMIs: Total UMI counts are very low (e.g., fewer than 100 to 500 UMIs), failing to reach the “knee” or “inflection point” on a barcode rank plot.
Homogeneous Profile: The expression profile across different empty droplets looks nearly identical because they all sample the same ambient liquid background
Low and Uniform Expression: Ambient RNA is present in virtually every droplet, including empty droplets and those containing healthy cells. It shows up as a low-level, pervasive background count across the entire dataset, usually contributing a few dozen to a few hundred UMIs per barcode.
Ambient RNA causes distinct cell types to falsely appear as though they express markers of an entirely different cell lineage.
- Example: If red blood cells lyse during processing, Hemoglobin RNA (Hbb-bs) spills into the ambient pool. As a result, every T-cell, B-cell, and macrophage droplet will show low-level expressions of Hemoglobin, even though these cells do not biologically produce it.
1.2.4 Ambient RNA contamination
cell-free background RNA that contaminates single-cell droplets. It originates from damaged or lysed cells that rupture during tissue dissociation, spilling their contents into the cell suspension before encapsulation.
Following are the commonly used quality metrics in single-cell RNA sequencing (scRNA-seq) analysis: - Detected features per cell - Total counts per cell - Mitochondrial gene expression percentage - Ribosomal gene expression percentage
1.2.5 Quality metrics
1.2.5.1 Total counts
1.2.5.2 Detected features
1.2.5.3 % Mitochondrial genes
1.2.5.4 % Ribosomal genes
1.2.6 Filtering Low-quality cells
- Dying cells with broken membranes
- Cells with a low number of detected genes
- Cells with a low count depth
- Cells with high mitochondrial content
- Contamination from ambient RNA (cell-free RNA in the solution)
- Problems: Ambient RNA contamination can lead to cell-type-specific marker gene granscripts being detectable in different cell populations together.
- Solutions:
- SoupX
- CellBender
- Uses the raw (unfiltered) and filtered matrices to identify empty droplets.
- Estimates the ambient mRNA expression profile from those empty droplets.
- Estimates (automatically or manually, often guided by known non-expressed genes or cluster markers) the contamination fraction ρ (fraction of UMIs in each cell coming from the ambient soup).
- Subtracts the estimated ambient contribution proportionally from each cell’s counts, producing a corrected count matrix that can plug directly into Seurat/Scanpy pipelines.
- Uses a deep generative model (variational autoencoder / Bayesian inference in Pyro) on the raw count matrix.
- Jointly learns the background noise profile across all droplets, distinguishes cell-containing from empty droplets (often recovering more cells than Cell Ranger), and produces denoised quantifications by subtracting the estimated ambient contribution.
- Explicitly models aspects of background noise and can handle both RNA and other modalities (e.g., CITE-seq).
- Computationally heavier (benefits from GPU) and tends to perform strongly in benchmarks for ambient removal while preserving endogenous signal
- Empty droplet and doublet/multiplet cells
- Problems: Doublets formed by different cell types are hard to annotate and can lead to wrong cell-type labels.
- Solutions:
- scDbFinder
- DoubletFinder
- Scrublet
- Low-quality cells are identified and filtered by manually setting thresholds
- Quality control is performed at the sample level as thresholds can vary substantially between samples
1.2.7 Normalization and Variance Stabilization
- Count normalization makes cellular profiles comparable.
- Subsequent variance stablization ensures that outlier profiles have limited effect on the overall data structure
- Scaling and centering of the data is performed to ensure that all genes contribute equally to downstream analyses
- Scran
- Normalization methods should be chosen carefully and based on the subsequent analysis tasks
1.2.8 Removing confounding sources of variation
- Batch effects
- Harmony: projects cells into a shared embedding where cells group by cell type rather than dataset-specific conditions
- scANVI:
- scVI
- scGen
- Scanorama
- Cell cycle effects
- built-in cell cycle labeling and correction functions in Scanpy or Seurat
- Tricycle
1.2.9 Reducing dimensionality
- PCA
- t-SNE
- UMAP
- PHATE
1.2.10 Cell Clustering
- Louvain
- Leiden
- Both methods are applied to the KNN graph computed on a low-dimensional representation of the data and can be run at different resolutions to control the number of identified clusters.
1.2.11 Cell Type Annotation
Annotation is the process of giving detected cell clusters a biological interpretation such as cell type.
A three-step approach is recommended that leverages automated annotation, followed by expert manual annotation and a last step of verification to obtain the ideal annotation result
- Classifier-based:
- CellTypist
- Clustifyr
- SingleR
- scmap
- scANVI
- scNym
- Reference mapping
- Azimuth
- Symphony
- scArches
- Manual annotation
- Marker gene expression
- Differential expression analysis (t-test, wilcoxon rank-sum test)
- Annotation results obtained with pre-trained classifiers are strongly affected by the classifier type and the quality of the training data used to create the classifier.
- Furthermore, it can be difficult to assess the resulting annotation without additionally inspecting individual markers
- The annotation should be verified by experts, especially for data sets with high complexity or studies that involve rare cell subpopulations for which references might not be available
1.2.12 Trajectory Inference
Trajectory Inference or pseudotime analysis orders cells along a trajectory based on similarities in their gene expression patterns.
- Slingshot
- PAGA
- RaceID/StemID
- Monocle
- CellRank
- scVelo
- Velocyto
1.2.13 Differential Expression Analysis
Single-cell specific methods were found to systematically underestimate the variance of gene expression and to be prone to wrongly labling highly expressed genes as differentially expressed.
Pueudo-bulk, sample-level view aggregates counts per sample-label * edgeR * DESeq2 * limma
- P values obtained with DGE tests over conditions must be corrected for multiple testing5,92 to obtain q values.
1.2.14 Gene Set Enrichment Analysis
Gene set enrichment analysis allows the summarization of many molecular insights into interpretable terms such as pathways
definied gene sets known to be involved through previous studies.
- Common database:
- MSigDB
- Gene Ontology
- KEGG
- Reactome
- Common methods for enrichement include:
- hypergeometric test
- GSEA
- GSVA
1.2.15 Cell Composition analysis
Compositional analysis addresses conditional changes not in the gene expression profile of a cell but instead in the relative abundance of different cell types in the form of compositional data.
- Poisson regression
- Wilson rank-sum test
- scDC
- scCODA
- tascCODA
- DA-seq
- MILO
1.2.16 Cell-cell communication analysis
Cells are in constant interaction with each other for organismal development and homeostasis. If this interaction is impaired, disease ensues. Cell–cell communication inference methods commonly use repositories of ligands, receptors and their interactions to predict interactions between annotated clusters.
- CellChat
- CellPhoneDB
- SingleCellSignalR
- LIANA
- Nichenet
- Cytotalk