Goals and approaches for each processing step for single-cell RNA sequencing data

Briefings in Bioinformatics ◽

10.1093/bib/bbaa314 ◽

2020 ◽

Author(s):

Zilong Zhang ◽

Feifei Cui ◽

Chunyu Wang ◽

Lingling Zhao ◽

Quan Zou

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Cellular Level ◽

Sequencing Data ◽

Analysis Tools ◽

Processing Step ◽

Study Gene Expression ◽

Single Cell Rna Sequencing ◽

Cell Data

Abstract Single-cell RNA sequencing (scRNA-seq) has enabled researchers to study gene expression at the cellular level. However, due to the extremely low levels of transcripts in a single cell and technical losses during reverse transcription, gene expression at a single-cell resolution is usually noisy and highly dimensional; thus, statistical analyses of single-cell data are a challenge. Although many scRNA-seq data analysis tools are currently available, a gold standard pipeline is not available for all datasets. Therefore, a general understanding of bioinformatics and associated computational issues would facilitate the selection of appropriate tools for a given set of data. In this review, we provide an overview of the goals and most popular computational analysis tools for the quality control, normalization, imputation, feature selection and dimension reduction of scRNA-seq data.

Download Full-text

G2S3: a gene graph-based imputation method for single-cell RNA sequencing data

10.1101/2020.04.01.020586 ◽

2020 ◽

Author(s):

Weimiao Wu ◽

Qile Dai ◽

Yunqing Liu ◽

Xiting Yan ◽

Zuoheng Wang

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Expression Profiles ◽

Gene Expression Profiles ◽

Sequencing Data ◽

High Data ◽

Study Gene Expression ◽

Single Cell Rna Sequencing ◽

Novel Method

AbstractSingle-cell RNA sequencing provides an opportunity to study gene expression at single-cell resolution. However, prevalent dropout events result in high data sparsity and noise that may obscure downstream analyses. We propose a novel method, G2S3, that imputes dropouts by borrowing information from adjacent genes in a sparse gene graph learned from gene expression profiles across cells. We applied G2S3 and other existing methods to seven single-cell datasets to compare their performance. Our results demonstrated that G2S3 is superior in recovering true expression levels, identifying cell subtypes, improving differential expression analyses, and recovering gene regulatory relationships, especially for mildly expressed genes.

Download Full-text

G2S3: A gene graph-based imputation method for single-cell RNA sequencing data

PLoS Computational Biology ◽

10.1371/journal.pcbi.1009029 ◽

2021 ◽

Vol 17 (5) ◽

pp. e1009029

Author(s):

Weimiao Wu ◽

Yunqing Liu ◽

Qile Dai ◽

Xiting Yan ◽

Zuoheng Wang

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Large Scale ◽

Expression Profiles ◽

Gene Expression Profiles ◽

Sequencing Data ◽

High Data ◽

Study Gene Expression ◽

Single Cell Rna Sequencing

Single-cell RNA sequencing technology provides an opportunity to study gene expression at single-cell resolution. However, prevalent dropout events result in high data sparsity and noise that may obscure downstream analyses in single-cell transcriptomic studies. We propose a new method, G2S3, that imputes dropouts by borrowing information from adjacent genes in a sparse gene graph learned from gene expression profiles across cells. We applied G2S3 and ten existing imputation methods to eight single-cell transcriptomic datasets and compared their performance. Our results demonstrated that G2S3 has superior overall performance in recovering gene expression, identifying cell subtypes, reconstructing cell trajectories, identifying differentially expressed genes, and recovering gene regulatory and correlation relationships. Moreover, G2S3 is computationally efficient for imputation in large-scale single-cell transcriptomic datasets.

Download Full-text

Single-Cell Transcriptome Analysis Reveals Dynamic Cell Populations and Differential Gene Expression Patterns in Control and Aneurysmal Human Aortic Tissue

Circulation ◽

10.1161/circulationaha.120.046528 ◽

2020 ◽

Vol 142 (14) ◽

pp. 1374-1388

Author(s):

Yanming Li ◽

Pingping Ren ◽

Ashley Dawson ◽

Hernan G. Vasquez ◽

Waleed Ageedi ◽

...

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Aortic Wall ◽

Genome Wide Association ◽

Aortic Tissue ◽

Sequencing Data ◽

Genome Wide ◽

Single Cell Rna Sequencing ◽

Differential Gene

Background: Ascending thoracic aortic aneurysm (ATAA) is caused by the progressive weakening and dilatation of the aortic wall and can lead to aortic dissection, rupture, and other life-threatening complications. To improve our understanding of ATAA pathogenesis, we aimed to comprehensively characterize the cellular composition of the ascending aortic wall and to identify molecular alterations in each cell population of human ATAA tissues. Methods: We performed single-cell RNA sequencing analysis of ascending aortic tissues from 11 study participants, including 8 patients with ATAA (4 women and 4 men) and 3 control subjects (2 women and 1 man). Cells extracted from aortic tissue were analyzed and categorized with single-cell RNA sequencing data to perform cluster identification. ATAA-related changes were then examined by comparing the proportions of each cell type and the gene expression profiles between ATAA and control tissues. We also examined which genes may be critical for ATAA by performing the integrative analysis of our single-cell RNA sequencing data with publicly available data from genome-wide association studies. Results: We identified 11 major cell types in human ascending aortic tissue; the high-resolution reclustering of these cells further divided them into 40 subtypes. Multiple subtypes were observed for smooth muscle cells, macrophages, and T lymphocytes, suggesting that these cells have multiple functional populations in the aortic wall. In general, ATAA tissues had fewer nonimmune cells and more immune cells, especially T lymphocytes, than control tissues did. Differential gene expression data suggested the presence of extensive mitochondrial dysfunction in ATAA tissues. In addition, integrative analysis of our single-cell RNA sequencing data with public genome-wide association study data and promoter capture Hi-C data suggested that the erythroblast transformation-specific related gene( ERG ) exerts an important role in maintaining normal aortic wall function. Conclusions: Our study provides a comprehensive evaluation of the cellular composition of the ascending aortic wall and reveals how the gene expression landscape is altered in human ATAA tissue. The information from this study makes important contributions to our understanding of ATAA formation and progression.

Download Full-text

Differential gene expression analysis in single-cell RNA sequencing data

2017 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) ◽

10.1109/bibm.2017.8217650 ◽

2017 ◽

Author(s):

Tianyu Wang ◽

Sheida Nabavi

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Differential Gene Expression ◽

Expression Analysis ◽

Gene Expression Analysis ◽

Sequencing Data ◽

Differential Gene Expression Analysis ◽

Single Cell Rna Sequencing ◽

Differential Gene

Download Full-text

Inferring the kinetics of stochastic gene expression from single-cell RNA-sequencing data

Genome Biology ◽

10.1186/gb-2013-14-1-r7 ◽

2013 ◽

Vol 14 (1) ◽

pp. R7 ◽

Cited By ~ 100

Author(s):

Jong Kim ◽

John C Marioni

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Stochastic Gene Expression ◽

Sequencing Data ◽

Single Cell Rna Sequencing ◽

Kinetics Of

Download Full-text

Abstract 4689: Subclone-specific evolution of tumor phenotypes – A framework to study subclone-specific gene expression from a combination of bulk DNA and single cell RNA sequencing data

10.1158/1538-7445.sabcs18-4689 ◽

2019 ◽

Author(s):

Yi Qiao ◽

Xiaomeng Huang ◽

Samuel Brady ◽

Andrea Bild ◽

David Bowtell ◽

...

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Specific Gene ◽

Sequencing Data ◽

Specific Gene Expression ◽

Single Cell Rna Sequencing ◽

Tumor Phenotypes

Download Full-text

DoubletFinder: Doublet detection in single-cell RNA sequencing data using artificial nearest neighbors

10.1101/352484 ◽

2018 ◽

Cited By ~ 17

Author(s):

Christopher S. McGinnis ◽

Lyndsay M. Murrow ◽

Zev J. Gartner

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

False Negative ◽

Droplet Microfluidics ◽

Cell Capture ◽

Sequencing Data ◽

Putative Gene ◽

Detection Tool ◽

Single Cell Rna Sequencing

SUMMARYSingle-cell RNA sequencing (scRNA-seq) using droplet microfluidics occasionally produces transcriptome data representing more than one cell. These technical artifacts are caused by cell doublets formed during cell capture and occur at a frequency proportional to the total number of sequenced cells. The presence of doublets can lead to spurious biological conclusions, which justifies the practice of sequencing fewer cells to limit doublet formation rates. Here, we present a computational doublet detection tool – DoubletFinder – that identifies doublets based solely on gene expression features. DoubletFinder infers the putative gene expression profile of real doublets by generating artificial doublets from existing scRNA-seq data. Neighborhood detection in gene expression space then identifies sequenced cells with increased probability of being doublets based on their proximity to artificial doublets. DoubletFinder robustly identifies doublets across scRNA-seq datasets with variable numbers of cells and sequencing depth, and predicts false-negative and false-positive doublets defined using conventional barcoding approaches. We anticipate that DoubletFinder will aid in scRNA-seq data analysis and will increase the throughput and accuracy of scRNA-seq experiments.

Download Full-text

SPsimSeq: semi-parametric simulation of bulk and single cell RNA sequencing data

10.1101/677740 ◽

2019 ◽

Cited By ~ 1

Author(s):

Alemu Takele Assefa ◽

Jo Vandesompele ◽

Olivier Thas

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Empirical Distribution ◽

Supplementary Information ◽

Rna Seq ◽

Sequencing Data ◽

Actual Distribution ◽

Wide Range ◽

Single Cell Rna Sequencing

SummarySPsimSeq is a semi-parametric simulation method for bulk and single cell RNA sequencing data. It simulates data from a good estimate of the actual distribution of a given real RNA-seq dataset. In contrast to existing approaches that assume a particular data distribution, our method constructs an empirical distribution of gene expression data from a given source RNA-seq experiment to faithfully capture the data characteristics of real data. Importantly, our method can be used to simulate a wide range of scenarios, such as single or multiple biological groups, systematic variations (e.g. confounding batch effects), and different sample sizes. It can also be used to simulate different gene expression units resulting from different library preparation protocols, such as read counts or UMI counts.Availability and implementationThe R package and associated documentation is available from https://github.com/CenterForStatistics-UGent/SPsimSeq.Supplementary informationSupplementary data are available at bioRχiv online.

Download Full-text

TSEE: an elastic embedding method to visualize the dynamic gene expression patterns of time series single-cell RNA sequencing data

BMC Genomics ◽

10.1186/s12864-019-5477-8 ◽

2019 ◽

Vol 20 (S2) ◽

Cited By ~ 5

Author(s):

Shaokun An ◽

Liang Ma ◽

Lin Wan

Keyword(s):

Gene Expression ◽

Time Series ◽

Single Cell ◽

Rna Sequencing ◽

Expression Patterns ◽

Gene Expression Patterns ◽

Sequencing Data ◽

Embedding Method ◽

Single Cell Rna Sequencing

Download Full-text

SigEMD: A powerful method for differential gene expression analysis in single-cell RNA sequencing data

Methods ◽

10.1016/j.ymeth.2018.04.017 ◽

2018 ◽

Vol 145 ◽

pp. 25-32 ◽

Cited By ~ 5

Author(s):

Tianyu Wang ◽

Sheida Nabavi

Keyword(s):

Gene Expression ◽

Single Cell ◽

Rna Sequencing ◽

Expression Analysis ◽

Gene Expression Analysis ◽

Powerful Method ◽

Sequencing Data ◽

Differential Gene Expression Analysis ◽

Single Cell Rna Sequencing ◽

Differential Gene

Download Full-text