Large-scale pathway-specific polygenic risk, transcriptomic community networks and functional inferences in Parkinson disease

ABSTRACTPolygenic inheritance plays a central role in Parkinson disease (PD). A priority in elucidating PD etiology lies in defining the biological basis of genetic risk. Unraveling how risk leads to disruption will yield disease-modifying therapeutic targets that may be effective. Here, we utilized a high-throughput and hypothesis-free approach to determine biological pathways underlying PD using the largest currently available cohorts of genetic data and gene expression data from International Parkinson’s Disease Genetics Consortium (IPDGC) and the Accelerating Medicines Partnership - Parkinson’s disease initiative (AMP-PD), among other sources. We placed these insights into a cellular context. We applied large-scale pathway-specific polygenic risk score (PRS) analyses to assess the role of common variation on PD risk in a cohort of 457,110 individuals by focusing on a compilation of 2,199 publicly annotated gene sets representative of curated pathways, of which we nominate 46 pathways associated with PD risk. We assessed the impact of rare variation on PD risk in an independent cohort of whole-genome sequencing data, including 4,331 individuals. We explored enrichment linked to expression cell specificity patterns using single-cell gene expression data and demonstrated a significant risk pattern for adult dopaminergic neurons, serotonergic neurons, and radial glia. Subsequently, we created a novel way of building de novo pathways by constructing a network expression community map using transcriptomic data derived from the blood of 1,612 PD patients, which revealed 54 connecting networks associated with PD. Our analyses highlight several promising pathways and genes for functional prioritization and provide a cellular context in which such work should be done.

Download Full-text

Graph Convolutional Network for Drug Response Prediction Using Gene Expression Data

Mathematics ◽

10.3390/math9070772 ◽

2021 ◽

Vol 9 (7) ◽

pp. 772

Author(s):

Seonghun Kim ◽

Seockhun Bae ◽

Yinhua Piao ◽

Kyuri Jo

Keyword(s):

Gene Expression ◽

Gene Expression Data ◽

Large Scale ◽

Drug Response ◽

Response Prediction ◽

Biological Data ◽

Expression Data ◽

Convolutional Network ◽

Essential Information ◽

Protein Protein Interaction

Genomic profiles of cancer patients such as gene expression have become a major source to predict responses to drugs in the era of personalized medicine. As large-scale drug screening data with cancer cell lines are available, a number of computational methods have been developed for drug response prediction. However, few methods incorporate both gene expression data and the biological network, which can harbor essential information about the underlying process of the drug response. We proposed an analysis framework called DrugGCN for prediction of Drug response using a Graph Convolutional Network (GCN). DrugGCN first generates a gene graph by combining a Protein-Protein Interaction (PPI) network and gene expression data with feature selection of drug-related genes, and the GCN model detects the local features such as subnetworks of genes that contribute to the drug response by localized filtering. We demonstrated the effectiveness of DrugGCN using biological data showing its high prediction accuracy among the competing methods.

Download Full-text

GENE DISCOVERY METHODS FROM LARGE-SCALE GENE EXPRESSION DATA

Quantum Bio-Informatics III ◽

10.1142/9789814304061_0040 ◽

2010 ◽

Author(s):

AKIFUMI SHIMIZU ◽

KENTARO YANO

Keyword(s):

Gene Expression ◽

Gene Expression Data ◽

Large Scale ◽

Gene Discovery ◽

Expression Data

Download Full-text

LSTrAP-Crowd: Prediction of novel components of bacterial ribosomes with crowd-sourced analysis of RNA sequencing data

10.1101/2020.04.20.005249 ◽

2020 ◽

Author(s):

Benedict Hew ◽

Qiao Wen Tan ◽

William Goh ◽

Jonathan Wei Xiong Ng ◽

Kenny Koh ◽

...

Keyword(s):

Gene Expression ◽

Protein Synthesis ◽

Rna Sequencing ◽

Gene Expression Data ◽

Large Scale ◽

Bacterial Resistance ◽

Expression Data ◽

Sequencing Data ◽

Novel Proteins ◽

Novel Antibiotics

AbstractBacterial resistance to antibiotics is a growing problem that is projected to cause more deaths than cancer in 2050. Consequently, novel antibiotics are urgently needed. Since more than half of the available antibiotics target the bacterial ribosomes, proteins that are involved in protein synthesis are thus prime targets for the development of novel antibiotics. However, experimental identification of these potential antibiotic target proteins can be labor-intensive and challenging, as these proteins are likely to be poorly characterized and specific to few bacteria. In order to identify these novel proteins, we established a Large-Scale Transcriptomic Analysis Pipeline in Crowd (LSTrAP-Crowd), where 285 individuals processed 26 terabytes of RNA-sequencing data of the 17 most notorious bacterial pathogens. In total, the crowd processed 26,269 RNA-seq experiments and used the data to construct gene co-expression networks, which were used to identify more than a hundred uncharacterized genes that were transcriptionally associated with protein synthesis. We provide the identity of these genes together with the processed gene expression data. The data can be used to identify other vulnerabilities or bacteria, while our approach demonstrates how the processing of gene expression data can be easily crowdsourced.

Download Full-text

Defining transcription modules using large-scale gene expression data

Bioinformatics ◽

10.1093/bioinformatics/bth166 ◽

2004 ◽

Vol 20 (13) ◽

pp. 1993-2003 ◽

Cited By ~ 216

Author(s):

J. Ihmels ◽

S. Bergmann ◽

N. Barkai

Keyword(s):

Gene Expression ◽

Gene Expression Data ◽

Large Scale ◽

Expression Data

Download Full-text

Large-Scale Integration of MicroRNA and Gene Expression Data for Identification of Enriched MicroRNA–mRNA Associations in Biological Systems

Methods in Molecular Biology - MicroRNAs and the Immune System ◽

10.1007/978-1-60761-811-9_20 ◽

2010 ◽

pp. 297-315 ◽

Cited By ~ 28

Author(s):

Preethi H. Gunaratne ◽

Chad J. Creighton ◽

Michael Watson ◽

Jayantha B. Tennakoon

Keyword(s):

Gene Expression ◽

Gene Expression Data ◽

Large Scale ◽

Biological Systems ◽

Expression Data ◽

Large Scale Integration ◽

Scale Integration

Download Full-text

Penetrance of Parkinson's Disease in LRRK2 p.G2019S Carriers Is Modified by a Polygenic Risk Score

Movement Disorders ◽

10.1002/mds.27974 ◽

2020 ◽

Vol 35 (5) ◽

pp. 774-780 ◽

Cited By ~ 7

Author(s):

Hirotaka Iwaki ◽

Cornelis Blauwendraat ◽

Mary B. Makarious ◽

Sara Bandrés‐Ciga ◽

Hampton L. Leonard ◽

...

Keyword(s):

Parkinson’S Disease ◽

Parkinson's Disease ◽

Risk Score ◽

Polygenic Risk Score ◽

Polygenic Risk

Download Full-text

Processing Large-Scale, High-Dimension Genetic and Gene Expression Data

Handbook on Analyzing Human Genetic Data ◽

10.1007/978-3-540-69264-5_11 ◽

2009 ◽

pp. 307-330

Author(s):

Cliona Molony ◽

Solveig K. Sieberts ◽

Eric E. Schadt

Keyword(s):

Gene Expression ◽

Gene Expression Data ◽

High Dimension ◽

Large Scale ◽

Expression Data

Download Full-text

Mining the Gene Expression Matrix: Inferring Gene Relationships from Large Scale Gene Expression Data

Information Processing in Cells and Tissues ◽

10.1007/978-1-4615-5345-8_22 ◽

1998 ◽

pp. 203-212 ◽

Cited By ~ 35

Author(s):

Patrik D’haeseleer ◽

Xiling Wen ◽

Stefanie Fuhrman ◽

Roland Somogyi

Keyword(s):

Gene Expression ◽

Gene Expression Data ◽

Large Scale ◽

Expression Data ◽

Gene Expression Matrix ◽

Expression Matrix

Download Full-text

3145 An Evaluation of Machine Learning and Traditional Statistical Methods for Discovery in Large-Scale Translational Data

Journal of Clinical and Translational Science ◽

10.1017/cts.2019.8 ◽

2019 ◽

Vol 3 (s1) ◽

pp. 2-2

Author(s):

Megan C Hollister ◽

Jeffrey D. Blume

Keyword(s):

Gene Expression ◽

Machine Learning ◽

Random Forest ◽

Gene Expression Data ◽

Large Scale ◽

Second Generation ◽

A Priori ◽

Expression Data ◽

P Values ◽

Machine Learning Methods

OBJECTIVES/SPECIFIC AIMS: To examine and compare the claims in Bzdok, Altman, and Brzywinski under a broader set of conditions by using unbiased methods of comparison. To explore how to accurately use various machine learning and traditional statistical methods in large-scale translational research by estimating their accuracy statistics. Then we will identify the methods with the best performance characteristics. METHODS/STUDY POPULATION: We conducted a simulation study with a microarray of gene expression data. We maintained the original structure proposed by Bzdok, Altman, and Brzywinski. The structure for gene expression data includes a total of 40 genes from 20 people, in which 10 people are phenotype positive and 10 are phenotype negative. In order to find a statistical difference 25% of the genes were set to be dysregulated across phenotype. This dysregulation forced the positive and negative phenotypes to have different mean population expressions. Additional variance was included to simulate genetic variation across the population. We also allowed for within person correlation across genes, which was not done in the original simulations. The following methods were used to determine the number of dysregulated genes in simulated data set: unadjusted p-values, Benjamini-Hochberg adjusted p-values, Bonferroni adjusted p-values, random forest importance levels, neural net prediction weights, and second-generation p-values. RESULTS/ANTICIPATED RESULTS: Results vary depending on whether a pre-specified significance level is used or the top 10 ranked values are taken. When all methods are given the same prior information of 10 dysregulated genes, the Benjamini-Hochberg adjusted p-values and the second-generation p-values generally outperform all other methods. We were not able to reproduce or validate the finding that random forest importance levels via a machine learning algorithm outperform classical methods. Almost uniformly, the machine learning methods did not yield improved accuracy statistics and they depend heavily on the a priori chosen number of dysregulated genes. DISCUSSION/SIGNIFICANCE OF IMPACT: In this context, machine learning methods do not outperform standard methods. Because of this and their additional complexity, machine learning approaches would not be preferable. Of all the approaches the second-generation p-value appears to offer significant benefit for the cost of a priori defining a region of trivially null effect sizes. The choice of an analysis method for large-scale translational data is critical to the success of any statistical investigation, and our simulations clearly highlight the various tradeoffs among the available methods.

Download Full-text

Bagging Statistical Network Inference from Large-Scale Gene Expression Data

PLoS ONE ◽

10.1371/journal.pone.0033624 ◽

2012 ◽

Vol 7 (3) ◽

pp. e33624 ◽

Cited By ~ 73

Author(s):

Ricardo de Matos Simoes ◽

Frank Emmert-Streib

Keyword(s):

Gene Expression ◽

Gene Expression Data ◽

Large Scale ◽

Network Inference ◽

Expression Data

Download Full-text