A general index for linear and nonlinear correlations for high dimensional genomic data

Abstract Background With the advance of high throughput sequencing, high-dimensional data are generated. Detecting dependence/correlation between these datasets is becoming one of most important issues in multi-dimensional data integration and co-expression network construction. RNA-sequencing data is widely used to construct gene regulatory networks. Such networks could be more accurate when methylation data, copy number aberration data and other types of data are introduced. Consequently, a general index for detecting relationships between high-dimensional data is indispensable. Results We proposed a Kernel-Based RV-coefficient, named KBRV, for testing both linear and nonlinear correlation between two matrices by introducing kernel functions into RV2 (the modified RV-coefficient). Permutation test and other validation methods were used on simulated data to test the significance and rationality of KBRV. In order to demonstrate the advantages of KBRV in constructing gene regulatory networks, we applied this index on real datasets (ovarian cancer datasets and exon-level RNA-Seq data in human myeloid differentiation) to illustrate its superiority over vector correlation. Conclusions We concluded that KBRV is an efficient index for detecting both linear and nonlinear relationships in high dimensional data. The correlation method for high dimensional data has possible applications in the construction of gene regulatory network.

Download Full-text

Multistability and Multicellularity: Cell Fates as High-Dimensional Attractors of Gene Regulatory Networks

Computational Systems Biology ◽

10.1016/b978-012088786-6/50033-2 ◽

2006 ◽

pp. 293-326 ◽

Cited By ~ 1

Author(s):

Sui Huang

Keyword(s):

Gene Regulatory Networks ◽

Regulatory Networks ◽

High Dimensional ◽

Cell Fates ◽

Gene Regulatory

Download Full-text

Independence screening for high dimensional nonlinear additive ODE models with applications to dynamic gene regulatory networks

Statistics in Medicine ◽

10.1002/sim.7669 ◽

2018 ◽

Vol 37 (17) ◽

pp. 2630-2644

Author(s):

Hongqi Xue ◽

Shuang Wu ◽

Yichao Wu ◽

Juan C. Ramirez Idarraga ◽

Hulin Wu

Keyword(s):

Gene Regulatory Networks ◽

Regulatory Networks ◽

High Dimensional ◽

Ode Models ◽

Gene Regulatory

Download Full-text

Learning Gene Regulatory Networks with High-Dimensional Heterogeneous Data

New Frontiers of Biostatistics and Bioinformatics - ICSA Book Series in Statistics ◽

10.1007/978-3-319-99389-8_15 ◽

2018 ◽

pp. 305-327

Author(s):

Bochao Jia ◽

Faming Liang

Keyword(s):

Gene Regulatory Networks ◽

Regulatory Networks ◽

Heterogeneous Data ◽

High Dimensional ◽

Gene Regulatory

Download Full-text

Interplay between Path and Speed in Decision Making by High-Dimensional Stochastic Gene Regulatory Networks

PLoS ONE ◽

10.1371/journal.pone.0040085 ◽

2012 ◽

Vol 7 (7) ◽

pp. e40085 ◽

Cited By ~ 8

Author(s):

Nuno R. Nené ◽

Alexey Zaikin

Keyword(s):

Decision Making ◽

Gene Regulatory Networks ◽

Regulatory Networks ◽

High Dimensional ◽

Gene Regulatory

Download Full-text

High-dimensional Bayesian network inference from systems genetics data using genetic node ordering

10.1101/501460 ◽

2018 ◽

Author(s):

Lingfei Wang ◽

Pieter Audenaert ◽

Tom Michoel

Keyword(s):

Genetic Variation ◽

Bayesian Network ◽

Gene Regulatory Networks ◽

Gene Networks ◽

Regulatory Networks ◽

Network Inference ◽

High Dimensional ◽

Systems Genetics ◽

Gene Regulatory ◽

Bayesian Network Inference

AbstractStudying the impact of genetic variation on gene regulatory networks is essential to understand the biological mechanisms by which genetic variation causes variation in phenotypes. Bayesian networks provide an elegant statistical approach for multi-trait genetic mapping and modelling causal trait relationships. However, inferring Bayesian gene networks from high-dimensional genetics and genomics data is challenging, because the number of possible networks scales super-exponentially with the number of nodes, and the computational cost of conventional Bayesian network inference methods quickly becomes prohibitive. We propose an alternative method to infer high-quality Bayesian gene networks that easily scales to thousands of genes. Our method first reconstructs a node ordering by conducting pairwise causal inference tests between genes, which then allows to infer a Bayesian network via a series of independent variable selection problems, one for each gene. We demonstrate using simulated and real systems genetics data that this results in a Bayesian network with equal, and sometimes better, likelihood than the conventional methods, while having a significantly higher over-lap with groundtruth networks and being orders of magnitude faster. Moreover our method allows for a unified false discovery rate control across genes and individual edges, and thus a rigorous and easily interpretable way for tuning the sparsity level of the inferred network. Bayesian network inference using pairwise node ordering is a highly efficient approach for reconstructing gene regulatory networks when prior information for the inclusion of edges exists or can be inferred from the available data.

Download Full-text