Predicting disease-causing variant combinations

Notwithstanding important advances in the context of single-variant pathogenicity identification, novel breakthroughs in discerning the origins of many rare diseases require methods able to identify more complex genetic models. We present here the Variant Combinations Pathogenicity Predictor (VarCoPP), a machine-learning approach that identifies pathogenic variant combinations in gene pairs (called digenic or bilocus variant combinations). We show that the results produced by this method are highly accurate and precise, an efficacy that is endorsed when validating the method on recently published independent disease-causing data. Confidence labels of 95% and 99% are identified, representing the probability of a bilocus combination being a true pathogenic result, providing geneticists with rational markers to evaluate the most relevant pathogenic combinations and limit the search space and time. Finally, the VarCoPP has been designed to act as an interpretable method that can provide explanations on why a bilocus combination is predicted as pathogenic and which biological information is important for that prediction. This work provides an important step toward the genetic understanding of rare diseases, paving the way to clinical knowledge and improved patient care.

Download Full-text

Predicting disease-causing variant combinations

10.1101/520353 ◽

2019 ◽

Cited By ~ 1

Author(s):

Sofia Papadimitriou ◽

Andrea Gazzo ◽

Nassim Versbraegen ◽

Charlotte Nachtegael ◽

Jan Aerts ◽

...

Keyword(s):

Machine Learning ◽

Patient Care ◽

Rare Diseases ◽

Search Space ◽

Pathogenic Variant ◽

Genetic Models ◽

Biological Information ◽

Learning Approach ◽

Clinical Knowledge ◽

Machine Learning Approach

ABSTRACTNotwithstanding important advances in the context of single-variant pathogenicity identification, novel breakthroughs in discerning the origins of many rare diseases require methods able to identify more complex genetic models. We present here the Variant Combinations Pathogenicity Predictor (VarCoPP), a machine-learning approach that identifies pathogenic variant combinations in gene pairs (bi-locus variant combinations). We show that the results produced by this method are highly accurate and precise, an efficacy that is endorsed when validating the method on recently published independent disease-causing data. Confidence labels of 95% and 99% are identified, representing the probability of a bi-locus combination being a true pathogenic result, providing geneticists with rational markers to evaluate the most relevant pathogenic combinations and limit the search space and time. Finally, VarCoPP has been designed to act as an interpretable method that can provide explanations on why a bi-locus combination is predicted as pathogenic and which biological information is important for that prediction. This work provides an important new step towards the genetic understanding of rare diseases, paving the way to new clinical knowledge and improved patient care.

Download Full-text

Identifying Rare Diseases from Behavioural Data: A Machine Learning Approach

2016 IEEE First International Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE) ◽

10.1109/chase.2016.7 ◽

2016 ◽

Cited By ~ 5

Author(s):

Haley MacLeod ◽

Shuo Yang ◽

Kim Oakes ◽

Kay Connelly ◽

Sriraam Natarajan

Keyword(s):

Machine Learning ◽

Rare Diseases ◽

Learning Approach ◽

Machine Learning Approach

Download Full-text

Constructing and Validating Geographically Refined HAZUS-MH4 Hurricane Wind Risk Models: A Machine Learning Approach

Advances in Hurricane Engineering ◽

10.1061/9780784412626.092 ◽

2012 ◽

Cited By ~ 2

Author(s):

D. Subramanian ◽

J. Salazar ◽

L. Duenas-Osorio ◽

R. Stein

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Risk Models ◽

Hurricane Wind ◽

Machine Learning Approach

Download Full-text

The impact of economic plans on the Chinese education system: a machine learning approach

CADMO ◽

10.3280/cad2018-001005 ◽

2018 ◽

pp. 37-49

Author(s):

Wenjun Lin ◽

Xuefu Xu ◽

Francesco Dell’Anna

Keyword(s):

Machine Learning ◽

Education System ◽

Learning Approach ◽

Chinese Education ◽

System A ◽

Machine Learning Approach ◽

The Impact

Download Full-text

Improving Bandwidth Utilization and Fairness between TCP Flows based on a Machine-learning Approach

IEEJ Transactions on Electronics Information and Systems ◽

10.1541/ieejeiss.133.1259 ◽

2013 ◽

Vol 133 (6) ◽

pp. 1259-1268

Author(s):

Akihiro Shiozu ◽

Syunji Yazaki ◽

K^|^ocirc;ki Abe

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Bandwidth Utilization ◽

Machine Learning Approach

Download Full-text

A Machine Learning Approach to Anaphora Resolution in Arabic

International Review on Computers and Software (IRECOS) ◽

10.15866/irecos.v9i12.4786 ◽

2014 ◽

Vol 9 (12) ◽

pp. 1956

Author(s):

Abdullatif Abolohom ◽

Nazlia Omar

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Anaphora Resolution ◽

Machine Learning Approach

Download Full-text

1552-P: Machine Learning Approach to Decision-Making for Initial Insulin Use in Japanese Patients with Type 2 Diabetes

Diabetes ◽

10.2337/db20-1552-p ◽

2020 ◽

Vol 69 (Supplement 1) ◽

pp. 1552-P

Author(s):

KAZUYA FUJIHARA ◽

MAYUKO H. YAMADA ◽

YASUHIRO MATSUBAYASHI ◽

MASAHIKO YAMAMOTO ◽

TOSHIHIRO IIZUKA ◽

...

Keyword(s):

Machine Learning ◽

Type 2 Diabetes ◽

Decision Making ◽

Japanese Patients ◽

Learning Approach ◽

Machine Learning Approach ◽

Insulin Use

Download Full-text

A Machine Learning Approach to Jet-Surface Interaction Noise Modeling

AIAA Scitech 2020 Forum ◽

10.2514/6.2020-1728 ◽

2020 ◽

Author(s):

Clifford A. Brown ◽

Jonny Dowdall ◽

Brian Whiteaker ◽

Lauren McIntyre

Keyword(s):

Machine Learning ◽

Surface Interaction ◽

Learning Approach ◽

Noise Modeling ◽

Machine Learning Approach

Download Full-text

A machine learning approach to evaluate IgE and IgG4 responses as patient-specific measure of exposure to sublingual allergen immunotherapy

10.26226/morressier.5afda3c8d64f25002cfc4098 ◽

2018 ◽

Author(s):

Thomas Stranzl

Keyword(s):

Machine Learning ◽

Allergen Immunotherapy ◽

Patient Specific ◽

Learning Approach ◽

Specific Measure ◽

Machine Learning Approach

Download Full-text

Mol2vec: Unsupervised Machine Learning Approach with Chemical Intuition

10.26434/chemrxiv.5513581.v1 ◽

2017 ◽

Author(s):

Sabrina Jaeger ◽

Simone Fulle ◽

Samo Turk

Keyword(s):

Machine Learning ◽

Language Processing ◽

Supervised Machine Learning ◽

Learning Approach ◽

Learning Approaches ◽

Unsupervised Machine Learning ◽

Feature Representations ◽

Machine Learning Approach ◽

The Individual ◽

Vector Representations

Inspired by natural language processing techniques we here introduce Mol2vec which is an unsupervised machine learning approach to learn vector representations of molecular substructures. Similarly, to the Word2vec models where vectors of closely related words are in close proximity in the vector space, Mol2vec learns vector representations of molecular substructures that are pointing in similar directions for chemically related substructures. Compounds can finally be encoded as vectors by summing up vectors of the individual substructures and, for instance, feed into supervised machine learning approaches to predict compound properties. The underlying substructure vector embeddings are obtained by training an unsupervised machine learning approach on a so-called corpus of compounds that consists of all available chemical matter. The resulting Mol2vec model is pre-trained once, yields dense vector representations and overcomes drawbacks of common compound feature representations such as sparseness and bit collisions. The prediction capabilities are demonstrated on several compound property and bioactivity data sets and compared with results obtained for Morgan fingerprints as reference compound representation. Mol2vec can be easily combined with ProtVec, which employs the same Word2vec concept on protein sequences, resulting in a proteochemometric approach that is alignment independent and can be thus also easily used for proteins with low sequence similarities.

Download Full-text