Integrating structure-based machine learning and co-evolution to investigate specificity in plant sesquiterpene synthases

Sesquiterpene synthases (STSs) catalyze the formation of a large class of plant volatiles called sesquiterpenes. While thousands of putative STS sequences from diverse plant species are available, only a small number of them have been functionally characterized. Sequence identity-based screening for desired enzymes, often used in biotechnological applications, is difficult to apply here as STS sequence similarity is strongly affected by species. This calls for more sophisticated computational methods for functionality prediction. We investigate the specificity of precursor cation formation in these elusive enzymes. By inspecting multi-product STSs, we demonstrate that STSs have a strong selectivity towards one precursor cation. We use a machine learning approach combining sequence and structure information to accurately predict precursor cation specificity for STSs across all plant species. We combine this with a co-evolutionary analysis on the wealth of uncharacterized putative STS sequences, to pinpoint residues and distant functional contacts influencing cation formation and reaction pathway selection. These structural factors can be used to predict and engineer enzymes with specific functions, as we demonstrate by predicting and characterizing two novel STSs from Citrus bergamia.

Download Full-text

Integrating structure-based machine learning and co-evolution to investigate specificity in plant sesquiterpene synthases

10.1101/2020.07.28.224527 ◽

2020 ◽

Author(s):

Janani Durairaj ◽

Elena Melillo ◽

Harro J Bouwmeester ◽

Jules Beekwilder ◽

Dick de Ridder ◽

...

Keyword(s):

Machine Learning ◽

Plant Species ◽

Sequence Data ◽

Reaction Pathway ◽

Sequence Similarity ◽

Homology Modelling ◽

Enzyme Engineering ◽

Evolutionary Analysis ◽

Structural Factors ◽

Terpene Synthases

AbstractSesquiterpene synthases (STSs) catalyze the formation of a large class of plant volatiles called sesquiterpenes. While thousands of putative STS sequences from diverse plant species are available, only a small number of them have been functionally characterized. Sequence identity-based screening for desired enzymes, often used in biotechnological applications, is difficult to apply here as STS sequence similarity is strongly affected by species. This calls for more sophisticated computational methods for functionality prediction. We investigate the specificity of precursor cation formation in these elusive enzymes. By inspecting multi-product STSs, we demonstrate that STSs have a strong selectivity towards one precursor cation. We use a machine learning approach combining sequence and structure information to accurately predict precursor cation specificity for STSs across all plant species. We combine this with a co-evolutionary analysis on the wealth of uncharacterized putative STS sequences, to pinpoint residues and distant functional contacts influencing cation formation and reaction pathway selection. These structural factors can be used to predict and engineer enzymes with specific functions, as we demonstrate by predicting and characterizing two novel STSs from Citrus bergamia.Author summaryPredicting enzyme function is a popular problem in the bioinformatics field that grows more pressing with the increase in protein sequences, and more attainable with the increase in experimentally characterized enzymes. Terpenes and terpenoids form the largest classes of natural products and find use in many drugs, flavouring agents, and perfumes. Terpene synthases catalyze the biosynthesis of terpenes via multiple cyclizations and carbocation rearrangements, generating a vast array of product skeletons. In this work, we present a three-pronged computational approach to predict carbocation specificity in sesquiterpene synthases, a subset of terpene synthases with one of the highest diversities of products. Using homology modelling, machine learning and co-evolutionary analysis, our approach combines sparse structural data, large amounts of uncharacterized sequence data, and the current set of experimentally characterized enzymes to provide insight into residues and structural regions that likely play a role in determining product specifcity. Similar techniques can be repurposed for function prediction and enzyme engineering in many other classes of enzymes.

Download Full-text

Recent progresses in the application of machine learning approach for predicting protein functional class independent of sequence similarity

PROTEOMICS ◽

10.1002/pmic.200500938 ◽

2006 ◽

Vol 6 (14) ◽

pp. 4023-4037 ◽

Cited By ~ 46

Author(s):

Lianyi Han ◽

Juan Cui ◽

Honghuang Lin ◽

Zhiliang Ji ◽

Zhiwei Cao ◽

...

Keyword(s):

Machine Learning ◽

Sequence Similarity ◽

Functional Class ◽

Learning Approach ◽

Machine Learning Approach

Download Full-text

Constructing and Validating Geographically Refined HAZUS-MH4 Hurricane Wind Risk Models: A Machine Learning Approach

Advances in Hurricane Engineering ◽

10.1061/9780784412626.092 ◽

2012 ◽

Cited By ~ 2

Author(s):

D. Subramanian ◽

J. Salazar ◽

L. Duenas-Osorio ◽

R. Stein

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Risk Models ◽

Hurricane Wind ◽

Machine Learning Approach

Download Full-text

The impact of economic plans on the Chinese education system: a machine learning approach

CADMO ◽

10.3280/cad2018-001005 ◽

2018 ◽

pp. 37-49

Author(s):

Wenjun Lin ◽

Xuefu Xu ◽

Francesco Dell’Anna

Keyword(s):

Machine Learning ◽

Education System ◽

Learning Approach ◽

Chinese Education ◽

System A ◽

Machine Learning Approach ◽

The Impact

Download Full-text

Improving Bandwidth Utilization and Fairness between TCP Flows based on a Machine-learning Approach

IEEJ Transactions on Electronics Information and Systems ◽

10.1541/ieejeiss.133.1259 ◽

2013 ◽

Vol 133 (6) ◽

pp. 1259-1268

Author(s):

Akihiro Shiozu ◽

Syunji Yazaki ◽

K^|^ocirc;ki Abe

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Bandwidth Utilization ◽

Machine Learning Approach

Download Full-text

A Machine Learning Approach to Anaphora Resolution in Arabic

International Review on Computers and Software (IRECOS) ◽

10.15866/irecos.v9i12.4786 ◽

2014 ◽

Vol 9 (12) ◽

pp. 1956

Author(s):

Abdullatif Abolohom ◽

Nazlia Omar

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Anaphora Resolution ◽

Machine Learning Approach

Download Full-text

1552-P: Machine Learning Approach to Decision-Making for Initial Insulin Use in Japanese Patients with Type 2 Diabetes

Diabetes ◽

10.2337/db20-1552-p ◽

2020 ◽

Vol 69 (Supplement 1) ◽

pp. 1552-P

Author(s):

KAZUYA FUJIHARA ◽

MAYUKO H. YAMADA ◽

YASUHIRO MATSUBAYASHI ◽

MASAHIKO YAMAMOTO ◽

TOSHIHIRO IIZUKA ◽

...

Keyword(s):

Machine Learning ◽

Type 2 Diabetes ◽

Decision Making ◽

Japanese Patients ◽

Learning Approach ◽

Machine Learning Approach ◽

Insulin Use

Download Full-text

A Brief Survey on Text Classification Using Various Machine Learning Techniques

International Journal of Advanced Research in Computer Science and Software Engineering ◽

10.23956/ijarcsse.v8i1.521 ◽

2018 ◽

Vol 8 (1) ◽

pp. 14

Author(s):

Padmavathi .S ◽

M. Chidambaram

Keyword(s):

Machine Learning ◽

Text Classification ◽

Fixed Number ◽

Machine Learning Techniques ◽

Online Information ◽

Rule Based ◽

Learning Techniques ◽

Machine Learning Approach ◽

Rule Based Approach

Text classification has grown into more significant in managing and organizing the text data due to tremendous growth of online information. It does classification of documents in to fixed number of predefined categories. Rule based approach and Machine learning approach are the two ways of text classification. In rule based approach, classification of documents is done based on manually defined rules. In Machine learning based approach, classification rules or classifier are defined automatically using example documents. It has higher recall and quick process. This paper shows an investigation on text classification utilizing different machine learning techniques.

Download Full-text