LINGO-DL: a text-based approach for molecular similarity searching

Molecular Similarity Searching with Different Similarity Coefficients and Different Molecular Descriptors

Recent Trends in Information and Communication Technology - Lecture Notes on Data Engineering and Communications Technologies ◽

10.1007/978-3-319-59427-9_5 ◽

2017 ◽

pp. 39-47

Author(s):

Fouaz Berrhail ◽

Hacene Belhadef ◽

Hamza Hentabli ◽

Faisal Saeed

Keyword(s):

Molecular Descriptors ◽

Molecular Similarity ◽

Similarity Searching ◽

Similarity Coefficients

Download Full-text

Molecular Similarity Searching Based on Deep Belief Networks with Different Molecular Descriptors

Proceedings of the 2020 2nd International Conference on Big Data Engineering and Technology ◽

10.1145/3378904.3378920 ◽

2020 ◽

Author(s):

Maged Nasser ◽

Naomie Salim ◽

Hentabli Hamza

Keyword(s):

Molecular Descriptors ◽

Molecular Similarity ◽

Similarity Searching ◽

Belief Networks ◽

Deep Belief Networks

Download Full-text

Molecular Similarity Searching Using Atom Environments, Information-Based Feature Selection, and a Naïve Bayesian Classifier

Journal of Chemical Information and Computer Sciences ◽

10.1021/ci034207y ◽

2004 ◽

Vol 44 (1) ◽

pp. 170-178 ◽

Cited By ~ 192

Author(s):

Andreas Bender ◽

Hamse Y. Mussa ◽

Robert C. Glen ◽

Stephan Reiling

Keyword(s):

Feature Selection ◽

Molecular Similarity ◽

Bayesian Classifier ◽

Similarity Searching ◽

Naïve Bayesian Classifier ◽

Naive Bayesian ◽

Naive Bayesian Classifier ◽

Naïve Bayesian

Download Full-text

Molecular Similarity Searching Method Based on Adaptive IR Technique

American Chemical Science Journal ◽

10.9734/acsj/2014/10175 ◽

2014 ◽

Vol 4 (6) ◽

pp. 787-797

Author(s):

Mohammed Binwahlan

Keyword(s):

Molecular Similarity ◽

Similarity Searching ◽

Searching Method

Download Full-text

Molecular Similarity Searching Using COSMO Screening Charges (COSMO/3PP)

Lecture Notes in Computer Science - Computational Life Sciences ◽

10.1007/11560500_16 ◽

2005 ◽

pp. 175-185 ◽

Cited By ~ 1

Author(s):

Andreas Bender ◽

Andreas Klamt ◽

Karin Wichmann ◽

Michael Thormann ◽

Robert C. Glen

Keyword(s):

Molecular Similarity ◽

Similarity Searching

Download Full-text

Improved Deep Learning Based Method for Molecular Similarity Searching Using Stack of Deep Belief Networks

Molecules ◽

10.3390/molecules26010128 ◽

2020 ◽

Vol 26 (1) ◽

pp. 128

Author(s):

Maged Nasser ◽

Naomie Salim ◽

Hentabli Hamza ◽

Faisal Saeed ◽

Idris Rabiu

Keyword(s):

Deep Learning ◽

Molecular Similarity ◽

Similarity Searching ◽

Reference Structure ◽

Belief Networks ◽

Deep Belief Networks ◽

Molecular Features ◽

Discovery Research ◽

Drug Discovery Research ◽

Similarity Method

Virtual screening (VS) is a computational practice applied in drug discovery research. VS is popularly applied in a computer-based search for new lead molecules based on molecular similarity searching. In chemical databases similarity searching is used to identify molecules that have similarities to a user-defined reference structure and is evaluated by quantitative measures of intermolecular structural similarity. Among existing approaches, 2D fingerprints are widely used. The similarity of a reference structure and a database structure is measured by the computation of association coefficients. In most classical similarity approaches, it is assumed that the molecular features in both biological and non-biologically-related activity carry the same weight. However, based on the chemical structure, it has been found that some distinguishable features are more important than others. Hence, this difference should be taken consideration by placing more weight on each important fragment. The main aim of this research is to enhance the performance of similarity searching by using multiple descriptors. In this paper, a deep learning method known as deep belief networks (DBN) has been used to reweight the molecule features. Several descriptors have been used for the MDL Drug Data Report (MDDR) dataset each of which represents different important features. The proposed method has been implemented with each descriptor individually to select the important features based on a new weight, with a lower error rate, and merging together all new features from all descriptors to produce a new descriptor for similarity searching. Based on the extensive experiments conducted, the results show that the proposed method outperformed several existing benchmark similarity methods, including Bayesian inference networks (BIN), the Tanimoto similarity method (TAN), adapted similarity measure of text processing (ASMTP) and the quantum-based similarity method (SQB). The results of this proposed multi-descriptor-based on Stack of deep belief networks method (SDBN) demonstrated a higher accuracy compared to existing methods on structurally heterogeneous datasets.

Download Full-text

Medicinal Chemistry Database GDBMedChem

10.26434/chemrxiv.7770809.v1 ◽

2019 ◽

Author(s):

Mahendra Awale ◽

Finton Sirockin ◽

Nikolaus Stiefl ◽

Jean-Louis Reymond

Keyword(s):

Small Molecules ◽

Natural Product ◽

Medicinal Chemistry ◽

3D Visualization ◽

Molecular Size ◽

Similarity Searching ◽

Complex Molecules ◽

Synthetic Accessibility ◽

Simple Chemical ◽

Reduced Complexity

<div>The generated database GDB17 enumerates 166.4 billion possible molecules up to 17 atoms of C, N, O, S and halogens following simple chemical stability and synthetic feasibility rules, however medicinal chemistry criteria are not taken into account. Here we applied rules inspired by medicinal chemistry to exclude problematic functional groups and complex molecules from GDB17, and sampled the resulting subset evenly across molecular size, stereochemistry and polarity to form GDBMedChem as a compact collection of 10 million small molecules.</div><div><br></div><div>This collection has reduced complexity and better synthetic accessibility than the entire GDB17 but retains higher sp 3 - carbon fraction and natural product likeness scores compared to known drugs. GDBMedChem molecules are more diverse and very different from known molecules in terms of substructures and represent an unprecedented source of diversity for drug design. GDBMedChem is available for 3D-visualization, similarity searching and for download at http://gdb.unibe.ch.</div>

Download Full-text

A Similarity Searching System for Biological Phenotype Images Using Deep Convolutional Encoder-decoder Architecture

Current Bioinformatics ◽

10.2174/1574893614666190204150109 ◽

2019 ◽

Vol 14 (7) ◽

pp. 628-639 ◽

Cited By ~ 10

Author(s):

Bizhi Wu ◽

Hangxiao Zhang ◽

Limei Lin ◽

Huiyuan Wang ◽

Yubang Gao ◽

...

Keyword(s):

Neural Network ◽

Retrieval System ◽

Sequence Similarity ◽

Local Alignment ◽

Similarity Searching ◽

Loss Of Function ◽

Biological Images ◽

The Neural Network ◽

Convolutional Autoencoder ◽

Biological Phenotype

Background: The BLAST (Basic Local Alignment Search Tool) algorithm has been widely used for sequence similarity searching. Analogously, the public phenotype images must be efficiently retrieved using biological images as queries and identify the phenotype with high similarity. Due to the accumulation of genotype-phenotype-mapping data, a system of searching for similar phenotypes is not available due to the bottleneck of image processing. Objective: In this study, we focus on the identification of similar query phenotypic images by searching the biological phenotype database, including information about loss-of-function and gain-of-function. Methods: We propose a deep convolutional autoencoder architecture to segment the biological phenotypic images and develop a phenotype retrieval system to enable a better understanding of genotype–phenotype correlation. Results: This study shows how deep convolutional autoencoder architecture can be trained on images from biological phenotypes to achieve state-of-the-art performance in a phenotypic images retrieval system. Conclusion: Taken together, the phenotype analysis system can provide further information on the correlation between genotype and phenotype. Additionally, it is obvious that the neural network model of image segmentation and the phenotype retrieval system is equally suitable for any species, which has enough phenotype images to train the neural network.

Download Full-text