RANGI: A Fast List-Colored Graph Motif Finding Algorithm

Abstract Motivation The availability of numerous ChIP-seq datasets for transcription factors (TF) has provided an unprecedented opportunity to identify all TF binding sites in genomes. However, the progress has been hindered by the lack of a highly efficient and accurate tool to find not only the target motifs, but also cooperative motifs in very big datasets. Results We herein present an ultrafast and accurate motif-finding algorithm, ProSampler, based on a novel numeration method and Gibbs sampler. ProSampler runs orders of magnitude faster than the fastest existing tools while often more accurately identifying motifs of both the target TFs and cooperators. Thus, ProSampler can greatly facilitate the efforts to identify the entire cis-regulatory code in genomes. Availability and implementation Source code and binaries are freely available for download at https://github.com/zhengchangsulab/prosampler. It was implemented in C++ and supported on Linux, macOS and MS Windows platforms. Supplementary information Supplementary materials are available at Bioinformatics online.

Download Full-text

An efficient motif finding algorithm for large DNA data sets

2014 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) ◽

10.1109/bibm.2014.6999191 ◽

2014 ◽

Cited By ~ 4

Author(s):

Qiang Yu ◽

Hongwei Huo ◽

Xiaoyang Chen ◽

Haitao Guo ◽

Jeffrey Scott Vitter ◽

...

Keyword(s):

Motif Finding ◽

Data Sets ◽

Motif Finding Algorithm

Download Full-text

A Motif Finding Algorithm Based on Color Coding Technology

Journal of Software ◽

10.1360/jos181298 ◽

2007 ◽

Vol 18 (4) ◽

pp. 1298 ◽

Cited By ~ 1

Author(s):

Jian-Xin WANG

Keyword(s):

Motif Finding ◽

Color Coding ◽

Motif Finding Algorithm

Download Full-text

A deterministic motif finding algorithm with application to the human genome

Bioinformatics ◽

10.1093/bioinformatics/btl037 ◽

2006 ◽

Vol 22 (9) ◽

pp. 1047-1054 ◽

Cited By ~ 11

Author(s):

Lawrence S Hon ◽

Ajay N Jain

Keyword(s):

Human Genome ◽

Motif Finding ◽

Motif Finding Algorithm

Download Full-text

A Fast Cluster Motif Finding Algorithm for ChIP-Seq Data Sets

BioMed Research International ◽

10.1155/2015/218068 ◽

2015 ◽

Vol 2015 ◽

pp. 1-10 ◽

Cited By ~ 4

Author(s):

Yipu Zhang ◽

Ping Wang

Keyword(s):

High Throughput ◽

Motif Discovery ◽

Large Scale ◽

High Throughput Sequencing ◽

Es Cells ◽

Motif Finding ◽

Data Sets ◽

Data Set ◽

Binding Motifs ◽

Motif Finding Algorithm

New high-throughput technique ChIP-seq, coupling chromatin immunoprecipitation experiment with high-throughput sequencing technologies, has extended the identification of binding locations of a transcription factor to the genome-wide regions. However, the most existing motif discovery algorithms are time-consuming and limited to identify binding motifs in ChIP-seq data which normally has the significant characteristics of large scale data. In order to improve the efficiency, we propose a fast cluster motif finding algorithm, named as FCmotif, to identify the(l, d)motifs in large scale ChIP-seq data set. It is inspired by the emerging substrings mining strategy to find the enriched substrings and then searching the neighborhood instances to construct PWM and cluster motifs in different length. FCmotif is not following the OOPS model constraint and can find long motifs. The effectiveness of proposed algorithm has been proved by experiments on the ChIP-seq data sets from mouse ES cells. The whole detection of the real binding motifs and processing of the full size data of several megabytes finished in a few minutes. The experimental results show that FCmotif has advantageous to deal with the(l, d)motif finding in the ChIP-seq data; meanwhile it also demonstrates better performance than other current widely-used algorithms such as MEME, Weeder, ChIPMunk, and DREME.

Download Full-text

Comparison of result differences in multiple implementations of a stochastic motif finding algorithm

2015 E-Health and Bioengineering Conference (EHB) ◽

10.1109/ehb.2015.7391409 ◽

2015 ◽

Author(s):

Mihai Isaroiu ◽

Luca Dan Serbanati

Keyword(s):

Motif Finding ◽

Motif Finding Algorithm

Download Full-text

An Identical String Motif Finding Algorithm Through Dynamic Programming

Practical Applications of Computational Biology and Bioinformatics, 13th International Conference - Advances in Intelligent Systems and Computing ◽

10.1007/978-3-030-23873-5_10 ◽

2019 ◽

pp. 78-86

Author(s):

Abdelmenem S. Elgabry ◽

Tahani M. Allam ◽

Mahmoud M. Fahmy

Keyword(s):

Dynamic Programming ◽

Motif Finding ◽

Motif Finding Algorithm

Download Full-text

Motif-Based Text Mining of Microbial Metagenome Redundancy Profiling Data for Disease Classification

BioMed Research International ◽

10.1155/2016/6598307 ◽

2016 ◽

Vol 2016 ◽

pp. 1-11 ◽

Cited By ~ 2

Author(s):

Yin Wang ◽

Rudong Li ◽

Yuhua Zhou ◽

Zongxin Ling ◽

Xiaokui Guo ◽

...

Keyword(s):

Dental Caries ◽

16S Rrna ◽

High Dimension ◽

Disease Classification ◽

Motif Finding ◽

Text Data ◽

Feature Spaces ◽

Efficiency And Reliability ◽

Motif Finding Algorithm ◽

Better Than

Background. Text data of 16S rRNA are informative for classifications of microbiota-associated diseases. However, the raw text data need to be systematically processed so that features for classification can be defined/extracted; moreover, the high-dimension feature spaces generated by the text data also pose an additional difficulty.Results. Here we present a Phylogenetic Tree-Based Motif Finding algorithm (PMF) to analyze 16S rRNA text data. By integrating phylogenetic rules and other statistical indexes for classification, we can effectively reduce the dimension of the large feature spaces generated by the text datasets. Using the retrieved motifs in combination with common classification methods, we can discriminate different samples of both pneumonia and dental caries better than other existing methods.Conclusions. We extend the phylogenetic approaches to perform supervised learning on microbiota text data to discriminate the pathological states for pneumonia and dental caries. The results have shown that PMF may enhance the efficiency and reliability in analyzing high-dimension text data.

Download Full-text