scholarly journals BLAMM: BLAS-based algorithm for finding position weight matrix occurrences in DNA sequences on CPUs and GPUs

2020 ◽  
Vol 21 (S2) ◽  
Author(s):  
Jan Fostier
2009 ◽  
Vol 25 (23) ◽  
pp. 3181-3182 ◽  
Author(s):  
J. Korhonen ◽  
P. Martinmaki ◽  
C. Pizzi ◽  
P. Rastas ◽  
E. Ukkonen

Author(s):  
Lee A Newberg ◽  
Lee Ann McCue ◽  
Charles E Lawrence

Approaches based upon sequence weights, to construct a position weight matrix of nucleotides from aligned inputs, are popular but little effort has been expended to measure their quality.We derive optimal sequence weights that minimize the sum of the variances of the estimators of base frequency parameters for sequences related by a phylogenetic tree. Using these we find that approaches based upon sequence weights can perform very poorly in comparison to approaches based upon a theoretically optimal maximum-likelihood method in the inference of the parameters of a position-weight matrix. Specifically, we find that among a collection of primate sequences, even an optimal sequences-weights approach is only 51% as efficient as the maximum-likelihood approach in inferences of base frequency parameters.We also show how to employ the variance estimators to obtain a greedy ordering of species for sequencing. Application of this ordering for the weighted estimators to a primate collection yields a curve with a long plateau that is not observed with maximum-likelihood estimators. This plateau indicates that the use of weighted estimators on these data seriously limits the utility of obtaining the sequences of more than two or three additional species.


Sign in / Sign up

Export Citation Format

Share Document