DNA sequence compression using the Burrows-Wheeler Transform

This paper proposes a seed based lossless compression algorithm to compress a DNA sequence which uses a substitution method that is similar to the LempelZiv compression scheme. The proposed method exploits the repetition structures that are inherent in DNA sequences by creating an offline dictionary which contains all such repeats along with the details of mismatches. By ensuring that only promising mismatches are allowed, the method achieves a compression ratio that is at par or better than the existing lossless DNA sequence compression algorithms.

Download Full-text

Self-configuration single particle optimizer for DNA sequence compression

Soft Computing ◽

10.1007/s00500-012-0939-9 ◽

2012 ◽

Vol 17 (4) ◽

pp. 675-682 ◽

Cited By ~ 3

Author(s):

Zhen Ji ◽

Jiarui Zhou ◽

Zexuan Zhu ◽

Siping Chen

Keyword(s):

Dna Sequence ◽

Single Particle ◽

Dna Sequence Compression ◽

Sequence Compression

Download Full-text

Towards Context-Aware DNA Sequence Compression Algorithms

2019 IEEE International Conference on Electrical, Computer and Communication Technologies (ICECCT) ◽

10.1109/icecct.2019.8869330 ◽

2019 ◽

Author(s):

Neha A.S. ◽

Salim A.

Keyword(s):

Dna Sequence ◽

Context Aware ◽

Compression Algorithms ◽

Dna Sequence Compression ◽

Sequence Compression

Download Full-text

Algorithm for DNA sequence compression based on prediction of mismatch bases and repeat location

2010 IEEE International Conference on Bioinformatics and Biomedicine Workshops (BIBMW) ◽

10.1109/bibmw.2010.5703941 ◽

2010 ◽

Cited By ~ 4

Author(s):

Kalyan Kumar Kaipa ◽

Ajit S Bopardikar ◽

Srikantha Abhilash ◽

Parthasarathy Venkataraman ◽

Kyusang Lee ◽

...

Keyword(s):

Dna Sequence ◽

Dna Sequence Compression ◽

Sequence Compression

Download Full-text

Efficient DNA sequence compression with neural networks

GigaScience ◽

10.1093/gigascience/giaa119 ◽

2020 ◽

Vol 9 (11) ◽

Cited By ~ 1

Author(s):

Milton Silva ◽

Diogo Pratas ◽

Armando J Pinho

Keyword(s):

Neural Networks ◽

Data Analysis ◽

Dna Sequence ◽

Dna Sequences ◽

Genomic Sequence ◽

State Of The Art ◽

The State ◽

Computational Time ◽

Dna Sequence Compression ◽

Sequence Compression

Abstract Background The increasing production of genomic data has led to an intensified need for models that can cope efficiently with the lossless compression of DNA sequences. Important applications include long-term storage and compression-based data analysis. In the literature, only a few recent articles propose the use of neural networks for DNA sequence compression. However, they fall short when compared with specific DNA compression tools, such as GeCo2. This limitation is due to the absence of models specifically designed for DNA sequences. In this work, we combine the power of neural networks with specific DNA models. For this purpose, we created GeCo3, a new genomic sequence compressor that uses neural networks for mixing multiple context and substitution-tolerant context models. Findings We benchmark GeCo3 as a reference-free DNA compressor in 5 datasets, including a balanced and comprehensive dataset of DNA sequences, the Y-chromosome and human mitogenome, 2 compilations of archaeal and virus genomes, 4 whole genomes, and 2 collections of FASTQ data of a human virome and ancient DNA. GeCo3 achieves a solid improvement in compression over the previous version (GeCo2) of $2.4\%$, $7.1\%$, $6.1\%$, $5.8\%$, and $6.0\%$, respectively. To test its performance as a reference-based DNA compressor, we benchmark GeCo3 in 4 datasets constituted by the pairwise compression of the chromosomes of the genomes of several primates. GeCo3 improves the compression in $12.4\%$, $11.7\%$, $10.8\%$, and $10.1\%$ over the state of the art. The cost of this compression improvement is some additional computational time (1.7–3 times slower than GeCo2). The RAM use is constant, and the tool scales efficiently, independently of the sequence size. Overall, these values outperform the state of the art. Conclusions GeCo3 is a genomic sequence compressor with a neural network mixing approach that provides additional gains over top specific genomic compressors. The proposed mixing method is portable, requiring only the probabilities of the models as inputs, providing easy adaptation to other data compressors or compression-based data analysis tools. GeCo3 is released under GPLv3 and is available for free download at https://github.com/cobilab/geco3.

Download Full-text

DNA sequence compression using the Burrows-Wheeler Transform

Polynomial Based Representation for DNA Sequence Compression and Search

SBVRLDNAComp:An Effective DNA Sequence Compression Algorithm

DNA Sequence Compression Using Adaptive Particle Swarm Optimization-Based Memetic Algorithm

DNA sequence compression within traditional text compression algorithms

An efficient normalized maximum likelihood algorithm for DNA sequence compression

An Optimal Seed Based Compression Algorithm for DNA Sequences

Self-configuration single particle optimizer for DNA sequence compression

Towards Context-Aware DNA Sequence Compression Algorithms

Algorithm for DNA sequence compression based on prediction of mismatch bases and repeat location

Efficient DNA sequence compression with neural networks

Export Citation Format