Using Apache Spark on genome assembly for scalable overlap-graph reduction

Abstract Background De novo genome assembly is a technique that builds the genome of a specimen using overlaps of genomic fragments without additional work with reference sequence. Sequence fragments (called reads) are assembled as contigs and scaffolds by the overlaps. The quality of the de novo assembly depends on the length and continuity of the assembly. To enable faster and more accurate assembly of species, existing sequencing techniques have been proposed, for example, high-throughput next-generation sequencing and long-reads-producing third-generation sequencing. However, these techniques require a large amounts of computer memory when very huge-size overlap graphs are resolved. Also, it is challenging for parallel computation. Results To address the limitations, we propose an innovative algorithmic approach, called Scalable Overlap-graph Reduction Algorithms (SORA). SORA is an algorithm package that performs string graph reduction algorithms by Apache Spark. The SORA’s implementations are designed to execute de novo genome assembly on either a single machine or a distributed computing platform. SORA efficiently compacts the number of edges on enormous graphing paths by adapting scalable features of graph processing libraries provided by Apache Spark, GraphX and GraphFrames. Conclusions We shared the algorithms and the experimental results at our project website, https://github.com/BioHPC/SORA. We evaluated SORA with the human genome samples. First, it processed a nearly one billion edge graph on a distributed cloud cluster. Second, it processed mid-to-small size graphs on a single workstation within a short time frame. Overall, SORA achieved the linear-scaling simulations for the increased computing instances.

Download Full-text

De novo Genome Assembly from Next-Generation Sequencing (NGS) Reads

Next-Generation Sequencing Data Analysis ◽

10.1201/b19532-11 ◽

2016 ◽

pp. 144-155

Keyword(s):

Next Generation Sequencing ◽

Genome Assembly ◽

De Novo ◽

Next Generation ◽

De Novo Genome Assembly ◽

Next Generation Sequencing Ngs ◽

Generation Sequencing

Download Full-text

De Novo Genome Assembly of Next-Generation Sequencing Data

Compendium of Plant Genomes - The Brassica rapa Genome ◽

10.1007/978-3-662-47901-8_4 ◽

2015 ◽

pp. 41-51

Author(s):

Min Liu ◽

Dongyuan Liu ◽

Hongkun Zheng

Keyword(s):

Next Generation Sequencing ◽

Genome Assembly ◽

De Novo ◽

Next Generation Sequencing Data ◽

Next Generation ◽

Sequencing Data ◽

De Novo Genome Assembly ◽

Generation Sequencing

Download Full-text

A Practical Comparison of De Novo Genome Assembly Software Tools for Next-Generation Sequencing Technologies

PLoS ONE ◽

10.1371/journal.pone.0017915 ◽

2011 ◽

Vol 6 (3) ◽

pp. e17915 ◽

Cited By ~ 144

Author(s):

Wenyu Zhang ◽

Jiajia Chen ◽

Yang Yang ◽

Yifei Tang ◽

Jing Shang ◽

...

Keyword(s):

Next Generation Sequencing ◽

Genome Assembly ◽

De Novo ◽

Software Tools ◽

Next Generation ◽

De Novo Genome Assembly ◽

Sequencing Technologies ◽

Generation Sequencing ◽

Assembly Software

Download Full-text

De Novo genome assembly for third generation sequencing data

Photonics Applications in Astronomy, Communications, Industry, and High-Energy Physics Experiments 2018 ◽

10.1117/12.2501543 ◽

2018 ◽

Author(s):

Robert M. Nowak ◽

Mateusz Forc ◽

Wiktor Kuśmirek

Keyword(s):

Genome Assembly ◽

De Novo ◽

Third Generation ◽

Sequencing Data ◽

De Novo Genome Assembly ◽

Third Generation Sequencing ◽

Generation Sequencing

Download Full-text

Subset selection of high-depth next generation sequencing reads for de novo genome assembly using MapReduce framework

BMC Genomics ◽

10.1186/1471-2164-16-s12-s9 ◽

2015 ◽

Vol 16 (Suppl 12) ◽

pp. S9 ◽

Cited By ~ 4

Author(s):

Chih-Hao Fang ◽

Yu-Jung Chang ◽

Wei-Chun Chung ◽

Ping-Heng Hsieh ◽

Chung-Yen Lin ◽

...

Keyword(s):

Next Generation Sequencing ◽

Genome Assembly ◽

De Novo ◽

Subset Selection ◽

Next Generation ◽

Mapreduce Framework ◽

De Novo Genome Assembly ◽

Generation Sequencing ◽

High Depth ◽

Selection Of

Download Full-text

Effects of GC Bias in Next-Generation-Sequencing Data on De Novo Genome Assembly

PLoS ONE ◽

10.1371/journal.pone.0062856 ◽

2013 ◽

Vol 8 (4) ◽

pp. e62856 ◽

Cited By ~ 121

Author(s):

Yen-Chun Chen ◽

Tsunglin Liu ◽

Chun-Hui Yu ◽

Tzen-Yuh Chiang ◽

Chi-Chuan Hwang

Keyword(s):

Next Generation Sequencing ◽

Genome Assembly ◽

De Novo ◽

Next Generation Sequencing Data ◽

Next Generation ◽

Sequencing Data ◽

De Novo Genome Assembly ◽

Gc Bias ◽

Generation Sequencing

Download Full-text

GraphSeq: Accelerating String Graph Construction for De Novo Assembly on Spark

10.1101/321729 ◽

2018 ◽

Author(s):

Chung-Tsai Su ◽

Ming-Tai Chang ◽

Yun-Chian Cheng ◽

Yun-Lung Li ◽

Yao-Ting Wang

Keyword(s):

Genome Assembly ◽

De Novo Assembly ◽

De Novo ◽

Data Representation ◽

Important Application ◽

Supplementary Information ◽

De Novo Genome Assembly ◽

String Graph ◽

Computing Framework ◽

Variant Identification

AbstractSummary: De novo genome assembly is an important application on both uncharacterized genome assembly and variant identification in a reference-unbiased way. In comparison with de Brujin graph, string graph is a lossless data representation for de novo assembly. However, string graph construction is computational intensive. We propose GraphSeq to accelerate string graph construction by leveraging the distributed computing framework.Availability and Implementation: GraphSeq is implemented with Scala on Spark and freely available at https://www.atgenomix.com/blog/graphseq.Supplementary information: Supplementary data are available at Bioinformatics online.

Download Full-text