Nanopore Long Read DNA Sequencing of Protozoan Parasites: Hybrid Genome Assembly of Trypanosoma cruzi

Comprehensive evaluation of non-hybrid genome assembly tools for third-generation PacBio long-read sequence data

Briefings in Bioinformatics ◽

10.1093/bib/bbx147 ◽

2017 ◽

Vol 20 (3) ◽

pp. 866-876 ◽

Cited By ~ 30

Author(s):

Vasanthan Jayakumar ◽

Yasubumi Sakakibara

Keyword(s):

Genome Assembly ◽

Comprehensive Evaluation ◽

Sequence Data ◽

Third Generation ◽

Hybrid Genome ◽

Long Read

Download Full-text

Long read and single molecule DNA sequencing simplifies genome assembly and TAL effector gene analysis of Xanthomonas translucens

BMC Genomics ◽

10.1186/s12864-015-2348-9 ◽

2016 ◽

Vol 17 (1) ◽

Cited By ~ 32

Author(s):

Zhao Peng ◽

Ying Hu ◽

Jingzhong Xie ◽

Neha Potnis ◽

Alina Akhunova ◽

...

Keyword(s):

Dna Sequencing ◽

Single Molecule ◽

Genome Assembly ◽

Gene Analysis ◽

Tal Effector ◽

Effector Gene ◽

Long Read ◽

Xanthomonas Translucens

Download Full-text

LazyB: fast and cheap genome assembly

Algorithms for Molecular Biology ◽

10.1186/s13015-021-00186-5 ◽

2021 ◽

Vol 16 (1) ◽

Author(s):

Thomas Gatter ◽

Sarah von Löhneysen ◽

Jörg Fallmann ◽

Polina Drozdova ◽

Tom Hartmann ◽

...

Keyword(s):

Genome Assembly ◽

Hybrid Methods ◽

Computational Effort ◽

Sequencing Data ◽

Long Reads ◽

Hybrid Genome ◽

Long Read ◽

Low Coverage ◽

Genome Assembler ◽

Large Genomes

Abstract Background Advances in genome sequencing over the last years have lead to a fundamental paradigm shift in the field. With steadily decreasing sequencing costs, genome projects are no longer limited by the cost of raw sequencing data, but rather by computational problems associated with genome assembly. There is an urgent demand for more efficient and and more accurate methods is particular with regard to the highly complex and often very large genomes of animals and plants. Most recently, “hybrid” methods that integrate short and long read data have been devised to address this need. Results is such a hybrid genome assembler. It has been designed specificially with an emphasis on utilizing low-coverage short and long reads. starts from a bipartite overlap graph between long reads and restrictively filtered short-read unitigs. This graph is translated into a long-read overlap graph G. Instead of the more conventional approach of removing tips, bubbles, and other local features, stepwisely extracts subgraphs whose global properties approach a disjoint union of paths. First, a consistently oriented subgraph is extracted, which in a second step is reduced to a directed acyclic graph. In the next step, properties of proper interval graphs are used to extract contigs as maximum weight paths. These path are translated into genomic sequences only in the final step. A prototype implementation of , entirely written in python, not only yields significantly more accurate assemblies of the yeast and fruit fly genomes compared to state-of-the-art pipelines but also requires much less computational effort. Conclusions is new low-cost genome assembler that copes well with large genomes and low coverage. It is based on a novel approach for reducing the overlap graph to a collection of paths, thus opening new avenues for future improvements. Availability The prototype is available at https://github.com/TGatter/LazyB.

Download Full-text

Amynthas corticis genome reveals molecular mechanisms behind global distribution

Communications Biology ◽

10.1038/s42003-021-01659-4 ◽

2021 ◽

Vol 4 (1) ◽

Author(s):

Xing Wang ◽

Yi Zhang ◽

Yufeng Zhang ◽

Mingming Kang ◽

Yuanbo Li ◽

...

Keyword(s):

Genome Assembly ◽

Molecular Mechanisms ◽

Gene Families ◽

The Body ◽

Gene Family Evolution ◽

Complex Environments ◽

Protein Coding ◽

Itraq Analysis ◽

Rdna Sequencing ◽

Long Read

AbstractEarthworms (Annelida: Crassiclitellata) are widely distributed around the world due to their ancient origination as well as adaptation and invasion after introduction into new habitats over the past few centuries. Herein, we report a 1.2 Gb complete genome assembly of the earthworm Amynthas corticis based on a strategy combining third-generation long-read sequencing and Hi-C mapping. A total of 29,256 protein-coding genes are annotated in this genome. Analysis of resequencing data indicates that this earthworm is a triploid species. Furthermore, gene family evolution analysis shows that comprehensive expansion of gene families in the Amynthas corticis genome has produced more defensive functions compared with other species in Annelida. Quantitative proteomic iTRAQ analysis shows that expression of 147 proteins changed in the body of Amynthas corticis and 16 S rDNA sequencing shows that abundance of 28 microorganisms changed in the gut of Amynthas corticis when the earthworm was incubated with pathogenic Escherichia coli O157:H7. Our genome assembly provides abundant and valuable resources for the earthworm research community, serving as a first step toward uncovering the mysteries of this species, and may provide molecular level indicators of its powerful defensive functions, adaptation to complex environments and invasion ability.

Download Full-text

Genome sequence resource of Phomopsis longicolla strain YC2-1, a fungal pathogen causing Phomopsis stem blight in soybean

Molecular Plant-Microbe Interactions ◽

10.1094/mpmi-12-20-0340-a ◽

2021 ◽

Author(s):

Xiaolin Zhao ◽

Zhichao Zhang ◽

Sujiao Zheng ◽

Wenwu Ye ◽

Xiaobo Zheng ◽

...

Keyword(s):

Genome Assembly ◽

Stem Canker ◽

Quality Data ◽

Phomopsis Longicolla ◽

Protein Coding ◽

Stem Blight ◽

A Genome ◽

Long Read ◽

Genomic Resource ◽

Blight Disease

Diaporthe-Phomopsis disease complex causes considerable yield losses in soybean production worldwide. As one of the major pathogens, Phomopsis longicolla T. W. Hobbs (syn. Diaporthe longicolla) is not only the primary agent of Phomopsis seed decay, but also one of the agents of Phomopsis pod and stem blight, and Phomopsis stem canker. We performed both PacBio long read sequencing and Illumina short read sequencing, and obtained a genome assembly for the P. longicolla strain YC2-1, which was isolated from soybean stem with Phomopsis stem blight disease. The 63.1 Mb genome assembly contains 87 scaffolds, with a minimum, maximum, and N50 scaffold length of 20 kb, 4.6 Mb, and 1.5 Mb respectively, and a total of 17,407 protein-coding genes. The high-quality data expand the genomic resource of P. longicolla species and will provide a solid foundation for a better understanding of their genetic diversity and pathogenic mechanisms.

Download Full-text

Chromosome-level assembly of Drosophila bifasciata reveals important karyotypic transition of the X chromosome

10.1101/847558 ◽

2019 ◽

Author(s):

Ryan Bracewell ◽

Anita Tran ◽

Kamalakar Chatla ◽

Doris Bachtrog

Keyword(s):

X Chromosome ◽

Genome Assembly ◽

De Novo ◽

Pericentromeric Region ◽

Species Group ◽

Chromosome 15 ◽

Protein Coding ◽

Protein Coding Genes ◽

Long Read ◽

Chromosome Level

ABSTRACTThe Drosophila obscura species group is one of the most studied clades of Drosophila and harbors multiple distinct karyotypes. Here we present a de novo genome assembly and annotation of D. bifasciata, a species which represents an important subgroup for which no high-quality chromosome-level genome assembly currently exists. We combined long-read sequencing (Nanopore) and Hi-C scaffolding to achieve a highly contiguous genome assembly approximately 193Mb in size, with repetitive elements constituting 30.1% of the total length. Drosophila bifasciata harbors four large metacentric chromosomes and the small dot, and our assembly contains each chromosome in a single scaffold, including the highly repetitive pericentromere, which were largely composed of Jockey and Gypsy transposable elements. We annotated a total of 12,821 protein-coding genes and comparisons of synteny with D. athabasca orthologs show that the large metacentric pericentromeric regions of multiple chromosomes are conserved between these species. Importantly, Muller A (X chromosome) was found to be metacentric in D. bifasciata and the pericentromeric region appears homologous to the pericentromeric region of the fused Muller A-AD (XL and XR) of pseudoobscura/affinis subgroup species. Our finding suggests a metacentric ancestral X fused to a telocentric Muller D and created the large neo-X (Muller A-AD) chromosome ∼15 MYA. We also confirm the fusion of Muller C and D in D. bifasciata and show that it likely involved a centromere-centromere fusion.

Download Full-text

Accurate long-read de novo assembly evaluation with Inspector

Genome Biology ◽

10.1186/s13059-021-02527-4 ◽

2021 ◽

Vol 22 (1) ◽

Author(s):

Yu Chen ◽

Yixin Zhang ◽

Amy Y. Wang ◽

Min Gao ◽

Zechen Chong

Keyword(s):

Genome Assembly ◽

De Novo Assembly ◽

In Silico ◽

Large Scale ◽

De Novo ◽

Small Scale ◽

De Novo Genome Assembly ◽

Consensus Sequences ◽

Assembly Evaluation ◽

Long Read

AbstractLong-read de novo genome assembly continues to advance rapidly. However, there is a lack of effective tools to accurately evaluate the assembly results, especially for structural errors. We present Inspector, a reference-free long-read de novo assembly evaluator which faithfully reports types of errors and their precise locations. Notably, Inspector can correct the assembly errors based on consensus sequences derived from raw reads covering erroneous regions. Based on in silico and long-read assembly results from multiple long-read data and assemblers, we demonstrate that in addition to providing generic metrics, Inspector can accurately identify both large-scale and small-scale assembly errors.

Download Full-text

Hybrid Genome Assembly of a Regenerative Chordate

The FASEB Journal ◽

10.1096/fasebj.2020.34.s1.09936 ◽

2020 ◽

Vol 34 (S1) ◽

pp. 1-1

Author(s):

Jack Thomas Sumner ◽

Sarah Wax ◽

Cassidy Andrasz ◽

Paul Anderson ◽

Elena Keeling ◽

...

Keyword(s):

Genome Assembly ◽

Hybrid Genome

Download Full-text

Long-read assembly and comparative evidence-based reanalysis of Cryptosporidium genome sequences reveals expanded transporter repertoire and duplication of entire chromosome ends including subtelomeric regions

Genome Research ◽

10.1101/gr.275325.121 ◽

2021 ◽

pp. gr.275325.121

Author(s):

Rodrigo P. Baptista ◽

Yiran Li ◽

Adam Sateriale ◽

Karen L. Brooks ◽

Alan Tracey ◽

...

Keyword(s):

Genome Assembly ◽

Reference Genome ◽

Diarrheal Disease ◽

Gene Copy ◽

Future Research ◽

Single Nucleotide Variants ◽

Cryptosporidium Hominis ◽

Entire Chromosome ◽

Long Read ◽

Subtelomeric Regions

Cryptosporidiosis is a leading cause of waterborne diarrheal disease globally and an important contributor to mortality in infants and the immunosuppressed. Despite its importance, the Cryptosporidium community has only had access to a good, but incomplete, Cryptosporidium parvum IOWA reference genome sequence. Incomplete reference sequences hamper annotation, experimental design and interpretation. We have generated a new C. parvum IOWA genome assembly supported by PacBio and Oxford Nanopore long-read technologies and a new comparative and consistent genome annotation for three closely related species C. parvum, Cryptosporidium hominis and Cryptosporidium tyzzeri. We made 1,926 C. parvum annotation updates based on experimental evidence. They include new transporters, ncRNAs, introns and altered gene structures. The new assembly and annotation revealed a complete Dnmt2 methylase ortholog. Comparative annotation between C. parvum, C. hominis and C. tyzzeri revealed that most "missing" orthologs are found suggesting that the biological differences between the species must result from gene copy number variation, differences in gene regulation and single nucleotide variants (SNVs). Using the new assembly and annotation as reference, 190 genes are identified as evolving under positive selection, including many not detected previously. The new C. parvum IOWA reference genome assembly is larger, gap free and lacks ambiguous bases. This chromosomal assembly recovers all 16 chromosome ends, 13 of which are contiguously assembled. The three remaining chromosome ends are provisionally placed. These ends represent duplication of entire chromosome ends including subtelomeric regions revealing a new level of genome plasticity that will both inform and impact future research.

Download Full-text

LRSim: a Linked Reads Simulator generating insights for better genome partitioning

10.1101/103549 ◽

2017 ◽

Cited By ~ 1

Author(s):

Ruibang Luo ◽

Fritz J. Sedlazeck ◽

Charlotte A. Darby ◽

Stephen M. Kelly ◽

Michael C. Schatz

Keyword(s):

Genome Assembly ◽

Fragment Length ◽

De Novo ◽

Library Preparation ◽

Haplotype Phasing ◽

Multiple Datasets ◽

Long Read ◽

Fine Control ◽

Informatics Tools ◽

Sequencing Process

AbstractMotivationLinked reads are a form of DNA sequencing commercialized by 10X Genomics that uses highly multiplexed barcoding within microdroplets to tag short reads to progenitor molecules. The linked reads, spanning tens to hundreds of kilobases, offer an alternative to long-read sequencing for de novo assembly, haplotype phasing and other applications. However, there is no available simulator, making it difficult to measure their capability or develop new informatics tools.ResultsOur analysis of 13 real linked read datasets revealed their characteristics of barcodes, molecules and partitions. Based on this, we introduce LRSim that simulates linked reads by emulating the library preparation and sequencing process with fine control of 1) the number of simulated variants; 2) the linked-read characteristics; and 3) the Illumina reads profile. We conclude from the phasing and genome assembly of multiple datasets, recommendations on coverage, fragment length, and partitioning when sequencing human and non-human genome.AvailabilityLRSIM is under MIT license and is freely available at https://github.com/aquaskyline/[email protected]

Download Full-text