Home LiteratureArticle Details
PMID: 28130360 Published · ppublish English Journal Article Research Support, U.S. Gov't, Non-P.H.S. Research Support, N.I.H., Intramural

Hybrid assembly of the large and highly repetitive genome of Aegilops tauschii, a progenitor of bread wheat, with the MaSuRCA mega-reads algorithm.

Genome research ·Vol. 27 ·No. 5 ·2017-00-00 ·Pages 787-792

Zimin AV, Puiu D, Luo MC, Zhu T, Koren S, Marçais G, Yorke JA, Dvořák J, Salzberg SL

Abstract

Long sequencing reads generated by single-molecule sequencing technology offer the possibility of dramatically improving the contiguity of genome assemblies. The biggest challenge today is that long reads have relatively high error rates, currently around 15%. The high error rates make it difficult to use this data alone, particularly with highly repetitive plant genomes. Errors in the raw data can lead to insertion or deletion errors (indels) in the consensus genome sequence, which in turn create significant problems for downstream analysis; for example, a single indel may shift the reading frame and incorrectly truncate a protein sequence. Here, we describe an algorithm that solves the high error rate problem by combining long, high-error reads with shorter but much more accurate Illumina sequencing reads, whose error rates average <1%. Our hybrid assembly algorithm combines these two types of reads to construct mega-reads, which are both long and accurate, and then assembles the mega-reads using the CABOG assembler, which was designed for long reads. We apply this technique to a large data set of Illumina and PacBio sequences from the species Aegilops tauschii, a large and extremely repetitive plant genome that has resisted previous attempts at assembly. We show that the resulting assembled contigs are far larger than in any previous assembly, with an N50 contig size of 486,807 nucleotides. We compare the contigs to independently produced optical maps to evaluate their large-scale accuracy, and to a set of high-quality bacterial artificial chromosome (BAC)-based assemblies to evaluate base-level accuracy.

MeSH Terms
Contig Mapping/methods,standards Genome Size Genome, Plant Genomics/methods,standards Poaceae/genetics Repetitive Sequences, Nucleic Acid Sequence Analysis, DNA/methods,standards Software
Authors & Affiliations
9 authors, click to expand affiliations / ORCID
Zimin Aleksey V
Center for Computational Biology, McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins School of Medicine, Baltimore, Maryland 21205, USA. | Institute for Physical Sciences and Technology, University of Maryland, College Park, Maryland 20742, USA.
Puiu Daniela
Center for Computational Biology, McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins School of Medicine, Baltimore, Maryland 21205, USA.
Luo Ming-Cheng
Department of Plant Sciences, University of California, Davis, California 95616, USA.
Zhu Tingting
Department of Plant Sciences, University of California, Davis, California 95616, USA.
Koren Sergey
National Human Genome Research Institute, National Institutes of Health, Bethesda, Maryland 20892, USA.
Marçais Guillaume
Institute for Physical Sciences and Technology, University of Maryland, College Park, Maryland 20742, USA. | Department of Computational Biology, Carnegie Mellon University, Pittsburgh, Pennsylvania 15213, USA.
Yorke James A
Institute for Physical Sciences and Technology, University of Maryland, College Park, Maryland 20742, USA. | Departments of Mathematics and Physics, University of Maryland, College Park, Maryland 20742, USA.
Dvořák Jan
Department of Plant Sciences, University of California, Davis, California 95616, USA.
Salzberg Steven L
Center for Computational Biology, McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins School of Medicine, Baltimore, Maryland 21205, USA. | Departments of Biomedical Engineering, Computer Science, and Biostatistics, Johns Hopkins University, Baltimore, Maryland 21218, USA.
References (20)
20 references, click to expand
  1. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler.
    Gigascience. 2012 Dec 27;1(1):18 PMID: 23587118
  2. LoRDEC: accurate and efficient long read error correction.
    Bioinformatics. 2014 Dec 15;30(24):3506-14 PMID: 25165095
  3. Aggressive assembly of pyrosequencing reads with mates.
    Bioinformatics. 2008 Dec 15;24(24):2818-24 PMID: 18952627
  4. proovread: large-scale high-accuracy PacBio correction through iterative short read consensus.
    Bioinformatics. 2014 Nov 1;30(21):3004-11 PMID: 25015988
  5. Fast algorithms for large-scale genome alignment and comparison.
    Nucleic Acids Res. 2002 Jun 1;30(11):2478-83 PMID: 12034836
  6. Assembling large genomes with single-molecule sequencing and locality-sensitive hashing.
    Nat Biotechnol. 2015 Jun;33(6):623-30 PMID: 26006009
  7. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data.
    Nat Methods. 2013 Jun;10(6):563-9 PMID: 23644548
  8. The MaSuRCA genome assembler.
    Bioinformatics. 2013 Nov 1;29(21):2669-77 PMID: 23990416
  9. Hybrid error correction and de novo assembly of single-molecule sequencing reads.
    Nat Biotechnol. 2012 Jul 01;30(7):693-700 PMID: 22750884
  10. Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences.
    Bioinformatics. 2016 Jul 15;32(14):2103-10 PMID: 27153593
  11. How important are transposons for plant evolution?
    Nat Rev Genet. 2013 Jan;14(1):49-61 PMID: 23247435
  12. Aegilops tauschii draft genome sequence reveals a gene repertoire for wheat adaptation.
    Nature. 2013 Apr 4;496(7443):91-5 PMID: 23535592
  13. A chromosome-based draft sequence of the hexaploid bread wheat (Triticum aestivum) genome.
    Science. 2014 Jul 18;345(6194):1251788 PMID: 25035500
  14. Analysis of tandem gene copies in maize chromosomal regions reconstructed from long sequence reads.
    Proc Natl Acad Sci U S A. 2016 Jul 19;113(29):7949-56 PMID: 27354512
  15. Versatile and open software for comparing large genomes.
    Genome Biol. 2004;5(2):R12 PMID: 14759262
  16. BioNano genome mapping of individual chromosomes supports physical mapping and sequence assembly in complex plant genomes.
    Plant Biotechnol J. 2016 Jul;14 (7):1523-31 PMID: 26801360
  17. Genome mapping on nanochannel arrays for structural variation analysis and sequence assembly.
    Nat Biotechnol. 2012 Aug;30(8):771-6 PMID: 22797562
  18. Rapid genome mapping in nanochannel arrays for highly complete and accurate de novo sequence assembly of the complex Aegilops tauschii genome.
    PLoS One. 2013;8(2):e55864 PMID: 23405223
  19. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation.
    Genome Res. 2017 May;27(5):722-736 PMID: 28298431
  20. De novo assembly of human genomes with massively parallel short read sequencing.
    Genome Res. 2010 Feb;20(2):265-72 PMID: 20019144
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1549-5469
Published
2017-00-00
Epub
2017-00-27
Pages
787-792
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC5411773
Subset
IM
Grants
NHGRI NIH HHS · R01 HG006677 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]