Home LiteratureArticle Details
PMID: 23990416 Published · ppublish English Evaluation Study Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

The MaSuRCA genome assembler.

Bioinformatics (Oxford, England) ·Vol. 29 ·No. 21 ·2013-11-01 ·Pages 2669-77

Zimin AV, Marçais G, Puiu D, Roberts M, Salzberg SL, Yorke JA

Abstract

Second-generation sequencing technologies produce high coverage of the genome by short reads at a low cost, which has prompted development of new assembly methods. In particular, multiple algorithms based on de Bruijn graphs have been shown to be effective for the assembly problem. In this article, we describe a new hybrid approach that has the computational efficiency of de Bruijn graph methods and the flexibility of overlap-based assembly strategies, and which allows variable read lengths while tolerating a significant level of sequencing error. Our method transforms large numbers of paired-end reads into a much smaller number of longer 'super-reads'. The use of super-reads allows us to assemble combinations of Illumina reads of differing lengths together with longer reads from 454 and Sanger sequencing technologies, making it one of the few assemblers capable of handling such mixtures. We call our system the Maryland Super-Read Celera Assembler (abbreviated MaSuRCA and pronounced 'mazurka'). We evaluate the performance of MaSuRCA against two of the most widely used assemblers for Illumina data, Allpaths-LG and SOAPdenovo2, on two datasets from organisms for which high-quality assemblies are available: the bacterium Rhodobacter sphaeroides and chromosome 16 of the mouse genome. We show that MaSuRCA performs on par or better than Allpaths-LG and significantly better than SOAPdenovo on these data, when evaluated against the finished sequence. We then show that MaSuRCA can significantly improve its assemblies when the original data are augmented with long reads. MaSuRCA is available as open-source code at ftp://ftp.genome.umd.edu/pub/MaSuRCA/. Previous (pre-publication) releases have been publicly available for over a year. [email protected]. Supplementary data are available at Bioinformatics online.

MeSH Terms
Algorithms Animals Genome, Bacterial Genomics/methods Mice Rhodobacter sphaeroides/genetics Sequence Analysis, DNA/methods Software
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Zimin Aleksey V
Institute for Physical Sciences and Technology, University of Maryland, College Park, MD 20742, USA, Center for Computational Biology, McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University School of Medicine, Baltimore, MD 21205, USA, Department of Mathematics and Department of Physics, University of Maryland, College Park, MD 20742, USA.
Marçais Guillaume
Puiu Daniela
Roberts Michael
Salzberg Steven L
Yorke James A
References (33)
33 references, click to expand
  1. The phusion assembler.
    Genome Res. 2003 Jan;13(1):81-90 PMID: 12529309
  2. Aggressive assembly of pyrosequencing reads with mates.
    Bioinformatics. 2008 Dec 15;24(24):2818-24 PMID: 18952627
  3. Versatile and open software for comparing large genomes.
    Genome Biol. 2004;5(2):R12 PMID: 14759262
  4. A new algorithm for DNA sequence assembly.
    J Comput Biol. 1995 Summer;2(2):291-306 PMID: 7497130
  5. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  6. SOAP: short oligonucleotide alignment program.
    Bioinformatics. 2008 Mar 1;24(5):713-4 PMID: 18227114
  7. GAGE-B: an evaluation of genome assemblers for bacterial organisms.
    Bioinformatics. 2013 Jul 15;29(14):1718-25 PMID: 23665771
  8. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  9. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers.
    Bioinformatics. 2011 Mar 15;27(6):764-70 PMID: 21217122
  10. GAGE: A critical evaluation of genome assemblies and assembly algorithms.
    Genome Res. 2012 Mar;22(3):557-67 PMID: 22147368
  11. QuorUM: An Error Corrector for Illumina Reads.
    PLoS One. 2015 Jun 17;10(6):e0130821 PMID: 26083032
  12. 1-Tuple DNA sequencing: computer analysis.
    J Biomol Struct Dyn. 1989 Aug;7(1):63-73 PMID: 2684223
  13. Short read fragment assembly of bacterial genomes.
    Genome Res. 2008 Feb;18(2):324-30 PMID: 18083777
  14. Bambus 2: scaffolding metagenomes.
    Bioinformatics. 2011 Nov 1;27(21):2964-71 PMID: 21926123
  15. Efficient de novo assembly of large genomes using compressed data structures.
    Genome Res. 2012 Mar;22(3):549-56 PMID: 22156294
  16. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler.
    Gigascience. 2012 Dec 27;1(1):18 PMID: 23587118
  17. De novo assembly of human genomes with massively parallel short read sequencing.
    Genome Res. 2010 Feb;20(2):265-72 PMID: 20019144
  18. PCAP: a whole-genome assembly program.
    Genome Res. 2003 Sep;13(9):2164-70 PMID: 12952883
  19. Genome analyses of three strains of Rhodobacter sphaeroides: evidence of rapid evolution of chromosome II.
    J Bacteriol. 2007 Mar;189(5):1914-21 PMID: 17172323
  20. Fast gapped-read alignment with Bowtie 2.
    Nat Methods. 2012 Mar 04;9(4):357-9 PMID: 22388286
  21. Using the miraEST assembler for reliable and automated mRNA transcript assembly and SNP detection in sequenced ESTs.
    Genome Res. 2004 Jun;14(6):1147-59 PMID: 15140833
  22. An Eulerian path approach to DNA fragment assembly.
    Proc Natl Acad Sci U S A. 2001 Aug 14;98(17):9748-53 PMID: 11504945
  23. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing.
    J Comput Biol. 2012 May;19(5):455-77 PMID: 22506599
  24. ARACHNE: a whole-genome shotgun assembler.
    Genome Res. 2002 Jan;12(1):177-89 PMID: 11779843
  25. The sequence of the human genome.
    Science. 2001 Feb 16;291(5507):1304-51 PMID: 11181995
  26. QUAST: quality assessment tool for genome assemblies.
    Bioinformatics. 2013 Apr 15;29(8):1072-5 PMID: 23422339
  27. High-quality draft assemblies of mammalian genomes from massively parallel sequence data.
    Proc Natl Acad Sci U S A. 2011 Jan 25;108(4):1513-8 PMID: 21187386
  28. Error correction of high-throughput sequencing datasets with non-uniform coverage.
    Bioinformatics. 2011 Jul 1;27(13):i137-41 PMID: 21685062
  29. Quake: quality-aware detection and correction of sequencing errors.
    Genome Biol. 2010;11(11):R116 PMID: 21114842
  30. ABySS: a parallel assembler for short read sequence data.
    Genome Res. 2009 Jun;19(6):1117-23 PMID: 19251739
  31. Assembly algorithms for next-generation sequencing data.
    Genomics. 2010 Jun;95(6):315-27 PMID: 20211242
  32. A whole-genome assembly of Drosophila.
    Science. 2000 Mar 24;287(5461):2196-204 PMID: 10731133
  33. Initial sequencing and comparative analysis of the mouse genome.
    Nature. 2002 Dec 5;420(6915):520-62 PMID: 12466850
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2013-11-01
Epub
2013-00-29
Pages
2669-77
Language
English
Region
England
NLM ID
9808944
PMCID
PMC3799473
Subset
IM
Grants
NHGRI NIH HHS · R01 HG006677 · United States
NHGRI NIH HHS · R01-HG002945 · United States
NHGRI NIH HHS · R01-HG006677 · United States
NHGRI NIH HHS · R21 HG006913 · United States
NHGRI NIH HHS · R01 HG002945 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]