Abstract
The MUMmer system and the genome sequence aligner nucmer included within it are among the most widely used alignment packages in genomics. Since the last major release of MUMmer version 3 in 2004, it has been applied to many types of problems including aligning whole genome sequences, aligning reads to a reference genome, and comparing different assemblies of the same genome. Despite its broad utility, MUMmer3 has limitations that can make it difficult to use for large genomes and for the very large sequence data sets that are common today. In this paper we describe MUMmer4, a substantially improved version of MUMmer that addresses genome size constraints by changing the 32-bit suffix tree data structure at the core of MUMmer to a 48-bit suffix array, and that offers improved speed through parallel processing of input query sequences. With a theoretical limit on the input size of 141Tbp, MUMmer4 can now work with input sequences of any biologically realistic length. We show that as a result of these enhancements, the nucmer program in MUMmer4 is easily able to handle alignments of large genomes; we illustrate this with an alignment of the human and chimpanzee genomes, which allows us to compute that the two species are 98% identical across 96% of their length. With the enhancements described here, MUMmer4 can also be used to efficiently align reads to reference genomes, although it is less sensitive and accurate than the dedicated read aligners. The nucmer aligner in MUMmer4 can now be called from scripting languages such as Perl, Python and Ruby. These improvements make MUMer4 one the most versatile genome alignment packages available.
MeSH Terms
Algorithms
Animals
Arabidopsis/genetics
Computational Biology/methods
Genome, Human
Genome, Plant
Genomics
Humans
Models, Theoretical
Pan troglodytes
Polymorphism, Single Nucleotide
Programming Languages
Sequence Alignment/methods
Sequence Analysis, DNA
Sequence Analysis, Protein
Software
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Marçais Guillaume
ORCID
Institute for Physical Science and Technology, University of Maryland, College Park, Maryland, United States of America. | Computational Biology Department, Carnegie Mellon University, Pittsburgh, Pennsylvania, United States of America.
Delcher Arthur L
Center for Computational Biology, Johns Hopkins School of Medicine, Baltimore, Maryland, United States of America.
Phillippy Adam M
National Human Genome Research Institute, Bethesda, Maryland, United States of America.
Coston Rachel
Center for Computational Biology, Johns Hopkins School of Medicine, Baltimore, Maryland, United States of America.
Salzberg Steven L
Center for Computational Biology, Johns Hopkins School of Medicine, Baltimore, Maryland, United States of America. | Departments of Biomedical Engineering, Computer Science, and Biostatistics, Johns Hopkins University, Baltimore, Maryland, United States of America.
Zimin Aleksey
ORCID
Institute for Physical Science and Technology, University of Maryland, College Park, Maryland, United States of America. | Center for Computational Biology, Johns Hopkins School of Medicine, Baltimore, Maryland, United States of America.
References (21)
21 references, click to expand
-
BLAT--the BLAST-like alignment tool.
Genome Res. 2002 Apr;12(4):656-64
PMID: 11932250
-
Fast gapped-read alignment with Bowtie 2.
Nat Methods. 2012 Mar 04;9(4):357-9
PMID: 22388286
-
PBSIM: PacBio reads simulator--toward accurate genome assembly.
Bioinformatics. 2013 Jan 1;29(1):119-21
PMID: 23129296
-
Alignment of whole genomes.
Nucleic Acids Res. 1999 Jun 1;27(11):2369-76
PMID: 10325427
-
Evidence for extensive horizontal gene transfer from the draft genome of a tardigrade.
Proc Natl Acad Sci U S A. 2015 Dec 29;112(52):15976-81
PMID: 26598659
-
No evidence for extensive horizontal gene transfer in the genome of the tardigrade Hypsibius dujardini.
Proc Natl Acad Sci U S A. 2016 May 3;113(18):5053-8
PMID: 27035985
-
Extensive sequencing of seven human genomes to characterize benchmark reference materials.
Sci Data. 2016 Jun 07;3:160025
PMID: 27271295
-
The Sequence Alignment/Map format and SAMtools.
Bioinformatics. 2009 Aug 15;25(16):2078-9
PMID: 19505943
-
Sequence of the Sugar Pine Megagenome.
Genetics. 2016 Dec;204(4):1613-1626
PMID: 27794028
-
Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly.
Genome Res. 2017 May;27(5):849-864
PMID: 28396521
-
Initial sequence of the chimpanzee genome and comparison with the human genome.
Nature. 2005 Sep 1;437(7055):69-87
PMID: 16136131
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
Versatile and open software for comparing large genomes.
Genome Biol. 2004;5(2):R12
PMID: 14759262
-
Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
Genome Biol. 2009;10(3):R25
PMID: 19261174
-
essaMEM: finding maximal exact matches using enhanced sparse suffix arrays.
Bioinformatics. 2013 Mar 15;29(6):802-4
PMID: 23349213
-
Mapping single molecule sequencing reads using basic local alignment with successive refinement (BLASR): application and theory.
BMC Bioinformatics. 2012 Sep 19;13:238
PMID: 22988817
-
The Arabidopsis lyrata genome sequence and the basis of rapid genome size change.
Nat Genet. 2011 May;43(5):476-81
PMID: 21478890
-
Analysis of the genome sequence of the flowering plant Arabidopsis thaliana.
Nature. 2000 Dec 14;408(6814):796-815
PMID: 11130711
-
Fast and accurate short read alignment with Burrows-Wheeler transform.
Bioinformatics. 2009 Jul 15;25(14):1754-60
PMID: 19451168
-
Mauve: multiple alignment of conserved genomic sequence with rearrangements.
Genome Res. 2004 Jul;14(7):1394-403
PMID: 15231754
-
Identification of common molecular subsequences.
J Mol Biol. 1981 Mar 25;147(1):195-7
PMID: 7265238