Home LiteratureArticle Details
PMID: 22192575 Published · epublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

MAKER2: an annotation pipeline and genome-database management tool for second-generation genome projects.

BMC bioinformatics ·Vol. 12 ·2011-12-22 ·Pages 491

Holt C, Yandell M

Abstract

Second-generation sequencing technologies are precipitating major shifts with regards to what kinds of genomes are being sequenced and how they are annotated. While the first generation of genome projects focused on well-studied model organisms, many of today's projects involve exotic organisms whose genomes are largely terra incognita. This complicates their annotation, because unlike first-generation projects, there are no pre-existing 'gold-standard' gene-models with which to train gene-finders. Improvements in genome assembly and the wide availability of mRNA-seq data are also creating opportunities to update and re-annotate previously published genome annotations. Today's genome projects are thus in need of new genome annotation tools that can meet the challenges and opportunities presented by second-generation sequencing technologies. We present MAKER2, a genome annotation and data management tool designed for second-generation genome projects. MAKER2 is a multi-threaded, parallelized application that can process second-generation datasets of virtually any size. We show that MAKER2 can produce accurate annotations for novel genomes where training-data are limited, of low quality or even non-existent. MAKER2 also provides an easy means to use mRNA-seq data to improve annotation quality; and it can use these data to update legacy annotations, significantly improving their quality. We also show that MAKER2 can evaluate the quality of genome annotations, and identify and prioritize problematic annotations for manual review. MAKER2 is the first annotation engine specifically designed for second-generation genome projects. MAKER2 scales to datasets of any size, requires little in the way of training data, and can use mRNA-seq data to improve annotation quality. It can also update and manage legacy genome annotation datasets.

MeSH Terms
Animals Databases, Genetic Genome High-Throughput Nucleotide Sequencing/methods Humans Molecular Sequence Annotation Plants/genetics Software
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Holt Carson
Eccles Institute of Human Genetics, University of Utah, Salt Lake City, Utah 84112, USA.
Yandell Mark
References (57)
57 references, click to expand
  1. Draft genome of the red harvester ant Pogonomyrmex barbatus.
    Proc Natl Acad Sci U S A. 2011 Apr 5;108(14):5667-72 PMID: 21282651
  2. NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D61-5 PMID: 17130148
  3. The Sequence Ontology: a tool for the unification of genome annotations.
    Genome Biol. 2005;6(5):R44 PMID: 15892872
  4. The genome of the fire ant Solenopsis invicta.
    Proc Natl Acad Sci U S A. 2011 Apr 5;108(14):5679-84 PMID: 21282665
  5. GenBank.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D21-5 PMID: 17202161
  6. Eval: a software package for analysis of genome annotations.
    BMC Bioinformatics. 2003 Oct 17;4:50 PMID: 14565849
  7. Comparison of mouse and human genomes followed by experimental verification yields an estimated 1,019 additional genes.
    Proc Natl Acad Sci U S A. 2003 Feb 4;100(3):1140-5 PMID: 12552088
  8. Pfam: clans, web tools and services.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D247-51 PMID: 16381856
  9. Quantitative measures for the management and comparison of annotated genomes.
    BMC Bioinformatics. 2009 Feb 23;10:67 PMID: 19236712
  10. Hymenoptera Genome Database: integrated community resources for insect species of the order Hymenoptera.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D658-62 PMID: 21071397
  11. Strong functional patterns in the evolution of eukaryotic genomes revealed by the reconstruction of ancestral protein domain repertoires.
    Genome Biol. 2011;12(1):R4 PMID: 21241503
  12. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  13. Genome sequence of the necrotrophic plant pathogen Pythium ultimum reveals original pathogenicity mechanisms and effector repertoire.
    Genome Biol. 2010;11(7):R73 PMID: 20626842
  14. Analysis of the genome sequence of the flowering plant Arabidopsis thaliana.
    Nature. 2000 Dec 14;408(6814):796-815 PMID: 11130711
  15. AphidBase: a centralized bioinformatic resource for annotation of the pea aphid genome.
    Insect Mol Biol. 2010 Mar;19 Suppl 2:5-12 PMID: 20482635
  16. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  17. The Pinus taeda genome is characterized by diverse and highly diverged repetitive sequences.
    BMC Genomics. 2010 Jul 07;11:420 PMID: 20609256
  18. Repbase Update, a database of eukaryotic repetitive elements.
    Cytogenet Genome Res. 2005;110(1-4):462-7 PMID: 16093699
  19. A draft sequence of the rice genome (Oryza sativa L. ssp. japonica).
    Science. 2002 Apr 5;296(5565):92-100 PMID: 11935018
  20. Genomic hotspots for adaptation: the population genetics of Müllerian mimicry in the Heliconius melpomene clade.
    PLoS Genet. 2010 Feb 05;6(2):e1000794 PMID: 20140188
  21. InterProScan: protein domains identifier.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W116-20 PMID: 15980438
  22. SmedGD: the Schmidtea mediterranea genome database.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D599-606 PMID: 17881371
  23. Galaxy: a platform for interactive large-scale genome analysis.
    Genome Res. 2005 Oct;15(10):1451-5 PMID: 16169926
  24. Gene identification in novel eukaryotic genomes by self-training algorithm.
    Nucleic Acids Res. 2005 Nov 28;33(20):6494-506 PMID: 16314312
  25. The genome sequence of the malaria mosquito Anopheles gambiae.
    Science. 2002 Oct 4;298(5591):129-49 PMID: 12364791
  26. Gene prediction with a hidden Markov model and a new intron submodel.
    Bioinformatics. 2003 Oct;19 Suppl 2:ii215-25 PMID: 14534192
  27. Insights into social insects from the genome of the honeybee Apis mellifera.
    Nature. 2006 Oct 26;443(7114):931-49 PMID: 17073008
  28. MAKER: an easy-to-use annotation pipeline designed for emerging model organism genomes.
    Genome Res. 2008 Jan;18(1):188-96 PMID: 18025269
  29. The genome sequence of the leaf-cutter ant Atta cephalotes reveals insights into its obligate symbiotic lifestyle.
    PLoS Genet. 2011 Feb 10;7(2):e1002007 PMID: 21347285
  30. Evaluation of gene structure prediction programs.
    Genomics. 1996 Jun 15;34(3):353-67 PMID: 8786136
  31. Functional and evolutionary insights from the genomes of three parasitoid Nasonia species.
    Science. 2010 Jan 15;327(5963):343-8 PMID: 20075255
  32. CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes.
    Bioinformatics. 2007 May 1;23(9):1061-7 PMID: 17332020
  33. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  34. Transcriptomic responses of the softwood-degrading white-rot fungus Phanerochaete carnosa during growth on coniferous and deciduous wood.
    Appl Environ Microbiol. 2011 May;77(10):3211-8 PMID: 21441342
  35. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  36. Genomic comparison of the ants Camponotus floridanus and Harpegnathos saltator.
    Science. 2010 Aug 27;329(5995):1068-71 PMID: 20798317
  37. Draft genome of the globally widespread and invasive Argentine ant (Linepithema humile).
    Proc Natl Acad Sci U S A. 2011 Apr 5;108(14):5673-8 PMID: 21282631
  38. dbEST--database for "expressed sequence tags".
    Nat Genet. 1993 Aug;4(4):332-3 PMID: 8401577
  39. The sequence of the human genome.
    Science. 2001 Feb 16;291(5507):1304-51 PMID: 11181995
  40. Detailed analysis of a contiguous 22-Mb region of the maize genome.
    PLoS Genet. 2009 Nov;5(11):e1000728 PMID: 19936048
  41. Genome sequence of the nematode C. elegans: a platform for investigating biology.
    Science. 1998 Dec 11;282(5396):2012-8 PMID: 9851916
  42. The genome sequence of Drosophila melanogaster.
    Science. 2000 Mar 24;287(5461):2185-95 PMID: 10731132
  43. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  44. EGASP: the human ENCODE Genome Annotation Assessment Project.
    Genome Biol. 2006;7 Suppl 1:S2.1-31 PMID: 16925836
  45. Ongoing and future developments at the Universal Protein Resource.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D214-9 PMID: 21051339
  46. Sequencing, mapping, and analysis of 27,455 maize full-length cDNAs.
    PLoS Genet. 2009 Nov;5(11):e1000740 PMID: 19936069
  47. A Chado case study: an ontology-based modular schema for representing genome-associated biological information.
    Bioinformatics. 2007 Jul 1;23(13):i337-46 PMID: 17646315
  48. The genome of the blood fluke Schistosoma mansoni.
    Nature. 2009 Jul 16;460(7253):352-8 PMID: 19606141
  49. The generic genome browser: a building block for a model organism system database.
    Genome Res. 2002 Oct;12(10):1599-610 PMID: 12368253
  50. Comparative genomics suggests that the fungal pathogen pneumocystis is an obligate parasite scavenging amino acids from its host's lungs.
    PLoS One. 2010 Dec 20;5(12):e15152 PMID: 21188143
  51. The genome sequence of Caenorhabditis briggsae: a platform for comparative genomics.
    PLoS Biol. 2003 Nov;1(2):E45 PMID: 14624247
  52. Characterization of a hotspot for mimicry: assembly of a butterfly wing transcriptome to genomic sequence at the HmYb/Sb locus.
    Mol Ecol. 2010 Mar;19 Suppl 1:240-54 PMID: 20331783
  53. Nematode.net update 2008: improvements enabling more efficient data mining and comparative nematode genomics.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D571-8 PMID: 18940860
  54. Genetic and physical maps of Saccharomyces cerevisiae.
    Nature. 1997 May 29;387(6632 Suppl):67-73 PMID: 9169866
  55. Gene finding in novel genomes.
    BMC Bioinformatics. 2004 May 14;5:59 PMID: 15144565
  56. nGASP--the nematode genome annotation assessment project.
    BMC Bioinformatics. 2008 Dec 19;9:549 PMID: 19099578
  57. Initial sequencing and comparative analysis of the mouse genome.
    Nature. 2002 Dec 5;420(6915):520-62 PMID: 12466850
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2011-12-22
Epub
2011-00-22
Pages
491
Language
English
Region
England
NLM ID
100965194
PMCID
PMC3280279
Subset
IM
Grants
NHGRI NIH HHS · R01-HG004694 · United States
NIGMS NIH HHS · T32-GM007464 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]