Home LiteratureArticle Details
PMID: 20534164 Published · epublish English Journal Article Research Support, N.I.H., Extramural

GIGA: a simple, efficient algorithm for gene tree inference in the genomic age.

BMC bioinformatics ·Vol. 11 ·2010-06-09 ·Pages 312

Thomas PD

Abstract

Phylogenetic relationships between genes are not only of theoretical interest: they enable us to learn about human genes through the experimental work on their relatives in numerous model organisms from bacteria to fruit flies and mice. Yet the most commonly used computational algorithms for reconstructing gene trees can be inaccurate for numerous reasons, both algorithmic and biological. Additional information beyond gene sequence data has been shown to improve the accuracy of reconstructions, though at great computational cost. We describe a simple, fast algorithm for inferring gene phylogenies, which makes use of information that was not available prior to the genomic age: namely, a reliable species tree spanning much of the tree of life, and knowledge of the complete complement of genes in a species' genome. The algorithm, called GIGA, constructs trees agglomeratively from a distance matrix representation of sequences, using simple rules to incorporate this genomic age information. GIGA makes use of a novel conceptualization of gene trees as being composed of orthologous subtrees (containing only speciation events), which are joined by other evolutionary events such as gene duplication or horizontal gene transfer. An important innovation in GIGA is that, at every step in the agglomeration process, the tree is interpreted/reinterpreted in terms of the evolutionary events that created it. Remarkably, GIGA performs well even when using a very simple distance metric (pairwise sequence differences) and no distance averaging over clades during the tree construction process. GIGA is efficient, allowing phylogenetic reconstruction of very large gene families and determination of orthologs on a large scale. It is exceptionally robust to adding more gene sequences, opening up the possibility of creating stable identifiers for referring to not only extant genes, but also their common ancestors. We compared trees produced by GIGA to those in the TreeFam database, and they were very similar in general, with most differences likely due to poor alignment quality. However, some remaining differences are algorithmic, and can be explained by the fact that GIGA tends to put a larger emphasis on minimizing gene duplication and deletion events.

MeSH Terms
Algorithms Animals Base Sequence Evolution, Molecular Gene Duplication Gene Transfer, Horizontal Genome Humans Mice Phylogeny
Authors & Affiliations
1 authors, click to expand affiliations / ORCID
Thomas Paul D
Evolutionary Systems Biology Group, SRI International, Menlo Park, CA, USA. [email protected]
References (38)
38 references, click to expand
  1. Construction of phylogenetic trees for proteins and nucleic acids: empirical evaluation of alternative matrix methods.
    J Mol Evol. 1978 Jun 20;11(2):129-42 PMID: 671561
  2. The net of life: reconstructing the microbial phylogenetic network.
    Genome Res. 2005 Jul;15(7):954-9 PMID: 15965028
  3. Predicting protein structure using hidden Markov models.
    Proteins. 1997;Suppl 1:134-9 PMID: 9485505
  4. Phylogenetic inference using whole genomes.
    Annu Rev Genomics Hum Genet. 2008;9:217-31 PMID: 18767964
  5. A hybrid micro-macroevolutionary approach to gene tree reconstruction.
    J Comput Biol. 2006 Mar;13(2):320-35 PMID: 16597243
  6. A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
    Syst Biol. 2003 Oct;52(5):696-704 PMID: 14530136
  7. Bayesian inference of phylogeny and its impact on evolutionary biology.
    Science. 2001 Dec 14;294(5550):2310-4 PMID: 11743192
  8. The Gene Ontology's Reference Genome Project: a unified framework for functional annotation across species.
    PLoS Comput Biol. 2009 Jul;5(7):e1000431 PMID: 19578431
  9. EGASP: the human ENCODE Genome Annotation Assessment Project.
    Genome Biol. 2006;7 Suppl 1:S2.1-31 PMID: 16925836
  10. Inferring phylogeny despite incomplete lineage sorting.
    Syst Biol. 2006 Feb;55(1):21-30 PMID: 16507521
  11. Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.
    Proc Natl Acad Sci U S A. 1999 Apr 13;96(8):4285-8 PMID: 10200254
  12. Phylogenetic and functional assessment of orthologs inference projects and methods.
    PLoS Comput Biol. 2009 Jan;5(1):e1000262 PMID: 19148271
  13. Automatic genome-wide reconstruction of phylogenetic gene trees.
    Bioinformatics. 2007 Jul 1;23(13):i549-58 PMID: 17646342
  14. The COG database: an updated version includes eukaryotes.
    BMC Bioinformatics. 2003 Sep 11;4:41 PMID: 12969510
  15. Model-based prediction of sequence alignment quality.
    Bioinformatics. 2008 Oct 1;24(19):2165-71 PMID: 18678587
  16. Detecting non-orthology in the COGs database and other approaches grouping orthologs using genome-specific best hits.
    Nucleic Acids Res. 2006 Jul 11;34(11):3309-16 PMID: 16835308
  17. EnsemblCompara GeneTrees: Complete, duplication-aware phylogenetic trees in vertebrates.
    Genome Res. 2009 Feb;19(2):327-35 PMID: 19029536
  18. PANTHER version 7: improved phylogenetic trees, orthologs and collaboration with the Gene Ontology Consortium.
    Nucleic Acids Res. 2010 Jan;38(Database issue):D204-10 PMID: 20015972
  19. Phylogenetic origins and adaptive evolution of avian and mammalian haemoglobin genes.
    Nature. 1982 Jul 15;298(5871):297-300 PMID: 6178039
  20. NOTUNG: a program for dating gene duplications and optimizing gene family trees.
    J Comput Biol. 2000;7(3-4):429-47 PMID: 11108472
  21. Inferring phylogenetic networks by the maximum parsimony criterion: a case study.
    Mol Biol Evol. 2007 Jan;24(1):324-37 PMID: 17068107
  22. The neighbor-joining method: a new method for reconstructing phylogenetic trees.
    Mol Biol Evol. 1987 Jul;4(4):406-25 PMID: 3447015
  23. Proof and evolutionary analysis of ancient genome duplication in the yeast Saccharomyces cerevisiae.
    Nature. 2004 Apr 8;428(6983):617-24 PMID: 15004568
  24. PAML 4: phylogenetic analysis by maximum likelihood.
    Mol Biol Evol. 2007 Aug;24(8):1586-91 PMID: 17483113
  25. The altered evolutionary trajectories of gene duplicates.
    Trends Genet. 2004 Nov;20(11):544-9 PMID: 15475113
  26. PhylomeDB: a database for genome-wide collections of gene phylogenies.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D491-6 PMID: 17962297
  27. Ribosomal RNA: a key to phylogeny.
    FASEB J. 1993 Jan;7(1):113-23 PMID: 8422957
  28. GeneTrees: a phylogenomics resource for prokaryotes.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D328-31 PMID: 17151073
  29. The tree versus the forest: the fungal tree of life and the topological diversity within the yeast phylome.
    PLoS One. 2009;4(2):e4357 PMID: 19190756
  30. TreeFam: 2008 Update.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D735-40 PMID: 18056084
  31. Widespread discordance of gene trees with species tree in Drosophila: evidence for incomplete lineage sorting.
    PLoS Genet. 2006 Oct 27;2(10):e173 PMID: 17132051
  32. Descent of mammalian alpha globin chain sequences investigated by the maximum parsimony method.
    J Mol Biol. 1972 Aug 21;69(2):249-78 PMID: 4627161
  33. Accurate gene-tree reconstruction by learning gene- and species-specific substitution rates across multiple complete genomes.
    Genome Res. 2007 Dec;17(12):1932-42 PMID: 17989260
  34. Optimal gene trees from sequences and species trees using a soft interpretation of parsimony.
    J Mol Evol. 2006 Aug;63(2):240-50 PMID: 16830091
  35. Phylogenetic identification of lateral genetic transfer events.
    BMC Evol Biol. 2006 Feb 11;6:15 PMID: 16472400
  36. Inferring trees.
    Methods Mol Biol. 2008;452:287-309 PMID: 18566770
  37. The rapid generation of mutation data matrices from protein sequences.
    Comput Appl Biosci. 1992 Jun;8(3):275-82 PMID: 1633570
  38. nGASP--the nematode genome annotation assessment project.
    BMC Bioinformatics. 2008 Dec 19;9:549 PMID: 19099578
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2010-06-09
Epub
2010-00-09
Pages
312
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2905364
Subset
IM
Grants
NIGMS NIH HHS · R01GM081084 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]