Abstract
Phylogenetic patterns show the presence or absence of certain genes or proteins in a set of species. They can also be used to determine sets of genes or proteins that occur only in certain evolutionary branches. Phylogenetic patterns analysis has routinely been applied to protein databases such as COG and OrthoMCL, but not upon gene databases. Here we present a tool named PhyloPat which allows the complete Ensembl gene database to be queried using phylogenetic patterns. PhyloPat is an easy-to-use webserver, which can be used to query the orthologies of all complete genomes within the EnsMart database using phylogenetic patterns. This enables the determination of sets of genes that occur only in certain evolutionary branches or even single species. We found in total 446,825 genes and 3,164,088 orthologous relationships within the EnsMart v40 database. We used a single linkage clustering algorithm to create 147,922 phylogenetic lineages, using every one of the orthologies provided by Ensembl. PhyloPat provides the possibility of querying with either binary phylogenetic patterns (created by checkboxes) or regular expressions. Specific branches of a phylogenetic tree of the 21 included species can be selected to create a branch-specific phylogenetic pattern. Users can also input a list of Ensembl or EMBL IDs to check which phylogenetic lineage any gene belongs to. The output can be saved in HTML, Excel or plain text format for further analysis. A link to the FatiGO web interface has been incorporated in the HTML output, creating easy access to functional information. Finally, lists of omnipresent, polypresent and oligopresent genes have been included. PhyloPat is the first tool to combine complete genome information with phylogenetic pattern querying. Since we used the orthologies generated by the accurate pipeline of Ensembl, the obtained phylogenetic lineages are reliable. The completeness and reliability of these phylogenetic lineages will further increase with the addition of newly found orthologous relationships within each new Ensembl release.
MeSH Terms
Algorithms
Base Sequence
Chromosome Mapping/methods
Conserved Sequence/genetics
Database Management Systems
Databases, Genetic
Eukaryotic Cells/physiology
Information Storage and Retrieval/methods
Molecular Sequence Data
Pattern Recognition, Automated/methods
Phylogeny
Sequence Alignment/methods
Sequence Analysis, DNA/methods
Sequence Homology, Nucleic Acid
Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Hulsen Tim
Centre for Molecular and Biomolecular Informatics (CMBI), Nijmegen Centre for Molecular Life Sciences (NCMLS), Radboud University Nijmegen, Nijmegen, The Netherlands.
[email protected]
de Vlieg Jacob
Groenen Peter M A
References (20)
20 references, click to expand
-
FatiGO: a web tool for finding significant associations of Gene Ontology terms with groups of genes.
Bioinformatics. 2004 Mar 1;20(4):578-80
PMID: 14990455
-
EnsMart: a generic system for fast and flexible access to biological data.
Genome Res. 2004 Jan;14(1):160-9
PMID: 14707178
-
The repertoire of G-protein-coupled receptors in fully sequenced genomes.
Mol Pharmacol. 2005 May;67(5):1414-25
PMID: 15687224
-
Tree pattern matching in phylogenetic trees: automatic search for orthologs or paralogs in homologous gene sequence databases.
Bioinformatics. 2005 Jun 1;21(11):2596-603
PMID: 15713731
-
The HUGO Gene Nomenclature Database, 2006 updates.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D319-21
PMID: 16381876
-
Ensembl 2006.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D556-61
PMID: 16381931
-
TreeFam: a curated database of phylogenetic trees of animal gene families.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D572-80
PMID: 16381935
-
Benchmarking ortholog identification methods using functional genomics data.
Genome Biol. 2006;7(4):R31
PMID: 16613613
-
A phylogenomic gene cluster resource: the Phylogenetically Inferred Groups (PhIGs) database.
BMC Bioinformatics. 2006;7:201
PMID: 16608522
-
No more than 14: the end of the amphioxus Hox cluster.
Int J Biol Sci. 2005;1(1):19-23
PMID: 15951846
-
OrthoMCL-DB: querying a comprehensive multi-species collection of ortholog groups.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D363-8
PMID: 16381887
-
Database resources of the National Center for Biotechnology Information.
Nucleic Acids Res. 2000 Jan 1;28(1):10-4
PMID: 10592169
-
Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Nat Genet. 2000 May;25(1):25-9
PMID: 10802651
-
Using the COG database to improve gene recognition in complete genomes.
Genetica. 2000;108(1):9-17
PMID: 11145426
-
SHOT: a web server for the construction of genome phylogenies.
Trends Genet. 2002 Mar;18(3):158-62
PMID: 11858840
-
TRANSFAC: transcriptional regulation, from patterns to profiles.
Nucleic Acids Res. 2003 Jan 1;31(1):374-8
PMID: 12520026
-
EPPS: mining the COG database by an extended phylogenetic patterns search.
Bioinformatics. 2003 Apr 12;19(6):784-5
PMID: 12691996
-
A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
Syst Biol. 2003 Oct;52(5):696-704
PMID: 14530136
-
Hox cluster duplications and the opportunity for evolutionary novelties.
Proc Natl Acad Sci U S A. 2003 Dec 9;100(25):14603-6
PMID: 14638945
-
MUSCLE: a multiple sequence alignment method with reduced time and space complexity.
BMC Bioinformatics. 2004 Aug 19;5:113
PMID: 15318951