Home LiteratureArticle Details
PMID: 16613613 Published · ppublish English Evaluation Study Journal Article Research Support, Non-U.S. Gov't

Benchmarking ortholog identification methods using functional genomics data.

Genome biology ·Vol. 7 ·No. 4 ·2006-00-00 ·Pages R31

Hulsen T, Huynen MA, de Vlieg J, Groenen PM

Abstract

The transfer of functional annotations from model organism proteins to human proteins is one of the main applications of comparative genomics. Various methods are used to analyze cross-species orthologous relationships according to an operational definition of orthology. Often the definition of orthology is incorrectly interpreted as a prediction of proteins that are functionally equivalent across species, while in fact it only defines the existence of a common ancestor for a gene in different species. However, it has been demonstrated that orthologs often reveal significant functional similarity. Therefore, the quality of the orthology prediction is an important factor in the transfer of functional annotations (and other related information). To identify protein pairs with the highest possible functional similarity, it is important to qualify ortholog identification methods. To measure the similarity in function of proteins from different species we used functional genomics data, such as expression data and protein interaction data. We tested several of the most popular ortholog identification methods. In general, we observed a sensitivity/selectivity trade-off: the functional similarity scores per orthologous pair of sequences become higher when the number of proteins included in the ortholog groups decreases. By combining the sensitivity and the selectivity into an overall score, we show that the InParanoid program is the best ortholog identification method in terms of identifying functionally equivalent proteins.

MeSH Terms
Algorithms Animals Databases, Genetic Evolution, Molecular Gene Expression Gene Order Genomics/methods Humans Mice Multigene Family Protein Interaction Mapping Proteins/genetics,metabolism,physiology Sequence Homology, Amino Acid Software
Chemicals
Proteins
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Hulsen Tim
Centre for Molecular and Biomolecular Informatics, Radboud University Nijmegen, Toernooiveld 1, Nijmegen, 6500 GL, The Netherlands. [email protected]
Huynen Martijn A
de Vlieg Jacob
Groenen Peter M A
References (36)
36 references, click to expand
  1. Structural divergence and distant relationships in proteins: evolution of the globins.
    Curr Opin Struct Biol. 2005 Jun;15(3):290-301 PMID: 15922591
  2. Progress in medical information management. Systematized nomenclature of medicine (SNOMED).
    JAMA. 1980 Feb 22-29;243(8):756-62 PMID: 6986000
  3. The neighbor-joining method: a new method for reconstructing phylogenetic trees.
    Mol Biol Evol. 1987 Jul;4(4):406-25 PMID: 3447015
  4. NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4 PMID: 15608248
  5. The InterPro Database, 2003 brings increased coverage and new features.
    Nucleic Acids Res. 2003 Jan 1;31(1):315-8 PMID: 12520011
  6. DIP, the Database of Interacting Proteins: a research tool for studying cellular networks of protein interactions.
    Nucleic Acids Res. 2002 Jan 1;30(1):303-5 PMID: 11752321
  7. Coevolution of gene expression among interacting proteins.
    Proc Natl Acad Sci U S A. 2004 Jun 15;101(24):9033-8 PMID: 15175431
  8. Expression divergence between duplicate genes.
    Trends Genet. 2005 Nov;21(11):602-7 PMID: 16140417
  9. EnsMart: a generic system for fast and flexible access to biological data.
    Genome Res. 2004 Jan;14(1):160-9 PMID: 14707178
  10. Assessing sequence comparison methods with reliable structurally identified distant evolutionary relationships.
    Proc Natl Acad Sci U S A. 1998 May 26;95(11):6073-8 PMID: 9600919
  11. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  12. RIO: analyzing proteomes by automated phylogenomics using resampled inference of orthologs.
    BMC Bioinformatics. 2002 May 16;3:14 PMID: 12028595
  13. Using orthologous and paralogous proteins to identify specificity determining residues.
    Genome Biol. 2002;3(3):PREPRINT0002 PMID: 11897020
  14. Predicting gene function by conserved co-expression.
    Trends Genet. 2003 May;19(5):238-42 PMID: 12711213
  15. Comparative genomics for reliable protein-function prediction from genomic data.
    Trends Genet. 2004 Aug;20(8):340-4 PMID: 15262404
  16. Automatic clustering of orthologs and in-paralogs from pairwise species comparisons.
    J Mol Biol. 2001 Dec 14;314(5):1041-52 PMID: 11743721
  17. The COG database: an updated version includes eukaryotes.
    BMC Bioinformatics. 2003 Sep 11;4:41 PMID: 12969510
  18. The Gene Ontology (GO) database and informatics resource.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D258-61 PMID: 14681407
  19. Evidence for 14 homeobox gene clusters in human genome ancestry.
    Curr Biol. 2000 Sep 7;10(17):1059-62 PMID: 10996074
  20. Significance of Z-value statistics of Smith-Waterman scores for protein alignments.
    Comput Chem. 1999 Jun 15;23(3-4):317-31 PMID: 10627144
  21. Phylogenomic inference of protein molecular function: advances and challenges.
    Bioinformatics. 2004 Jan 22;20(2):170-9 PMID: 14734307
  22. Expression and function of conserved nuclear receptor genes in Caenorhabditis elegans.
    Dev Biol. 2004 Feb 15;266(2):399-416 PMID: 14738886
  23. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  24. Sm and Sm-like proteins assemble in two related complexes of deep evolutionary origin.
    EMBO J. 1999 Jun 15;18(12):3451-62 PMID: 10369684
  25. The Ensembl automatic gene annotation system.
    Genome Res. 2004 May;14(5):942-50 PMID: 15123590
  26. OrthoMCL: identification of ortholog groups for eukaryotic genomes.
    Genome Res. 2003 Sep;13(9):2178-89 PMID: 12952885
  27. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.
    Nucleic Acids Res. 2003 Jan 1;31(1):365-70 PMID: 12520024
  28. Distinguishing homologous from analogous proteins.
    Syst Zool. 1970 Jun;19(2):99-113 PMID: 5449325
  29. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  30. The COG database: a tool for genome-scale analysis of protein functions and evolution.
    Nucleic Acids Res. 2000 Jan 1;28(1):33-6 PMID: 10592175
  31. An efficient algorithm for large-scale detection of protein families.
    Nucleic Acids Res. 2002 Apr 1;30(7):1575-84 PMID: 11917018
  32. HCOP: the HGNC comparison of orthology predictions search tool.
    Mamm Genome. 2005 Nov;16(11):827-8 PMID: 16284797
  33. OrthoMCL-DB: querying a comprehensive multi-species collection of ortholog groups.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D363-8 PMID: 16381887
  34. A gene-coexpression network for global discovery of conserved genetic modules.
    Science. 2003 Oct 10;302(5643):249-55 PMID: 12934013
  35. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  36. Measuring genome evolution.
    Proc Natl Acad Sci U S A. 1998 May 26;95(11):5849-56 PMID: 9600883
Article Info
Journal
Genome biology
Abbr.
Genome Biol
ISSN
1474-760X
Published
2006-00-00
Epub
2006-00-13
Pages
R31
Language
English
Region
England
NLM ID
100960660
PMCID
PMC1557999
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]