Home LiteratureArticle Details
PMID: 15491499 Published · epublish English Comparative Study Journal Article Research Support, U.S. Gov't, Non-P.H.S. Research Support, U.S. Gov't, P.H.S. Validation Study

Information assessment on predicting protein-protein interactions.

BMC bioinformatics ·Vol. 5 ·2004-10-18 ·Pages 154

Lin N, Wu B, Jansen R, Gerstein M, Zhao H

Abstract

Identifying protein-protein interactions is fundamental for understanding the molecular machinery of the cell. Proteome-wide studies of protein-protein interactions are of significant value, but the high-throughput experimental technologies suffer from high rates of both false positive and false negative predictions. In addition to high-throughput experimental data, many diverse types of genomic data can help predict protein-protein interactions, such as mRNA expression, localization, essentiality, and functional annotation. Evaluations of the information contributions from different evidences help to establish more parsimonious models with comparable or better prediction accuracy, and to obtain biological insights of the relationships between protein-protein interactions and other genomic information. Our assessment is based on the genomic features used in a Bayesian network approach to predict protein-protein interactions genome-wide in yeast. In the special case, when one does not have any missing information about any of the features, our analysis shows that there is a larger information contribution from the functional-classification than from expression correlations or essentiality. We also show that in this case alternative models, such as logistic regression and random forest, may be more effective than Bayesian networks for predicting interactions. In the restricted problem posed by the complete-information subset, we identified that the MIPS and Gene Ontology (GO) functional similarity datasets as the dominating information contributors for predicting the protein-protein interactions under the framework proposed by Jansen et al. Random forests based on the MIPS and GO information alone can give highly accurate classifications. In this particular subset of complete information, adding other genomic data does little for improving predictions. We also found that the data discretizations used in the Bayesian methods decreased classification performance.

MeSH Terms
Artificial Intelligence Databases, Protein Fungal Proteins/metabolism Genome, Fungal Genomics/methods Logistic Models Models, Genetic Predictive Value of Tests Protein Interaction Mapping/methods,statistics & numerical data Proteome/metabolism
Chemicals
Fungal Proteins Proteome
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Lin Nan
Department of Mathematics, Washington University in St. Louis, St. Louis, MO 63130, USA. [email protected] <[email protected]>
Wu Baolin
Jansen Ronald
Gerstein Mark
Zhao Hongyu
References (14)
14 references, click to expand
  1. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  2. Functional discovery via a compendium of expression profiles.
    Cell. 2000 Jul 7;102(1):109-26 PMID: 10929718
  3. A comprehensive two-hybrid analysis to explore the yeast protein interactome.
    Proc Natl Acad Sci U S A. 2001 Apr 10;98(8):4569-74 PMID: 11283351
  4. MIPS: a database for genomes and protein sequences.
    Nucleic Acids Res. 2002 Jan 1;30(1):31-4 PMID: 11752246
  5. DIP, the Database of Interacting Proteins: a research tool for studying cellular networks of protein interactions.
    Nucleic Acids Res. 2002 Jan 1;30(1):303-5 PMID: 11752321
  6. Functional organization of the yeast proteome by systematic analysis of protein complexes.
    Nature. 2002 Jan 10;415(6868):141-7 PMID: 11805826
  7. Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.
    Proc Natl Acad Sci U S A. 1999 Apr 13;96(8):4285-8 PMID: 10200254
  8. Subcellular localization of the yeast proteome.
    Genes Dev. 2002 Mar 15;16(6):707-19 PMID: 11914276
  9. Assessing experimentally derived interactions in a small world.
    Proc Natl Acad Sci U S A. 2003 Apr 15;100(8):4372-6 PMID: 12676999
  10. A Bayesian networks approach for predicting protein-protein interactions from genomic data.
    Science. 2003 Oct 17;302(5644):449-53 PMID: 14564010
  11. A protein interaction map of Drosophila melanogaster.
    Science. 2003 Dec 5;302(5651):1727-36 PMID: 14605208
  12. Gaining confidence in high-throughput protein interaction networks.
    Nat Biotechnol. 2004 Jan;22(1):78-85 PMID: 14704708
  13. A genome-wide transcriptional analysis of the mitotic cell cycle.
    Mol Cell. 1998 Jul;2(1):65-73 PMID: 9702192
  14. Systematic identification of protein complexes in Saccharomyces cerevisiae by mass spectrometry.
    Nature. 2002 Jan 10;415(6868):180-3 PMID: 11805837
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2004-10-18
Epub
2004-00-18
Pages
154
Language
English
Region
England
NLM ID
100965194
PMCID
PMC529436
Subset
IM
Grants
NIGMS NIH HHS · R01 GM059507 · United States
NIGMS NIH HHS · GM 59507 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]