Abstract
As we are moving into the post genome-sequencing era, various high-throughput experimental techniques have been developed to characterize biological systems on the genomic scale. Discovering new biological knowledge from the high-throughput biological data is a major challenge to bioinformatics today. To address this challenge, we developed a Bayesian statistical method together with Boltzmann machine and simulated annealing for protein functional annotation in the yeast Saccharomyces cerevisiae through integrating various high-throughput biological data, including yeast two-hybrid data, protein complexes and microarray gene expression profiles. In our approach, we quantified the relationship between functional similarity and high-throughput data, and coded the relationship into 'functional linkage graph', where each node represents one protein and the weight of each edge is characterized by the Bayesian probability of function similarity between two proteins. We also integrated the evolution information and protein subcellular localization information into the prediction. Based on our method, 1802 out of 2280 unannotated proteins in yeast were assigned functions systematically.
MeSH Terms
Bayes Theorem
Computational Biology/methods
Gene Expression Profiling
Genome, Fungal
Genomics/methods
Saccharomyces cerevisiae/genetics,metabolism
Saccharomyces cerevisiae Proteins/analysis,genetics,physiology
Two-Hybrid System Techniques
Chemicals
Saccharomyces cerevisiae Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Chen Yu
UT-ORNL Graduate School of Genome Science and Technology, Oak Ridge, TN, USA.
Xu Dong
References (31)
31 references, click to expand
-
Functional organization of the yeast proteome by systematic analysis of protein complexes.
Nature. 2002 Jan 10;415(6868):141-7
PMID: 11805826
-
Predicting protein function from protein/protein interaction data: a probabilistic approach.
Bioinformatics. 2003;19 Suppl 1:i197-204
PMID: 12855458
-
Life with 6000 genes.
Science. 1996 Oct 25;274(5287):546, 563-7
PMID: 8849441
-
Cluster analysis and display of genome-wide expression patterns.
Proc Natl Acad Sci U S A. 1998 Dec 8;95(25):14863-8
PMID: 9843981
-
Improved tools for biological sequence comparison.
Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8
PMID: 3162770
-
Whole-genome annotation by using evidence integration in functional-linkage networks.
Proc Natl Acad Sci U S A. 2004 Mar 2;101(9):2888-93
PMID: 14981259
-
Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.
Proc Natl Acad Sci U S A. 1999 Apr 13;96(8):4285-8
PMID: 10200254
-
The two-hybrid system: a method to identify and clone genes for proteins that interact with a protein of interest.
Proc Natl Acad Sci U S A. 1991 Nov 1;88(21):9578-82
PMID: 1946372
-
Global analysis of protein localization in budding yeast.
Nature. 2003 Oct 16;425(6959):686-91
PMID: 14562095
-
Predicting subcellular localization via protein motif co-occurrence.
Genome Res. 2004 Oct;14(10A):1957-66
PMID: 15466294
-
Genomic expression programs in the response of yeast cells to environmental changes.
Mol Biol Cell. 2000 Dec;11(12):4241-57
PMID: 11102521
-
A Bayesian framework for combining heterogeneous data sources for gene function prediction (in Saccharomyces cerevisiae).
Proc Natl Acad Sci U S A. 2003 Jul 8;100(14):8348-53
PMID: 12826619
-
MIPS: a database for genomes and protein sequences.
Nucleic Acids Res. 2002 Jan 1;30(1):31-4
PMID: 11752246
-
Assessment of prediction accuracy of protein function from protein--protein interaction data.
Yeast. 2001 Apr;18(6):523-31
PMID: 11284008
-
A Bayesian networks approach for predicting protein-protein interactions from genomic data.
Science. 2003 Oct 17;302(5644):449-53
PMID: 14564010
-
Predicting protein complex membership using probabilistic network reliability.
Genome Res. 2004 Jun;14(6):1170-5
PMID: 15140827
-
Systematic identification of protein complexes in Saccharomyces cerevisiae by mass spectrometry.
Nature. 2002 Jan 10;415(6868):180-3
PMID: 11805837
-
Global protein function prediction from protein-protein interaction networks.
Nat Biotechnol. 2003 Jun;21(6):697-700
PMID: 12740586
-
PSORT-B: Improving protein subcellular localization prediction for Gram-negative bacteria.
Nucleic Acids Res. 2003 Jul 1;31(13):3613-7
PMID: 12824378
-
The constraints protein-protein interactions place on sequence divergence.
J Mol Biol. 2002 Nov 29;324(3):399-407
PMID: 12445777
-
Computational analyses of high-throughput protein-protein interaction data.
Curr Protein Pept Sci. 2003 Jun;4(3):159-81
PMID: 12769716
-
Cellular function prediction and biological pathway discovery in Arabidopsis thaliana using microarray data.
Conf Proc IEEE Eng Med Biol Soc. 2004;4:2881-4
PMID: 17270879
-
A comprehensive analysis of protein-protein interactions in Saccharomyces cerevisiae.
Nature. 2000 Feb 10;403(6770):623-7
PMID: 10688190
-
Knowledge-based analysis of microarray gene expression data by using support vector machines.
Proc Natl Acad Sci U S A. 2000 Jan 4;97(1):262-7
PMID: 10618406
-
Understanding protein dispensability through machine-learning analysis of high-throughput data.
Bioinformatics. 2005 Mar 1;21(5):575-81
PMID: 15479713
-
Toward a protein-protein interaction map of the budding yeast: A comprehensive system to examine two-hybrid interactions in all possible combinations between the yeast proteins.
Proc Natl Acad Sci U S A. 2000 Feb 1;97(3):1143-7
PMID: 10655498
-
A combined algorithm for genome-wide prediction of protein function.
Nature. 1999 Nov 4;402(6757):83-6
PMID: 10573421
-
A network of protein-protein interactions in yeast.
Nat Biotechnol. 2000 Dec;18(12):1257-61
PMID: 11101803
-
Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Nat Genet. 2000 May;25(1):25-9
PMID: 10802651
-
Predicting gene function in Saccharomyces cerevisiae.
Bioinformatics. 2003 Oct;19 Suppl 2:ii42-9
PMID: 14534170
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694