Home LiteratureArticle Details
PMID: 17597916 Published · epublish English Journal Article

Functional annotation of hypothetical proteins - A review.

Bioinformation ·Vol. 1 ·No. 8 ·2006-12-29 ·Pages 335-8

Sivashankari S, Shanmughavel P

Abstract

The complete human genome sequences in the public database provide ways to understand the blue print of life. As of June 29, 2006, 27 archaeal, 326 bacterial and 21 eukaryotes is complete genomes are available and the sequencing for 316 bacterial, 24 archaeal, 126 eukaryotic genomes are in progress. The traditional biochemical/molecular experiments can assign accurate functions for genes in these genomes. However, the process is time-consuming and costly. Despite several efforts, only 50-60 % of genes have been annotated in most completely sequenced genomes. Automated genome sequence analysis and annotation may provide ways to understand genomes. Thus, determination of protein function is one of the challenging problems of the post-genome era. This demands bioinformatics to predict functions of un-annotated protein sequences by developing efficient tools. Here, we discuss some of the recent and popular approaches developed in Bioinformatics to predict functions for hypothetical proteins.

Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Sivashankari Selvarajan
Department of Bioinformatics, Kongunadu arts and science college, Coimbatore - 641029, India.
Shanmughavel Piramanayagam
References (31)
31 references, click to expand
  1. Coexpression analysis of human genes across many microarray data sets.
    Genome Res. 2004 Jun;14(6):1085-94 PMID: 15173114
  2. STRING: a database of predicted functional associations between proteins.
    Nucleic Acids Res. 2003 Jan 1;31(1):258-61 PMID: 12519996
  3. HemK, a class of protein methyl transferase with similarity to DNA methyl transferases, methylates polypeptide chain release factors, and hemK knockout induces defects in translational termination.
    Proc Natl Acad Sci U S A. 2002 Feb 5;99(3):1473-8 PMID: 11805295
  4. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  5. Detecting protein function and protein-protein interactions from genome sequences.
    Science. 1999 Jul 30;285(5428):751-3 PMID: 10427000
  6. Pigeon-liver NAD kinase. The structural and kinetic basis of regulation of NADPH.
    Eur J Biochem. 1975 Jul 1;55(2):475-83 PMID: 239
  7. A network of protein-protein interactions in yeast.
    Nat Biotechnol. 2000 Dec;18(12):1257-61 PMID: 11101803
  8. Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.
    Proc Natl Acad Sci U S A. 1999 Apr 13;96(8):4285-8 PMID: 10200254
  9. The hemK gene in Escherichia coli encodes the N(5)-glutamine methyltransferase that modifies peptide release factors.
    EMBO J. 2002 Feb 15;21(4):769-78 PMID: 11847124
  10. Assessment of prediction accuracy of protein function from protein--protein interaction data.
    Yeast. 2001 Apr;18(6):523-31 PMID: 11284008
  11. Errors in genome annotation.
    Trends Genet. 1999 Apr;15(4):132-3 PMID: 10203816
  12. Biogenesis of Fe-S cluster by the bacterial Suf system: SufS and SufE form a new type of cysteine desulfurase.
    J Biol Chem. 2003 Oct 3;278(40):38352-9 PMID: 12876288
  13. Predicting protein function by genomic context: quantitative evaluation and qualitative inferences.
    Genome Res. 2000 Aug;10(8):1204-10 PMID: 10958638
  14. Iterated profile searches with PSI-BLAST--a tool for discovery in protein databases.
    Trends Biochem Sci. 1998 Nov;23(11):444-7 PMID: 9852764
  15. The methylator meets the terminator.
    Proc Natl Acad Sci U S A. 2002 Feb 5;99(3):1104-6 PMID: 11830650
  16. Prediction of protein function using protein-protein interaction data.
    J Comput Biol. 2003;10(6):947-60 PMID: 14980019
  17. The COG database: new developments in phylogenetic classification of proteins from complete genomes.
    Nucleic Acids Res. 2001 Jan 1;29(1):22-8 PMID: 11125040
  18. Rosetta Stone proteins: "chance and necessity"?
    Genome Biol. 2002;3(2):INTERACTIONS1001 PMID: 11864366
  19. Prediction of protein function using protein-protein interaction data.
    Proc IEEE Comput Soc Bioinform Conf. 2002;1:197-206 PMID: 15838136
  20. The Protein Data Bank: a computer-based archival file for macromolecular structures.
    J Mol Biol. 1977 May 25;112(3):535-42 PMID: 875032
  21. Cluster analysis and display of genome-wide expression patterns.
    Proc Natl Acad Sci U S A. 1998 Dec 8;95(25):14863-8 PMID: 9843981
  22. Identification and functional analysis of 'hypothetical' genes expressed in Haemophilus influenzae.
    Nucleic Acids Res. 2004 Apr 30;32(8):2353-61 PMID: 15121896
  23. Isolation and characterization of yeast nicotinamide adenine dinucleotide kinase.
    Biochim Biophys Acta. 1979 May 10;568(1):205-14 PMID: 221029
  24. 'Conserved hypothetical' proteins: prioritization of targets for experimental study.
    Nucleic Acids Res. 2004 Oct 12;32(18):5452-63 PMID: 15479782
  25. Visualizing associations between genome sequences and gene expression data using genome-mean expression profiles.
    Bioinformatics. 2001;17 Suppl 1:S49-55 PMID: 11472992
  26. Powers and pitfalls in sequence analysis: the 70% hurdle.
    Genome Res. 2000 Apr;10(4):398-400 PMID: 10779480
  27. Structure-based assignment of the biochemical function of a hypothetical protein: a test case of structural genomics.
    Proc Natl Acad Sci U S A. 1998 Dec 22;95(26):15189-93 PMID: 9860944
  28. The use of gene clusters to infer functional coupling.
    Proc Natl Acad Sci U S A. 1999 Mar 16;96(6):2896-901 PMID: 10077608
  29. Molecular characterization of Escherichia coli NAD kinase.
    Eur J Biochem. 2001 Aug;268(15):4359-65 PMID: 11488932
  30. A Bayesian framework for combining heterogeneous data sources for gene function prediction (in Saccharomyces cerevisiae).
    Proc Natl Acad Sci U S A. 2003 Jul 8;100(14):8348-53 PMID: 12826619
  31. Prediction of protein function from protein sequence and structure.
    Q Rev Biophys. 2003 Aug;36(3):307-40 PMID: 15029827
Article Info
Journal
Bioinformation
Abbr.
Bioinformation
ISSN
0973-2063
Published
2006-12-29
Epub
2006-00-29
Pages
335-8
Language
English
Region
Singapore
NLM ID
101258255
PMCID
PMC1891709
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]