Home LiteratureArticle Details
PMID: 15576349 Published · epublish English Evaluation Study Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

EFICAz: a comprehensive approach for accurate genome-scale enzyme function inference.

Nucleic acids research ·Vol. 32 ·No. 21 ·2004-00-00 ·Pages 6226-39

Tian W, Arakaki AK, Skolnick J

Abstract

EFICAz (Enzyme Function Inference by Combined Approach) is an automatic engine for large-scale enzyme function inference that combines predictions from four different methods developed and optimized to achieve high prediction accuracy: (i) recognition of functionally discriminating residues (FDRs) in enzyme families obtained by a Conservation-controlled HMM Iterative procedure for Enzyme Family classification (CHIEFc), (ii) pairwise sequence comparison using a family specific Sequence Identity Threshold, (iii) recognition of FDRs in Multiple Pfam enzyme families, and (iv) recognition of multiple Prosite patterns of high specificity. For FDR (i.e. conserved positions in an enzyme family that discriminate between true and false members of the family) identification, we have developed an Evolutionary Footprinting method that uses evolutionary information from homofunctional and heterofunctional multiple sequence alignments associated with an enzyme family. The FDRs show a significant correlation with annotated active site residues. In a jackknife test, EFICAz shows high accuracy (92%) and sensitivity (82%) for predicting four EC digits in testing sequences that are <40% identical to any member of the corresponding training set. Applied to Escherichia coli genome, EFICAz assigns more detailed enzymatic function than KEGG, and generates numerous novel predictions.

MeSH Terms
Amino Acid Sequence Binding Sites Computational Biology/methods Databases, Protein Enzymes/classification,genetics,physiology Escherichia coli/enzymology,genetics Genome, Bacterial Genomics/methods Software
Chemicals
Enzymes
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Tian Weidong
Center of Excellence in Bioinformatics, University at Buffalo, The State University of New York,901 Washington Street, Buffalo, NY 14203-1199, USA.
Arakaki Adrian K
Skolnick Jeffrey
References (43)
43 references, click to expand
  1. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  2. Local alignment statistics.
    Methods Enzymol. 1996;266:460-80 PMID: 8743700
  3. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  4. Errors in genome annotation.
    Trends Genet. 1999 Apr;15(4):132-3 PMID: 10203816
  5. From protein structure to function.
    Curr Opin Struct Biol. 1999 Jun;9(3):374-82 PMID: 10361094
  6. Flexible sequence similarity searching with the FASTA3 program package.
    Methods Mol Biol. 2000;132:185-219 PMID: 10547837
  7. KEGG: kyoto encyclopedia of genes and genomes.
    Nucleic Acids Res. 2000 Jan 1;28(1):27-30 PMID: 10592173
  8. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  9. PRINTS-S: the database formerly known as PRINTS.
    Nucleic Acids Res. 2000 Jan 1;28(1):225-7 PMID: 10592232
  10. Increased coverage of protein families with the blocks database servers.
    Nucleic Acids Res. 2000 Jan 1;28(1):228-30 PMID: 10592233
  11. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  12. The ENZYME database in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):304-5 PMID: 10592255
  13. Analysis and prediction of functional sub-types from protein sequence alignments.
    J Mol Biol. 2000 Oct 13;303(1):61-76 PMID: 11021970
  14. Human disease genes.
    Nature. 2001 Feb 15;409(6822):853-5 PMID: 11237009
  15. Pathogenesis and evolution of virulence in enteropathogenic and enterohemorrhagic Escherichia coli.
    J Clin Invest. 2001 Mar;107(5):539-48 PMID: 11238553
  16. ConSurf: an algorithmic tool for the identification of functional regions in proteins by surface mapping of phylogenetic information.
    J Mol Biol. 2001 Mar 16;307(1):447-63 PMID: 11243830
  17. Three-dimensional cluster analysis identifies interfaces and functional residue clusters in proteins.
    J Mol Biol. 2001 Apr 13;307(5):1487-502 PMID: 11292355
  18. Intrinsic errors in genome annotation.
    Trends Genet. 2001 Aug;17(8):429-31 PMID: 11485799
  19. Utilization of L-ascorbate by Escherichia coli K-12: assignments of functions to products of the yjf-sga and yia-sgb operons.
    J Bacteriol. 2002 Jan;184(1):302-6 PMID: 11741871
  20. The Pfam protein families database.
    Nucleic Acids Res. 2002 Jan 1;30(1):276-80 PMID: 11752314
  21. Divergent evolution of enzymatic function: mechanistically diverse superfamilies and functionally distinct suprafamilies.
    Annu Rev Biochem. 2001;70:209-46 PMID: 11395407
  22. Quod erat demonstrandum? The mystery of experimental validation of apparently erroneous computational analyses of protein sequences.
    Genome Biol. 2001;2(12):RESEARCH0051 PMID: 11790254
  23. The past, present and future of genome-wide re-annotation.
    Genome Biol. 2002;3(2):COMMENT2001 PMID: 11864365
  24. Plasticity of enzyme active sites.
    Trends Biochem Sci. 2002 Aug;27(8):419-26 PMID: 12151227
  25. PROSITE: a documented database using patterns and profiles as motif descriptors.
    Brief Bioinform. 2002 Sep;3(3):265-74 PMID: 12230035
  26. Defining the mandate of proteomics in the post-genomics era: workshop report.
    Mol Cell Proteomics. 2002 Oct;1(10):763-80 PMID: 12438560
  27. Automatic methods for predicting functionally important residues.
    J Mol Biol. 2003 Feb 28;326(4):1289-302 PMID: 12589769
  28. Definitions of enzyme function for the structural genomics era.
    Curr Opin Chem Biol. 2003 Apr;7(2):230-7 PMID: 12714057
  29. How well is enzyme function conserved as a function of pairwise sequence identity?
    J Mol Biol. 2003 Oct 31;333(4):863-82 PMID: 14568541
  30. Enzyme-specific profiles for genome annotation: PRIAM.
    Nucleic Acids Res. 2003 Nov 15;31(22):6633-9 PMID: 14602924
  31. Automatic prediction of protein function.
    Cell Mol Life Sci. 2003 Dec;60(12):2637-50 PMID: 14685688
  32. Identifying latent enzyme activities: substrate ambiguity within modern bacterial sugar kinases.
    Biochemistry. 2004 Jun 1;43(21):6387-92 PMID: 15157072
  33. A structural study for the optimisation of functional motifs encoded in protein sequences.
    BMC Bioinformatics. 2004 Apr 30;5:50 PMID: 15119965
  34. YbdK is a carboxylate-amine ligase with a gamma-glutamyl:Cysteine ligase activity: crystal structure and enzymatic assays.
    Proteins. 2004 Aug 1;56(2):376-83 PMID: 15211520
  35. Glycolate oxidoreductase in Escherichia coli.
    Biochim Biophys Acta. 1972 May 25;267(2):227-37 PMID: 4557653
  36. Comparison of the predicted and observed secondary structure of T4 phage lysozyme.
    Biochim Biophys Acta. 1975 Oct 20;405(2):442-51 PMID: 1180967
  37. Optimal alignments in linear space.
    Comput Appl Biosci. 1988 Mar;4(1):11-7 PMID: 3382986
  38. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  39. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  40. A method to predict functional residues in proteins.
    Nat Struct Biol. 1995 Feb;2(2):171-8 PMID: 7749921
  41. glc locus of Escherichia coli: characterization of genes encoding the subunits of glycolate oxidase and the glc regulator protein.
    J Bacteriol. 1996 Apr;178(7):2051-9 PMID: 8606183
  42. An evolutionary trace method defines binding surfaces common to protein families.
    J Mol Biol. 1996 Mar 29;257(2):342-58 PMID: 8609628
  43. Predicting functions from protein sequences--where are the bottlenecks?
    Nat Genet. 1998 Apr;18(4):313-8 PMID: 9537411
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2004-00-00
Epub
2004-00-01
Pages
6226-39
Language
English
Region
England
NLM ID
0411011
PMCID
PMC535665
Subset
IM
Grants
NIGMS NIH HHS · R01 GM048835 · United States
NIAID NIH HHS · U54 AI057158 · United States
NIGMS NIH HHS · GM-48835 · United States
NIAID NIH HHS · U54AI-057158 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]