Home LiteratureArticle Details
PMID: 15189570 Published · epublish English Journal Article Research Support, U.S. Gov't, Non-P.H.S. Research Support, U.S. Gov't, P.H.S.

A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.

BMC bioinformatics ·Vol. 5 ·2004-06-09 ·Pages 76

Green ML, Karp PD

Abstract

The PathoLogic program constructs Pathway/Genome databases by using a genome's annotation to predict the set of metabolic pathways present in an organism. PathoLogic determines the set of reactions composing those pathways from the enzymes annotated in the organism's genome. Most annotation efforts fail to assign function to 40-60% of sequences. In addition, large numbers of sequences may have non-specific annotations (e.g., thiolase family protein). Pathway holes occur when a genome appears to lack the enzymes needed to catalyze reactions in a pathway. If a protein has not been assigned a specific function during the annotation process, any reaction catalyzed by that protein will appear as a missing enzyme or pathway hole in a Pathway/Genome database. We have developed a method that efficiently combines homology and pathway-based evidence to identify candidates for filling pathway holes in Pathway/Genome databases. Our program not only identifies potential candidate sequences for pathway holes, but combines data from multiple, heterogeneous sources to assess the likelihood that a candidate has the required function. Our algorithm emulates the manual sequence annotation process, considering not only evidence from homology searches, but also considering evidence from genomic context (i.e., is the gene part of an operon?) and functional context (e.g., are there functionally-related genes nearby in the genome?) to determine the posterior belief that a candidate has the required function. The method can be applied across an entire metabolic pathway network and is generally applicable to any pathway database. The program uses a set of sequences encoding the required activity in other genomes to identify candidate proteins in the genome of interest, and then evaluates each candidate by using a simple Bayes classifier to determine the probability that the candidate has the desired function. We achieved 71% precision at a probability threshold of 0.9 during cross-validation using known reactions in computationally-predicted pathway databases. After applying our method to 513 pathway holes in 333 pathways from three Pathway/Genome databases, we increased the number of complete pathways by 42%. We made putative assignments to 46% of the holes, including annotation of 17 sequences of previously unknown function. Our pathway hole filler can be used not only to increase the utility of Pathway/Genome databases to both experimental and computational researchers, but also to improve predictions of protein function.

MeSH Terms
Amino Acid Oxidoreductases/metabolism Bacterial Proteins/metabolism,physiology Bayes Theorem Carbon-Nitrogen Ligases with Glutamine as Amide-N-Donor/metabolism Caulobacter crescentus/enzymology,metabolism Computational Biology Databases, Factual Escherichia coli Proteins Models, Statistical Multienzyme Complexes/chemistry,metabolism Mycobacterium tuberculosis/enzymology,metabolism Nicotinamide-Nucleotide Adenylyltransferase/metabolism Predictive Value of Tests Pyridines/metabolism Software Software Validation Vibrio cholerae/enzymology,metabolism
Chemicals
Bacterial Proteins Escherichia coli Proteins Multienzyme Complexes Pyridines Amino Acid Oxidoreductases L-aspartate oxidase, E coli Nicotinamide-Nucleotide Adenylyltransferase nicotinic acid mononucleotide adenylyltransferase Carbon-Nitrogen Ligases with Glutamine as Amide-N-Donor NAD+ synthase (glutamine-hydrolysing)
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Green Michelle L
Bioinformatics Research Group, SRI International, 333 Ravenswood Ave, Menlo Park, CA 94025, USA. [email protected]
Karp Peter D
References (30)
30 references, click to expand
  1. The TIGRFAMs database of protein families.
    Nucleic Acids Res. 2003 Jan 1;31(1):371-3 PMID: 12520025
  2. Shotgun: getting more from sequence similarity searches.
    Bioinformatics. 1999 Sep;15(9):729-40 PMID: 10498773
  3. DNA sequence of both chromosomes of the cholera pathogen Vibrio cholerae.
    Nature. 2000 Aug 3;406(6795):477-83 PMID: 10952301
  4. The COG database: new developments in phylogenetic classification of proteins from complete genomes.
    Nucleic Acids Res. 2001 Jan 1;29(1):22-8 PMID: 11125040
  5. Complete genome sequence of Caulobacter crescentus.
    Proc Natl Acad Sci U S A. 2001 Mar 27;98(7):4136-41 PMID: 11259647
  6. The EcoCyc Database.
    Nucleic Acids Res. 2002 Jan 1;30(1):56-8 PMID: 11752253
  7. The Pathway Tools software.
    Bioinformatics. 2002;18 Suppl 1:S225-32 PMID: 12169551
  8. GenBank.
    Nucleic Acids Res. 2003 Jan 1;31(1):23-7 PMID: 12519940
  9. The Protein Information Resource.
    Nucleic Acids Res. 2003 Jan 1;31(1):345-7 PMID: 12520019
  10. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.
    Nucleic Acids Res. 2003 Jan 1;31(1):365-70 PMID: 12520024
  11. Missing genes in metabolic pathways: a comparative genomics approach.
    Curr Opin Chem Biol. 2003 Apr;7(2):238-51 PMID: 12714058
  12. An expanded genome-scale model of Escherichia coli K-12 (iJR904 GSM/GPR).
    Genome Biol. 2003;4(9):R54 PMID: 12952533
  13. Enzyme-specific profiles for genome annotation: PRIAM.
    Nucleic Acids Res. 2003 Nov 15;31(22):6633-9 PMID: 14602924
  14. L-Aspartate oxidase, a newly discovered enzyme of Escherichia coli, is the B protein of quinolinate synthetase.
    J Biol Chem. 1982 Jan 25;257(2):626-32 PMID: 7033218
  15. Molecular biology of pyridine nucleotide biosynthesis in Escherichia coli. Cloning and characterization of quinolinate synthesis genes nadA and nadB.
    Eur J Biochem. 1988 Aug 1;175(2):221-8 PMID: 2841129
  16. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  17. Hidden Markov models in computational biology. Applications to protein modeling.
    J Mol Biol. 1994 Feb 4;235(5):1501-31 PMID: 8107089
  18. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  19. Fitting a mixture model by expectation maximization to discover motifs in biopolymers.
    Proc Int Conf Intell Syst Mol Biol. 1994;2:28-36 PMID: 7584402
  20. Hidden Markov models for sequence analysis: extension and analysis of the basic method.
    Comput Appl Biosci. 1996 Apr;12(2):95-107 PMID: 8744772
  21. Hidden Markov models.
    Curr Opin Struct Biol. 1996 Jun;6(3):361-5 PMID: 8804822
  22. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  23. Assessing sequence comparison methods with reliable structurally identified distant evolutionary relationships.
    Proc Natl Acad Sci U S A. 1998 May 26;95(11):6073-8 PMID: 9600919
  24. Statistics of large-scale sequence searching.
    Bioinformatics. 1998;14(3):279-84 PMID: 9614271
  25. The MTCY428.08 gene of Mycobacterium tuberculosis codes for NAD+ synthetase.
    J Bacteriol. 1998 Jun;180(12):3218-21 PMID: 9620974
  26. Deciphering the biology of Mycobacterium tuberculosis from the complete genome sequence.
    Nature. 1998 Jun 11;393(6685):537-44 PMID: 9634230
  27. Combining evidence using p-values: application to sequence homology searches.
    Bioinformatics. 1998;14(1):48-54 PMID: 9520501
  28. Hidden Markov models for detecting remote protein homologies.
    Bioinformatics. 1998;14(10):846-56 PMID: 9927713
  29. Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.
    Proc Natl Acad Sci U S A. 1999 Apr 13;96(8):4285-8 PMID: 10200254
  30. Microbial genomes and "missing" enzymes: redefining biochemical pathways.
    Arch Microbiol. 1999 Nov;172(5):269-79 PMID: 10550468
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2004-06-09
Epub
2004-00-09
Pages
76
Language
English
Region
England
NLM ID
100965194
PMCID
PMC446185
Subset
IM
Grants
PHS HHS · 07033 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]