Home LiteratureArticle Details
PMID: 17584497 Published · epublish English Journal Article Research Support, U.S. Gov't, Non-P.H.S.

Optimization based automated curation of metabolic reconstructions.

BMC bioinformatics ·Vol. 8 ·2007-06-20 ·Pages 212

Satish Kumar V, Dasika MS, Maranas CD

Abstract

Currently, there exists tens of different microbial and eukaryotic metabolic reconstructions (e.g., Escherichia coli, Saccharomyces cerevisiae, Bacillus subtilis) with many more under development. All of these reconstructions are inherently incomplete with some functionalities missing due to the lack of experimental and/or homology information. A key challenge in the automated generation of genome-scale reconstructions is the elucidation of these gaps and the subsequent generation of hypotheses to bridge them. In this work, an optimization based procedure is proposed to identify and eliminate network gaps in these reconstructions. First we identify the metabolites in the metabolic network reconstruction which cannot be produced under any uptake conditions and subsequently we identify the reactions from a customized multi-organism database that restores the connectivity of these metabolites to the parent network using four mechanisms. This connectivity restoration is hypothesized to take place through four mechanisms: a) reversing the directionality of one or more reactions in the existing model, b) adding reaction from another organism to provide functionality absent in the existing model, c) adding external transport mechanisms to allow for importation of metabolites in the existing model and d) restore flow by adding intracellular transport reactions in multi-compartment models. We demonstrate this procedure for the genome- scale reconstruction of Escherichia coli and also Saccharomyces cerevisiae wherein compartmentalization of intra-cellular reactions results in a more complex topology of the metabolic network. We determine that about 10% of metabolites in E. coli and 30% of metabolites in S. cerevisiae cannot carry any flux. Interestingly, the dominant flow restoration mechanism is directionality reversals of existing reactions in the respective models. We have proposed systematic methods to identify and fill gaps in genome-scale metabolic reconstructions. The identified gaps can be filled both by making modifications in the existing model and by adding missing reactions by reconciling multi-organism databases of reactions with existing genome-scale models. Computational results provide a list of hypotheses to be queried further and tested experimentally.

MeSH Terms
Algorithms Bacteria/metabolism Bacterial Proteins/metabolism Computer Simulation Gene Expression/physiology Models, Biological Signal Transduction/physiology
Chemicals
Bacterial Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Satish Kumar Vinay
Department of Industrial and Manufacturing Engineering, The Pennsylvania State University, University Park, PA 16802, USA. [email protected] <[email protected]>
Dasika Madhukar S
Maranas Costas D
References (27)
27 references, click to expand
  1. In silico predictions of Escherichia coli metabolic capabilities are consistent with experimental data.
    Nat Biotechnol. 2001 Feb;19(2):125-30 PMID: 11175725
  2. Predicting genes for orphan metabolic activities using phylogenetic profiles.
    Genome Biol. 2006;7(2):R17 PMID: 16507154
  3. Identification of the human methylmalonyl-CoA racemase gene based on the analysis of prokaryotic gene arrangements. Implications for decoding the human genome.
    J Biol Chem. 2001 Oct 5;276(40):37194-8 PMID: 11481338
  4. Computational method to assign microbial genes to pathways.
    J Cell Biochem Suppl. 2001;Suppl 37:106-9 PMID: 11842435
  5. Identifying metabolic enzymes with multiple types of association evidence.
    BMC Bioinformatics. 2006;7:177 PMID: 16571130
  6. Accelerating the reconstruction of genome-scale metabolic networks.
    BMC Bioinformatics. 2006;7:296 PMID: 16772023
  7. Systems approach to refining genome annotation.
    Proc Natl Acad Sci U S A. 2006 Nov 14;103(46):17480-4 PMID: 17088549
  8. Systematic assignment of thermodynamic constraints in metabolic network models.
    BMC Bioinformatics. 2006;7:512 PMID: 17123434
  9. Identification of the tRNA-dihydrouridine synthase family.
    J Biol Chem. 2002 Jul 12;277(28):25090-5 PMID: 11983710
  10. Genome sequence of Yersinia pestis KIM.
    J Bacteriol. 2002 Aug;184(16):4601-11 PMID: 12142430
  11. The Pathway Tools software.
    Bioinformatics. 2002;18 Suppl 1:S225-32 PMID: 12169551
  12. Complete genome sequence and comparative genomics of Shigella flexneri serotype 2a strain 2457T.
    Infect Immun. 2003 May;71(5):2775-86 PMID: 12704152
  13. Missing genes in metabolic pathways: a comparative genomics approach.
    Curr Opin Chem Biol. 2003 Apr;7(2):238-51 PMID: 12714058
  14. An expanded genome-scale model of Escherichia coli K-12 (iJR904 GSM/GPR).
    Genome Biol. 2003;4(9):R54 PMID: 12952533
  15. Metabolic networks: enzyme function and metabolite structure.
    Curr Opin Struct Biol. 2004 Jun;14(3):300-6 PMID: 15193309
  16. Reconstruction and validation of Saccharomyces cerevisiae iND750, a fully compartmentalized genome-scale metabolic model.
    Genome Res. 2004 Jul;14(7):1298-309 PMID: 15197165
  17. A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.
    BMC Bioinformatics. 2004 Jun 9;5:76 PMID: 15189570
  18. Filling gaps in a metabolic network using expression information.
    Bioinformatics. 2004 Aug 4;20 Suppl 1:i178-85 PMID: 15262797
  19. Metabolism and evolution of Haemophilus influenzae deduced from a whole-genome comparison with Escherichia coli.
    Curr Biol. 1996 Mar 1;6(3):279-91 PMID: 8805245
  20. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  21. The complete genome sequence of Escherichia coli K-12.
    Science. 1997 Sep 5;277(5331):1453-62 PMID: 9278503
  22. A novel methyltransferase catalyzes the methyl esterification of trans-aconitate in Escherichia coli.
    J Biol Chem. 1999 May 7;274(19):13470-9 PMID: 10224113
  23. EcoCyc: a comprehensive database resource for Escherichia coli.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D334-7 PMID: 15608210
  24. The Genomes On Line Database (GOLD) v.2: a monitor of genome projects worldwide.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D332-4 PMID: 16381880
  25. MetaCyc: a multiorganism database of metabolic pathways and enzymes.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D511-6 PMID: 16381923
  26. Genome-scale thermodynamic analysis of Escherichia coli metabolism.
    Biophys J. 2006 Feb 15;90(4):1453-61 PMID: 16299075
  27. Distinct reactions catalyzed by bacterial and yeast trans-aconitate methyltransferases.
    Biochemistry. 2001 Feb 20;40(7):2210-9 PMID: 11329290
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2007-06-20
Epub
2007-00-20
Pages
212
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1933441
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]