Home LiteratureArticle Details
PMID: 15312235 Published · epublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't

IdentiCS--identification of coding sequence and in silico reconstruction of the metabolic network directly from unannotated low-coverage bacterial genome sequence.

BMC bioinformatics ·Vol. 5 ·2004-08-16 ·Pages 112

Sun J, Zeng AP

Abstract

A necessary step for a genome level analysis of the cellular metabolism is the in silico reconstruction of the metabolic network from genome sequences. The available methods are mainly based on the annotation of genome sequences including two successive steps, the prediction of coding sequences (CDS) and their function assignment. The annotation process takes time. The available methods often encounter difficulties when dealing with unfinished error-containing genomic sequence. In this work a fast method is proposed to use unannotated genome sequence for predicting CDSs and for an in silico reconstruction of metabolic networks. Instead of using predicted genes or CDSs to query public databases, entries from public DNA or protein databases are used as queries to search a local database of the unannotated genome sequence to predict CDSs. Functions are assigned to the predicted CDSs simultaneously. The well-annotated genome of Salmonella typhimurium LT2 is used as an example to demonstrate the applicability of the method. 97.7% of the CDSs in the original annotation are correctly identified. The use of SWISS-PROT-TrEMBL databases resulted in an identification of 98.9% of CDSs that have EC-numbers in the published annotation. Furthermore, two versions of sequences of the bacterium Klebsiella pneumoniae with different genome coverage (3.9 and 7.9 fold, respectively) are examined. The results suggest that a 3.9-fold coverage of the bacterial genome could be sufficiently used for the in silico reconstruction of the metabolic network. Compared to other gene finding methods such as CRITICA our method is more suitable for exploiting sequences of low genome coverage. Based on the new method, a program called IdentiCS (Identification of Coding Sequences from Unfinished Genome Sequences) is delivered that combines the identification of CDSs with the reconstruction, comparison and visualization of metabolic networks (free to download at http://genome.gbf.de/bioinformatics/index.html). The reversed querying process and the program IdentiCS allow a fast and adequate prediction protein coding sequences and reconstruction of the potential metabolic network from low coverage genome sequences of bacteria. The new method can accelerate the use of genomic data for studying cellular metabolism.

MeSH Terms
Base Sequence/genetics Computational Biology/methods DNA, Bacterial/genetics Genome, Bacterial Klebsiella pneumoniae/genetics,metabolism Open Reading Frames/genetics Salmonella typhimurium/genetics,metabolism
Chemicals
DNA, Bacterial
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Sun Jibin
Department of Genome Analysis, GBF-German Research Center for Biotechnology, Mascheroder Weg 1, Braunschweig, 38124, Germany. [email protected]
Zeng An-Ping
References (22)
22 references, click to expand
  1. Flexible sequence similarity searching with the FASTA3 program package.
    Methods Mol Biol. 2000;132:185-219 PMID: 10547837
  2. Phylogenetic comparison of metabolic capacities of organisms at genome level.
    Mol Phylogenet Evol. 2004 Apr;31(1):204-13 PMID: 15019620
  3. KEGG: kyoto encyclopedia of genes and genomes.
    Nucleic Acids Res. 2000 Jan 1;28(1):27-30 PMID: 10592173
  4. The EcoCyc and MetaCyc databases.
    Nucleic Acids Res. 2000 Jan 1;28(1):56-9 PMID: 10592180
  5. WIT: integrated system for high-throughput genome sequence analysis and metabolic reconstruction.
    Nucleic Acids Res. 2000 Jan 1;28(1):123-5 PMID: 10592199
  6. Integrated genomic and proteomic analyses of a systematically perturbed metabolic network.
    Science. 2001 May 4;292(5518):929-34 PMID: 11340206
  7. GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions.
    Nucleic Acids Res. 2001 Jun 15;29(12):2607-18 PMID: 11410670
  8. Complete genome sequence of Salmonella enterica serovar Typhimurium LT2.
    Nature. 2001 Oct 25;413(6858):852-6 PMID: 11677609
  9. The EcoCyc Database.
    Nucleic Acids Res. 2002 Jan 1;30(1):56-8 PMID: 11752253
  10. Evaluation of gene structure prediction programs.
    Genomics. 1996 Jun 15;34(3):353-67 PMID: 8786136
  11. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  12. MPW: the Metabolic Pathways Database.
    Nucleic Acids Res. 1998 Jan 1;26(1):43-5 PMID: 9407141
  13. CRITICA: coding region identification tool invoking comparative analysis.
    Mol Biol Evol. 1999 Apr;16(4):512-24 PMID: 10331277
  14. The PROSITE database, its status in 2002.
    Nucleic Acids Res. 2002 Jan 1;30(1):235-8 PMID: 11752303
  15. The Pfam protein families database.
    Nucleic Acids Res. 2002 Jan 1;30(1):276-80 PMID: 11752314
  16. PathFinder: reconstruction and dynamic visualization of metabolic pathways.
    Bioinformatics. 2002 Jan;18(1):124-9 PMID: 11836220
  17. The InterPro Database, 2003 brings increased coverage and new features.
    Nucleic Acids Res. 2003 Jan 1;31(1):315-8 PMID: 12520011
  18. Reconstruction of metabolic networks from genome data and analysis of their global structure for various organisms.
    Bioinformatics. 2003 Jan 22;19(2):270-7 PMID: 12538249
  19. ZCURVE: a new system for recognizing protein-coding genes in bacterial and archaeal genomes.
    Nucleic Acids Res. 2003 Mar 15;31(6):1780-9 PMID: 12626720
  20. The connectivity structure, giant strong component and centrality of metabolic networks.
    Bioinformatics. 2003 Jul 22;19(11):1423-30 PMID: 12874056
  21. BRENDA, the enzyme database: updates and major new developments.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D431-3 PMID: 14681450
  22. Improved microbial gene identification with GLIMMER.
    Nucleic Acids Res. 1999 Dec 1;27(23):4636-41 PMID: 10556321
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2004-08-16
Epub
2004-00-16
Pages
112
Language
English
Region
England
NLM ID
100965194
PMCID
PMC514700
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]