Abstract
A necessary step for a genome level analysis of the cellular metabolism is the in silico reconstruction of the metabolic network from genome sequences. The available methods are mainly based on the annotation of genome sequences including two successive steps, the prediction of coding sequences (CDS) and their function assignment. The annotation process takes time. The available methods often encounter difficulties when dealing with unfinished error-containing genomic sequence. In this work a fast method is proposed to use unannotated genome sequence for predicting CDSs and for an in silico reconstruction of metabolic networks. Instead of using predicted genes or CDSs to query public databases, entries from public DNA or protein databases are used as queries to search a local database of the unannotated genome sequence to predict CDSs. Functions are assigned to the predicted CDSs simultaneously. The well-annotated genome of Salmonella typhimurium LT2 is used as an example to demonstrate the applicability of the method. 97.7% of the CDSs in the original annotation are correctly identified. The use of SWISS-PROT-TrEMBL databases resulted in an identification of 98.9% of CDSs that have EC-numbers in the published annotation. Furthermore, two versions of sequences of the bacterium Klebsiella pneumoniae with different genome coverage (3.9 and 7.9 fold, respectively) are examined. The results suggest that a 3.9-fold coverage of the bacterial genome could be sufficiently used for the in silico reconstruction of the metabolic network. Compared to other gene finding methods such as CRITICA our method is more suitable for exploiting sequences of low genome coverage. Based on the new method, a program called IdentiCS (Identification of Coding Sequences from Unfinished Genome Sequences) is delivered that combines the identification of CDSs with the reconstruction, comparison and visualization of metabolic networks (free to download at http://genome.gbf.de/bioinformatics/index.html). The reversed querying process and the program IdentiCS allow a fast and adequate prediction protein coding sequences and reconstruction of the potential metabolic network from low coverage genome sequences of bacteria. The new method can accelerate the use of genomic data for studying cellular metabolism.
MeSH Terms
Base Sequence/genetics
Computational Biology/methods
DNA, Bacterial/genetics
Genome, Bacterial
Klebsiella pneumoniae/genetics,metabolism
Open Reading Frames/genetics
Salmonella typhimurium/genetics,metabolism
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Sun Jibin
Department of Genome Analysis, GBF-German Research Center for Biotechnology, Mascheroder Weg 1, Braunschweig, 38124, Germany.
[email protected]
Zeng An-Ping
References (22)
22 references, click to expand
-
Flexible sequence similarity searching with the FASTA3 program package.
Methods Mol Biol. 2000;132:185-219
PMID: 10547837
-
Phylogenetic comparison of metabolic capacities of organisms at genome level.
Mol Phylogenet Evol. 2004 Apr;31(1):204-13
PMID: 15019620
-
KEGG: kyoto encyclopedia of genes and genomes.
Nucleic Acids Res. 2000 Jan 1;28(1):27-30
PMID: 10592173
-
The EcoCyc and MetaCyc databases.
Nucleic Acids Res. 2000 Jan 1;28(1):56-9
PMID: 10592180
-
WIT: integrated system for high-throughput genome sequence analysis and metabolic reconstruction.
Nucleic Acids Res. 2000 Jan 1;28(1):123-5
PMID: 10592199
-
Integrated genomic and proteomic analyses of a systematically perturbed metabolic network.
Science. 2001 May 4;292(5518):929-34
PMID: 11340206
-
GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions.
Nucleic Acids Res. 2001 Jun 15;29(12):2607-18
PMID: 11410670
-
Complete genome sequence of Salmonella enterica serovar Typhimurium LT2.
Nature. 2001 Oct 25;413(6858):852-6
PMID: 11677609
-
The EcoCyc Database.
Nucleic Acids Res. 2002 Jan 1;30(1):56-8
PMID: 11752253
-
Evaluation of gene structure prediction programs.
Genomics. 1996 Jun 15;34(3):353-67
PMID: 8786136
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
MPW: the Metabolic Pathways Database.
Nucleic Acids Res. 1998 Jan 1;26(1):43-5
PMID: 9407141
-
CRITICA: coding region identification tool invoking comparative analysis.
Mol Biol Evol. 1999 Apr;16(4):512-24
PMID: 10331277
-
The PROSITE database, its status in 2002.
Nucleic Acids Res. 2002 Jan 1;30(1):235-8
PMID: 11752303
-
The Pfam protein families database.
Nucleic Acids Res. 2002 Jan 1;30(1):276-80
PMID: 11752314
-
PathFinder: reconstruction and dynamic visualization of metabolic pathways.
Bioinformatics. 2002 Jan;18(1):124-9
PMID: 11836220
-
The InterPro Database, 2003 brings increased coverage and new features.
Nucleic Acids Res. 2003 Jan 1;31(1):315-8
PMID: 12520011
-
Reconstruction of metabolic networks from genome data and analysis of their global structure for various organisms.
Bioinformatics. 2003 Jan 22;19(2):270-7
PMID: 12538249
-
ZCURVE: a new system for recognizing protein-coding genes in bacterial and archaeal genomes.
Nucleic Acids Res. 2003 Mar 15;31(6):1780-9
PMID: 12626720
-
The connectivity structure, giant strong component and centrality of metabolic networks.
Bioinformatics. 2003 Jul 22;19(11):1423-30
PMID: 12874056
-
BRENDA, the enzyme database: updates and major new developments.
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D431-3
PMID: 14681450
-
Improved microbial gene identification with GLIMMER.
Nucleic Acids Res. 1999 Dec 1;27(23):4636-41
PMID: 10556321