Abstract
The PEDANT genome database (http://pedant.gsf.de) provides exhaustive automatic analysis of genomic sequences by a large variety of established bioinformatics tools through a comprehensive Web-based user interface. One hundred and seventy seven completely sequenced and unfinished genomes have been processed so far, including large eukaryotic genomes (mouse, human) published recently. In this contribution, we describe the current status of the PEDANT database and novel analytical features added to the PEDANT server in 2002. Those include: (i) integration with the BioRS data retrieval system which allows fast text queries, (ii) pre-computed sequence clusters in each complete genome, (iii) a comprehensive set of tools for genome comparison, including genome comparison tables and protein function prediction based on genomic context, and (iv) computation and visualization of protein-protein interaction (PPI) networks based on experimental data. The availability of functional and structural predictions for 650 000 genomic proteins in well organized form makes PEDANT a useful resource for both functional and structural genomics.
MeSH Terms
Animals
Computational Biology
Databases, Genetic
Gene Duplication
Genome
Genomics
Humans
Information Storage and Retrieval
Mice
Protein Interaction Mapping
Proteins/chemistry,physiology
Saccharomyces cerevisiae Proteins/metabolism
Sequence Analysis, Protein
Sequence Homology
Chemicals
Proteins
Saccharomyces cerevisiae Proteins
Authors & Affiliations
15 authors, click to expand affiliations / ORCID
Frishman Dmitrij
Institute for Bioinformatics, GSF - National Research Center for Environment and Health, Ingolstädter Landstrasse 1, 85764 Neueherberg, Germany.
[email protected]
Mokrejs Martin
Kosykh Denis
Kastenmüller Gabi
Kolesov Grigory
Zubrzycki Igor
Gruber Christian
Geier Birgitta
Kaps Andreas
Albermann Kaj
Volz Andreas
Wagner Christian
Fellenberg Matthias
Heumann Klaus
Mewes Hans-Werner
References (27)
27 references, click to expand
-
The protein information resource (PIR).
Nucleic Acids Res. 2000 Jan 1;28(1):41-4
PMID: 10592177
-
Combining diverse evidence for gene recognition in completely sequenced bacterial genomes.
Nucleic Acids Res. 1998 Jun 15;26(12):2941-7
PMID: 9611239
-
A comprehensive analysis of protein-protein interactions in Saccharomyces cerevisiae.
Nature. 2000 Feb 10;403(6770):623-7
PMID: 10688190
-
IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices.
Bioinformatics. 1999 Dec;15(12):1000-11
PMID: 10745990
-
Integrative analysis of protein interaction data.
Proc Int Conf Intell Syst Mol Biol. 2000;8:152-61
PMID: 10977076
-
The COG database: new developments in phylogenetic classification of proteins from complete genomes.
Nucleic Acids Res. 2001 Jan 1;29(1):22-8
PMID: 11125040
-
The InterPro database, an integrated documentation resource for protein families, domains and functional sites.
Nucleic Acids Res. 2001 Jan 1;29(1):37-40
PMID: 11125043
-
Profile hidden Markov models.
Bioinformatics. 1998;14(9):755-63
PMID: 9918945
-
Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.
Proc Natl Acad Sci U S A. 1999 Apr 13;96(8):4285-8
PMID: 10200254
-
Blocks+: a non-redundant database of protein alignment blocks derived from multiple compilations.
Bioinformatics. 1999 Jun;15(6):471-9
PMID: 10383472
-
A rapid classification protocol for the CATH Domain Database to support structural genomics.
Nucleic Acids Res. 2001 Jan 1;29(1):223-7
PMID: 11125098
-
Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes.
J Mol Biol. 2001 Jan 19;305(3):567-80
PMID: 11152613
-
Functional and structural genomics using PEDANT.
Bioinformatics. 2001 Jan;17(1):44-57
PMID: 11222261
-
Sources of systematic error in functional annotation of genomes: domain rearrangement, non-orthologous gene displacement and operon disruption.
In Silico Biol. 1998;1(1):55-67
PMID: 11471243
-
SNAPping up functionally related genes based on context information: a colinearity-free approach.
J Mol Biol. 2001 Aug 24;311(4):639-56
PMID: 11518521
-
GenBank.
Nucleic Acids Res. 2002 Jan 1;30(1):17-20
PMID: 11752243
-
The EMBL Nucleotide Sequence Database.
Nucleic Acids Res. 2002 Jan 1;30(1):21-6
PMID: 11752244
-
MIPS: a database for genomes and protein sequences.
Nucleic Acids Res. 2002 Jan 1;30(1):31-4
PMID: 11752246
-
The PROSITE database, its status in 2002.
Nucleic Acids Res. 2002 Jan 1;30(1):235-8
PMID: 11752303
-
SCOP database in 2002: refinements accommodate structural genomics.
Nucleic Acids Res. 2002 Jan 1;30(1):264-7
PMID: 11752311
-
The Pfam protein families database.
Nucleic Acids Res. 2002 Jan 1;30(1):276-80
PMID: 11752314
-
Knowledge-based selection of targets for structural genomics.
Protein Eng. 2002 Mar;15(3):169-83
PMID: 11932488
-
SNAPper: gene order predicts gene function.
Bioinformatics. 2002 Jul;18(7):1017-9
PMID: 12117803
-
Predicting coiled coils from protein sequences.
Science. 1991 May 24;252(5009):1162-4
PMID: 2031185
-
Overview of the yeast genome.
Nature. 1997 May 29;387(6632 Suppl):7-65
PMID: 9169865
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
The Protein Data Bank.
Nucleic Acids Res. 2000 Jan 1;28(1):235-42
PMID: 10592235