Abstract
The SUPERFAMILY database provides protein domain assignments, at the SCOP 'superfamily' level, for the predicted protein sequences in over 400 completed genomes. A superfamily groups together domains of different families which have a common evolutionary ancestor based on structural, functional and sequence data. SUPERFAMILY domain assignments are generated using an expert curated set of profile hidden Markov models. All models and structural assignments are available for browsing and download from http://supfam.org. The web interface includes services such as domain architectures and alignment details for all protein assignments, searchable domain combinations, domain occurrence network visualization, detection of over- or under-represented superfamilies for a given genome by comparison with other genomes, assignment of manually submitted sequences and keyword searches. In this update we describe the SUPERFAMILY database and outline two major developments: (i) incorporation of family level assignments and (ii) a superfamily-level functional annotation. The SUPERFAMILY database can be used for general protein evolution and superfamily-specific studies, genomic annotation, and structural genomics target suggestion and assessment.
MeSH Terms
Databases, Protein
Genomics
Internet
Protein Structure, Tertiary/genetics,physiology
Proteins/classification
User-Computer Interface
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Wilson Derek
MRC Laboratory of Molecular Biology, Hills Road, Cambridge CB2 2QH, UK.
[email protected]
Madera Martin
Vogel Christine
Chothia Cyrus
Gough Julian
References (22)
22 references, click to expand
-
The Universal Protein Resource (UniProt): an expanding universe of protein information.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D187-91
PMID: 16381842
-
Improved profile HMM performance by assessment of critical algorithmic features in SAM and HMMER.
BMC Bioinformatics. 2005;6:99
PMID: 15831105
-
Gene3D: modelling protein structure, function and evolution.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D281-4
PMID: 16381865
-
Ensembl 2006.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D556-61
PMID: 16381931
-
DBD: a transcription factor prediction database.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D74-81
PMID: 16381970
-
Protein family expansions and biological complexity.
PLoS Comput Biol. 2006 May;2(5):e48
PMID: 16733546
-
Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure.
J Mol Biol. 2001 Nov 2;313(4):903-19
PMID: 11697912
-
Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Nat Genet. 2000 May;25(1):25-9
PMID: 10802651
-
Genomic scale sub-family assignment of protein domains.
Nucleic Acids Res. 2006;34(13):3625-33
PMID: 16877569
-
A comparison of profile hidden Markov model procedures for remote homology detection.
Nucleic Acids Res. 2002 Oct 1;30(19):4321-8
PMID: 12364612
-
The COG database: an updated version includes eukaryotes.
BMC Bioinformatics. 2003 Sep 11;4:41
PMID: 12969510
-
SCOP database in 2004: refinements integrate structure and sequence family data.
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D226-9
PMID: 14681400
-
The SUPERFAMILY database in 2004: additions and improvements.
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D235-9
PMID: 14681402
-
Supra-domains: evolutionary units larger than single protein domains.
J Mol Biol. 2004 Feb 20;336(3):809-23
PMID: 15095989
-
CNplot: visualizing pre-clustered networks.
Bioinformatics. 2004 Jun 12;20(9):1455-6
PMID: 14871859
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
Profile hidden Markov models.
Bioinformatics. 1998;14(9):755-63
PMID: 9918945
-
Hidden Markov models for detecting remote protein homologies.
Bioinformatics. 1998;14(10):846-56
PMID: 9927713
-
InterPro, progress and status in 2005.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D201-5
PMID: 15608177
-
The RCSB Protein Data Bank: a redesigned query system and relational database based on the mmCIF schema.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D233-7
PMID: 15608185
-
The relationship between domain duplication and recombination.
J Mol Biol. 2005 Feb 11;346(1):355-65
PMID: 15663950
-
Pfam: clans, web tools and services.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D247-51
PMID: 16381856