Abstract
Pfam is a large collection of protein multiple sequence alignments and profile hidden Markov models. Pfam is available on the World Wide Web in the UK at http://www.sanger.ac.uk/Software/Pfam/, in Sweden at http://www.cgb.ki.se/Pfam/, in France at http://pfam.jouy.inra.fr/ and in the US at http://pfam.wustl.edu/. The latest version (6.6) of Pfam contains 3071 families, which match 69% of proteins in SWISS-PROT 39 and TrEMBL 14. Structural data, where available, have been utilised to ensure that Pfam families correspond with structural domains, and to improve domain-based annotation. Predictions of non-domain regions are now also included. In addition to secondary structure, Pfam multiple sequence alignments now contain active site residue mark-up. New search tools, including taxonomy search and domain query, greatly add to the functionality and usability of the Pfam resource.
MeSH Terms
Animals
Binding Sites
Computer Graphics
Databases, Protein
Evolution, Molecular
Genome
Humans
Information Storage and Retrieval
Internet
Macromolecular Substances
Markov Chains
Phylogeny
Protein Structure, Secondary
Protein Structure, Tertiary
Proteins/chemistry,genetics,physiology
Sequence Alignment
Chemicals
Macromolecular Substances
Proteins
Authors & Affiliations
10 authors, click to expand affiliations / ORCID
Bateman Alex
Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SA, UK.
[email protected]
Birney Ewan
Cerruti Lorenzo
Durbin Richard
Etwiller Laurence
Eddy Sean R
Griffiths-Jones Sam
Howe Kevin L
Marshall Mhairi
Sonnhammer Erik L L
References (16)
16 references, click to expand
-
ProDom and ProDom-CG: tools for protein domain analysis and whole genome comparisons.
Nucleic Acids Res. 2000 Jan 1;28(1):267-9
PMID: 10592243
-
Machine learning approaches for the prediction of signal peptides and other protein sorting signals.
Protein Eng. 1999 Jan;12(1):3-9
PMID: 10065704
-
TIGRFAMs: a protein family resource for the functional identification of proteins.
Nucleic Acids Res. 2001 Jan 1;29(1):41-3
PMID: 11125044
-
PDBsum: summaries and analyses of PDB structures.
Nucleic Acids Res. 2001 Jan 1;29(1):221-2
PMID: 11125097
-
Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes.
J Mol Biol. 2001 Jan 19;305(3):567-80
PMID: 11152613
-
InterPro--an integrated documentation resource for protein families, domains and functional sites.
Bioinformatics. 2000 Dec;16(12):1145-50
PMID: 11159333
-
Initial sequencing and analysis of the human genome.
Nature. 2001 Feb 15;409(6822):860-921
PMID: 11237011
-
NIFAS: visual analysis of domain evolution in proteins.
Bioinformatics. 2001 Apr;17(4):343-8
PMID: 11301303
-
The Protein Data Bank: a computer-based archival file for macromolecular structures.
J Mol Biol. 1977 May 25;112(3):535-42
PMID: 875032
-
Predicting coiled coils from protein sequences.
Science. 1991 May 24;252(5009):1162-4
PMID: 2031185
-
SCOP: a structural classification of proteins database for the investigation of sequences and structures.
J Mol Biol. 1995 Apr 7;247(4):536-40
PMID: 7723011
-
Pfam: a comprehensive database of protein domain families based on seed alignments.
Proteins. 1997 Jul;28(3):405-20
PMID: 9223186
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1999.
Nucleic Acids Res. 1999 Jan 1;27(1):49-54
PMID: 9847139
-
SMART: identification and annotation of domains from signalling and extracellular protein sequences.
Nucleic Acids Res. 1999 Jan 1;27(1):229-32
PMID: 9847187
-
The genome sequence of Drosophila melanogaster.
Science. 2000 Mar 24;287(5461):2185-95
PMID: 10731132