Abstract
The search for amino acid sequence homologies can be a powerful tool for predicting protein structure. Discovered sequence homologies are currently used in predicting the function of oncogene proteins. To sharpen this tool, we investigated the structural significance of short sequence homologies by searching proteins of known three-dimensional structure for subsequence identities. In 62 proteins with 10,000 residues, we found that the longest isolated homologies between unrelated proteins are five residues long. In 6 (out of 25) cases we saw surprising structural adaptability: the same five residues are part of an alpha-helix in one protein and part of a beta-strand in another protein. These examples show quantitatively that pentapeptide structure within a protein is strongly dependent on sequence context, a fact essentially ignored in most protein structure prediction methods: just considering the local sequence of five residues is not sufficient to predict correctly the local conformation (secondary structure). Cooperativity of length six or longer must be taken into account. Also, we are warned that in the growing practice of comparing a new protein sequence with a data base of known sequences, finding an identical pentapeptide sequence between two proteins is not a significant indication of structural similarity or of evolutionary kinship.
MeSH Terms
Amino Acid Sequence
Animals
Carbonic Anhydrases
Humans
Models, Molecular
Oligopeptides
Prealbumin
Probability
Protein Conformation
Proteins
Structure-Activity Relationship
Chemicals
Oligopeptides
Prealbumin
Proteins
Carbonic Anhydrases
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Kabsch W
Sander C
References (13)
13 references, click to expand
-
Molecular palaeogenetics: amino acid sequence homology in ribonuclease and lysozyme.
Comp Biochem Physiol. 1967 Nov;23(2):383-406
PMID: 6080502
-
An evaluation of the relatedness of proteins based on comparison of amino acid sequences.
J Mol Biol. 1970 Jun 28;50(3):617-39
PMID: 4097749
-
The Protein Data Bank: a computer-based archival file for macromolecular structures.
J Mol Biol. 1977 May 25;112(3):535-42
PMID: 875032
-
Structure of erythrocruorin in different ligand states refined at 1.4 A resolution.
J Mol Biol. 1979 Jan 25;127(3):309-38
PMID: 430568
-
Similar amino acid sequences: chance or common ancestry?
Science. 1981 Oct 9;214(4517):149-59
PMID: 7280687
-
Homology between human bladder carcinoma oncogene product and mitochondrial ATP-synthase.
Nature. 1983 Jan 20;301(5897):262-4
PMID: 6296696
-
Theory of protein secondary structure and algorithm of its prediction.
Biopolymers. 1983 Jan;22(1):15-25
PMID: 6673754
-
Predicted nucleotide-binding properties of p21 protein and its cancer-associated variant.
Nature. 1983 Apr 28;302(5911):842-4
PMID: 6843652
-
How good are predictions of protein secondary structure?
FEBS Lett. 1983 May 8;155(2):179-82
PMID: 6852232
-
Simian sarcoma virus onc gene, v-sis, is derived from the gene (or genes) encoding a platelet-derived growth factor.
Science. 1983 Jul 15;221(4607):275-7
PMID: 6304883
-
Platelet-derived growth factor is structurally related to the putative transforming protein p28sis of simian sarcoma virus.
Nature. 1983 Jul 7-13;304(5921):35-9
PMID: 6306471
-
Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features.
Biopolymers. 1983 Dec;22(12):2577-637
PMID: 6667333
-
Prediction of super-secondary structure in proteins.
Nature. 1983 Feb 10;301(5900):540-2
PMID: 6823335