Abstract
Hidden Markov model (HMM) techniques are used to model families of biological sequences. A smooth and convergent algorithm is introduced to iteratively adapt the transition and emission parameters of the models from the examples in a given family. The HMM approach is applied to three protein families: globins, immunoglobulins, and kinases. In all cases, the models derived capture the important statistical characteristics of the family and can be used for a number of tasks, including multiple alignments, motif detection, and classification. For K sequences of average length N, this approach yields an effective multiple-alignment algorithm which requires O(KN2) operations, linear in the number of sequences.
MeSH Terms
Algorithms
Amino Acid Sequence
Animals
Globins/genetics
Humans
Immunoglobulins/genetics
Markov Chains
Models, Genetic
Molecular Sequence Data
Protein Kinases/genetics
Proteins/genetics
Sequence Alignment/methods,statistics & numerical data
Sequence Homology, Amino Acid
Chemicals
Immunoglobulins
Proteins
Globins
Protein Kinases
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Baldi P
Division of Biology, California Institute of Technology, Pasadena 91125.
Chauvin Y
Hunkapiller T
McClure M A
References (18)
18 references, click to expand
-
Efficient methods for multiple sequence alignment with guaranteed error bounds.
Bull Math Biol. 1993 Jan;55(1):141-54
PMID: 7680269
-
Dual-specificity protein kinases: will any hydroxyl do?
Trends Biochem Sci. 1992 Mar;17(3):114-9
PMID: 1412695
-
A general method applicable to the search for similarities in the amino acid sequence of two proteins.
J Mol Biol. 1970 Mar;48(3):443-53
PMID: 5420325
-
Evolutionary processes and evolutionary noise at the molecular level. I. Functional density in proteins.
J Mol Evol. 1976 Apr 9;7(3):167-83
PMID: 933174
-
Similar amino acid sequences: chance or common ancestry?
Science. 1981 Oct 9;214(4517):149-59
PMID: 7280687
-
Establishing homologies in protein sequences.
Methods Enzymol. 1983;91:524-45
PMID: 6855599
-
A thousand and one protein kinases.
Cell. 1987 Sep 11;50(6):823-9
PMID: 3113737
-
Determinants of a protein fold. Unique features of the globin amino acid sequences.
J Mol Biol. 1987 Jul 5;196(1):199-216
PMID: 3656444
-
Stochastic models for heterogeneous DNA sequences.
Bull Math Biol. 1989;51(1):79-94
PMID: 2706403
-
An expectation maximization (EM) algorithm for the identification and characterization of common sites in unaligned biopolymer sequences.
Proteins. 1990;7(1):41-51
PMID: 2184437
-
Motif recognition and alignment for many sequences by comparison of dot-matrices.
J Mol Biol. 1991 Mar 5;218(1):33-43
PMID: 1900535
-
Crystal structure of the catalytic subunit of cyclic adenosine monophosphate-dependent protein kinase.
Science. 1991 Jul 26;253(5018):407-14
PMID: 1862342
-
An evolutionary model for maximum likelihood alignment of DNA sequences.
J Mol Evol. 1991 Aug;33(2):114-24
PMID: 1920447
-
Protein kinase catalytic domain sequence database: identification of conserved features of primary structure and classification of family members.
Methods Enzymol. 1991;200:38-62
PMID: 1956325
-
Expectation maximization algorithm for identifying protein-binding sites with variable lengths from unaligned DNA fragments.
J Mol Biol. 1992 Jan 5;223(1):159-70
PMID: 1731067
-
A survey of multiple sequence comparison methods.
Bull Math Biol. 1992 Jul;54(4):563-98
PMID: 1591533
-
CLUSTAL V: improved software for multiple sequence alignment.
Comput Appl Biosci. 1992 Apr;8(2):189-91
PMID: 1591615
-
Hidden Markov models in computational biology. Applications to protein modeling.
J Mol Biol. 1994 Feb 4;235(5):1501-31
PMID: 8107089