Home LiteratureArticle Details
PMID: 17267437 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

PROMALS: towards accurate multiple sequence alignments of distantly related proteins.

Bioinformatics (Oxford, England) ·Vol. 23 ·No. 7 ·2007-04-01 ·Pages 802-8

Pei J, Grishin NV

Abstract

Accurate multiple sequence alignments are essential in protein structure modeling, functional prediction and efficient planning of experiments. Although the alignment problem has attracted considerable attention, preparation of high-quality alignments for distantly related sequences remains a difficult task. We developed PROMALS, a multiple alignment method that shows promising results for protein homologs with sequence identity below 10%, aligning close to half of the amino acid residues correctly on average. This is about three times more accurate than traditional pairwise sequence alignment methods. PROMALS algorithm derives its strength from several sources: (i) sequence database searches to retrieve additional homologs; (ii) accurate secondary structure prediction; (iii) a hidden Markov model that uses a novel combined scoring of amino acids and secondary structures; (iv) probabilistic consistency-based scoring applied to progressive alignment of profiles. Compared to the best alignment methods that do not use secondary structure prediction and database searches (e.g. MUMMALS, ProbCons and MAFFT), PROMALS is up to 30% more accurate, with improvement being most prominent for highly divergent homologs. Compared to SPEM and HHalign, which also employ database searches and secondary structure prediction, PROMALS shows an accuracy improvement of several percent. The PROMALS web server is available at: http://prodata.swmed.edu/promals/. Supplementary data are available at Bioinformatics online.

MeSH Terms
Algorithms Amino Acid Sequence Conserved Sequence Molecular Sequence Data Proteins/chemistry Reproducibility of Results Sensitivity and Specificity Sequence Alignment/methods Sequence Analysis, Protein/methods Sequence Homology, Amino Acid Software
Chemicals
Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Pei Jimin
Howard Hughes Medical Institute, The University of Texas Southwestern Medical Center at Dallas, 6001 Forest Park Road, Dallas, TX 75390-9050, USA. [email protected]
Grishin Nick V
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2007-04-01
Epub
2007-00-31
Pages
802-8
Language
English
Region
England
NLM ID
9808944
Subset
IM
Grants
NIGMS NIH HHS · GM67165 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]