Home LiteratureArticle Details
PMID: 17062146 Published · epublish English Comparative Study Evaluation Study Journal Article Research Support, Non-U.S. Gov't

The accuracy of several multiple sequence alignment programs for proteins.

BMC bioinformatics ·Vol. 7 ·2006-10-24 ·Pages 471

Nuin PA, Wang Z, Tillier ER

Abstract

There have been many algorithms and software programs implemented for the inference of multiple sequence alignments of protein and DNA sequences. The "true" alignment is usually unknown due to the incomplete knowledge of the evolutionary history of the sequences, making it difficult to gauge the relative accuracy of the programs. We tested nine of the most often used protein alignment programs and compared their results using sequences generated with the simulation software Simprot which creates known alignments under realistic and controlled evolutionary scenarios. We have simulated more than 30,000 alignment sets using various evolutionary histories in order to define strengths and weaknesses of each program tested. We found that alignment accuracy is extremely dependent on the number of insertions and deletions in the sequences, and that indel size has a weaker effect. We also considered benchmark alignments from the latest version of BAliBASE and the results relative to BAliBASE- and Simprot-generated data sets were consistent in most cases. Our results indicate that employing Simprot's simulated sequences allows the creation of a more flexible and broader range of alignment classes than the usual methods for alignment accuracy assessment. Simprot also allows for a quick and efficient analysis of a wider range of possible evolutionary histories that might not be present in currently available alignment sets. Among the nine programs tested, the iterative approach available in Mafft (L-INS-i) and ProbCons were consistently the most accurate, with Mafft being the faster of the two.

MeSH Terms
Amino Acid Sequence Computational Biology Computer Simulation Databases, Protein Gene Deletion Mutation Protein Conformation Proteins/chemistry,genetics Sequence Alignment/methods,standards Software
Chemicals
Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Nuin Paulo A S
Division of Cancer Genomics and Proteomics, Ontario Cancer Institute, University Health Network, 101 College St, M5G 1L7, Toronto, Ontario, Canada. [email protected]
Wang Zhouzhi
Tillier Elisabeth R M
References (36)
36 references, click to expand
  1. Maximum-likelihood estimation of phylogeny from DNA sequences when substitution rates differ over sites.
    Mol Biol Evol. 1993 Nov;10(6):1396-401 PMID: 8277861
  2. Rose: generating sequence families.
    Bioinformatics. 1998;14(2):157-63 PMID: 9545448
  3. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  4. DIALIGN-T: an improved algorithm for segment-based multiple sequence alignment.
    BMC Bioinformatics. 2005;6:66 PMID: 15784139
  5. A generalized affine gap model significantly improves protein sequence alignment accuracy.
    Proteins. 2005 Feb 1;58(2):329-38 PMID: 15562515
  6. Distribution of Indel lengths.
    Proteins. 2001 Oct 1;45(1):102-4 PMID: 11536366
  7. A weighting system and algorithm for aligning many phylogenetically related sequences.
    Comput Appl Biosci. 1995 Oct;11(5):543-51 PMID: 8590178
  8. MUSCLE: a multiple sequence alignment method with reduced time and space complexity.
    BMC Bioinformatics. 2004 Aug 19;5:113 PMID: 15318951
  9. ProbCons: Probabilistic consistency-based multiple sequence alignment.
    Genome Res. 2005 Feb;15(2):330-40 PMID: 15687296
  10. MySSP: non-stationary evolutionary sequence simulation, including indels.
    Evol Bioinform Online. 2007 Feb 26;1:81-3 PMID: 19325855
  11. Multiple DNA and protein sequence alignment based on segment-to-segment comparison.
    Proc Natl Acad Sci U S A. 1996 Oct 29;93(22):12098-103 PMID: 8901539
  12. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform.
    Nucleic Acids Res. 2002 Jul 15;30(14):3059-66 PMID: 12136088
  13. A transition probability model for amino acid substitutions from blocks.
    J Comput Biol. 2003;10(6):997-1010 PMID: 14980022
  14. Multiple sequence alignment accuracy and evolutionary distance estimation.
    BMC Bioinformatics. 2005;6:278 PMID: 16305750
  15. BAliBASE: a benchmark alignment database for the evaluation of multiple alignment programs.
    Bioinformatics. 1999 Jan;15(1):87-8 PMID: 10068696
  16. BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark.
    Proteins. 2005 Oct 1;61(1):127-36 PMID: 16044462
  17. SIMPROT: using an empirically determined indel distribution in simulations of protein evolution.
    BMC Bioinformatics. 2005;6:236 PMID: 16188037
  18. MAFFT version 5: improvement in accuracy of multiple sequence alignment.
    Nucleic Acids Res. 2005;33(2):511-8 PMID: 15661851
  19. Database of homology-derived protein structures and the structural meaning of sequence alignment.
    Proteins. 1991;9(1):56-68 PMID: 2017436
  20. Quality assessment of multiple alignment programs.
    FEBS Lett. 2002 Oct 2;529(1):126-30 PMID: 12354624
  21. Evaluation of protein multiple alignments by SAM-T99 using the BAliBASE multiple alignment test set.
    Bioinformatics. 2001 Aug;17(8):713-20 PMID: 11524372
  22. A comparison of scoring functions for protein sequence profile alignment.
    Bioinformatics. 2004 May 22;20(8):1301-8 PMID: 14962936
  23. Kalign--an accurate and fast multiple sequence alignment algorithm.
    BMC Bioinformatics. 2005;6:298 PMID: 16343337
  24. CASA: a server for the critical assessment of protein sequence alignment accuracy.
    Bioinformatics. 2002 Mar;18(3):496-7 PMID: 11934755
  25. DIALIGN 2: improvement of the segment-to-segment approach to multiple sequence alignment.
    Bioinformatics. 1999 Mar;15(3):211-8 PMID: 10222408
  26. Evolutionary distance estimation and fidelity of pair wise sequence alignment.
    BMC Bioinformatics. 2005;6:102 PMID: 15840174
  27. Comprehensive study on iterative algorithms of multiple sequence alignment.
    Comput Appl Biosci. 1995 Feb;11(1):13-8 PMID: 7796270
  28. The Pfam protein families database.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D138-41 PMID: 14681378
  29. SABmark--a benchmark for sequence alignment that covers the entire known fold space.
    Bioinformatics. 2005 Apr 1;21(7):1267-8 PMID: 15333456
  30. Large-scale comparison of protein sequence alignment algorithms with structure alignments.
    Proteins. 2000 Jul 1;40(1):6-22 PMID: 10813826
  31. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  32. DNA assembly with gaps (Dawg): simulating sequence evolution.
    Bioinformatics. 2005 Nov 1;21 Suppl 3:iii31-8 PMID: 16306390
  33. A space-efficient algorithm for local similarities.
    Comput Appl Biosci. 1990 Oct;6(4):373-81 PMID: 2257499
  34. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  35. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  36. Multiple sequence alignment using partial order graphs.
    Bioinformatics. 2002 Mar;18(3):452-64 PMID: 11934745
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2006-10-24
Epub
2006-00-24
Pages
471
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1633746
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]