Home LiteratureArticle Details
PMID: 19648142 Published · ppublish English Journal Article Review

Upcoming challenges for multiple sequence alignment methods in the high-throughput era.

Bioinformatics (Oxford, England) ·Vol. 25 ·No. 19 ·2009-10-01 ·Pages 2455-65

Kemena C, Notredame C

Abstract

This review focuses on recent trends in multiple sequence alignment tools. It describes the latest algorithmic improvements including the extension of consistency-based methods to the problem of template-based multiple sequence alignments. Some results are presented suggesting that template-based methods are significantly more accurate than simpler alternative methods. The validation of existing methods is also discussed at length with the detailed description of recent results and some suggestions for future validation strategies. The last part of the review addresses future challenges for multiple sequence alignment methods in the genomic era, most notably the need to cope with very large sequences, the need to integrate large amounts of experimental data, the need to accurately align non-coding and non-transcribed sequences and finally, the need to integrate many alternative methods and approaches.

MeSH Terms
Algorithms Amino Acid Sequence Computational Biology/methods Genome Molecular Sequence Data Phylogeny Sequence Alignment/methods Sequence Analysis, Protein
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Kemena Carsten
Centre For Genomic Regulation, Pompeus Fabre University, Carrer del Doctor Aiguader 88, 08003 Barcelona, Spain.
Notredame Cedric
References (73)
73 references, click to expand
  1. HOMSTRAD: recent developments of the Homologous Protein Structure Alignment Database.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D203-7 PMID: 14681395
  2. Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project.
    Nature. 2007 Jun 14;447(7146):799-816 PMID: 17571346
  3. Automated server predictions in CASP7.
    Proteins. 2007;69 Suppl 8:68-82 PMID: 17894354
  4. Fast embedding methods for clustering tens of thousands of sequences.
    Comput Biol Chem. 2008 Aug;32(4):282-6 PMID: 18450519
  5. PROMALS: towards accurate multiple sequence alignments of distantly related proteins.
    Bioinformatics. 2007 Apr 1;23(7):802-8 PMID: 17267437
  6. Multiple sequence alignments.
    Curr Opin Struct Biol. 2005 Jun;15(3):261-6 PMID: 15963889
  7. Multiple protein sequence alignment.
    Curr Opin Struct Biol. 2008 Jun;18(3):382-6 PMID: 18485694
  8. Phylogeny-aware gap placement prevents errors in sequence alignment and evolutionary analysis.
    Science. 2008 Jun 20;320(5883):1632-5 PMID: 18566285
  9. A small molecule-kinase interaction map for clinical kinase inhibitors.
    Nat Biotechnol. 2005 Mar;23(3):329-36 PMID: 15711537
  10. Segment-based multiple sequence alignment.
    Bioinformatics. 2008 Aug 15;24(16):i187-92 PMID: 18689823
  11. Efficient pairwise RNA structure prediction and alignment using sequence alignment constraints.
    BMC Bioinformatics. 2006 Sep 04;7:400 PMID: 16952317
  12. Motif recognition and alignment for many sequences by comparison of dot-matrices.
    J Mol Biol. 1991 Mar 5;218(1):33-43 PMID: 1900535
  13. SPEM: improving multiple sequence alignment with sequence profiles and predicted secondary structures.
    Bioinformatics. 2005 Sep 15;21(18):3615-21 PMID: 16020471
  14. On the complexity of multiple sequence alignment.
    J Comput Biol. 1994 Winter;1(4):337-48 PMID: 8790475
  15. Comprehensive evaluation of protein structure alignment methods: scoring by geometric measures.
    J Mol Biol. 2005 Mar 4;346(4):1173-88 PMID: 15701525
  16. BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark.
    Proteins. 2005 Oct 1;61(1):127-36 PMID: 16044462
  17. Recent evolutions of multiple sequence alignment algorithms.
    PLoS Comput Biol. 2007 Aug;3(8):e123 PMID: 17784778
  18. Kalign--an accurate and fast multiple sequence alignment algorithm.
    BMC Bioinformatics. 2005 Dec 12;6:298 PMID: 16343337
  19. PRALINE: a multiple sequence alignment toolbox that integrates homology-extended and secondary structure information.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W289-94 PMID: 15980472
  20. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  21. Prediction of function divergence in protein families using the substitution rate variation parameter alpha.
    Mol Biol Evol. 2006 Jul;23(7):1406-13 PMID: 16672285
  22. Multiple alignment by aligning alignments.
    Bioinformatics. 2007 Jul 1;23(13):i559-68 PMID: 17646343
  23. M-Coffee: combining multiple sequence alignment methods with T-Coffee.
    Nucleic Acids Res. 2006 Mar 23;34(6):1692-9 PMID: 16556910
  24. ProbCons: Probabilistic consistency-based multiple sequence alignment.
    Genome Res. 2005 Feb;15(2):330-40 PMID: 15687296
  25. The alignment of sets of sequences and the construction of phyletic trees: an integrated method.
    J Mol Evol. 1984;20(2):175-86 PMID: 6433036
  26. MUMMALS: multiple sequence alignment improved by using hidden Markov models with local structural information.
    Nucleic Acids Res. 2006;34(16):4364-74 PMID: 16936316
  27. Evaluation of iterative alignment algorithms for multiple alignment.
    Bioinformatics. 2005 Apr 15;21(8):1408-14 PMID: 15564300
  28. 3DCoffee: combining protein sequences and structures within multiple sequence alignments.
    J Mol Biol. 2004 Jul 2;340(2):385-95 PMID: 15201059
  29. Expresso: automatic incorporation of structural information in multiple sequence alignments using 3D-Coffee.
    Nucleic Acids Res. 2006 Jul 1;34(Web Server issue):W604-8 PMID: 16845081
  30. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  31. APDB: a novel measure for benchmarking sequence alignment methods without reference alignments.
    Bioinformatics. 2003;19 Suppl 1:i215-21 PMID: 12855461
  32. 1000 Genomes project.
    Nat Biotechnol. 2008 Mar;26(3):256 PMID: 18327223
  33. Comparative analysis of multiple protein-sequence alignment methods.
    Mol Biol Evol. 1994 Jul;11(4):571-92 PMID: 8078398
  34. BAliBASE: a benchmark alignment database for the evaluation of multiple alignment programs.
    Bioinformatics. 1999 Jan;15(1):87-8 PMID: 10068696
  35. Analysis and comparison of benchmarks for multiple sequence alignment.
    In Silico Biol. 2006;6(4):321-39 PMID: 16922695
  36. Sequence progressive alignment, a framework for practical large-scale probabilistic consistency alignment.
    Bioinformatics. 2009 Feb 1;25(3):295-301 PMID: 19056777
  37. Quality assessment of multiple alignment programs.
    FEBS Lett. 2002 Oct 2;529(1):126-30 PMID: 12354624
  38. Protein structure alignment by incremental combinatorial extension (CE) of the optimal path.
    Protein Eng. 1998 Sep;11(9):739-47 PMID: 9796821
  39. The iRMSD: a local measure of sequence alignment accuracy using structural information.
    Bioinformatics. 2006 Jul 15;22(14):e35-9 PMID: 16873492
  40. Multiple sequence alignment.
    Curr Opin Struct Biol. 2006 Jun;16(3):368-73 PMID: 16679011
  41. R-Coffee: a method for multiple alignment of non-coding RNA.
    Nucleic Acids Res. 2008 May;36(9):e52 PMID: 18420654
  42. SAGA: sequence alignment by genetic algorithm.
    Nucleic Acids Res. 1996 Apr 15;24(8):1515-24 PMID: 8628686
  43. A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
    Syst Biol. 2003 Oct;52(5):696-704 PMID: 14530136
  44. Significant improvement in accuracy of multiple protein sequence alignments by iterative refinement as assessed by reference to structural alignments.
    J Mol Biol. 1996 Dec 13;264(4):823-38 PMID: 8980688
  45. Consistency of optimal sequence alignments.
    Bull Math Biol. 1990;52(4):509-25 PMID: 1697773
  46. An enhanced RNA alignment benchmark for sequence alignment programs.
    Algorithms Mol Biol. 2006 Oct 24;1:19 PMID: 17062125
  47. Local RNA base pairing probabilities in large sequences.
    Bioinformatics. 2006 Mar 1;22(5):614-5 PMID: 16368769
  48. DIALIGN-TX: greedy and progressive approaches for segment-based multiple sequence alignment.
    Algorithms Mol Biol. 2008 May 27;3:6 PMID: 18505568
  49. DIALIGN-T: an improved algorithm for segment-based multiple sequence alignment.
    BMC Bioinformatics. 2005 Mar 22;6:66 PMID: 15784139
  50. MARNA: multiple alignment and consensus structure prediction of RNAs based on sequence structure comparisons.
    Bioinformatics. 2005 Aug 15;21(16):3352-9 PMID: 15972285
  51. Dali: a network tool for protein structure comparison.
    Trends Biochem Sci. 1995 Nov;20(11):478-80 PMID: 8578593
  52. Identification of protein sequence homology by consensus template alignment.
    J Mol Biol. 1986 Mar 20;188(2):233-58 PMID: 3088284
  53. OXBench: a benchmark for evaluation of protein multiple sequence alignment accuracy.
    BMC Bioinformatics. 2003 Oct 10;4:47 PMID: 14552658
  54. Alignment uncertainty and genomic analysis.
    Science. 2008 Jan 25;319(5862):473-6 PMID: 18218900
  55. A second generation human haplotype map of over 3.1 million SNPs.
    Nature. 2007 Oct 18;449(7164):851-61 PMID: 17943122
  56. A databank (3D-ali) collecting related protein sequences and structures.
    Protein Eng. 1996 Mar;9(3):249-51 PMID: 8736491
  57. MUSCLE: a multiple sequence alignment method with reduced time and space complexity.
    BMC Bioinformatics. 2004 Aug 19;5:113 PMID: 15318951
  58. Multiple sequence alignment using partial order graphs.
    Bioinformatics. 2002 Mar;18(3):452-64 PMID: 11934745
  59. PCMA: fast and accurate multiple sequence alignment based on profile consistency.
    Bioinformatics. 2003 Feb 12;19(3):427-8 PMID: 12584134
  60. Target selection and deselection at the Berkeley Structural Genomics Center.
    Proteins. 2006 Feb 1;62(2):356-70 PMID: 16276528
  61. Multiple DNA and protein sequence alignment based on segment-to-segment comparison.
    Proc Natl Acad Sci U S A. 1996 Oct 29;93(22):12098-103 PMID: 8901539
  62. CaspR: a web server for automated molecular replacement using homology modelling.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W606-9 PMID: 15215460
  63. Generating benchmarks for multiple sequence alignments and phylogenetic reconstructions.
    Proc Int Conf Intell Syst Mol Biol. 1997;5:303-6 PMID: 9322053
  64. A simple genetic algorithm for multiple sequence alignment.
    Genet Mol Res. 2007 Oct 05;6(4):964-82 PMID: 18058716
  65. A tabu search algorithm for post-processing multiple sequence alignment.
    J Bioinform Comput Biol. 2005 Feb;3(1):145-56 PMID: 15751117
  66. Automatic assessment of alignment quality.
    Nucleic Acids Res. 2005 Dec 16;33(22):7120-8 PMID: 16361270
  67. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  68. Recent developments in the MAFFT multiple sequence alignment program.
    Brief Bioinform. 2008 Jul;9(4):286-98 PMID: 18372315
  69. SeqAn an efficient, generic C++ library for sequence analysis.
    BMC Bioinformatics. 2008 Jan 09;9:11 PMID: 18184432
  70. SABmark--a benchmark for sequence alignment that covers the entire known fold space.
    Bioinformatics. 2005 Apr 1;21(7):1267-8 PMID: 15333456
  71. MUSCLE: multiple sequence alignment with high accuracy and high throughput.
    Nucleic Acids Res. 2004 Mar 19;32(5):1792-7 PMID: 15034147
  72. Compression-based classification of biological sequences and structures via the Universal Similarity Metric: experimental assessment.
    BMC Bioinformatics. 2007 Jul 13;8:252 PMID: 17629909
  73. PROMALS3D: a tool for multiple protein sequence and structure alignments.
    Nucleic Acids Res. 2008 Apr;36(7):2295-300 PMID: 18287115
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2009-10-01
Epub
2009-00-30
Pages
2455-65
Language
English
Region
England
NLM ID
9808944
PMCID
PMC2752613
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]