Home LiteratureArticle Details
PMID: 19478997 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

Fast statistical alignment.

PLoS computational biology ·Vol. 5 ·No. 5 ·2009-05-00 ·Pages e1000392

Bradley RK, Roberts A, Smoot M, Juvekar S, Do J, Dewey C, Holmes I, Pachter L

Abstract

We describe a new program for the alignment of multiple biological sequences that is both statistically motivated and fast enough for problem sizes that arise in practice. Our Fast Statistical Alignment program is based on pair hidden Markov models which approximate an insertion/deletion process on a tree and uses a sequence annealing algorithm to combine the posterior probabilities estimated from these models into a multiple alignment. FSA uses its explicit statistical model to produce multiple alignments which are accompanied by estimates of the alignment accuracy and uncertainty for every column and character of the alignment--previously available only with alignment programs which use computationally-expensive Markov Chain Monte Carlo approaches--yet can align thousands of long sequences. Moreover, FSA utilizes an unsupervised query-specific learning procedure for parameter estimation which leads to improved accuracy on benchmark reference alignments in comparison to existing programs. The centroid alignment approach taken by FSA, in combination with its learning procedure, drastically reduces the amount of false-positive alignment on biological data in comparison to that given by other methods. The FSA program and a companion visualization tool for exploring uncertainty in alignments can be used via a web interface at http://orangutan.math.berkeley.edu/fsa/, and the source code is available at http://fsa.sourceforge.net/.

MeSH Terms
Algorithms Amino Acid Sequence Animals Artificial Intelligence Base Sequence Data Interpretation, Statistical Databases, Genetic Humans Markov Chains Models, Genetic Molecular Sequence Data Sensitivity and Specificity Sequence Alignment/methods Sequence Analysis Software
Authors & Affiliations
8 authors, click to expand affiliations / ORCID
Bradley Robert K
Department of Mathematics, University of California Berkeley, Berkeley, California, United States of America. [email protected]
Roberts Adam
Smoot Michael
Juvekar Sudeep
Do Jaeyoung
Dewey Colin
Holmes Ian
Pachter Lior
References (53)
53 references, click to expand
  1. Using guide trees to construct multiple-sequence evolutionary HMMs.
    Bioinformatics. 2003;19 Suppl 1:i147-57 PMID: 12855451
  2. An efficient algorithm for statistical multiple alignment on arbitrary phylogenetic trees.
    J Comput Biol. 2003;10(6):869-89 PMID: 14980015
  3. HMMoC--a compiler for hidden Markov models.
    Bioinformatics. 2007 Sep 15;23(18):2485-7 PMID: 17623703
  4. Greengenes, a chimera-checked 16S rRNA gene database and workbench compatible with ARB.
    Appl Environ Microbiol. 2006 Jul;72(7):5069-72 PMID: 16820507
  5. MUMMALS: multiple sequence alignment improved by using hidden Markov models with local structural information.
    Nucleic Acids Res. 2006;34(16):4364-74 PMID: 16936316
  6. Transducers: an emerging probabilistic framework for modeling indels on trees.
    Bioinformatics. 2007 Dec 1;23(23):3258-62 PMID: 17804440
  7. Enredo and Pecan: genome-wide mammalian consistency-based multiple alignment with paralogs.
    Genome Res. 2008 Nov;18(11):1814-28 PMID: 18849524
  8. Comprehensive study on iterative algorithms of multiple sequence alignment.
    Comput Appl Biosci. 1995 Feb;11(1):13-8 PMID: 7796270
  9. Uncertainty in homology inferences: assessing and improving genomic sequence alignment.
    Genome Res. 2008 Feb;18(2):298-309 PMID: 18073381
  10. Probalign: multiple sequence alignment using partition function posterior probabilities.
    Bioinformatics. 2006 Nov 15;22(22):2715-21 PMID: 16954142
  11. Phylogeny-aware gap placement prevents errors in sequence alignment and evolutionary analysis.
    Science. 2008 Jun 20;320(5883):1632-5 PMID: 18566285
  12. Direct evidence of extensive diversity of HIV-1 in Kinshasa by 1960.
    Nature. 2008 Oct 2;455(7213):661-4 PMID: 18833279
  13. Segment-based multiple sequence alignment.
    Bioinformatics. 2008 Aug 15;24(16):i187-92 PMID: 18689823
  14. Efficient pairwise RNA structure prediction and alignment using sequence alignment constraints.
    BMC Bioinformatics. 2006 Sep 04;7:400 PMID: 16952317
  15. Clustal W and Clustal X version 2.0.
    Bioinformatics. 2007 Nov 1;23(21):2947-8 PMID: 17846036
  16. Multiple alignment by sequence annealing.
    Bioinformatics. 2007 Jan 15;23(2):e24-9 PMID: 17237099
  17. Tools for simulating evolution of aligned genomic regions with integrated parameter estimation.
    Genome Biol. 2008 Oct 08;9(10):R147 PMID: 18840304
  18. Specific alignment of structured RNA: stochastic grammars and sequence annealing.
    Bioinformatics. 2008 Dec 1;24(23):2677-83 PMID: 18796475
  19. MAVID: constrained ancestral alignment of multiple sequences.
    Genome Res. 2004 Apr;14(4):693-9 PMID: 15060012
  20. MORPH: probabilistic alignment combined with hidden Markov models of cis-regulatory modules.
    PLoS Comput Biol. 2007 Nov;3(11):e216 PMID: 17997594
  21. BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark.
    Proteins. 2005 Oct 1;61(1):127-36 PMID: 16044462
  22. Versatile and open software for comparing large genomes.
    Genome Biol. 2004;5(2):R12 PMID: 14759262
  23. Sequencing and comparison of yeast species to identify genes and regulatory elements.
    Nature. 2003 May 15;423(6937):241-54 PMID: 12748633
  24. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  25. Automated generation of heuristics for biological sequence comparison.
    BMC Bioinformatics. 2005 Feb 15;6:31 PMID: 15713233
  26. Aligning multiple genomic sequences with the threaded blockset aligner.
    Genome Res. 2004 Apr;14(4):708-15 PMID: 15060014
  27. ProbCons: Probabilistic consistency-based multiple sequence alignment.
    Genome Res. 2005 Feb;15(2):330-40 PMID: 15687296
  28. The Jalview Java alignment editor.
    Bioinformatics. 2004 Feb 12;20(3):426-7 PMID: 14960472
  29. Evolution of genes and genomes on the Drosophila phylogeny.
    Nature. 2007 Nov 8;450(7167):203-18 PMID: 17994087
  30. Statistical alignment: computational properties, homology testing and goodness-of-fit.
    J Mol Biol. 2000 Sep 8;302(1):265-79 PMID: 10964574
  31. Multiple sequence alignment.
    Curr Opin Struct Biol. 2006 Jun;16(3):368-73 PMID: 16679011
  32. An enhanced RNA alignment benchmark for sequence alignment programs.
    Algorithms Mol Biol. 2006 Oct 24;1:19 PMID: 17062125
  33. DIALIGN-TX: greedy and progressive approaches for segment-based multiple sequence alignment.
    Algorithms Mol Biol. 2008 May 27;3:6 PMID: 18505568
  34. TEXshade: shading and labeling of multiple sequence alignments using LATEX2 epsilon.
    Bioinformatics. 2000 Feb;16(2):135-9 PMID: 10842735
  35. The European ribosomal RNA database.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D101-3 PMID: 14681368
  36. Alignment uncertainty and genomic analysis.
    Science. 2008 Jan 25;319(5862):473-6 PMID: 18218900
  37. MUSCLE: a multiple sequence alignment method with reduced time and space complexity.
    BMC Bioinformatics. 2004 Aug 19;5:113 PMID: 15318951
  38. Multiple sequence alignment using partial order graphs.
    Bioinformatics. 2002 Mar;18(3):452-64 PMID: 11934745
  39. The CHAOS/DIALIGN WWW server for multiple alignment of genomic sequences.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W41-4 PMID: 15215346
  40. BAli-Phy: simultaneous Bayesian inference of alignment and phylogeny.
    Bioinformatics. 2006 Aug 15;22(16):2047-8 PMID: 16679334
  41. An algorithm for statistical alignment of sequences related by a binary tree.
    Pac Symp Biocomput. 2001;:179-90 PMID: 11262938
  42. Indel-based evolutionary distance and mouse-human divergence.
    Genome Res. 2004 Aug;14(8):1610-6 PMID: 15289479
  43. LAGAN and Multi-LAGAN: efficient tools for large-scale multiple alignment of genomic DNA.
    Genome Res. 2003 Apr;13(4):721-31 PMID: 12654723
  44. Multiple DNA and protein sequence alignment based on segment-to-segment comparison.
    Proc Natl Acad Sci U S A. 1996 Oct 29;93(22):12098-103 PMID: 8901539
  45. DNA assembly with gaps (Dawg): simulating sequence evolution.
    Bioinformatics. 2005 Nov 1;21 Suppl 3:iii31-8 PMID: 16306390
  46. Probabilistic phylogenetic inference with insertions and deletions.
    PLoS Comput Biol. 2008 Sep 19;4(9):e1000172 PMID: 18787703
  47. Rfam: annotating non-coding RNAs in complete genomes.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D121-4 PMID: 15608160
  48. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  49. Recent developments in the MAFFT multiple sequence alignment program.
    Brief Bioinform. 2008 Jul;9(4):286-98 PMID: 18372315
  50. Dynamic programming alignment accuracy.
    J Comput Biol. 1998 Fall;5(3):493-504 PMID: 9773345
  51. SABmark--a benchmark for sequence alignment that covers the entire known fold space.
    Bioinformatics. 2005 Apr 1;21(7):1267-8 PMID: 15333456
  52. StatAlign: an extendable software package for joint Bayesian estimation of alignments and evolutionary trees.
    Bioinformatics. 2008 Oct 15;24(20):2403-4 PMID: 18753153
  53. Comparison of genomic DNA sequences: solved and unsolved problems.
    Bioinformatics. 2001 May;17(5):391-7 PMID: 11331233
Article Info
Journal
PLoS computational biology
Abbr.
PLoS Comput Biol
ISSN
1553-7358
Published
2009-05-00
Epub
2009-00-29
Pages
e1000392
Language
English
Region
United States
NLM ID
101238922
PMCID
PMC2684580
Subset
IM
Grants
NIGMS NIH HHS · R01 GM076705 · United States
NIGMS NIH HHS · 1R01GM076705 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]