Home LiteratureArticle Details
PMID: 15804354 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Bayesian coestimation of phylogeny and sequence alignment.

BMC bioinformatics ·Vol. 6 ·2005-04-01 ·Pages 83

Lunter G, Miklós I, Drummond A, Jensen JL, Hein J

Abstract

Two central problems in computational biology are the determination of the alignment and phylogeny of a set of biological sequences. The traditional approach to this problem is to first build a multiple alignment of these sequences, followed by a phylogenetic reconstruction step based on this multiple alignment. However, alignment and phylogenetic inference are fundamentally interdependent, and ignoring this fact leads to biased and overconfident estimations. Whether the main interest be in sequence alignment or phylogeny, a major goal of computational biology is the co-estimation of both. We developed a fully Bayesian Markov chain Monte Carlo method for coestimating phylogeny and sequence alignment, under the Thorne-Kishino-Felsenstein model of substitution and single nucleotide insertion-deletion (indel) events. In our earlier work, we introduced a novel and efficient algorithm, termed the "indel peeling algorithm", which includes indels as phylogenetically informative evolutionary events, and resembles Felsenstein's peeling algorithm for substitutions on a phylogenetic tree. For a fixed alignment, our extension analytically integrates out both substitution and indel events within a proper statistical model, without the need for data augmentation at internal tree nodes, allowing for efficient sampling of tree topologies and edge lengths. To additionally sample multiple alignments, we here introduce an efficient partial Metropolized independence sampler for alignments, and combine these two algorithms into a fully Bayesian co-estimation procedure for the alignment and phylogeny problem. Our approach results in estimates for the posterior distribution of evolutionary rate parameters, for the maximum a-posteriori (MAP) phylogenetic tree, and for the posterior decoding alignment. Estimates for the evolutionary tree and multiple alignment are augmented with confidence estimates for each node height and alignment column. Our results indicate that the patterns in reliability broadly correspond to structural features of the proteins, and thus provides biologically meaningful information which is not existent in the usual point-estimate of the alignment. Our methods can handle input data of moderate size (10-20 protein sequences, each 100-200 bp), which we analyzed overnight on a standard 2 GHz personal computer. Joint analysis of multiple sequence alignment, evolutionary trees and additional evolutionary parameters can be now done within a single coherent statistical framework.

MeSH Terms
Algorithms Amino Acid Sequence Animals Base Sequence Bayes Theorem Computational Biology/methods Computer Simulation Evolution, Molecular Gene Deletion Humans Likelihood Functions Markov Chains Models, Genetic Models, Statistical Molecular Sequence Data Monte Carlo Method Mutation Myoglobin/chemistry Phylogeny Polymorphism, Single Nucleotide/genetics Sequence Alignment Sequence Analysis, DNA Sequence Homology, Amino Acid Software Species Specificity
Chemicals
Myoglobin
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Lunter Gerton
Department of Statistics, University of Oxford, 1 South Parks Road, Oxford OX1 3TG, UK. [email protected]
Miklós István
Drummond Alexei
Jensen Jens Ledet
Hein Jotun
References (26)
26 references, click to expand
  1. A molecular phylogeny of reptiles.
    Science. 1999 Feb 12;283(5404):998-1001 PMID: 9974396
  2. Dynamic programming alignment accuracy.
    J Comput Biol. 1998 Fall;5(3):493-504 PMID: 9773345
  3. An integrated framework for the inference of viral population history from reconstructed genealogies.
    Genetics. 2000 Jul;155(3):1429-37 PMID: 10880500
  4. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  5. Statistical alignment: computational properties, homology testing and goodness-of-fit.
    J Mol Biol. 2000 Sep 8;302(1):265-79 PMID: 10964574
  6. An algorithm for statistical alignment of sequences related by a binary tree.
    Pac Symp Biocomput. 2001;:179-90 PMID: 11262938
  7. Molecular phylogenetics: state-of-the-art methods for looking into the past.
    Trends Genet. 2001 May;17(5):262-72 PMID: 11335036
  8. MRBAYES: Bayesian inference of phylogenetic trees.
    Bioinformatics. 2001 Aug;17(8):754-5 PMID: 11524383
  9. Evolutionary HMMs: a Bayesian approach to multiple alignment.
    Bioinformatics. 2001 Sep;17(9):803-20 PMID: 11590097
  10. Assessing variability by joint sampling of alignments and mutation rates.
    J Mol Evol. 2001 Dec;53(6):660-9 PMID: 11677626
  11. Estimating mutation parameters, population history and genealogy simultaneously from temporally spaced sequence data.
    Genetics. 2002 Jul;161(3):1307-20 PMID: 12136032
  12. An improved algorithm for statistical alignment of sequences related by a star tree.
    Bull Math Biol. 2002 Jul;64(4):771-9 PMID: 12216420
  13. Statistical alignment based on fragment insertion and deletion models.
    Bioinformatics. 2003 Mar 1;19(4):490-9 PMID: 12611804
  14. The epidemiology and iatrogenic transmission of hepatitis C virus in Egypt: a Bayesian coalescent approach.
    Mol Biol Evol. 2003 Mar;20(3):381-7 PMID: 12644558
  15. Recursions for statistical multiple alignment.
    Proc Natl Acad Sci U S A. 2003 Dec 9;100(25):14960-5 PMID: 14657378
  16. An efficient algorithm for statistical multiple alignment on arbitrary phylogenetic trees.
    J Comput Biol. 2003;10(6):869-89 PMID: 14980015
  17. A "Long Indel" model for evolutionary sequence alignment.
    Mol Biol Evol. 2004 Mar;21(3):529-40 PMID: 14694074
  18. Evolution of 5S RNA and the non-randomness of base replacement.
    Nat New Biol. 1973 Oct 24;245(147):232-4 PMID: 4201431
  19. Evolutionary trees from DNA sequences: a maximum likelihood approach.
    J Mol Evol. 1981;17(6):368-76 PMID: 7288891
  20. Maximum likelihood alignment of DNA sequences.
    J Mol Biol. 1986 Jul 20;190(2):159-65 PMID: 3641921
  21. An evolutionary model for maximum likelihood alignment of DNA sequences.
    J Mol Evol. 1991 Aug;33(2):114-24 PMID: 1920447
  22. Estimating effective population size from samples of sequences: inefficiency of pairwise and segregating sites as compared to phylogenetic estimates.
    Genet Res. 1992 Apr;59(2):139-47 PMID: 1628818
  23. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  24. Estimating effective population size and mutation rate from sequence data using Metropolis-Hastings sampling.
    Genetics. 1995 Aug;140(4):1421-30 PMID: 7498781
  25. Genealogical inference from microsatellite data.
    Genetics. 1998 Sep;150(1):499-510 PMID: 9725864
  26. A probabilistic model for the evolution of RNA structure.
    BMC Bioinformatics. 2004 Oct 26;5:166 PMID: 15507142
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2005-04-01
Epub
2005-00-01
Pages
83
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1087833
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]