Abstract
Two central problems in computational biology are the determination of the alignment and phylogeny of a set of biological sequences. The traditional approach to this problem is to first build a multiple alignment of these sequences, followed by a phylogenetic reconstruction step based on this multiple alignment. However, alignment and phylogenetic inference are fundamentally interdependent, and ignoring this fact leads to biased and overconfident estimations. Whether the main interest be in sequence alignment or phylogeny, a major goal of computational biology is the co-estimation of both. We developed a fully Bayesian Markov chain Monte Carlo method for coestimating phylogeny and sequence alignment, under the Thorne-Kishino-Felsenstein model of substitution and single nucleotide insertion-deletion (indel) events. In our earlier work, we introduced a novel and efficient algorithm, termed the "indel peeling algorithm", which includes indels as phylogenetically informative evolutionary events, and resembles Felsenstein's peeling algorithm for substitutions on a phylogenetic tree. For a fixed alignment, our extension analytically integrates out both substitution and indel events within a proper statistical model, without the need for data augmentation at internal tree nodes, allowing for efficient sampling of tree topologies and edge lengths. To additionally sample multiple alignments, we here introduce an efficient partial Metropolized independence sampler for alignments, and combine these two algorithms into a fully Bayesian co-estimation procedure for the alignment and phylogeny problem. Our approach results in estimates for the posterior distribution of evolutionary rate parameters, for the maximum a-posteriori (MAP) phylogenetic tree, and for the posterior decoding alignment. Estimates for the evolutionary tree and multiple alignment are augmented with confidence estimates for each node height and alignment column. Our results indicate that the patterns in reliability broadly correspond to structural features of the proteins, and thus provides biologically meaningful information which is not existent in the usual point-estimate of the alignment. Our methods can handle input data of moderate size (10-20 protein sequences, each 100-200 bp), which we analyzed overnight on a standard 2 GHz personal computer. Joint analysis of multiple sequence alignment, evolutionary trees and additional evolutionary parameters can be now done within a single coherent statistical framework.
MeSH Terms
Algorithms
Amino Acid Sequence
Animals
Base Sequence
Bayes Theorem
Computational Biology/methods
Computer Simulation
Evolution, Molecular
Gene Deletion
Humans
Likelihood Functions
Markov Chains
Models, Genetic
Models, Statistical
Molecular Sequence Data
Monte Carlo Method
Mutation
Myoglobin/chemistry
Phylogeny
Polymorphism, Single Nucleotide/genetics
Sequence Alignment
Sequence Analysis, DNA
Sequence Homology, Amino Acid
Software
Species Specificity
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Lunter Gerton
Department of Statistics, University of Oxford, 1 South Parks Road, Oxford OX1 3TG, UK.
[email protected]
Miklós István
Drummond Alexei
Jensen Jens Ledet
Hein Jotun
References (26)
26 references, click to expand
-
A molecular phylogeny of reptiles.
Science. 1999 Feb 12;283(5404):998-1001
PMID: 9974396
-
Dynamic programming alignment accuracy.
J Comput Biol. 1998 Fall;5(3):493-504
PMID: 9773345
-
An integrated framework for the inference of viral population history from reconstructed genealogies.
Genetics. 2000 Jul;155(3):1429-37
PMID: 10880500
-
T-Coffee: A novel method for fast and accurate multiple sequence alignment.
J Mol Biol. 2000 Sep 8;302(1):205-17
PMID: 10964570
-
Statistical alignment: computational properties, homology testing and goodness-of-fit.
J Mol Biol. 2000 Sep 8;302(1):265-79
PMID: 10964574
-
An algorithm for statistical alignment of sequences related by a binary tree.
Pac Symp Biocomput. 2001;:179-90
PMID: 11262938
-
Molecular phylogenetics: state-of-the-art methods for looking into the past.
Trends Genet. 2001 May;17(5):262-72
PMID: 11335036
-
MRBAYES: Bayesian inference of phylogenetic trees.
Bioinformatics. 2001 Aug;17(8):754-5
PMID: 11524383
-
Evolutionary HMMs: a Bayesian approach to multiple alignment.
Bioinformatics. 2001 Sep;17(9):803-20
PMID: 11590097
-
Assessing variability by joint sampling of alignments and mutation rates.
J Mol Evol. 2001 Dec;53(6):660-9
PMID: 11677626
-
Estimating mutation parameters, population history and genealogy simultaneously from temporally spaced sequence data.
Genetics. 2002 Jul;161(3):1307-20
PMID: 12136032
-
An improved algorithm for statistical alignment of sequences related by a star tree.
Bull Math Biol. 2002 Jul;64(4):771-9
PMID: 12216420
-
Statistical alignment based on fragment insertion and deletion models.
Bioinformatics. 2003 Mar 1;19(4):490-9
PMID: 12611804
-
The epidemiology and iatrogenic transmission of hepatitis C virus in Egypt: a Bayesian coalescent approach.
Mol Biol Evol. 2003 Mar;20(3):381-7
PMID: 12644558
-
Recursions for statistical multiple alignment.
Proc Natl Acad Sci U S A. 2003 Dec 9;100(25):14960-5
PMID: 14657378
-
An efficient algorithm for statistical multiple alignment on arbitrary phylogenetic trees.
J Comput Biol. 2003;10(6):869-89
PMID: 14980015
-
A "Long Indel" model for evolutionary sequence alignment.
Mol Biol Evol. 2004 Mar;21(3):529-40
PMID: 14694074
-
Evolution of 5S RNA and the non-randomness of base replacement.
Nat New Biol. 1973 Oct 24;245(147):232-4
PMID: 4201431
-
Evolutionary trees from DNA sequences: a maximum likelihood approach.
J Mol Evol. 1981;17(6):368-76
PMID: 7288891
-
Maximum likelihood alignment of DNA sequences.
J Mol Biol. 1986 Jul 20;190(2):159-65
PMID: 3641921
-
An evolutionary model for maximum likelihood alignment of DNA sequences.
J Mol Evol. 1991 Aug;33(2):114-24
PMID: 1920447
-
Estimating effective population size from samples of sequences: inefficiency of pairwise and segregating sites as compared to phylogenetic estimates.
Genet Res. 1992 Apr;59(2):139-47
PMID: 1628818
-
CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
Nucleic Acids Res. 1994 Nov 11;22(22):4673-80
PMID: 7984417
-
Estimating effective population size and mutation rate from sequence data using Metropolis-Hastings sampling.
Genetics. 1995 Aug;140(4):1421-30
PMID: 7498781
-
Genealogical inference from microsatellite data.
Genetics. 1998 Sep;150(1):499-510
PMID: 9725864
-
A probabilistic model for the evolution of RNA structure.
BMC Bioinformatics. 2004 Oct 26;5:166
PMID: 15507142