Home LiteratureArticle Details
PMID: 19087239 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Pitfalls of the most commonly used models of context dependent substitution.

Biology direct ·Vol. 3 ·2008-12-16 ·Pages 52

Lindsay H, Yap VB, Ying H, Huttley GA

Abstract

Neighboring nucleotides exert a striking influence on mutation, with the hypermutability of CpG dinucleotides in many genomes being an exemplar. Among the approaches employed to measure the relative importance of sequence neighbors on molecular evolution have been continuous-time Markov process models for substitutions that treat sequences as a series of independent tuples. The most widely used examples are the codon substitution models. We evaluated the suitability of derivatives of the nucleotide frequency weighted (hereafter NF) and tuple frequency weighted (hereafter TF) models for measuring sequence context dependent substitution. Critical properties we address are their relationships to an independent nucleotide process and the robustness of parameter estimation to changes in sequence composition. We then consider the impact on inference concerning dinucleotide substitution processes from application of these two forms to intron sequence alignments from primates. We prove that the NF form always nests the independent nucleotide process and that this is not true for the TF form. As a consequence, using TF to study context effects can be misleading, which is shown by both theoretical calculations and simulations. We describe a simple example where a context parameter estimated under TF is confounded with composition terms unless all sequence states are equi-frequent. We illustrate this for the dinucleotide case by simulation under a nucleotide model, showing that the TF form identifies a CpG effect when none exists. Our analysis of primate introns revealed that the effect of nucleotide neighbors is over-estimated under TF compared with NF. Parameter estimates for a number of contexts are also strikingly discordant between the two model forms. Our results establish that the NF form should be used for analysis of independent-tuple context dependent processes. Although neighboring effects in general are still important, prominent influences such as the elevated CpG transversion rate previously identified using the TF form are an artifact. Our results further suggest as few as 5 parameters may account for approximately 85% of neighboring nucleotide influence.

MeSH Terms
Amino Acid Substitution/genetics Animals CpG Islands/genetics Introns/genetics Likelihood Functions Models, Genetic Nucleotides/genetics Primates/genetics Sequence Alignment
Chemicals
Nucleotides
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Lindsay Helen
Computational Genomics Laboratory, John Curtin School of Medical Research, The Australian National University, Canberra, Australia. [email protected]
Yap Von Bing
Ying Hua
Huttley Gavin A
References (34)
34 references, click to expand
  1. A codon-based model of nucleotide substitution for protein-coding DNA sequences.
    Mol Biol Evol. 1994 Sep;11(5):725-36 PMID: 7968486
  2. Likelihood models for detecting positively selected amino acid sites and applications to the HIV-1 envelope gene.
    Genetics. 1998 Mar;148(3):929-36 PMID: 9539414
  3. Evolutionary trees from DNA sequences: a maximum likelihood approach.
    J Mol Evol. 1981;17(6):368-76 PMID: 7288891
  4. CpG-rich islands and the function of DNA methylation.
    Nature. 1986 May 15-21;321(6067):209-13 PMID: 2423876
  5. DNA repair in an active gene: removal of pyrimidine dimers from the DHFR gene of CHO cells is much more efficient than in the genome overall.
    Cell. 1985 Feb;40(2):359-69 PMID: 3838150
  6. Distinct changes of genomic biases in nucleotide substitution at the time of Mammalian radiation.
    Mol Biol Evol. 2003 Nov;20(11):1887-96 PMID: 12885958
  7. Nucleotide excision repair activity varies among murine spermatogenic cell types.
    Biol Reprod. 2005 Jul;73(1):123-30 PMID: 15758148
  8. Phylogenetic estimation of context-dependent substitution rates by maximum likelihood.
    Mol Biol Evol. 2004 Mar;21(3):468-88 PMID: 14660683
  9. Theoretical analysis of mutation hotspots and their DNA sequence context specificity.
    Mutat Res. 2003 Sep;544(1):65-85 PMID: 12888108
  10. Structure and function of eukaryotic DNA methyltransferases.
    Curr Top Dev Biol. 2004;60:55-89 PMID: 15094296
  11. A simple model based on mutation and selection explains trends in codon and amino-acid usage and GC composition within and across genomes.
    Genome Biol. 2001;2(4):RESEARCH0010 PMID: 11305938
  12. MrBayes 3: Bayesian phylogenetic inference under mixed models.
    Bioinformatics. 2003 Aug 12;19(12):1572-4 PMID: 12912839
  13. Performance of maximum parsimony and likelihood phylogenetics when evolution is heterogeneous.
    Nature. 2004 Oct 21;431(7011):980-4 PMID: 15496922
  14. HyPhy: hypothesis testing using phylogenies.
    Bioinformatics. 2005 Mar 1;21(5):676-9 PMID: 15509596
  15. PyEvolve: a toolkit for statistical modelling of molecular evolution.
    BMC Bioinformatics. 2004 Jan 05;5:1 PMID: 14706121
  16. The CpG dinucleotide and human genetic disease.
    Hum Genet. 1988 Feb;78(2):151-5 PMID: 3338800
  17. Transcription-associated mutational asymmetry in mammalian evolution.
    Nat Genet. 2003 Apr;33(4):514-7 PMID: 12612582
  18. Ensembl 2006.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D556-61 PMID: 16381931
  19. Evolutionary analyses of DNA sequences subject to constraints of secondary structure.
    Genetics. 1995 Mar;139(3):1429-39 PMID: 7768450
  20. Neighboring-nucleotide effects on the rates of germ-line single-base-pair substitution in human genes.
    Am J Hum Genet. 1998 Aug;63(2):474-88 PMID: 9683596
  21. Molecular basis of base substitution hotspots in Escherichia coli.
    Nature. 1978 Aug 24;274(5673):775-80 PMID: 355893
  22. Modeling the impact of DNA methylation on the evolution of BRCA1 in mammals.
    Mol Biol Evol. 2004 Sep;21(9):1760-8 PMID: 15190129
  23. Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.
    Genome Res. 2005 Aug;15(8):1034-50 PMID: 16024819
  24. A dependent-rates model and an MCMC-based methodology for the maximum-likelihood analysis of sequences with overlapping reading frames.
    Mol Biol Evol. 2001 May;18(5):763-76 PMID: 11319261
  25. A stochastic model for the evolution of autocorrelated DNA sequences.
    Mol Phylogenet Evol. 1994 Sep;3(3):240-7 PMID: 7529616
  26. Mutagenic specificity of ultraviolet light.
    J Mol Biol. 1985 Mar 5;182(1):45-65 PMID: 3923204
  27. Bayesian Markov chain Monte Carlo sequence analysis reveals varying neutral substitution patterns in mammalian evolution.
    Proc Natl Acad Sci U S A. 2004 Sep 28;101(39):13994-4001 PMID: 15292512
  28. A new method for calculating evolutionary substitution rates.
    J Mol Evol. 1984;20(1):86-93 PMID: 6429346
  29. PAML: a program package for phylogenetic analysis by maximum likelihood.
    Comput Appl Biosci. 1997 Oct;13(5):555-6 PMID: 9367129
  30. Large-scale analyses of synonymous substitution rates can be sensitive to assumptions about the process of mutation.
    Gene. 2006 Aug 15;378:58-64 PMID: 16797879
  31. From context-dependence of mutations to molecular mechanisms of mutagenesis.
    Pac Symp Biocomput. 2005;:409-20 PMID: 15759646
  32. Maximum-likelihood estimation of phylogeny from DNA sequences when substitution rates differ over sites.
    Mol Biol Evol. 1993 Nov;10(6):1396-401 PMID: 8277861
  33. PyCogent: a toolkit for making sense from sequence.
    Genome Biol. 2007;8(8):R171 PMID: 17708774
  34. A likelihood approach for comparing synonymous and nonsynonymous nucleotide substitution rates, with application to the chloroplast genome.
    Mol Biol Evol. 1994 Sep;11(5):715-24 PMID: 7968485
Article Info
Journal
Biology direct
Abbr.
Biol Direct
ISSN
1745-6150
Published
2008-12-16
Epub
2008-00-16
Pages
52
Language
English
Region
England
NLM ID
101258412
PMCID
PMC2628887
Subset
IM
Corrections
ErratumIn
-
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]