Home LiteratureArticle Details
PMID: 19772557 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

BayesPeak: Bayesian analysis of ChIP-seq data.

BMC bioinformatics ·Vol. 10 ·2009-09-21 ·Pages 299

Spyrou C, Stark R, Lynch AG, Tavaré S

Abstract

High-throughput sequencing technology has become popular and widely used to study protein and DNA interactions. Chromatin immunoprecipitation, followed by sequencing of the resulting samples, produces large amounts of data that can be used to map genomic features such as transcription factor binding sites and histone modifications. Our proposed statistical algorithm, BayesPeak, uses a fully Bayesian hidden Markov model to detect enriched locations in the genome. The structure accommodates the natural features of the Solexa/Illumina sequencing data and allows for overdispersion in the abundance of reads in different regions. Moreover, a control sample can be incorporated in the analysis to account for experimental and sequence biases. Markov chain Monte Carlo algorithms are applied to estimate the posterior distributions of the model parameters, and posterior probabilities are used to detect the sites of interest. We have presented a flexible approach for identifying peaks from ChIP-seq reads, suitable for use on both transcription factor binding and histone modification data. Our method estimates probabilities of enrichment that can be used in downstream analysis. The method is assessed using experimentally verified data and is shown to provide high-confidence calls with low false positive rates.

MeSH Terms
Bayes Theorem Binding Sites Chromatin Immunoprecipitation/methods Computational Biology/methods DNA/chemistry,metabolism Proteins/metabolism
Chemicals
Proteins DNA
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Spyrou Christiana
Statistical Laboratory, Centre for Mathematical Sciences, Wilberforce Road, Cambridge, UK. [email protected]
Stark Rory
Lynch Andy G
Tavaré Simon
References (27)
27 references, click to expand
  1. Hierarchical hidden Markov model with application to joint analysis of ChIP-chip and ChIP-seq data.
    Bioinformatics. 2009 Jul 15;25(14):1715-21 PMID: 19447789
  2. A new generation of JASPAR, the open-access repository for transcription factor binding site profiles.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D95-7 PMID: 16381983
  3. Species-specific transcription in mice carrying human chromosome 21.
    Science. 2008 Oct 17;322(5900):434-8 PMID: 18787134
  4. Transcript mapping with high-density oligonucleotide tiling arrays.
    Bioinformatics. 2006 Aug 15;22(16):1963-70 PMID: 16787969
  5. Genome-wide analysis of protein-DNA interactions.
    Annu Rev Genomics Hum Genet. 2006;7:81-102 PMID: 16722805
  6. Genome-wide maps of chromatin state in pluripotent and lineage-committed cells.
    Nature. 2007 Aug 2;448(7153):553-60 PMID: 17603471
  7. A hidden Markov model approach for determining expression from genomic tiling micro arrays.
    BMC Bioinformatics. 2006 May 03;7:239 PMID: 16672042
  8. Comparative genomics modeling of the NRSF/REST repressor network: from single conserved sites to genome-wide repertoire.
    Genome Res. 2006 Oct;16(10):1208-21 PMID: 16963704
  9. Genome-scale validation of deep-sequencing libraries.
    PLoS One. 2008;3(11):e3713 PMID: 19002256
  10. Empirical methods for controlling false positives and estimating confidence in ChIP-Seq peaks.
    BMC Bioinformatics. 2008 Dec 05;9:523 PMID: 19061503
  11. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  12. Detection of functional DNA motifs via statistical over-representation.
    Nucleic Acids Res. 2004 Feb 26;32(4):1372-81 PMID: 14988425
  13. Genome-wide mapping of in vivo protein-DNA interactions.
    Science. 2007 Jun 8;316(5830):1497-502 PMID: 17540862
  14. FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology.
    Bioinformatics. 2008 Aug 1;24(15):1729-30 PMID: 18599518
  15. Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data.
    Nat Methods. 2008 Sep;5(9):829-34 PMID: 19160518
  16. Parameter estimation for robust HMM analysis of ChIP-chip data.
    BMC Bioinformatics. 2008 Aug 18;9:343 PMID: 18706106
  17. A supervised hidden markov model framework for efficiently segmenting tiling array data in transcriptional and chIP-chip experiments: systematically incorporating validated biological knowledge.
    Bioinformatics. 2006 Dec 15;22(24):3016-24 PMID: 17038339
  18. Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data.
    Nucleic Acids Res. 2008 Sep;36(16):5221-31 PMID: 18684996
  19. ChIP-seq: welcome to the new frontier.
    Nat Methods. 2007 Aug;4(8):613-4 PMID: 17664943
  20. TileMap: create chromosomal map of tiling array hybridizations.
    Bioinformatics. 2005 Sep 15;21(18):3629-36 PMID: 16046496
  21. A hidden Markov model for analyzing ChIP-chip experiments on genome tiling arrays and its application to p53 binding sequences.
    Bioinformatics. 2005 Jun;21 Suppl 1:i274-82 PMID: 15961467
  22. Design and analysis of ChIP-seq experiments for DNA-binding proteins.
    Nat Biotechnol. 2008 Dec;26(12):1351-9 PMID: 19029915
  23. PeakSeq enables systematic scoring of ChIP-seq experiments relative to controls.
    Nat Biotechnol. 2009 Jan;27(1):66-75 PMID: 19122651
  24. Tissue-specific transcriptional regulation has diverged significantly between human and mouse.
    Nat Genet. 2007 Jun;39(6):730-2 PMID: 17529977
  25. Model-based analysis of ChIP-Seq (MACS).
    Genome Biol. 2008;9(9):R137 PMID: 18798982
  26. The transcriptional program controlled by the stem cell leukemia gene Scl/Tal1 during early embryonic hematopoietic development.
    Blood. 2009 May 28;113(22):5456-65 PMID: 19346495
  27. Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing.
    Nat Methods. 2007 Aug;4(8):651-7 PMID: 17558387
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2009-09-21
Epub
2009-00-21
Pages
299
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2760534
Subset
IM
Grants
Cancer Research UK · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]