Abstract
We present a mixture model-based analysis for identifying differences in the distribution of RNA polymerase II (Pol II) in transcribed regions, measured using ChIP-seq (chromatin immunoprecipitation following massively parallel sequencing technology). The statistical model assumes that the number of Pol II-targeted sequences contained within each genomic region follows a Poisson distribution. A Poisson mixture model was then developed to distinguish Pol II binding changes in transcribed region using an empirical approach and an expectation-maximization (EM) algorithm developed for estimation and inference. In order to achieve a global maximum in the M-step, a particle swarm optimization (PSO) was implemented. We applied this model to Pol II binding data generated from hormone-dependent MCF7 breast cancer cells and antiestrogen-resistant MCF7 breast cancer cells before and after treatment with 17beta-estradiol (E2). We determined that in the hormone-dependent cells, approximately 9.9% (2527) genes showed significant changes in Pol II binding after E2 treatment. However, only approximately 0.7% (172) genes displayed significant Pol II binding changes in E2-treated antiestrogen-resistant cells. These results show that a Poisson mixture model can be used to analyze ChIP-seq data.
MeSH Terms
Algorithms
Bayes Theorem
Cell Line, Tumor
Chromatin Immunoprecipitation
Estradiol/pharmacology
Genome, Human
Humans
Models, Statistical
Neoplasms, Hormone-Dependent/genetics,metabolism
Oligonucleotide Array Sequence Analysis/methods
Poisson Distribution
Protein Binding
RNA Polymerase II/genetics,metabolism
Chemicals
Estradiol
RNA Polymerase II
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Feng Weixing
Division of Biostatistics, Indiana University School of Medicine, Indianapolis, IN 46202, USA.
[email protected]
Liu Yunlong
Wu Jiejun
Nephew Kenneth P
Huang Tim H M
Li Lang
References (10)
10 references, click to expand
-
Genome-wide analysis of estrogen receptor binding sites.
Nat Genet. 2006 Nov;38(11):1289-97
PMID: 17013392
-
Genome-wide mapping of in vivo protein-DNA interactions.
Science. 2007 Jun 8;316(5830):1497-502
PMID: 17540862
-
Assessing gene significance from cDNA microarray expression data via mixed models.
J Comput Biol. 2001;8(6):625-37
PMID: 11747616
-
Gamma-Normal-Gamma mixture model for detecting differentially methylated loci in three breast cancer cell lines.
Cancer Inform. 2007 Feb 07;3:43-54
PMID: 19455234
-
Analysis of variance for gene expression microarray data.
J Comput Biol. 2000;7(6):819-37
PMID: 11382364
-
Epigenetic hypothesis tests for methylation and acetylation in a triple microarray system.
J Comput Biol. 2005 Apr;12(3):370-90
PMID: 15857248
-
A mixture model-based discriminate analysis for identifying ordered transcription factor binding site pairs in gene promoters directly regulated by estrogen receptor-alpha.
Bioinformatics. 2006 Sep 15;22(18):2210-6
PMID: 16809387
-
On differential variability of expression ratios: improving statistical inference about gene expression changes from microarray data.
J Comput Biol. 2001;8(1):37-52
PMID: 11339905
-
Chromatin immunoprecipitation and microarray-based analysis of protein location.
Nat Protoc. 2006;1(2):729-48
PMID: 17406303
-
Diverse gene expression and DNA methylation profiles correlate with differential adaptation of breast cancer cells to the antiestrogens tamoxifen and fulvestrant.
Cancer Res. 2006 Dec 15;66(24):11954-66
PMID: 17178894