Abstract
Chromatin immunoprecipitation (ChIP), coupled with massively parallel short-read sequencing (seq) is used to probe chromatin dynamics. Although there are many algorithms to call peaks from ChIP-seq datasets, most are tuned either to handle punctate sites, such as transcriptional factor binding sites, or broad regions, such as histone modification marks; few can do both. Other algorithms are limited in their configurability, performance on large data sets, and ability to distinguish closely-spaced peaks. In this paper, we introduce PeakRanger, a peak caller software package that works equally well on punctate and broad sites, can resolve closely-spaced peaks, has excellent performance, and is easily customized. In addition, PeakRanger can be run in a parallel cloud computing environment to obtain extremely high performance on very large data sets. We present a series of benchmarks to evaluate PeakRanger against 10 other peak callers, and demonstrate the performance of PeakRanger on both real and synthetic data sets. We also present real world usages of PeakRanger, including peak-calling in the modENCODE project. Compared to other peak callers tested, PeakRanger offers improved resolution in distinguishing extremely closely-spaced peaks. PeakRanger has above-average spatial accuracy in terms of identifying the precise location of binding events. PeakRanger also has excellent sensitivity and specificity in all benchmarks evaluated. In addition, PeakRanger offers significant improvements in run time when running on a single processor system, and very marked improvements when allowed to take advantage of the MapReduce parallel environment offered by a cloud computing resource. PeakRanger can be downloaded at the official site of modENCODE project: http://www.modencode.org/software/ranger/
MeSH Terms
Algorithms
Base Sequence
Chromatin/chemistry
Chromatin Assembly and Disassembly
Chromatin Immunoprecipitation/methods,standards
Histone Code
Protein Binding
Sensitivity and Specificity
Sequence Analysis, DNA/methods,standards
Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Feng Xin
Department of Biomedical Engineering, Stony Brook University, Stony Brook, NY 11794, USA.
[email protected]
Grossman Robert
Stein Lincoln
References (34)
34 references, click to expand
-
Combining evidence using p-values: application to sequence homology searches.
Bioinformatics. 1998;14(1):48-54
PMID: 9520501
-
F-Seq: a feature density estimator for high-throughput sequence tags.
Bioinformatics. 2008 Nov 1;24(21):2537-8
PMID: 18784119
-
Genomic binding sites of the yeast cell-cycle transcription factors SBF and MBF.
Nature. 2001 Jan 25;409(6819):533-8
PMID: 11206552
-
Genome-wide histone acetylation data improve prediction of mammalian transcription factor binding sites.
Bioinformatics. 2010 Sep 1;26(17):2071-5
PMID: 20663846
-
Genome-wide maps of chromatin state in pluripotent and lineage-committed cells.
Nature. 2007 Aug 2;448(7153):553-60
PMID: 17603471
-
Genome-wide location and function of DNA binding proteins.
Science. 2000 Dec 22;290(5500):2306-9
PMID: 11125145
-
Unlocking the secrets of the genome.
Nature. 2009 Jun 18;459(7249):927-30
PMID: 19536255
-
Nucleosome dynamics define transcriptional enhancers.
Nat Genet. 2010 Apr;42(4):343-7
PMID: 20208536
-
High-resolution profiling of histone methylations in the human genome.
Cell. 2007 May 18;129(4):823-37
PMID: 17512414
-
Model-based analysis of ChIP-Seq (MACS).
Genome Biol. 2008;9(9):R137
PMID: 18798982
-
Empirical methods for controlling false positives and estimating confidence in ChIP-Seq peaks.
BMC Bioinformatics. 2008 Dec 05;9:523
PMID: 19061503
-
Histone modifications at human enhancers reflect global cell-type-specific gene expression.
Nature. 2009 May 7;459(7243):108-12
PMID: 19295514
-
Computation for ChIP-seq and RNA-seq studies.
Nat Methods. 2009 Nov;6(11 Suppl):S22-32
PMID: 19844228
-
Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing.
Nat Methods. 2007 Aug;4(8):651-7
PMID: 17558387
-
Discovering homotypic binding events at high spatial resolution.
Bioinformatics. 2010 Dec 15;26(24):3028-34
PMID: 20966006
-
Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data.
Nat Methods. 2008 Sep;5(9):829-34
PMID: 19160518
-
A blind deconvolution approach to high-resolution mapping of transcription factor binding sites from ChIP-seq data.
Genome Biol. 2009;10(12):R142
PMID: 20028542
-
Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
Genome Biol. 2009;10(3):R25
PMID: 19261174
-
An integrated software system for analyzing ChIP-chip and ChIP-seq data.
Nat Biotechnol. 2008 Nov;26(11):1293-300
PMID: 18978777
-
Sole-Search: an integrated analysis program for peak detection and functional annotation using ChIP-seq data.
Nucleic Acids Res. 2010 Jan;38(3):e13
PMID: 19906703
-
ChIP-seq: advantages and challenges of a maturing technology.
Nat Rev Genet. 2009 Oct;10(10):669-80
PMID: 19736561
-
Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data.
Nucleic Acids Res. 2008 Sep;36(16):5221-31
PMID: 18684996
-
Design and analysis of ChIP-seq experiments for DNA-binding proteins.
Nat Biotechnol. 2008 Dec;26(12):1351-9
PMID: 19029915
-
PeakSeq enables systematic scoring of ChIP-seq experiments relative to controls.
Nat Biotechnol. 2009 Jan;27(1):66-75
PMID: 19122651
-
HPeak: an HMM-based algorithm for defining read-enriched regions in ChIP-Seq data.
BMC Bioinformatics. 2010 Jul 02;11:369
PMID: 20598134
-
The case for cloud computing in genome informatics.
Genome Biol. 2010;11(5):207
PMID: 20441614
-
A clustering approach for identification of enriched domains from histone modification ChIP-Seq data.
Bioinformatics. 2009 Aug 1;25(15):1952-8
PMID: 19505939
-
Mapping and quantifying mammalian transcriptomes by RNA-Seq.
Nat Methods. 2008 Jul;5(7):621-8
PMID: 18516045
-
Extracting transcription factor targets from ChIP-Seq data.
Nucleic Acids Res. 2009 Sep;37(17):e113
PMID: 19553195
-
Evaluation of algorithm performance in ChIP-seq peak detection.
PLoS One. 2010 Jul 08;5(7):e11471
PMID: 20628599
-
Integrative analysis of the Caenorhabditis elegans genome by the modENCODE project.
Science. 2010 Dec 24;330(6012):1775-87
PMID: 21177976
-
Genome-wide mapping of in vivo protein-DNA interactions.
Science. 2007 Jun 8;316(5830):1497-502
PMID: 17540862
-
FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology.
Bioinformatics. 2008 Aug 1;24(15):1729-30
PMID: 18599518
-
The Sequence Alignment/Map format and SAMtools.
Bioinformatics. 2009 Aug 15;25(16):2078-9
PMID: 19505943