Home LiteratureArticle Details
PMID: 21554709 Published · epublish English Journal Article Research Support, U.S. Gov't, Non-P.H.S.

PeakRanger: a cloud-enabled peak caller for ChIP-seq data.

BMC bioinformatics ·Vol. 12 ·2011-05-09 ·Pages 139

Feng X, Grossman R, Stein L

Abstract

Chromatin immunoprecipitation (ChIP), coupled with massively parallel short-read sequencing (seq) is used to probe chromatin dynamics. Although there are many algorithms to call peaks from ChIP-seq datasets, most are tuned either to handle punctate sites, such as transcriptional factor binding sites, or broad regions, such as histone modification marks; few can do both. Other algorithms are limited in their configurability, performance on large data sets, and ability to distinguish closely-spaced peaks. In this paper, we introduce PeakRanger, a peak caller software package that works equally well on punctate and broad sites, can resolve closely-spaced peaks, has excellent performance, and is easily customized. In addition, PeakRanger can be run in a parallel cloud computing environment to obtain extremely high performance on very large data sets. We present a series of benchmarks to evaluate PeakRanger against 10 other peak callers, and demonstrate the performance of PeakRanger on both real and synthetic data sets. We also present real world usages of PeakRanger, including peak-calling in the modENCODE project. Compared to other peak callers tested, PeakRanger offers improved resolution in distinguishing extremely closely-spaced peaks. PeakRanger has above-average spatial accuracy in terms of identifying the precise location of binding events. PeakRanger also has excellent sensitivity and specificity in all benchmarks evaluated. In addition, PeakRanger offers significant improvements in run time when running on a single processor system, and very marked improvements when allowed to take advantage of the MapReduce parallel environment offered by a cloud computing resource. PeakRanger can be downloaded at the official site of modENCODE project: http://www.modencode.org/software/ranger/

MeSH Terms
Algorithms Base Sequence Chromatin/chemistry Chromatin Assembly and Disassembly Chromatin Immunoprecipitation/methods,standards Histone Code Protein Binding Sensitivity and Specificity Sequence Analysis, DNA/methods,standards Software
Chemicals
Chromatin
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Feng Xin
Department of Biomedical Engineering, Stony Brook University, Stony Brook, NY 11794, USA. [email protected]
Grossman Robert
Stein Lincoln
References (34)
34 references, click to expand
  1. Combining evidence using p-values: application to sequence homology searches.
    Bioinformatics. 1998;14(1):48-54 PMID: 9520501
  2. F-Seq: a feature density estimator for high-throughput sequence tags.
    Bioinformatics. 2008 Nov 1;24(21):2537-8 PMID: 18784119
  3. Genomic binding sites of the yeast cell-cycle transcription factors SBF and MBF.
    Nature. 2001 Jan 25;409(6819):533-8 PMID: 11206552
  4. Genome-wide histone acetylation data improve prediction of mammalian transcription factor binding sites.
    Bioinformatics. 2010 Sep 1;26(17):2071-5 PMID: 20663846
  5. Genome-wide maps of chromatin state in pluripotent and lineage-committed cells.
    Nature. 2007 Aug 2;448(7153):553-60 PMID: 17603471
  6. Genome-wide location and function of DNA binding proteins.
    Science. 2000 Dec 22;290(5500):2306-9 PMID: 11125145
  7. Unlocking the secrets of the genome.
    Nature. 2009 Jun 18;459(7249):927-30 PMID: 19536255
  8. Nucleosome dynamics define transcriptional enhancers.
    Nat Genet. 2010 Apr;42(4):343-7 PMID: 20208536
  9. High-resolution profiling of histone methylations in the human genome.
    Cell. 2007 May 18;129(4):823-37 PMID: 17512414
  10. Model-based analysis of ChIP-Seq (MACS).
    Genome Biol. 2008;9(9):R137 PMID: 18798982
  11. Empirical methods for controlling false positives and estimating confidence in ChIP-Seq peaks.
    BMC Bioinformatics. 2008 Dec 05;9:523 PMID: 19061503
  12. Histone modifications at human enhancers reflect global cell-type-specific gene expression.
    Nature. 2009 May 7;459(7243):108-12 PMID: 19295514
  13. Computation for ChIP-seq and RNA-seq studies.
    Nat Methods. 2009 Nov;6(11 Suppl):S22-32 PMID: 19844228
  14. Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing.
    Nat Methods. 2007 Aug;4(8):651-7 PMID: 17558387
  15. Discovering homotypic binding events at high spatial resolution.
    Bioinformatics. 2010 Dec 15;26(24):3028-34 PMID: 20966006
  16. Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data.
    Nat Methods. 2008 Sep;5(9):829-34 PMID: 19160518
  17. A blind deconvolution approach to high-resolution mapping of transcription factor binding sites from ChIP-seq data.
    Genome Biol. 2009;10(12):R142 PMID: 20028542
  18. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  19. An integrated software system for analyzing ChIP-chip and ChIP-seq data.
    Nat Biotechnol. 2008 Nov;26(11):1293-300 PMID: 18978777
  20. Sole-Search: an integrated analysis program for peak detection and functional annotation using ChIP-seq data.
    Nucleic Acids Res. 2010 Jan;38(3):e13 PMID: 19906703
  21. ChIP-seq: advantages and challenges of a maturing technology.
    Nat Rev Genet. 2009 Oct;10(10):669-80 PMID: 19736561
  22. Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data.
    Nucleic Acids Res. 2008 Sep;36(16):5221-31 PMID: 18684996
  23. Design and analysis of ChIP-seq experiments for DNA-binding proteins.
    Nat Biotechnol. 2008 Dec;26(12):1351-9 PMID: 19029915
  24. PeakSeq enables systematic scoring of ChIP-seq experiments relative to controls.
    Nat Biotechnol. 2009 Jan;27(1):66-75 PMID: 19122651
  25. HPeak: an HMM-based algorithm for defining read-enriched regions in ChIP-Seq data.
    BMC Bioinformatics. 2010 Jul 02;11:369 PMID: 20598134
  26. The case for cloud computing in genome informatics.
    Genome Biol. 2010;11(5):207 PMID: 20441614
  27. A clustering approach for identification of enriched domains from histone modification ChIP-Seq data.
    Bioinformatics. 2009 Aug 1;25(15):1952-8 PMID: 19505939
  28. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  29. Extracting transcription factor targets from ChIP-Seq data.
    Nucleic Acids Res. 2009 Sep;37(17):e113 PMID: 19553195
  30. Evaluation of algorithm performance in ChIP-seq peak detection.
    PLoS One. 2010 Jul 08;5(7):e11471 PMID: 20628599
  31. Integrative analysis of the Caenorhabditis elegans genome by the modENCODE project.
    Science. 2010 Dec 24;330(6012):1775-87 PMID: 21177976
  32. Genome-wide mapping of in vivo protein-DNA interactions.
    Science. 2007 Jun 8;316(5830):1497-502 PMID: 17540862
  33. FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology.
    Bioinformatics. 2008 Aug 1;24(15):1729-30 PMID: 18599518
  34. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2011-05-09
Epub
2011-00-09
Pages
139
Language
English
Region
England
NLM ID
100965194
PMCID
PMC3103446
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]