Home LiteratureArticle Details
PMID: 19844228 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't Review

Computation for ChIP-seq and RNA-seq studies.

Nature methods ·Vol. 6 ·No. 11 Suppl ·2009-11-00 ·Pages S22-32

Pepke S, Wold B, Mortazavi A

Abstract

Genome-wide measurements of protein-DNA interactions and transcriptomes are increasingly done by deep DNA sequencing methods (ChIP-seq and RNA-seq). The power and richness of these counting-based measurements comes at the cost of routinely handling tens to hundreds of millions of reads. Whereas early adopters necessarily developed their own custom computer code to analyze the first ChIP-seq and RNA-seq datasets, a new generation of more sophisticated algorithms and software tools are emerging to assist in the analysis phase of these projects. Here we describe the multilayered analyses of ChIP-seq and RNA-seq datasets, discuss the software packages currently available to perform tasks at each layer and describe some upcoming challenges and features for future analysis tools. We also discuss how software choices and uses are affected by specific aspects of the underlying biology and data structure, including genome size, positional clustering of transcription factor binding sites, transcript discovery and expression quantification.

MeSH Terms
Algorithms Base Sequence Chromatin Immunoprecipitation/methods Gene Expression Profiling/methods RNA/analysis,genetics Sequence Alignment/methods Sequence Analysis, DNA/methods
Chemicals
RNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Pepke Shirley
Center for Advanced Computing Research, California Institute of Technology, Pasadena, California, USA.
Wold Barbara
Mortazavi Ali
References (47)
47 references, click to expand
  1. RNA Pol II accumulates at promoters of growth genes during developmental arrest.
    Science. 2009 Apr 3;324(5923):92-4 PMID: 19251593
  2. Variance stabilization applied to microarray data calibration and to the quantification of differential expression.
    Bioinformatics. 2002;18 Suppl 1:S96-104 PMID: 12169536
  3. Genome-wide maps of chromatin state in pluripotent and lineage-committed cells.
    Nature. 2007 Aug 2;448(7153):553-60 PMID: 17603471
  4. Dynamic repertoire of a eukaryotic transcriptome surveyed at single-nucleotide resolution.
    Nature. 2008 Jun 26;453(7199):1239-43 PMID: 18488015
  5. Sequence census methods for functional genomics.
    Nat Methods. 2008 Jan;5(1):19-21 PMID: 18165803
  6. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  7. ChromaSig: a probabilistic approach to finding common chromatin signatures in the human genome.
    PLoS Comput Biol. 2008 Oct;4(10):e1000201 PMID: 18927605
  8. Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing.
    Nat Methods. 2007 Aug;4(8):651-7 PMID: 17558387
  9. FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology.
    Bioinformatics. 2008 Aug 1;24(15):1729-30 PMID: 18599518
  10. Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data.
    Nat Methods. 2008 Sep;5(9):829-34 PMID: 19160518
  11. Detection of single nucleotide variations in expressed exons of the human genome using RNA-Seq.
    Nucleic Acids Res. 2009 Sep;37(16):e106 PMID: 19528076
  12. Next-generation DNA sequencing of paired-end tags (PET) for transcriptome and genome analyses.
    Genome Res. 2009 Apr;19(4):521-32 PMID: 19339662
  13. F-Seq: a feature density estimator for high-throughput sequence tags.
    Bioinformatics. 2008 Nov 1;24(21):2537-8 PMID: 18784119
  14. Digital transcriptome profiling using selective hexamer priming for cDNA synthesis.
    Nat Methods. 2009 Sep;6(9):647-9 PMID: 19668204
  15. A global view of gene activity and alternative splicing by deep sequencing of the human transcriptome.
    Science. 2008 Aug 15;321(5891):956-60 PMID: 18599741
  16. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  17. Transcript length bias in RNA-seq data confounds systems biology.
    Biol Direct. 2009 Apr 16;4:14 PMID: 19371405
  18. An integrated software system for analyzing ChIP-chip and ChIP-seq data.
    Nat Biotechnol. 2008 Nov;26(11):1293-300 PMID: 18978777
  19. Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project.
    Nature. 2007 Jun 14;447(7146):799-816 PMID: 17571346
  20. Statistical inferences for isoform expression in RNA-Seq.
    Bioinformatics. 2009 Apr 15;25(8):1026-32 PMID: 19244387
  21. Comparative analysis of processed pseudogenes in the mouse and human genomes.
    Trends Genet. 2004 Feb;20(2):62-7 PMID: 14746985
  22. Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data.
    Nucleic Acids Res. 2008 Sep;36(16):5221-31 PMID: 18684996
  23. Highly integrated single-base resolution maps of the epigenome in Arabidopsis.
    Cell. 2008 May 2;133(3):523-36 PMID: 18423832
  24. Design and analysis of ChIP-seq experiments for DNA-binding proteins.
    Nat Biotechnol. 2008 Dec;26(12):1351-9 PMID: 19029915
  25. PeakSeq enables systematic scoring of ChIP-seq experiments relative to controls.
    Nat Biotechnol. 2009 Jan;27(1):66-75 PMID: 19122651
  26. Genome-wide identification of human RNA editing sites by parallel DNA capturing and sequencing.
    Science. 2009 May 29;324(5931):1210-3 PMID: 19478186
  27. Alternative isoform regulation in human tissue transcriptomes.
    Nature. 2008 Nov 27;456(7221):470-6 PMID: 18978772
  28. Annotating genomes with massive-scale RNA sequencing.
    Genome Biol. 2008;9(12):R175 PMID: 19087247
  29. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  30. De novo transcriptome assembly with ABySS.
    Bioinformatics. 2009 Nov 1;25(21):2872-7 PMID: 19528083
  31. Stem cell transcriptome profiling via massive-scale mRNA sequencing.
    Nat Methods. 2008 Jul;5(7):613-9 PMID: 18516046
  32. SOAP: short oligonucleotide alignment program.
    Bioinformatics. 2008 Mar 1;24(5):713-4 PMID: 18227114
  33. The transcriptional landscape of the yeast genome defined by RNA sequencing.
    Science. 2008 Jun 6;320(5881):1344-9 PMID: 18451266
  34. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  35. High-resolution profiling of histone methylations in the human genome.
    Cell. 2007 May 18;129(4):823-37 PMID: 17512414
  36. An HMM approach to genome-wide identification of differential histone modification sites from ChIP-seq data.
    Bioinformatics. 2008 Oct 15;24(20):2344-9 PMID: 18667444
  37. Model-based analysis of ChIP-Seq (MACS).
    Genome Biol. 2008;9(9):R137 PMID: 18798982
  38. Empirical methods for controlling false positives and estimating confidence in ChIP-Seq peaks.
    BMC Bioinformatics. 2008 Dec 05;9:523 PMID: 19061503
  39. A hierarchical Bayesian model for comparing transcriptomes at the individual transcript isoform level.
    Nucleic Acids Res. 2009 Jun;37(10):e75 PMID: 19417075
  40. Chromosome Conformation Capture Carbon Copy (5C): a massively parallel solution for mapping interactions between genomic elements.
    Genome Res. 2006 Oct;16(10):1299-309 PMID: 16954542
  41. How to map billions of short reads onto genomes.
    Nat Biotechnol. 2009 May;27(5):455-7 PMID: 19430453
  42. A clustering approach for identification of enriched domains from histone modification ChIP-Seq data.
    Bioinformatics. 2009 Aug 1;25(15):1952-8 PMID: 19505939
  43. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  44. Extracting transcription factor targets from ChIP-Seq data.
    Nucleic Acids Res. 2009 Sep;37(17):e113 PMID: 19553195
  45. Optimal spliced alignments of short sequence reads.
    Bioinformatics. 2008 Aug 15;24(16):i174-80 PMID: 18689821
  46. Genome-wide mapping of in vivo protein-DNA interactions.
    Science. 2007 Jun 8;316(5830):1497-502 PMID: 17540862
  47. RNA-MATE: a recursive mapping strategy for high-throughput RNA-sequencing data.
    Bioinformatics. 2009 Oct 1;25(19):2615-6 PMID: 19648138
Article Info
Journal
Nature methods
Abbr.
Nat Methods
ISSN
1548-7105
Published
2009-11-00
Pages
S22-32
Language
English
Region
United States
NLM ID
101215604
PMCID
PMC4121056
Subset
IM
Grants
NHGRI NIH HHS · U54 HG004576 · United States
NHGRI NIH HHS · U54 HG004576-02 · United States
NHGRI NIH HHS · U54HG004576 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]