Home LiteratureArticle Details
PMID: 21176179 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Review

From RNA-seq reads to differential expression results.

Genome biology ·Vol. 11 ·No. 12 ·2010-00-00 ·Pages 220

Oshlack A, Robinson MD, Young MD

Abstract

Many methods and tools are available for preprocessing high-throughput RNA sequencing data and detecting differential expression.

MeSH Terms
Chromosome Mapping Computational Biology Gene Expression Gene Expression Profiling High-Throughput Nucleotide Sequencing/methods Humans Microarray Analysis/methods RNA/genetics Sequence Alignment Sequence Analysis, RNA/methods Software
Chemicals
RNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Oshlack Alicia
Bioinformatics Division, Walter and Eliza Hall Institute, 1G Royal Parade, Parkville 3052, Australia. [email protected]
Robinson Mark D
Young Matthew D
References (88)
88 references, click to expand
  1. mrsFAST: a cache-oblivious algorithm for short-read mapping.
    Nat Methods. 2010 Aug;7(8):576-7 PMID: 20676076
  2. Global and unbiased detection of splice junctions from RNA-seq data.
    Genome Biol. 2010;11(3):R34 PMID: 20236510
  3. BFAST: an alignment tool for large scale genome resequencing.
    PLoS One. 2009 Nov 11;4(11):e7767 PMID: 19907642
  4. Computational solutions to large-scale data management and analysis.
    Nat Rev Genet. 2010 Sep;11(9):647-57 PMID: 20717155
  5. MapSplice: accurate mapping of RNA-seq reads for splice junction discovery.
    Nucleic Acids Res. 2010 Oct;38(18):e178 PMID: 20802226
  6. Next-generation genomics: an integrative approach.
    Nat Rev Genet. 2010 Jul;11(7):476-86 PMID: 20531367
  7. Understanding mechanisms underlying human gene expression variation with RNA sequencing.
    Nature. 2010 Apr 1;464(7289):768-72 PMID: 20220758
  8. A two-parameter generalized Poisson model to improve the analysis of RNA-seq data.
    Nucleic Acids Res. 2010 Sep;38(17):e170 PMID: 20671027
  9. Conserved developmental transcriptomes in evolutionarily divergent species.
    Genome Biol. 2010;11(3):R35 PMID: 20236529
  10. De novo assembly and analysis of RNA-seq data.
    Nat Methods. 2010 Nov;7(11):909-12 PMID: 20935650
  11. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  12. Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing.
    Nat Genet. 2008 Dec;40(12):1413-5 PMID: 18978789
  13. ChIP-Seq of transcription factors predicts absolute and differential gene expression in embryonic stem cells.
    Proc Natl Acad Sci U S A. 2009 Dec 22;106(51):21521-6 PMID: 19995984
  14. Large-scale detection and analysis of RNA editing in grape mtDNA by RNA deep-sequencing.
    Nucleic Acids Res. 2010 Aug;38(14):4755-67 PMID: 20385587
  15. RazerS--fast read mapping with sensitivity control.
    Genome Res. 2009 Sep;19(9):1646-54 PMID: 19592482
  16. A global view of gene activity and alternative splicing by deep sequencing of the human transcriptome.
    Science. 2008 Aug 15;321(5891):956-60 PMID: 18599741
  17. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  18. Transcript length bias in RNA-seq data confounds systems biology.
    Biol Direct. 2009 Apr 16;4:14 PMID: 19371405
  19. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  20. Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.
    BMC Bioinformatics. 2010 Feb 18;11:94 PMID: 20167110
  21. Statistical inferences for isoform expression in RNA-Seq.
    Bioinformatics. 2009 Apr 15;25(8):1026-32 PMID: 19244387
  22. RNA-Seq gene expression estimation with read mapping uncertainty.
    Bioinformatics. 2010 Feb 15;26(4):493-500 PMID: 20022975
  23. Statistical design and analysis of RNA sequencing data.
    Genetics. 2010 Jun;185(2):405-16 PMID: 20439781
  24. Small-sample estimation of negative binomial dispersion, with applications to SAGE data.
    Biostatistics. 2008 Apr;9(2):321-32 PMID: 17728317
  25. A large genome center's improvements to the Illumina sequencing system.
    Nat Methods. 2008 Dec;5(12):1005-10 PMID: 19034268
  26. Methods for analyzing deep sequencing expression data: constructing the human and mouse promoterome with deepCAGE data.
    Genome Biol. 2009;10(7):R79 PMID: 19624849
  27. A survey of sequence alignment algorithms for next-generation sequencing.
    Brief Bioinform. 2010 Sep;11(5):473-83 PMID: 20460430
  28. Comparison of sequencing-based methods to profile DNA methylation and identification of monoallelic epigenetic modifications.
    Nat Biotechnol. 2010 Oct;28(10):1097-105 PMID: 20852635
  29. SOAP2: an improved ultrafast tool for short read alignment.
    Bioinformatics. 2009 Aug 1;25(15):1966-7 PMID: 19497933
  30. Computational analysis of whole-genome differential allelic expression data in human.
    PLoS Comput Biol. 2010 Jul 08;6(7):e1000849 PMID: 20628616
  31. CloudBurst: highly sensitive read mapping with MapReduce.
    Bioinformatics. 2009 Jun 1;25(11):1363-9 PMID: 19357099
  32. Solving the riddle of the bright mismatches: labeling and effective binding in oligonucleotide arrays.
    Phys Rev E Stat Nonlin Soft Matter Phys. 2003 Jul;68(1 Pt 1):011906 PMID: 12935175
  33. A comparison of massively parallel nucleotide sequencing with oligonucleotide microarrays for global transcription profiling.
    BMC Genomics. 2010 May 05;11:282 PMID: 20444259
  34. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  35. Modeling non-uniformity in short-read rates in RNA-Seq data.
    Genome Biol. 2010;11(5):R50 PMID: 20459815
  36. Gene ontology analysis for RNA-seq: accounting for selection bias.
    Genome Biol. 2010;11(2):R14 PMID: 20132535
  37. Transcriptome genetics using second generation sequencing in a Caucasian population.
    Nature. 2010 Apr 1;464(7289):773-7 PMID: 20220756
  38. Next-generation DNA sequencing.
    Nat Biotechnol. 2008 Oct;26(10):1135-45 PMID: 18846087
  39. High-resolution analysis of parent-of-origin allelic expression in the mouse brain.
    Science. 2010 Aug 6;329(5992):643-8 PMID: 20616232
  40. DAVID: Database for Annotation, Visualization, and Integrated Discovery.
    Genome Biol. 2003;4(5):P3 PMID: 12734009
  41. SHRiMP: accurate mapping of short color-space reads.
    PLoS Comput Biol. 2009 May;5(5):e1000386 PMID: 19461883
  42. Aberrant luminal progenitors as the candidate target population for basal tumor development in BRCA1 mutation carriers.
    Nat Med. 2009 Aug;15(8):907-13 PMID: 19648928
  43. mRNA-seq with agnostic splice site discovery for nervous system transcriptomics tested in chronic pain.
    Genome Res. 2010 Jun;20(6):847-60 PMID: 20452967
  44. Detection of splice junctions from paired-end RNA-seq data by SpliceMap.
    Nucleic Acids Res. 2010 Aug;38(14):4570-8 PMID: 20371516
  45. Molecular mechanisms of ethanol-induced pathogenesis revealed by RNA-sequencing.
    PLoS Pathog. 2010 Apr 01;6(4):e1000834 PMID: 20368969
  46. Annotating genomes with massive-scale RNA sequencing.
    Genome Biol. 2008;9(12):R175 PMID: 19087247
  47. Transcriptome sequencing to detect gene fusions in cancer.
    Nature. 2009 Mar 5;458(7234):97-101 PMID: 19136943
  48. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  49. Alternative expression analysis by RNA sequencing.
    Nat Methods. 2010 Oct;7(10):843-7 PMID: 20835245
  50. A scaling normalization method for differential expression analysis of RNA-seq data.
    Genome Biol. 2010;11(3):R25 PMID: 20196867
  51. Stem cell transcriptome profiling via massive-scale mRNA sequencing.
    Nat Methods. 2008 Jul;5(7):613-9 PMID: 18516046
  52. Moderated statistical tests for assessing differences in tag abundance.
    Bioinformatics. 2007 Nov 1;23(21):2881-7 PMID: 17881408
  53. SOAP: short oligonucleotide alignment program.
    Bioinformatics. 2008 Mar 1;24(5):713-4 PMID: 18227114
  54. Human DNA methylomes at base resolution show widespread epigenomic differences.
    Nature. 2009 Nov 19;462(7271):315-22 PMID: 19829295
  55. The GNUMAP algorithm: unbiased probabilistic mapping of oligonucleotides from next-generation sequencing.
    Bioinformatics. 2010 Jan 1;26(1):38-45 PMID: 19861355
  56. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  57. Differential expression analysis for sequence count data.
    Genome Biol. 2010;11(10):R106 PMID: 20979621
  58. Deep sequencing-based expression analysis shows major advances in robustness, resolution and inter-lab portability over five microarray platforms.
    Nucleic Acids Res. 2008 Dec;36(21):e141 PMID: 18927111
  59. RNA-Seq: a revolutionary tool for transcriptomics.
    Nat Rev Genet. 2009 Jan;10(1):57-63 PMID: 19015660
  60. Sense from sequence reads: methods for alignment and assembly.
    Nat Methods. 2009 Nov;6(11 Suppl):S6-S12 PMID: 19844229
  61. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  62. Limitations and possibilities of small RNA digital gene expression profiling.
    Nat Methods. 2009 Jul;6(7):474-6 PMID: 19564845
  63. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  64. Fast and SNP-tolerant detection of complex variants and splicing in short reads.
    Bioinformatics. 2010 Apr 1;26(7):873-81 PMID: 20147302
  65. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles.
    Proc Natl Acad Sci U S A. 2005 Oct 25;102(43):15545-50 PMID: 16199517
  66. Estimating accuracy of RNA-Seq and microarrays with proteomics.
    BMC Genomics. 2009 Apr 16;10:161 PMID: 19371429
  67. A mutation accumulation assay reveals a broad capacity for rapid evolution of gene expression.
    Nature. 2005 Nov 10;438(7065):220-3 PMID: 16281035
  68. Biases in Illumina transcriptome sequencing caused by random hexamer priming.
    Nucleic Acids Res. 2010 Jul;38(12):e131 PMID: 20395217
  69. Using the miraEST assembler for reliable and automated mRNA transcript assembly and SNP detection in sequenced ESTs.
    Genome Res. 2004 Jun;14(6):1147-59 PMID: 15140833
  70. DEGseq: an R package for identifying differentially expressed genes from RNA-seq data.
    Bioinformatics. 2010 Jan 1;26(1):136-8 PMID: 19855105
  71. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  72. Mapping short DNA sequencing reads and calling variants using mapping quality scores.
    Genome Res. 2008 Nov;18(11):1851-8 PMID: 18714091
  73. Close association of RNA polymerase II and many transcription factors with Pol III genes.
    Proc Natl Acad Sci U S A. 2010 Feb 23;107(8):3639-44 PMID: 20139302
  74. A statistical method for the detection of alternative splicing using RNA-seq.
    PLoS One. 2010 Jan 08;5(1):e8529 PMID: 20072613
  75. baySeq: empirical Bayesian methods for identifying differential expression in sequence count data.
    BMC Bioinformatics. 2010 Aug 10;11:422 PMID: 20698981
  76. Computation for ChIP-seq and RNA-seq studies.
    Nat Methods. 2009 Nov;6(11 Suppl):S22-32 PMID: 19844228
  77. Cloud-scale RNA-sequencing differential expression analysis with Myrna.
    Genome Biol. 2010;11(8):R83 PMID: 20701754
  78. Optimal spliced alignments of short sequence reads.
    Bioinformatics. 2008 Aug 15;24(16):i174-80 PMID: 18689821
  79. RNA-MATE: a recursive mapping strategy for high-throughput RNA-sequencing data.
    Bioinformatics. 2009 Oct 1;25(19):2615-6 PMID: 19648138
  80. KEGG: kyoto encyclopedia of genes and genomes.
    Nucleic Acids Res. 2000 Jan 1;28(1):27-30 PMID: 10592173
  81. Stochastic models inspired by hybridization theory for short oligonucleotide arrays.
    J Comput Biol. 2005 Jul-Aug;12(6):882-93 PMID: 16108723
  82. Transcriptome-wide identification of novel imprinted genes in neonatal mouse brain.
    PLoS One. 2008;3(12):e3839 PMID: 19052635
  83. Determination of tag density required for digital transcriptome analysis: application to an androgen-sensitive prostate cancer model.
    Proc Natl Acad Sci U S A. 2008 Dec 23;105(51):20179-84 PMID: 19088194
  84. The Tasmanian devil transcriptome reveals Schwann cell origins of a clonally transmissible cancer.
    Science. 2010 Jan 1;327(5961):84-7 PMID: 20044575
  85. ABySS: a parallel assembler for short read sequence data.
    Genome Res. 2009 Jun;19(6):1117-23 PMID: 19251739
  86. Integrative analysis of the melanoma transcriptome.
    Genome Res. 2010 Apr;20(4):413-27 PMID: 20179022
  87. PerM: efficient mapping of short sequencing reads with periodic full sensitive spaced seeds.
    Bioinformatics. 2009 Oct 1;25(19):2514-21 PMID: 19675096
  88. mRNA-Seq whole-transcriptome analysis of a single cell.
    Nat Methods. 2009 May;6(5):377-82 PMID: 19349980
Article Info
Journal
Genome biology
Abbr.
Genome Biol
ISSN
1474-760X
Published
2010-00-00
Epub
2010-00-22
Pages
220
Language
English
Region
England
NLM ID
100960660
PMCID
PMC3046478
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]