Home LiteratureArticle Details
PMID: 21623353 Published · ppublish English Journal Article Review

Computational methods for transcriptome annotation and quantification using RNA-seq.

Nature methods ·Vol. 8 ·No. 6 ·2011-06-00 ·Pages 469-77

Garber M, Grabherr MG, Guttman M, Trapnell C

Abstract

High-throughput RNA sequencing (RNA-seq) promises a comprehensive picture of the transcriptome, allowing for the complete annotation and quantification of all genes and their isoforms across samples. Realizing this promise requires increasingly complex computational methods. These computational challenges fall into three main categories: (i) read mapping, (ii) transcriptome reconstruction and (iii) expression quantification. Here we explain the major conceptual and practical challenges, and the general classes of solutions for each category. Finally, we highlight the interdependence between these categories and discuss the benefits for different biological applications.

MeSH Terms
Animals Computational Biology/methods Gene Expression Profiling/statistics & numerical data Genomics/statistics & numerical data High-Throughput Nucleotide Sequencing/statistics & numerical data Humans Sequence Alignment/statistics & numerical data Sequence Analysis, RNA/statistics & numerical data
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Garber Manuel
Broad Institute of Massachusetts Institute of Technology and Harvard, Cambridge, Massachusetts, USA. [email protected]
Grabherr Manfred G
Guttman Mitchell
Trapnell Cole
References (78)
78 references, click to expand
  1. RNA-seq analysis of 2 closely related leukemia clones that differ in their self-renewal capacity.
    Blood. 2011 Jan 13;117(2):e27-38 PMID: 20980679
  2. Linking promoters to functional transcripts in small samples with nanoCAGE and CAGEscan.
    Nat Methods. 2010 Jul;7(7):528-34 PMID: 20543846
  3. BFAST: an alignment tool for large scale genome resequencing.
    PLoS One. 2009 Nov 11;4(11):e7767 PMID: 19907642
  4. Integrative analysis of the melanoma transcriptome.
    Genome Res. 2010 Apr;20(4):413-27 PMID: 20179022
  5. Optimization of de novo transcriptome assembly from next-generation sequencing data.
    Genome Res. 2010 Oct;20(10):1432-40 PMID: 20693479
  6. Comprehensive comparative analysis of strand-specific RNA sequencing methods.
    Nat Methods. 2010 Sep;7(9):709-15 PMID: 20711195
  7. De novo assembly and analysis of RNA-seq data.
    Nat Methods. 2010 Nov;7(11):909-12 PMID: 20935650
  8. 1-Tuple DNA sequencing: computer analysis.
    J Biomol Struct Dyn. 1989 Aug;7(1):63-73 PMID: 2684223
  9. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  10. The landscape of C. elegans 3'UTRs.
    Science. 2010 Jul 23;329(5990):432-5 PMID: 20522740
  11. A global view of gene activity and alternative splicing by deep sequencing of the human transcriptome.
    Science. 2008 Aug 15;321(5891):956-60 PMID: 18599741
  12. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  13. Transcript length bias in RNA-seq data confounds systems biology.
    Biol Direct. 2009 Apr 16;4:14 PMID: 19371405
  14. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  15. Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.
    BMC Bioinformatics. 2010 Feb 18;11:94 PMID: 20167110
  16. Statistical inferences for isoform expression in RNA-Seq.
    Bioinformatics. 2009 Apr 15;25(8):1026-32 PMID: 19244387
  17. RNA-Seq gene expression estimation with read mapping uncertainty.
    Bioinformatics. 2010 Feb 15;26(4):493-500 PMID: 20022975
  18. Analysis and management of microarray gene expression data.
    Curr Protoc Mol Biol. 2007 Jan;Chapter 19:Unit 19.6 PMID: 18265395
  19. Effect of read-mapping biases on detecting allele-specific expression from RNA-sequencing data.
    Bioinformatics. 2009 Dec 15;25(24):3207-12 PMID: 19808877
  20. SOAP2: an improved ultrafast tool for short read alignment.
    Bioinformatics. 2009 Aug 1;25(15):1966-7 PMID: 19497933
  21. Using quality scores and longer reads improves accuracy of Solexa read mapping.
    BMC Bioinformatics. 2008 Feb 28;9:128 PMID: 18307793
  22. Highly integrated single-base resolution maps of the epigenome in Arabidopsis.
    Cell. 2008 May 2;133(3):523-36 PMID: 18423832
  23. Accurate quantification of transcriptome from RNA-Seq data by effective length normalization.
    Nucleic Acids Res. 2011 Jan;39(2):e9 PMID: 21059678
  24. SHRiMP: accurate mapping of short color-space reads.
    PLoS Comput Biol. 2009 May;5(5):e1000386 PMID: 19461883
  25. Formation, regulation and evolution of Caenorhabditis elegans 3'UTRs.
    Nature. 2011 Jan 6;469(7328):97-101 PMID: 21085120
  26. Revealing global regulatory features of mammalian alternative splicing using a quantitative microarray platform.
    Mol Cell. 2004 Dec 22;16(6):929-41 PMID: 15610736
  27. Current-generation high-throughput sequencing: deepening insights into mammalian transcriptomes.
    Genes Dev. 2009 Jun 15;23(12):1379-86 PMID: 19528315
  28. Detection of splice junctions from paired-end RNA-seq data by SpliceMap.
    Nucleic Acids Res. 2010 Aug;38(14):4570-8 PMID: 20371516
  29. Alternative isoform regulation in human tissue transcriptomes.
    Nature. 2008 Nov 27;456(7221):470-6 PMID: 18978772
  30. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies.
    Nucleic Acids Res. 2003 Oct 1;31(19):5654-66 PMID: 14500829
  31. Annotating genomes with massive-scale RNA sequencing.
    Genome Biol. 2008;9(12):R175 PMID: 19087247
  32. Quantitative monitoring of gene expression patterns with a complementary DNA microarray.
    Science. 1995 Oct 20;270(5235):467-70 PMID: 7569999
  33. Transcriptome sequencing to detect gene fusions in cancer.
    Nature. 2009 Mar 5;458(7234):97-101 PMID: 19136943
  34. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  35. De novo transcriptome assembly with ABySS.
    Bioinformatics. 2009 Nov 1;25(21):2872-7 PMID: 19528083
  36. Alternative expression analysis by RNA sequencing.
    Nat Methods. 2010 Oct;7(10):843-7 PMID: 20835245
  37. A scaling normalization method for differential expression analysis of RNA-seq data.
    Genome Biol. 2010;11(3):R25 PMID: 20196867
  38. Expression of 24,426 human alternative splicing events and predicted cis regulation in 48 tissues and cell lines.
    Nat Genet. 2008 Dec;40(12):1416-25 PMID: 18978788
  39. Stem cell transcriptome profiling via massive-scale mRNA sequencing.
    Nat Methods. 2008 Jul;5(7):613-9 PMID: 18516046
  40. Genome of the marsupial Monodelphis domestica reveals innovation in non-coding sequences.
    Nature. 2007 May 10;447(7141):167-77 PMID: 17495919
  41. SOAP: short oligonucleotide alignment program.
    Bioinformatics. 2008 Mar 1;24(5):713-4 PMID: 18227114
  42. SeqMap: mapping massive amount of oligonucleotides to the genome.
    Bioinformatics. 2008 Oct 15;24(20):2395-6 PMID: 18697769
  43. Moderated statistical tests for assessing differences in tag abundance.
    Bioinformatics. 2007 Nov 1;23(21):2881-7 PMID: 17881408
  44. Scaffolding a Caenorhabditis nematode genome with RNA-seq.
    Genome Res. 2010 Dec;20(12):1740-7 PMID: 20980554
  45. A practical false discovery rate approach to identifying patterns of differential expression in microarray data.
    Bioinformatics. 2005 Jun 1;21(11):2684-90 PMID: 15797908
  46. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  47. Differential expression analysis for sequence count data.
    Genome Biol. 2010;11(10):R106 PMID: 20979621
  48. The transcriptional landscape of the yeast genome defined by RNA sequencing.
    Science. 2008 Jun 6;320(5881):1344-9 PMID: 18451266
  49. RNA-Seq: a revolutionary tool for transcriptomics.
    Nat Rev Genet. 2009 Jan;10(1):57-63 PMID: 19015660
  50. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  51. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  52. Fast and SNP-tolerant detection of complex variants and splicing in short reads.
    Bioinformatics. 2010 Apr 1;26(7):873-81 PMID: 20147302
  53. Targeting a complex transcriptome: the construction of the mouse full-length cDNA encyclopedia.
    Genome Res. 2003 Jun;13(6B):1273-89 PMID: 12819125
  54. Sex-specific and lineage-specific alternative splicing in primates.
    Genome Res. 2010 Feb;20(2):180-9 PMID: 20009012
  55. Identification of human chromosome 22 transcribed sequences with ORF expressed sequence tags.
    Proc Natl Acad Sci U S A. 2000 Nov 7;97(23):12690-3 PMID: 11070084
  56. Ab initio reconstruction of cell type-specific transcriptomes in mouse reveals the conserved multi-exonic structure of lincRNAs.
    Nat Biotechnol. 2010 May;28(5):503-10 PMID: 20436462
  57. GMAP: a genomic mapping and alignment program for mRNA and EST sequences.
    Bioinformatics. 2005 May 1;21(9):1859-75 PMID: 15728110
  58. Ab initio construction of a eukaryotic transcriptome by massively parallel mRNA sequencing.
    Proc Natl Acad Sci U S A. 2009 Mar 3;106(9):3264-9 PMID: 19208812
  59. DEGseq: an R package for identifying differentially expressed genes from RNA-seq data.
    Bioinformatics. 2010 Jan 1;26(1):136-8 PMID: 19855105
  60. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  61. Mapping short DNA sequencing reads and calling variants using mapping quality scores.
    Genome Res. 2008 Nov;18(11):1851-8 PMID: 18714091
  62. BLAT--the BLAST-like alignment tool.
    Genome Res. 2002 Apr;12(4):656-64 PMID: 11932250
  63. Complementary DNA sequencing: expressed sequence tags and human genome project.
    Science. 1991 Jun 21;252(5013):1651-6 PMID: 2047873
  64. Computation for ChIP-seq and RNA-seq studies.
    Nat Methods. 2009 Nov;6(11 Suppl):S22-32 PMID: 19844228
  65. GASSST: global alignment short sequence search tool.
    Bioinformatics. 2010 Oct 15;26(20):2534-40 PMID: 20739310
  66. Cloud-scale RNA-sequencing differential expression analysis with Myrna.
    Genome Biol. 2010;11(8):R83 PMID: 20701754
  67. An encyclopedia of mouse genes.
    Nat Genet. 1999 Feb;21(2):191-4 PMID: 9988271
  68. Large-scale transcriptional activity in chromosomes 21 and 22.
    Science. 2002 May 3;296(5569):916-9 PMID: 11988577
  69. Significance analysis of microarrays applied to the ionizing radiation response.
    Proc Natl Acad Sci U S A. 2001 Apr 24;98(9):5116-21 PMID: 11309499
  70. MapSplice: accurate mapping of RNA-seq reads for splice junction discovery.
    Nucleic Acids Res. 2010 Oct;38(18):e178 PMID: 20802226
  71. Using the Velvet de novo assembler for short-read sequencing technologies.
    Curr Protoc Bioinformatics. 2010 Sep;Chapter 11:Unit 11.5 PMID: 20836074
  72. Optimal spliced alignments of short sequence reads.
    Bioinformatics. 2008 Aug 15;24(16):i174-80 PMID: 18689821
  73. Molecular classification of cancer: class discovery and class prediction by gene expression monitoring.
    Science. 1999 Oct 15;286(5439):531-7 PMID: 10521349
  74. Isoform abundance inference provides a more accurate estimation of gene expression levels in RNA-seq.
    J Bioinform Comput Biol. 2010 Dec;8 Suppl 1:177-92 PMID: 21155027
  75. Chromatin signature reveals over a thousand highly conserved large non-coding RNAs in mammals.
    Nature. 2009 Mar 12;458(7235):223-7 PMID: 19182780
  76. RNA-MATE: a recursive mapping strategy for high-throughput RNA-sequencing data.
    Bioinformatics. 2009 Oct 1;25(19):2615-6 PMID: 19648138
  77. Stampy: a statistical algorithm for sensitive and fast mapping of Illumina sequence reads.
    Genome Res. 2011 Jun;21(6):936-9 PMID: 20980556
  78. Next is now: new technologies for sequencing of genomes, transcriptomes, and beyond.
    Curr Opin Plant Biol. 2009 Apr;12(2):107-18 PMID: 19157957
Article Info
Journal
Nature methods
Abbr.
Nat Methods
ISSN
1548-7105
Published
2011-06-00
Epub
2011-00-27
Pages
469-77
Language
English
Region
United States
NLM ID
101215604
Subset
IM
Corrections
ErratumIn
-
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]