Home LiteratureArticle Details
PMID: 27107712 Published · epublish English Evaluation Study Journal Article

A benchmark for RNA-seq quantification pipelines.

Genome biology ·Vol. 17 ·2016-04-23 ·Pages 74

Teng M, Love MI, Davis CA, Djebali S, Dobin A, Graveley BR, Li S, Mason CE, Olson S, Pervouchine D, Sloan CA, Wei X, Zhan L, Irizarry RA

Abstract

Obtaining RNA-seq measurements involves a complex data analytical process with a large number of competing algorithms as options. There is much debate about which of these methods provides the best approach. Unfortunately, it is currently difficult to evaluate their performance due in part to a lack of sensitive assessment metrics. We present a series of statistical summaries and plots to evaluate the performance in terms of specificity and sensitivity, available as a R/Bioconductor package ( http://bioconductor.org/packages/rnaseqcomp ). Using two independent datasets, we assessed seven competing pipelines. Performance was generally poor, with two methods clearly underperforming and RSEM slightly outperforming the rest.

MeSH Terms
Algorithms Animals Humans Reference Values Sensitivity and Specificity Sequence Analysis, RNA/methods,standards
Authors & Affiliations
14 authors, click to expand affiliations / ORCID
Teng Mingxiang
Department of Biostatistics and Computational Biology, Dana-Farber Cancer Institute, 450 Brookline Avenue, Boston, MA, 02215, USA. | Department of Biostatistics, Harvard TH Chan School of Public Health, 677 Huntington Avenue, Boston, MA, 02115, USA. | School of Computer Science and Technology, Harbin Institute of Technology, Harbin, China.
Love Michael I
Department of Biostatistics and Computational Biology, Dana-Farber Cancer Institute, 450 Brookline Avenue, Boston, MA, 02215, USA. | Department of Biostatistics, Harvard TH Chan School of Public Health, 677 Huntington Avenue, Boston, MA, 02115, USA.
Davis Carrie A
Functional Genomics Group, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY, 11724, USA.
Djebali Sarah
Bioinformatics and Genomics Programme, Centre for Genomic Regulation (CRG) and UPF, Doctor Aiguader, 88, Barcelona, 08003, Spain.
Dobin Alexander
Functional Genomics Group, Cold Spring Harbor Laboratory, 1 Bungtown Road, Cold Spring Harbor, NY, 11724, USA.
Graveley Brenton R
Department of Genetics and Genome Sciences, Institute for System Genomics, UConn Health Center, Farmington, CT, 06030, USA.
Li Sheng
Department of Physiology and Biophysics, Weill Cornell Medical College, New York, New York, USA.
Mason Christopher E
Department of Physiology and Biophysics, Weill Cornell Medical College, New York, New York, USA.
Olson Sara
Department of Genetics and Genome Sciences, Institute for System Genomics, UConn Health Center, Farmington, CT, 06030, USA.
Pervouchine Dmitri
Bioinformatics and Genomics Programme, Centre for Genomic Regulation (CRG) and UPF, Doctor Aiguader, 88, Barcelona, 08003, Spain.
Sloan Cricket A
Department of Genetics, Stanford University, 300 Pasteur Drive, MC-5477, Stanford, CA, 94305, USA.
Wei Xintao
Department of Genetics and Genome Sciences, Institute for System Genomics, UConn Health Center, Farmington, CT, 06030, USA.
Zhan Lijun
Department of Genetics and Genome Sciences, Institute for System Genomics, UConn Health Center, Farmington, CT, 06030, USA.
Irizarry Rafael A
Department of Biostatistics and Computational Biology, Dana-Farber Cancer Institute, 450 Brookline Avenue, Boston, MA, 02215, USA. [email protected]. | Department of Biostatistics, Harvard TH Chan School of Public Health, 677 Huntington Avenue, Boston, MA, 02115, USA. [email protected].
References (36)
36 references, click to expand
  1. Exploration, normalization, and summaries of high density oligonucleotide array probe level data.
    Biostatistics. 2003 Apr;4(2):249-64 PMID: 12925520
  2. A benchmark for Affymetrix GeneChip expression measures.
    Bioinformatics. 2004 Feb 12;20(3):323-31 PMID: 14960458
  3. The ENCODE (ENCyclopedia Of DNA Elements) Project.
    Science. 2004 Oct 22;306(5696):636-40 PMID: 15499007
  4. Multiple-laboratory comparison of microarray platforms.
    Nat Methods. 2005 May;2(5):345-50 PMID: 15846361
  5. Comparison of Affymetrix GeneChip expression measures.
    Bioinformatics. 2006 Apr 1;22(7):789-94 PMID: 16410320
  6. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  7. Consolidated strategy for the analysis of microarray spike-in data.
    Nucleic Acids Res. 2008 Oct;36(17):e108 PMID: 18676452
  8. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  9. Transcriptome genetics using second generation sequencing in a Caucasian population.
    Nature. 2010 Apr 1;464(7289):773-7 PMID: 20220756
  10. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  11. Tackling the widespread and critical impact of batch effects in high-throughput data.
    Nat Rev Genet. 2010 Oct;11(10):733-9 PMID: 20838408
  12. Mapping and analysis of chromatin state dynamics in nine human cell types.
    Nature. 2011 May 5;473(7345):43-9 PMID: 21441907
  13. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome.
    BMC Bioinformatics. 2011 Aug 04;12:323 PMID: 21816040
  14. Synthetic spike-in standards for RNA-seq experiments.
    Genome Res. 2011 Sep;21(9):1543-51 PMID: 21816910
  15. The self-assessment trap: can we all be better than average?
    Mol Syst Biol. 2011 Oct 11;7:537 PMID: 21988833
  16. Fast gapped-read alignment with Bowtie 2.
    Nat Methods. 2012 Mar 04;9(4):357-9 PMID: 22388286
  17. GENCODE: the reference human genome annotation for The ENCODE Project.
    Genome Res. 2012 Sep;22(9):1760-74 PMID: 22955987
  18. Direct isolation and RNA-seq reveal environment-dependent properties of engrafted neural stem/progenitor cells.
    Nat Commun. 2012;3:1140 PMID: 23072808
  19. Revisiting global gene expression analysis.
    Cell. 2012 Oct 26;151(3):476-82 PMID: 23101621
  20. STAR: ultrafast universal RNA-seq aligner.
    Bioinformatics. 2013 Jan 1;29(1):15-21 PMID: 23104886
  21. Streaming fragment assignment for real-time analysis of sequencing experiments.
    Nat Methods. 2013 Jan;10(1):71-3 PMID: 23160280
  22. TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions.
    Genome Biol. 2013 Apr 25;14(4):R36 PMID: 23618408
  23. Human housekeeping genes, revisited.
    Trends Genet. 2013 Oct;29(10):569-74 PMID: 23810203
  24. Transcriptome and genome sequencing uncovers functional variation in humans.
    Nature. 2013 Sep 26;501(7468):506-11 PMID: 24037378
  25. From single-cell to cell-pool transcriptomes: stochasticity in gene expression and RNA splicing.
    Genome Res. 2014 Mar;24(3):496-510 PMID: 24299736
  26. voom: Precision weights unlock linear model analysis tools for RNA-seq read counts.
    Genome Biol. 2014 Feb 03;15(2):R29 PMID: 24485249
  27. Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms.
    Nat Biotechnol. 2014 May;32(5):462-4 PMID: 24752080
  28. Normalization of RNA-seq data using factor analysis of control genes or samples.
    Nat Biotechnol. 2014 Sep;32(9):896-902 PMID: 25150836
  29. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2.
    Genome Biol. 2014;15(12):550 PMID: 25516281
  30. limma powers differential expression analyses for RNA-sequencing and microarray studies.
    Nucleic Acids Res. 2015 Apr 20;43(7):e47 PMID: 25605792
  31. Polyester: simulating RNA-seq datasets with differential transcript expression.
    Bioinformatics. 2015 Sep 1;31(17):2778-84 PMID: 25926345
  32. Comparative assessment of methods for the computational inference of transcript isoform abundance from RNA-seq data.
    Genome Biol. 2015 Jul 23;16:150 PMID: 26201343
  33. Analyzing a portion of the ROC curve.
    Med Decis Making. 1989 Jul-Sep;9(3):190-5 PMID: 2668680
  34. Isoform prefiltering improves performance of count-based methods for analysis of differential transcript usage.
    Genome Biol. 2016 Jan 26;17:12 PMID: 26813113
  35. Near-optimal probabilistic RNA-seq quantification.
    Nat Biotechnol. 2016 May;34(5):525-7 PMID: 27043002
  36. Modeling of RNA-seq fragment sequence bias reduces systematic errors in transcript abundance estimation.
    Nat Biotechnol. 2016 Dec;34(12):1287-1291 PMID: 27669167
Article Info
Journal
Genome biology
Abbr.
Genome Biol
ISSN
1474-760X
Published
2016-04-23
Epub
2016-00-23
Pages
74
Language
English
Region
England
NLM ID
100960660
PMCID
PMC4842274
Subset
IM
Grants
NCI NIH HHS · T32 CA009337 · United States
NHGRI NIH HHS · U54 HG007005 · United States
Corrections
ErratumIn
ErratumIn
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]