Home LiteratureArticle Details
PMID: 27669167 Published · ppublish English Journal Article

Modeling of RNA-seq fragment sequence bias reduces systematic errors in transcript abundance estimation.

Nature biotechnology ·Vol. 34 ·No. 12 ·2016-12-00 ·Pages 1287-1291

Love MI, Hogenesch JB, Irizarry RA

Abstract

We find that current computational methods for estimating transcript abundance from RNA-seq data can lead to hundreds of false-positive results. We show that these systematic errors stem largely from a failure to model fragment GC content bias. Sample-specific biases associated with fragment sequence features lead to misidentification of transcript isoforms. We introduce alpine, a method for estimating sample-specific bias-corrected transcript abundance. By incorporating fragment sequence features, alpine greatly increases the accuracy of transcript abundance estimates, enabling a fourfold reduction in the number of false positives for reported changes in expression compared with Cufflinks. Using simulated data, we also show that alpine retains the ability to discover true positives, similar to other approaches. The method is available as an R/Bioconductor package that includes data visualization tools useful for bias discovery.

MeSH Terms
Algorithms Artifacts Base Composition/genetics Computer Simulation Models, Genetic Models, Statistical RNA/genetics Reproducibility of Results Sensitivity and Specificity Sequence Analysis, RNA Software Transcription Factors/genetics
Chemicals
Transcription Factors RNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Love Michael I ORCID
Department of Biostatistics and Computational Biology, Dana-Farber Cancer Institute, Boston, Massachusetts, USA. | Department of Biostatistics, Harvard TH Chan School of Public Health, Boston, Massachusetts, USA.
Hogenesch John B ORCID
Department of Pharmacology, Institute for Translational Medicine and Therapeutics, University of Pennsylvania School of Medicine, Philadelphia, Pennsylvania, USA.
Irizarry Rafael A
Department of Biostatistics and Computational Biology, Dana-Farber Cancer Institute, Boston, Massachusetts, USA. | Department of Biostatistics, Harvard TH Chan School of Public Health, Boston, Massachusetts, USA.
References (34)
34 references, click to expand
  1. RNA-Seq gene expression estimation with read mapping uncertainty.
    Bioinformatics. 2010 Feb 15;26(4):493-500 PMID: 20022975
  2. Biases in Illumina transcriptome sequencing caused by random hexamer priming.
    Nucleic Acids Res. 2010 Jul;38(12):e131 PMID: 20395217
  3. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  4. Modeling non-uniformity in short-read rates in RNA-Seq data.
    Genome Biol. 2010;11(5):R50 PMID: 20459815
  5. Differential expression analysis for sequence count data.
    Genome Biol. 2010;11(10):R106 PMID: 20979621
  6. Analysis and design of RNA sequencing experiments for identifying isoform regulation.
    Nat Methods. 2010 Dec;7(12):1009-15 PMID: 21057496
  7. Analyzing and minimizing PCR amplification bias in Illumina sequencing libraries.
    Genome Biol. 2011;12(2):R18 PMID: 21338519
  8. Improving RNA-Seq expression estimates by correcting for fragment bias.
    Genome Biol. 2011;12(3):R22 PMID: 21410973
  9. Estimation of alternative splicing isoform frequencies from RNA-Seq data.
    Algorithms Mol Biol. 2011 Apr 19;6(1):9 PMID: 21504602
  10. Bias detection and correction in RNA-Sequencing data.
    BMC Bioinformatics. 2011 Jul 19;12:290 PMID: 21771300
  11. Sparse linear modeling of next-generation mRNA sequencing (RNA-Seq) data for isoform discovery and abundance estimation.
    Proc Natl Acad Sci U S A. 2011 Dec 13;108(50):19867-72 PMID: 22135461
  12. GC-content normalization for RNA-Seq data.
    BMC Bioinformatics. 2011 Dec 17;12:480 PMID: 22177264
  13. Removing technical variability in RNA-seq data using conditional quantile normalization.
    Biostatistics. 2012 Apr;13(2):204-16 PMID: 22285995
  14. Summarizing and correcting the GC content bias in high-throughput sequencing.
    Nucleic Acids Res. 2012 May;40(10):e72 PMID: 22323520
  15. Using probabilistic estimation of expression residuals (PEER) to obtain increased power and interpretability of gene expression analyses.
    Nat Protoc. 2012 Feb 16;7(3):500-7 PMID: 22343431
  16. Transcriptome assembly and isoform expression level estimation from biased RNA-Seq reads.
    Bioinformatics. 2012 Nov 15;28(22):2914-21 PMID: 23060617
  17. STAR: ultrafast universal RNA-seq aligner.
    Bioinformatics. 2013 Jan 1;29(1):15-21 PMID: 23104886
  18. TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions.
    Genome Biol. 2013 Apr 25;14(4):R36 PMID: 23618408
  19. Software for computing and annotating genomic ranges.
    PLoS Comput Biol. 2013;9(8):e1003118 PMID: 23950696
  20. Transcriptome and genome sequencing uncovers functional variation in humans.
    Nature. 2013 Sep 26;501(7468):506-11 PMID: 24037378
  21. Reproducibility of high-throughput mRNA and small RNA sequencing across laboratories.
    Nat Biotechnol. 2013 Nov;31(11):1015-22 PMID: 24037425
  22. Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms.
    Nat Biotechnol. 2014 May;32(5):462-4 PMID: 24752080
  23. IVT-seq reveals extreme bias in RNA sequencing.
    Genome Biol. 2014 Jun 30;15(6):R86 PMID: 24981968
  24. Multi-platform assessment of transcriptome profiling using RNA-seq in the ABRF next-generation sequencing study.
    Nat Biotechnol. 2014 Sep;32(9):915-925 PMID: 25150835
  25. Normalization of RNA-seq data using factor analysis of control genes or samples.
    Nat Biotechnol. 2014 Sep;32(9):896-902 PMID: 25150836
  26. Detecting and correcting systematic variation in large-scale RNA sequencing data.
    Nat Biotechnol. 2014 Sep;32(9):888-95 PMID: 25150837
  27. A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium.
    Nat Biotechnol. 2014 Sep;32(9):903-14 PMID: 25150838
  28. svaseq: removing batch effects and other unwanted noise from sequencing data.
    Nucleic Acids Res. 2014 Dec 1;42(21):null PMID: 25294822
  29. Quantitative visualization of alternative exon expression from RNA-seq data.
    Bioinformatics. 2015 Jul 15;31(14):2400-2 PMID: 25617416
  30. Orchestrating high-throughput genomic analysis with Bioconductor.
    Nat Methods. 2015 Feb;12(2):115-21 PMID: 25633503
  31. Polyester: simulating RNA-seq datasets with differential transcript expression.
    Bioinformatics. 2015 Sep 1;31(17):2778-84 PMID: 25926345
  32. Hidden genes in birds.
    Genome Biol. 2015 Aug 18;16:164 PMID: 26283656
  33. Benchmark analysis of algorithms for determining and quantifying full-length mRNA splice forms from RNA-seq data.
    Bioinformatics. 2015 Dec 15;31(24):3938-45 PMID: 26338770
  34. Near-optimal probabilistic RNA-seq quantification.
    Nat Biotechnol. 2016 May;34(5):525-7 PMID: 27043002
Article Info
Journal
Nature biotechnology
Abbr.
Nat Biotechnol
ISSN
1546-1696
Published
2016-12-00
Epub
2016-00-26
Pages
1287-1291
Language
English
Region
United States
NLM ID
9604648
PMCID
PMC5143225
Subset
IM
Grants
NIGMS NIH HHS · R01 GM083084 · United States
NHGRI NIH HHS · R01 HG005220 · United States
NINDS NIH HHS · R01 NS054794 · United States
NCI NIH HHS · T32 CA009337 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]