Home LiteratureArticle Details
PMID: 20022975 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

RNA-Seq gene expression estimation with read mapping uncertainty.

Bioinformatics (Oxford, England) ·Vol. 26 ·No. 4 ·2010-02-15 ·Pages 493-500

Li B, Ruotti V, Stewart RM, Thomson JA, Dewey CN

Abstract

RNA-Seq is a promising new technology for accurately measuring gene expression levels. Expression estimation with RNA-Seq requires the mapping of relatively short sequencing reads to a reference genome or transcript set. Because reads are generally shorter than transcripts from which they are derived, a single read may map to multiple genes and isoforms, complicating expression analyses. Previous computational methods either discard reads that map to multiple locations or allocate them to genes heuristically. We present a generative statistical model and associated inference methods that handle read mapping uncertainty in a principled manner. Through simulations parameterized by real RNA-Seq data, we show that our method is more accurate than previous methods. Our improved accuracy is the result of handling read mapping uncertainty with a statistical model and the estimation of gene expression levels as the sum of isoform expression levels. Unlike previous methods, our method is capable of modeling non-uniform read distributions. Simulations with our method indicate that a read length of 20-25 bases is optimal for gene-level expression estimation from mouse and maize RNA-Seq data when sequencing throughput is fixed.

MeSH Terms
Algorithms Animals Base Sequence Computational Biology/methods Databases, Genetic Gene Expression Gene Expression Profiling Genome Mice Sequence Analysis, RNA/methods Software Zea mays/genetics
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Li Bo
Department of Computer Sciences, University of Wisconsin, Madison, WI 53706, USA.
Ruotti Victor
Stewart Ron M
Thomson James A
Dewey Colin N
References (15)
15 references, click to expand
  1. Profiling the HeLa S3 transcriptome using randomly primed cDNA and massively parallel short-read sequencing.
    Biotechniques. 2008 Jul;45(1):81-94 PMID: 18611170
  2. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
  3. A strategy of DNA sequencing employing computer programs.
    Nucleic Acids Res. 1979 Jun 11;6(7):2601-10 PMID: 461197
  4. Highly integrated single-base resolution maps of the epigenome in Arabidopsis.
    Cell. 2008 May 2;133(3):523-36 PMID: 18423832
  5. Statistical modeling of sequencing errors in SAGE libraries.
    Bioinformatics. 2004 Aug 4;20 Suppl 1:i31-9 PMID: 15262778
  6. Stem cell transcriptome profiling via massive-scale mRNA sequencing.
    Nat Methods. 2008 Jul;5(7):613-9 PMID: 18516046
  7. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  8. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  9. Cross-hybridization modeling on Affymetrix exon arrays.
    Bioinformatics. 2008 Dec 15;24(24):2887-93 PMID: 18984598
  10. A rescue strategy for multimapping short sequence tags refines surveys of transcriptional activity by CAGE.
    Genomics. 2008 Mar;91(3):281-8 PMID: 18178374
  11. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  12. The UCSC Known Genes.
    Bioinformatics. 2006 May 1;22(9):1036-46 PMID: 16500937
  13. Statistical inferences for isoform expression in RNA-Seq.
    Bioinformatics. 2009 Apr 15;25(8):1026-32 PMID: 19244387
  14. The transcriptional landscape of the yeast genome defined by RNA sequencing.
    Science. 2008 Jun 6;320(5881):1344-9 PMID: 18451266
  15. RNA-Seq: a revolutionary tool for transcriptomics.
    Nat Rev Genet. 2009 Jan;10(1):57-63 PMID: 19015660
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2010-02-15
Epub
2009-00-18
Pages
493-500
Language
English
Region
England
NLM ID
9808944
PMCID
PMC2820677
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]