Home LiteratureArticle Details
PMID: 23888185 Published · ppublish English Journal Article

Simultaneous isoform discovery and quantification from RNA-seq.

Statistics in biosciences ·Vol. 5 ·No. 1 ·2013-05-01 ·Pages 100-118

Hiller D, Wong WH

Abstract

RNA sequencing is a recent technology which has seen an explosion of methods addressing all levels of analysis, from read mapping to transcript assembly to differential expression modeling. In particular the discovery of isoforms at the transcript assembly stage is a complex problem and current approaches suffer from various limitations. For instance, many approaches use graphs to construct a minimal set of isoforms which covers the observed reads, then perform a separate algorithm to quantify the isoforms, which can result in a loss of power. Current methods also use ad-hoc solutions to deal with the vast number of possible isoforms which can be constructed from a given set of reads. Finally, while the need of taking into account features such as read pairing and sampling rate of reads has been acknowledged, most existing methods do not seamlessly integrate these features as part of the model. We present Montebello, an integrated statistical approach which performs simultaneous isoform discovery and quantification by using a Monte Carlo simulation to find the most likely isoform composition leading to a set of observed reads. We compare Montebello to Cufflinks, a popular isoform discovery approach, on a simulated data set and on 46.3 million brain reads from an Illumina tissue panel. On this data set Montebello appears to offer a modest improvement over Cufflinks when considering discovery and parsimony metrics. In addition Montebello mitigates specific difficulties inherent in the Cufflinks approach. Finally, Montebello can be fine-tuned depending on the type of solution desired.

Keywords
Algorithms Alternative Splicing Isoform Discovery Monte Carlo RNA-Seq
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Hiller David
Center for Epigenetics, Johns Hopkins School of Medicine, 855 N. Wolfe St., Rangos 570, Baltimore, MD 21205 [email protected].
Wong Wing Hung
References (35)
35 references, click to expand
  1. Differential expression in RNA-seq: a matter of depth.
    Genome Res. 2011 Dec;21(12):2213-23 PMID: 21903743
  2. Comparative analysis of RNA-Seq alignment algorithms and the RNA-Seq unified mapper (RUM).
    Bioinformatics. 2011 Sep 15;27(18):2518-28 PMID: 21775302
  3. Full-length transcriptome assembly from RNA-Seq data without a reference genome.
    Nat Biotechnol. 2011 May 15;29(7):644-52 PMID: 21572440
  4. Unproductive splicing of SR genes associated with highly conserved and ultraconserved DNA elements.
    Nature. 2007 Apr 19;446(7138):926-9 PMID: 17361132
  5. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome.
    BMC Bioinformatics. 2011 Aug 04;12:323 PMID: 21816040
  6. Using Poisson mixed-effects model to quantify transcript-level gene expression in RNA-Seq.
    Bioinformatics. 2012 Jan 1;28(1):63-8 PMID: 22072384
  7. Ab initio reconstruction of cell type-specific transcriptomes in mouse reveals the conserved multi-exonic structure of lincRNAs.
    Nat Biotechnol. 2010 May;28(5):503-10 PMID: 20436462
  8. Generating consensus sequences from partial order multiple sequence alignment graphs.
    Bioinformatics. 2003 May 22;19(8):999-1008 PMID: 12761063
  9. Improving RNA-Seq expression estimates by correcting for fragment bias.
    Genome Biol. 2011;12(3):R22 PMID: 21410973
  10. Next-generation transcriptome assembly.
    Nat Rev Genet. 2011 Sep 07;12(10):671-82 PMID: 21897427
  11. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
  12. Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing.
    Nat Genet. 2008 Dec;40(12):1413-5 PMID: 18978789
  13. IsoformEx: isoform level gene expression estimation using weighted non-negative least squares from mRNA-Seq data.
    BMC Bioinformatics. 2011 Jul 27;12:305 PMID: 21794104
  14. Splicing graphs and EST assembly problem.
    Bioinformatics. 2002;18 Suppl 1:S181-8 PMID: 12169546
  15. A powerful and flexible approach to the analysis of RNA sequence count data.
    Bioinformatics. 2011 Oct 1;27(19):2672-8 PMID: 21810900
  16. Identifiability of isoform deconvolution from junction arrays and RNA-Seq.
    Bioinformatics. 2009 Dec 1;25(23):3056-9 PMID: 19762346
  17. The UCSC Known Genes.
    Bioinformatics. 2006 May 1;22(9):1036-46 PMID: 16500937
  18. Statistical inferences for isoform expression in RNA-Seq.
    Bioinformatics. 2009 Apr 15;25(8):1026-32 PMID: 19244387
  19. Analysis and design of RNA sequencing experiments for identifying isoform regulation.
    Nat Methods. 2010 Dec;7(12):1009-15 PMID: 21057496
  20. Modeling non-uniformity in short-read rates in RNA-Seq data.
    Genome Biol. 2010;11(5):R50 PMID: 20459815
  21. Accurate quantification of transcriptome from RNA-Seq data by effective length normalization.
    Nucleic Acids Res. 2011 Jan;39(2):e9 PMID: 21059678
  22. MATS: a Bayesian framework for flexible detection of differential alternative splicing from RNA-Seq data.
    Nucleic Acids Res. 2012 Apr;40(8):e61 PMID: 22266656
  23. Detection of splice junctions from paired-end RNA-seq data by SpliceMap.
    Nucleic Acids Res. 2010 Aug;38(14):4570-8 PMID: 20371516
  24. Alternative isoform regulation in human tissue transcriptomes.
    Nature. 2008 Nov 27;456(7221):470-6 PMID: 18978772
  25. NSMAP: a method for spliced isoforms identification and quantification from RNA-Seq.
    BMC Bioinformatics. 2011 May 16;12:162 PMID: 21575225
  26. SPACE: an algorithm to predict and quantify alternatively spliced isoforms using microarrays.
    Genome Biol. 2008;9(2):R46 PMID: 18312629
  27. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  28. Gene structure-based splice variant deconvolution using a microarray platform.
    Bioinformatics. 2003;19 Suppl 1:i315-22 PMID: 12855476
  29. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  30. IsoLasso: a LASSO regression approach to RNA-Seq based transcriptome assembly.
    J Comput Biol. 2011 Nov;18(11):1693-707 PMID: 21951053
  31. Sparse linear modeling of next-generation mRNA sequencing (RNA-Seq) data for isoform discovery and abundance estimation.
    Proc Natl Acad Sci U S A. 2011 Dec 13;108(50):19867-72 PMID: 22135461
  32. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  33. baySeq: empirical Bayesian methods for identifying differential expression in sequence count data.
    BMC Bioinformatics. 2010 Aug 10;11:422 PMID: 20698981
  34. Statistical Modeling of RNA-Seq Data.
    Stat Sci. 2011 Feb;26(1): PMID: 24307754
  35. An expectation-maximization algorithm for probabilistic reconstructions of full-length isoforms from splice graphs.
    Nucleic Acids Res. 2006 Jun 06;34(10):3150-60 PMID: 16757580
Article Info
Journal
Statistics in biosciences
Abbr.
Stat Biosci
ISSN
1867-1764
Published
2013-05-01
Pages
100-118
Language
English
Region
United States
NLM ID
101498115
PMCID
PMC3718502
Grants
NHGRI NIH HHS · R01 HG004634 · United States
NHGRI NIH HHS · R01 HG005220 · United States
NHGRI NIH HHS · R01 HG005717 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]