Abstract
The utility of genome assemblies does not only rely on the quality of the assembled genome sequence, but also on the quality of the gene annotations. The Pacific Biosciences Iso-Seq technology is a powerful support for accurate eukaryotic gene model annotation as it allows for direct readout of full-length cDNA sequences without the need for noisy short read-based transcript assembly. We propose the implementation of the TeloPrime Full Length cDNA Amplification kit to the Pacific Biosciences Iso-Seq technology in order to enrich for genuine full-length transcripts in the cDNA libraries. We provide evidence that TeloPrime outperforms the commonly used SMARTer PCR cDNA Synthesis Kit in identifying transcription start and end sites in Arabidopsis thaliana. Furthermore, we show that TeloPrime-based Pacific Biosciences Iso-Seq can be successfully applied to the polyploid genome of bread wheat (Triticum aestivum) not only to efficiently annotate gene models, but also to identify novel transcription sites, gene homeologs, splicing isoforms and previously unidentified gene loci.
MeSH Terms
Arabidopsis/genetics
Computer Systems
DNA, Complementary/genetics
Data Curation
Databases, Genetic
Gene Library
Genome, Plant
RNA, Messenger/genetics,metabolism
Sequence Analysis, DNA/methods
Transcription Initiation Site
Triticum/genetics
Chemicals
DNA, Complementary
RNA, Messenger
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Cartolano Maria
Department of Plant Developmental Biology, Max Planck Institute for Plant Breeding Research, Cologne, Germany.
Huettel Bruno
Max Planck-Genome-centre Cologne, Max Planck Institute for Plant Breeding Research, Cologne, Germany.
Hartwig Benjamin
Department of Plant Developmental Biology, Max Planck Institute for Plant Breeding Research, Cologne, Germany.
Reinhardt Richard
Max Planck-Genome-centre Cologne, Max Planck Institute for Plant Breeding Research, Cologne, Germany.
Schneeberger Korbinian
Department of Plant Developmental Biology, Max Planck Institute for Plant Breeding Research, Cologne, Germany.
References (25)
25 references, click to expand
-
RNA-Seq: a revolutionary tool for transcriptomics.
Nat Rev Genet. 2009 Jan;10(1):57-63
PMID: 19015660
-
Assessment of transcript reconstruction methods for RNA-seq.
Nat Methods. 2013 Dec;10(12):1177-84
PMID: 24185837
-
Genomics as the key to unlocking the polyploid potential of wheat.
New Phytol. 2015 Dec;208(4):1008-22
PMID: 26108556
-
Amplification of cDNA ends based on template-switching effect and step-out PCR.
Nucleic Acids Res. 1999 Mar 15;27(6):1558-60
PMID: 10037822
-
No assembly required: Full-length MHC class I allele discovery by PacBio circular consensus sequencing.
Hum Immunol. 2015 Dec;76(12):891-6
PMID: 26028281
-
MIPS PlantsDB: a database framework for comparative plant genome research.
Nucleic Acids Res. 2013 Jan;41(Database issue):D1144-51
PMID: 23203886
-
Next-generation transcriptome assembly.
Nat Rev Genet. 2011 Sep 07;12(10):671-82
PMID: 21897427
-
The Arabidopsis Information Resource (TAIR): improved gene annotation and new tools.
Nucleic Acids Res. 2012 Jan;40(Database issue):D1202-10
PMID: 22140109
-
Analysis of the bread wheat genome using whole-genome shotgun sequencing.
Nature. 2012 Nov 29;491(7426):705-10
PMID: 23192148
-
Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
Nat Biotechnol. 2010 May;28(5):511-5
PMID: 20436464
-
Full-length transcriptome sequences and splice variants obtained by a combination of sequencing platforms applied to different root tissues of Salvia miltiorrhiza and tanshinone biosynthesis.
Plant J. 2015 Jun;82(6):951-61
PMID: 25912611
-
Long-read sequencing of chicken transcripts and identification of new transcript isoforms.
PLoS One. 2014 Apr 15;9(4):e94650
PMID: 24736250
-
Reconstruction of the synthetic W7984 x Opata M85 wheat reference population.
Genome. 2011 Nov;54(11):875-82
PMID: 21999208
-
GMAP: a genomic mapping and alignment program for mRNA and EST sequences.
Bioinformatics. 2005 May 1;21(9):1859-75
PMID: 15728110
-
A chromosome-based draft sequence of the hexaploid bread wheat (Triticum aestivum) genome.
Science. 2014 Jul 18;345(6194):1251788
PMID: 25035500
-
Widespread Polycistronic Transcripts in Fungi Revealed by Single-Molecule mRNA Sequencing.
PLoS One. 2015 Jul 15;10(7):e0132628
PMID: 26177194
-
Structural and functional partitioning of bread wheat chromosome 3B.
Science. 2014 Jul 18;345(6194):1249721
PMID: 25035497
-
PacBio sequencing of gene families - a case study with wheat gluten genes.
Gene. 2014 Jan 10;533(2):541-6
PMID: 24144842
-
A whole-genome shotgun approach for assembling and anchoring the hexaploid bread wheat genome.
Genome Biol. 2015 Jan 31;16:26
PMID: 25637298
-
Exploiting single-molecule transcript sequencing for eukaryotic gene prediction.
Genome Biol. 2015 Sep 02;16:184
PMID: 26328666
-
Normalizing cDNA libraries.
Curr Protoc Mol Biol. 2010 Apr;Chapter 5:Unit 5.12.1-27
PMID: 20373503
-
Analysis of the genome sequence of the flowering plant Arabidopsis thaliana.
Nature. 2000 Dec 14;408(6814):796-815
PMID: 11130711
-
PacBio Sequencing and Its Applications.
Genomics Proteomics Bioinformatics. 2015 Oct;13(5):278-89
PMID: 26542840
-
Improving the Annotation of Arabidopsis lyrata Using RNA-Seq Data.
PLoS One. 2015 Sep 18;10(9):e0137391
PMID: 26382944
-
Anchoring and ordering NGS contig assemblies by population sequencing (POPSEQ).
Plant J. 2013 Nov;76(4):718-27
PMID: 23998490