Abstract
Ensembl gene annotation provides a comprehensive catalog of transcripts aligned to the reference sequence. It relies on publicly available species-specific and orthologous transcripts plus their inferred protein sequence. The accuracy of gene models is improved by increasing the species-specific component that can be cost-effectively achieved using RNA-seq. Two zebrafish gene annotations are presented in Ensembl version 62 built on the Zv9 reference sequence. Firstly, RNA-seq data from five tissues and seven developmental stages were assembled into 25,748 gene models. A 3'-end capture and sequencing protocol was developed to predict the 3' ends of transcripts, and 46.1% of the original models were subsequently refined. Secondly, a standard Ensembl genebuild, incorporating carefully filtered elements from the RNA-seq-only build, followed by a merge with the manually curated VEGA database, produced a comprehensive annotation of 26,152 genes represented by 51,569 transcripts. The RNA-seq-only and the Ensembl/VEGA genebuilds contribute contrasting elements to the final genebuild. The RNA-seq genebuild was used to adjust intron/exon boundaries of orthologous defined models, confirm their expression, and improve 3' untranslated regions. Importantly, the inferred protein alignments within the Ensembl genebuild conferred proof of model contiguity for the RNA-seq models. The zebrafish gene annotation has been enhanced by the incorporation of RNA-seq data and the pipeline will be used for other organisms. Organisms with little species-specific cDNA data will generally benefit the most.
MeSH Terms
3' Untranslated Regions
Animals
Computational Biology/methods
DNA, Complementary
Databases, Nucleic Acid
Exons
Genomics/methods
Introns
Male
Models, Genetic
Molecular Sequence Annotation
RNA/chemistry,genetics
Transcription, Genetic
Zebrafish/genetics
Chemicals
3' Untranslated Regions
DNA, Complementary
RNA
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Collins John E
Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridgeshire, CB10 1SA, United Kingdom.
[email protected]
White Simon
Searle Stephen M J
Stemple Derek L
References (31)
31 references, click to expand
-
Sequencing of cDNA using anchored oligo dT primers.
Nucleic Acids Res. 1993 Aug 11;21(16):3915-6
PMID: 8367318
-
Conserved function of lincRNAs in vertebrate embryonic development despite rapid sequence evolution.
Cell. 2011 Dec 23;147(7):1537-50
PMID: 22196729
-
Comprehensive polyadenylation site maps in yeast and human reveal pervasive alternative polyadenylation.
Cell. 2010 Dec 10;143(6):1018-29
PMID: 21145465
-
Stem cell transcriptome profiling via massive-scale mRNA sequencing.
Nat Methods. 2008 Jul;5(7):613-9
PMID: 18516046
-
Ensembl 2011.
Nucleic Acids Res. 2011 Jan;39(Database issue):D800-6
PMID: 21045057
-
Mapping and quantifying mammalian transcriptomes by RNA-Seq.
Nat Methods. 2008 Jul;5(7):621-8
PMID: 18516045
-
The vertebrate genome annotation (Vega) database.
Nucleic Acids Res. 2008 Jan;36(Database issue):D753-60
PMID: 18003653
-
The Sequence Alignment/Map format and SAMtools.
Bioinformatics. 2009 Aug 15;25(16):2078-9
PMID: 19505943
-
Optimization of de novo transcriptome assembly from next-generation sequencing data.
Genome Res. 2010 Oct;20(10):1432-40
PMID: 20693479
-
De novo assembly and analysis of RNA-seq data.
Nat Methods. 2010 Nov;7(11):909-12
PMID: 20935650
-
Genome sequencing and analysis of the Tasmanian devil and its transmissible cancer.
Cell. 2012 Feb 17;148(4):780-91
PMID: 22341448
-
The completion of the Mammalian Gene Collection (MGC).
Genome Res. 2009 Dec;19(12):2324-33
PMID: 19767417
-
Fast and accurate short read alignment with Burrows-Wheeler transform.
Bioinformatics. 2009 Jul 15;25(14):1754-60
PMID: 19451168
-
The landscape of C. elegans 3'UTRs.
Science. 2010 Jul 23;329(5990):432-5
PMID: 20522740
-
A global view of gene activity and alternative splicing by deep sequencing of the human transcriptome.
Science. 2008 Aug 15;321(5891):956-60
PMID: 18599741
-
Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
Nat Biotechnol. 2010 May;28(5):511-5
PMID: 20436464
-
NCBI Reference Sequences: current status, policy and new initiatives.
Nucleic Acids Res. 2009 Jan;37(Database issue):D32-6
PMID: 18927115
-
Automated generation of heuristics for biological sequence comparison.
BMC Bioinformatics. 2005 Feb 15;6:31
PMID: 15713233
-
Systematic identification of long noncoding RNAs expressed during zebrafish embryogenesis.
Genome Res. 2012 Mar;22(3):577-91
PMID: 22110045
-
Dynamic repertoire of a eukaryotic transcriptome surveyed at single-nucleotide resolution.
Nature. 2008 Jun 26;453(7199):1239-43
PMID: 18488015
-
Alternative isoform regulation in human tissue transcriptomes.
Nature. 2008 Nov 27;456(7221):470-6
PMID: 18978772
-
The Ensembl automatic gene annotation system.
Genome Res. 2004 May;14(5):942-50
PMID: 15123590
-
Annotating genomes with massive-scale RNA sequencing.
Genome Biol. 2008;9(12):R175
PMID: 19087247
-
Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
Genome Res. 2008 May;18(5):821-9
PMID: 18349386
-
The transcriptional landscape of the yeast genome defined by RNA sequencing.
Science. 2008 Jun 6;320(5881):1344-9
PMID: 18451266
-
RNA-Seq: a revolutionary tool for transcriptomics.
Nat Rev Genet. 2009 Jan;10(1):57-63
PMID: 19015660
-
The Universal Protein Resource (UniProt) in 2010.
Nucleic Acids Res. 2010 Jan;38(Database issue):D142-8
PMID: 19843607
-
Noncanonical transcript forms in yeast and their regulation during environmental stress.
RNA. 2010 Jun;16(6):1256-67
PMID: 20421314
-
Ab initio reconstruction of cell type-specific transcriptomes in mouse reveals the conserved multi-exonic structure of lincRNAs.
Nat Biotechnol. 2010 May;28(5):503-10
PMID: 20436462
-
Accurate whole human genome sequencing using reversible terminator chemistry.
Nature. 2008 Nov 6;456(7218):53-9
PMID: 18987734
-
Ab initio construction of a eukaryotic transcriptome by massively parallel mRNA sequencing.
Proc Natl Acad Sci U S A. 2009 Mar 3;106(9):3264-9
PMID: 19208812