Abstract
Generation of large mate-pair libraries is necessary for de novo genome assembly but the procedure is complex and time-consuming. Furthermore, in some complex genomes, it is hard to increase the N50 length even with large mate-pair libraries, which leads to low transcript coverage. Thus, it is necessary to develop other simple scaffolding approaches, to at least solve the elongation of transcribed fragments. We describe L_RNA_scaffolder, a novel genome scaffolding method that uses long transcriptome reads to order, orient and combine genomic fragments into larger sequences. To demonstrate the accuracy of the method, the zebrafish genome was scaffolded. With expanded human transcriptome data, the N50 of human genome was doubled and L_RNA_scaffolder out-performed most scaffolding results by existing scaffolders which employ mate-pair libraries. In these two examples, the transcript coverage was almost complete, especially for long transcripts. We applied L_RNA_scaffolder to the highly polymorphic pearl oyster draft genome and the gene model length significantly increased. The simplicity and high-throughput of RNA-seq data makes this approach suitable for genome scaffolding. L_RNA_scaffolder is available at http://www.fishbrowser.org/software/L_RNA_scaffolder.
MeSH Terms
Animals
Genome, Human
Genomics/methods
Humans
Pinctada/genetics
RNA/genetics
Sequence Alignment
Sequence Analysis, DNA/methods
Software
Transcriptome
Zebrafish/genetics
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Xue Wei
The Centre for Applied Aquatic Genomics, Chinese Academy of Fishery Sciences, Beijing 100141, China.
[email protected].
Li Jiong-Tang
Zhu Ya-Ping
Hou Guang-Yuan
Kong Xiang-Fei
Kuang You-Yi
Sun Xiao-Wen
References (30)
30 references, click to expand
-
Draft genome of the pearl oyster Pinctada fucata: a platform for understanding bivalve biology.
DNA Res. 2012 Apr;19(2):117-30
PMID: 22315334
-
SolexaQA: At-a-glance quality assessment of Illumina second-generation sequencing data.
BMC Bioinformatics. 2010 Sep 27;11:485
PMID: 20875133
-
Updating benchtop sequencing performance comparison.
Nat Biotechnol. 2013 Apr;31(4):294-6
PMID: 23563421
-
Generation of long insert pairs using a Cre-LoxP Inverse PCR approach.
PLoS One. 2012;7(1):e29437
PMID: 22253722
-
SOPRA: Scaffolding algorithm for paired reads via statistical optimization.
BMC Bioinformatics. 2010 Jun 24;11:345
PMID: 20576136
-
Ab initio gene finding in Drosophila genomic DNA.
Genome Res. 2000 Apr;10(4):516-22
PMID: 10779491
-
Assembly of the working draft of the human genome with GigAssembler.
Genome Res. 2001 Sep;11(9):1541-8
PMID: 11544197
-
Genomewide characterization of non-polyadenylated RNAs.
Genome Biol. 2011;12(2):R16
PMID: 21324177
-
De novo assembly of human genomes with massively parallel short read sequencing.
Genome Res. 2010 Feb;20(2):265-72
PMID: 20019144
-
Ensembl 2011.
Nucleic Acids Res. 2011 Jan;39(Database issue):D800-6
PMID: 21045057
-
Assessing the gene space in draft genomes.
Nucleic Acids Res. 2009 Jan;37(1):289-97
PMID: 19042974
-
BLAT--the BLAST-like alignment tool.
Genome Res. 2002 Apr;12(4):656-64
PMID: 11932250
-
Opera: reconstructing optimal genomic scaffolds with high-throughput paired-end sequences.
J Comput Biol. 2011 Nov;18(11):1681-91
PMID: 21929371
-
The zebrafish reference genome sequence and its relationship to the human genome.
Nature. 2013 Apr 25;496(7446):498-503
PMID: 23594743
-
The complexity of the mammalian transcriptome.
J Physiol. 2006 Sep 1;575(Pt 2):321-32
PMID: 16857706
-
Improving PacBio long read accuracy by short read alignment.
PLoS One. 2012;7(10):e46679
PMID: 23056399
-
Paired-end sequencing of long-range DNA fragments for de novo assembly of large, complex Mammalian genomes by direct intra-molecule ligation.
PLoS One. 2012;7(9):e46211
PMID: 23029438
-
Scaffolding pre-assembled contigs using SSPACE.
Bioinformatics. 2011 Feb 15;27(4):578-9
PMID: 21149342
-
SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler.
Gigascience. 2012 Dec 27;1(1):18
PMID: 23587118
-
The oyster genome reveals stress adaptation and complexity of shell formation.
Nature. 2012 Oct 4;490(7418):49-54
PMID: 22992520
-
Full-length transcriptome assembly from RNA-Seq data without a reference genome.
Nat Biotechnol. 2011 May 15;29(7):644-52
PMID: 21572440
-
Fast scaffolding with small independent mixed integer programs.
Bioinformatics. 2011 Dec 1;27(23):3259-65
PMID: 21998153
-
Transcriptome sequencing to detect gene fusions in cancer.
Nature. 2009 Mar 5;458(7234):97-101
PMID: 19136943
-
A neoplastic gene fusion mimics trans-splicing of RNAs in normal human cells.
Science. 2008 Sep 5;321(5894):1357-61
PMID: 18772439
-
The reality of pervasive transcription.
PLoS Biol. 2011 Jul;9(7):e1000625; discussion e1001102
PMID: 21765801
-
GAGE: A critical evaluation of genome assemblies and assembly algorithms.
Genome Res. 2012 Mar;22(3):557-67
PMID: 22147368
-
Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data.
Nat Methods. 2013 Jun;10(6):563-9
PMID: 23644548
-
The UCSC Genome Browser database: extensions and updates 2011.
Nucleic Acids Res. 2012 Jan;40(Database issue):D918-23
PMID: 22086951
-
Paired-end sequencing of Fosmid libraries by Illumina.
Genome Res. 2012 Nov;22(11):2241-9
PMID: 22800726
-
ABySS: a parallel assembler for short read sequence data.
Genome Res. 2009 Jun;19(6):1117-23
PMID: 19251739