Home LiteratureArticle Details
PMID: 26244089 Published · epublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads.

GigaScience ·Vol. 4 ·2015-00-00 ·Pages 35

Warren RL, Yang C, Vandervalk BP, Behsaz B, Lagman A, Jones SJ, Birol I

Abstract

Owing to the complexity of the assembly problem, we do not yet have complete genome sequences. The difficulty in assembling reads into finished genomes is exacerbated by sequence repeats and the inability of short reads to capture sufficient genomic information to resolve those problematic regions. In this regard, established and emerging long read technologies show great promise, but their current associated higher error rates typically require computational base correction and/or additional bioinformatics pre-processing before they can be of value. We present LINKS, the Long Interval Nucleotide K-mer Scaffolder algorithm, a method that makes use of the sequence properties of nanopore sequence data and other error-containing sequence data, to scaffold high-quality genome assemblies, without the need for read alignment or base correction. Here, we show how the contiguity of an ABySS Escherichia coli K-12 genome assembly can be increased greater than five-fold by the use of beta-released Oxford Nanopore Technologies Ltd. long reads and how LINKS leverages long-range information in Saccharomyces cerevisiae W303 nanopore reads to yield assemblies whose resulting contiguity and correctness are on par with or better than that of competing applications. We also present the re-scaffolding of the colossal white spruce (Picea glauca) draft assembly (PG29, 20 Gbp) and demonstrate how LINKS scales to larger genomes. This study highlights the present utility of nanopore reads for genome scaffolding in spite of their current limitations, which are expected to diminish as the nanopore sequencing technology advances. We expect LINKS to have broad utility in harnessing the potential of long reads in connecting high-quality sequences of small and large genome assembly drafts.

Keywords
Genome assembly LINKS Nanopore sequencing Next-generation sequencing Scaffolding
MeSH Terms
Genome Sequence Alignment
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Warren René L
BC Cancer Agency, Michael Smith Genome Sciences Centre, Vancouver, British Columbia V5Z 4S6 Canada.
Yang Chen
BC Cancer Agency, Michael Smith Genome Sciences Centre, Vancouver, British Columbia V5Z 4S6 Canada.
Vandervalk Benjamin P
BC Cancer Agency, Michael Smith Genome Sciences Centre, Vancouver, British Columbia V5Z 4S6 Canada.
Behsaz Bahar
BC Cancer Agency, Michael Smith Genome Sciences Centre, Vancouver, British Columbia V5Z 4S6 Canada.
Lagman Albert
BC Cancer Agency, Michael Smith Genome Sciences Centre, Vancouver, British Columbia V5Z 4S6 Canada.
Jones Steven J M
BC Cancer Agency, Michael Smith Genome Sciences Centre, Vancouver, British Columbia V5Z 4S6 Canada.
Birol Inanç
BC Cancer Agency, Michael Smith Genome Sciences Centre, Vancouver, British Columbia V5Z 4S6 Canada.
References (25)
25 references, click to expand
  1. Assembling millions of short DNA sequences using SSAKE.
    Bioinformatics. 2007 Feb 15;23(4):500-1 PMID: 17158514
  2. Sealer: a scalable gap-closing application for finishing draft genomes.
    BMC Bioinformatics. 2015 Jul 25;16:230 PMID: 26209068
  3. Improved data analysis for the MinION nanopore sequencer.
    Nat Methods. 2015 Apr;12(4):351-6 PMID: 25686389
  4. Adaptive seeds tame genomic sequence comparison.
    Genome Res. 2011 Mar;21(3):487-93 PMID: 21209072
  5. Oxford Nanopore sequencing, hybrid error correction, and de novo assembly of a eukaryotic genome.
    Genome Res. 2015 Nov;25(11):1750-6 PMID: 26447147
  6. One chromosome, one contig: complete microbial genomes from long-read sequencing and assembly.
    Curr Opin Microbiol. 2015 Feb;23:110-20 PMID: 25461581
  7. Assembling the 20 Gb white spruce (Picea glauca) genome from whole-genome shotgun sequencing data.
    Bioinformatics. 2013 Jun 15;29(12):1492-7 PMID: 23698863
  8. Origins of the E. coli strain causing an outbreak of hemolytic-uremic syndrome in Germany.
    N Engl J Med. 2011 Aug 25;365(8):709-17 PMID: 21793740
  9. SSPACE-LongRead: scaffolding bacterial draft genomes using long read sequence information.
    BMC Bioinformatics. 2014 Jun 20;15:211 PMID: 24950923
  10. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data.
    Nat Methods. 2013 Jun;10(6):563-9 PMID: 23644548
  11. A whole-genome assembly of Drosophila.
    Science. 2000 Mar 24;287(5461):2196-204 PMID: 10731133
  12. A reference bacterial genome dataset generated on the MinION™ portable single-molecule nanopore sequencer.
    Gigascience. 2014 Oct 20;3:22 PMID: 25386338
  13. Assembling large genomes with single-molecule sequencing and locality-sensitive hashing.
    Nat Biotechnol. 2015 Jun;33(6):623-30 PMID: 26006009
  14. Continuous base identification for single-molecule nanopore DNA sequencing.
    Nat Nanotechnol. 2009 Apr;4(4):265-70 PMID: 19350039
  15. Improved white spruce (Picea glauca) genome assemblies and annotation of large gene families of conifer terpenoid and phenolic defense metabolism.
    Plant J. 2015 Jul;83(2):189-212 PMID: 26017574
  16. Genome assembly using Nanopore-guided long and error-free DNA reads.
    BMC Genomics. 2015 Apr 20;16:327 PMID: 25927464
  17. Assemblathon 1: a competitive assessment of de novo short read assembly methods.
    Genome Res. 2011 Dec;21(12):2224-41 PMID: 21926179
  18. Poretools: a toolkit for analyzing nanopore sequence data.
    Bioinformatics. 2014 Dec 1;30(23):3399-401 PMID: 25143291
  19. QUAST: quality assessment tool for genome assemblies.
    Bioinformatics. 2013 Apr 15;29(8):1072-5 PMID: 23422339
  20. High-quality draft assemblies of mammalian genomes from massively parallel sequence data.
    Proc Natl Acad Sci U S A. 2011 Jan 25;108(4):1513-8 PMID: 21187386
  21. LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads.
    Gigascience. 2015 Aug 04;4:35 PMID: 26244089
  22. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement.
    PLoS One. 2014 Nov 19;9(11):e112963 PMID: 25409509
  23. A complete bacterial genome assembled de novo using only nanopore sequencing data.
    Nat Methods. 2015 Aug;12(8):733-5 PMID: 26076426
  24. ABySS: a parallel assembler for short read sequence data.
    Genome Res. 2009 Jun;19(6):1117-23 PMID: 19251739
  25. MinION nanopore sequencing identifies the position and structure of a bacterial antibiotic resistance island.
    Nat Biotechnol. 2015 Mar;33(3):296-300 PMID: 25485618
Article Info
Journal
GigaScience
Abbr.
Gigascience
ISSN
2047-217X
Published
2015-00-00
Epub
2015-00-04
Pages
35
Language
English
Region
United States
NLM ID
101596872
PMCID
PMC4524009
Subset
IM
Grants
NHGRI NIH HHS · R01 HG007182 · United States
NHGRI NIH HHS · R01HG007182 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]