Home LiteratureArticle Details
PMID: 25461763 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

PERGA: a paired-end read guided de novo assembler for extending contigs using SVM and look ahead approach.

PloS one ·Vol. 9 ·No. 12 ·2014-00-00 ·Pages e114253

Zhu X, Leung HC, Chin FY, Yiu SM, Quan G, Liu B, Wang Y

Abstract

Since the read lengths of high throughput sequencing (HTS) technologies are short, de novo assembly which plays significant roles in many applications remains a great challenge. Most of the state-of-the-art approaches base on de Bruijn graph strategy and overlap-layout strategy. However, these approaches which depend on k-mers or read overlaps do not fully utilize information of paired-end and single-end reads when resolving branches. Since they treat all single-end reads with overlapped length larger than a fix threshold equally, they fail to use the more confident long overlapped reads for assembling and mix up with the relative short overlapped reads. Moreover, these approaches have not been special designed for handling tandem repeats (repeats occur adjacently in the genome) and they usually break down the contigs near the tandem repeats. We present PERGA (Paired-End Reads Guided Assembler), a novel sequence-reads-guided de novo assembly approach, which adopts greedy-like prediction strategy for assembling reads to contigs and scaffolds using paired-end reads and different read overlap size ranging from Omax to Omin to resolve the gaps and branches. By constructing a decision model using machine learning approach based on branch features, PERGA can determine the correct extension in 99.7% of cases. When the correct extension cannot be determined, PERGA will try to extend the contig by all feasible extensions and determine the correct extension by using look-ahead approach. Many difficult-resolved branches are due to tandem repeats which are close in the genome. PERGA detects such different copies of the repeats to resolve the branches to make the extension much longer and more accurate. We evaluated PERGA on both Illumina real and simulated datasets ranging from small bacterial genomes to large human chromosome, and it constructed longer and more accurate contigs and scaffolds than other state-of-the-art assemblers. PERGA can be freely downloaded at https://github.com/hitbio/PERGA.

MeSH Terms
High-Throughput Nucleotide Sequencing Microsatellite Repeats Support Vector Machine
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Zhu Xiao
Center for Bioinformatics, School of Computer Science and Technology, Harbin Institute of Technology, Harbin, Heilongjiang, China.
Leung Henry C M
Department of Computer Science, University of Hong Kong, Hong Kong.
Chin Francis Y L
Department of Computer Science, University of Hong Kong, Hong Kong.
Yiu Siu Ming
Department of Computer Science, University of Hong Kong, Hong Kong.
Quan Guangri
National Pilot School of Software, Harbin Institute of Technology, Weihai, Shandong, China.
Liu Bo
Center for Bioinformatics, School of Computer Science and Technology, Harbin Institute of Technology, Harbin, Heilongjiang, China.
Wang Yadong
Center for Bioinformatics, School of Computer Science and Technology, Harbin Institute of Technology, Harbin, Heilongjiang, China.
References (30)
30 references, click to expand
  1. De novo bacterial genome sequencing: millions of very short reads assembled on a desktop computer.
    Genome Res. 2008 May;18(5):802-9 PMID: 18332092
  2. Assembling millions of short DNA sequences using SSAKE.
    Bioinformatics. 2007 Feb 15;23(4):500-1 PMID: 17158514
  3. SHARCGS, a fast and highly accurate short-read assembly algorithm for de novo genomic sequencing.
    Genome Res. 2007 Nov;17(11):1697-706 PMID: 17908823
  4. Extending assembly of short DNA sequences to handle error.
    Bioinformatics. 2007 Nov 1;23(21):2942-4 PMID: 17893086
  5. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  6. The MaSuRCA genome assembler.
    Bioinformatics. 2013 Nov 1;29(21):2669-77 PMID: 23990416
  7. A whole-genome assembly of Drosophila.
    Science. 2000 Mar 24;287(5461):2196-204 PMID: 10731133
  8. Short read fragment assembly of bacterial genomes.
    Genome Res. 2008 Feb;18(2):324-30 PMID: 18083777
  9. An Eulerian path approach to DNA fragment assembly.
    Proc Natl Acad Sci U S A. 2001 Aug 14;98(17):9748-53 PMID: 11504945
  10. Exploring single-sample SNP and INDEL calling with whole-genome de novo assembly.
    Bioinformatics. 2012 Jul 15;28(14):1838-44 PMID: 22569178
  11. IDBA-UD: a de novo assembler for single-cell and metagenomic sequencing data with highly uneven depth.
    Bioinformatics. 2012 Jun 1;28(11):1420-8 PMID: 22495754
  12. GemSIM: general, error-model based simulator of next-generation sequencing data.
    BMC Genomics. 2012;13:74 PMID: 22336055
  13. Efficient de novo assembly of large genomes using compressed data structures.
    Genome Res. 2012 Mar;22(3):549-56 PMID: 22156294
  14. GAGE: A critical evaluation of genome assemblies and assembly algorithms.
    Genome Res. 2012 Mar;22(3):557-67 PMID: 22147368
  15. Repetitive DNA and next-generation sequencing: computational challenges and solutions.
    Nat Rev Genet. 2012 Jan;13(1):36-46 PMID: 22124482
  16. ngs_backbone: a pipeline for read cleaning, mapping and SNP calling using next generation sequence.
    BMC Genomics. 2011;12:285 PMID: 21635747
  17. Quake: quality-aware detection and correction of sequencing errors.
    Genome Biol. 2010;11(11):R116 PMID: 21114842
  18. Optimization of de novo transcriptome assembly from next-generation sequencing data.
    Genome Res. 2010 Oct;20(10):1432-40 PMID: 20693479
  19. Assembly of large genomes using second-generation sequencing.
    Genome Res. 2010 Sep;20(9):1165-73 PMID: 20508146
  20. De novo assembly of human genomes with massively parallel short read sequencing.
    Genome Res. 2010 Feb;20(2):265-72 PMID: 20019144
  21. The sequence and de novo assembly of the giant panda genome.
    Nature. 2010 Jan 21;463(7279):311-7 PMID: 20010809
  22. Sense from sequence reads: methods for alignment and assembly.
    Nat Methods. 2009 Nov;6(11 Suppl):S6-S12 PMID: 19844229
  23. ABySS: a parallel assembler for short read sequence data.
    Genome Res. 2009 Jun;19(6):1117-23 PMID: 19251739
  24. Aggressive assembly of pyrosequencing reads with mates.
    Bioinformatics. 2008 Dec 15;24(24):2818-24 PMID: 18952627
  25. Accurate whole human genome sequencing using reversible terminator chemistry.
    Nature. 2008 Nov 6;456(7218):53-9 PMID: 18987734
  26. Next-generation DNA sequencing.
    Nat Biotechnol. 2008 Oct;26(10):1135-45 PMID: 18846087
  27. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  28. ALLPATHS: de novo assembly of whole-genome shotgun microreads.
    Genome Res. 2008 May;18(5):810-20 PMID: 18340039
  29. Accurate multiplex polony sequencing of an evolved bacterial genome.
    Science. 2005 Sep 9;309(5741):1728-32 PMID: 16081699
  30. Genome sequencing in microfabricated high-density picolitre reactors.
    Nature. 2005 Sep 15;437(7057):376-80 PMID: 16056220
Article Info
Journal
PloS one
Abbr.
PLoS One
ISSN
1932-6203
Published
2014-00-00
Epub
2014-00-02
Pages
e114253
Language
English
Region
United States
NLM ID
101285081
PMCID
PMC4252104
Subset
IM
Grants
NIAAA NIH HHS · R01 AA020404 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]