Home LiteratureArticle Details
PMID: 28821237 Published · epublish English Journal Article

Evaluation of the impact of Illumina error correction tools on de novo genome assembly.

BMC bioinformatics ·Vol. 18 ·No. 1 ·2017-08-18 ·Pages 374

Heydari M, Miclotte G, Demeester P, Van de Peer Y, Fostier J

Abstract

Recently, many standalone applications have been proposed to correct sequencing errors in Illumina data. The key idea is that downstream analysis tools such as de novo genome assemblers benefit from a reduced error rate in the input data. Surprisingly, a systematic validation of this assumption using state-of-the-art assembly methods is lacking, even for recently published methods. For twelve recent Illumina error correction tools (EC tools) we evaluated both their ability to correct sequencing errors and their ability to improve de novo genome assembly in terms of contig size and accuracy. We confirm that most EC tools reduce the number of errors in sequencing data without introducing many new errors. However, we found that many EC tools suffer from poor performance in certain sequence contexts such as regions with low coverage or regions that contain short repeated or low-complexity sequences. Reads overlapping such regions are often ill-corrected in an inconsistent manner, leading to breakpoints in the resulting assemblies that are not present in assemblies obtained from uncorrected data. Resolving this systematic flaw in future EC tools could greatly improve the applicability of such tools.

Keywords
Error correction Genome assembly Illumina Next-generation sequencing
MeSH Terms
Algorithms Animals Bacteria/genetics Caenorhabditis elegans/genetics DNA/chemistry,metabolism Drosophila/genetics Genome High-Throughput Nucleotide Sequencing Humans Sequence Alignment Sequence Analysis, DNA
Chemicals
DNA
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Heydari Mahdi
Department of Information Technology, Ghent University-imec, IDLab, Ghent, B-9052, Belgium. | Bioinformatics Institute Ghent, Ghent, B-9052, Belgium.
Miclotte Giles
Department of Information Technology, Ghent University-imec, IDLab, Ghent, B-9052, Belgium. | Bioinformatics Institute Ghent, Ghent, B-9052, Belgium.
Demeester Piet
Department of Information Technology, Ghent University-imec, IDLab, Ghent, B-9052, Belgium. | Bioinformatics Institute Ghent, Ghent, B-9052, Belgium.
Van de Peer Yves
Center for Plant Systems Biology, VIB, Ghent, B-9052, Belgium. | Department of Plant Biotechnology and Bioinformatics, Ghent University, Ghent, B-9052, Belgium. | Bioinformatics Institute Ghent, Ghent, B-9052, Belgium. | Department of Genetics, Genome Research Institute, University of Pretoria, Pretoria, South Africa.
Fostier Jan ORCID
Department of Information Technology, Ghent University-imec, IDLab, Ghent, B-9052, Belgium. [email protected]. | Bioinformatics Institute Ghent, Ghent, B-9052, Belgium. [email protected].
References (35)
35 references, click to expand
  1. Quake: quality-aware detection and correction of sequencing errors.
    Genome Biol. 2010;11(11):R116 PMID: 21114842
  2. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler.
    Gigascience. 2012 Dec 27;1(1):18 PMID: 23587118
  3. Aggressive assembly of pyrosequencing reads with mates.
    Bioinformatics. 2008 Dec 15;24(24):2818-24 PMID: 18952627
  4. Alignment of whole genomes.
    Nucleic Acids Res. 1999 Jun 1;27(11):2369-76 PMID: 10325427
  5. BayesHammer: Bayesian clustering for error correction in single-cell sequencing.
    BMC Genomics. 2013;14 Suppl 1:S7 PMID: 23368723
  6. Lighter: fast and memory-efficient sequencing error correction without counting.
    Genome Biol. 2014;15(11):509 PMID: 25398208
  7. How to apply de Bruijn graphs to genome assembly.
    Nat Biotechnol. 2011 Nov 08;29(11):987-91 PMID: 22068540
  8. A survey of error-correction methods for next-generation sequencing.
    Brief Bioinform. 2013 Jan;14(1):56-66 PMID: 22492192
  9. ACE: accurate correction of errors using K-mer tries.
    Bioinformatics. 2015 Oct 1;31(19):3216-8 PMID: 26026137
  10. Blue: correcting sequencing errors using consensus and context.
    Bioinformatics. 2014 Oct;30(19):2723-32 PMID: 24919879
  11. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  12. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers.
    Bioinformatics. 2011 Mar 15;27(6):764-70 PMID: 21217122
  13. QuorUM: An Error Corrector for Illumina Reads.
    PLoS One. 2015 Jun 17;10(6):e0130821 PMID: 26083032
  14. BLESS: bloom filter-based error correction solution for high-throughput sequencing reads.
    Bioinformatics. 2014 May 15;30(10):1354-62 PMID: 24451628
  15. Denoising DNA deep sequencing data-high-throughput sequencing errors and their correction.
    Brief Bioinform. 2016 Jan;17 (1):154-79 PMID: 26026159
  16. Efficient de novo assembly of large genomes using compressed data structures.
    Genome Res. 2012 Mar;22(3):549-56 PMID: 22156294
  17. Gossamer--a resource-efficient de novo assembler.
    Bioinformatics. 2012 Jul 15;28(14):1937-8 PMID: 22611131
  18. RACER: Rapid and accurate correction of errors in reads.
    Bioinformatics. 2013 Oct 1;29(19):2490-3 PMID: 23853064
  19. BFC: correcting Illumina sequencing errors.
    Bioinformatics. 2015 Sep 1;31(17):2885-7 PMID: 25953801
  20. ABySS: a parallel assembler for short read sequence data.
    Genome Res. 2009 Jun;19(6):1117-23 PMID: 19251739
  21. Evaluation of genomic high-throughput sequencing data generated on Illumina HiSeq and genome analyzer systems.
    Genome Biol. 2011 Nov 08;12(11):R112 PMID: 22067484
  22. EC: an efficient error correction algorithm for short reads.
    BMC Bioinformatics. 2015;16 Suppl 17:S2 PMID: 26678663
  23. Comprehensive variation discovery in single human genomes.
    Nat Genet. 2014 Dec;46(12):1350-5 PMID: 25326702
  24. Trowel: a fast and accurate error correction module for Illumina sequencing reads.
    Bioinformatics. 2014 Nov 15;30(22):3264-5 PMID: 25075116
  25. QUAST: quality assessment tool for genome assemblies.
    Bioinformatics. 2013 Apr 15;29(8):1072-5 PMID: 23422339
  26. Pollux: platform independent error correction of single and mixed genomes.
    BMC Bioinformatics. 2015 Jan 16;16:10 PMID: 25592313
  27. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  28. Musket: a multistage k-mer spectrum-based error corrector for Illumina sequence data.
    Bioinformatics. 2013 Feb 1;29(3):308-15 PMID: 23202746
  29. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing.
    J Comput Biol. 2012 May;19(5):455-77 PMID: 22506599
  30. Characterizing and measuring bias in sequence data.
    Genome Biol. 2013 May 29;14(5):R51 PMID: 23718773
  31. BLESS 2: accurate, memory-efficient and fast error correction method.
    Bioinformatics. 2016 Aug 1;32(15):2369-71 PMID: 27153708
  32. Fiona: a parallel and automatic strategy for read error correction.
    Bioinformatics. 2014 Sep 1;30(17):i356-63 PMID: 25161220
  33. ART: a next-generation sequencing read simulator.
    Bioinformatics. 2012 Feb 15;28(4):593-4 PMID: 22199392
  34. Karect: accurate correction of substitution, insertion and deletion errors for next-generation sequencing data.
    Bioinformatics. 2015 Nov 1;31(21):3421-8 PMID: 26177965
  35. Correcting Illumina data.
    Brief Bioinform. 2015 Jul;16(4):588-99 PMID: 25183248
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2017-08-18
Epub
2017-00-18
Pages
374
Language
English
Region
England
NLM ID
100965194
PMCID
PMC5563063
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]