Home LiteratureArticle Details
PMID: 24376861 Published · epublish English Comparative Study Evaluation Study Journal Article Research Support, Non-U.S. Gov't

An extensive evaluation of read trimming effects on Illumina NGS data analysis.

PloS one ·Vol. 8 ·No. 12 ·2013-00-00 ·Pages e85024

Del Fabbro C, Scalabrin S, Morgante M, Giorgi FM

Abstract

Next Generation Sequencing is having an extremely strong impact in biological and medical research and diagnostics, with applications ranging from gene expression quantification to genotyping and genome reconstruction. Sequencing data is often provided as raw reads which are processed prior to analysis 1 of the most used preprocessing procedures is read trimming, which aims at removing low quality portions while preserving the longest high quality part of a NGS read. In the current work, we evaluate nine different trimming algorithms in four datasets and three common NGS-based applications (RNA-Seq, SNP calling and genome assembly). Trimming is shown to increase the quality and reliability of the analysis, with concurrent gains in terms of execution time and computational resources needed.

MeSH Terms
Algorithms Computational Biology/methods High-Throughput Nucleotide Sequencing/methods Research Design/standards
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Del Fabbro Cristian
Institute of Applied Genomics, Udine, Italy.
Scalabrin Simone
IGA Technology Services, Udine, Italy.
Morgante Michele
Institute of Applied Genomics, Udine, Italy.
Giorgi Federico M
Institute of Applied Genomics, Udine, Italy ; Center for Computational Biology and Bioinformatics, Columbia University, New York, New York, United States of America.
References (40)
40 references, click to expand
  1. De novo assembly and analysis of RNA-seq data.
    Nat Methods. 2010 Nov;7(11):909-12 PMID: 20935650
  2. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  3. ConDeTri--a content dependent read trimmer for Illumina data.
    PLoS One. 2011;6(10):e26314 PMID: 22039460
  4. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks.
    Nat Protoc. 2012 Mar 01;7(3):562-78 PMID: 22383036
  5. SEAL: a distributed short read mapping and duplicate removal tool.
    Bioinformatics. 2011 Aug 1;27(15):2159-60 PMID: 21697132
  6. HapCUT: an efficient and accurate algorithm for the haplotype assembly problem.
    Bioinformatics. 2008 Aug 15;24(16):i153-9 PMID: 18689818
  7. De novo transcriptome assembly in chili pepper (Capsicum frutescens) to identify genes involved in the biosynthesis of capsaicinoids.
    PLoS One. 2013;8(1):e48156 PMID: 23349661
  8. Using MUMmer to identify similar regions in large sequence sets.
    Curr Protoc Bioinformatics. 2003 Feb;Chapter 10:Unit 10.3 PMID: 18428693
  9. Prions are a common mechanism for phenotypic inheritance in wild yeasts.
    Nature. 2012 Feb 15;482(7385):363-8 PMID: 22337056
  10. The Arabidopsis Information Resource (TAIR): improved gene annotation and new tools.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D1202-10 PMID: 22140109
  11. VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing.
    Genome Res. 2012 Mar;22(3):568-76 PMID: 22300766
  12. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  13. A comprehensive comparison of RNA-Seq-based transcriptome analysis from reads to differential gene expression and cross-comparison with microarrays: a case study in Saccharomyces cerevisiae.
    Nucleic Acids Res. 2012 Nov 1;40(20):10084-97 PMID: 22965124
  14. RobiNA: a user-friendly, integrated software solution for RNA-Seq-based transcriptomics.
    Nucleic Acids Res. 2012 Jul;40(Web Server issue):W622-7 PMID: 22684630
  15. Absolute quantification of somatic DNA alterations in human cancer.
    Nat Biotechnol. 2012 May;30(5):413-21 PMID: 22544022
  16. Microevolution of a zoonotic Helicobacter population colonizing the stomach of a human host before and after failed treatment.
    Genome Biol Evol. 2012;4(12):1310-5 PMID: 23196968
  17. The UCSC Genome Browser database: update 2011.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D876-82 PMID: 20959295
  18. ChimeraScan: a tool for identifying chimeric transcription in sequencing data.
    Bioinformatics. 2011 Oct 15;27(20):2903-4 PMID: 21840877
  19. Full-length transcriptome assembly from RNA-Seq data without a reference genome.
    Nat Biotechnol. 2011 May 15;29(7):644-52 PMID: 21572440
  20. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants.
    Nucleic Acids Res. 2010 Apr;38(6):1767-71 PMID: 20015970
  21. SolexaQA: At-a-glance quality assessment of Illumina second-generation sequencing data.
    BMC Bioinformatics. 2010 Sep 27;11:485 PMID: 20875133
  22. A human gut microbial gene catalogue established by metagenomic sequencing.
    Nature. 2010 Mar 4;464(7285):59-65 PMID: 20203603
  23. Saccharomyces Genome Database: the genomics resource of budding yeast.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D700-5 PMID: 22110037
  24. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  25. The sequence and de novo assembly of the giant panda genome.
    Nature. 2010 Jan 21;463(7279):311-7 PMID: 20010809
  26. Comparative study of RNA-seq- and microarray-derived coexpression networks in Arabidopsis thaliana.
    Bioinformatics. 2013 Mar 15;29(6):717-24 PMID: 23376351
  27. The sequence read archive.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D19-21 PMID: 21062823
  28. RNA-Seq: a revolutionary tool for transcriptomics.
    Nat Rev Genet. 2009 Jan;10(1):57-63 PMID: 19015660
  29. Base-calling of automated sequencer traces using phred. II. Error probabilities.
    Genome Res. 1998 Mar;8(3):186-94 PMID: 9521922
  30. The variant call format and VCFtools.
    Bioinformatics. 2011 Aug 1;27(15):2156-8 PMID: 21653522
  31. Symptomatic atherosclerosis is associated with an altered gut metagenome.
    Nat Commun. 2012;3:1245 PMID: 23212374
  32. DNA methylome analysis using short bisulfite sequencing data.
    Nat Methods. 2012 Jan 30;9(2):145-51 PMID: 22290186
  33. Fast identification and removal of sequence contamination from genomic and metagenomic datasets.
    PLoS One. 2011 Mar 09;6(3):e17288 PMID: 21408061
  34. The tomato genome sequence provides insights into fleshy fruit evolution.
    Nature. 2012 May 30;485(7400):635-41 PMID: 22660326
  35. The high-quality draft genome of peach (Prunus persica) identifies unique patterns of genetic diversity, domestication and genome evolution.
    Nat Genet. 2013 May;45(5):487-94 PMID: 23525075
  36. High-quality draft assemblies of mammalian genomes from massively parallel sequence data.
    Proc Natl Acad Sci U S A. 2011 Jan 25;108(4):1513-8 PMID: 21187386
  37. Next-generation sequencing in the clinic: are we ready?
    Nat Rev Genet. 2012 Nov;13(11):818-24 PMID: 23076269
  38. GAM-NGS: genomic assemblies merger for next generation sequencing.
    BMC Bioinformatics. 2013;14 Suppl 7:S6 PMID: 23815503
  39. Quake: quality-aware detection and correction of sequencing errors.
    Genome Biol. 2010;11(11):R116 PMID: 21114842
  40. ABySS: a parallel assembler for short read sequence data.
    Genome Res. 2009 Jun;19(6):1117-23 PMID: 19251739
Article Info
Journal
PloS one
Abbr.
PLoS One
ISSN
1932-6203
Published
2013-00-00
Epub
2013-00-23
Pages
e85024
Language
English
Region
United States
NLM ID
101285081
PMCID
PMC3871669
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]