Abstract
Adapter trimming is a prerequisite step for analyzing next-generation sequencing (NGS) data when the reads are longer than the target DNA/RNA fragments. Although typically used in small RNA sequencing, adapter trimming is also used widely in other applications, such as genome DNA sequencing and transcriptome RNA/cDNA sequencing, where fragments shorter than a read are sometimes obtained because of the limitations of NGS protocols. For the newly emerged Nextera long mate-pair (LMP) protocol, junction adapters are located in the middle of all properly constructed fragments; hence, adapter trimming is essential to gain the correct paired reads. However, our investigations have shown that few adapter trimming tools meet both efficiency and accuracy requirements simultaneously. The performances of these tools can be even worse for paired-end and/or mate-pair sequencing. To improve the efficiency of adapter trimming, we devised a novel algorithm, the bit-masked k-difference matching algorithm, which has O(kn) expected time with O(m) space, where k is the maximum number of differences allowed, n is the read length, and m is the adapter length. This algorithm makes it possible to fully enumerate all candidates that meet a specified threshold, e.g. error ratio, within a short period of time. To improve the accuracy of this algorithm, we designed a simple and easy-to-explain statistical scoring scheme to evaluate candidates in the pattern matching step. We also devised scoring schemes to fully exploit the paired-end/mate-pair information when it is applicable. All these features have been implemented in an industry-standard tool named Skewer (https://sourceforge.net/projects/skewer). Experiments on simulated data, real data of small RNA sequencing, paired-end RNA sequencing, and Nextera LMP sequencing showed that Skewer outperforms all other similar tools that have the same utility. Further, Skewer is considerably faster than other tools that have comparative accuracies; namely, one times faster for single-end sequencing, more than 12 times faster for paired-end sequencing, and 49% faster for LMP sequencing. Skewer achieved as yet unmatched accuracies for adapter trimming with low time bound.
MeSH Terms
Algorithms
Animals
Arabidopsis
Caenorhabditis elegans
Drosophila
High-Throughput Nucleotide Sequencing/methods
Humans
Sequence Analysis, RNA
Software
Time Factors
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Jiang Hongshan
Institute of Plant Quarantine Research, Chinese Academy of Inspection and Quarantine, Huixinli 241, Beijing, 100029 China.
[email protected].
Lei Rong
Ding Shou-Wei
Zhu Shuifang
References (20)
20 references, click to expand
-
RobiNA: a user-friendly, integrated software solution for RNA-Seq-based transcriptomics.
Nucleic Acids Res. 2012 Jul;40(Web Server issue):W622-7
PMID: 22684630
-
Refined DNase-seq protocol and data analysis reveals intrinsic bias in transcription factor footprint identification.
Nat Methods. 2014 Jan;11(1):73-78
PMID: 24317252
-
NextClip: an analysis and read preparation tool for Nextera Long Mate Pair libraries.
Bioinformatics. 2014 Feb 15;30(4):566-8
PMID: 24297520
-
ART: a next-generation sequencing read simulator.
Bioinformatics. 2012 Feb 15;28(4):593-4
PMID: 22199392
-
TopHat: discovering splice junctions with RNA-Seq.
Bioinformatics. 2009 May 1;25(9):1105-11
PMID: 19289445
-
Fast gapped-read alignment with Bowtie 2.
Nat Methods. 2012 Mar 04;9(4):357-9
PMID: 22388286
-
SeqTrim: a high-throughput pipeline for pre-processing any type of sequence read.
BMC Bioinformatics. 2010 Jan 20;11:38
PMID: 20089148
-
Whole-genome sequencing and variant discovery in C. elegans.
Nat Methods. 2008 Feb;5(2):183-8
PMID: 18204455
-
AdapterRemoval: easy cleaning of next-generation sequencing reads.
BMC Res Notes. 2012 Jul 02;5:337
PMID: 22748135
-
FLEXBAR-Flexible Barcode and Adapter Processing for Next-Generation Sequencing Platforms.
Biology (Basel). 2012 Dec 14;1(3):895-905
PMID: 24832523
-
A general method applicable to the search for similarities in the amino acid sequence of two proteins.
J Mol Biol. 1970 Mar;48(3):443-53
PMID: 5420325
-
Btrim: a fast, lightweight adapter and quality trimming program for next-generation sequencing technologies.
Genomics. 2011 Aug;98(2):152-3
PMID: 21651976
-
AlienTrimmer: a tool to quickly and accurately trim off multiple short contaminant sequences from high-throughput sequencing reads.
Genomics. 2013 Nov-Dec;102(5-6):500-6
PMID: 23912058
-
Dynamic expression of small non-coding RNAs, including novel microRNAs and piRNAs/21U-RNAs, during Caenorhabditis elegans development.
Genome Biol. 2009;10(5):R54
PMID: 19460142
-
TagCleaner: Identification and removal of tag sequences from genomic and metagenomic datasets.
BMC Bioinformatics. 2010 Jun 23;11:341
PMID: 20573248
-
Identification of common molecular subsequences.
J Mol Biol. 1981 Mar 25;147(1):195-7
PMID: 7265238
-
TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions.
Genome Biol. 2013 Apr 25;14(4):R36
PMID: 23618408
-
The FlyBase database of the Drosophila genome projects and community literature.
Nucleic Acids Res. 2003 Jan 1;31(1):172-5
PMID: 12519974
-
ABySS: a parallel assembler for short read sequence data.
Genome Res. 2009 Jun;19(6):1117-23
PMID: 19251739
-
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PLoS One. 2013 Dec 23;8(12):e85024
PMID: 24376861