Home LiteratureArticle Details
PMID: 26803159 Published · ppublish English

ParDRe: faster parallel duplicated reads removal tool for sequencing studies.

Bioinformatics (Oxford, England) ·Vol. 32 ·No. 10 ·0000-00-00

González-Domínguez Jorge, Schmidt Bertil

Abstract

Current next generation sequencing technologies often generate duplicated or near-duplicated reads that (depending on the application scenario) do not provide any interesting biological information but can increase memory requirements and computational time of downstream analysis. In this work we present ParDRe, a de novo parallel tool to remove duplicated and near-duplicated reads through the clustering of Single-End or Paired-End sequences from fasta or fastq files. It uses a novel bitwise approach to compare the suffixes of DNA strings and employs hybrid MPI/multithreading to reduce runtime on multicore systems. We show that ParDRe is up to 27.29 times faster than Fulcrum (a representative state-of-the-art tool) on a platform with two 8-core Sandy-Bridge processors.,Source code in C ++ and MPI running on Linux systems as well as a reference manual are available at https://sourceforge.net/projects/pardre/,[email protected].

Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
Published
0000-00-00
Indexed
2016-05-18
Updated
2016-05-18
Language
English
Country/Region
England
NLM ID
9808944
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]