Home LiteratureArticle Details
PMID: 20089148 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

SeqTrim: a high-throughput pipeline for pre-processing any type of sequence read.

BMC bioinformatics ·Vol. 11 ·2010-01-20 ·Pages 38

Falgueras J, Lara AJ, Fernández-Pozo N, Cantón FR, Pérez-Trabado G, Claros MG

Abstract

High-throughput automated sequencing has enabled an exponential growth rate of sequencing data. This requires increasing sequence quality and reliability in order to avoid database contamination with artefactual sequences. The arrival of pyrosequencing enhances this problem and necessitates customisable pre-processing algorithms. SeqTrim has been implemented both as a Web and as a standalone command line application. Already-published and newly-designed algorithms have been included to identify sequence inserts, to remove low quality, vector, adaptor, low complexity and contaminant sequences, and to detect chimeric reads. The availability of several input and output formats allows its inclusion in sequence processing workflows. Due to its specific algorithms, SeqTrim outperforms other pre-processors implemented as Web services or standalone applications. It performs equally well with sequences from EST libraries, SSH libraries, genomic DNA libraries and pyrosequencing reads and does not lead to over-trimming. SeqTrim is an efficient pipeline designed for pre-processing of any type of sequence read, including next-generation sequencing. It is easily configurable and provides a friendly interface that allows users to know what happened with sequences at every pre-processing stage, and to verify pre-processing of an individual sequence if desired. The recommended pipeline reveals more information about each sequence than previously described pre-processors and can discard more sequencing or experimental artefacts.

MeSH Terms
Base Sequence Computational Biology/methods Databases, Genetic Expressed Sequence Tags Gene Library Sequence Analysis, DNA/methods Software
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Falgueras Juan
Departamento de Lenguajes y Ciencias de la Computación, Universidad de Málaga, Málaga, Spain.
Lara Antonio J
Fernández-Pozo Noé
Cantón Francisco R
Pérez-Trabado Guillermo
Claros M Gonzalo
References (16)
16 references, click to expand
  1. Identifying adaptor contamination when mining DNA sequence data.
    Biotechniques. 2004 Aug;37(2):194, 196, 198 PMID: 15335207
  2. An optimized procedure greatly improves EST vector contamination removal.
    BMC Genomics. 2007 Nov 13;8:416 PMID: 17997864
  3. A new DNA sequence assembly program.
    Nucleic Acids Res. 1995 Dec 25;23(24):4992-9 PMID: 8559656
  4. An optimized protocol for analysis of EST sequences.
    Nucleic Acids Res. 2000 Sep 15;28(18):3657-65 PMID: 10982889
  5. Figaro: a novel statistical method for vector sequence removal.
    Bioinformatics. 2008 Feb 15;24(4):462-7 PMID: 18202027
  6. EST2uni: an open, parallel tool for automated EST analysis and database creation, with a data mining web interface and microarray expression data integration.
    BMC Bioinformatics. 2008 Jan 07;9:5 PMID: 18179701
  7. Base-calling of automated sequencer traces using phred. II. Error probabilities.
    Genome Res. 1998 Mar;8(3):186-94 PMID: 9521922
  8. Base-calling of automated sequencer traces using phred. I. Accuracy assessment.
    Genome Res. 1998 Mar;8(3):175-85 PMID: 9521921
  9. LUCY2: an interactive DNA sequence quality trimming and vector removal tool.
    Bioinformatics. 2004 Nov 1;20(16):2865-6 PMID: 15130926
  10. ESTAnnotator: A tool for high throughput EST annotation.
    Nucleic Acids Res. 2003 Jul 1;31(13):3716-9 PMID: 12824401
  11. Establishing a method of vector contamination identification in database sequences.
    Bioinformatics. 1999 Feb;15(2):106-10 PMID: 10089195
  12. ESTprep: preprocessing cDNA sequence reads.
    Bioinformatics. 2003 Jul 22;19(11):1318-24 PMID: 12874042
  13. ESTpass: a web-based server for processing and annotating expressed sequence tag (EST) sequences.
    Nucleic Acids Res. 2007 Jul;35(Web Server issue):W159-62 PMID: 17526512
  14. Repbase update: a database and an electronic journal of repetitive elements.
    Trends Genet. 2000 Sep;16(9):418-20 PMID: 10973072
  15. EGassembler: online bioinformatics service for large-scale processing, clustering and assembling ESTs and genomic DNA fragments.
    Nucleic Acids Res. 2006 Jul 1;34(Web Server issue):W459-62 PMID: 16845049
  16. ESTExplorer: an expressed sequence tag (EST) assembly and annotation platform.
    Nucleic Acids Res. 2007 Jul;35(Web Server issue):W143-7 PMID: 17545197
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2010-01-20
Epub
2010-00-20
Pages
38
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2832897
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]