Home LiteratureArticle Details
PMID: 17997864 Published · epublish English Journal Article

An optimized procedure greatly improves EST vector contamination removal.

BMC genomics ·Vol. 8 ·2007-11-13 ·Pages 416

Chen YA, Lin CC, Wang CD, Wu HB, Hwang PI

Abstract

The enormous amount of sequence data available in the public domain database has been a gold mine for researchers exploring various themes in life sciences, and hence the quality of such data is of serious concern to researchers. Removal of vector contamination is one of the most significant operations to obtain accurate sequence data containing only a cDNA insert from the basecalls output by an automatic DNA sequencer. Popular bioinformatics programs to accomplish vector trimming include LUCY, cross_match and SeqClean. In a recent study, where the program SeqClean was used to remove vector contamination from our test set of EST data compiled through various library construction systems, however, a significant number of errors remained after preliminary trimming. These errors were later almost completely corrected by simply using a re-linearized form of the cloning vector to compare against the target ESTs. The modified trimming procedure for SeqClean was also compared with the trimming efficiency of the other two popular programs, LUCY2, and cross_match. Using SeqClean with a re-linearized form of the cloning vector significantly surpassed the other two programs in all tested conditions, while the performance of the other two programs was not influenced by the modified procedure. Vector contamination in dbEST was also investigated in this study: 2203 out of the 48212 ESTs sampled from dbEST (2007-04-18 freeze) were found to match sequences in UNIVEC. Vector contamination remains a serious concern to the data quality in the public sequence database nowadays. Based on the results presented here, we feel that our modified procedure with SeqClean should be recommended to all researchers for the task of vector removal from EST or genomic sequences.

MeSH Terms
Base Sequence Expressed Sequence Tags Genetic Vectors Molecular Sequence Data
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Chen Yi-An
Bioinformatics Core Laboratory, Agricultural Biotechnology Research Center, Academia Sinica, Taipei, Taiwan. [email protected]
Lin Chang-Chun
Wang Chin-Di
Wu Huan-Bin
Hwang Pei-Ing
References (18)
18 references, click to expand
  1. Bioinformatics of the Paracoccidioides brasiliensis EST Project.
    Genet Mol Res. 2005 Jun 30;4(2):203-15 PMID: 16110442
  2. DNA sequence quality trimming and vector removal.
    Bioinformatics. 2001 Dec;17(12):1093-104 PMID: 11751217
  3. Quality control in databanks for molecular biology.
    Bioessays. 2000 Nov;22(11):1024-34 PMID: 11056479
  4. Base-calling of automated sequencer traces using phred. II. Error probabilities.
    Genome Res. 1998 Mar;8(3):186-94 PMID: 9521922
  5. dbEST--database for "expressed sequence tags".
    Nat Genet. 1993 Aug;4(4):332-3 PMID: 8401577
  6. Complementary DNA sequencing: expressed sequence tags and human genome project.
    Science. 1991 Jun 21;252(5013):1651-6 PMID: 2047873
  7. Cleaning the GenBank Arabidopsis thaliana data set.
    Nucleic Acids Res. 1996 Jan 15;24(2):316-20 PMID: 8628656
  8. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  9. A strategy for assembling the maize (Zea mays L.) genome.
    Bioinformatics. 2004 Jan 22;20(2):140-7 PMID: 14734303
  10. EST data suggest that poplar is an ancient polyploid.
    New Phytol. 2005 Jul;167(1):165-70 PMID: 15948839
  11. A RAPID algorithm for sequence database comparisons: application to the identification of vector contamination in the EMBL databases.
    Bioinformatics. 1999 Feb;15(2):111-21 PMID: 10089196
  12. Go hunting in sequence databases but watch out for the traps.
    Trends Genet. 1996 Oct;12(10):425-7 PMID: 8909140
  13. Profile and analysis of gene expression changes during early development in germinating spores of Ceratopteris richardii.
    Plant Physiol. 2005 Jul;138(3):1734-45 PMID: 15965014
  14. Establishing a method of vector contamination identification in database sequences.
    Bioinformatics. 1999 Feb;15(2):106-10 PMID: 10089195
  15. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  16. Corruption of genomic databases with anomalous sequence.
    Nucleic Acids Res. 1992 Jun 11;20(11):2741-7 PMID: 1614861
  17. Base-calling of automated sequencer traces using phred. I. Accuracy assessment.
    Genome Res. 1998 Mar;8(3):175-85 PMID: 9521921
  18. PartiGene--constructing partial genomes.
    Bioinformatics. 2004 Jun 12;20(9):1398-404 PMID: 14988115
Article Info
Journal
BMC genomics
Abbr.
BMC Genomics
ISSN
1471-2164
Published
2007-11-13
Epub
2007-00-13
Pages
416
Language
English
Region
England
NLM ID
100965258
PMCID
PMC2194723
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]