Home LiteratureArticle Details
PMID: 21210985 Published · epublish English Evaluation Study Journal Article Research Support, N.I.H., Intramural Research Support, Non-U.S. Gov't

WordSeeker: concurrent bioinformatics software for discovering genome-wide patterns and word-based genomic signatures.

BMC bioinformatics ·Vol. 11 Suppl 12 ·2010-12-21 ·Pages S6

Lichtenberg J, Kurz K, Liang X, Al-ouran R, Neiman L, Nau LJ, Welch JD, Jacox E, Bitterman T, Ecker K, Elnitski L, Drews F, Lee SS, Welch LR

Abstract

An important focus of genomic science is the discovery and characterization of all functional elements within genomes. In silico methods are used in genome studies to discover putative regulatory genomic elements (called words or motifs). Although a number of methods have been developed for motif discovery, most of them lack the scalability needed to analyze large genomic data sets. This manuscript presents WordSeeker, an enumerative motif discovery toolkit that utilizes multi-core and distributed computational platforms to enable scalable analysis of genomic data. A controller task coordinates activities of worker nodes, each of which (1) enumerates a subset of the DNA word space and (2) scores words with a distributed Markov chain model. A comprehensive suite of performance tests was conducted to demonstrate the performance, speedup and efficiency of WordSeeker. The scalability of the toolkit enabled the analysis of the entire genome of Arabidopsis thaliana; the results of the analysis were integrated into The Arabidopsis Gene Regulatory Information Server (AGRIS). A public version of WordSeeker was deployed on the Glenn cluster at the Ohio Supercomputer Center. WordSeeker effectively utilizes concurrent computing platforms to enable the identification of putative functional elements in genomic data sets. This capability facilitates the analysis of the large quantity of sequenced genomic data.

MeSH Terms
Algorithms Arabidopsis/genetics DNA/chemistry Genome, Plant Genomics/methods Markov Chains Regulatory Sequences, Nucleic Acid Sequence Analysis, DNA Software
Chemicals
DNA
Authors & Affiliations
14 authors, click to expand affiliations / ORCID
Lichtenberg Jens
Bioinformatics Laboratory, School of EECS, Ohio University, Athens, Ohio 45701, USA. [email protected]
Kurz Kyle
Liang Xiaoyu
Al-ouran Rami
Neiman Lev
Nau Lee J
Welch Joshua D
Jacox Edwin
Bitterman Thomas
Ecker Klaus
Elnitski Laura
Drews Frank
Lee Stephen Sauchi
Welch Lonnie R
References (19)
19 references, click to expand
  1. WordSpy: identifying transcription factor binding motifs by building a dictionary and learning a grammar.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W412-6 PMID: 15980501
  2. Sole-Search: an integrated analysis program for peak detection and functional annotation using ChIP-seq data.
    Nucleic Acids Res. 2010 Jan;38(3):e13 PMID: 19906703
  3. Finding composite regulatory patterns in DNA sequences.
    Bioinformatics. 2002;18 Suppl 1:S354-63 PMID: 12169566
  4. Short blocks from the noncoding parts of the human genome have instances within nearly all known genes and relate to biological processes.
    Proc Natl Acad Sci U S A. 2006 Apr 25;103(17):6605-10 PMID: 16636294
  5. AGRIS: Arabidopsis gene regulatory information server, an information resource of Arabidopsis cis-regulatory elements and transcription factors.
    BMC Bioinformatics. 2003 Jun 23;4:25 PMID: 12820902
  6. Separating real motifs from their artifacts.
    Bioinformatics. 2001;17 Suppl 1:S30-8 PMID: 11472990
  7. Exceptional motifs in different Markov chain models for a statistical analysis of DNA sequences.
    J Comput Biol. 1995 Fall;2(3):417-37 PMID: 8521272
  8. Combinatorial pattern discovery in biological sequences: The TEIRESIAS algorithm.
    Bioinformatics. 1998;14(1):55-67 PMID: 9520502
  9. REPuter: the manifold applications of repeat analysis on a genomic scale.
    Nucleic Acids Res. 2001 Nov 15;29(22):4633-42 PMID: 11713313
  10. A steganalysis-based approach to comprehensive identification and characterization of functional regulatory elements.
    Genome Biol. 2006;7(6):R49 PMID: 16787547
  11. Seeder: discriminative seeding DNA motif discovery.
    Bioinformatics. 2008 Oct 15;24(20):2303-7 PMID: 18718942
  12. Assessing computational tools for the discovery of transcription factor binding sites.
    Nat Biotechnol. 2005 Jan;23(1):137-44 PMID: 15637633
  13. Word-based characterization of promoters involved in human DNA repair pathways.
    BMC Genomics. 2009 Jul 07;10 Suppl 1:S18 PMID: 19594877
  14. Efficient detection of unusual words.
    J Comput Biol. 2000 Feb-Apr;7(1-2):71-94 PMID: 10890389
  15. The word landscape of the non-coding segments of the Arabidopsis thaliana genome.
    BMC Genomics. 2009 Oct 08;10:463 PMID: 19814816
  16. YMF: A program for discovery of novel transcription factor binding sites by statistical overrepresentation.
    Nucleic Acids Res. 2003 Jul 1;31(13):3586-8 PMID: 12824371
  17. Coding DNA repeated throughout intergenic regions of the Arabidopsis thaliana genome: evolutionary footprints of RNA silencing.
    Mol Biosyst. 2009 Dec;5(12):1679-87 PMID: 19452047
  18. The ENCODE (ENCyclopedia Of DNA Elements) Project.
    Science. 2004 Oct 22;306(5696):636-40 PMID: 15499007
  19. Weeder Web: discovery of transcription factor binding sites in a set of sequences from co-regulated genes.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W199-203 PMID: 15215380
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2010-12-21
Epub
2010-00-21
Pages
S6
Language
English
Region
England
NLM ID
100965194
PMCID
PMC3040532
Subset
IM
Grants
Intramural NIH HHS · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]