Home LiteratureArticle Details
PMID: 15290780 Published · ppublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

CUBIC: identification of regulatory binding sites through data clustering.

Journal of bioinformatics and computational biology ·Vol. 1 ·No. 1 ·2003-04-00 ·Pages 21-40

Olman V, Xu D, Xu Y

Abstract

Transcription factor binding sites are short fragments in the upstream regions of genes, to which transcription factors bind to regulate the transcription of genes into mRNA. Computational identification of transcription factor binding sites remains an unsolved challenging problem though a great amount of effort has been put into the study of this problem. We have recently developed a novel technique for identification of binding sites from a set of upstream regions of genes, that could possibly be transcriptionally co-regulated and hence might share similar transcription factor binding sites. By utilizing two key features of such binding sites (i.e. their high sequence similarities and their relatively high frequencies compared to other sequence fragments), we have formulated this problem as a cluster identification problem. That is to identify and extract data clusters from a noisy background. While the classical data clustering problem (partitioning a data set into clusters sharing common or similar features) has been extensively studied, there is no general algorithm for solving the problem of identifying data clusters from a noisy background. In this paper, we present a novel algorithm for solving such a problem. We have proved that a cluster identification problem, under our definition, can be rigorously and efficiently solved through searching for substrings with special properties in a linear sequence. We have also developed a method for assessing the statistical significance of each identified cluster, which can be used to rule out accidental data clusters. We have implemented the cluster identification algorithm and the statistical significance analysis method as a computer software CUBIC. Extensive testing on CUBIC has been carried out. We present here a few applications of CUBIC on challenging cases of binding site identification.

MeSH Terms
Algorithms Base Sequence Binding Sites/genetics Cluster Analysis Computational Biology DNA/genetics,metabolism Humans Models, Genetic Models, Statistical RNA, Messenger/genetics,metabolism Saccharomyces cerevisiae/genetics,metabolism Software Transcription Factors/metabolism Transcription, Genetic
Chemicals
RNA, Messenger Transcription Factors DNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Olman Victor
Protein Informatics Group, Life Sciences Division, Oak Ridge National Laboratory, Oak Ridge, TN 37831-6480, USA.
Xu Dong
Xu Ying
Article Info
Journal
Journal of bioinformatics and computational biology
Abbr.
J Bioinform Comput Biol
ISSN
0219-7200
Published
2003-04-00
Pages
21-40
Language
English
Region
Singapore
NLM ID
101187344
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]