Home LiteratureArticle Details
PMID: 15252200 Published · ppublish English Journal Article

MotifPrototyper: a Bayesian profile model for motif families.

Xing EP, Karp RM

Abstract

In this article, we address the problem of modeling generic features of structurally but not textually related DNA motifs, that is, motifs whose consensus sequences are entirely different but nevertheless share "metasequence features" reflecting similarities in the DNA-binding domains of their associated protein recognizers. We present MotifPrototyper, a profile Bayesian model that can capture structural properties typical of particular families of motifs. Each family corresponds to transcription regulatory proteins with similar types of structural signatures in their DNA-binding domains. We show how to train MotifPrototypers from biologically identified motifs categorized according to the TRANSFAC categorization of transcription factors and present empirical results of motif classification, motif parameter estimation, and de novo motif detection by using the learned profile models.

MeSH Terms
Amino Acid Motifs Base Sequence Bayes Theorem DNA/chemistry,genetics Models, Genetic Models, Molecular Nucleic Acid Conformation Protein Conformation Transcription Factors/chemistry,genetics,metabolism
Chemicals
Transcription Factors DNA
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Xing Eric P
Computer Science Division, University of California, Berkeley, CA 94720, USA. [email protected]
Karp Richard M
References (7)
7 references, click to expand
  1. TRANSFAC: an integrated system for gene expression regulation.
    Nucleic Acids Res. 2000 Jan 1;28(1):316-9 PMID: 10592259
  2. Classifying G-protein coupled receptors with support vector machines.
    Bioinformatics. 2002 Jan;18(1):147-59 PMID: 11836223
  3. Logos: a modular bayesian model for de novo motif detection.
    J Bioinform Comput Biol. 2004 Mar;2(1):127-54 PMID: 15272436
  4. Specificity, free energy and information content in protein-DNA interactions.
    Trends Biochem Sci. 1998 Mar;23(3):109-13 PMID: 9581503
  5. Hidden Markov models in computational biology. Applications to protein modeling.
    J Mol Biol. 1994 Feb 4;235(5):1501-31 PMID: 8107089
  6. Dirichlet mixtures: a method for improved detection of weak but significant protein sequence homology.
    Comput Appl Biosci. 1996 Aug;12(4):327-45 PMID: 8902360
  7. An expectation maximization (EM) algorithm for the identification and characterization of common sites in unaligned biopolymer sequences.
    Proteins. 1990;7(1):41-51 PMID: 2184437
Article Info
Journal
Proceedings of the National Academy of Sciences of the United States of America
Abbr.
Proc Natl Acad Sci U S A
ISSN
0027-8424
Published
2004-07-20
Epub
2004-00-13
Pages
10523-8
Language
English
Region
United States
NLM ID
7505876
PMCID
PMC489970
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]