Home LiteratureArticle Details
PMID: 12962546 Published · ppublish English Comparative Study Journal Article Research Support, U.S. Gov't, P.H.S.

Genome-wide prediction, display and refinement of binding sites with information theory-based models.

BMC bioinformatics ·Vol. 4 ·2003-09-08 ·Pages 38

Gadiraju S, Vyhlidal CA, Leeder JS, Rogan PK

Abstract

We present Delila-genome, a software system for identification, visualization and analysis of protein binding sites in complete genome sequences. Binding sites are predicted by scanning genomic sequences with information theory-based (or user-defined) weight matrices. Matrices are refined by adding experimentally-defined binding sites to published binding sites. Delila-Genome was used to examine the accuracy of individual information contents of binding sites detected with refined matrices as a measure of the strengths of the corresponding protein-nucleic acid interactions. The software can then be used to predict novel sites by rescanning the genome with the refined matrices. Parameters for genome scans are entered using a Java-based GUI interface and backend scripts in Perl. Multi-processor CPU load-sharing minimized the average response time for scans of different chromosomes. Scans of human genome assemblies required 4-6 hours for transcription factor binding sites and 10-19 hours for splice sites, respectively, on 24- and 3-node Mosix and Beowulf clusters. Individual binding sites are displayed either as high-resolution sequence walkers or in low-resolution custom tracks in the UCSC genome browser. For large datasets, we applied a data reduction strategy that limited displays of binding sites exceeding a threshold information content to specific chromosomal regions within or adjacent to genes. An HTML document is produced listing binding sites ranked by binding site strength or chromosomal location hyperlinked to the UCSC custom track, other annotation databases and binding site sequences. Post-genome scan tools parse binding site annotations of selected chromosome intervals and compare the results of genome scans using different weight matrices. Comparisons of multiple genome scans can display binding sites that are unique to each scan and identify sites with significantly altered binding strengths. Delila-Genome was used to scan the human genome sequence with information weight matrices of transcription factor binding sites, including PXR/RXRalpha, AHR and NF-kappaB p50/p65, and matrices for RNA binding sites including splice donor, acceptor, and SC35 recognition sites. Comparisons of genome scans with the original and refined PXR/RXRalpha information weight matrices indicate that the refined model more accurately predicts the strengths of known binding sites and is more sensitive for detection of novel binding sites.

MeSH Terms
Binding Sites/genetics Computer Graphics/instrumentation Computer Systems E-Box Elements/genetics Genome, Human Humans Information Theory Models, Genetic Predictive Value of Tests Programming Languages RNA Splice Sites/genetics Response Elements/genetics User-Computer Interface
Chemicals
RNA Splice Sites
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Gadiraju Sashidhar
Laboratory of Human Molecular Genetics, Children's Mercy Hospital and Clinics, School of Medicine, and School of Interdisciplinary Computer Science and Engineering University of Missouri-Kansas City, Kansas City, MO 64108 USA. [email protected]
Vyhlidal Carrie A
Leeder J Steven
Rogan Peter K
References (16)
16 references, click to expand
  1. Measuring molecular information.
    J Theor Biol. 1999 Nov 7;201(1):87-92 PMID: 10534438
  2. OxyR and SoxRS regulation of fur.
    J Bacteriol. 1999 Aug;181(15):4639-43 PMID: 10419964
  3. Monitoring expression of genes involved in drug metabolism and toxicology using DNA microarrays.
    Physiol Genomics. 2001 Apr 27;5(4):161-70 PMID: 11328961
  4. Anatomy of Escherichia coli ribosome binding sites.
    J Mol Biol. 2001 Oct 12;313(1):215-28 PMID: 11601857
  5. Rifampin is a selective, pleiotropic inducer of drug metabolism genes in human hepatocytes: studies with cDNA and oligonucleotide expression arrays.
    J Pharmacol Exp Ther. 2001 Dec;299(3):849-57 PMID: 11714868
  6. Exploiting transcription factor binding site clustering to identify cis-regulatory modules involved in pattern formation in the Drosophila genome.
    Proc Natl Acad Sci U S A. 2002 Jan 22;99(2):757-62 PMID: 11805330
  7. SCORE: a computational approach to the identification of cis-regulatory modules and target genes in whole-genome sequence data. Site clustering over random expectation.
    Proc Natl Acad Sci U S A. 2002 Jul 23;99(15):9888-93 PMID: 12107285
  8. Information theory-based analysis of CYP2C19, CYP2D6 and CYP3A5 splicing mutations.
    Pharmacogenetics. 2003 Apr;13(4):207-18 PMID: 12668917
  9. A design for computer nucleic-acid-sequence storage, retrieval, and manipulation.
    Nucleic Acids Res. 1982 May 11;10(9):3013-24 PMID: 7099972
  10. Features of spliceosome evolution and function inferred from an analysis of the information at human splice sites.
    J Mol Biol. 1992 Dec 20;228(4):1124-36 PMID: 1474582
  11. Sequence walkers: a graphical method to display how binding proteins interact with DNA or RNA sequences.
    Nucleic Acids Res. 1997 Nov 1;25(21):4408-15 PMID: 9336476
  12. Information analysis of Fis binding sites.
    Nucleic Acids Res. 1997 Dec 15;25(24):4994-5002 PMID: 9396807
  13. Information content of individual genetic sequences.
    J Theor Biol. 1997 Dec 21;189(4):427-41 PMID: 9446751
  14. Information analysis of human splice site mutations.
    Hum Mutat. 1998;12(3):153-71 PMID: 9711873
  15. Using sequence logos and information analysis of Lrp DNA binding sites to investigate discrepancies between natural selection and SELEX.
    Nucleic Acids Res. 1999 Feb 1;27(3):882-7 PMID: 9889287
  16. Characterization of human RNA splice signals by iterative functional selection of splice sites.
    RNA. 2000 Apr;6(4):528-44 PMID: 10786844
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2003-09-08
Epub
2003-00-08
Pages
38
Language
English
Region
England
NLM ID
100965194
PMCID
PMC200970
Subset
IM
Grants
NIEHS NIH HHS · R01 ES010855 · United States
NIEHS NIH HHS · ES 10855 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]