Abstract
We present an algorithm that extracts the binding sites (represented by position-specific weight matrices) for many different transcription factors from the regulatory regions of a genome, without the need for delineating groups of coregulated genes. The algorithm uses the fact that many DNA-binding proteins in bacteria bind to a bipartite motif with two short segments more conserved than the intervening region. It identifies all statistically significant patterns of the form W(1)N(x)W(2), where W(1) and W(2) are two short oligonucleotides separated by x arbitrary bases, and groups them into clusters of similar patterns. These clusters are then used to derive quantitative recognition profiles of putative regulatory proteins. For a given cluster, the algorithm finds the matching sequences plus the flanking regions in the genome and performs a multiple sequence alignment to derive position-specific weight matrices. We have analyzed the Escherichia coli genome with this algorithm and found approximately 1,500 significant patterns, which give rise to approximately 160 distinct position-specific weight matrices. A fraction of these matrices match the binding sites of one-third of the approximately 60 characterized transcription factors with high statistical significance. Many of the remaining matrices are likely to describe binding sites and regulons of uncharacterized transcription factors. The significance of these matrices was evaluated by their specificity, the location of the predicted sites, and the biological functions of the corresponding regulons, allowing us to suggest putative regulatory functions. The algorithm is efficient for analyzing newly sequenced bacterial genomes for which little is known about transcriptional regulation.
MeSH Terms
Algorithms
Bacterial Proteins/metabolism
Base Sequence
Binding Sites
DNA, Bacterial
Genome, Bacterial
Transcription Factors/metabolism
Chemicals
Bacterial Proteins
DNA, Bacterial
Transcription Factors
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Li Hao
Department of Biochemistry, University of California, 513 Parnassus Avenue, San Francisco, CA 94143, USA.
[email protected]
Rhodius Virgil
Gross Carol
Siggia Eric D
References (26)
26 references, click to expand
-
Building a dictionary for genomes: identification of presumptive regulatory sites by statistical analysis.
Proc Natl Acad Sci U S A. 2000 Aug 29;97(18):10096-100
PMID: 10944202
-
Conservation of DNA regulatory motifs and discovery of new motifs in microbial genomes.
Genome Res. 2000 Jun;10(6):744-57
PMID: 10854408
-
Nitrogen regulatory protein C-controlled genes of Escherichia coli: scavenging as a defense against nitrogen limitation.
Proc Natl Acad Sci U S A. 2000 Dec 19;97(26):14674-9
PMID: 11121068
-
RegulonDB (version 3.2): transcriptional regulation and operon organization in Escherichia coli K-12.
Nucleic Acids Res. 2001 Jan 1;29(1):72-4
PMID: 11125053
-
Phylogenetic footprinting of transcription factor binding sites in proteobacterial genomes.
Nucleic Acids Res. 2001 Feb 1;29(3):774-82
PMID: 11160901
-
Comparative gene expression profiles following UV exposure in wild-type and SOS-deficient Escherichia coli.
Genetics. 2001 May;158(1):41-64
PMID: 11333217
-
MultiFun, a multifunctional classification scheme for Escherichia coli K-12 gene products.
Microb Comp Genomics. 2000;5(4):205-22
PMID: 11471834
-
The evolution of DNA regulatory regions for proteo-gamma bacteria by interspecies comparisons.
Genome Res. 2002 Feb;12(2):298-308
PMID: 11827949
-
Probabilistic clustering of sequences: inferring new bacterial regulons by comparative genomics.
Proc Natl Acad Sci U S A. 2002 May 28;99(11):7323-8
PMID: 12032281
-
Comparison of the consensus sequence flanking translational start sites in Drosophila and vertebrates.
Nucleic Acids Res. 1987 Feb 25;15(4):1353-61
PMID: 3822832
-
Selection of DNA binding sites by regulatory proteins. Statistical-mechanical theory and application to operators and promoters.
J Mol Biol. 1987 Feb 20;193(4):723-50
PMID: 3612791
-
Identifying protein-binding sites from unaligned DNA fragments.
Proc Natl Acad Sci U S A. 1989 Feb;86(4):1183-7
PMID: 2919167
-
Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignment.
Science. 1993 Oct 8;262(5131):208-14
PMID: 8211139
-
The value of prior knowledge in discovering motifs with MEME.
Proc Int Conf Intell Syst Mol Biol. 1995;3:21-9
PMID: 7584439
-
The complete genome sequence of Escherichia coli K-12.
Science. 1997 Sep 5;277(5331):1453-62
PMID: 9278503
-
Extracting regulatory sites from the upstream region of yeast genes by computational analysis of oligonucleotide frequencies.
J Mol Biol. 1998 Sep 4;281(5):827-42
PMID: 9719638
-
Finding DNA regulatory motifs within unaligned noncoding sequences clustered by whole-genome mRNA quantitation.
Nat Biotechnol. 1998 Oct;16(10):939-45
PMID: 9788350
-
A comprehensive library of DNA-binding site matrices for 55 proteins applied to the complete Escherichia coli K-12 genome.
J Mol Biol. 1998 Nov 27;284(2):241-54
PMID: 9813115
-
Cluster analysis and display of genome-wide expression patterns.
Proc Natl Acad Sci U S A. 1998 Dec 8;95(25):14863-8
PMID: 9843981
-
The functional and regulatory roles of sigma factors in transcription.
Cold Spring Harb Symp Quant Biol. 1998;63:141-55
PMID: 10384278
-
Identifying DNA and protein patterns with statistically significant alignments of multiple sequences.
Bioinformatics. 1999 Jul-Aug;15(7-8):563-77
PMID: 10487864
-
Clustering gene expression patterns.
J Comput Biol. 1999 Fall-Winter;6(3-4):281-97
PMID: 10582567
-
A web site for the computational analysis of yeast regulatory sequences.
Yeast. 2000 Jan 30;16(2):177-87
PMID: 10641039
-
Inferring regulatory elements from a whole genome. An analysis of Helicobacter pylori sigma(80) family of promoter signals.
J Mol Biol. 2000 Mar 24;297(2):335-53
PMID: 10715205
-
Discovering regulatory elements in non-coding sequences by analysis of spaced dyads.
Nucleic Acids Res. 2000 Apr 15;28(8):1808-18
PMID: 10734201
-
A statistical method for finding transcription factor binding sites.
Proc Int Conf Intell Syst Mol Biol. 2000;8:344-54
PMID: 10977095