Home LiteratureArticle Details
PMID: 1304897 Published · ppublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't

Comprehensive sequence analysis of the 182 predicted open reading frames of yeast chromosome III.

Protein science : a publication of the Protein Society ·Vol. 1 ·No. 12 ·1992-12-00 ·Pages 1677-90

Bork P, Ouzounis C, Sander C, Scharf M, Schneider R, Sonnhammer E

Abstract

With the completion of the first phase of the European yeast genome sequencing project, the complete DNA sequence of chromosome III of Saccharomyces cerevisiae has become available (Oliver, S. G., et al., 1992, Nature 357, 38-46). We have tested the predictive power of computer sequence analysis of the 176 probable protein products of this chromosome, after exclusion of six problem cases. When the results of database similarity searches are pooled with prior knowledge, a likely function can be assigned to 42% of the proteins, and a predicted three-dimensional structure to a third of these (14% of the total). The function of the remaining 58% remains to be determined. Of these, about one-third have one or more probable transmembrane segments. Among the most interesting proteins with predicted functions are a new member of the type X polymerase family, a transcription factor with an N-terminal DNA-binding domain related to GAL4, a "fork head" DNA-binding domain previously known only in Drosophila and in mammals, and a putative methyltransferase. Our analysis increased the number of known significant sequence similarities on chromosome III by 13, to now 67. Although the near 40% success rate of identifying unknown protein function by sequence analysis is surprisingly high, the information gap between known protein sequences and unknown function is expected to widen and become a major bottleneck of genome projects in the near future. Based on the experience gained in this test study, we suggest that the development of an automated computer workbench for protein sequence analysis must be an important item in genome projects.

MeSH Terms
Acetolactate Synthase/genetics Amino Acid Sequence Animals Cattle Chromosome Mapping Chromosomes, Fungal DNA, Fungal/genetics DNA-Directed DNA Polymerase/genetics Enzymes/genetics Fungal Proteins/genetics Humans Molecular Sequence Data Open Reading Frames Rats Saccharomyces cerevisiae/genetics Sequence Homology, Amino Acid
Chemicals
DNA, Fungal Enzymes Fungal Proteins Acetolactate Synthase DNA-Directed DNA Polymerase
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Bork P
European Molecular Biology Laboratory, Heidelberg, Germany.
Ouzounis C
Sander C
Scharf M
Schneider R
Sonnhammer E
References (38)
38 references, click to expand
  1. Yeast gene SRP1 (serine-rich protein). Intragenic repeat structure and identification of a family of SRP1-related DNA sequences.
    J Mol Biol. 1988 Aug 5;202(3):455-70 PMID: 3139887
  2. ERS1 a seven transmembrane domain protein from Saccharomyces cerevisiae.
    Nucleic Acids Res. 1990 Apr 25;18(8):2177 PMID: 2186379
  3. Proteins.
    Sci Am. 1985 Oct;253(4):88-99 PMID: 4071032
  4. An improved algorithm for matching biological sequences.
    J Mol Biol. 1982 Dec 15;162(3):705-8 PMID: 7166760
  5. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  6. A simple method for displaying the hydropathic character of a protein.
    J Mol Biol. 1982 May 5;157(1):105-32 PMID: 7108955
  7. The fork head domain: a novel DNA binding motif of eukaryotic transcription factors?
    Cell. 1990 Nov 2;63(3):455-6 PMID: 2225060
  8. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  9. Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes.
    Proc Natl Acad Sci U S A. 1990 Mar;87(6):2264-8 PMID: 2315319
  10. Cystic fibrosis. A transport problem?
    Nature. 1990 Jul 26;346(6282):312-3 PMID: 2374603
  11. The Protein Data Bank: a computer-based archival file for macromolecular structures.
    J Mol Biol. 1977 May 25;112(3):535-42 PMID: 875032
  12. Methods and algorithms for statistical analysis of protein sequences.
    Proc Natl Acad Sci U S A. 1992 Mar 15;89(6):2002-6 PMID: 1549558
  13. A new algorithm for best subsequence alignments with application to tRNA-rRNA comparisons.
    J Mol Biol. 1987 Oct 20;197(4):723-8 PMID: 2448477
  14. Isolation, expression and phylogenetic inheritance of an acetolactate synthase gene from Brassica napus.
    Mol Gen Genet. 1989 Nov;219(3):413-20 PMID: 2482934
  15. Sequence of the D-aspartyl/L-isoaspartyl protein methyltransferase from human erythrocytes. Common sequence motifs for protein, DNA, RNA, and small molecule S-adenosylmethionine-dependent methyltransferases.
    J Biol Chem. 1989 Nov 25;264(33):20131-9 PMID: 2684970
  16. Sequence motifs characteristic of DNA[cytosine-N4]methyltransferases: similarity to adenine and cytosine-C5 DNA-methylases.
    Nucleic Acids Res. 1989 Dec 11;17(23):9823-32 PMID: 2690010
  17. Pilin expression in Neisseria gonorrhoeae is under both positive and negative transcriptional control.
    EMBO J. 1988 Dec 20;7(13):4367-78 PMID: 2854063
  18. Nucleotide sequence characterization of Ty 1-17, a class II transposon from yeast.
    Nucleic Acids Res. 1985 Sep 25;13(18):6679-93 PMID: 2997719
  19. DNA recognition by GAL4: structure of a protein-DNA complex.
    Nature. 1992 Apr 2;356(6368):408-14 PMID: 1557122
  20. The complete DNA sequence of yeast chromosome III.
    Nature. 1992 May 7;357(6373):38-46 PMID: 1574125
  21. Sequence of the HMR region on chromosome III of Saccharomyces cerevisiae.
    Yeast. 1992 Mar;8(3):215-22 PMID: 1574927
  22. Chance and statistical significance in protein and DNA sequence analysis.
    Science. 1992 Jul 3;257(5066):39-49 PMID: 1621093
  23. The complete sequence of a 10.8 kb segment distal of SUF2 on the right arm of chromosome III from Saccharomyces cerevisiae reveals seven open reading frames including the RVS161, ADP1 and PGK genes.
    Yeast. 1992 May;8(5):409-17 PMID: 1626432
  24. What's in a genome?
    Nature. 1992 Jul 23;358(6384):287 PMID: 1641000
  25. The WD-40 repeat.
    FEBS Lett. 1992 Jul 28;307(2):131-4 PMID: 1644165
  26. The nucleotide sequence of a third cyclophilin-homologous gene from Saccharomyces cerevisiae.
    Yeast. 1991 Dec;7(9):971-9 PMID: 1803821
  27. Compilation and alignment of DNA polymerase sequences.
    Nucleic Acids Res. 1991 Aug 11;19(15):4045-57 PMID: 1870963
  28. TIP 1, a cold shock-inducible gene of Saccharomyces cerevisiae.
    J Biol Chem. 1991 Sep 15;266(26):17537-44 PMID: 1894636
  29. Cloning of a cellular factor, interleukin binding factor, that binds to NFAT-like motifs in the human immunodeficiency virus long terminal repeat.
    Proc Natl Acad Sci U S A. 1991 Sep 1;88(17):7739-43 PMID: 1909027
  30. The complete sequence of the 8.2 kb segment left of MAT on chromosome III reveals five ORFs, including a gene for a yeast ribokinase.
    Yeast. 1990 Nov-Dec;6(6):521-34 PMID: 1964349
  31. Database algorithm for generating protein backbone and side-chain co-ordinates from a C alpha trace application to model building and detection of co-ordinate errors.
    J Mol Biol. 1991 Mar 5;218(1):183-94 PMID: 2002501
  32. Database of homology-derived protein structures and the structural meaning of sequence alignment.
    Proteins. 1991;9(1):56-68 PMID: 2017436
  33. Predicting coiled coils from protein sequences.
    Science. 1991 May 24;252(5009):1162-4 PMID: 2031185
  34. Aspartic acid residues at positions 190 and 192 of rat DNA polymerase beta are involved in primer binding.
    Biochemistry. 1991 May 28;30(21):5286-92 PMID: 2036395
  35. The PIR protein sequence database.
    Nucleic Acids Res. 1991 Apr 25;19 Suppl:2231-36 PMID: 2041808
  36. PROSITE: a dictionary of sites and patterns in proteins.
    Nucleic Acids Res. 1991 Apr 25;19 Suppl:2241-5 PMID: 2041810
  37. The SWISS-PROT protein sequence data bank.
    Nucleic Acids Res. 1991 Apr 25;19 Suppl:2247-9 PMID: 2041811
  38. Structural principles of parallel beta-barrels in proteins.
    Proc Natl Acad Sci U S A. 1988 May;85(10):3338-42 PMID: 3368445
Article Info
Journal
Protein science : a publication of the Protein Society
Abbr.
Protein Sci
ISSN
0961-8368
Published
1992-12-00
Pages
1677-90
Language
English
Region
United States
NLM ID
9211750
PMCID
PMC2142145
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]