Home LiteratureArticle Details
PMID: 1480466 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S. Research Support, U.S. Gov't, P.H.S. Review

Assessment of protein coding measures.

Nucleic acids research ·Vol. 20 ·No. 24 ·1992-12-25 ·Pages 6441-50

Fickett JW, Tung CS

Abstract

A number of methods for recognizing protein coding genes in DNA sequence have been published over the last 13 years, and new, more comprehensive algorithms, drawing on the repertoire of existing techniques, continue to be developed. To optimize continued development, it is valuable to systematically review and evaluate published techniques. At the core of most gene recognition algorithms is one or more coding measures--functions which produce, given any sample window of sequence, a number or vector intended to measure the degree to which a sample sequence resembles a window of 'typical' exonic DNA. In this paper we review and synthesize the underlying coding measures from published algorithms. A standardized benchmark is described, and each of the measures is evaluated according to this benchmark. Our main conclusion is that a very simple and obvious measure--counting oligomers--is more effective than any of the more sophisticated measures. Different measures contain different information. However there is a great deal of redundancy in the current suite of measures. We show that in future development of gene recognition algorithms, attention can probably be limited to six of the twenty or so measures proposed to date.

MeSH Terms
Algorithms Base Composition Base Sequence Codon/genetics DNA/genetics Exons Fourier Analysis Genes Genetic Techniques Humans Proteins/genetics
Chemicals
Codon Proteins DNA
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Fickett J W
Theoretical Biology and Biophysics Group, Los Alamos National Laboratory, NM 87545.
Tung C S
References (40)
40 references, click to expand
  1. Oligopeptide biases in protein sequences and their use in predicting protein coding regions in nucleotide sequences.
    Proteins. 1988;4(2):99-122 PMID: 3227018
  2. Database bias and the identification of protein coding sequences.
    DNA. 1987 Oct;6(5):493-5 PMID: 3677996
  3. Probability of coding of a DNA sequence: an algorithm to predict translated reading frames from their thermodynamic characteristics.
    Nucleic Acids Res. 1986 Jan 10;14(1):127-35 PMID: 3753761
  4. A method to locate protein coding sequences in DNA of prokaryotic systems.
    Nucleic Acids Res. 1985 Jan 11;13(1):185-94 PMID: 3839071
  5. The relationship between base composition and codon usage in bacterial genes and its use for the simple and reliable identification of protein-coding sequences.
    Gene. 1984 Oct;30(1-3):157-66 PMID: 6096212
  6. A prevalent persistent global nonrandomness that distinguishes coding and non-coding eucaryotic nuclear DNA sequences.
    J Mol Evol. 1983;19(2):122-33 PMID: 6571217
  7. Method to determine the reading frame of a protein from the purine/pyrimidine genome sequence and its possible evolutionary justification.
    Proc Natl Acad Sci U S A. 1981 Mar;78(3):1596-600 PMID: 6940175
  8. Recognition of protein coding regions in DNA sequences.
    Nucleic Acids Res. 1982 Sep 11;10(17):5303-18 PMID: 7145702
  9. Identifying coding exons by similarity search: alu-derived and other potentially misleading protein sequences.
    Genomics. 1992 Apr;12(4):838-41 PMID: 1572661
  10. gm: a practical tool for automating DNA sequence analysis.
    Comput Appl Biosci. 1990 Jul;6(3):263-70 PMID: 2242161
  11. Translation framing code and frame-monitoring mechanism as suggested by the analysis of mRNA and 16 S rRNA nucleotide sequences.
    J Mol Biol. 1987 Apr 20;194(4):643-52 PMID: 2443708
  12. The frequency of oligonucleotides in mammalian genic regions.
    Comput Appl Biosci. 1989 Feb;5(1):33-40 PMID: 2924169
  13. Distribution and evolution of sequence characteristics in the E. coli genome.
    J Biomol Struct Dyn. 1986 Oct;4(2):291-307 PMID: 3078231
  14. Statistical method for predicting protein coding regions in nucleic acid sequences.
    Comput Appl Biosci. 1987 Nov;3(4):287-95 PMID: 3134115
  15. Computer methods for analyzing sequence recognition of nucleic acids.
    Annu Rev Biophys Biophys Chem. 1988;17:241-63 PMID: 3293587
  16. [Statistical characteristics in primary structures of functional regions of Escherichia coli genome. II. Non-stationary Markov chains].
    Mol Biol (Mosk). 1986 Jul-Aug;20(4):1024-33 PMID: 3531811
  17. [Statistical characteristics of primary structures of the functional regions of the Escherichia coli genome. III. Computer recognition of coding regions].
    Mol Biol (Mosk). 1986 Sep-Oct;20(5):1390-8 PMID: 3534549
  18. Periodicities in introns.
    Nucleic Acids Res. 1987 Sep 25;15(18):7581-92 PMID: 3658704
  19. A measure of DNA periodicity.
    J Theor Biol. 1986 Feb 7;118(3):295-300 PMID: 3713213
  20. Heuristic informational analysis of sequences.
    Nucleic Acids Res. 1986 Jan 10;14(1):179-96 PMID: 3753763
  21. New statistical approach to discriminate between protein coding and non-coding regions in DNA sequences and its evaluation.
    J Theor Biol. 1986 May 21;120(2):223-36 PMID: 3784581
  22. Delineation of coding areas in DNA sequences through assignment of codon probabilities.
    J Biomol Struct Dyn. 1985 Dec;3(3):543-9 PMID: 3917036
  23. Measurements of the effects that coding for a protein has on a DNA sequence and their use for finding genes.
    Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):551-67 PMID: 6364041
  24. The coding function of nucleotide sequences can be discerned by statistical analysis.
    J Theor Biol. 1981 Feb 7;88(3):409-20 PMID: 6456380
  25. The codon preference plot: graphic analysis of protein coding sequences and prediction of gene expression.
    Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):539-49 PMID: 6694906
  26. Codon preference and its use in identifying protein coding regions in long DNA sequences.
    Nucleic Acids Res. 1982 Jan 11;10(1):141-56 PMID: 7063399
  27. Distance, size and shape.
    Ann Eugen. 1954 Mar;18(4):337-43 PMID: 13149002
  28. The C. elegans genome sequencing project: a beginning.
    Nature. 1992 Mar 5;356(6364):37-41 PMID: 1538779
  29. The complete DNA sequence of yeast chromosome III.
    Nature. 1992 May 7;357(6373):38-46 PMID: 1574125
  30. K-tuple frequency analysis: from intron/exon discrimination to T-cell epitope mapping.
    Methods Enzymol. 1990;183:237-52 PMID: 1690334
  31. Electronic data publishing and GenBank.
    Science. 1991 May 31;252(5010):1273-7 PMID: 1925538
  32. The EMBL Data Library.
    Nucleic Acids Res. 1992 May 11;20 Suppl:2071-4 PMID: 1598236
  33. Prediction of gene structure.
    J Mol Biol. 1992 Jul 5;226(1):141-57 PMID: 1619647
  34. Locating protein-coding regions in human DNA sequences by a multiple sensor-neural network approach.
    Proc Natl Acad Sci U S A. 1991 Dec 15;88(24):11261-5 PMID: 1763041
  35. Finding protein coding regions in genomic sequences.
    Methods Enzymol. 1990;183:163-80 PMID: 2314274
  36. DISTAN--a program which detects significant distances between short oligonucleotides.
    Comput Appl Biosci. 1987 Sep;3(3):193-201 PMID: 2455588
  37. Construction of a facsimile data set for large genome sequence analysis.
    Genomics. 1990 Sep;8(1):71-82 PMID: 2081603
  38. Nucleotide distribution and the recognition of coding regions in DNA sequences: an information theory approach.
    J Theor Biol. 1985 Nov 7;117(1):127-36 PMID: 3001434
  39. [Statistical characteristics in primary structures of functional regions of Escherichia coli genome. I. Frequency characteristics].
    Mol Biol (Mosk). 1986 Jul-Aug;20(4):1014-23 PMID: 3531810
  40. The codon Adaptation Index--a measure of directional synonymous codon usage bias, and its potential applications.
    Nucleic Acids Res. 1987 Feb 11;15(3):1281-95 PMID: 3547335
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
0305-1048
Published
1992-12-25
Pages
6441-50
Language
English
Region
England
NLM ID
0411011
PMCID
PMC334555
Subset
IM
Grants
NIGMS NIH HHS · GM-37812 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]