Home LiteratureArticle Details
PMID: 10764574 Published · ppublish English Journal Article Research Support, U.S. Gov't, Non-P.H.S.

Gene structure prediction by spliced alignment of genomic DNA with protein sequences: increased accuracy by differential splice site scoring.

Journal of molecular biology ·Vol. 297 ·No. 5 ·2000-04-14 ·Pages 1075-85

Usuka J, Brendel V

Abstract

Gene identification in genomic DNA from eukaryotes is complicated by the vast combinatorial possibilities of potential exon assemblies. If the gene encodes a protein that is closely related to known proteins, gene identification is aided by matching similarity of potential translation products to those target proteins. The genomic DNA and protein sequences can be aligned directly by scoring the implied residues of in-frame nucleotide triplets against the protein residues in conventional ways, while allowing for long gaps in the alignment corresponding to introns in the genomic DNA. We describe a novel method for such spliced alignment. The method derives an optimal alignment based on scoring for both sequence similarity of the predicted gene product to the protein sequence and intrinsic splice site strength of the predicted introns. Application of the method to a representative set of 50 known genes from Arabidopsis thaliana showed significant improvement in prediction accuracy compared to previous spliced alignment methods. The method is also more accurate than ab initio gene prediction methods, provided sufficiently close target proteins are available. In view of the fast growth of public sequence repositories, we argue that close targets will be available for the majority of novel genes, making spliced alignment an excellent practical tool for high-throughput automated genome annotation.

MeSH Terms
Algorithms Amino Acid Sequence Arabidopsis/genetics Automation/methods Base Sequence Bias Codon/genetics Computational Biology/methods,statistics & numerical data Exons/genetics Genes, Plant/genetics Genome, Plant Introns/genetics Molecular Sequence Data Nucleotides/genetics Plant Proteins/chemistry,genetics RNA Splicing/genetics Regulatory Sequences, Nucleic Acid/genetics Reproducibility of Results Sensitivity and Specificity Sequence Alignment/methods,statistics & numerical data Sequence Homology, Amino Acid Software
Chemicals
Codon Nucleotides Plant Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Usuka J
Department of Chemistry, Stanford University, Stanford, CA, 94305, USA.
Brendel V
Article Info
Journal
Journal of molecular biology
Abbr.
J Mol Biol
ISSN
0022-2836
Published
2000-04-14
Pages
1075-85
Language
English
Region
England
NLM ID
2985088R
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]