Home LiteratureArticle Details
PMID: 8670622 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Identification of functional elements in unaligned nucleic acid sequences by a novel tuple search algorithm.

Computer applications in the biosciences : CABIOS ·Vol. 12 ·No. 1 ·1996-02-00 ·Pages 71-80

Wolfertstetter F, Frech K, Herrmann G, Werner T

Abstract

We present an algorithm to identify potential functional elements like protein binding sites in DNA sequences, solely from nucleotide sequence data. Prerequisites are a set of at least seven not closely related sequences with a common biological function which is correlated to one or more unknown sequence elements present in most but not necessarily all of the sequences. The algorithm is based on a search for n-tuples which occur at least in a minimum percentage of the sequences with no or one mismatch, which may be at any position of the tuple. In contrast to functional tuples, random tuples show no preferred pattern of mismatch locations within the tuple nor is the conservation extended beyond the tuple. Both features of functional tuples are used to eliminate random tuples. Selection is carried out by maximization of the information content first for the n-tuple, then for a region containing the tuple and finally for the complete binding site. Further matches are found in an additional selection step, using the ConsInd method previously described. The algorithm is capable of identifying and delimiting elements (e.g. protein binding sites) represented by single short cores (e.g. TATA box) in sets of unaligned sequences of about 500 nucleotides using no information other than the nucleotide sequences. Furthermore, we show its ability to identify multiple elements in a set of complete LTR sequences (more than 600 nucleotides per sequence).

MeSH Terms
Algorithms Base Sequence Consensus Sequence Conserved Sequence DNA/chemistry,genetics Databases, Factual Evaluation Studies as Topic Molecular Sequence Data Sequence Analysis, DNA/methods,statistics & numerical data Software
Chemicals
DNA
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Wolfertstetter F
Institut für Säugetiergenetik, GSF-Forschungszentrum für Umwelt und Gesundheit GmbH, Oberschleiáheim, Germany.
Frech K
Herrmann G
Werner T
Article Info
Journal
Computer applications in the biosciences : CABIOS
Abbr.
Comput Appl Biosci
ISSN
0266-7061
Published
1996-02-00
Pages
71-80
Language
English
Region
England
NLM ID
8511758
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]