Home LiteratureArticle Details
PMID: 6572363 Published · ppublish English Journal Article

Rapid similarity searches of nucleic acid and protein data banks.

Wilbur WJ, Lipman DJ

Abstract

With the development of large data banks of protein and nucleic acid sequences, the need for efficient methods of searching such banks for sequences similar to a given sequence has become evident. We present an algorithm for the global comparison of sequences based on matching k-tuples of sequence elements for a fixed k. The method results in substantial reduction in the time required to search a data bank when compared with prior techniques of similarity analysis, with minimal loss in sensitivity. The algorithm has also been adapted, in a separate implementation, to produce rigorous sequence alignments. Currently, using the DEC KL-10 system, we can compare all sequences in the entire Protein Data Bank of the National Biomedical Research Foundation with a 350-residue query sequence in less than 3 min and carry out a similar analysis with a 500-base query sequence against all eukaryotic sequences in the Los Alamos Nucleic Acid Data Base in less than 2 min.

MeSH Terms
Amino Acid Sequence Base Sequence Computers Nucleic Acids/genetics Proteins/genetics
Chemicals
Nucleic Acids Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Wilbur W J
Lipman D J
References (11)
11 references, click to expand
  1. Comparative biosequence metrics.
    J Mol Evol. 1981;18(1):38-46 PMID: 7334527
  2. Pattern recognition in genetic sequences.
    Proc Natl Acad Sci U S A. 1979 Jul;76(7):3041 PMID: 16592667
  3. Efficient algorithms for folding and comparing nucleic acid sequences.
    Nucleic Acids Res. 1982 Jan 11;10(1):197-206 PMID: 6174935
  4. Matching sequences under deletion-insertion constraints.
    Proc Natl Acad Sci U S A. 1972 Jan;69(1):4-6 PMID: 4500555
  5. Enhanced graphic matrix analysis of nucleic acid and protein sequences.
    Proc Natl Acad Sci U S A. 1981 Dec;78(12):7665-9 PMID: 6801656
  6. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  7. Pattern recognition in nucleic acid sequences. I. A general method for finding local homologies and symmetries.
    Nucleic Acids Res. 1982 Jan 11;10(1):247-63 PMID: 6801626
  8. An improved method of testing for evolutionary homology.
    J Mol Biol. 1966 Mar;16(1):9-16 PMID: 5917736
  9. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  10. Computer analysis of nucleic acid regulatory sequences.
    Proc Natl Acad Sci U S A. 1977 Oct;74(10):4401-5 PMID: 270683
  11. Viral src gene products are related to the catalytic chain of mammalian cAMP-dependent protein kinase.
    Proc Natl Acad Sci U S A. 1982 May;79(9):2836-9 PMID: 6283546
Article Info
Journal
Proceedings of the National Academy of Sciences of the United States of America
Abbr.
Proc Natl Acad Sci U S A
ISSN
0027-8424
Published
1983-02-00
Pages
726-30
Language
English
Region
United States
NLM ID
7505876
PMCID
PMC393452
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]