Home LiteratureArticle Details
PMID: 11591649 Published · ppublish English Journal Article

SSAHA: a fast search method for large DNA databases.

Genome research ·Vol. 11 ·No. 10 ·2001-10-00 ·Pages 1725-9

Ning Z, Cox AJ, Mullikin JC

Abstract

We describe an algorithm, SSAHA (Sequence Search and Alignment by Hashing Algorithm), for performing fast searches on databases containing multiple gigabases of DNA. Sequences in the database are preprocessed by breaking them into consecutive k-tuples of k contiguous bases and then using a hash table to store the position of each occurrence of each k-tuple. Searching for a query sequence in the database is done by obtaining from the hash table the "hits" for each k-tuple in the query sequence and then performing a sort on the results. We discuss the effect of the tuple length k on the search speed, memory usage, and sensitivity of the algorithm and present the results of computational experiments which show that SSAHA can be three to four orders of magnitude faster than BLAST or FASTA, while requiring less memory than suffix tree methods. The SSAHA algorithm is used for high-throughput single nucleotide polymorphism (SNP) detection and very large scale sequence assembly. Also, it provides Web-based sequence search facilities for Ensembl projects.

MeSH Terms
Algorithms Base Composition Base Sequence DNA/genetics Database Management Systems/statistics & numerical data Databases, Factual Sensitivity and Specificity Sequence Alignment Software/statistics & numerical data
Chemicals
DNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Ning Z
Informatics Division, The Sanger Centre, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SA, UK.
Cox A J
Mullikin J C
References (13)
13 references, click to expand
  1. A greedy algorithm for aligning DNA sequences.
    J Comput Biol. 2000 Feb-Apr;7(1-2):203-14 PMID: 10890397
  2. An SNP map of the human genome generated by reduced representation shotgun sequencing.
    Nature. 2000 Sep 28;407(6803):513-6 PMID: 11029002
  3. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  4. A map of human genome sequence variation containing 1.42 million single nucleotide polymorphisms.
    Nature. 2001 Feb 15;409(6822):928-33 PMID: 11237013
  5. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  6. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  7. Alignment of whole genomes.
    Nucleic Acids Res. 1999 Jun 1;27(11):2369-76 PMID: 10325427
  8. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  9. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  10. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  11. Tandem repeats finder: a program to analyze DNA sequences.
    Nucleic Acids Res. 1999 Jan 15;27(2):573-80 PMID: 9862982
  12. A RAPID algorithm for sequence database comparisons: application to the identification of vector contamination in the EMBL databases.
    Bioinformatics. 1999 Feb;15(2):111-21 PMID: 10089196
  13. Rapid and sensitive protein similarity searches.
    Science. 1985 Mar 22;227(4693):1435-41 PMID: 2983426
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2001-10-00
Pages
1725-9
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC311141
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]