Home LiteratureArticle Details
PMID: 18450519 Published · ppublish English Journal Article

Fast embedding methods for clustering tens of thousands of sequences.

Computational biology and chemistry ·Vol. 32 ·No. 4 ·2008-08-00 ·Pages 282-6

Blackshields G, Larkin M, Wallace IM, Wilm A, Higgins DG

Abstract

Most sequence clustering methods require a full distance matrix to be computed between all pairs of sequences. This requires computer memory and time proportional to N(2) for N sequences. For small N or say up to 10000 or so, this can be accomplished in reasonable times for sequences of moderate length. For very large N, however, this becomes increasingly prohibitive. In this paper, we have tested variations on a class of published embedding methods that have been designed for clustering large numbers of complex objects where the individual distance calculations are expensive. These methods involve embedding the sequences in a space where the similarities within a set of sequences can be closely approximated without having to compute all pair-wise distances. We show how this approach greatly reduces computation time and memory requirements for clustering large numbers of sequences and demonstrate the quality of the clusterings by benchmarking them as guide trees for multiple alignments. Source code is available on request from the authors.

MeSH Terms
Algorithms Amino Acid Sequence Cluster Analysis Databases, Protein Molecular Sequence Data Proteins/chemistry,genetics Sequence Alignment/methods Sequence Homology, Amino Acid Time Factors
Chemicals
Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Blackshields Gordon
UCD Conway Institute of Biomolecular and Biomedical Sciences, University College Dublin, Dublin 4, Ireland. [email protected]
Larkin Mark
Wallace Iain M
Wilm Andreas
Higgins Desmond G
Article Info
Journal
Computational biology and chemistry
Abbr.
Comput Biol Chem
ISSN
1476-928X
Published
2008-08-00
Epub
2008-00-26
Pages
282-6
Language
English
Region
England
NLM ID
101157394
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]