Home LiteratureArticle Details
PMID: 11836214 Published · ppublish English Journal Article Research Support, U.S. Gov't, P.H.S.

Tolerating some redundancy significantly speeds up clustering of large protein databases.

Bioinformatics (Oxford, England) ·Vol. 18 ·No. 1 ·2002-01-00 ·Pages 77-82

Li W, Jaroszewski L, Godzik A

Abstract

Sequence clustering replaces groups of similar sequences in a database with single representatives. Clustering large protein databases like the NCBI Non-Redundant database (NR) using even the best currently available clustering algorithms is very time-consuming and only practical at relatively high sequence identity thresholds. Our previous program, CD-HI, clustered NR at 90% identity in approximately 1 h and at 75% identity in approximately 1 day on a 1 GHz Linux PC (Li et al., Bioinformatics, 17, 282, 2001); however even faster clustering speed is needed because the size of protein databases are rapidly growing and many applications desire a lower attainable thresholds. For our previous algorithm (CD-HI), we have employed short-word filters to speed up the clustering. In this paper, we show that tolerating some redundancy makes for more efficient use of these short-word filters and increases the program's speed 100 times. Our new program implements this technique and clusters NR at 70% identity within 2 h, and at 50% identity in approximately 5 days. Although some redundancy is present after clustering, our new program's results only differ from our previous program's by less than 0.4%.

MeSH Terms
Algorithms Cluster Analysis Computational Biology Database Management Systems Databases, Protein Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Li Weizhong
The Burnham Institute, 10901 N. Torrey Pines Road, La Jolla, CA 92037, USA. [email protected]
Jaroszewski Lukasz
Godzik Adam
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4803
Published
2002-01-00
Pages
77-82
Language
English
Region
England
NLM ID
9808944
Subset
IM
Grants
NIGMS NIH HHS · GM60049 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]