Home LiteratureArticle Details
PMID: 20053844 Published · ppublish English Comparative Study Journal Article Research Support, N.I.H., Extramural

CD-HIT Suite: a web server for clustering and comparing biological sequences.

Bioinformatics (Oxford, England) ·Vol. 26 ·No. 5 ·2010-03-01 ·Pages 680-2

Huang Y, Niu B, Gao Y, Fu L, Li W

Abstract

CD-HIT is a widely used program for clustering and comparing large biological sequence datasets. In order to further assist the CD-HIT users, we significantly improved this program with more functions and better accuracy, scalability and flexibility. Most importantly, we developed a new web server, CD-HIT Suite, for clustering a user-uploaded sequence dataset or comparing it to another dataset at different identity levels. Users can now interactively explore the clusters within web browsers. We also provide downloadable clusters for several public databases (NCBI NR, Swissprot and PDB) at different identity levels. Free access at http://cd-hit.org

MeSH Terms
Cluster Analysis Computational Biology/methods Databases, Genetic Internet Sequence Alignment Sequence Analysis Software User-Computer Interface
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Huang Ying
California Institute for Telecommunications and Information Technology, University of California San Diego, La Jolla, CA, USA.
Niu Beifang
Gao Ying
Fu Limin
Li Weizhong
References (9)
9 references, click to expand
  1. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.
    Bioinformatics. 2006 Jul 1;22(13):1658-9 PMID: 16731699
  2. UniRef: comprehensive and non-redundant UniProt reference clusters.
    Bioinformatics. 2007 May 15;23(10):1282-8 PMID: 17379688
  3. The Sorcerer II Global Ocean Sampling expedition: expanding the universe of protein families.
    PLoS Biol. 2007 Mar;5(3):e16 PMID: 17355171
  4. Tolerating some redundancy significantly speeds up clustering of large protein databases.
    Bioinformatics. 2002 Jan;18(1):77-82 PMID: 11836214
  5. SMART 6: recent updates and new developments.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D229-32 PMID: 18978020
  6. A core gut microbiome in obese and lean twins.
    Nature. 2009 Jan 22;457(7228):480-4 PMID: 19043404
  7. Gene identification and protein classification in microbial metagenomic sequence data via incremental clustering.
    BMC Bioinformatics. 2008 Apr 10;9:182 PMID: 18402669
  8. Probing metagenomics by rapid cluster analysis of very large datasets.
    PLoS One. 2008;3(10):e3375 PMID: 18846219
  9. Clustering of highly homologous sequences to reduce the size of large protein databases.
    Bioinformatics. 2001 Mar;17(3):282-3 PMID: 11294794
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2010-03-01
Epub
2010-00-06
Pages
680-2
Language
English
Region
England
NLM ID
9808944
PMCID
PMC2828112
Subset
IM
Grants
NCRR NIH HHS · 1R01RR025030 · United States
NCRR NIH HHS · R01 RR025030-02 · United States
NCRR NIH HHS · R01 RR025030-02S1 · United States
NCRR NIH HHS · R01 RR025030-03 · United States
NCRR NIH HHS · R01 RR025030 · United States
NCRR NIH HHS · R01 RR025030-01 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]