Home LiteratureArticle Details
PMID: 18959783 Published · epublish English Journal Article Research Support, Non-U.S. Gov't Validation Study

Computational cluster validation for microarray data analysis: experimental assessment of Clest, Consensus Clustering, Figure of Merit, Gap Statistics and Model Explorer.

BMC bioinformatics ·Vol. 9 ·2008-10-29 ·Pages 462

Giancarlo R, Scaturro D, Utro F

Abstract

Inferring cluster structure in microarray datasets is a fundamental task for the so-called -omic sciences. It is also a fundamental question in Statistics, Data Analysis and Classification, in particular with regard to the prediction of the number of clusters in a dataset, usually established via internal validation measures. Despite the wealth of internal measures available in the literature, new ones have been recently proposed, some of them specifically for microarray data. We consider five such measures: Clest, Consensus (Consensus Clustering), FOM (Figure of Merit), Gap (Gap Statistics) and ME (Model Explorer), in addition to the classic WCSS (Within Cluster Sum-of-Squares) and KL (Krzanowski and Lai index). We perform extensive experiments on six benchmark microarray datasets, using both Hierarchical and K-means clustering algorithms, and we provide an analysis assessing both the intrinsic ability of a measure to predict the correct number of clusters in a dataset and its merit relative to the other measures. We pay particular attention both to precision and speed. Moreover, we also provide various fast approximation algorithms for the computation of Gap, FOM and WCSS. The main result is a hierarchy of those measures in terms of precision and speed, highlighting some of their merits and limitations not reported before in the literature. Based on our analysis, we draw several conclusions for the use of those internal measures on microarray data. We report the main ones. Consensus is by far the best performer in terms of predictive power and remarkably algorithm-independent. Unfortunately, on large datasets, it may be of no use because of its non-trivial computer time demand (weeks on a state of the art PC). FOM is the second best performer although, quite surprisingly, it may not be competitive in this scenario: it has essentially the same predictive power of WCSS but it is from 6 to 100 times slower in time, depending on the dataset. The approximation algorithms for the computation of FOM, Gap and WCSS perform very well, i.e., they are faster while still granting a very close approximation of FOM and WCSS. The approximation algorithm for the computation of Gap deserves to be singled-out since it has a predictive power far better than Gap, it is competitive with the other measures, but it is at least two order of magnitude faster in time with respect to Gap. Another important novel conclusion that can be drawn from our analysis is that all the measures we have considered show severe limitations on large datasets, either due to computational demand (Consensus, as already mentioned, Clest and Gap) or to lack of precision (all of the other measures, including their approximations). The software and datasets are available under the GNU GPL on the supplementary material web page.

MeSH Terms
Algorithms Animals Benchmarking/methods Cluster Analysis Computational Biology/methods Databases, Genetic Humans Oligonucleotide Array Sequence Analysis/methods Software Statistics as Topic/methods
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Giancarlo Raffaele
Dipartimento di Matematica ed Applicazioni, Universitá di Palermo, Palermo, Italy. [email protected]
Scaturro Davide
Utro Filippo
References (14)
14 references, click to expand
  1. Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling.
    Nature. 2000 Feb 3;403(6769):503-11 PMID: 10676951
  2. An algorithm for clustering cDNA fingerprints.
    Genomics. 2000 Jun 15;66(3):249-56 PMID: 10873379
  3. Validating clustering for gene expression data.
    Bioinformatics. 2001 Apr;17(4):309-18 PMID: 11301299
  4. A stability based method for discovering structure in clustered data.
    Pac Symp Biocomput. 2002;:6-17 PMID: 11928511
  5. A prediction-based resampling method for estimating the number of clusters in a dataset.
    Genome Biol. 2002 Jun 25;3(7):RESEARCH0036 PMID: 12184810
  6. Comparisons and validation of statistical clustering techniques for microarray gene expression data.
    Bioinformatics. 2003 Mar 1;19(4):459-66 PMID: 12611800
  7. Computational cluster validation in post-genomic data analysis.
    Bioinformatics. 2005 Aug 1;21(15):3201-12 PMID: 15914541
  8. GenClust: a genetic algorithm for clustering gene expression data.
    BMC Bioinformatics. 2005 Dec 07;6:289 PMID: 16336639
  9. Are clusters found in one dataset present in another dataset?
    Biostatistics. 2007 Jan;8(1):9-31 PMID: 16613834
  10. Evaluation of gene-expression clustering via mutual information distance measure.
    BMC Bioinformatics. 2007 Mar 30;8:111 PMID: 17397530
  11. Determining the number of clusters using the weighted gap statistic.
    Biometrics. 2007 Dec;63(4):1031-7 PMID: 17425640
  12. Replicating Cluster Analysis: Method, Consistency, and Validity.
    Multivariate Behav Res. 1989 Apr 1;24(2):147-61 PMID: 26755276
  13. Large-scale temporal gene expression mapping of central nervous system development.
    Proc Natl Acad Sci U S A. 1998 Jan 6;95(1):334-9 PMID: 9419376
  14. Comprehensive identification of cell cycle-regulated genes of the yeast Saccharomyces cerevisiae by microarray hybridization.
    Mol Biol Cell. 1998 Dec;9(12):3273-97 PMID: 9843569
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2008-10-29
Epub
2008-00-29
Pages
462
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2657801
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]