Home LiteratureArticle Details
PMID: 16945146 Published · epublish English Evaluation Study Journal Article Research Support, U.S. Gov't, Non-P.H.S.

Methods for evaluating clustering algorithms for gene expression data using a reference set of functional classes.

BMC bioinformatics ·Vol. 7 ·2006-08-31 ·Pages 397

Datta S, Datta S

Abstract

A cluster analysis is the most commonly performed procedure (often regarded as a first step) on a set of gene expression profiles. In most cases, a post hoc analysis is done to see if the genes in the same clusters can be functionally correlated. While past successes of such analyses have often been reported in a number of microarray studies (most of which used the standard hierarchical clustering, UPGMA, with one minus the Pearson's correlation coefficient as a measure of dissimilarity), often times such groupings could be misleading. More importantly, a systematic evaluation of the entire set of clusters produced by such unsupervised procedures is necessary since they also contain genes that are seemingly unrelated or may have more than one common function. Here we quantify the performance of a given unsupervised clustering algorithm applied to a given microarray study in terms of its ability to produce biologically meaningful clusters using a reference set of functional classes. Such a reference set may come from prior biological knowledge specific to a microarray study or may be formed using the growing databases of gene ontologies (GO) for the annotated genes of the relevant species. In this paper, we introduce two performance measures for evaluating the results of a clustering algorithm in its ability to produce biologically meaningful clusters. The first measure is a biological homogeneity index (BHI). As the name suggests, it is a measure of how biologically homogeneous the clusters are. This can be used to quantify the performance of a given clustering algorithm such as UPGMA in grouping genes for a particular data set and also for comparing the performance of a number of competing clustering algorithms applied to the same data set. The second performance measure is called a biological stability index (BSI). For a given clustering algorithm and an expression data set, it measures the consistency of the clustering algorithm's ability to produce biologically meaningful clusters when applied repeatedly to similar data sets. A good clustering algorithm should have high BHI and moderate to high BSI. We evaluated the performance of ten well known clustering algorithms on two gene expression data sets and identified the optimal algorithm in each case. The first data set deals with SAGE profiles of differentially expressed tags between normal and ductal carcinoma in situ samples of breast cancer patients. The second data set contains the expression profiles over time of positively expressed genes (ORF's) during sporulation of budding yeast. Two separate choices of the functional classes were used for this data set and the results were compared for consistency. Functional information of annotated genes available from various GO databases mined using ontology tools can be used to systematically judge the results of an unsupervised clustering algorithm as applied to a gene expression data set in clustering genes. This information could be used to select the right algorithm from a class of clustering algorithms for the given data set.

MeSH Terms
Algorithms Artificial Intelligence Benchmarking/methods Cluster Analysis Databases, Protein Gene Expression Profiling/methods,standards Information Storage and Retrieval/methods Multigene Family/genetics Oligonucleotide Array Sequence Analysis/methods,standards Pattern Recognition, Automated/methods Reference Values
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Datta Susmita
Department of Bioinformatics and Biostatistics, University of Louisville, Louisville, KY 40202, USA. [email protected]
Datta Somnath
References (19)
19 references, click to expand
  1. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  2. Computational cluster validation in post-genomic data analysis.
    Bioinformatics. 2005 Aug 1;21(15):3201-12 PMID: 15914541
  3. Computational analysis of microarray data.
    Nat Rev Genet. 2001 Jun;2(6):418-27 PMID: 11389458
  4. Bootstrapping cluster analysis: assessing the reliability of conclusions from microarray experiments.
    Proc Natl Acad Sci U S A. 2001 Jul 31;98(16):8961-5 PMID: 11470909
  5. A prediction-based resampling method for estimating the number of clusters in a dataset.
    Genome Biol. 2002 Jun 25;3(7):RESEARCH0036 PMID: 12184810
  6. Judging the quality of gene expression-based clustering methods using gene annotation.
    Genome Res. 2002 Oct;12(10):1574-81 PMID: 12368250
  7. Comparisons and validation of statistical clustering techniques for microarray gene expression data.
    Bioinformatics. 2003 Mar 1;19(4):459-66 PMID: 12611800
  8. TM4: a free, open-source system for microarray data management and analysis.
    Biotechniques. 2003 Feb;34(2):374-8 PMID: 12613259
  9. FunSpec: a web-based cluster interpreter for yeast.
    BMC Bioinformatics. 2002 Nov 13;3:35 PMID: 12431279
  10. Scoring clustering solutions by their biological relevance.
    Bioinformatics. 2003 Dec 12;19(18):2381-9 PMID: 14668221
  11. A graph-theoretic modeling on GO space for biological interpretation of gene clusters.
    Bioinformatics. 2004 Feb 12;20(3):381-8 PMID: 14960465
  12. FatiGO: a web tool for finding significant associations of Gene Ontology terms with groups of genes.
    Bioinformatics. 2004 Mar 1;20(4):578-80 PMID: 14990455
  13. Selection of informative clusters from hierarchical cluster tree with gene classes.
    BMC Bioinformatics. 2004 Mar 25;5:32 PMID: 15043761
  14. Transcriptomic changes in human breast cancer progression as determined by serial analysis of gene expression.
    Breast Cancer Res. 2004;6(5):R499-513 PMID: 15318932
  15. The FunCat, a functional annotation scheme for systematic classification of proteins from whole genomes.
    Nucleic Acids Res. 2004;32(18):5539-45 PMID: 15486203
  16. Phylogenetic reconstruction using an unsupervised growing neural network that adopts the topology of a phylogenetic tree.
    J Mol Evol. 1997 Feb;44(2):226-33 PMID: 9069183
  17. The transcriptional program of sporulation in budding yeast.
    Science. 1998 Oct 23;282(5389):699-705 PMID: 9784122
  18. A knowledge-driven approach to cluster validity assessment.
    Bioinformatics. 2005 May 15;21(10):2546-7 PMID: 15713738
  19. Validating clustering for gene expression data.
    Bioinformatics. 2001 Apr;17(4):309-18 PMID: 11301299
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2006-08-31
Epub
2006-00-31
Pages
397
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1590054
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]