Abstract
Numerous feature selection methods have been applied to the identification of differentially expressed genes in microarray data. These include simple fold change, classical t-statistic and moderated t-statistics. Even though these methods return gene lists that are often dissimilar, few direct comparisons of these exist. We present an empirical study in which we compare some of the most commonly used feature selection methods. We apply these to 9 publicly available datasets, and compare, both the gene lists produced and how these perform in class prediction of test datasets. In this study, we compared the efficiency of the feature selection methods; significance analysis of microarrays (SAM), analysis of variance (ANOVA), empirical bayes t-statistic, template matching, maxT, between group analysis (BGA), Area under the receiver operating characteristic (ROC) curve, the Welch t-statistic, fold change, rank products, and sets of randomly selected genes. In each case these methods were applied to 9 different binary (two class) microarray datasets. Firstly we found little agreement in gene lists produced by the different methods. Only 8 to 21% of genes were in common across all 10 feature selection methods. Secondly, we evaluated the class prediction efficiency of each gene list in training and test cross-validation using four supervised classifiers. We report that the choice of feature selection method, the number of genes in the genelist, the number of cases (samples) and the noise in the dataset, substantially influence classification success. Recommendations are made for choice of feature selection. Area under a ROC curve performed well with datasets that had low levels of noise and large sample size. Rank products performs well when datasets had low numbers of samples or high levels of noise. The Empirical bayes t-statistic performed well across a range of sample sizes.
MeSH Terms
Algorithms
Cell Line, Tumor
Computational Biology/methods
Data Interpretation, Statistical
Databases, Genetic
Gene Expression Profiling/methods
Gene Expression Regulation, Neoplastic
Humans
Models, Statistical
Oligonucleotide Array Sequence Analysis
Pattern Recognition, Automated
ROC Curve
Sample Size
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Jeffery Ian B
Bioinformatics, Conway Institute, University College Dublin, Belfield, Dublin 4, Ireland.
[email protected]
Higgins Desmond G
Culhane Aedín C
References (29)
29 references, click to expand
-
Differential gene expression detection using penalized linear regression models: the improved SAM statistics.
Bioinformatics. 2005 Apr 15;21(8):1565-71
PMID: 15598833
-
Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays.
Proc Natl Acad Sci U S A. 1999 Jun 8;96(12):6745-50
PMID: 10359783
-
Gene expression profile of adult T-cell acute lymphocytic leukemia identifies distinct subsets of patients with different response to therapy and survival.
Blood. 2004 Apr 1;103(7):2771-8
PMID: 14684422
-
Selecting differentially expressed genes from microarray experiments.
Biometrics. 2003 Mar;59(1):133-42
PMID: 12762450
-
Gene expression correlates of clinical prostate cancer behavior.
Cancer Cell. 2002 Mar;1(2):203-9
PMID: 12086878
-
Analysis of strain and regional variation in gene expression in mouse brain.
Genome Biol. 2001;2(10):RESEARCH0042
PMID: 11597334
-
Rank Difference Analysis of Microarrays (RDAM), a novel approach to statistical analysis of microarray expression profiling data.
BMC Bioinformatics. 2004 Oct 11;5:148
PMID: 15476558
-
Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning.
Nat Med. 2002 Jan;8(1):68-74
PMID: 11786909
-
Preferred analysis methods for Affymetrix GeneChips revealed by a wholly defined control dataset.
Genome Biol. 2005;6(2):R16
PMID: 15693945
-
Between-group analysis of microarray data.
Bioinformatics. 2002 Dec;18(12):1600-8
PMID: 12490444
-
Support vector machine classification and validation of cancer tissue samples using microarray expression data.
Bioinformatics. 2000 Oct;16(10):906-14
PMID: 11120680
-
Microarray-based gene expression profiling of hematologic malignancies: basic concepts and clinical applications.
Blood Rev. 2005 Jul;19(4):223-34
PMID: 15784300
-
ROC curves are a suitable and flexible tool for the analysis of gene expression profiles.
Cytogenet Genome Res. 2003;101(1):90-1
PMID: 14571143
-
Bioconductor: open software development for computational biology and bioinformatics.
Genome Biol. 2004;5(10):R80
PMID: 15461798
-
Rank products: a simple, yet powerful, new method to detect differentially regulated genes in replicated microarray experiments.
FEBS Lett. 2004 Aug 27;573(1-3):83-92
PMID: 15327980
-
A comparative review of statistical methods for discovering differentially expressed genes in replicated microarray experiments.
Bioinformatics. 2002 Apr;18(4):546-54
PMID: 12016052
-
Linear models and empirical bayes methods for assessing differential expression in microarray experiments.
Stat Appl Genet Mol Biol. 2004;3:Article3
PMID: 16646809
-
Improved statistical inference from DNA microarray data using analysis of variance and a Bayesian statistical framework. Analysis of global gene expression in Escherichia coli K12.
J Biol Chem. 2001 Jun 8;276(23):19937-44
PMID: 11259426
-
A comprehensive evaluation of multicategory classification methods for microarray gene expression cancer diagnosis.
Bioinformatics. 2005 Mar 1;21(5):631-43
PMID: 15374862
-
Significance analysis of ROC indices for comparing diagnostic markers: applications to gene microarray data.
J Biopharm Stat. 2004 Nov;14(4):985-1003
PMID: 15587976
-
Prediction of central nervous system embryonal tumour outcome based on gene expression.
Nature. 2002 Jan 24;415(6870):436-42
PMID: 11807556
-
The role of the Wnt-signaling antagonist DKK1 in the development of osteolytic lesions in multiple myeloma.
N Engl J Med. 2003 Dec 25;349(26):2483-94
PMID: 14695408
-
Rank-based methods as a non-parametric alternative of the T-statistic for the analysis of biological microarray data.
J Bioinform Comput Biol. 2005 Oct;3(5):1171-89
PMID: 16278953
-
The limit fold change model: a practical approach for selecting differentially expressed genes from microarray data.
BMC Bioinformatics. 2002 Jun 21;3:17
PMID: 12095422
-
Significance analysis of microarrays applied to the ionizing radiation response.
Proc Natl Acad Sci U S A. 2001 Apr 24;98(9):5116-21
PMID: 11309499
-
Gene expression-based classification of malignant gliomas correlates better with survival than histological classification.
Cancer Res. 2003 Apr 1;63(7):1602-7
PMID: 12670911
-
MADE4: an R package for multivariate analysis of gene expression data.
Bioinformatics. 2005 Jun 1;21(11):2789-90
PMID: 15797915
-
Molecular classification of cancer: class discovery and class prediction by gene expression monitoring.
Science. 1999 Oct 15;286(5439):531-7
PMID: 10521349
-
Exploration, normalization, and summaries of high density oligonucleotide array probe level data.
Biostatistics. 2003 Apr;4(2):249-64
PMID: 12925520