Abstract
Variable importance measures for random forests have been receiving increased attention as a means of variable selection in many classification tasks in bioinformatics and related scientific fields, for instance to select a subset of genetic markers relevant for the prediction of a certain disease. We show that random forest variable importance measures are a sensible means for variable selection in many applications, but are not reliable in situations where potential predictor variables vary in their scale of measurement or their number of categories. This is particularly important in genomics and computational biology, where predictors often include variables of different types, for example when predictors include both sequence data and continuous variables such as folding energy, or when amino acid sequence data show different numbers of categories. Simulation studies are presented illustrating that, when random forest variable importance measures are used with data of varying types, the results are misleading because suboptimal predictor variables may be artificially preferred in variable selection. The two mechanisms underlying this deficiency are biased variable selection in the individual classification trees used to build the random forest on one hand, and effects induced by bootstrap sampling with replacement on the other hand. We propose to employ an alternative implementation of random forests, that provides unbiased variable selection in the individual classification trees. When this method is applied using subsampling without replacement, the resulting variable importance measures can be used reliably for variable selection even in situations where the potential predictor variables vary in their scale of measurement or their number of categories. The usage of both random forest algorithms and their variable importance measures in the R system for statistical computing is illustrated and documented thoroughly in an application re-analyzing data from a study on RNA editing. Therefore the suggested method can be applied straightforwardly by scientists in bioinformatics research.
MeSH Terms
Algorithms
Bias
Computational Biology/methods
Computer Simulation
Data Interpretation, Statistical
Genomics/methods
Models, Biological
Models, Statistical
Population Dynamics
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Strobl Carolin
Institut für Statistik, Ludwig-Maximilians-Universität München, Ludwigstr, 33, 80539 München, Germany.
[email protected]
Boulesteix Anne-Laure
Zeileis Achim
Hothorn Torsten
References (16)
16 references, click to expand
-
Screening large-scale association study data: exploiting interactions using random forests.
BMC Genet. 2004;5:32
PMID: 15588316
-
Evaluation of different biological data and computational classification methods for use in protein interaction prediction.
Proteins. 2006 May 15;63(3):490-500
PMID: 16450363
-
Maximally selected chi-square statistics for ordinal variables.
Biom J. 2006 Jun;48(3):451-62
PMID: 16845908
-
Relating HIV-1 sequence variation to replication capacity via trees and forests.
Stat Appl Genet Mol Biol. 2004;3:Article2; discussion article 7, article 9
PMID: 16646798
-
Simple statistical models predict C-to-U edited sites in plant mitochondrial RNA.
BMC Bioinformatics. 2004 Sep 16;5:132
PMID: 15373947
-
Few amino acid positions in rpoB are associated with most of the rifampin resistance in Mycobacterium tuberculosis.
BMC Bioinformatics. 2004 Sep 28;5:137
PMID: 15453919
-
Tumor classification by tissue microarray profiling: random forest clustering applied to renal cell carcinoma.
Mod Pathol. 2005 Apr;18(4):547-57
PMID: 15529185
-
Gene selection and classification of microarray data using random forest.
BMC Bioinformatics. 2006;7:3
PMID: 16398926
-
Identifying SNPs predictive of phenotype using random forests.
Genet Epidemiol. 2005 Feb;28(2):171-82
PMID: 15593090
-
Development of linear, ensemble, and nonlinear models for the prediction and interpretation of the biological activity of a set of PDGFR inhibitors.
J Chem Inf Comput Sci. 2004 Nov-Dec;44(6):2179-89
PMID: 15554688
-
Prediction of clinical drug efficacy by classification of drug-induced genomic expression profiles in vitro.
Proc Natl Acad Sci U S A. 2003 Aug 5;100(16):9608-13
PMID: 12869696
-
A comparative study of discriminating human heart failure etiology using gene expression profiles.
BMC Bioinformatics. 2005;6:205
PMID: 16120216
-
Random forest: a classification and regression tool for compound classification and QSAR modeling.
J Chem Inf Comput Sci. 2003 Nov-Dec;43(6):1947-58
PMID: 14632445
-
Short-term prediction of mortality in patients with systemic lupus erythematosus: classification of outcomes using random forests.
Arthritis Rheum. 2006 Feb 15;55(1):74-80
PMID: 16463416
-
Maximally selected chi-square statistics and binary splits of nominal variables.
Biom J. 2006 Aug;48(5):838-48
PMID: 17094347
-
The challenge for genetic epidemiologists: how to analyze large numbers of SNPs in relation to complex diseases.
BMC Genet. 2006 Apr 21;7:23
PMID: 16630340