Home LiteratureArticle Details
PMID: 17254353 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Bias in random forest variable importance measures: illustrations, sources and a solution.

BMC bioinformatics ·Vol. 8 ·2007-01-25 ·Pages 25

Strobl C, Boulesteix AL, Zeileis A, Hothorn T

Abstract

Variable importance measures for random forests have been receiving increased attention as a means of variable selection in many classification tasks in bioinformatics and related scientific fields, for instance to select a subset of genetic markers relevant for the prediction of a certain disease. We show that random forest variable importance measures are a sensible means for variable selection in many applications, but are not reliable in situations where potential predictor variables vary in their scale of measurement or their number of categories. This is particularly important in genomics and computational biology, where predictors often include variables of different types, for example when predictors include both sequence data and continuous variables such as folding energy, or when amino acid sequence data show different numbers of categories. Simulation studies are presented illustrating that, when random forest variable importance measures are used with data of varying types, the results are misleading because suboptimal predictor variables may be artificially preferred in variable selection. The two mechanisms underlying this deficiency are biased variable selection in the individual classification trees used to build the random forest on one hand, and effects induced by bootstrap sampling with replacement on the other hand. We propose to employ an alternative implementation of random forests, that provides unbiased variable selection in the individual classification trees. When this method is applied using subsampling without replacement, the resulting variable importance measures can be used reliably for variable selection even in situations where the potential predictor variables vary in their scale of measurement or their number of categories. The usage of both random forest algorithms and their variable importance measures in the R system for statistical computing is illustrated and documented thoroughly in an application re-analyzing data from a study on RNA editing. Therefore the suggested method can be applied straightforwardly by scientists in bioinformatics research.

MeSH Terms
Algorithms Bias Computational Biology/methods Computer Simulation Data Interpretation, Statistical Genomics/methods Models, Biological Models, Statistical Population Dynamics
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Strobl Carolin
Institut für Statistik, Ludwig-Maximilians-Universität München, Ludwigstr, 33, 80539 München, Germany. [email protected]
Boulesteix Anne-Laure
Zeileis Achim
Hothorn Torsten
References (16)
16 references, click to expand
  1. Screening large-scale association study data: exploiting interactions using random forests.
    BMC Genet. 2004;5:32 PMID: 15588316
  2. Evaluation of different biological data and computational classification methods for use in protein interaction prediction.
    Proteins. 2006 May 15;63(3):490-500 PMID: 16450363
  3. Maximally selected chi-square statistics for ordinal variables.
    Biom J. 2006 Jun;48(3):451-62 PMID: 16845908
  4. Relating HIV-1 sequence variation to replication capacity via trees and forests.
    Stat Appl Genet Mol Biol. 2004;3:Article2; discussion article 7, article 9 PMID: 16646798
  5. Simple statistical models predict C-to-U edited sites in plant mitochondrial RNA.
    BMC Bioinformatics. 2004 Sep 16;5:132 PMID: 15373947
  6. Few amino acid positions in rpoB are associated with most of the rifampin resistance in Mycobacterium tuberculosis.
    BMC Bioinformatics. 2004 Sep 28;5:137 PMID: 15453919
  7. Tumor classification by tissue microarray profiling: random forest clustering applied to renal cell carcinoma.
    Mod Pathol. 2005 Apr;18(4):547-57 PMID: 15529185
  8. Gene selection and classification of microarray data using random forest.
    BMC Bioinformatics. 2006;7:3 PMID: 16398926
  9. Identifying SNPs predictive of phenotype using random forests.
    Genet Epidemiol. 2005 Feb;28(2):171-82 PMID: 15593090
  10. Development of linear, ensemble, and nonlinear models for the prediction and interpretation of the biological activity of a set of PDGFR inhibitors.
    J Chem Inf Comput Sci. 2004 Nov-Dec;44(6):2179-89 PMID: 15554688
  11. Prediction of clinical drug efficacy by classification of drug-induced genomic expression profiles in vitro.
    Proc Natl Acad Sci U S A. 2003 Aug 5;100(16):9608-13 PMID: 12869696
  12. A comparative study of discriminating human heart failure etiology using gene expression profiles.
    BMC Bioinformatics. 2005;6:205 PMID: 16120216
  13. Random forest: a classification and regression tool for compound classification and QSAR modeling.
    J Chem Inf Comput Sci. 2003 Nov-Dec;43(6):1947-58 PMID: 14632445
  14. Short-term prediction of mortality in patients with systemic lupus erythematosus: classification of outcomes using random forests.
    Arthritis Rheum. 2006 Feb 15;55(1):74-80 PMID: 16463416
  15. Maximally selected chi-square statistics and binary splits of nominal variables.
    Biom J. 2006 Aug;48(5):838-48 PMID: 17094347
  16. The challenge for genetic epidemiologists: how to analyze large numbers of SNPs in relation to complex diseases.
    BMC Genet. 2006 Apr 21;7:23 PMID: 16630340
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2007-01-25
Epub
2007-00-25
Pages
25
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1796903
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]