Home LiteratureArticle Details
PMID: 21572976 Published · ppublish English Journal Article

A Selective Overview of Variable Selection in High Dimensional Feature Space.

Statistica Sinica ·Vol. 20 ·No. 1 ·2010-01-00 ·Pages 101-148

Fan J, Lv J

Abstract

High dimensional statistical problems arise from diverse fields of scientific research and technological development. Variable selection plays a pivotal role in contemporary statistical learning and scientific discoveries. The traditional idea of best subset selection methods, which can be regarded as a specific form of penalized likelihood, is computationally too expensive for many modern statistical applications. Other forms of penalized likelihood methods have been successfully developed over the last decade to cope with high dimensionality. They have been widely applied for simultaneously selecting important variables and estimating their effects in high dimensional statistical inference. In this article, we present a brief account of the recent developments of theory, methods, and implementations for high dimensional variable selection. What limits of the dimensionality such methods can handle, what the role of penalty functions is, and what the statistical properties are rapidly drive the advances of the field. The properties of non-concave penalized likelihood and its roles in high dimensional statistical modeling are emphasized. We also review some recent advances in ultra-high dimensional variable selection, with emphasis on independence screening and two-scale methods.

Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Fan Jianqing
Frederick L. Moore '18 Professor of Finance, Department of Operations Research and Financial Engineering, Princeton University, Princeton, NJ 08544, USA ( [email protected] ).
Lv Jinchi
References (20)
20 references, click to expand
  1. Penalized Composite Quasi-Likelihood for Ultrahigh-Dimensional Variable Selection.
    J R Stat Soc Series B Stat Methodol. 2011 Jun;73(3):325-349 PMID: 21589849
  2. One-step Sparse Estimates in Nonconcave Penalized Likelihood Models.
    Ann Stat. 2008 Aug 1;36(4):1509-1533 PMID: 19823597
  3. An Equivalence Between Sparse Approximation and Support Vector Machines.
    Neural Comput. 1998 Jul 28;10(6):1455-80 PMID: 9698353
  4. Impossibility of successful classification when useful features are rare and weak.
    Proc Natl Acad Sci U S A. 2009 Jun 2;106(22):8859-64 PMID: 19447927
  5. Higher criticism thresholding: Optimal feature selection when useful features are rare and weak.
    Proc Natl Acad Sci U S A. 2008 Sep 30;105(39):14790-5 PMID: 18815365
  6. Sparse inverse covariance estimation with the graphical lasso.
    Biostatistics. 2008 Jul;9(3):432-41 PMID: 18079126
  7. Effective dimension reduction methods for tumor classification using gene expression data.
    Bioinformatics. 2003 Mar 22;19(5):563-70 PMID: 12651713
  8. Diagnosis of multiple cancer types by shrunken centroids of gene expression.
    Proc Natl Acad Sci U S A. 2002 May 14;99(10):6567-72 PMID: 12011421
  9. Dimension reduction strategies for analyzing global gene expression data with a response.
    Math Biosci. 2002 Mar;176(1):123-44 PMID: 11867087
  10. Tuning parameter selectors for the smoothly clipped absolute deviation method.
    Biometrika. 2007 Aug 1;94(3):553-568 PMID: 19343105
  11. Singular value decomposition regression models for classification of tumors from microarray experiments.
    Pac Symp Biocomput. 2002;:18-29 PMID: 11928474
  12. Non-Concave Penalized Likelihood with NP-Dimensionality.
    IEEE Trans Inf Theory. 2011 Aug;57(8):5467-5484 PMID: 22287795
  13. Statistical significance for genomewide studies.
    Proc Natl Acad Sci U S A. 2003 Aug 5;100(16):9440-5 PMID: 12883005
  14. Linear regression and two-class classification with gene expression data.
    Bioinformatics. 2003 Nov 1;19(16):2072-8 PMID: 14594712
  15. Statistical analysis of DNA microarray data in cancer research.
    Clin Cancer Res. 2006 Aug 1;12(15):4469-73 PMID: 16899590
  16. Variable Selection using MM Algorithms.
    Ann Stat. 2005;33(4):1617-1642 PMID: 19458786
  17. Graphical methods for class prediction using dimension reduction techniques on DNA microarray data.
    Bioinformatics. 2003 Jul 1;19(10):1252-8 PMID: 12835269
  18. Optimally sparse representation in general (nonorthogonal) dictionaries via l minimization.
    Proc Natl Acad Sci U S A. 2003 Mar 4;100(5):2197-202 PMID: 16576749
  19. High Dimensional Classification Using Features Annealed Independence Rules.
    Ann Stat. 2008;36(6):2605-2637 PMID: 19169416
  20. PLS dimension reduction for classification with microarray data.
    Stat Appl Genet Mol Biol. 2004;3:Article33 PMID: 16646813
Article Info
Journal
Statistica Sinica
Abbr.
Stat Sin
ISSN
1017-0405
Published
2010-01-00
Pages
101-148
Language
English
Region
China (Republic : 1949- )
NLM ID
101473244
PMCID
PMC3092303
Grants
NIGMS NIH HHS · R01 GM072611 · United States
NIGMS NIH HHS · R01 GM072611-05 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]