Home LiteratureArticle Details
PMID: 20460310 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Independent filtering increases detection power for high-throughput experiments.

Bourgon R, Gentleman R, Huber W

Abstract

With high-dimensional data, variable-by-variable statistical testing is often used to select variables whose behavior differs across conditions. Such an approach requires adjustment for multiple testing, which can result in low statistical power. A two-stage approach that first filters variables by a criterion independent of the test statistic, and then only tests variables which pass the filter, can provide higher power. We show that use of some filter/test statistics pairs presented in the literature may, however, lead to loss of type I error control. We describe other pairs which avoid this problem. In an application to microarray data, we found that gene-by-gene filtering by overall variance followed by a t-test increased the number of discoveries by 50%. We also show that this particular statistic pair induces a lower bound on fold-change among the set of discoveries. Independent filtering-using filter/test pairs that are independent under the null hypothesis but correlated under the alternative-is a general approach that can substantially increase the efficiency of experiments.

MeSH Terms
Algorithms Biometry/methods Computational Biology Models, Genetic
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Bourgon Richard
European Bioinformatics Institute, Cambridge CB10 1SD, UK.
Gentleman Robert
Huber Wolfgang
References (17)
17 references, click to expand
  1. Exploration, normalization, and summaries of high density oligonucleotide array probe level data.
    Biostatistics. 2003 Apr;4(2):249-64 PMID: 12925520
  2. A method to increase the power of multiple testing procedures through sample splitting.
    Stat Appl Genet Mol Biol. 2006;5:Article19 PMID: 17049030
  3. Genome-Wide Significance Levels and Weighted Hypothesis Testing.
    Stat Sci. 2009 Nov;24(4):398-413 PMID: 20711421
  4. HIGH DIMENSIONAL VARIABLE SELECTION.
    Ann Stat. 2009 Jan 1;37(5A):2178-2201 PMID: 19784398
  5. Gene expression profile of adult T-cell acute lymphocytic leukemia identifies distinct subsets of patients with different response to therapy and survival.
    Blood. 2004 Apr 1;103(7):2771-8 PMID: 14684422
  6. Filtering for increased power for microarray data analysis.
    BMC Bioinformatics. 2009 Jan 08;10:11 PMID: 19133141
  7. Significance analysis of microarrays applied to the ionizing radiation response.
    Proc Natl Acad Sci U S A. 2001 Apr 24;98(9):5116-21 PMID: 11309499
  8. Linear models and empirical bayes methods for assessing differential expression in microarray experiments.
    Stat Appl Genet Mol Biol. 2004;3:Article3 PMID: 16646809
  9. A class comparison method with filtering-enhanced variable selection for high-dimensional data sets.
    Stat Med. 2008 Dec 10;27(28):5834-49 PMID: 18781559
  10. Filtering genes for cluster and network analysis.
    BMC Bioinformatics. 2009 Jun 23;10:193 PMID: 19549335
  11. Effects of filtering by Present call on analysis of microarray experiments.
    BMC Bioinformatics. 2006 Jan 31;7:49 PMID: 16448562
  12. Moderated statistical tests for assessing differences in tag abundance.
    Bioinformatics. 2007 Nov 1;23(21):2881-7 PMID: 17881408
  13. Gene expression profiles of B-lineage adult acute lymphocytic leukemia reveal genetic patterns that identify lineage derivation and distinct mechanisms of transformation.
    Clin Cancer Res. 2005 Oct 15;11(20):7209-19 PMID: 16243790
  14. Bioconductor: open software development for computational biology and bioinformatics.
    Genome Biol. 2004;5(10):R80 PMID: 15461798
  15. Discussion of "Sure Independence Screening for Ultra-High Dimensional Feature Space.
    J R Stat Soc Series B Stat Methodol. 2008;70(5):903 PMID: 19603084
  16. Analysis of variance for gene expression microarray data.
    J Comput Biol. 2000;7(6):819-37 PMID: 11382364
  17. I/NI-calls for the exclusion of non-informative genes: a highly effective filtering tool for microarray data.
    Bioinformatics. 2007 Nov 1;23(21):2897-902 PMID: 17921172
Article Info
Journal
Proceedings of the National Academy of Sciences of the United States of America
Abbr.
Proc Natl Acad Sci U S A
ISSN
1091-6490
Published
2010-05-25
Epub
2010-00-11
Pages
9546-51
Language
English
Region
United States
NLM ID
7505876
PMCID
PMC2906865
Subset
IM
Corrections
CommentIn
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]