Abstract
We introduce a new Empirical Bayes approach for large-scale hypothesis testing, including estimating false discovery rates (FDRs), and effect sizes. This approach has two key differences from existing approaches to FDR analysis. First, it assumes that the distribution of the actual (unobserved) effects is unimodal, with a mode at 0. This "unimodal assumption" (UA), although natural in many contexts, is not usually incorporated into standard FDR analysis, and we demonstrate how incorporating it brings many benefits. Specifically, the UA facilitates efficient and robust computation-estimating the unimodal distribution involves solving a simple convex optimization problem-and enables more accurate inferences provided that it holds. Second, the method takes as its input two numbers for each test (an effect size estimate and corresponding standard error), rather than the one number usually used ($p$ value or $z$ score). When available, using two numbers instead of one helps account for variation in measurement precision across tests. It also facilitates estimation of effects, and unlike standard FDR methods, our approach provides interval estimates (credible regions) for each effect in addition to measures of significance. To provide a bridge between interval estimates and significance measures, we introduce the term "local false sign rate" to refer to the probability of getting the sign of an effect wrong and argue that it is a superior measure of significance than the local FDR because it is both more generally applicable and can be more robustly estimated. Our methods are implemented in an R package ashr available from http://github.com/stephens999/ashr.
Keywords
Empirical Bayes
False discovery rates
Multiple testing
Shrinkage
Unimodal
MeSH Terms
Bayes Theorem
Data Interpretation, Statistical
Humans
Models, Statistical
Authors & Affiliations
1 authors, click to expand affiliations / ORCID
Stephens Matthew
References (18)
18 references, click to expand
-
Genomic control for association studies.
Biometrics. 1999 Dec;55(4):997-1004
PMID: 11315092
-
Principal components analysis corrects for stratification in genome-wide association studies.
Nat Genet. 2006 Aug;38(8):904-9
PMID: 16862161
-
Variance adaptive shrinkage (vash): flexible empirical Bayes estimation of variances.
Bioinformatics. 2016 Nov 15;32(22):3428-3434
PMID: 27436563
-
Association mapping in structured populations.
Am J Hum Genet. 2000 Jul;67(1):170-81
PMID: 10827107
-
Using control genes to correct for unwanted variation in microarray data.
Biostatistics. 2012 Jul;13(3):539-52
PMID: 22101192
-
Multidimensional local false discovery rate for microarray studies.
Bioinformatics. 2006 Mar 1;22(5):556-65
PMID: 16368770
-
Empirical Bayes screening of many p-values with applications to microarray studies.
Bioinformatics. 2005 May 1;21(9):1987-94
PMID: 15691856
-
Improving accuracy of genomic predictions within and between dairy cattle breeds with imputed high-density single nucleotide polymorphism panels.
J Dairy Sci. 2012 Jul;95(7):4114-29
PMID: 22720968
-
On parametric empirical Bayes methods for comparing multiple groups using replicated gene expression profiles.
Stat Med. 2003 Dec 30;22(24):3899-914
PMID: 14673946
-
Empirical-Bayes adjustments for multiple comparisons are sometimes useful.
Epidemiology. 1991 Jul;2(4):244-51
PMID: 1912039
-
Linear models and empirical bayes methods for assessing differential expression in microarray experiments.
Stat Appl Genet Mol Biol. 2004;3:Article3
PMID: 16646809
-
Empirical bayes methods and false discovery rates for microarrays.
Genet Epidemiol. 2002 Jun;23(1):70-86
PMID: 12112249
-
Detecting differential gene expression with a semiparametric hierarchical mixture method.
Biostatistics. 2004 Apr;5(2):155-76
PMID: 15054023
-
Capturing heterogeneity in gene expression studies by surrogate variable analysis.
PLoS Genet. 2007 Sep;3(9):1724-35
PMID: 17907809
-
Practical issues in imputation-based association mapping.
PLoS Genet. 2008 Dec;4(12):e1000279
PMID: 19057666
-
SURE Estimates for a Heteroscedastic Hierarchical Model.
J Am Stat Assoc. 2012 Dec;107(500):1465-1479
PMID: 25301976
-
Bayesian Semiparametric Density Deconvolution in the Presence of Conditionally Heteroscedastic Measurement Errors.
J Comput Graph Stat. 2014 Oct 1;23(4):1101-1125
PMID: 25378893
-
Bayes factors for genome-wide association studies: comparison with P-values.
Genet Epidemiol. 2009 Jan;33(1):79-86
PMID: 18642345