Home LiteratureArticle Details
PMID: 22549815 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

Reprioritizing genetic associations in hit regions using LASSO-based resample model averaging.

Genetic epidemiology ·Vol. 36 ·No. 5 ·2012-07-00 ·Pages 451-62

Valdar W, Sabourin J, Nobel A, Holmes CC

Abstract

Significance testing one SNP at a time has proven useful for identifying genomic regions that harbor variants affecting human disease. But after an initial genome scan has identified a "hit region" of association, single-locus approaches can falter. Local linkage disequilibrium (LD) can make both the number of underlying true signals and their identities ambiguous. Simultaneous modeling of multiple loci should help. However, it is typically applied ad hoc: conditioning on the top SNPs, with limited exploration of the model space and no assessment of how sensitive model choice was to sampling variability. Formal alternatives exist but are seldom used. Bayesian variable selection is coherent but requires specifying a full joint model, including priors on parameters and the model space. Penalized regression methods (e.g., LASSO) appear promising but require calibration, and, once calibrated, lead to a choice of SNPs that can be misleadingly decisive. We present a general method for characterizing uncertainty in model choice that is tailored to reprioritizing SNPs within a hit region under strong LD. Our method, LASSO local automatic regularization resample model averaging (LLARRMA), combines LASSO shrinkage with resample model averaging and multiple imputation, estimating for each SNP the probability that it would be included in a multi-SNP model in alternative realizations of the data. We apply LLARRMA to simulations based on case-control genome-wide association studies data, and find that when there are several causal loci and strong LD, LLARRMA identifies a set of candidates that is enriched for true signals relative to single locus analysis and to the recently proposed method of Stability Selection.

MeSH Terms
Algorithms Bayes Theorem Calibration Case-Control Studies Chromosome Mapping Computer Simulation Genome-Wide Association Study/methods Genotype Humans Models, Genetic Models, Statistical Models, Theoretical Molecular Epidemiology/methods ROC Curve Regression Analysis
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Valdar William
Department of Genetics, and Lineberger Comprehensive Cancer Center, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina 27599-7265, USA. [email protected]
Sabourin Jeremy
Nobel Andrew
Holmes Christopher C
References (35)
35 references, click to expand
  1. Stability selection for genome-wide association.
    Genet Epidemiol. 2011 Nov;35(7):722-8 PMID: 22009793
  2. Prioritizing GWAS results: A review of statistical methods and recommendations for their application.
    Am J Hum Genet. 2010 Jan;86(1):6-22 PMID: 20074509
  3. Remapping the insulin gene/IDDM2 locus in type 1 diabetes.
    Diabetes. 2004 Jul;53(7):1884-9 PMID: 15220214
  4. Bayesian statistical methods for genetic association studies.
    Nat Rev Genet. 2009 Oct;10(10):681-90 PMID: 19763151
  5. SNP selection in genome-wide and candidate gene studies via penalized logistic regression.
    Genet Epidemiol. 2010 Dec;34(8):879-91 PMID: 21104890
  6. Discussion of "Sure Independence Screening for Ultra-High Dimensional Feature Space.
    J R Stat Soc Series B Stat Methodol. 2008;70(5):903 PMID: 19603084
  7. A unified stepwise regression procedure for evaluating the relative effects of polymorphisms within a gene using case/control or family data: application to HLA in type 1 diabetes.
    Am J Hum Genet. 2002 Jan;70(1):124-41 PMID: 11719900
  8. Mining gold dust under the genome wide significance level: a two-stage approach to analysis of GWAS.
    Genet Epidemiol. 2011 Feb;35(2):111-8 PMID: 21254218
  9. PLINK: a tool set for whole-genome association and population-based linkage analyses.
    Am J Hum Genet. 2007 Sep;81(3):559-75 PMID: 17701901
  10. BAYESIAN MODEL SEARCH AND MULTILEVEL INFERENCE FOR SNP ASSOCIATION STUDIES.
    Ann Appl Stat. 2010 Sep 1;4(3):1342-1364 PMID: 21179394
  11. Selecting SNPs in two-stage analysis of disease association data: a model-free approach.
    Ann Hum Genet. 2000 Sep;64(Pt 5):413-7 PMID: 11281279
  12. Association screening of common and rare genetic variants by penalized regression.
    Bioinformatics. 2010 Oct 1;26(19):2375-82 PMID: 20693321
  13. Imputation-based analysis of association studies: candidate regions and quantitative traits.
    PLoS Genet. 2007 Jul;3(7):e114 PMID: 17676998
  14. Regularization Paths for Generalized Linear Models via Coordinate Descent.
    J Stat Softw. 2010;33(1):1-22 PMID: 20808728
  15. Genome-wide genetic association of complex traits in heterogeneous stock mice.
    Nat Genet. 2006 Aug;38(8):879-87 PMID: 16832355
  16. Association of the T-cell regulatory gene CTLA4 with susceptibility to autoimmune disease.
    Nature. 2003 May 29;423(6939):506-11 PMID: 12724780
  17. A novel bayesian graphical model for genome-wide multi-SNP association mapping.
    Genet Epidemiol. 2012 Jan;36(1):36-47 PMID: 22127647
  18. Predicting unobserved phenotypes for complex traits from whole-genome SNP data.
    PLoS Genet. 2008 Oct;4(10):e1000231 PMID: 18949033
  19. New approaches to population stratification in genome-wide association studies.
    Nat Rev Genet. 2010 Jul;11(7):459-63 PMID: 20548291
  20. Identification of association between disease and multiple markers via sparse partial least-squares regression.
    Genet Epidemiol. 2011 Sep;35(6):479-86 PMID: 21678491
  21. Automated variable selection methods for logistic regression produced unstable models for predicting acute myocardial infarction mortality.
    J Clin Epidemiol. 2004 Nov;57(11):1138-46 PMID: 15567629
  22. A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.
    PLoS Genet. 2009 Jun;5(6):e1000529 PMID: 19543373
  23. MaCH: using sequence and genotype data to estimate haplotypes and unobserved genotypes.
    Genet Epidemiol. 2010 Dec;34(8):816-34 PMID: 21058334
  24. A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase.
    Am J Hum Genet. 2006 Apr;78(4):629-44 PMID: 16532393
  25. Simultaneous analysis of all SNPs in genome-wide and re-sequencing association studies.
    PLoS Genet. 2008 Jul 25;4(7):e1000130 PMID: 18654633
  26. Multilocus association testing with penalized regression.
    Genet Epidemiol. 2011 Dec;35(8):755-65 PMID: 21922539
  27. A genome-wide association study identifies new psoriasis susceptibility loci and an interaction between HLA-C and ERAP1.
    Nat Genet. 2010 Nov;42(11):985-90 PMID: 20953190
  28. A comparison of approaches to account for uncertainty in analysis of imputed genotypes.
    Genet Epidemiol. 2011 Feb;35(2):102-10 PMID: 21254217
  29. A tutorial on statistical methods for population association studies.
    Nat Rev Genet. 2006 Oct;7(10):781-91 PMID: 16983374
  30. Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls.
    Nature. 2007 Jun 7;447(7145):661-78 PMID: 17554300
  31. Mapping in structured populations by resample model averaging.
    Genetics. 2009 Aug;182(4):1263-77 PMID: 19474203
  32. Inference of population structure using multilocus genotype data.
    Genetics. 2000 Jun;155(2):945-59 PMID: 10835412
  33. Genome-wide association analysis by lasso penalized logistic regression.
    Bioinformatics. 2009 Mar 15;25(6):714-21 PMID: 19176549
  34. Comparison of statistical tests for disease association with rare variants.
    Genet Epidemiol. 2011 Nov;35(7):606-19 PMID: 21769936
  35. Finding the missing heritability of complex diseases.
    Nature. 2009 Oct 8;461(7265):747-53 PMID: 19812666
Article Info
Journal
Genetic epidemiology
Abbr.
Genet Epidemiol
ISSN
1098-2272
Published
2012-07-00
Epub
2012-00-30
Pages
451-62
Language
English
Region
United States
NLM ID
8411723
PMCID
PMC3470705
Subset
IM
Grants
NIMH NIH HHS · R01 MH090936 · United States
NCI NIH HHS · P50 CA058223 · United States
Medical Research Council · G0701612 · United Kingdom
NIMH NIH HHS · MH090936-02 · United States
Medical Research Council · MC_UP_A390_1107 · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]