Home LiteratureArticle Details
PMID: 25357204 Published · epublish English Journal Article Research Support, N.I.H., Extramural

Integrating functional data to prioritize causal variants in statistical fine-mapping studies.

PLoS genetics ·Vol. 10 ·No. 10 ·2014-10-00 ·Pages e1004722

Kichaev G, Yang WY, Lindstrom S, Hormozdiari F, Eskin E, Price AL, Kraft P, Pasaniuc B

Abstract

Standard statistical approaches for prioritization of variants for functional testing in fine-mapping studies either use marginal association statistics or estimate posterior probabilities for variants to be causal under simplifying assumptions. Here, we present a probabilistic framework that integrates association strength with functional genomic annotation data to improve accuracy in selecting plausible causal variants for functional validation. A key feature of our approach is that it empirically estimates the contribution of each functional annotation to the trait of interest directly from summary association statistics while allowing for multiple causal variants at any risk locus. We devise efficient algorithms that estimate the parameters of our model across all risk loci to further increase performance. Using simulations starting from the 1000 Genomes data, we find that our framework consistently outperforms the current state-of-the-art fine-mapping methods, reducing the number of variants that need to be selected to capture 90% of the causal variants from an average of 13.3 to 10.4 SNPs per locus (as compared to the next-best performing strategy). Furthermore, we introduce a cost-to-benefit optimization framework for determining the number of variants to be followed up in functional assays and assess its performance using real and simulation data. We validate our findings using a large scale meta-analysis of four blood lipids traits and find that the relative probability for causality is increased for variants in exons and transcription start sites and decreased in repressed genomic regions at the risk loci of these traits. Using these highly predictive, trait-specific functional annotations, we estimate causality probabilities across all traits and variants, reducing the size of the 90% confidence set from an average of 17.5 to 13.5 variants per locus in this data.

MeSH Terms
Algorithms Chromosome Mapping/methods Genome-Wide Association Study/methods Humans Linkage Disequilibrium Models, Theoretical Polymorphism, Single Nucleotide/genetics
Authors & Affiliations
8 authors, click to expand affiliations / ORCID
Kichaev Gleb
Bioinformatics Interdepartmental Program, University of California Los Angeles, Los Angeles, California, United States of America.
Yang Wen-Yun
Department of Computer Science, University of California Los Angeles, Los Angeles, California, United States of America.
Lindstrom Sara
Program in Genetic Epidemiology and Statistical Genetics, Harvard School of Public Health, Boston, Massachusetts, United States of America.
Hormozdiari Farhad
Department of Computer Science, University of California Los Angeles, Los Angeles, California, United States of America.
Eskin Eleazar
Bioinformatics Interdepartmental Program, University of California Los Angeles, Los Angeles, California, United States of America; Department of Computer Science, University of California Los Angeles, Los Angeles, California, United States of America; Department of Human Genetics, David Geffen School of Medicine, University of California Los Angeles, Los Angeles, California, United States of America.
Price Alkes L
Program in Genetic Epidemiology and Statistical Genetics, Harvard School of Public Health, Boston, Massachusetts, United States of America; Department of Biostatistics, Harvard School of Public Health, Boston, Massachusetts, United States of America.
Kraft Peter
Program in Genetic Epidemiology and Statistical Genetics, Harvard School of Public Health, Boston, Massachusetts, United States of America; Department of Biostatistics, Harvard School of Public Health, Boston, Massachusetts, United States of America.
Pasaniuc Bogdan
Bioinformatics Interdepartmental Program, University of California Los Angeles, Los Angeles, California, United States of America; Department of Human Genetics, David Geffen School of Medicine, University of California Los Angeles, Los Angeles, California, United States of America; Department of Pathology and Laboratory Medicine, David Geffen School of Medicine, University of California Los Angeles, Los Angeles, California, United States of America.
References (39)
39 references, click to expand
  1. Linkage disequilibrium patterns of the human genome across populations.
    Hum Mol Genet. 2003 Apr 1;12(7):771-6 PMID: 12651872
  2. Imputation-based analysis of association studies: candidate regions and quantitative traits.
    PLoS Genet. 2007 Jul;3(7):e114 PMID: 17676998
  3. Dissecting the regulatory architecture of gene expression QTLs.
    Genome Biol. 2012 Jan 31;13(1):R7 PMID: 22293038
  4. HAPGEN2: simulation of multiple disease SNPs.
    Bioinformatics. 2011 Aug 15;27(16):2304-5 PMID: 21653516
  5. Fine-mapping the genetic association of the major histocompatibility complex in multiple sclerosis: HLA and non-HLA effects.
    PLoS Genet. 2013 Nov;9(11):e1003926 PMID: 24278027
  6. Joint analysis of functional genomic data and genome-wide association studies of 18 human traits.
    Am J Hum Genet. 2014 Apr 3;94(4):559-73 PMID: 24702953
  7. Bayesian refinement of association signals for 14 loci in 3 common diseases.
    Nat Genet. 2012 Dec;44(12):1294-301 PMID: 23104008
  8. Genome-wide trans-ancestry meta-analysis provides insight into the genetic architecture of type 2 diabetes susceptibility.
    Nat Genet. 2014 Mar;46(3):234-44 PMID: 24509480
  9. Systematic localization of common disease-associated variation in regulatory DNA.
    Science. 2012 Sep 7;337(6099):1190-5 PMID: 22955828
  10. A novel algorithm for simultaneous SNP selection in high-dimensional genome-wide association studies.
    BMC Bioinformatics. 2012 Oct 31;13:284 PMID: 23113980
  11. Learning a prior on regulatory potential from eQTL data.
    PLoS Genet. 2009 Jan;5(1):e1000358 PMID: 19180192
  12. Integrated enrichment analysis of variants and pathways in genome-wide association studies indicates central role for IL-2 signaling genes in type 1 diabetes, and cytokine signaling genes in Crohn's disease.
    PLoS Genet. 2013;9(10):e1003770 PMID: 24098138
  13. Analysis of immune-related loci identifies 48 new susceptibility variants for multiple sclerosis.
    Nat Genet. 2013 Nov;45(11):1353-60 PMID: 24076602
  14. Fine-scale mapping of the FGFR2 breast cancer risk locus: putative functional variants differentially bind FOXA1 and E2F1.
    Am J Hum Genet. 2013 Dec 5;93(6):1046-60 PMID: 24290378
  15. Fast and accurate imputation of summary statistics enhances evidence of functional enrichment.
    Bioinformatics. 2014 Oct 15;30(20):2906-14 PMID: 24990607
  16. Fine-mapping identifies multiple prostate cancer risk loci at 5p15, one of which associates with TERT expression.
    Hum Mol Genet. 2013 Jun 15;22(12):2520-8 PMID: 23535824
  17. ITPA gene variants protect against anaemia in patients treated for chronic hepatitis C.
    Nature. 2010 Mar 18;464(7287):405-8 PMID: 20173735
  18. Identifying causal variants at loci with multiple signals of association.
    Genetics. 2014 Oct;198(2):497-508 PMID: 25104515
  19. Conditional and joint multiple-SNP analysis of GWAS summary statistics identifies additional variants influencing complex traits.
    Nat Genet. 2012 Mar 18;44(4):369-75, S1-3 PMID: 22426310
  20. Potential etiologic and functional implications of genome-wide association loci for human diseases and traits.
    Proc Natl Acad Sci U S A. 2009 Jun 9;106(23):9362-7 PMID: 19474294
  21. Dense genotyping identifies and localizes multiple common and rare variant association signals in celiac disease.
    Nat Genet. 2011 Nov 06;43(12):1193-201 PMID: 22057235
  22. Systematic functional regulatory assessment of disease-associated variants.
    Proc Natl Acad Sci U S A. 2013 Jun 4;110(23):9607-12 PMID: 23690573
  23. Chromatin marks identify critical cell types for fine mapping complex trait variants.
    Nat Genet. 2013 Feb;45(2):124-30 PMID: 23263488
  24. Partitioning heritability of regulatory and cell-type-specific variants across 11 common diseases.
    Am J Hum Genet. 2014 Nov 6;95(5):535-52 PMID: 25439723
  25. The accessible chromatin landscape of the human genome.
    Nature. 2012 Sep 6;489(7414):75-82 PMID: 22955617
  26. An integrated encyclopedia of DNA elements in the human genome.
    Nature. 2012 Sep 6;489(7414):57-74 PMID: 22955616
  27. Dense fine-mapping study identifies new susceptibility loci for primary biliary cirrhosis.
    Nat Genet. 2012 Oct;44(10):1137-41 PMID: 22961000
  28. Using chromatin marks to interpret and localize genetic associations to complex human traits and diseases.
    Curr Opin Genet Dev. 2013 Dec;23(6):635-41 PMID: 24287333
  29. Re-ranking sequencing variants in the post-GWAS era for accurate causal variant identification.
    PLoS Genet. 2013;9(8):e1003609 PMID: 23950724
  30. Biological, clinical and population relevance of 95 loci for blood lipids.
    Nature. 2010 Aug 5;466(7307):707-13 PMID: 20686565
  31. Trans-ethnic fine-mapping of lipid loci identifies population-specific signals and allelic heterogeneity that increases the trait variance explained.
    PLoS Genet. 2013 Mar;9(3):e1003379 PMID: 23555291
  32. Leveraging genetic variability across populations for the identification of causal variants.
    Am J Hum Genet. 2010 Jan;86(1):23-33 PMID: 20085711
  33. Integrative variable selection via Bayesian model uncertainty.
    Stat Med. 2013 Dec 10;32(28):4938-53 PMID: 23824835
  34. Rapid and accurate multiple testing correction and power estimation for millions of correlated markers.
    PLoS Genet. 2009 Apr;5(4):e1000456 PMID: 19381255
  35. Reprioritizing genetic associations in hit regions using LASSO-based resample model averaging.
    Genet Epidemiol. 2012 Jul;36(5):451-62 PMID: 22549815
  36. FGFR2 variants and breast cancer risk: fine-scale mapping using African American studies and analysis of chromatin conformation.
    Hum Mol Genet. 2009 May 1;18(9):1692-703 PMID: 19223389
  37. So many correlated tests, so little time! Rapid adjustment of P values for multiple correlated tests.
    Am J Hum Genet. 2007 Dec;81(6):1158-68 PMID: 17966093
  38. Hierarchical Bayes prioritization of marker associations from a genome-wide association scan for further investigation.
    Genet Epidemiol. 2007 Dec;31(8):871-82 PMID: 17654612
  39. Evaluating the power to discriminate between highly correlated SNPs in genetic association studies.
    Genet Epidemiol. 2010 Jul;34(5):463-8 PMID: 20583289
Article Info
Journal
PLoS genetics
Abbr.
PLoS Genet
ISSN
1553-7404
Published
2014-10-00
Epub
2014-00-30
Pages
e1004722
Language
English
Region
United States
NLM ID
101239074
PMCID
PMC4214605
Subset
IM
Grants
NCI NIH HHS · R21 CA182821 · United States
NCI NIH HHS · R21-CA182821 · United States
NCI NIH HHS · R03 CA162200 · United States
NIGMS NIH HHS · R01 GM053275 · United States
NIGMS NIH HHS · R01-GM053275 · United States
NCI NIH HHS · R03-CA162200 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]