Abstract
By aggregating data for complex traits in a biologically meaningful way, gene and gene-set analysis constitute a valuable addition to single-marker analysis. However, although various methods for gene and gene-set analysis currently exist, they generally suffer from a number of issues. Statistical power for most methods is strongly affected by linkage disequilibrium between markers, multi-marker associations are often hard to detect, and the reliance on permutation to compute p-values tends to make the analysis computationally very expensive. To address these issues we have developed MAGMA, a novel tool for gene and gene-set analysis. The gene analysis is based on a multiple regression model, to provide better statistical performance. The gene-set analysis is built as a separate layer around the gene analysis for additional flexibility. This gene-set analysis also uses a regression structure to allow generalization to analysis of continuous properties of genes and simultaneous analysis of multiple gene sets and other gene properties. Simulations and an analysis of Crohn's Disease data are used to evaluate the performance of MAGMA and to compare it to a number of other gene and gene-set analysis tools. The results show that MAGMA has significantly more power than other tools for both the gene and the gene-set analysis, identifying more genes and gene sets associated with Crohn's Disease while maintaining a correct type 1 error rate. Moreover, the MAGMA analysis of the Crohn's Disease data was found to be considerably faster as well.
MeSH Terms
Computational Biology/methods
Computer Simulation
Crohn Disease/genetics
Databases, Genetic
Genome-Wide Association Study/methods
Humans
Models, Genetic
Software
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
de Leeuw Christiaan A
Department of Complex Trait Genetics, Center for Neurogenomics and Cognitive Research, VU University Amsterdam, Amsterdam, The Netherlands; Institute for Computing and Information Sciences, Radboud University Nijmegen, Nijmegen, The Netherlands.
Mooij Joris M
Informatics Institute, University of Amsterdam, Amsterdam, The Netherlands.
Heskes Tom
Institute for Computing and Information Sciences, Radboud University Nijmegen, Nijmegen, The Netherlands.
Posthuma Danielle
Department of Complex Trait Genetics, Center for Neurogenomics and Cognitive Research, VU University Amsterdam, Amsterdam, The Netherlands; Department of Clinical Genetics, VU University Medical Centre Amsterdam, Neuroscience Campus Amsterdam, The Netherlands.
References (22)
22 references, click to expand
-
Finding the missing heritability of complex diseases.
Nature. 2009 Oct 8;461(7265):747-53
PMID: 19812666
-
A versatile gene-based test for genome-wide association studies.
Am J Hum Genet. 2010 Jul 9;87(1):139-45
PMID: 20598278
-
Common inherited variation in mitochondrial genes is not enriched for associations with type 2 diabetes or related glycemic traits.
PLoS Genet. 2010 Aug;6(8). pii: e1001058. doi: 10.1371/journal.pgen.1001058
PMID: 20714348
-
Integrating common and rare genetic variation in diverse human populations.
Nature. 2010 Sep 2;467(7311):52-8
PMID: 20811451
-
Hundreds of variants clustered in genomic loci and biological pathways affect human height.
Nature. 2010 Oct 14;467(7317):832-8
PMID: 20881960
-
Association analyses of 249,796 individuals reveal 18 new loci associated with body mass index.
Nat Genet. 2010 Nov;42(11):937-48
PMID: 20935630
-
Data quality control in genetic case-control association studies.
Nat Protoc. 2010 Sep;5(9):1564-73
PMID: 21085122
-
Genome-wide meta-analysis increases to 71 the number of confirmed Crohn's disease susceptibility loci.
Nat Genet. 2010 Dec;42(12):1118-25
PMID: 21102463
-
Estimating missing heritability for disease from genome-wide association studies.
Am J Hum Genet. 2011 Mar 11;88(3):294-305
PMID: 21376301
-
Pathway-based approaches for analysis of genomewide association studies.
Am J Hum Genet. 2007 Dec;81(6):1278-83
PMID: 17966091
-
Gene set analysis of genome-wide association studies: methodological issues and perspectives.
Genomics. 2011 Jul;98(1):1-8
PMID: 21565265
-
Five years of GWAS discovery.
Am J Hum Genet. 2012 Jan 13;90(1):7-24
PMID: 22243964
-
INRICH: interval-based enrichment analysis for genome-wide association studies.
Bioinformatics. 2012 Jul 1;28(13):1797-9
PMID: 22513993
-
Permutation-based approaches do not adequately allow for linkage disequilibrium in gene-wide multi-locus association analysis.
Eur J Hum Genet. 2012 Aug;20(8):890-6
PMID: 22317971
-
HYST: a hybrid set-based test for genome-wide association studies, with application to protein-protein interaction-based association analysis.
Am J Hum Genet. 2012 Sep 7;91(3):478-88
PMID: 22958900
-
Functional gene group analysis identifies synaptic gene groups as risk factor for schizophrenia.
Mol Psychiatry. 2012 Oct;17(10):996-1006
PMID: 21931320
-
An integrated map of genetic variation from 1,092 human genomes.
Nature. 2012 Nov 1;491(7422):56-65
PMID: 23128226
-
Genome-wide association analysis identifies 13 new risk loci for schizophrenia.
Nat Genet. 2013 Oct;45(10):1150-9
PMID: 23974872
-
Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles.
Proc Natl Acad Sci U S A. 2005 Oct 25;102(43):15545-50
PMID: 16199517
-
Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls.
Nature. 2007 Jun 7;447(7145):661-78
PMID: 17554300
-
PLINK: a tool set for whole-genome association and population-based linkage analyses.
Am J Hum Genet. 2007 Sep;81(3):559-75
PMID: 17701901
-
Gene ontology analysis of GWA study data sets provides insights into the biology of bipolar disorder.
Am J Hum Genet. 2009 Jul;85(1):13-24
PMID: 19539887