Abstract
Population stratification has long been recognized as a confounding factor in genetic association studies. Estimated ancestries, derived from multi-locus genotype data, can be used to perform a statistical correction for population stratification. One popular technique for estimation of ancestry is the model-based approach embodied by the widely applied program structure. Another approach, implemented in the program EIGENSTRAT, relies on Principal Component Analysis rather than model-based estimation and does not directly deliver admixture fractions. EIGENSTRAT has gained in popularity in part owing to its remarkable speed in comparison to structure. We present a new algorithm and a program, ADMIXTURE, for model-based estimation of ancestry in unrelated individuals. ADMIXTURE adopts the likelihood model embedded in structure. However, ADMIXTURE runs considerably faster, solving problems in minutes that take structure hours. In many of our experiments, we have found that ADMIXTURE is almost as fast as EIGENSTRAT. The runtime improvements of ADMIXTURE rely on a fast block relaxation scheme using sequential quadratic programming for block updates, coupled with a novel quasi-Newton acceleration of convergence. Our algorithm also runs faster and with greater accuracy than the implementation of an Expectation-Maximization (EM) algorithm incorporated in the program FRAPPE. Our simulations show that ADMIXTURE's maximum likelihood estimates of the underlying admixture coefficients and ancestral allele frequencies are as accurate as structure's Bayesian estimates. On real-world data sets, ADMIXTURE's estimates are directly comparable to those from structure and EIGENSTRAT. Taken together, our results show that ADMIXTURE's computational speed opens up the possibility of using a much larger set of markers in model-based ancestry estimation and that its estimates are suitable for use in correcting for population stratification in association studies.
MeSH Terms
Algorithms
Computational Biology
Europe/ethnology
Gene Frequency
Genetic Association Studies
Genetics, Population
Genotype
Humans
Inflammatory Bowel Diseases/ethnology,genetics
Jews/ethnology
Likelihood Functions
Models, Genetic
Polymorphism, Single Nucleotide
Software
Time Factors
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Alexander David H
Department of Biomathematics, University of California at Los Angeles, Los Angeles, California 90095, USA.
[email protected]
Novembre John
Lange Kenneth
References (20)
20 references, click to expand
-
Estimating local ancestry in admixed populations.
Am J Hum Genet. 2008 Feb;82(2):290-303
PMID: 18252211
-
Worldwide human relationships inferred from genome-wide patterns of variation.
Science. 2008 Feb 22;319(5866):1100-4
PMID: 18292342
-
The New York Cancer Project: rationale, organization, design, and baseline characteristics.
J Urban Health. 2004 Jun;81(2):301-10
PMID: 15136663
-
A haplotype map of the human genome.
Nature. 2005 Oct 27;437(7063):1299-320
PMID: 16255080
-
Population subdivision with respect to multiple alleles.
Ann Hum Genet. 1969 Jul;33(1):23-9
PMID: 5821316
-
Case-control studies of association in structured or admixed populations.
Theor Popul Biol. 2001 Nov;60(3):227-37
PMID: 11855957
-
Genes mirror geography within Europe.
Nature. 2008 Nov 6;456(7218):98-101
PMID: 18758442
-
Discerning the ancestry of European Americans in genetic association studies.
PLoS Genet. 2008 Jan;4(1):e236
PMID: 18208327
-
Population structure and eigenanalysis.
PLoS Genet. 2006 Dec;2(12):e190
PMID: 17194218
-
Gm3;5,13,14 and type 2 diabetes mellitus: an association in American Indians with genetic admixture.
Am J Hum Genet. 1988 Oct;43(4):520-6
PMID: 3177389
-
Principal components analysis corrects for stratification in genome-wide association studies.
Nat Genet. 2006 Aug;38(8):904-9
PMID: 16862161
-
Estimation of individual admixture: analytical and study design considerations.
Genet Epidemiol. 2005 May;28(4):289-301
PMID: 15712363
-
Methods for high-density admixture mapping of disease genes.
Am J Hum Genet. 2004 May;74(5):979-1000
PMID: 15088269
-
Reconstructing genetic ancestry blocks in admixed individuals.
Am J Hum Genet. 2006 Jul;79(1):1-12
PMID: 16773560
-
Interpreting principal component analyses of spatial population genetic variation.
Nat Genet. 2008 May;40(5):646-9
PMID: 18425127
-
On the inference of ancestries in admixed populations.
Genome Res. 2008 Apr;18(4):668-75
PMID: 18353809
-
Inference of population structure using multilocus genotype data.
Genetics. 2000 Jun;155(2):945-59
PMID: 10835412
-
Genotype, haplotype and copy-number variation in worldwide human populations.
Nature. 2008 Feb 21;451(7181):998-1003
PMID: 18288195
-
The effects of human population structure on large genetic association studies.
Nat Genet. 2004 May;36(5):512-7
PMID: 15052271
-
Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies.
Genetics. 2003 Aug;164(4):1567-87
PMID: 12930761