Abstract
Haplotypes are an important resource for a large number of applications in human genetics, but computationally inferred haplotypes are subject to switch errors that decrease their utility. The accuracy of computationally inferred haplotypes increases with sample size, and although ever larger genotypic data sets are being generated, the fact that existing methods require substantial computational resources limits their applicability to data sets containing tens or hundreds of thousands of samples. Here, we present HAPI-UR (haplotype inference for unrelated samples), an algorithm that is designed to handle unrelated and/or trio and duo family data, that has accuracy comparable to or greater than existing methods, and that is computationally efficient and can be applied to 100,000 samples or more. We use HAPI-UR to phase a data set with 58,207 samples and show that it achieves practical runtime and that switch errors decrease with sample size even with the use of samples from multiple ethnicities. Using a data set with 16,353 samples, we compare HAPI-UR to Beagle, MaCH, IMPUTE2, and SHAPEIT and show that HAPI-UR runs 18× faster than all methods and has a lower switch-error rate than do other methods except for Beagle; with the use of consensus phasing, running HAPI-UR three times gives a slightly lower switch-error rate than Beagle does and is more than six times faster. We demonstrate results similar to those from Beagle on another data set with a higher marker density. Lastly, we show that HAPI-UR has better runtime scaling properties than does Beagle so that for larger data sets, HAPI-UR will be practical and will have an even larger runtime advantage. HAPI-UR is available online (see Web Resources).
MeSH Terms
Algorithms
Computational Biology/methods
Genetics
Haplotypes/genetics
Humans
Internet
Research Design
Sample Size
Software
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Williams Amy L
Department of Genetics, Harvard Medical School, Boston, MA 02115, USA; Broad Institute of Harvard and MIT, Cambridge, MA 02142, USA.
[email protected]
Patterson Nick
Glessner Joseph
Hakonarson Hakon
Reich David
References (30)
30 references, click to expand
-
Haplotype phasing: existing methods and new developments.
Nat Rev Genet. 2011 Sep 16;12(10):703-14
PMID: 21921926
-
Detection of sharing by descent, long-range phasing and haplotype imputation.
Nat Genet. 2008 Sep;40(9):1068-75
PMID: 19165921
-
Genome-wide copy number variation study associates metabotropic glutamate receptor gene networks with attention deficit hyperactivity disorder.
Nat Genet. 2011 Dec 04;44(1):78-84
PMID: 22138692
-
Integrating common and rare genetic variation in diverse human populations.
Nature. 2010 Sep 2;467(7311):52-8
PMID: 20811451
-
Loci on 20q13 and 21q22 are associated with pediatric-onset inflammatory bowel disease.
Nat Genet. 2008 Oct;40(10):1211-5
PMID: 18758464
-
A map of recent positive selection in the human genome.
PLoS Biol. 2006 Mar;4(3):e72
PMID: 16494531
-
Genotype imputation for genome-wide association studies.
Nat Rev Genet. 2010 Jul;11(7):499-511
PMID: 20517342
-
Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls.
Nature. 2007 Jun 7;447(7145):661-78
PMID: 17554300
-
Genome-wide association study of ulcerative colitis identifies three new susceptibility loci, including the HNF4A region.
Nat Genet. 2009 Dec;41(12):1330-4
PMID: 19915572
-
Genotype-imputation accuracy across worldwide human populations.
Am J Hum Genet. 2009 Feb;84(2):235-50
PMID: 19215730
-
Genotype imputation with thousands of genomes.
G3 (Bethesda). 2011 Nov;1(6):457-70
PMID: 22384356
-
Accounting for decay of linkage disequilibrium in haplotype inference and missing-data imputation.
Am J Hum Genet. 2005 Mar;76(3):449-62
PMID: 15700229
-
Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering.
Am J Hum Genet. 2007 Nov;81(5):1084-97
PMID: 17924348
-
Autism genome-wide copy number variation reveals ubiquitin and neuronal genes.
Nature. 2009 May 28;459(7246):569-73
PMID: 19404257
-
Imputation of low-frequency variants using the HapMap3 benefits from large, diverse reference sets.
Eur J Hum Genet. 2011 Jun;19(6):662-6
PMID: 21364697
-
Potential etiologic and functional implications of genome-wide association loci for human diseases and traits.
Proc Natl Acad Sci U S A. 2009 Jun 9;106(23):9362-7
PMID: 19474294
-
A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.
PLoS Genet. 2009 Jun;5(6):e1000529
PMID: 19543373
-
Haplotype inference in random population samples.
Am J Hum Genet. 2002 Nov;71(5):1129-37
PMID: 12386835
-
MaCH: using sequence and genotype data to estimate haplotypes and unobserved genotypes.
Genet Epidemiol. 2010 Dec;34(8):816-34
PMID: 21058334
-
A genome-wide association study identifies KIAA0350 as a type 1 diabetes gene.
Nature. 2007 Aug 2;448(7153):591-4
PMID: 17632545
-
Detecting recent positive selection in the human genome from haplotype structure.
Nature. 2002 Oct 24;419(6909):832-7
PMID: 12397357
-
A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase.
Am J Hum Genet. 2006 Apr;78(4):629-44
PMID: 16532393
-
Whole population, genome-wide mapping of hidden relatedness.
Genome Res. 2009 Feb;19(2):318-26
PMID: 18971310
-
Modeling linkage disequilibrium and identifying recombination hotspots using single-nucleotide polymorphism data.
Genetics. 2003 Dec;165(4):2213-33
PMID: 14704198
-
Haplotype-resolved genome sequencing of a Gujarati Indian individual.
Nat Biotechnol. 2011 Jan;29(1):59-63
PMID: 21170042
-
Sensitive detection of chromosomal segments of distinct ancestry in admixed populations.
PLoS Genet. 2009 Jun;5(6):e1000519
PMID: 19543370
-
A comparison of phasing algorithms for trios and unrelated individuals.
Am J Hum Genet. 2006 Mar;78(3):437-50
PMID: 16465620
-
Principal component analysis of genetic data.
Nat Genet. 2008 May;40(5):491-2
PMID: 18443580
-
Effect of genetic divergence in identifying ancestral origin using HAPAA.
Genome Res. 2008 Apr;18(4):676-82
PMID: 18353807
-
A linear complexity phasing method for thousands of genomes.
Nat Methods. 2011 Dec 04;9(2):179-81
PMID: 22138821