Home LiteratureArticle Details
PMID: 22883141 Published · ppublish English Comparative Study Evaluation Study Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

Phasing of many thousands of genotyped samples.

American journal of human genetics ·Vol. 91 ·No. 2 ·2012-08-10 ·Pages 238-51

Williams AL, Patterson N, Glessner J, Hakonarson H, Reich D

Abstract

Haplotypes are an important resource for a large number of applications in human genetics, but computationally inferred haplotypes are subject to switch errors that decrease their utility. The accuracy of computationally inferred haplotypes increases with sample size, and although ever larger genotypic data sets are being generated, the fact that existing methods require substantial computational resources limits their applicability to data sets containing tens or hundreds of thousands of samples. Here, we present HAPI-UR (haplotype inference for unrelated samples), an algorithm that is designed to handle unrelated and/or trio and duo family data, that has accuracy comparable to or greater than existing methods, and that is computationally efficient and can be applied to 100,000 samples or more. We use HAPI-UR to phase a data set with 58,207 samples and show that it achieves practical runtime and that switch errors decrease with sample size even with the use of samples from multiple ethnicities. Using a data set with 16,353 samples, we compare HAPI-UR to Beagle, MaCH, IMPUTE2, and SHAPEIT and show that HAPI-UR runs 18× faster than all methods and has a lower switch-error rate than do other methods except for Beagle; with the use of consensus phasing, running HAPI-UR three times gives a slightly lower switch-error rate than Beagle does and is more than six times faster. We demonstrate results similar to those from Beagle on another data set with a higher marker density. Lastly, we show that HAPI-UR has better runtime scaling properties than does Beagle so that for larger data sets, HAPI-UR will be practical and will have an even larger runtime advantage. HAPI-UR is available online (see Web Resources).

MeSH Terms
Algorithms Computational Biology/methods Genetics Haplotypes/genetics Humans Internet Research Design Sample Size Software
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Williams Amy L
Department of Genetics, Harvard Medical School, Boston, MA 02115, USA; Broad Institute of Harvard and MIT, Cambridge, MA 02142, USA. [email protected]
Patterson Nick
Glessner Joseph
Hakonarson Hakon
Reich David
References (30)
30 references, click to expand
  1. Haplotype phasing: existing methods and new developments.
    Nat Rev Genet. 2011 Sep 16;12(10):703-14 PMID: 21921926
  2. Detection of sharing by descent, long-range phasing and haplotype imputation.
    Nat Genet. 2008 Sep;40(9):1068-75 PMID: 19165921
  3. Genome-wide copy number variation study associates metabotropic glutamate receptor gene networks with attention deficit hyperactivity disorder.
    Nat Genet. 2011 Dec 04;44(1):78-84 PMID: 22138692
  4. Integrating common and rare genetic variation in diverse human populations.
    Nature. 2010 Sep 2;467(7311):52-8 PMID: 20811451
  5. Loci on 20q13 and 21q22 are associated with pediatric-onset inflammatory bowel disease.
    Nat Genet. 2008 Oct;40(10):1211-5 PMID: 18758464
  6. A map of recent positive selection in the human genome.
    PLoS Biol. 2006 Mar;4(3):e72 PMID: 16494531
  7. Genotype imputation for genome-wide association studies.
    Nat Rev Genet. 2010 Jul;11(7):499-511 PMID: 20517342
  8. Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls.
    Nature. 2007 Jun 7;447(7145):661-78 PMID: 17554300
  9. Genome-wide association study of ulcerative colitis identifies three new susceptibility loci, including the HNF4A region.
    Nat Genet. 2009 Dec;41(12):1330-4 PMID: 19915572
  10. Genotype-imputation accuracy across worldwide human populations.
    Am J Hum Genet. 2009 Feb;84(2):235-50 PMID: 19215730
  11. Genotype imputation with thousands of genomes.
    G3 (Bethesda). 2011 Nov;1(6):457-70 PMID: 22384356
  12. Accounting for decay of linkage disequilibrium in haplotype inference and missing-data imputation.
    Am J Hum Genet. 2005 Mar;76(3):449-62 PMID: 15700229
  13. Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering.
    Am J Hum Genet. 2007 Nov;81(5):1084-97 PMID: 17924348
  14. Autism genome-wide copy number variation reveals ubiquitin and neuronal genes.
    Nature. 2009 May 28;459(7246):569-73 PMID: 19404257
  15. Imputation of low-frequency variants using the HapMap3 benefits from large, diverse reference sets.
    Eur J Hum Genet. 2011 Jun;19(6):662-6 PMID: 21364697
  16. Potential etiologic and functional implications of genome-wide association loci for human diseases and traits.
    Proc Natl Acad Sci U S A. 2009 Jun 9;106(23):9362-7 PMID: 19474294
  17. A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.
    PLoS Genet. 2009 Jun;5(6):e1000529 PMID: 19543373
  18. Haplotype inference in random population samples.
    Am J Hum Genet. 2002 Nov;71(5):1129-37 PMID: 12386835
  19. MaCH: using sequence and genotype data to estimate haplotypes and unobserved genotypes.
    Genet Epidemiol. 2010 Dec;34(8):816-34 PMID: 21058334
  20. A genome-wide association study identifies KIAA0350 as a type 1 diabetes gene.
    Nature. 2007 Aug 2;448(7153):591-4 PMID: 17632545
  21. Detecting recent positive selection in the human genome from haplotype structure.
    Nature. 2002 Oct 24;419(6909):832-7 PMID: 12397357
  22. A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase.
    Am J Hum Genet. 2006 Apr;78(4):629-44 PMID: 16532393
  23. Whole population, genome-wide mapping of hidden relatedness.
    Genome Res. 2009 Feb;19(2):318-26 PMID: 18971310
  24. Modeling linkage disequilibrium and identifying recombination hotspots using single-nucleotide polymorphism data.
    Genetics. 2003 Dec;165(4):2213-33 PMID: 14704198
  25. Haplotype-resolved genome sequencing of a Gujarati Indian individual.
    Nat Biotechnol. 2011 Jan;29(1):59-63 PMID: 21170042
  26. Sensitive detection of chromosomal segments of distinct ancestry in admixed populations.
    PLoS Genet. 2009 Jun;5(6):e1000519 PMID: 19543370
  27. A comparison of phasing algorithms for trios and unrelated individuals.
    Am J Hum Genet. 2006 Mar;78(3):437-50 PMID: 16465620
  28. Principal component analysis of genetic data.
    Nat Genet. 2008 May;40(5):491-2 PMID: 18443580
  29. Effect of genetic divergence in identifying ancestral origin using HAPAA.
    Genome Res. 2008 Apr;18(4):676-82 PMID: 18353807
  30. A linear complexity phasing method for thousands of genomes.
    Nat Methods. 2011 Dec 04;9(2):179-81 PMID: 22138821
Article Info
Journal
American journal of human genetics
Abbr.
Am J Hum Genet
ISSN
1537-6605
Published
2012-08-10
Pages
238-51
Language
English
Region
United States
NLM ID
0370475
PMCID
PMC3415548
Subset
IM
Grants
NIGMS NIH HHS · R01 GM100233 · United States
NHGRI NIH HHS · F32 HG005944 · United States
Wellcome Trust · 076113 · United Kingdom
NHGRI NIH HHS · F32HG005944 · United States
Wellcome Trust · 085475 · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]