Home LiteratureArticle Details
PMID: 35410384 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, N.I.H., Extramural

Pangenome-based genome inference allows efficient and accurate genotyping across a wide spectrum of variant classes.

Nature genetics ·Vol. 54 ·No. 4 ·2022-00-00 ·Pages 518-525

Ebler J, Ebert P, Clarke WE, Rausch T, Audano PA, Houwaart T, Mao Y, Korbel JO, Eichler EE, Zody MC, Dilthey AT, Marschall T

Abstract

Typical genotyping workflows map reads to a reference genome before identifying genetic variants. Generating such alignments introduces reference biases and comes with substantial computational burden. Furthermore, short-read lengths limit the ability to characterize repetitive genomic regions, which are particularly challenging for fast k-mer-based genotypers. In the present study, we propose a new algorithm, PanGenie, that leverages a haplotype-resolved pangenome reference together with k-mer counts from short-read sequencing data to genotype a wide spectrum of genetic variation-a process we refer to as genome inference. Compared with mapping-based approaches, PanGenie is more than 4 times faster at 30-fold coverage and achieves better genotype concordances for almost all variant types and coverages tested. Improvements are especially pronounced for large insertions (≥50 bp) and variants in repetitive regions, enabling the inclusion of these classes of variants in genome-wide association studies. PanGenie efficiently leverages the increasing amount of haplotype-resolved assemblies to unravel the functional impact of previously inaccessible variants while being faster compared with alignment-based workflows.

MeSH Terms
Algorithms Genetic Variation Genome, Human/genetics Genome-Wide Association Study Genomics/methods Genotype High-Throughput Nucleotide Sequencing Humans Sequence Analysis, DNA
Authors & Affiliations
12 authors, click to expand affiliations / ORCID
Ebler Jana ORCID
Institute for Medical Biometry and Bioinformatics, Medical Faculty, Heinrich Heine University Düsseldorf, Düsseldorf, Germany.
Ebert Peter ORCID
Institute for Medical Biometry and Bioinformatics, Medical Faculty, Heinrich Heine University Düsseldorf, Düsseldorf, Germany.
Clarke Wayne E
New York Genome Center, New York, NY, USA.
Rausch Tobias ORCID
European Molecular Biology Laboratory, Genome Biology Unit, Heidelberg, Germany. | European Molecular Biology Laboratory, GeneCore, Heidelberg, Germany.
Audano Peter A
Department of Genome Sciences, University of Washington School of Medicine, Seattle, WA, USA.
Houwaart Torsten ORCID
Institute of Medical Microbiology and Hospital Hygiene, Heinrich Heine University Düsseldorf, Düsseldorf, Germany.
Mao Yafei ORCID
Department of Genome Sciences, University of Washington School of Medicine, Seattle, WA, USA.
Korbel Jan O ORCID
European Molecular Biology Laboratory, Genome Biology Unit, Heidelberg, Germany.
Eichler Evan E ORCID
Department of Genome Sciences, University of Washington School of Medicine, Seattle, WA, USA. | Howard Hughes Medical Institute, University of Washington, Seattle, WA, USA.
Zody Michael C ORCID
New York Genome Center, New York, NY, USA.
Dilthey Alexander T
Institute of Medical Microbiology and Hospital Hygiene, Heinrich Heine University Düsseldorf, Düsseldorf, Germany. | Institute of Medical Statistics and Computational Biology, University of Cologne, Cologne, Germany. | Cologne Excellence Cluster on Cellular Stress Responses in Aging-Associated Diseases, University of Cologne, Cologne, Germany.
Marschall Tobias ORCID
Institute for Medical Biometry and Bioinformatics, Medical Faculty, Heinrich Heine University Düsseldorf, Düsseldorf, Germany. [email protected].
References (71)
71 references, click to expand
  1. Genome measures used for quality control are dependent on gene function and ancestry.
    Bioinformatics. 2015 Feb 1;31(3):318-23 PMID: 25297068
  2. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype.
    Nat Biotechnol. 2019 Aug;37(8):907-915 PMID: 31375807
  3. Integrating mapping-, assembly- and haplotype-based approaches for calling variants in clinical sequencing applications.
    Nat Genet. 2014 Aug;46(8):912-918 PMID: 25017105
  4. The variant call format and VCFtools.
    Bioinformatics. 2011 Aug 1;27(15):2156-8 PMID: 21653522
  5. An open resource for accurately benchmarking small variant and reference calls.
    Nat Biotechnol. 2019 May;37(5):561-566 PMID: 30936564
  6. A linear complexity phasing method for thousands of genomes.
    Nat Methods. 2011 Dec 04;9(2):179-81 PMID: 22138821
  7. Genome-wide association study of CNVs in 16,000 cases of eight common diseases and 3,000 shared controls.
    Nature. 2010 Apr 1;464(7289):713-20 PMID: 20360734
  8. The structure, function and evolution of a complete human chromosome 8.
    Nature. 2021 May;593(7857):101-107 PMID: 33828295
  9. DELLY: structural variant discovery by integrated paired-end and split-read analysis.
    Bioinformatics. 2012 Sep 15;28(18):i333-i339 PMID: 22962449
  10. Genotype imputation with thousands of genomes.
    G3 (Bethesda). 2011 Nov;1(6):457-70 PMID: 22384356
  11. Fast and accurate genomic analyses using genome graphs.
    Nat Genet. 2019 Feb;51(2):354-362 PMID: 30643257
  12. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  13. Strand-seq enables reliable separation of long reads by chromosome via expectation maximization.
    Bioinformatics. 2018 Jul 1;34(13):i115-i123 PMID: 29949971
  14. High-Accuracy HLA Type Inference from Whole-Genome Sequencing Data Using Population Reference Graphs.
    PLoS Comput Biol. 2016 Oct 28;12(10):e1005151 PMID: 27792722
  15. Paragraph: a graph-based structural variant genotyper for short-read sequence data.
    Genome Biol. 2019 Dec 19;20(1):291 PMID: 31856913
  16. High frequencies of de novo CNVs in bipolar disorder and schizophrenia.
    Neuron. 2011 Dec 22;72(6):951-63 PMID: 22196331
  17. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios.
    Cell. 2022 Sep 1;185(18):3426-3440.e19 PMID: 36055201
  18. Multiple recurrent de novo CNVs, including duplications of the 7q11.23 Williams syndrome region, are strongly associated with autism.
    Neuron. 2011 Jun 9;70(5):863-85 PMID: 21658581
  19. An integrated map of structural variation in 2,504 human genomes.
    Nature. 2015 Oct 1;526(7571):75-81 PMID: 26432246
  20. dbSNP: the NCBI database of genetic variation.
    Nucleic Acids Res. 2001 Jan 1;29(1):308-11 PMID: 11125122
  21. Three-stage quality control strategies for DNA re-sequencing data.
    Brief Bioinform. 2014 Nov;15(6):879-89 PMID: 24067931
  22. A One-Penny Imputed Genome from Next-Generation Reference Panels.
    Am J Hum Genet. 2018 Sep 6;103(3):338-348 PMID: 30100085
  23. Improved genome inference in the MHC using a population reference graph.
    Nat Genet. 2015 Jun;47(6):682-8 PMID: 25915597
  24. Next-generation genotype imputation service and methods.
    Nat Genet. 2016 Oct;48(10):1284-1287 PMID: 27571263
  25. The Human Pangenome Project: a global resource to map genomic diversity.
    Nature. 2022 Apr;604(7906):437-446 PMID: 35444317
  26. Immune diversity sheds light on missing variation in worldwide genetic diversity panels.
    PLoS One. 2018 Oct 26;13(10):e0206512 PMID: 30365549
  27. HLA*LA-HLA typing from linearly projected graph alignments.
    Bioinformatics. 2019 Nov 1;35(21):4394-4396 PMID: 30942877
  28. High-resolution comparative analysis of great ape genomes.
    Science. 2018 Jun 8;360(6393): PMID: 29880660
  29. Histo-blood group gene polymorphisms as potential genetic modifiers of infection and cystic fibrosis lung disease severity.
    PLoS One. 2009;4(1):e4270 PMID: 19169360
  30. IPD--the Immuno Polymorphism Database.
    Nucleic Acids Res. 2010 Jan;38(Database issue):D863-9 PMID: 19875415
  31. The UCSC Table Browser data retrieval tool.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D493-6 PMID: 14681465
  32. Rare structural variants disrupt multiple genes in neurodevelopmental pathways in schizophrenia.
    Science. 2008 Apr 25;320(5875):539-43 PMID: 18369103
  33. DNA-based methods in the immunohematology reference laboratory.
    Transfus Apher Sci. 2011 Feb;44(1):65-72 PMID: 21257350
  34. A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.
    PLoS Genet. 2009 Jun;5(6):e1000529 PMID: 19543373
  35. Graphtyper enables population-scale genotyping using pangenome graphs.
    Nat Genet. 2017 Nov;49(11):1654-1660 PMID: 28945251
  36. GraphTyper2 enables population-scale genotyping of structural variation using pangenome graphs.
    Nat Commun. 2019 Nov 27;10(1):5402 PMID: 31776332
  37. Rare chromosomal deletions and duplications in attention-deficit hyperactivity disorder: a genome-wide analysis.
    Lancet. 2010 Oct 23;376(9750):1401-8 PMID: 20888040
  38. Genotype Imputation with Millions of Reference Samples.
    Am J Hum Genet. 2016 Jan 7;98(1):116-26 PMID: 26748515
  39. A genome-wide association study identifies protein quantitative trait loci (pQTLs).
    PLoS Genet. 2008 May 09;4(5):e1000072 PMID: 18464913
  40. A structural variation reference for medical and population genetics.
    Nature. 2020 May;581(7809):444-451 PMID: 32461652
  41. A framework for variation discovery and genotyping using next-generation DNA sequencing data.
    Nat Genet. 2011 May;43(5):491-8 PMID: 21478889
  42. Modeling linkage disequilibrium and identifying recombination hotspots using single-nucleotide polymorphism data.
    Genetics. 2003 Dec;165(4):2213-33 PMID: 14704198
  43. De novo assembly of haplotype-resolved genomes with trio binning.
    Nat Biotechnol. 2018 Oct 22;: PMID: 30346939
  44. Expectations and blind spots for structural variation detection from long-read assemblies and short-read genome sequencing technologies.
    Am J Hum Genet. 2021 May 6;108(5):919-928 PMID: 33789087
  45. Genotyping structural variants in pangenome graphs using the vg toolkit.
    Genome Biol. 2020 Feb 12;21(1):35 PMID: 32051000
  46. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019.
    Nucleic Acids Res. 2019 Jan 8;47(D1):D1005-D1012 PMID: 30445434
  47. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm.
    Nat Methods. 2021 Feb;18(2):170-175 PMID: 33526886
  48. Chromosome-scale, haplotype-resolved assembly of human genomes.
    Nat Biotechnol. 2021 Mar;39(3):309-312 PMID: 33288905
  49. Fully phased human genome assembly without parental data using single-cell strand sequencing and long reads.
    Nat Biotechnol. 2021 Mar;39(3):302-308 PMID: 33288906
  50. A global reference for human genetic variation.
    Nature. 2015 Oct 1;526(7571):68-74 PMID: 26432245
  51. PLINK: a tool set for whole-genome association and population-based linkage analyses.
    Am J Hum Genet. 2007 Sep;81(3):559-75 PMID: 17701901
  52. Genotype calling and phasing using next-generation sequencing reads and a haplotype scaffold.
    Bioinformatics. 2013 Jan 1;29(1):84-91 PMID: 23093610
  53. Haplotype-resolved diverse human genomes and integrated analysis of structural variation.
    Science. 2021 Apr 2;372(6537): PMID: 33632895
  54. Extensive sequencing of seven human genomes to characterize benchmark reference materials.
    Sci Data. 2016 Jun 07;3:160025 PMID: 27271295
  55. Multi-platform discovery of haplotype-resolved structural variation in human genomes.
    Nat Commun. 2019 Apr 16;10(1):1784 PMID: 30992455
  56. Curated variation benchmarks for challenging medically relevant autosomal genes.
    Nat Biotechnol. 2022 May;40(5):672-680 PMID: 35132260
  57. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers.
    Bioinformatics. 2011 Mar 15;27(6):764-70 PMID: 21217122
  58. SpeedSeq: ultra-fast personal genome analysis and interpretation.
    Nat Methods. 2015 Oct;12(10):966-8 PMID: 26258291
  59. Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes.
    Genome Res. 2017 Feb;27(2):300-309 PMID: 27986821
  60. Accurate genotyping across variant classes and lengths using variant graphs.
    Nat Genet. 2018 Jul;50(7):1054-1059 PMID: 29915429
  61. Integrating long-range connectivity information into de Bruijn graphs.
    Bioinformatics. 2018 Aug 1;34(15):2556-2565 PMID: 29554215
  62. De novo assembly and genotyping of variants using colored de Bruijn graphs.
    Nat Genet. 2012 Jan 08;44(2):226-32 PMID: 22231483
  63. Pangenomics enables genotyping of known structural variants in 5202 diverse genomes.
    Science. 2021 Dec 17;374(6574):abg8871 PMID: 34914532
  64. The ENCODE (ENCyclopedia Of DNA Elements) Project.
    Science. 2004 Oct 22;306(5696):636-40 PMID: 15499007
  65. Minimap2: pairwise alignment for nucleotide sequences.
    Bioinformatics. 2018 Sep 15;34(18):3094-3100 PMID: 29750242
  66. Strong association of de novo copy number mutations with autism.
    Science. 2007 Apr 20;316(5823):445-9 PMID: 17363630
  67. Toward fast and accurate SNP genotyping from whole genome sequencing data for bedside diagnostics.
    Bioinformatics. 2019 Feb 1;35(3):415-420 PMID: 30032192
  68. Phenotypic impact of genomic structural variation: insights from and for human disease.
    Nat Rev Genet. 2013 Feb;14(2):125-38 PMID: 23329113
  69. A synthetic-diploid benchmark for accurate variant-calling evaluation.
    Nat Methods. 2018 Aug;15(8):595-597 PMID: 30013044
  70. HLA diversity in the 1000 genomes dataset.
    PLoS One. 2014 Jul 02;9(7):e97282 PMID: 24988075
  71. Fast genotyping of known SNPs through approximate k-mer matching.
    Bioinformatics. 2016 Sep 1;32(17):i538-i544 PMID: 27587672
Article Info
Journal
Nature genetics
Abbr.
Nat Genet
ISSN
1546-1718
Published
2022-00-00
Epub
2022-00-11
Pages
518-525
Language
English
Region
United States
NLM ID
9216904
PMCID
PMC9005351
Subset
IM
Grants
NHGRI NIH HHS · U01 HG010973 · United States
NHGRI NIH HHS · R01 HG010169 · United States
NHGRI NIH HHS · U24 HG007497 · United States
NHGRI NIH HHS · R01 HG002385 · United States
Howard Hughes Medical Institute · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]