Home LiteratureArticle Details
PMID: 36055201 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios.

Cell ·Vol. 185 ·No. 18 ·2022-00-01 ·Pages 3426-3440.e19

Byrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, Corvelo A, Clarke WE, Musunuri R, Nagulapalli K, Fairley S, Runnels A, Winterkorn L, Lowy E, Human Genome Structural Variation Consortium, Paul Flicek, Germer S, Brand H, Hall IM, Talkowski ME, Narzisi G, Zody MC

Abstract

The 1000 Genomes Project (1kGP) is the largest fully open resource of whole-genome sequencing (WGS) data consented for public distribution without access or use restrictions. The final, phase 3 release of the 1kGP included 2,504 unrelated samples from 26 populations and was based primarily on low-coverage WGS. Here, we present a high-coverage 3,202-sample WGS 1kGP resource, which now includes 602 complete trios, sequenced to a depth of 30X using Illumina. We performed single-nucleotide variant (SNV) and short insertion and deletion (INDEL) discovery and generated a comprehensive set of structural variants (SVs) by integrating multiple analytic methods through a machine learning model. We show gains in sensitivity and precision of variant calls compared to phase 3, especially among rare SNVs as well as INDELs and SVs spanning frequency spectrum. We also generated an improved reference imputation panel, making variants discovered here accessible for association studies.

Keywords
1000 Genomes Project INDEL SNV population genetics reference imputation panel structural variation trio sequencing whole-genome sequencing
MeSH Terms
Female Genome, Human High-Throughput Nucleotide Sequencing/methods Humans INDEL Mutation Male Polymorphism, Single Nucleotide Whole Genome Sequencing
Authors & Affiliations
22 authors, click to expand affiliations / ORCID
Byrska-Bishop Marta
New York Genome Center, New York, NY 10013, USA. Electronic address: [email protected].
Evani Uday S
New York Genome Center, New York, NY 10013, USA.
Zhao Xuefang
Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA 02142, USA; Center for Genomic Medicine, Massachusetts General Hospital, Boston, MA 02114, USA; Department of Neurology, Massachusetts General Hospital and Harvard Medical School, Boston, MA 02114, USA.
Basile Anna O
New York Genome Center, New York, NY 10013, USA.
Abel Haley J
McDonnell Genome Institute, Washington University School of Medicine, St. Louis, MO 63108, USA; Department of Medicine, Washington University School of Medicine, St. Louis, MO 63110, USA.
Regier Allison A
McDonnell Genome Institute, Washington University School of Medicine, St. Louis, MO 63108, USA; Department of Medicine, Washington University School of Medicine, St. Louis, MO 63110, USA.
Corvelo André
New York Genome Center, New York, NY 10013, USA.
Clarke Wayne E
New York Genome Center, New York, NY 10013, USA; Outlier Informatics Inc., Saskatoon, SK S7H 1L4, Canada.
Musunuri Rajeeva
New York Genome Center, New York, NY 10013, USA.
Nagulapalli Kshithija
New York Genome Center, New York, NY 10013, USA.
Fairley Susan
European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, UK.
Runnels Alexi
New York Genome Center, New York, NY 10013, USA.
Winterkorn Lara
New York Genome Center, New York, NY 10013, USA.
Lowy Ernesto
European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, UK.
Human Genome Structural Variation Consortium
Paul Flicek
European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, UK.
Germer Soren
New York Genome Center, New York, NY 10013, USA.
Brand Harrison
Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA 02142, USA; Center for Genomic Medicine, Massachusetts General Hospital, Boston, MA 02114, USA; Department of Neurology, Massachusetts General Hospital and Harvard Medical School, Boston, MA 02114, USA; Stanley Center for Psychiatric Research, Broad Institute of MIT and Harvard, Cambridge, MA 02142, USA.
Hall Ira M
McDonnell Genome Institute, Washington University School of Medicine, St. Louis, MO 63108, USA; Department of Medicine, Washington University School of Medicine, St. Louis, MO 63110, USA; Center for Genomic Health, Yale University School of Medicine, New Haven, CT 06510, USA; Department of Genetics, Yale University School of Medicine, New Haven, CT 06520, USA.
Talkowski Michael E
Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA 02142, USA; Center for Genomic Medicine, Massachusetts General Hospital, Boston, MA 02114, USA; Department of Neurology, Massachusetts General Hospital and Harvard Medical School, Boston, MA 02114, USA; Stanley Center for Psychiatric Research, Broad Institute of MIT and Harvard, Cambridge, MA 02142, USA.
Narzisi Giuseppe
New York Genome Center, New York, NY 10013, USA.
Zody Michael C
New York Genome Center, New York, NY 10013, USA. Electronic address: [email protected].
Investigators
21 investigators, click to expand
Eichler Evan E
Korbel Jan O
Lee Charles
Marschall Tobias
Devine Scott E
Harvey William T
Zhou Weichen
Mills Ryan E
Rausch Tobias
Kumar Sushant
Alkan Can
Hormozdiari Fereydoun
Chong Zechen
Chen Yu
Yang Xiaofei
Lin Jiadong
Gerstein Mark B
Kai Ye
Zhu Qihui
Yilmaz Feyza
Xiao Chunlin
Conflict of Interest

Declaration of interests E.E.E. is a scientific advisory board (SAB) member of Variant Bio, Inc. P.F. is an SAB member of Fabric Genomics, Inc., and Eagle Genomics, Ltd.

References (72)
72 references, click to expand
  1. Wham: Identifying Structural Variants of Biological Consequence.
    PLoS Comput Biol. 2015 Dec 01;11(12):e1004572 PMID: 26625158
  2. A reference panel of 64,976 haplotypes for genotype imputation.
    Nat Genet. 2016 Oct;48(10):1279-83 PMID: 27548312
  3. Applications of the 1000 Genomes Project resources.
    Brief Funct Genomics. 2017 May 1;16(3):163-170 PMID: 27436001
  4. An open resource for accurately benchmarking small variant and reference calls.
    Nat Biotechnol. 2019 May;37(5):561-566 PMID: 30936564
  5. eQTL mapping identifies insertion- and deletion-specific eQTLs in multiple tissues.
    Nat Commun. 2015 May 08;6:6821 PMID: 25951796
  6. Paragraph: a graph-based structural variant genotyper for short-read sequence data.
    Genome Biol. 2019 Dec 19;20(1):291 PMID: 31856913
  7. svtools: population-scale analysis of structural variation.
    Bioinformatics. 2019 Nov 1;35(22):4782-4787 PMID: 31218349
  8. Parental influence on human germline de novo mutations in 1,548 trios from Iceland.
    Nature. 2017 Sep 28;549(7673):519-522 PMID: 28959963
  9. An integrated map of genetic variation from 1,092 human genomes.
    Nature. 2012 Nov 1;491(7422):56-65 PMID: 23128226
  10. Benchmarking challenging small variants with linked and long reads.
    Cell Genom. 2022 May;2(5): PMID: 36452119
  11. A map of human genome variation from population-scale sequencing.
    Nature. 2010 Oct 28;467(7319):1061-73 PMID: 20981092
  12. An integrated map of structural variation in 2,504 human genomes.
    Nature. 2015 Oct 1;526(7571):75-81 PMID: 26432246
  13. MAFFT multiple sequence alignment software version 7: improvements in performance and usability.
    Mol Biol Evol. 2013 Apr;30(4):772-80 PMID: 23329690
  14. de novo variant calling identifies cancer mutation signatures in the 1000 Genomes Project.
    Hum Mutat. 2022 Dec;43(12):1979-1993 PMID: 36054329
  15. Profiling the genome-wide landscape of tandem repeat expansions.
    Nucleic Acids Res. 2019 Sep 5;47(15):e90 PMID: 31194863
  16. Genomic Patterns of De Novo Mutation in Simplex Autism.
    Cell. 2017 Oct 19;171(3):710-722.e12 PMID: 28965761
  17. African genetic diversity: implications for human demographic history, modern human origins, and complex disease mapping.
    Annu Rev Genomics Hum Genet. 2008;9:403-33 PMID: 18593304
  18. The origin, evolution, and functional impact of short insertion-deletion variants identified in 179 human genomes.
    Genome Res. 2013 May;23(5):749-61 PMID: 23478400
  19. Discovery and Fine-Mapping of Glycaemic and Obesity-Related Trait Loci Using High-Density Imputation.
    PLoS Genet. 2015 Jul 01;11(7):e1005230 PMID: 26132169
  20. The Ensembl Variant Effect Predictor.
    Genome Biol. 2016 Jun 06;17(1):122 PMID: 27268795
  21. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data.
    Genome Res. 2010 Sep;20(9):1297-303 PMID: 20644199
  22. Functional annotation of noncoding sequence variants.
    Nat Methods. 2014 Mar;11(3):294-6 PMID: 24487584
  23. The Mobile Element Locator Tool (MELT): population-scale mobile element discovery and biology.
    Genome Res. 2017 Nov;27(11):1916-1929 PMID: 28855259
  24. Integrating common and rare genetic variation in diverse human populations.
    Nature. 2010 Sep 2;467(7311):52-8 PMID: 20811451
  25. The mutational constraint spectrum quantified from variation in 141,456 humans.
    Nature. 2020 May;581(7809):434-443 PMID: 32461654
  26. Second-generation PLINK: rising to the challenge of larger and richer datasets.
    Gigascience. 2015 Feb 25;4:7 PMID: 25722852
  27. A recurrence-based approach for validating structural variation using long-read sequencing technology.
    Gigascience. 2017 Aug 1;6(8):1-9 PMID: 28873962
  28. The sequences of 150,119 genomes in the UK Biobank.
    Nature. 2022 Jul;607(7920):732-740 PMID: 35859178
  29. A reference data set of 5.4 million phased human variants validated by genetic inheritance from sequencing a three-generation 17-member pedigree.
    Genome Res. 2017 Jan;27(1):157-164 PMID: 27903644
  30. A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.
    PLoS Genet. 2009 Jun;5(6):e1000529 PMID: 19543373
  31. A general approach for haplotype phasing across the full spectrum of relatedness.
    PLoS Genet. 2014 Apr 17;10(4):e1004234 PMID: 24743097
  32. Twelve years of SAMtools and BCFtools.
    Gigascience. 2021 Feb 16;10(2): PMID: 33590861
  33. LUMPY: a probabilistic framework for structural variant discovery.
    Genome Biol. 2014 Jun 26;15(6):R84 PMID: 24970577
  34. cn.MOPS: mixture of Poissons for discovering copy number variations in next-generation sequencing data with a low false discovery rate.
    Nucleic Acids Res. 2012 May;40(9):e69 PMID: 22302147
  35. The Simons Genome Diversity Project: 300 genomes from 142 diverse populations.
    Nature. 2016 Oct 13;538(7624):201-206 PMID: 27654912
  36. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications.
    Bioinformatics. 2016 Apr 15;32(8):1220-2 PMID: 26647377
  37. A structural variation reference for medical and population genetics.
    Nature. 2020 May;581(7809):444-451 PMID: 32461652
  38. Functional equivalence of genome sequencing analysis pipelines enables harmonized variant calling across human genetics projects.
    Nat Commun. 2018 Oct 2;9(1):4038 PMID: 30279509
  39. Genome-wide association study identifies three novel loci for type 2 diabetes.
    Hum Mol Genet. 2014 Jan 1;23(1):239-46 PMID: 23945395
  40. Expectations and blind spots for structural variation detection from long-read assemblies and short-read genome sequencing technologies.
    Am J Hum Genet. 2021 May 6;108(5):919-928 PMID: 33789087
  41. Integrative annotation of variants from 1092 humans: application to cancer genomics.
    Science. 2013 Oct 04;342(6154):1235587 PMID: 24092746
  42. Detecting and estimating contamination of human DNA samples in sequencing and array-based genotype data.
    Am J Hum Genet. 2012 Nov 2;91(5):839-48 PMID: 23103226
  43. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data.
    Bioinformatics. 2011 Nov 1;27(21):2987-93 PMID: 21903627
  44. The variant call format and VCFtools.
    Bioinformatics. 2011 Aug 1;27(15):2156-8 PMID: 21653522
  45. A note on exact tests of Hardy-Weinberg equilibrium.
    Am J Hum Genet. 2005 May;76(5):887-93 PMID: 15789306
  46. Best practices for benchmarking germline small-variant calls in human genomes.
    Nat Biotechnol. 2019 May;37(5):555-560 PMID: 30858580
  47. Transcriptome and genome sequencing uncovers functional variation in humans.
    Nature. 2013 Sep 26;501(7468):506-11 PMID: 24037378
  48. Rate of de novo mutations and the importance of father's age to disease risk.
    Nature. 2012 Aug 23;488(7412):471-5 PMID: 22914163
  49. A comprehensive 1,000 Genomes-based genome-wide association meta-analysis of coronary artery disease.
    Nat Genet. 2015 Oct;47(10):1121-1130 PMID: 26343387
  50. A general framework for estimating the relative pathogenicity of human genetic variants.
    Nat Genet. 2014 Mar;46(3):310-5 PMID: 24487276
  51. A global reference for human genetic variation.
    Nature. 2015 Oct 1;526(7571):68-74 PMID: 26432245
  52. Haplotype-resolved diverse human genomes and integrated analysis of structural variation.
    Science. 2021 Apr 2;372(6537): PMID: 33632895
  53. CNVnator: an approach to discover, genotype, and characterize typical and atypical CNVs from family and population genome sequencing.
    Genome Res. 2011 Jun;21(6):974-84 PMID: 21324876
  54. Reference-based phasing using the Haplotype Reference Consortium panel.
    Nat Genet. 2016 Nov;48(11):1443-1448 PMID: 27694958
  55. STRetch: detecting and discovering pathogenic short tandem repeat expansions.
    Genome Biol. 2018 Aug 21;19(1):121 PMID: 30129428
  56. Mapping and characterization of structural variation in 17,795 human genomes.
    Nature. 2020 Jul;583(7814):83-89 PMID: 32460305
  57. Multi-platform discovery of haplotype-resolved structural variation in human genomes.
    Nat Commun. 2019 Apr 16;10(1):1784 PMID: 30992455
  58. The International Genome Sample Resource (IGSR) collection of open human genomic variation resources.
    Nucleic Acids Res. 2020 Jan 8;48(D1):D941-D947 PMID: 31584097
  59. dbSNP-database for single nucleotide polymorphisms and other classes of minor genetic variation.
    Genome Res. 1999 Aug;9(8):677-9 PMID: 10447503
  60. SpeedSeq: ultra-fast personal genome analysis and interpretation.
    Nat Methods. 2015 Oct;12(10):966-8 PMID: 26258291
  61. An analytical framework for whole-genome sequence association studies and its implications for autism spectrum disorder.
    Nat Genet. 2018 May;50(5):727-736 PMID: 29700473
  62. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
  63. CrossMap: a versatile tool for coordinate conversion between genome assemblies.
    Bioinformatics. 2014 Apr 1;30(7):1006-7 PMID: 24351709
  64. Pangenome-based genome inference allows efficient and accurate genotyping across a wide spectrum of variant classes.
    Nat Genet. 2022 Apr;54(4):518-525 PMID: 35410384
  65. Robust relationship inference in genome-wide association studies.
    Bioinformatics. 2010 Nov 15;26(22):2867-73 PMID: 20926424
  66. Deep sequencing of 10,000 human genomes.
    Proc Natl Acad Sci U S A. 2016 Oct 18;113(42):11901-11906 PMID: 27702888
  67. BEDTools: a flexible suite of utilities for comparing genomic features.
    Bioinformatics. 2010 Mar 15;26(6):841-2 PMID: 20110278
  68. A linear complexity phasing method for thousands of genomes.
    Nat Methods. 2011 Dec 04;9(2):179-81 PMID: 22138821
  69. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program.
    Nature. 2021 Feb;590(7845):290-299 PMID: 33568819
  70. ExpansionHunter: a sequence-graph-based tool to analyze variation in short tandem repeat regions.
    Bioinformatics. 2019 Nov 1;35(22):4754-4756 PMID: 31134279
  71. Fine mapping of the celiac disease-associated LPP locus reveals a potential functional variant.
    Hum Mol Genet. 2014 May 1;23(9):2481-9 PMID: 24334606
  72. Accurate, scalable and integrative haplotype estimation.
    Nat Commun. 2019 Nov 28;10(1):5436 PMID: 31780650
Article Info
Journal
Cell
Abbr.
Cell
ISSN
1097-4172
Published
2022-00-01
Pages
3426-3440.e19
Language
English
Region
United States
NLM ID
0413066
PMCID
PMC9439720
Subset
IM
Grants
NHGRI NIH HHS · R01 HG002898 · United States
NICHD NIH HHS · R03 HD099547 · United States
NIGMS NIH HHS · R35 GM138212 · United States
NHGRI NIH HHS · UM1 HG008895 · United States
NHGRI NIH HHS · UM1 HG008901 · United States
NICHD NIH HHS · R01 HD081256 · United States
NIMH NIH HHS · R56 MH115957 · United States
Wellcome Trust · United Kingdom
NHGRI NIH HHS · U24 HG007497 · United States
NCI NIH HHS · R21 CA259309 · United States
NIMH NIH HHS · R01 MH115957 · United States
NCI NIH HHS · R01 CA261934 · United States
NHGRI NIH HHS · UM1 HG008853 · United States
Corrections
CommentIn
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]