Home LiteratureArticle Details
PMID: 22125226 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, N.I.H., Intramural

Pitfalls of merging GWAS data: lessons learned in the eMERGE network and quality control procedures to maintain high data quality.

Genetic epidemiology ·Vol. 35 ·No. 8 ·2011-12-00 ·Pages 887-98

Zuvich RL, Armstrong LL, Bielinski SJ, Bradford Y, Carlson CS, Crawford DC, Crenshaw AT, de Andrade M, Doheny KF, Haines JL, Hayes MG, Jarvik GP, Jiang L, Kullo IJ, Li R, Ling H, Manolio TA, Matsumoto ME, McCarty CA, McDavid AN, Mirel DB, Olson LM, Paschall JE, Pugh EW, Rasmussen LV, Rasmussen-Torvik LJ, Turner SD, Wilke RA, Ritchie MD

Abstract

Genome-wide association studies (GWAS) are a useful approach in the study of the genetic components of complex phenotypes. Aside from large cohorts, GWAS have generally been limited to the study of one or a few diseases or traits. The emergence of biobanks linked to electronic medical records (EMRs) allows the efficient reuse of genetic data to yield meaningful genotype-phenotype associations for multiple phenotypes or traits. Phase I of the electronic MEdical Records and GEnomics (eMERGE-I) Network is a National Human Genome Research Institute-supported consortium composed of five sites to perform various genetic association studies using DNA repositories and EMR systems. Each eMERGE site has developed EMR-based algorithms to comprise a core set of 14 phenotypes for extraction of study samples from each site's DNA repository. Each eMERGE site selected samples for a specific phenotype, and these samples were genotyped at either the Broad Institute or at the Center for Inherited Disease Research using the Illumina Infinium BeadChip technology. In all, approximately 17,000 samples from across the five sites were genotyped. A unified quality control (QC) pipeline was developed by the eMERGE Genomics Working Group and used to ensure thorough cleaning of the data. This process includes examination of sample and marker quality and various batch effects. Upon completion of the genotyping and QC analyses for each site's primary study, eMERGE Coordinating Center merged the datasets from all five sites. This larger merged dataset reentered the established eMERGE QC pipeline. Based on lessons learned during the process, additional analyses and QC checkpoints were added to the pipeline to ensure proper merging. Here, we explore the challenges associated with combining datasets from different genotyping centers and describe the expansion to eMERGE QC pipeline for merged datasets. These additional steps will be useful as the eMERGE project expands to include additional sites in eMERGE-II, and also serve as a starting point for investigators merging multiple genotype datasets accessible through the National Center for Biotechnology Information in the database of Genotypes and Phenotypes. Our experience demonstrates that merging multiple datasets after additional QC can be an efficient use of genotype data despite new challenges that appear in the process.

MeSH Terms
Algorithms Electronic Health Records Genome-Wide Association Study/standards Genotype Humans National Human Genome Research Institute (U.S.) Phenotype Quality Control United States
Authors & Affiliations
29 authors, click to expand affiliations / ORCID
Zuvich Rebecca L
Center for Human Genetics Research, Department of Molecular Physiology and Biophysics, Vanderbilt University, Nashville, TN, USA.
Armstrong Loren L
Bielinski Suzette J
Bradford Yuki
Carlson Christopher S
Crawford Dana C
Crenshaw Andrew T
de Andrade Mariza
Doheny Kimberly F
Haines Jonathan L
Hayes M Geoffrey
Jarvik Gail P
Jiang Lan
Kullo Iftikhar J
Li Rongling
Ling Hua
Manolio Teri A
Matsumoto Martha E
McCarty Catherine A
McDavid Andrew N
Mirel Daniel B
Olson Lana M
Paschall Justin E
Pugh Elizabeth W
Rasmussen Luke V
Rasmussen-Torvik Laura J
Turner Stephen D
Wilke Russell A
Ritchie Marylyn D
References (28)
28 references, click to expand
  1. Complement factor H polymorphism in age-related macular degeneration.
    Science. 2005 Apr 15;308(5720):385-9 PMID: 15761122
  2. The NCBI dbGaP database of genotypes and phenotypes.
    Nat Genet. 2007 Oct;39(10):1181-6 PMID: 17898773
  3. Genomewide association studies and assessment of the risk of disease.
    N Engl J Med. 2010 Jul 8;363(2):166-76 PMID: 20647212
  4. The eMERGE Network: a consortium of biorepositories linked to electronic medical records data for conducting genomic studies.
    BMC Med Genomics. 2011 Jan 26;4:13 PMID: 21269473
  5. A genome-wide association study of red blood cell traits using the electronic medical record.
    PLoS One. 2010 Sep 28;5(9): PMID: 20927387
  6. Knowledge-driven multi-locus analysis reveals gene-gene interactions influencing HDL cholesterol level in two independent EMR-linked biobanks.
    PLoS One. 2011 May 11;6(5):e19586 PMID: 21589926
  7. A tutorial on statistical methods for population association studies.
    Nat Rev Genet. 2006 Oct;7(10):781-91 PMID: 16983374
  8. Editorial expression of concern.
    Science. 2010 Nov 12;330(6006):912 PMID: 21071647
  9. Are rare variants responsible for susceptibility to complex diseases?
    Am J Hum Genet. 2001 Jul;69(1):124-37 PMID: 11404818
  10. Biological, clinical and population relevance of 95 loci for blood lipids.
    Nature. 2010 Aug 5;466(7307):707-13 PMID: 20686565
  11. The Next PAGE in understanding complex traits: design for the analysis of Population Architecture Using Genetics and Epidemiology (PAGE) Study.
    Am J Epidemiol. 2011 Oct 1;174(7):849-59 PMID: 21836165
  12. Quality control and quality assurance in genotypic data for genome-wide association studies.
    Genet Epidemiol. 2010 Sep;34(6):591-602 PMID: 20718045
  13. Development of a large-scale de-identified DNA biobank to enable personalized medicine.
    Clin Pharmacol Ther. 2008 Sep;84(3):362-9 PMID: 18500243
  14. PLINK: a tool set for whole-genome association and population-based linkage analyses.
    Am J Hum Genet. 2007 Sep;81(3):559-75 PMID: 17701901
  15. Principal components analysis corrects for stratification in genome-wide association studies.
    Nat Genet. 2006 Aug;38(8):904-9 PMID: 16862161
  16. Assessing the accuracy of observer-reported ancestry in a biorepository linked to electronic medical records.
    Genet Med. 2010 Oct;12(10):648-50 PMID: 20733501
  17. Identification of genomic predictors of atrioventricular conduction: using electronic medical records as a tool for genome science.
    Circulation. 2010 Nov 16;122(20):2016-21 PMID: 21041692
  18. Leveraging informatics for genetic studies: use of the electronic medical record to enable a genome-wide association study of peripheral arterial disease.
    J Am Med Inform Assoc. 2010 Sep-Oct;17(5):568-74 PMID: 20819866
  19. On the allelic spectrum of human disease.
    Trends Genet. 2001 Sep;17(9):502-10 PMID: 11525833
  20. The allelic architecture of human disease genes: common disease-common variant...or not?
    Hum Mol Genet. 2002 Oct 1;11(20):2417-23 PMID: 12351577
  21. Quality control procedures for genome-wide association studies.
    Curr Protoc Hum Genet. 2011 Jan;Chapter 1:Unit1.19 PMID: 21234875
  22. Postassociation cleaning using linkage disequilibrium information.
    Genet Epidemiol. 2011 Jan;35(1):1-10 PMID: 21181893
  23. Inference of population structure using multilocus genotype data.
    Genetics. 2000 Jun;155(2):945-59 PMID: 10835412
  24. Spoiling the whole bunch: quality control aimed at preserving the integrity of high-throughput genotyping.
    Am J Hum Genet. 2010 Jul 9;87(1):123-8 PMID: 20598280
  25. Validity of racial/ethnic classifications in medical records data: an exploratory study.
    Am J Public Health. 2003 Jul;93(7):1084-6 PMID: 12835189
  26. Genetic signatures of exceptional longevity in humans.
    Science. 2010 Jul 1;2010: PMID: 20595579
  27. PhenX: a toolkit for interdisciplinary genetics research.
    Curr Opin Lipidol. 2010 Apr;21(2):136-40 PMID: 20154612
  28. Robust replication of genotype-phenotype associations across multiple diseases in an electronic medical record.
    Am J Hum Genet. 2010 Apr 9;86(4):560-72 PMID: 20362271
Article Info
Journal
Genetic epidemiology
Abbr.
Genet Epidemiol
ISSN
1098-2272
Published
2011-12-00
Pages
887-98
Language
English
Region
United States
NLM ID
8411723
PMCID
PMC3592376
Subset
IM
Grants
NHGRI NIH HHS · U01HG004438 · United States
NIGMS NIH HHS · T32 GM080178 · United States
NHGRI NIH HHS · U01HG004610 · United States
NHGRI NIH HHS · U01 HG004603 · United States
NHGRI NIH HHS · U01HG04603 · United States
NHGRI NIH HHS · U01HG004609 · United States
NHGRI NIH HHS · U01 HG004609 · United States
NHGRI NIH HHS · U01 HG004599 · United States
NHGRI NIH HHS · U01HG004608 · United States
NHGRI NIH HHS · U01 HG004608 · United States
NHGRI NIH HHS · U01 HG006375 · United States
NHGRI NIH HHS · U01 HG004438 · United States
NHGRI NIH HHS · U01HG04599 · United States
Intramural NIH HHS · United States
NLM NIH HHS · R01 LM010040 · United States
NHGRI NIH HHS · U01 HG004610 · United States
NLM NIH HHS · R01LM010040 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]