Home LiteratureArticle Details
PMID: 23103226 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

Detecting and estimating contamination of human DNA samples in sequencing and array-based genotype data.

American journal of human genetics ·Vol. 91 ·No. 5 ·2012-11-02 ·Pages 839-48

Jun G, Flickinger M, Hetrick KN, Romm JM, Doheny KF, Abecasis GR, Boehnke M, Kang HM

Abstract

DNA sample contamination is a serious problem in DNA sequencing studies and may result in systematic genotype misclassification and false positive associations. Although methods exist to detect and filter out cross-species contamination, few methods to detect within-species sample contamination are available. In this paper, we describe methods to identify within-species DNA sample contamination based on (1) a combination of sequencing reads and array-based genotype data, (2) sequence reads alone, and (3) array-based genotype data alone. Analysis of sequencing reads allows contamination detection after sequence data is generated but prior to variant calling; analysis of array-based genotype data allows contamination detection prior to generation of costly sequence data. Through a combination of analysis of in silico and experimentally contaminated samples, we show that our methods can reliably detect and estimate levels of contamination as low as 1%. We evaluate the impact of DNA contamination on genotype accuracy and propose effective strategies to screen for and prevent DNA contamination in sequencing studies.

MeSH Terms
DNA Contamination Diabetes Mellitus, Type 2/diagnosis,genetics Genotype Humans Sequence Analysis, DNA
Authors & Affiliations
8 authors, click to expand affiliations / ORCID
Jun Goo
Department of Biostatistics and Center for Statistical Genetics, School of Public Health, University of Michigan, Ann Arbor, MI 48109, USA.
Flickinger Matthew
Hetrick Kurt N
Romm Jane M
Doheny Kimberly F
Abecasis Gonçalo R
Boehnke Michael
Kang Hyun Min
References (10)
10 references, click to expand
  1. A map of human genome variation from population-scale sequencing.
    Nature. 2010 Oct 28;467(7319):1061-73 PMID: 20981092
  2. Increasing power for tests of genetic association in the presence of phenotype and/or genotype error by use of double-sampling.
    Stat Appl Genet Mol Biol. 2004;3:Article26 PMID: 16646805
  3. The metabochip, a custom genotyping array for genetic studies of metabolic, cardiovascular, and anthropometric traits.
    PLoS Genet. 2012;8(8):e1002793 PMID: 22876189
  4. Integrated genotype calling and association analysis of SNPs, common copy number polymorphisms and rare CNVs.
    Nat Genet. 2008 Oct;40(10):1253-60 PMID: 18776909
  5. Low-coverage sequencing: implications for design of complex trait association studies.
    Genome Res. 2011 Jun;21(6):940-51 PMID: 21460063
  6. Simultaneous genotype calling and haplotype phasing improves genotype accuracy and reduces false-positive associations for genome-wide association studies.
    Am J Hum Genet. 2009 Dec;85(6):847-61 PMID: 19931040
  7. ContEst: estimating cross-contamination of human samples in next-generation sequencing data.
    Bioinformatics. 2011 Sep 15;27(18):2601-2 PMID: 21803805
  8. PennCNV: an integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data.
    Genome Res. 2007 Nov;17(11):1665-74 PMID: 17921354
  9. GenoSNP: a variational Bayes within-sample SNP genotyping algorithm that does not require a reference population.
    Bioinformatics. 2008 Oct 1;24(19):2209-14 PMID: 18653518
  10. Fast identification and removal of sequence contamination from genomic and metagenomic datasets.
    PLoS One. 2011 Mar 09;6(3):e17288 PMID: 21408061
Article Info
Journal
American journal of human genetics
Abbr.
Am J Hum Genet
ISSN
1537-6605
Published
2012-11-02
Epub
2012-00-25
Pages
839-48
Language
English
Region
United States
NLM ID
0370475
PMCID
PMC3487130
Subset
IM
Grants
NIMH NIH HHS · MH084698 · United States
NIMH NIH HHS · R01 MH084698 · United States
NHGRI NIH HHS · T32 HG000040 · United States
NHGRI NIH HHS · R56 HG000376 · United States
NIDDK NIH HHS · R13 DK088398 · United States
NHGRI NIH HHS · HHSN268200782096C · United States
NHGRI NIH HHS · HG006513 · United States
NHGRI NIH HHS · HG005214 · United States
NIDDK NIH HHS · P30 DK020572 · United States
NHGRI NIH HHS · R01 HG007022 · United States
NHLBI NIH HHS · R01 HL117626 · United States
NHGRI NIH HHS · HG000376 · United States
NIDDK NIH HHS · DK088398 · United States
NHGRI NIH HHS · U01 HG006513 · United States
NHGRI NIH HHS · U01 HG005214 · United States
NHLBI NIH HHS · HHSN268201100011I · United States
NHLBI NIH HHS · HHSN268201100011C · United States
NHGRI NIH HHS · R01 HG000376 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]