Abstract
Determination of copy number variants (CNVs) inferred in genome wide single nucleotide polymorphism arrays has shown increasing utility in genetic variant disease associations. Several CNV detection methods are available, but differences in CNV call thresholds and characteristics exist. We evaluated the relative performance of seven methods: circular binary segmentation, CNVFinder, cnvPartition, gain and loss of DNA, Nexus algorithms, PennCNV and QuantiSNP. Tested data included real and simulated Illumina HumHap 550 data from the Singapore cohort study of the risk factors for Myopia (SCORM) and simulated data from Affymetrix 6.0 and platform-independent distributions. The normalized singleton ratio (NSR) is proposed as a metric for parameter optimization before enacting full analysis. We used 10 SCORM samples for optimizing parameter settings for each method and then evaluated method performance at optimal parameters using 100 SCORM samples. The statistical power, false positive rates, and receiver operating characteristic (ROC) curve residuals were evaluated by simulation studies. Optimal parameters, as determined by NSR and ROC curve residuals, were consistent across datasets. QuantiSNP outperformed other methods based on ROC curve residuals over most datasets. Nexus Rank and SNPRank have low specificity and high power. Nexus Rank calls oversized CNVs. PennCNV detects one of the fewest numbers of CNVs.
MeSH Terms
Algorithms
Computer Simulation
DNA Copy Number Variations
Female
Humans
Male
Myopia/genetics
Oligonucleotide Array Sequence Analysis
Polymorphism, Single Nucleotide
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Dellinger Andrew E
Center for Human Genetics, Duke University Medical Center, Durham, NC 27710, USA.
Saw Seang-Mei
Goh Liang K
Seielstad Mark
Young Terri L
Li Yi-Ju
References (18)
18 references, click to expand
-
High resolution array-CGH analysis of single cells.
Nucleic Acids Res. 2007;35(3):e15
PMID: 17178751
-
QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
Nucleic Acids Res. 2007;35(6):2013-25
PMID: 17341461
-
Copy-number variation in sporadic amyotrophic lateral sclerosis: a genome-wide screen.
Lancet Neurol. 2008 Apr;7(4):319-26
PMID: 18313986
-
Accurate and reliable high-throughput detection of copy number variation in the human genome.
Genome Res. 2006 Dec;16(12):1566-74
PMID: 17122085
-
Structural variation in the human genome.
Nat Rev Genet. 2006 Feb;7(2):85-97
PMID: 16418744
-
Population analysis of large copy number variants and hotspots of human genetic disease.
Am J Hum Genet. 2009 Feb;84(2):148-61
PMID: 19166990
-
Circular binary segmentation for the analysis of array-based DNA copy number data.
Biostatistics. 2004 Oct;5(4):557-72
PMID: 15475419
-
NCBI GEO: archive for high-throughput functional genomic data.
Nucleic Acids Res. 2009 Jan;37(Database issue):D885-90
PMID: 18940857
-
Detection of large-scale variation in the human genome.
Nat Genet. 2004 Sep;36(9):949-51
PMID: 15286789
-
Genotype, haplotype and copy-number variation in worldwide human populations.
Nature. 2008 Feb 21;451(7181):998-1003
PMID: 18288195
-
Recurrent CNVs disrupt three candidate genes in schizophrenia patients.
Am J Hum Genet. 2008 Oct;83(4):504-10
PMID: 18940311
-
Global variation in copy number in the human genome.
Nature. 2006 Nov 23;444(7118):444-54
PMID: 17122850
-
Analysis of array CGH data: from signal ratio to gain and loss of DNA regions.
Bioinformatics. 2004 Dec 12;20(18):3413-22
PMID: 15381628
-
High-resolution mapping of copy-number alterations with massively parallel sequencing.
Nat Methods. 2009 Jan;6(1):99-103
PMID: 19043412
-
Autism genome-wide copy number variation reveals ubiquitin and neuronal genes.
Nature. 2009 May 28;459(7246):569-73
PMID: 19404257
-
Impact of whole genome amplification on analysis of copy number variants.
Nucleic Acids Res. 2008 Aug;36(13):e80
PMID: 18559357
-
Comparative analysis of algorithms for identifying amplifications and deletions in array CGH data.
Bioinformatics. 2005 Oct 1;21(19):3763-70
PMID: 16081473
-
PennCNV: an integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data.
Genome Res. 2007 Nov;17(11):1665-74
PMID: 17921354