Abstract
In a mouse intercross with more than 500 animals and genome-wide gene expression data on six tissues, we identified a high proportion (18%) of sample mix-ups in the genotype data. Local expression quantitative trait loci (eQTL; genetic loci influencing gene expression) with extremely large effect were used to form a classifier to predict an individual's eQTL genotype based on expression data alone. By considering multiple eQTL and their related transcripts, we identified numerous individuals whose predicted eQTL genotypes (based on their expression data) did not match their observed genotypes, and then went on to identify other individuals whose genotypes did match the predicted eQTL genotypes. The concordance of predictions across six tissues indicated that the problem was due to mix-ups in the genotypes (although we further identified a small number of sample mix-ups in each of the six panels of gene expression microarrays). Consideration of the plate positions of the DNA samples indicated a number of off-by-one and off-by-two errors, likely the result of pipetting errors. Such sample mix-ups can be a problem in any genetic study, but eQTL data allow us to identify, and even correct, such problems. Our methods have been implemented in an R package, R/lineup.
Keywords
eQTL
genetical genomics
microarrays
mislabeling errors
quality control
MeSH Terms
Animals
Chromosome Mapping
Computational Biology/methods
Gene Expression
Gene Expression Profiling
Genome-Wide Association Study
Genomics/methods
Genotype
Lod Score
Mice
Phenotype
Quantitative Trait Loci
Transcriptome
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Broman Karl W
ORCID
Department of Biostatistics and Medical Informatics, University of Wisconsin, Madison, Wisconsin 53706
[email protected].
Keller Mark P
Department of Biochemistry, University of Wisconsin, Madison, Wisconsin 53706.
Broman Aimee Teo
ORCID
Department of Biostatistics and Medical Informatics, University of Wisconsin, Madison, Wisconsin 53706.
Kendziorski Christina
Department of Biostatistics and Medical Informatics, University of Wisconsin, Madison, Wisconsin 53706.
Yandell Brian S
ORCID
Department of Statistics, University of Wisconsin, Madison, Wisconsin 53706 Department of Horticulture, University of Wisconsin, Madison, Wisconsin 53706.
Sen Śaunak
ORCID
Department of Epidemiology and Biostatistics, University of California, San Francisco, California 94107.
Attie Alan D
Department of Biochemistry, University of Wisconsin, Madison, Wisconsin 53706.
References (22)
22 references, click to expand
-
Mapping quantitative trait loci in the case of a spike in the phenotype distribution.
Genetics. 2003 Mar;163(3):1169-75
PMID: 12663553
-
Run batch effects potentially compromise the usefulness of genomic signatures for ovarian cancer.
J Clin Oncol. 2008 Mar 1;26(7):1186-7; author reply 1187-8
PMID: 18309960
-
Maintaining data integrity in microarray data management.
Biotechnol Bioeng. 2003 Dec 30;84(7):795-800
PMID: 14708120
-
Mapping mendelian factors underlying quantitative traits using RFLP linkage maps.
Genetics. 1989 Jan;121(1):185-99
PMID: 2563713
-
Mapping quantitative trait loci for complex binary diseases using line crosses.
Genetics. 1996 Jul;143(3):1417-24
PMID: 8807312
-
A gene expression network model of type 2 diabetes links cell cycle regulation in islets with diabetes susceptibility.
Genome Res. 2008 May;18(5):706-16
PMID: 18347327
-
Genetic analysis of human traits in vitro: drug response and gene expression in lymphoblastoid cell lines.
PLoS Genet. 2008 Nov;4(11):e1000287
PMID: 19043577
-
Methods for labeling error detection in microarrays based on the effect of data perturbation on the regression model.
Bioinformatics. 2009 Oct 15;25(20):2708-14
PMID: 19661242
-
Utilization of AFFX spike-in control probes to monitor sample identity throughout Affymetrix GeneChip Array processing.
Biotechniques. 2010 May;48(5):371-8
PMID: 20569210
-
MixupMapper: correcting sample mix-ups in genome-wide datasets increases power to detect small genetic effects.
Bioinformatics. 2011 Aug 1;27(15):2104-11
PMID: 21653519
-
Detecting sample misidentifications in genetic association studies.
Stat Appl Genet Mol Biol. 2012;11(3):Article 13
PMID: 22611595
-
Bayesian method to predict individual SNP genotypes from gene expression data.
Nat Genet. 2012 May;44(5):603-8
PMID: 22484626
-
Calling sample mix-ups in cancer population studies.
PLoS One. 2012;7(8):e41815
PMID: 22912679
-
Detecting and estimating contamination of human DNA samples in sequencing and array-based genotype data.
Am J Hum Genet. 2012 Nov 2;91(5):839-48
PMID: 23103226
-
Classification of mislabelled microarrays using robust sparse logistic regression.
Bioinformatics. 2013 Apr 1;29(7):870-7
PMID: 23418189
-
Stocks for detecting linkage in the mouse, and the theory of their design.
J Genet. 1951 Jan;50(2):307-23
PMID: 24539711
-
'The 39 steps' in gene expression profiling: critical issues and proposed best practices for microarray experiments.
Drug Discov Today. 2005 Sep 1;10(17):1175-82
PMID: 16182210
-
A simple regression method for mapping quantitative trait loci in line crosses using flanking markers.
Heredity (Edinb). 1992 Oct;69(4):315-24
PMID: 16718932
-
Single nucleotide polymorphism profiling assay to exclude serum sample mix-up.
Vox Sang. 2007 Feb;92(2):148-53
PMID: 17298578
-
Single nucleotide polymorphism profiling assay to confirm the identity of human tissues.
J Mol Diagn. 2007 Apr;9(2):205-13
PMID: 17384212
-
A gene expression bar code for microarray data.
Nat Methods. 2007 Nov;4(11):911-3
PMID: 17906632
-
R/qtl: QTL mapping in experimental crosses.
Bioinformatics. 2003 May 1;19(7):889-90
PMID: 12724300