Home LiteratureArticle Details
PMID: 27746907 Published · epublish English Journal Article

Whose sample is it anyway? Widespread misannotation of samples in transcriptomics studies.

F1000Research ·Vol. 5 ·2016-00-00 ·Pages 2103

Toker L, Feng M, Pavlidis P

Abstract

Concern about the reproducibility and reliability of biomedical research has been rising. An understudied issue is the prevalence of sample mislabeling, one impact of which would be invalid comparisons. We studied this issue in a corpus of human transcriptomics studies by comparing the provided annotations of sex to the expression levels of sex-specific genes. We identified apparent mislabeled samples in 46% of the datasets studied, yielding a 99% confidence lower-bound estimate for all studies of 33%. In a separate analysis of a set of datasets concerning a single cohort of subjects, 2/4 had mislabeled samples, indicating laboratory mix-ups rather than data recording errors. While the number of mixed-up samples per study was generally small, because our method can only identify a subset of potential mix-ups, our estimate is conservative for the breadth of the problem. Our findings emphasize the need for more stringent sample tracking, and that re-users of published data must be alert to the possibility of annotation and labelling errors.

Keywords
Transcriptomics data quality gene expression misannotation mislabeling reproducibility
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Toker Lilah
Department of Psychiatry, University of British Columbia, Vancouver, V6T 2A1, Canada; Michael Smith Laboratories, University of British Columbia, Vancouver, V6T 1Z4, Canada.
Feng Min
Department of Psychiatry, University of British Columbia, Vancouver, V6T 2A1, Canada; Michael Smith Laboratories, University of British Columbia, Vancouver, V6T 1Z4, Canada; Graduate Program in Genome Sciences and Technology, University of British Columbia, Vancouver, V5Z 4S6, Canada.
Pavlidis Paul
Department of Psychiatry, University of British Columbia, Vancouver, V6T 2A1, Canada; Michael Smith Laboratories, University of British Columbia, Vancouver, V6T 1Z4, Canada.
References (17)
17 references, click to expand
  1. Expression and function of a large non-coding RNA gene XIST in human cancer.
    World J Surg. 2011 Aug;35(8):1751-6 PMID: 21212949
  2. A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium.
    Nat Biotechnol. 2014 Sep;32(9):903-14 PMID: 25150838
  3. Reproducibility in science: improving the standard for basic and preclinical research.
    Circ Res. 2015 Jan 2;116(1):116-26 PMID: 25552691
  4. How common is intersex? a response to Anne Fausto-Sterling.
    J Sex Res. 2002 Aug;39(3):174-8 PMID: 12476264
  5. Bioconductor: open software development for computational biology and bioinformatics.
    Genome Biol. 2004;5(10):R80 PMID: 15461798
  6. Gemma: a resource for the reuse, sharing and meta-analysis of expression profiling data.
    Bioinformatics. 2012 Sep 1;28(17):2272-3 PMID: 22782548
  7. Cost-effective prediction of gender-labeling errors and estimation of gender-labeling error rates in candidate-gene association studies.
    Front Genet. 2011 Jun 15;2:31 PMID: 22303327
  8. The Doppelgänger Effect: Hidden Duplicates in Databases of Transcriptome Profiles.
    J Natl Cancer Inst. 2016 Jul 05;108(11): PMID: 27381624
  9. Amelogenin-based sex identification as a strategy to control the identity of DNA samples in genetic association studies.
    Pharmacogenomics. 2010 Mar;11(3):449-57 PMID: 20235797
  10. arrayQualityMetrics--a bioconductor package for quality assessment of microarray data.
    Bioinformatics. 2009 Feb 1;25(3):415-6 PMID: 19106121
  11. Orchestrating high-throughput genomic analysis with Bioconductor.
    Nat Methods. 2015 Feb;12(2):115-21 PMID: 25633503
  12. Network-based metaanalysis identifies HNF4A and PTBP1 as longitudinally dynamic biomarkers for Parkinson's disease.
    Proc Natl Acad Sci U S A. 2015 Feb 17;112(7):2257-62 PMID: 25646437
  13. Tackling the widespread and critical impact of batch effects in high-throughput data.
    Nat Rev Genet. 2010 Oct;11(10 ):733-9 PMID: 20838408
  14. Metaanalysis of flawed expression profiling data leading to erroneous Parkinson's biomarker identification.
    Proc Natl Acad Sci U S A. 2015 Jul 14;112(28):E3637 PMID: 26106167
  15. Reproducibility: A tragedy of errors.
    Nature. 2016 Feb 4;530(7588):27-9 PMID: 26842041
  16. Identification of sample annotation errors in gene expression datasets.
    Arch Toxicol. 2015 Dec;89(12):2265-72 PMID: 26608184
  17. NCBI GEO standards and services for microarray data.
    Nat Biotechnol. 2006 Dec;24(12):1471-2 PMID: 17160034
Article Info
Journal
F1000Research
Abbr.
F1000Res
ISSN
2046-1402
Published
2016-00-00
Epub
2016-00-30
Pages
2103
Language
English
Region
England
NLM ID
101594320
PMCID
PMC5034794
Grants
NIGMS NIH HHS · R01 GM076990 · United States
Corrections
CommentIn
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]