Home LiteratureArticle Details
PMID: 24699258 Published · epublish English Journal Article

Waste not, want not: why rarefying microbiome data is inadmissible.

PLoS computational biology ·Vol. 10 ·No. 4 ·2014-04-00 ·Pages e1003531

McMurdie PJ, Holmes S

Abstract

Current practice in the normalization of microbiome count data is inefficient in the statistical sense. For apparently historical reasons, the common approach is either to use simple proportions (which does not address heteroscedasticity) or to use rarefying of counts, even though both of these approaches are inappropriate for detection of differentially abundant species. Well-established statistical theory is available that simultaneously accounts for library size differences and biological variability using an appropriate mixture model. Moreover, specific implementations for DNA sequencing read count data (based on a Negative Binomial model for instance) are already available in RNA-Seq focused R packages such as edgeR and DESeq. Here we summarize the supporting statistical theory and use simulations and empirical data to demonstrate substantial improvements provided by a relevant mixture model framework over simple proportions or rarefying. We show how both proportions and rarefied counts result in a high rate of false positives in tests for species that are differentially abundant across sample classes. Regarding microbiome sample-wise clustering, we also show that the rarefying procedure often discards samples that can be accurately clustered by alternative methods. We further compare different Negative Binomial methods with a recently-described zero-inflated Gaussian mixture, implemented in a package called metagenomeSeq. We find that metagenomeSeq performs well when there is an adequate number of biological replicates, but it nevertheless tends toward a higher false positive rate. Based on these results and well-established statistical theory, we advocate that investigators avoid rarefying altogether. We have provided microbiome-specific extensions to these tools in the R package, phyloseq.

MeSH Terms
Microbiota Models, Theoretical Sequence Analysis, DNA Sequence Analysis, RNA
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
McMurdie Paul J
Statistics Department, Stanford University, Stanford, California, United States of America.
Holmes Susan
Statistics Department, Stanford University, Stanford, California, United States of America.
References (52)
52 references, click to expand
  1. Statistical methods for detecting differentially abundant features in clinical metagenomic samples.
    PLoS Comput Biol. 2009 Apr;5(4):e1000352 PMID: 19360128
  2. Advancing our understanding of the human microbiome using QIIME.
    Methods Enzymol. 2013;531:371-444 PMID: 24060131
  3. A guide to enterotypes across the human body: meta-analysis of microbial community structures in human microbiome datasets.
    PLoS Comput Biol. 2013;9(1):e1002863 PMID: 23326225
  4. Evaluating different approaches that test whether microbial communities have the same structure.
    ISME J. 2008 Mar;2(3):265-75 PMID: 18239608
  5. Reduced incidence of Prevotella and other fermenters in intestinal microflora of autistic children.
    PLoS One. 2013 Jul 03;8(7):e68322 PMID: 23844187
  6. Metagenomics: genomic analysis of microbial communities.
    Annu Rev Genet. 2004;38:525-52 PMID: 15568985
  7. The seasonal structure of microbial communities in the Western English Channel.
    Environ Microbiol. 2009 Dec;11(12):3132-9 PMID: 19659500
  8. A molecular view of microbial diversity and the biosphere.
    Science. 1997 May 2;276(5313):734-40 PMID: 9115194
  9. A comparison of methods for differential expression analysis of RNA-seq data.
    BMC Bioinformatics. 2013 Mar 09;14:91 PMID: 23497356
  10. UniFrac: a new phylogenetic method for comparing microbial communities.
    Appl Environ Microbiol. 2005 Dec;71(12):8228-35 PMID: 16332807
  11. Computational meta'omics for microbial community studies.
    Mol Syst Biol. 2013 May 14;9:666 PMID: 23670539
  12. Disordered microbial communities in the upper respiratory tract of cigarette smokers.
    PLoS One. 2010 Dec 20;5(12):e15216 PMID: 21188149
  13. Exploring microbial diversity and taxonomy using SSU rRNA hypervariable tag sequencing.
    PLoS Genet. 2008 Nov;4(11):e1000255 PMID: 19023400
  14. A comprehensive comparison of RNA-Seq-based transcriptome analysis from reads to differential gene expression and cross-comparison with microarrays: a case study in Saccharomyces cerevisiae.
    Nucleic Acids Res. 2012 Nov 1;40(20):10084-97 PMID: 22965124
  15. The effects of circumcision on the penis microbiome.
    PLoS One. 2010 Jan 06;5(1):e8422 PMID: 20066050
  16. Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.
    BMC Bioinformatics. 2010 Feb 18;11:94 PMID: 20167110
  17. Shrinkage estimation of dispersion in Negative Binomial models for RNA-seq experiments with small sample size.
    Bioinformatics. 2013 May 15;29(10):1275-82 PMID: 23589650
  18. Independent filtering increases detection power for high-throughput experiments.
    Proc Natl Acad Sci U S A. 2010 May 25;107(21):9546-51 PMID: 20460310
  19. An invitation to reproducible computational research.
    Biostatistics. 2010 Jul;11(3):385-8 PMID: 20538873
  20. Small-sample estimation of negative binomial dispersion, with applications to SAGE data.
    Biostatistics. 2008 Apr;9(2):321-32 PMID: 17728317
  21. Introducing mothur: open-source, platform-independent, community-supported software for describing and comparing microbial communities.
    Appl Environ Microbiol. 2009 Dec;75(23):7537-41 PMID: 19801464
  22. phyloseq: an R package for reproducible interactive analysis and graphics of microbiome census data.
    PLoS One. 2013 Apr 22;8(4):e61217 PMID: 23630581
  23. Architectural design influences the diversity and structure of the built environment microbiome.
    ISME J. 2012 Aug;6(8):1469-79 PMID: 22278670
  24. TCC: an R package for comparing tag count data with robust normalization strategies.
    BMC Bioinformatics. 2013 Jul 09;14:219 PMID: 23837715
  25. Human gut microbiome viewed across age and geography.
    Nature. 2012 May 09;486(7402):222-7 PMID: 22699611
  26. Fast UniFrac: facilitating high-throughput phylogenetic analyses of microbial communities including analysis of pyrosequencing and PhyloChip data.
    ISME J. 2010 Jan;4(1):17-27 PMID: 19710709
  27. Composition of the adult digestive tract bacterial microbiome based on seven mouth surfaces, tonsils, throat and stool samples.
    Genome Biol. 2012 Jun 14;13(6):R42 PMID: 22698087
  28. Next-generation DNA sequencing.
    Nat Biotechnol. 2008 Oct;26(10):1135-45 PMID: 18846087
  29. Error-correcting barcoded primers for pyrosequencing hundreds of samples in multiplex.
    Nat Methods. 2008 Mar;5(3):235-7 PMID: 18264105
  30. Global patterns of 16S rRNA diversity at a depth of millions of sequences per sample.
    Proc Natl Acad Sci U S A. 2011 Mar 15;108 Suppl 1:4516-22 PMID: 20534432
  31. The application of rarefaction techniques to molecular inventories of microbial diversity.
    Methods Enzymol. 2005;397:292-308 PMID: 16260298
  32. Reproducible research in computational science.
    Science. 2011 Dec 2;334(6060):1226-7 PMID: 22144613
  33. Microarray data analysis: from disarray to consolidation and consensus.
    Nat Rev Genet. 2006 Jan;7(1):55-65 PMID: 16369572
  34. Linking long-term dietary patterns with gut microbial enterotypes.
    Science. 2011 Oct 7;334(6052):105-8 PMID: 21885731
  35. ROCR: visualizing classifier performance in R.
    Bioinformatics. 2005 Oct 15;21(20):3940-1 PMID: 16096348
  36. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  37. Differential expression analysis for sequence count data.
    Genome Biol. 2010;11(10):R106 PMID: 20979621
  38. High throughput sequencing methods and analysis for microbiome research.
    J Microbiol Methods. 2013 Dec;95(3):401-14 PMID: 24029734
  39. Bioconductor: open software development for computational biology and bioinformatics.
    Genome Biol. 2004;5(10):R80 PMID: 15461798
  40. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  41. Differential abundance analysis for microbial marker-gene surveys.
    Nat Methods. 2013 Dec;10(12):1200-2 PMID: 24076764
  42. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  43. DFI: gene feature discovery in RNA-seq experiments from multiple sources.
    BMC Genomics. 2012;13 Suppl 8:S11 PMID: 23281963
  44. baySeq: empirical Bayesian methods for identifying differential expression in sequence count data.
    BMC Bioinformatics. 2010 Aug 10;11:422 PMID: 20698981
  45. Accurate taxonomy assignments from 16S rRNA sequences produced by highly parallel pyrosequencers.
    Nucleic Acids Res. 2008 Oct;36(18):e120 PMID: 18723574
  46. Diversity, distribution and sources of bacteria in residential kitchens.
    Environ Microbiol. 2013 Feb;15(2):588-96 PMID: 23171378
  47. The expanding scope of DNA sequencing.
    Nat Biotechnol. 2012 Nov;30(11):1084-94 PMID: 23138308
  48. Identifying differential expression in multiple SAGE libraries: an overdispersed log-linear model approach.
    BMC Bioinformatics. 2005 Jun 29;6:165 PMID: 15987513
  49. UniFrac: an effective distance metric for microbial community comparison.
    ISME J. 2011 Feb;5(2):169-72 PMID: 20827291
  50. High-density microarray of small-subunit ribosomal DNA probes.
    Appl Environ Microbiol. 2002 May;68(5):2535-41 PMID: 11976131
  51. Quantitative and qualitative beta diversity measures lead to different insights into factors that structure microbial communities.
    Appl Environ Microbiol. 2007 Mar;73(5):1576-85 PMID: 17220268
  52. QIIME allows analysis of high-throughput community sequencing data.
    Nat Methods. 2010 May;7(5):335-6 PMID: 20383131
Article Info
Journal
PLoS computational biology
Abbr.
PLoS Comput Biol
ISSN
1553-7358
Published
2014-04-00
Epub
2014-00-03
Pages
e1003531
Language
English
Region
United States
NLM ID
101238922
PMCID
PMC3974642
Subset
IM
Grants
NIGMS NIH HHS · R01 GM086884 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]