Scinovex
article Open AccessTop 1% cited

Waste Not, Want Not: Why Rarefying Microbiome Data Is Inadmissible

PLoS Computational Biology · 2014 · Vol. 10(4) · pp. e1003531–e1003531
Paul J. McMurdieSusan Holmes

Abstract

Current practice in the normalization of microbiome count data is inefficient in the statistical sense. For apparently historical reasons, the common approach is either to use simple proportions (which does not address heteroscedasticity) or to use rarefying of counts, even though both of these approaches are inappropriate for detection of differentially abundant species. Well-established statistical theory is available that simultaneously accounts for library size differences and biological variability using an appropriate mixture model. Moreover, specific implementations for DNA sequencing read count data (based on a Negative Binomial model for instance) are already available in RNA-Seq focused R packages such as edgeR and DESeq. Here we summarize the supporting statistical theory and use simulations and empirical data to demonstrate substantial improvements provided by a relevant mixture model framework over simple proportions or rarefying. We show how both proportions and rarefied counts result in a high rate of false positives in tests for species that are differentially abundant across sample classes. Regarding microbiome sample-wise clustering, we also show that the rarefying procedure often discards samples that can be accurately clustered by alternative methods. We further compare different Negative Binomial methods with a recently-described zero-inflated Gaussian mixture, implemented in a package called metagenomeSeq. We find that metagenomeSeq performs well when there is an adequate number of biological replicates, but it nevertheless tends toward a higher false positive rate. Based on these results and well-established statistical theory, we advocate that investigators avoid rarefying altogether. We have provided microbiome-specific extensions to these tools in the R package, phyloseq.

Gut microbiota and healthBayesian Methods and Mixture ModelsMetabolomics and Mass Spectrometry StudiesNegative binomial distributionHeteroscedasticityCount dataMicrobiomeComputer scienceNormalization (sociology)R packageMixture modelSample size determinationStatistical model

MeSH terms

Models, TheoreticalSequence Analysis, DNASequence Analysis, RNAMicrobiota

Funding

  • National Institutes of Health
Citations
3,022
FWCI
74.74
field-weighted impact
References
93
Percentile
100%
vs. same field & year
Citations per year
Cited by
Microbiome Datasets Are Compositional: And This Is Not Optional
Frontiers in Microbiology · 2017 · 2,931 citations
Multivariable association discovery in population-scale meta-omics studies
PLoS Computational Biology · 2021 · 2,289 citations
Rarefaction, Alpha Diversity, and Statistics
Frontiers in Microbiology · 2019 · 679 citations
Best practices for analysing microbiomes
Nature Reviews Microbiology · 2018 · 1,823 citations
Sparse and Compositionally Robust Inference of Microbial Ecological Networks
PLoS Computational Biology · 2015 · 1,814 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.