The human microbiome has become one of the most intensively studied frontiers in biomedical science, with researchers racing to identify which gut and body-surface microbes transmit the effects of diet, drugs, and environmental exposures onto health outcomes ranging from inflammation to metabolic disease. A central statistical tool in this hunt is mediation analysis, which asks whether a microbe sits on the causal pathway between an exposure and a disease endpoint. If a particular bacterial taxon genuinely carries the influence of, say, a high-fat diet onto insulin resistance, then that microbe becomes an attractive therapeutic target. But a new methodological study published in the journal Microbiome warns that many of the tools researchers rely on for this task may be quietly producing far more false discoveries than their advertised error rates suggest, and the team behind the work has built a remedy designed to restore trust in mediator discovery.
The study, led by Qiyu Wang and Yiluan Li of the University of Science and Technology of China together with Yunfei Peng and Zheng-Zheng Tang of the University of Wisconsin-Madison, systematically benchmarked nine existing microbiome mediation methods across a broad sweep of simulation settings and a real data application. Their central finding is stark: across the simulation settings examined, many of the methods exhibited inflated error rates, and the inflation grew worse as compositional effects intensified. In other words, the very property that makes microbiome data distinctive is the property that undermines the statistical guarantees these methods are supposed to provide.
To understand why, one has to appreciate a technical subtlety that plagues nearly all microbiome sequencing. Standard sequencing workflows do not directly measure how many cells of each taxon are present in a sample. Instead, they yield relative abundances, which describe each taxon’s share of the total sequenced material. Absolute abundances, the actual quantities that usually matter for mechanistic interpretation, are not directly observed. Because every relative abundance depends on all the others through the constraint that shares must sum to one, changing the true amount of one taxon mechanically alters the measured proportions of every other taxon in the sample, even when nothing biological has happened to them. This is the compositional effect, and it is well known to cause trouble in differential abundance testing.
What the new study establishes is that the same compositional distortion systematically corrupts mediation analysis as well. Mediation tests evaluate whether the effect of an exposure on an outcome travels through a candidate taxon, typically by modeling the exposure’s effect on the microbe and the microbe’s effect on the outcome and then testing the product of these pathways. When the taxon measurements are relative rather than absolute, spurious correlations can be induced between taxa and outcomes simply because of shifts elsewhere in the community. The result is statistical signals that look like genuine mediation but are in fact artifacts of the compositional structure of the data. Despite documented compositionality-induced false positives in other microbiome analyses, systematic benchmarking of microbiome mediation methods had remained scarce, leaving practitioners without a clear picture of which tools, if any, controlled error reliably.
A second innovation of the study lies in how the simulations were constructed. Synthetic data used to evaluate bioinformatic methods can easily be unrealistic, and a benchmark is only as convincing as the data it rests on. The team anchored their simulations to an experimentally derived absolute abundance template dataset, generating synthetic absolute abundance profiles that better preserve the empirical structure of real microbiome communities than commonly used simulators. This matters because the severity of compositional effects depends on realistic properties such as the distribution of taxon proportions and the correlation structure among taxa, so a simulator that smooths these features away can dramatically understate the problems that arise in practice. By grounding the simulations in real experimental measurements, the researchers aimed to make their benchmark reflect the conditions analysts actually face.
The results of that benchmark were sobering. Across the simulation settings, the nine methods varied widely in how well they controlled false discoveries, with a substantial number showing inflation of error rates beyond their nominal levels, particularly as compositional effects grew stronger. The pattern carried over into real data. In their application to a real microbiome dataset, the researchers observed that the methods which had exhibited severe inflation in simulations also returned substantially larger sets of candidate mediators, a hallmark of methods that are too permissive and flag taxa indiscriminately. For a field where mediator lists often drive downstream experiments and therapeutic hypotheses, such inflated mediator sets represent more than a statistical nuisance; they can send entire research programs chasing signals that do not exist.
The remedy proposed by the team is a new method called CAMRA, which is designed to infer and test mediation effects at the absolute abundance level while working from the relative abundance data that standard sequencing pipelines produce. Rather than treating the observed relative abundances as if they were the quantities of mechanistic interest, CAMRA attempts to account for the compositional relationship between what is measured and what is biologically meaningful, and it performs its mediation testing on that inferred absolute scale. The goal is to align the statistical target of the analysis with the scale on which microbial mechanisms operate, thereby removing the source of the spurious associations that plague relative abundance level testing.
According to the study, CAMRA delivered on both fronts that matter for practitioners. In simulations, the method improved calibration of the false discovery rate, meaning that when it reported a controlled error level, the actual rate of false discoveries conformed to that promise. It also improved power for taxon level mediator discovery, detecting genuine mediators more effectively than the approaches that either drowned in false positives or, in guarding against them, sacrificed sensitivity. Equally important for real world adoption, CAMRA maintained favorable runtime, so the gains in statistical reliability do not come at a computational cost that would make the method impractical for the high dimensional taxon tables typical of microbiome studies.
The implications reach well beyond the immediate technical community. Identifying microbial taxa that mediate the effects of exposures on health outcomes can elucidate causal pathways and suggest therapeutic targets, which is precisely the kind of translational promise that has drawn enormous investment into microbiome research. If the mediation analyses underlying such claims are systematically inflated, therapeutic leads may be built on statistical mirages, and failed validation studies down the line may be misattributed to biology rather than to methodology. Well-calibrated inference, the authors argue, is a prerequisite for reliable microbial mediator discovery, and their benchmark provides the field with its first systematic map of where current methods stand.
For working microbiome scientists, the practical guidance emerging from this study is to treat mediator lists produced by relative abundance level tests with heightened caution, especially in communities where compositional effects are strong, and to consider absolute abundance targeted approaches such as CAMRA when the goal is taxon level mediator discovery. The study also underscores the value of rigorous, simulation grounded benchmarking as a discipline in its own right: by anchoring synthetic data to experimental templates and evaluating error control explicitly, methodologists can surface failure modes that individual applications would never reveal. As microbiome mediation analysis continues to move from exploratory statistics toward the justification of therapeutic interventions, the difference between a well calibrated test and an inflated one may ultimately determine which microbial targets turn out to be real.
Subject of Research: Statistical error control and compositional bias in microbiome mediation analysis
Article Title: Error control for microbiome mediation analysis: benchmarking and remedy
Article References: Wang, Q., Li, Y., Peng, Y., & Tang, Z.-Z. (2026). Error control for microbiome mediation analysis: benchmarking and remedy. Microbiome. https://doi.org/10.1186/s40168-026-02527-1
Image Credits: AI Generated
DOI: 10.1186/s40168-026-02527-1
Keywords: microbiome, mediation analysis, relative abundance, absolute abundance, compositional effects, false discovery rate, benchmarking, CAMRA, statistical methods, microbial taxa, causal inference, bioinformatics
Cite Scienmag News
Morgan Morrow. (September 25, 2026). Scientists expose hidden false positives in microbiome mediation studies. Scienmag. https://scienmag.com/scientists-expose-hidden-false-positives-in-microbiome-mediation-studies/
Morgan Morrow. "Scientists expose hidden false positives in microbiome mediation studies." Scienmag, 25 September 2026, https://scienmag.com/scientists-expose-hidden-false-positives-in-microbiome-mediation-studies/. Accessed 25 September 2026.
Morgan Morrow. "Scientists expose hidden false positives in microbiome mediation studies." Scienmag. September 25, 2026. https://scienmag.com/scientists-expose-hidden-false-positives-in-microbiome-mediation-studies/

