Bayes factors have become one of the most fashionable tools in the statistical toolbox of psychology, neuroscience, and the social sciences. Unlike the familiar p-value, a Bayes factor promises a continuous measure of evidence: it tells you, in a single number, how strongly your data favor one hypothesis over another, and it can even quantify evidence in favor of a null hypothesis that an effect is truly absent. But a new simulation study published in Behavior Research Methods delivers a sobering message for the thousands of researchers who rely on this method every year. Daniel J. Schad of HMU Health and Medical University in Potsdam and Martin Modrák of Charles University in Prague show that the Bayes factors produced by widely used software can be accurate, but only under specific conditions, and that a small warning message buried in the software output is the dividing line between trustworthy evidence and potentially misleading numbers.
The heart of the problem is mathematical. For the simple textbook models used in introductory statistics courses, Bayes factors can be computed exactly. But real experiments in cognitive science typically involve linear mixed-effects models with random effects for subjects and items, hierarchical structure, and multiple fixed effects. For these models, the Bayes factor requires evaluating an integral over all possible parameter values, an integral that has no closed-form solution. Researchers therefore depend on numerical approximations. One popular strategy, implemented in the R package bridgesampling, first estimates the posterior distribution of each competing model using Markov-chain Monte Carlo (MCMC) sampling and then uses those posterior samples to approximate the marginal likelihood, the quantity whose ratio defines the Bayes factor. Every step of this pipeline introduces potential error, and until now there has been no systematic way to know whether the final number on the screen corresponds to the true Bayes factor or to a biased approximation of it.
Schad and Modrák attacked this question with a technique called marginal simulation-based calibration, or marginal SBC, which the authors had previously developed. The logic is elegant. The researchers define a prior probability for each competing hypothesis, for example a 50 percent chance that the null model and the alternative model are each true. They then repeat a simulation many times: sample a model from that prior, sample parameter values from the parameter priors, generate artificial data from the sampled model, and compute the Bayes factor for the simulated data using the method under scrutiny. If the Bayes factor computation is correct, then the average posterior model probability recovered across all the simulations should match the prior model probability that was used to generate the data. If the average posterior drifts away from the prior, the estimator is biased, either liberally, exaggerating evidence for effects, or conservatively, understating it. The authors supplemented this test with reliability diagrams, which visualize whether posterior probabilities correctly predict which model actually generated each simulated dataset.
The team applied this procedure to three experimental designs that dominate cognitive psychology and psycholinguistics. The first was a 2×2 repeated measures design inspired by the two-step decision-making task used to study model-based and model-free reinforcement learning, with random effects for subjects only. The second was a Latin square design with a single two-level fixed factor and crossed random effects for both subjects and items, analyzed with a Bayesian generalized linear mixed model with a lognormal likelihood. The third was a full 2×2 Latin square design with crossed random effects for subjects and items, testing two main effects and their interaction. In each case, the null model was created by setting one fixed effect to exactly zero, and the full model was compared against it using the brms package for model fitting in combination with bridgesampling for the Bayes factor computation, using the warp3 variant of the bridge sampler and 10,000 MCMC iterations with 2,000 warm-up iterations.
The results split cleanly along one diagnostic line: whether or not the bridge sampling algorithm issued a warning message. In design 1, the algorithm flagged 122 of 1,200 model fits, roughly 20 percent, with the message that the log marginal likelihood could not be estimated within the maximum number of iterations and that the estimate might be more variable than usual. For the simulations without warnings, the calibration checks were reassuring. Bayesian t-tests comparing the average posterior model probability against the prior found default Bayes factors exceeding 10 in favor of no difference, and the reliability diagrams showed posterior probabilities aligned with the true generating model. In other words, when the software stayed silent, the Bayes factors were accurate. When a warning appeared, however, the picture darkened: the reliability diagrams deviated visibly from the diagonal, indicating miscalibrated and biased estimates.
The pattern repeated across the other designs, sometimes more dramatically. In the Latin square design with crossed random effects, only 3 percent of fits produced warnings, and the Bayes factors proved accurate. But in the full 2×2 design with crossed random effects, warnings appeared in 37 percent of simulations, and the consequences were serious. When the prior probability of the null hypothesis was set high at 0.8, the biased estimates pushed posterior model probabilities for the alternative hypothesis too high, producing liberal tests that could declare evidence for an effect even when the true effect was zero. When the researchers flipped the setup and gave the alternative hypothesis a prior probability of 0.8, the bias reversed direction, yielding conservative tests that could hide genuine effects. The culprit was not a systematic upward or downward distortion but extreme noise: the variance of the estimated log marginal likelihoods ballooned in the warned simulations, and floor and ceiling effects then pushed posterior probabilities toward 0.5, making the resulting Bayes factors essentially uninformative.
The authors also probed whether the problem could simply be computed away. Increasing the MCMC sample size from 10,000 to 40,000 iterations reduced the warning rate only marginally, from 37 percent to 32.5 percent, a change well within expected sampling variability. Running 1,000 SBC simulations instead of 200, however, sharpened the conclusions considerably: with that statistical power, the evidence for accuracy in warning-free simulations grew strong, with default Bayes factors above 10 favoring no bias, while the evidence for bias in warned simulations became overwhelming, with Bayes factors exceeding 10,000. A re-analysis across varying numbers of SBC runs showed that evidence for bias in warned simulations emerged quickly, with fewer than 100 simulation runs, whereas evidence for accuracy accumulated more slowly. The practical lesson is that researchers can detect dangerous Bayes factors relatively cheaply, but confirming that a computation is clean requires a somewhat larger investment of simulations.
What causes the warnings in the first place? The authors found that warned datasets were precisely those that forced Stan’s Hamiltonian Monte Carlo sampler to take very small steps and explore deep trajectories, indicating a highly irregular posterior geometry. Problematic cases tended to involve very small fitted residual standard deviations or random effect standard deviations close to zero, creating weak non-identifiability between the overall effects and the subject-level random effects. In such situations, the authors suggest, reparameterizing the model can help: dropping the random effect for the interaction, or imposing sum-to-zero constraints on random intercepts and slopes, resolved the warnings in some cases. But these fixes are not guaranteed, and the authors are candid that researchers facing persistent warnings may need to turn to alternative approaches altogether, such as the Savage-Dickey method, the BayesFactor package, region of practical equivalence decisions, or posterior credibility intervals, after verifying that those alternatives are themselves accurate for the analysis at hand.
The broader message is both a validation and a caution. On the one hand, the study provides real evidence that brms and bridgesampling deliver accurate Bayes factors for the factorial designs that dominate experimental psychology and psycholinguistics, provided the computation proceeds without complaints. On the other hand, it demolishes the comfortable assumption that software output can be taken at face value. A warning message that many users skim past, or dismiss as a minor numerical hiccup, turns out to mark the boundary between calibrated evidence and numbers that are biased, highly variable, and potentially capable of steering scientific conclusions in the wrong direction. The authors argue that marginal SBC should become a routine part of the Bayesian workflow: before trusting Bayes factors from any particular combination of design, model, and priors, researchers should check the computation’s accuracy through simulation. It is a modest additional burden, but one that could spare the field a great deal of misplaced confidence in numbers that look precise and are anything but.
Subject of Research: Accuracy of Bayes factor-based null hypothesis tests in linear mixed-effects models assessed with simulation-based calibration
Article Title: How accurate are Bayes factor-based null hypothesis tests? A simulation study
Article References: Schad, D. J., & Modrák, M. (2026). How accurate are Bayes factor-based null hypothesis tests? A simulation study. Behavior Research Methods, 58(10), Article 282. https://doi.org/10.3758/s13428-026-03168-w
Image Credits: AI Generated
DOI: 10.3758/s13428-026-03168-w
Keywords: Bayes factors, null hypothesis testing, simulation-based calibration, bridge sampling, linear mixed-effects models, brms, Bayesian statistics, marginal likelihood, psycholinguistics, factorial designs, MCMC, statistical software
Cite Scienmag News
Glenn Wilkins. (October 1, 2026). Warning Signs That Your Bayes Factor May Be Lying to You. Scienmag. https://scienmag.com/warning-signs-that-your-bayes-factor-may-be-lying-to-you/
Glenn Wilkins. "Warning Signs That Your Bayes Factor May Be Lying to You." Scienmag, 1 October 2026, https://scienmag.com/warning-signs-that-your-bayes-factor-may-be-lying-to-you/. Accessed 1 October 2026.
Glenn Wilkins. "Warning Signs That Your Bayes Factor May Be Lying to You." Scienmag. October 1, 2026. https://scienmag.com/warning-signs-that-your-bayes-factor-may-be-lying-to-you/

