Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Psychology & Psychiatry

Warning Signs That Your Bayes Factor May Be Lying to You

October 1, 2026
in Psychology & Psychiatry
Glenn Wilkins
By Glenn Wilkins Scienmag Editorial Profile - Clinical Psychology
Reading Time: 6 mins read
0
Warning Signs That Your Bayes Factor May Be Lying to You

Warning Signs That Your Bayes Factor May Be Lying to You

Warning Signs That Your Bayes Factor May Be Lying to You

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Bayes factors have become one of the most fashionable tools in the statistical toolbox of psychology, neuroscience, and the social sciences. Unlike the familiar p-value, a Bayes factor promises a continuous measure of evidence: it tells you, in a single number, how strongly your data favor one hypothesis over another, and it can even quantify evidence in favor of a null hypothesis that an effect is truly absent. But a new simulation study published in Behavior Research Methods delivers a sobering message for the thousands of researchers who rely on this method every year. Daniel J. Schad of HMU Health and Medical University in Potsdam and Martin Modrák of Charles University in Prague show that the Bayes factors produced by widely used software can be accurate, but only under specific conditions, and that a small warning message buried in the software output is the dividing line between trustworthy evidence and potentially misleading numbers.

The heart of the problem is mathematical. For the simple textbook models used in introductory statistics courses, Bayes factors can be computed exactly. But real experiments in cognitive science typically involve linear mixed-effects models with random effects for subjects and items, hierarchical structure, and multiple fixed effects. For these models, the Bayes factor requires evaluating an integral over all possible parameter values, an integral that has no closed-form solution. Researchers therefore depend on numerical approximations. One popular strategy, implemented in the R package bridgesampling, first estimates the posterior distribution of each competing model using Markov-chain Monte Carlo (MCMC) sampling and then uses those posterior samples to approximate the marginal likelihood, the quantity whose ratio defines the Bayes factor. Every step of this pipeline introduces potential error, and until now there has been no systematic way to know whether the final number on the screen corresponds to the true Bayes factor or to a biased approximation of it.

Schad and Modrák attacked this question with a technique called marginal simulation-based calibration, or marginal SBC, which the authors had previously developed. The logic is elegant. The researchers define a prior probability for each competing hypothesis, for example a 50 percent chance that the null model and the alternative model are each true. They then repeat a simulation many times: sample a model from that prior, sample parameter values from the parameter priors, generate artificial data from the sampled model, and compute the Bayes factor for the simulated data using the method under scrutiny. If the Bayes factor computation is correct, then the average posterior model probability recovered across all the simulations should match the prior model probability that was used to generate the data. If the average posterior drifts away from the prior, the estimator is biased, either liberally, exaggerating evidence for effects, or conservatively, understating it. The authors supplemented this test with reliability diagrams, which visualize whether posterior probabilities correctly predict which model actually generated each simulated dataset.

The team applied this procedure to three experimental designs that dominate cognitive psychology and psycholinguistics. The first was a 2×2 repeated measures design inspired by the two-step decision-making task used to study model-based and model-free reinforcement learning, with random effects for subjects only. The second was a Latin square design with a single two-level fixed factor and crossed random effects for both subjects and items, analyzed with a Bayesian generalized linear mixed model with a lognormal likelihood. The third was a full 2×2 Latin square design with crossed random effects for subjects and items, testing two main effects and their interaction. In each case, the null model was created by setting one fixed effect to exactly zero, and the full model was compared against it using the brms package for model fitting in combination with bridgesampling for the Bayes factor computation, using the warp3 variant of the bridge sampler and 10,000 MCMC iterations with 2,000 warm-up iterations.

The results split cleanly along one diagnostic line: whether or not the bridge sampling algorithm issued a warning message. In design 1, the algorithm flagged 122 of 1,200 model fits, roughly 20 percent, with the message that the log marginal likelihood could not be estimated within the maximum number of iterations and that the estimate might be more variable than usual. For the simulations without warnings, the calibration checks were reassuring. Bayesian t-tests comparing the average posterior model probability against the prior found default Bayes factors exceeding 10 in favor of no difference, and the reliability diagrams showed posterior probabilities aligned with the true generating model. In other words, when the software stayed silent, the Bayes factors were accurate. When a warning appeared, however, the picture darkened: the reliability diagrams deviated visibly from the diagonal, indicating miscalibrated and biased estimates.

The pattern repeated across the other designs, sometimes more dramatically. In the Latin square design with crossed random effects, only 3 percent of fits produced warnings, and the Bayes factors proved accurate. But in the full 2×2 design with crossed random effects, warnings appeared in 37 percent of simulations, and the consequences were serious. When the prior probability of the null hypothesis was set high at 0.8, the biased estimates pushed posterior model probabilities for the alternative hypothesis too high, producing liberal tests that could declare evidence for an effect even when the true effect was zero. When the researchers flipped the setup and gave the alternative hypothesis a prior probability of 0.8, the bias reversed direction, yielding conservative tests that could hide genuine effects. The culprit was not a systematic upward or downward distortion but extreme noise: the variance of the estimated log marginal likelihoods ballooned in the warned simulations, and floor and ceiling effects then pushed posterior probabilities toward 0.5, making the resulting Bayes factors essentially uninformative.

The authors also probed whether the problem could simply be computed away. Increasing the MCMC sample size from 10,000 to 40,000 iterations reduced the warning rate only marginally, from 37 percent to 32.5 percent, a change well within expected sampling variability. Running 1,000 SBC simulations instead of 200, however, sharpened the conclusions considerably: with that statistical power, the evidence for accuracy in warning-free simulations grew strong, with default Bayes factors above 10 favoring no bias, while the evidence for bias in warned simulations became overwhelming, with Bayes factors exceeding 10,000. A re-analysis across varying numbers of SBC runs showed that evidence for bias in warned simulations emerged quickly, with fewer than 100 simulation runs, whereas evidence for accuracy accumulated more slowly. The practical lesson is that researchers can detect dangerous Bayes factors relatively cheaply, but confirming that a computation is clean requires a somewhat larger investment of simulations.

What causes the warnings in the first place? The authors found that warned datasets were precisely those that forced Stan’s Hamiltonian Monte Carlo sampler to take very small steps and explore deep trajectories, indicating a highly irregular posterior geometry. Problematic cases tended to involve very small fitted residual standard deviations or random effect standard deviations close to zero, creating weak non-identifiability between the overall effects and the subject-level random effects. In such situations, the authors suggest, reparameterizing the model can help: dropping the random effect for the interaction, or imposing sum-to-zero constraints on random intercepts and slopes, resolved the warnings in some cases. But these fixes are not guaranteed, and the authors are candid that researchers facing persistent warnings may need to turn to alternative approaches altogether, such as the Savage-Dickey method, the BayesFactor package, region of practical equivalence decisions, or posterior credibility intervals, after verifying that those alternatives are themselves accurate for the analysis at hand.

The broader message is both a validation and a caution. On the one hand, the study provides real evidence that brms and bridgesampling deliver accurate Bayes factors for the factorial designs that dominate experimental psychology and psycholinguistics, provided the computation proceeds without complaints. On the other hand, it demolishes the comfortable assumption that software output can be taken at face value. A warning message that many users skim past, or dismiss as a minor numerical hiccup, turns out to mark the boundary between calibrated evidence and numbers that are biased, highly variable, and potentially capable of steering scientific conclusions in the wrong direction. The authors argue that marginal SBC should become a routine part of the Bayesian workflow: before trusting Bayes factors from any particular combination of design, model, and priors, researchers should check the computation’s accuracy through simulation. It is a modest additional burden, but one that could spare the field a great deal of misplaced confidence in numbers that look precise and are anything but.

Subject of Research: Accuracy of Bayes factor-based null hypothesis tests in linear mixed-effects models assessed with simulation-based calibration

Article Title: How accurate are Bayes factor-based null hypothesis tests? A simulation study

Article References: Schad, D. J., & Modrák, M. (2026). How accurate are Bayes factor-based null hypothesis tests? A simulation study. Behavior Research Methods, 58(10), Article 282. https://doi.org/10.3758/s13428-026-03168-w

Image Credits: AI Generated

DOI: 10.3758/s13428-026-03168-w

Keywords: Bayes factors, null hypothesis testing, simulation-based calibration, bridge sampling, linear mixed-effects models, brms, Bayesian statistics, marginal likelihood, psycholinguistics, factorial designs, MCMC, statistical software

Cite Scienmag News

Glenn Wilkins. (October 1, 2026). Warning Signs That Your Bayes Factor May Be Lying to You. Scienmag. https://scienmag.com/warning-signs-that-your-bayes-factor-may-be-lying-to-you/

Glenn Wilkins. "Warning Signs That Your Bayes Factor May Be Lying to You." Scienmag, 1 October 2026, https://scienmag.com/warning-signs-that-your-bayes-factor-may-be-lying-to-you/. Accessed 1 October 2026.

Glenn Wilkins. "Warning Signs That Your Bayes Factor May Be Lying to You." Scienmag. October 1, 2026. https://scienmag.com/warning-signs-that-your-bayes-factor-may-be-lying-to-you/

Tags: Bayes factor interpretationBayes factorsBayesian statisticsBayesian statistics in psychology and neuroscienceBayesian vs p-value comparisonbest practices for Bayesian hypothesis testingbridge samplingbrmsfactorial designshierarchical models in Bayesian analysisissues with Bayesian software outputslimitations of Bayes factors in complex modelslinear mixed-effects modelsmarginal likelihoodMCMCnull hypothesis evidence quantificationnull hypothesis testingpitfalls of Bayesian statisticspsycholinguisticsreliability of Bayesian evidencesimulation studies on Bayes factorssimulation-based calibrationstatistical softwarewarning signs in Bayesian model evidence
Share26Tweet16
Previous Post

EU Adopts Landmark Rules for Gene-Edited Plants After Two-Decade Wait

Next Post

AI and Non-Destructive Spectroscopy Set to Replace Century-Old Antioxidant Tests

Related Posts

Nurses Split Into Two Thriving Profiles, and Calling Predicts Which One They Land In
Psychology & Psychiatry

Nurses Split Into Two Thriving Profiles, and Calling Predicts Which One They Land In

October 1, 2026
Erectile Dysfunction Drug Linked to Rare Acute Psychiatric Syndrome in Case Report
Psychology & Psychiatry

Erectile Dysfunction Drug Linked to Rare Acute Psychiatric Syndrome in Case Report

October 1, 2026
Coffee Shop Culture Boosts Youth Social Lives, But Only Meaningful Engagement Lifts Wellbeing
Psychology & Psychiatry

Coffee Shop Culture Boosts Youth Social Lives, But Only Meaningful Engagement Lifts Wellbeing

October 1, 2026
When a Depression Questionnaire Gets Lost in Translation: Rural Indian Women Read the EPDS Through Daily Life
Psychology & Psychiatry

When a Depression Questionnaire Gets Lost in Translation: Rural Indian Women Read the EPDS Through Daily Life

October 1, 2026
Homophobic Violence Linked to Suicidal Thoughts in New Study of LGBTQIA+ Adults
Psychology & Psychiatry

Homophobic Violence Linked to Suicidal Thoughts in New Study of LGBTQIA+ Adults

October 1, 2026
Brain Injuries and Chronic Stress May Drive Memory Loss in Abuse Survivors
Psychology & Psychiatry

Brain Injuries and Chronic Stress May Drive Memory Loss in Abuse Survivors

October 1, 2026
Next Post
AI and Non-Destructive Spectroscopy Set to Replace Century-Old Antioxidant Tests

AI and Non-Destructive Spectroscopy Set to Replace Century-Old Antioxidant Tests

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Droplet Digital PCR Speeds Pathogen Detection in Febrile Leukopenia Patients
  • PET/MR Scan Predicts Prostate Cancer Relapse Before Surgery
  • AI and Non-Destructive Spectroscopy Set to Replace Century-Old Antioxidant Tests
  • Warning Signs That Your Bayes Factor May Be Lying to You

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading