A new study in Nature Mental Health warns that a common promise of machine learning—precise predictions from noisy clinical data—may be overstated when psychological symptom scores fluctuate for reasons unrelated to treatment. Researchers led by Radua and colleagues examined how “regression to the mean” can artificially boost the apparent accuracy of models forecasting changes in symptoms over time.
In many mental health trials, individuals do not simply move in one direction; their scores naturally drift. Some participants begin an assessment with unusually high or low symptom levels, driven by temporary circumstances, measurement noise, or short-term variation. Even without any therapeutic effect, their later scores are statistically likely to move toward average values. This phenomenon is known as regression to the mean.
The team showed that prediction pipelines that learn patterns from past data can inadvertently capitalize on this statistical tendency. When a model sees extreme baseline symptoms, it may predict improvement (or worsening) largely because extreme observations are, by definition, followed by more average outcomes. In such cases, high performance metrics can reflect the mathematics of sampling rather than an understanding of underlying clinical mechanisms.
To evaluate this risk, the authors focused on how machine learning estimates symptom change. They analyzed whether “accuracy” remained robust when accounting for the tendency of repeated measurements to cluster near the population average. Their findings suggest that, without careful controls, models can look impressively accurate while delivering limited real-world clinical value.
The paper underscores that symptom prediction is especially vulnerable because psychiatric measures are often noisy and frequently subject to ceiling and floor effects, varying engagement, and context-dependent reporting. Together, these factors can amplify apparent predictability that stems from regression to the mean.
For clinicians and developers, the study offers a practical warning: evaluation protocols should distinguish true prognostic signals from statistical artifacts. That means testing against baselines that reflect expected drift, using appropriate validation strategies, and interpreting performance in relation to the stability and distribution of symptom measures.
The results also highlight a broader lesson for viral science news in clinical AI: “better accuracy” is not synonymous with “better decision-making.” Models can improve prediction scores while still failing to identify who will benefit meaningfully from an intervention.
Ultimately, this work calls for more rigorous benchmarking of mental health prediction models. By explicitly addressing regression to the mean, future systems can be assessed on whether they capture treatment-relevant dynamics rather than predictable statistical rebound.
Subject of Research: Machine learning prediction of symptom change in mental health; regression to the mean in clinical trial data.
Article Title: Regression to the mean inflates accuracy in machine learning prediction of symptom change.
Article References: Radua, J., Oliva, V., Vieta, E. et al. Regression to the mean inflates accuracy in machine learning prediction of symptom change. Nat. Mental Health (2026). https://doi.org/10.1038/s44220-026-00686-6
Image Credits: AI Generated
DOI: 10.1038/s44220-026-00686-6
Keywords: Regression to the mean; machine learning; symptom change prediction; mental health; clinical data; validation; statistical bias

