Non-suicidal self-injury, the deliberate harming of one’s own body without suicidal intent, remains one of the most troubling and least well-understood behaviors of adolescence. It leaves scars on skin and psyche alike, strains families, and overwhelms school counselors and clinicians who often have no systematic way to know which young people are most at risk. A new systematic review and meta-analysis published in BMC Psychiatry by a team of researchers based at Nanjing Medical University and its affiliated Nanjing Brain Hospital takes stock of a rapidly growing field: the statistical and machine-learning models built to forecast which adolescents will engage in self-injury. The verdict is a careful mix of encouragement and caution. The models, on average, perform reasonably well, but the science behind them is far weaker than their headline numbers suggest.
The research team, led by Tong Xiao and Yiru Wang, who contributed equally to the work, with Yanhong Zhang as corresponding author, searched ten databases spanning the English-language and Chinese biomedical literature, including PubMed, Embase, Web of Science, CINAHL, the Cochrane Library, PsycINFO, and four major Chinese databases. Their search covered everything published from database inception through August 15, 2025. From that sweep they identified 27 studies containing 58 distinct prediction models for adolescent non-suicidal self-injury. Seventeen of the source studies were conducted in hospital settings and ten in community settings, a split that matters because the predictors available and the base rates of self-injury can differ substantially between a psychiatric clinic and a school population.
The centerpiece of the analysis is the pooled area under the receiver operating characteristic curve, or AUC, the most widely used summary of how well a model separates those who will experience an outcome from those who will not. An AUC of 0.5 reflects performance no better than a coin flip, while 1.0 represents perfect discrimination. Across the 17 models for which training-set data could be pooled, the meta-analysis found a combined AUC of 0.850, with a 95 percent confidence interval of 0.810 to 0.890. For the far smaller set of five models evaluated on independent validation data, the pooled AUC was 0.840, with a wider interval of 0.760 to 0.930. In conventional terms, that places these tools in the moderate to good range of predictive performance, comparable to risk models used in other areas of medicine.
But the authors are quick to point out that a strong AUC on a training set is not the same as a trustworthy clinical instrument. Only a handful of the 58 models had ever been externally validated, meaning tested on data not used to build them, and the confidence interval for the validation estimate is wide enough to encompass everything from mediocre to excellent performance. Discrimination is only half the story in prediction research. A model must also be well calibrated, meaning that when it says a patient has a 30 percent risk, roughly 30 percent of similar patients should actually experience the outcome. Calibration was rarely assessed or reported in the studies the team reviewed, a gap that has plagued prediction modeling across medicine and that the authors single out as a priority for future work.
Methodological quality emerged as the study’s most sobering finding. Using the PROBAST + AI tool, a risk-of-bias instrument designed specifically for prediction models, including those built with artificial intelligence, the reviewers judged 24 of the 27 studies to be of low overall quality and 23 to be at high risk of bias. Two studies were rated at high risk of concerns regarding applicability, meaning their data or predictors might not transfer well to the populations clinicians actually want to serve. Common sources of bias in prediction research include small sample sizes relative to the number of predictors, inappropriate handling of missing data, optimistic performance estimates from evaluating models on the same data used to develop them, and selective reporting of the best-performing model among many that were tried.
What drives the risk, according to the models themselves? Across the included studies, a consistent cluster of predictors surfaced again and again. Childhood trauma was among the most powerful signals, echoing decades of research linking early adversity to self-directed harm. Depressive mood and anxiety, the emotional workhorses of adolescent risk screening, appeared frequently, as did a history of prior non-suicidal self-injury, which is unsurprising given that past behavior is one of the strongest predictors of future behavior in nearly every domain of mental health. Sleep disorders, stressful life events, and female gender rounded out the list of common predictors. The recurrence of these variables across diverse settings suggests a reasonably stable risk architecture, even if the models that combine them vary widely in construction.
The implications for clinical practice are double-edged. On one hand, the finding that pooled AUCs hover around 0.85 suggests that a data-informed approach to identifying at-risk adolescents is feasible. Schools, pediatric clinics, and emergency departments could, in principle, deploy brief screening tools that combine a small number of well-validated predictors to flag young people who warrant deeper assessment. On the other hand, the authors conclude that the existing models perform poorly in methodological quality and clinical applicability, and they explicitly call for future research to strengthen validation and calibration of existing models and to improve methodological rigor in line with PROBAST criteria. Until that happens, they caution, these tools should not be treated as ready-made clinical decision instruments.
The review also highlights a structural weakness of the field: fragmentation. With 58 models across 27 studies, many built on small, single-center samples and using different predictor sets, statistical techniques, and outcome definitions, there is little opportunity to compare approaches head to head or to pool their strengths. The authors suggest that the path forward lies less in building yet more novel models and more in rigorously validating, calibrating, and refining the ones that already exist, ideally in large, diverse, prospectively collected cohorts. Reporting standards such as TRIPOD, the transparency checklist for prediction studies, and the PROBAST framework for assessing bias offer a shared language for judging whether a model deserves clinical trust.
For a behavior as hidden and as consequential as adolescent self-injury, the promise of early, data-driven identification is genuinely compelling. Non-suicidal self-injury is associated with serious adverse effects on individuals, families, and society, and it frequently goes undetected until it has become entrenched. A validated screening model could shift the clinical posture from reaction to prevention, directing scarce mental health resources toward the adolescents who need them most. This meta-analysis provides the clearest picture yet of what that future might look like: a field with real predictive signal, a recognizable set of risk factors, and performance statistics that look respectable on paper, but one that still must clear the harder hurdles of external validation, honest calibration, and methodological discipline before its algorithms earn a place in the clinic.
The study, funded by the Nanjing Health Science and Technology Development Special Fund Project and published open access, arrives at a moment when machine-learning tools are being proposed for nearly every corner of mental health care. Its message generalizes well beyond self-injury. Impressive AUCs are easy to produce and hard to trust; the difference lies in validation, calibration, and transparent reporting. For the researchers, clinicians, and families hoping that prediction models can help protect vulnerable teenagers, the new review offers both a map of what has been achieved and a candid account of how much work remains before those achievements can be safely put to use.
Subject of Research: Risk prediction models for non-suicidal self-injury among adolescents
Article Title: Risk prediction models of non-suicidal self-injury among adolescents: a systematic review and meta-analysis
Article References: Risk prediction models of non-suicidal self-injury among adolescents: a systematic review and meta-analysis. (n.d.). https://doi.org/10.1186/s12888-026-08719-1
Image Credits: AI Generated
DOI: 10.1186/s12888-026-08719-1
Keywords: non-suicidal self-injury, adolescents, risk prediction models, systematic review, meta-analysis, machine learning, PROBAST, AUC, childhood trauma, depression, anxiety, mental health screening
Cite Scienmag News
Glenn Wilkins. (October 3, 2026). Can Algorithms Predict Which Teens Will Self-Harm? A sweeping review finds promise and pitfalls. Scienmag. https://scienmag.com/can-algorithms-predict-which-teens-will-self-harm-a-sweeping-review-finds-promise-and-pitfalls/
Glenn Wilkins. "Can Algorithms Predict Which Teens Will Self-Harm? A sweeping review finds promise and pitfalls." Scienmag, 3 October 2026, https://scienmag.com/can-algorithms-predict-which-teens-will-self-harm-a-sweeping-review-finds-promise-and-pitfalls/. Accessed 3 October 2026.
Glenn Wilkins. "Can Algorithms Predict Which Teens Will Self-Harm? A sweeping review finds promise and pitfalls." Scienmag. October 3, 2026. https://scienmag.com/can-algorithms-predict-which-teens-will-self-harm-a-sweeping-review-finds-promise-and-pitfalls/

