<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>suicide prevention using artificial intelligence &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/suicide-prevention-using-artificial-intelligence/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 06 Sep 2026 15:17:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>suicide prevention using artificial intelligence &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Meta-analysis evaluates machine learning accuracy in predicting suicide risk</title>
		<link>https://scienmag.com/meta-analysis-evaluates-machine-learning-accuracy-in-predicting-suicide-risk/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 06 Sep 2026 15:17:06 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[accuracy of suicide risk forecasting algorithms]]></category>
		<category><![CDATA[area under ROC curve for suicide forecasting]]></category>
		<category><![CDATA[bedside measures for suicide risk evaluation]]></category>
		<category><![CDATA[challenges in machine learning-based suicide forecasting]]></category>
		<category><![CDATA[challenges in predicting suicidal behavior]]></category>
		<category><![CDATA[clinical evaluation of suicide risk algorithms]]></category>
		<category><![CDATA[clinical meta-analysis of suicide prediction models]]></category>
		<category><![CDATA[ensemble machine learning for mental health assessment]]></category>
		<category><![CDATA[ensemble machine learning in mental health]]></category>
		<category><![CDATA[importance of post-test probabilities in clinical decision-making]]></category>
		<category><![CDATA[limitations of static risk factors in suicide prediction]]></category>
		<category><![CDATA[Machine learning in suicide risk prediction]]></category>
		<category><![CDATA[machine learning limitations in suicide prevention]]></category>
		<category><![CDATA[machine learning model performance metrics]]></category>
		<category><![CDATA[Machine learning suicide risk prediction]]></category>
		<category><![CDATA[meta-analysis of suicide prediction models]]></category>
		<category><![CDATA[predictive values in mental health risk assessment]]></category>
		<category><![CDATA[predictive values in suicide risk assessment]]></category>
		<category><![CDATA[sensitivity and specificity in suicide prediction]]></category>
		<category><![CDATA[suicide prevention using artificial intelligence]]></category>
		<category><![CDATA[suicide risk prediction metrics]]></category>
		<category><![CDATA[systematic review of AI models in psychiatry]]></category>
		<category><![CDATA[systematic review of mental health risk prediction tools]]></category>
		<guid isPermaLink="false">https://scienmag.com/meta-analysis-evaluates-machine-learning-accuracy-in-predicting-suicide-risk/</guid>

					<description><![CDATA[Suicide remains one of the most difficult outcomes in medicine to predict, and for decades clinicians have relied on judgment, checklists and static risk factors that perform only marginally better than chance. A new systematic review and meta-analysis published in Annals of General Psychiatry now offers the most clinically interpretable picture to date of how [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Suicide remains one of the most difficult outcomes in medicine to predict, and for decades clinicians have relied on judgment, checklists and static risk factors that perform only marginally better than chance. A new systematic review and meta-analysis published in Annals of General Psychiatry now offers the most clinically interpretable picture to date of how machine learning models actually perform when tasked with forecasting suicidal ideation, suicide attempts and suicide mortality. The findings are striking in their potential and sobering in their caveats: ensemble models that combine multiple algorithms achieved an area under the receiver operating characteristic curve of 0.95, yet their pooled sensitivity was only 0.50, meaning that even the best-performing approaches in the literature miss roughly half of true positive cases. The analysis, conducted by researchers at Kurdistan University of Medical Sciences in Sanandaj, Iran, and registered prospectively in PROSPERO under number CRD420251159042, shifts the field&#8217;s focus away from raw discrimination statistics and toward measures that matter at the bedside, including sensitivity, specificity, positive and negative predictive values, likelihood ratios and post-test probabilities.</p>
<p>The urgency of the work is difficult to overstate. Traditional suicide risk assessment has a well-documented failure pattern: meta-analytic evidence spanning fifty years of research, notably the 2017 analysis by Franklin and colleagues, concluded that conventional quantitative models perform barely better than random chance. Up to 95 percent of patients classified as high risk do not die by suicide, while many of those who do were previously deemed low risk. This misclassification problem is compounded by the low base rate of suicide itself, which mathematically suppresses positive predictive value even for reasonably accurate tests. Machine learning, by contrast, can model complex, non-linear, high-dimensional relationships among predictors drawn from electronic health records, psychological assessments, clinical transcripts, speech acoustics and social media text, without assuming linearity or independence among variables. That structural flexibility is precisely why enthusiasm for algorithmic suicide prediction has surged since 2021, and why a rigorous, diagnostic-accuracy-style synthesis of the evidence became necessary.</p>
<p>To build that synthesis, the research team followed the PRISMA-DTA reporting guidelines and searched PubMed, Embase, PsycINFO and Web of Science for studies published between January 2010 and December 2024, a window chosen deliberately to coincide with the modern machine learning era and the broader availability of digitized health records. Eligible studies used single-gate designs, enrolled at least 100 participants, included a minimum of six months of follow-up so that prediction genuinely preceded outcome, and reported at least one formal diagnostic accuracy metric with internal or external validation. Reference standards had to be validated measures, such as the Columbia-Suicide Severity Rating Scale or Beck Scale for Suicide Ideation for ideation, documented attempts in registries or electronic records, and mortality confirmed through national death registries or coroner reports. From 500 screened records, 22 studies survived the full screening cascade, with two reviewers independently extracting data and disagreements resolved by consensus or third-reviewer adjudication. Risk of bias was assessed with the QUADAS-2 tool across four domains: patient selection, index test, reference standard, and flow and timing.</p>
<p>The methodological heart of the analysis is a bivariate random-effects model that pools sensitivity and specificity jointly, accounting for the inherent correlation between them, rather than averaging AUC values as many earlier reviews did. Summary receiver operating characteristic curves captured discriminative ability, while Fagan&#8217;s nomograms translated pooled likelihood ratios into post-test probabilities at a pre-test risk of 25 percent, a figure chosen to represent clinically meaningful moderate risk. The team pre-specified a taxonomy of algorithm families: logistic regression as the linear baseline, support vector machines as margin-based kernel methods, random forests as bagged tree ensembles, gradient boosting and XGBoost as boosted trees, artificial neural networks as deep learning models, and study-level stacking, blending or voting meta-models as the ensemble category. Notably, random forests and gradient boosting were excluded from the ensemble category because their learning strategies are intrinsically ensemble-based, a decision that avoids double counting and preserves clean contrasts between families.</p>
<p>The results reveal a clear hierarchy. Ensemble models topped the table with a pooled AUC of 0.95 (95 percent CI 0.92 to 0.96), specificity of 0.97 (95 percent CI 0.95 to 0.98) but sensitivity of only 0.50 (95 percent CI 0.29 to 0.71), a positive likelihood ratio of 9.2 and a negative likelihood ratio of 0.19. In clinical terms, a patient at 25 percent pre-test probability who tests positive on an ensemble model has an 88 percent post-test probability of a suicide-related outcome, the largest diagnostic shift observed. XGBoost followed closely with an AUC of 0.94, sensitivity of 0.71 and specificity of 0.85, lifting post-test probability from 25 to 86 percent with a positive likelihood ratio of 8.9. Artificial neural networks matched the 0.95 AUC but with wider confidence intervals, moderate heterogeneity in specificity and a pooled sensitivity of just 0.50 against a striking specificity of 0.97. Random forests achieved an AUC of 0.92 with sensitivity of 0.59 and specificity of 0.91, translating a positive result into an 81 percent post-test probability.</p>
<p>Support vector machines delivered the most balanced profile, with an AUC of 0.89 and nearly symmetrical pooled sensitivity and specificity of 0.77 each, yielding a post-test probability of 72 percent after a positive test and a reduction to 9 percent after a negative one. Gradient boosting showed the highest sensitivity of any family at 0.79 but the weakest specificity at 0.71, producing an AUC of 0.89 and a post-test probability of 71 percent. Logistic regression, the traditional statistical workhorse, trailed the field with an AUC of 0.86, sensitivity of just 0.54 and specificity of 0.87, raising post-test probability to only 64 percent. This pattern aligns with a well-known literature, including the 2019 systematic review by Christodoulou and colleagues, which found no consistent performance benefit of machine learning over logistic regression in clinical prediction. Linear models struggle to capture the complex interactions that characterize suicide risk, but the authors caution that differences in study design, feature selection and outcome definitions also contribute, and that direct head-to-head comparisons within individual studies were rare.</p>
<p>The asymmetry between high specificity and low sensitivity in the top-performing models carries an important clinical trade-off. In settings where minimizing false negatives is paramount, as it often is in suicide prevention, a model that confidently rules in risk but misses half of true cases has limited standalone utility. The authors frame the ensemble results as descriptive findings rather than evidence of superiority, noting that the apparent consistency of ensemble performance may reflect the ability of stacking and voting strategies to combine complementary strengths of individual algorithms, but that variations in outcome definitions, predictor sets and validation strategies across primary studies limit generalization. The QUADAS-2 assessment found generally low risk of bias for the index tests and reference standards but more pronounced variability in patient selection and flow and timing, with several studies raising concerns about sample representativeness and follow-up consistency. Threshold variation across studies, a recognized driver of heterogeneity in diagnostic test accuracy meta-analyses, was accommodated by the hierarchical models rather than treated as mere noise.</p>
<p>Beyond raw accuracy, the analysis probes whether these models are ready for deployment, and the answer is a careful no. The reviewers found that only a minority of primary studies applied formal bias-assessment tools such as PROBAST or QUADAS-2, and that data leakage, suboptimal feature selection and the absence of external validation were frequently unaddressed, all of which can inflate reported performance. Interpretability poses a parallel challenge: complex ensemble and neural network approaches operate as black boxes, raising concerns about trust, accountability and ethical deployment in psychiatry. Methods such as SHAP and LIME, which attribute predictions to individual features, were highlighted as essential adjuncts for clinician understanding and ethically informed use. On the public health side, the authors emphasize that the low base rate of suicide fundamentally limits positive predictive value at the population level, meaning that even highly specific models would generate overwhelming numbers of false positives if applied broadly. They suggest that ML tools may be most valuable in resource-constrained settings for focused screening and follow-up, always as complements to, never replacements for, clinical judgment.</p>
<p>The review is candid about its own limitations. Substantial heterogeneity in populations, data sources and model specifications could not be resolved through the planned subgroup analyses and meta-regression because primary studies rarely reported stratified results. Most critically, outcome granularity was lacking: sensitivity, specificity and AUC estimates were seldom separated for suicidal ideation, attempts and mortality, so the pooled estimates blend heterogeneous outcomes, and the pre-specified threshold of at least three studies per outcome for outcome-specific pooling was never met. Diagnosis-stratified analyses, which would have distinguished performance in depression, bipolar disorder, psychosis or personality disorders, were likewise impossible. This means the encouraging headline numbers cannot be assumed to generalize to suicide mortality specifically or to any single diagnostic group, a caveat the authors state forcefully in their conclusion.</p>
<p>What the analysis ultimately delivers is a template for what rigorous evaluation of clinical AI should look like: longitudinal designs in which predictors genuinely precede outcomes, bivariate pooling of paired sensitivity and specificity rather than AUCs alone, likelihood ratios and post-test probabilities that clinicians can act on, and transparent handling of bias and threshold effects. The message for the field is twofold. Machine learning has clearly outgrown the near-chance performance of traditional risk prediction, with several algorithm families demonstrating strong discriminative power in diagnostically diverse, longitudinal settings. Yet the gap between statistical performance and clinical readiness remains wide, bridged only by rigorous internal and external validation, calibration, outcome-resolved evaluation and attention to equity, transparency and ethics. Future research, the authors conclude, must stratify validation by diagnosis and outcome type, particularly mortality, before any of these promising models can be trusted to support life-or-death decisions.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Diagnostic accuracy of machine learning models for predicting suicide-related outcomes, including suicidal ideation, suicide attempts and suicide mortality</p>
<p><strong>Article Title:</strong> Diagnostic accuracy of machine learning approaches for suicide-related outcomes: a meta-analysis</p>
<p><strong>Article References:</strong> Kohnepoushi, P., Afraie, M., Rahmani, H., Seyedoshohadaei, S. A., Ghadirzadeh, B., &amp; Moradi, Y. (2026). Diagnostic accuracy of machine learning approaches for suicide‑related outcomes: a meta‑analysis. <em>Annals of General Psychiatry, 25</em>(1), Article 52. <a href="https://doi.org/10.1186/s12991-026-00671-4" target="_blank" rel="noopener noreferrer">https://doi.org/10.1186/s12991-026-00671-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12991-026-00671-4" target="_blank" rel="noopener noreferrer">10.1186/s12991-026-00671-4</a></p>
<p><strong>Keywords:</strong> machine learning, suicidal ideation, suicide attempts, suicide mortality, diagnostic accuracy, meta-analysis, sensitivity, specificity, ensemble models, XGBoost, likelihood ratios, suicide risk prediction</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">188792</post-id>	</item>
	</channel>
</rss>
