<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>natural language processing for Persian texts &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/natural-language-processing-for-persian-texts/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 10 Oct 2026 23:38:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>natural language processing for Persian texts &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Spot Hidden Suicide Risk in Persian Tweets</title>
		<link>https://scienmag.com/ai-learns-to-spot-hidden-suicide-risk-in-persian-tweets/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 10 Oct 2026 23:38:36 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI-based suicide risk assessment]]></category>
		<category><![CDATA[automated suicidal ideation identification]]></category>
		<category><![CDATA[clinical annotation]]></category>
		<category><![CDATA[early warning systems for suicidal behavior]]></category>
		<category><![CDATA[expert-annotated Persian tweet dataset]]></category>
		<category><![CDATA[Iran]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for mental health]]></category>
		<category><![CDATA[multilingual suicide prevention technologies]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing for Persian texts]]></category>
		<category><![CDATA[Persian language]]></category>
		<category><![CDATA[Persian Twitter mental health surveillance]]></category>
		<category><![CDATA[public mental health]]></category>
		<category><![CDATA[real-time mental health monitoring]]></category>
		<category><![CDATA[scalable mental health tools in Iran]]></category>
		<category><![CDATA[social media analytics for mental health]]></category>
		<category><![CDATA[social media surveillance]]></category>
		<category><![CDATA[suicidal ideation]]></category>
		<category><![CDATA[Suicide detection in Persian social media]]></category>
		<category><![CDATA[Suicide Prevention]]></category>
		<category><![CDATA[text classification]]></category>
		<category><![CDATA[TRIPOD+AI]]></category>
		<category><![CDATA[Twitter]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=260326</guid>

					<description><![CDATA[Researchers have built the first large-scale, psychiatrist-annotated Persian Twitter corpus for suicidal ideation and shown that machine-learning classifiers detect personal suicidal distress with high accuracy even when tweets contain no explicit suicide-related terms.]]></description>
										<content:encoded><![CDATA[<p>Suicide is one of the most urgent and least visible public-health crises in Iran and the wider Middle East, and clinicians there have long lacked scalable tools for detecting suicidal distress in Persian-speaking populations before it is too late. A new study published in Discover Mental Health offers a striking technological answer: researchers have built the first large-scale, expert-annotated Persian Twitter corpus for suicidal ideation and shown that relatively simple machine-learning classifiers can reliably distinguish tweets expressing personal suicidal thoughts from ordinary content, even when the tweets contain no explicit suicide-related vocabulary at all. The work, led by Seyed-Ali Sadegh-Zadeh of the University of Staffordshire together with collaborators at Amirkabir University of Technology, Tehran University of Medical Sciences, Urmia University of Medical Sciences, Iran University of Medical Sciences and the Brain and Cognition Clinic in Tehran, may lay the groundwork for real-time mental-health surveillance across the Persian-speaking world.</p>
<p>The scale of the dataset is what sets the study apart. Between January 2020 and December 2021, the team retrieved Persian tweets through Twitter&#8217;s Academic Research API using a psychiatrist-curated lexicon of 86 suicide-related and neutral query terms, deployed in two parallel retrieval streams of comparable size in a balanced case-control design. After language verification, deduplication and relevance filtering, the tweets were annotated for the presence of personal suicidal ideation by two board-certified psychiatrists. The annotation protocol was finalised in an independent double-annotation pilot phase, in which the two psychiatrists reached a Cohen&#8217;s kappa of 0.81, a level of agreement generally considered almost perfect. The final corpus comprised 197,270 tweets, split almost exactly in half: 98,669 tweets expressing suicidal ideation and 98,601 non-suicidal tweets.</p>
<p>On the computational side, the researchers deliberately chose interpretable, well-understood methods rather than opaque deep-learning architectures. They trained three classic classifiers, a linear support vector machine, logistic regression and a Bernoulli Naive Bayes model, on TF-IDF features built from word one- to three-grams and capped at 10,000 features. Evaluation was carried out with seeded stratified 10-fold cross-validation, supplemented by a keyword-masking ablation, a class-weighting sensitivity analysis and a probability-calibration analysis. Crucially, the entire computational workflow has been publicly released with a fixed random seed, so that the numbers can be audited and reproduced by anyone, although exact numerical regeneration requires access to the restricted corpus itself, a restriction imposed to protect the privacy of social media users.</p>
<p>The headline results are impressive. Logistic regression and the linear SVM performed equivalently, each achieving an F1 score of 0.933 with a standard deviation of just 0.001 across folds, and precision-recall area-under-curve values of 0.980 and 0.979 respectively. These figures indicate that the models can separate suicidal from non-suicidal tweets with very high fidelity in this corpus. But the most revealing comparison came against a keyword-lexicon baseline, the kind of simple rule-based filter that many surveillance systems still rely on. The lexicon approach achieved high precision of 0.970, meaning that when it flagged a tweet it was almost always right, yet it missed 62 percent of suicidal tweets, with a recall of only 0.378.</p>
<p>That gap carries a profound implication: most expressions of suicidal ideation in the corpus do not contain explicit suicide-related query terms at all. People in acute psychological distress often write about hopelessness, exhaustion, loneliness and despair in everyday language, without ever naming suicide. A keyword filter scanning for the word itself will therefore pass silently over the majority of at-risk messages. Machine-learning classifiers, by contrast, learn the broader contextual and linguistic signature of suicidal distress, picking up on combinations of words, phrases and grammatical patterns that no hand-crafted lexicon can anticipate. In a resource-constrained mental-health system, the difference between catching 38 percent of cases and catching 93 percent, as measured by the F1 metric&#8217;s balance of precision and recall, could translate into thousands of people identified for support who would otherwise have gone unnoticed.</p>
<p>To test whether the models were genuinely learning contextual cues rather than simply memorising the query terms used to collect the data, the team ran a keyword-masking ablation. They removed all 44 suicide-related query terms and their morphological variants from the test tweets before classification. Even under this handicap, the SVM still achieved an F1 score of 0.898, demonstrating that the classifiers rely substantially on linguistic information beyond the lexicon. This ablation matters for scientific credibility as well as performance: because the corpus was assembled using suicide-related search terms, there was a real risk of circularity, where the model might simply detect the presence of the very keywords that defined the dataset. The masking experiment largely rules that out, showing that the learned decision boundary reflects the texture of suicidal language itself.</p>
<p>The researchers also examined whether the models&#8217; probability outputs could be trusted as actual risk estimates, a question that becomes critical if such tools are ever used to triage human review. The logistic regression model proved well calibrated, with a Brier score of 0.052 and a Cox calibration slope of 1.22, values indicating that a predicted probability of, say, 0.8 corresponds closely to the true proportion of suicidal tweets among cases predicted at that level. Well-calibrated probabilities make a model interpretable and actionable: a health authority could set a threshold that balances the burden of false alarms against the cost of missed cases, rather than treating the model as an inscrutable black box. The study was conducted and reported in accordance with the TRIPOD+AI statement for prediction model studies, and a completed checklist is provided in the supplementary material, reflecting a growing standard for transparency in clinical artificial intelligence.</p>
<p>The ethical architecture of the study is as carefully constructed as its statistical one. The protocol was approved by the Research Ethics Committee of Iran University of Medical Sciences in December 2021, in accordance with the Declaration of Helsinki. Informed consent was waived because the study was observational and retrospective, relied exclusively on publicly available posts, and could not practicably be conducted with individual consent across a dataset of roughly 197,000 tweets; no user was contacted, followed or subjected to any intervention at any stage. To protect privacy, the archived analysis dataset retains only preprocessed tweet text and the annotation label, with no tweet identifiers, usernames, profile information, geolocation data or timestamps stored. The authors also declare no competing interests, and the research received no specific grant from any funding agency in the public, commercial or not-for-profit sectors.</p>
<p>For all its promise, the team is explicit about the limits of what they have built. The balanced case-control design, in which suicidal and non-suicidal tweets appear in equal proportions, is ideal for training and benchmarking but wildly unrealistic for deployment, where suicidal ideation is rare among the general stream of posts. The authors caution that performance must be re-validated under realistic, low-prevalence conditions before any operational use, because precision inevitably degrades as the base rate falls. They also stress that any deployment must be embedded within human-oversight frameworks and appropriate ethical safeguards, with algorithms flagging content for trained professionals rather than making autonomous decisions about individuals. Social media platforms themselves present moving targets, as data access policies shift and platform populations change, and a model trained on 2020-2021 tweets may need updating as language and online culture evolve.</p>
<p>Nevertheless, the study marks a genuine milestone for low-resource-language mental-health informatics. Most suicide-detection research to date has been conducted in English, leaving the world&#8217;s hundreds of millions of Persian speakers, and much of the Middle East more broadly, without tools calibrated to their language and culture. By releasing a transparent, seeded, publicly inspectable pipeline alongside the methodology, the researchers have established a dedicated framework that other teams can extend, critique and adapt. If the promised re-validation under realistic conditions holds up, and if deployment proceeds under genuine human oversight, systems of this kind could give public-health authorities in Iran and across the Persian-speaking world something they have never had before: an early-warning capability that hears the quiet, unspoken distress of people who never use the word for what they are contemplating.</p>
<p><strong>Subject of Research:</strong> Machine learning detection of suicidal ideation in Persian-language social media posts using a clinically annotated corpus</p>
<p><strong>Article Title:</strong> Machine learning detection of suicidal ideation in Persian language tweets using a large scale clinically annotated corpus</p>
<p><strong>Article References:</strong> Sadegh-Zadeh, S.-A., Nazari, M.-J., Khalilian, E., Mamalo, A. S., Anoosheh, S., Mousavi, S.-Y., &amp; Shalbafan, M. (2026). Machine learning detection of suicidal ideation in Persian language tweets using a large scale clinically annotated corpus. <em>Discover Mental Health</em>. <a href="https://doi.org/10.1007/s44192-026-00602-5" rel="noopener noreferrer">https://doi.org/10.1007/s44192-026-00602-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44192-026-00602-5" rel="noopener noreferrer">10.1007/s44192-026-00602-5</a></p>
<p><strong>Keywords:</strong> suicidal ideation, machine learning, Persian language, Twitter, social media surveillance, clinical annotation, natural language processing, suicide prevention, Iran, public mental health, text classification, TRIPOD+AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">260326</post-id>	</item>
	</channel>
</rss>
