Saturday, October 10, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Psychology & Psychiatry

AI Learns to Spot Hidden Suicide Risk in Persian Tweets

October 10, 2026
in Psychology & Psychiatry
Glenn Wilkins
By Glenn Wilkins Scienmag Editorial Profile - Clinical Psychology
Reading Time: 5 mins read
0
AI Learns to Spot Hidden Suicide Risk in Persian Tweets

AI Learns to Spot Hidden Suicide Risk in Persian Tweets

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Suicide is one of the most urgent and least visible public-health crises in Iran and the wider Middle East, and clinicians there have long lacked scalable tools for detecting suicidal distress in Persian-speaking populations before it is too late. A new study published in Discover Mental Health offers a striking technological answer: researchers have built the first large-scale, expert-annotated Persian Twitter corpus for suicidal ideation and shown that relatively simple machine-learning classifiers can reliably distinguish tweets expressing personal suicidal thoughts from ordinary content, even when the tweets contain no explicit suicide-related vocabulary at all. The work, led by Seyed-Ali Sadegh-Zadeh of the University of Staffordshire together with collaborators at Amirkabir University of Technology, Tehran University of Medical Sciences, Urmia University of Medical Sciences, Iran University of Medical Sciences and the Brain and Cognition Clinic in Tehran, may lay the groundwork for real-time mental-health surveillance across the Persian-speaking world.

The scale of the dataset is what sets the study apart. Between January 2020 and December 2021, the team retrieved Persian tweets through Twitter’s Academic Research API using a psychiatrist-curated lexicon of 86 suicide-related and neutral query terms, deployed in two parallel retrieval streams of comparable size in a balanced case-control design. After language verification, deduplication and relevance filtering, the tweets were annotated for the presence of personal suicidal ideation by two board-certified psychiatrists. The annotation protocol was finalised in an independent double-annotation pilot phase, in which the two psychiatrists reached a Cohen’s kappa of 0.81, a level of agreement generally considered almost perfect. The final corpus comprised 197,270 tweets, split almost exactly in half: 98,669 tweets expressing suicidal ideation and 98,601 non-suicidal tweets.

On the computational side, the researchers deliberately chose interpretable, well-understood methods rather than opaque deep-learning architectures. They trained three classic classifiers, a linear support vector machine, logistic regression and a Bernoulli Naive Bayes model, on TF-IDF features built from word one- to three-grams and capped at 10,000 features. Evaluation was carried out with seeded stratified 10-fold cross-validation, supplemented by a keyword-masking ablation, a class-weighting sensitivity analysis and a probability-calibration analysis. Crucially, the entire computational workflow has been publicly released with a fixed random seed, so that the numbers can be audited and reproduced by anyone, although exact numerical regeneration requires access to the restricted corpus itself, a restriction imposed to protect the privacy of social media users.

The headline results are impressive. Logistic regression and the linear SVM performed equivalently, each achieving an F1 score of 0.933 with a standard deviation of just 0.001 across folds, and precision-recall area-under-curve values of 0.980 and 0.979 respectively. These figures indicate that the models can separate suicidal from non-suicidal tweets with very high fidelity in this corpus. But the most revealing comparison came against a keyword-lexicon baseline, the kind of simple rule-based filter that many surveillance systems still rely on. The lexicon approach achieved high precision of 0.970, meaning that when it flagged a tweet it was almost always right, yet it missed 62 percent of suicidal tweets, with a recall of only 0.378.

That gap carries a profound implication: most expressions of suicidal ideation in the corpus do not contain explicit suicide-related query terms at all. People in acute psychological distress often write about hopelessness, exhaustion, loneliness and despair in everyday language, without ever naming suicide. A keyword filter scanning for the word itself will therefore pass silently over the majority of at-risk messages. Machine-learning classifiers, by contrast, learn the broader contextual and linguistic signature of suicidal distress, picking up on combinations of words, phrases and grammatical patterns that no hand-crafted lexicon can anticipate. In a resource-constrained mental-health system, the difference between catching 38 percent of cases and catching 93 percent, as measured by the F1 metric’s balance of precision and recall, could translate into thousands of people identified for support who would otherwise have gone unnoticed.

To test whether the models were genuinely learning contextual cues rather than simply memorising the query terms used to collect the data, the team ran a keyword-masking ablation. They removed all 44 suicide-related query terms and their morphological variants from the test tweets before classification. Even under this handicap, the SVM still achieved an F1 score of 0.898, demonstrating that the classifiers rely substantially on linguistic information beyond the lexicon. This ablation matters for scientific credibility as well as performance: because the corpus was assembled using suicide-related search terms, there was a real risk of circularity, where the model might simply detect the presence of the very keywords that defined the dataset. The masking experiment largely rules that out, showing that the learned decision boundary reflects the texture of suicidal language itself.

The researchers also examined whether the models’ probability outputs could be trusted as actual risk estimates, a question that becomes critical if such tools are ever used to triage human review. The logistic regression model proved well calibrated, with a Brier score of 0.052 and a Cox calibration slope of 1.22, values indicating that a predicted probability of, say, 0.8 corresponds closely to the true proportion of suicidal tweets among cases predicted at that level. Well-calibrated probabilities make a model interpretable and actionable: a health authority could set a threshold that balances the burden of false alarms against the cost of missed cases, rather than treating the model as an inscrutable black box. The study was conducted and reported in accordance with the TRIPOD+AI statement for prediction model studies, and a completed checklist is provided in the supplementary material, reflecting a growing standard for transparency in clinical artificial intelligence.

The ethical architecture of the study is as carefully constructed as its statistical one. The protocol was approved by the Research Ethics Committee of Iran University of Medical Sciences in December 2021, in accordance with the Declaration of Helsinki. Informed consent was waived because the study was observational and retrospective, relied exclusively on publicly available posts, and could not practicably be conducted with individual consent across a dataset of roughly 197,000 tweets; no user was contacted, followed or subjected to any intervention at any stage. To protect privacy, the archived analysis dataset retains only preprocessed tweet text and the annotation label, with no tweet identifiers, usernames, profile information, geolocation data or timestamps stored. The authors also declare no competing interests, and the research received no specific grant from any funding agency in the public, commercial or not-for-profit sectors.

For all its promise, the team is explicit about the limits of what they have built. The balanced case-control design, in which suicidal and non-suicidal tweets appear in equal proportions, is ideal for training and benchmarking but wildly unrealistic for deployment, where suicidal ideation is rare among the general stream of posts. The authors caution that performance must be re-validated under realistic, low-prevalence conditions before any operational use, because precision inevitably degrades as the base rate falls. They also stress that any deployment must be embedded within human-oversight frameworks and appropriate ethical safeguards, with algorithms flagging content for trained professionals rather than making autonomous decisions about individuals. Social media platforms themselves present moving targets, as data access policies shift and platform populations change, and a model trained on 2020-2021 tweets may need updating as language and online culture evolve.

Nevertheless, the study marks a genuine milestone for low-resource-language mental-health informatics. Most suicide-detection research to date has been conducted in English, leaving the world’s hundreds of millions of Persian speakers, and much of the Middle East more broadly, without tools calibrated to their language and culture. By releasing a transparent, seeded, publicly inspectable pipeline alongside the methodology, the researchers have established a dedicated framework that other teams can extend, critique and adapt. If the promised re-validation under realistic conditions holds up, and if deployment proceeds under genuine human oversight, systems of this kind could give public-health authorities in Iran and across the Persian-speaking world something they have never had before: an early-warning capability that hears the quiet, unspoken distress of people who never use the word for what they are contemplating.

Subject of Research: Machine learning detection of suicidal ideation in Persian-language social media posts using a clinically annotated corpus

Article Title: Machine learning detection of suicidal ideation in Persian language tweets using a large scale clinically annotated corpus

Article References: Sadegh-Zadeh, S.-A., Nazari, M.-J., Khalilian, E., Mamalo, A. S., Anoosheh, S., Mousavi, S.-Y., & Shalbafan, M. (2026). Machine learning detection of suicidal ideation in Persian language tweets using a large scale clinically annotated corpus. Discover Mental Health. https://doi.org/10.1007/s44192-026-00602-5

Image Credits: AI Generated

DOI: 10.1007/s44192-026-00602-5

Keywords: suicidal ideation, machine learning, Persian language, Twitter, social media surveillance, clinical annotation, natural language processing, suicide prevention, Iran, public mental health, text classification, TRIPOD+AI

Cite Scienmag News

Glenn Wilkins. (October 10, 2026). AI Learns to Spot Hidden Suicide Risk in Persian Tweets. Scienmag. https://scienmag.com/ai-learns-to-spot-hidden-suicide-risk-in-persian-tweets/

Glenn Wilkins. "AI Learns to Spot Hidden Suicide Risk in Persian Tweets." Scienmag, 10 October 2026, https://scienmag.com/ai-learns-to-spot-hidden-suicide-risk-in-persian-tweets/. Accessed 10 October 2026.

Glenn Wilkins. "AI Learns to Spot Hidden Suicide Risk in Persian Tweets." Scienmag. October 10, 2026. https://scienmag.com/ai-learns-to-spot-hidden-suicide-risk-in-persian-tweets/

Tags: AI-based suicide risk assessmentautomated suicidal ideation identificationclinical annotationearly warning systems for suicidal behaviorexpert-annotated Persian tweet datasetIranMachine learningmachine learning for mental healthmultilingual suicide prevention technologiesnatural language processingnatural language processing for Persian textsPersian languagePersian Twitter mental health surveillancepublic mental healthreal-time mental health monitoringscalable mental health tools in Iransocial media analytics for mental healthsocial media surveillancesuicidal ideationSuicide detection in Persian social mediaSuicide Preventiontext classificationTRIPOD+AITwitter
Share26Tweet16
Previous Post

Why Malawi’s Groundwater Data Keeps Vanishing: A Systemic Diagnosis

Next Post

Mixing Secrets: How Carbon Black Distribution Shapes Solid-State Battery Performance

Related Posts

Journalists on the Brink: Study Maps Social Dysfunction Among Latin American Media Workers
Psychology & Psychiatry

Journalists on the Brink: Study Maps Social Dysfunction Among Latin American Media Workers

October 10, 2026
Heavy Thinking Eases Emotional Conflict in Non-Social Anxiety, but Not Social Anxiety
Psychology & Psychiatry

Heavy Thinking Eases Emotional Conflict in Non-Social Anxiety, but Not Social Anxiety

October 10, 2026
How YouTube Tutorials and Evening Art Classes Build Confidence in Adult Learners
Psychology & Psychiatry

How YouTube Tutorials and Evening Art Classes Build Confidence in Adult Learners

October 10, 2026
Eye Injections for Vision Loss Linked to Sudden Psychosis in One Elderly Patient
Psychology & Psychiatry

Eye Injections for Vision Loss Linked to Sudden Psychosis in One Elderly Patient

October 10, 2026
A Decade of Data Reveals the Hidden Burden of Substance-Induced Mental Disorders in US Treatment Centers
Psychology & Psychiatry

A Decade of Data Reveals the Hidden Burden of Substance-Induced Mental Disorders in US Treatment Centers

October 10, 2026
Teachers’ Inner Resources Fuel Commitment Differently for Men and Women
Psychology & Psychiatry

Teachers’ Inner Resources Fuel Commitment Differently for Men and Women

October 10, 2026
Next Post
Mixing Secrets: How Carbon Black Distribution Shapes Solid-State Battery Performance

Mixing Secrets: How Carbon Black Distribution Shapes Solid-State Battery Performance

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Hidden Metabolic Trick Lets Myeloma Cells Steal Fuel to Survive a Key Cancer Drug
  • Mixing Secrets: How Carbon Black Distribution Shapes Solid-State Battery Performance
  • AI Learns to Spot Hidden Suicide Risk in Persian Tweets
  • Why Malawi’s Groundwater Data Keeps Vanishing: A Systemic Diagnosis

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading