Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media

October 1, 2026
in Technology and Engineering
Teresa Odom
By Teresa Odom Scienmag Editorial Profile - Machine Learning
Reading Time: 5 mins read
0
Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media

Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media

Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

A team of researchers in Bangladesh has built one of the most ambitious tools yet for detecting religious hostility online, and the results upend a common assumption in modern artificial intelligence. In a study published in Discover Artificial Intelligence, Riad Hossain of East Delta University and colleagues from Chittagong University of Engineering and Technology introduce the Bangla Religious Aggression Comments dataset, or BRAC, a corpus of 20,000 manually annotated social media comments. Alongside the dataset, they present a detection system that pairs old-fashioned statistical features with an ensemble of tree-based classifiers, and it beats sophisticated Transformer models on the central task of spotting religious aggression.

The motivation is starkly practical. Bangladesh has repeatedly witnessed how inflammatory online content can ignite real-world violence. In October 2021, viral Facebook posts alleging a Quran desecration preceded attacks on Hindu homes and temples during Durga Puja celebrations. Five years earlier, the Nasirnagar violence followed an allegedly defamatory Facebook post and left Hindu neighborhoods devastated. In a country where religion is deeply woven into cultural identity and where more than 230 million people speak Bangla worldwide, the absence of effective automated moderation in that language represents a dangerous gap. Moderation systems trained on English-centric data routinely miss context-dependent insults, religious innuendos, and culturally grounded metaphors of hate.

What makes BRAC distinctive is its two-tier annotation scheme. Each comment carries a binary label marking it as aggressive or non-aggressive, and aggressive comments are further tagged with the specific religion being targeted: Muslim, Hindu, Christian, or Buddhist. The researchers argue that this fine-grained view matters because different communities face different forms of online hostility. Comments attacking Hindus during religious festivals often invoke derogatory stereotypes about idol worship, while aggression toward Muslims may center on religious attire or practices. Knowing which group is under attack allows platforms and policymakers to track rising hostility toward particular communities and potentially intervene before violence erupts.

Building the dataset demanded unusual rigor. The team collected comments from Facebook and YouTube, drawing aggressive material from threads with hostile religious discussions and non-aggressive material from news posts and respectful discussions of festivals, interfaith dialogue, and cultural events. Only threads with at least ten user reactions were included, and no keyword pre-filtering was applied, preserving the natural distribution of online discourse. Five trained annotators from diverse academic and residential backgrounds labeled the data under detailed guidelines. Agreement was measured with Cohen’s Kappa, reaching 0.91 for the binary aggression task and 0.84 for the harder religion-target task, figures the authors describe as high reliability. Roughly 500 comments lacking a clear religious target, along with duplicates and irrelevant entries, were discarded to produce the final corpus, which is perfectly balanced between aggressive and non-aggressive samples.

The technical heart of the study lies in its hybrid feature design. The researchers constructed two manually curated lexicons drawn exclusively from the training data: a religion lexicon of 1,273 root terms covering religion names, communities, practices, and institutions, and an offensive lexicon of 1,376 root terms representing abusive and derogatory expressions. These were expanded with common Bangla social media variants, including inflected forms, informal spellings, and repeated-character emphasis. Each comment was then represented by TF-IDF vectors, which capture the statistical importance of words, concatenated with counts of religion-specific and offensive terms. The authors illustrate the effect with a comment translating to Muslims must be destroyed: the words for Muslims and destroy may carry modest TF-IDF scores on their own, but the handcrafted flags boost their weight dramatically, making the aggressive intent unmistakable to the classifier.

After evaluating every pairwise combination of candidate learners, including logistic regression, support vector machines, decision trees, random forests, and XGBoost, the team settled on a soft-voting ensemble of XGBoost and Random Forest as its proposed model. The results were striking. For binary aggression detection, the ensemble achieved the highest accuracy of any model tested, 96.73 percent, while BanglaBERT, a Transformer pretrained specifically for Bangla, attained the best precision at 95.55 percent and the best F1-score at 96.06 percent. For the finer task of identifying the targeted religion, the two paradigms were nearly tied: the ensemble edged ahead in accuracy at 94.67 percent, while BanglaBERT led in F1-score at 94.81 percent. A two-proportion z-test confirmed that the ensemble’s accuracy gains over strong published baselines were statistically significant for aggression detection.

The authors offer a mathematical explanation for why sparse lexical features can outperform dense neural embeddings on moderately sized datasets. Transformer models encode text through attention-weighted averaging of token embeddings, a process that can dilute the contribution of rare but highly discriminative words such as explicit slurs or religion markers. Tree-based ensembles, by contrast, operate in a sparse feature space where each nonzero dimension corresponds to a linguistically meaningful signal, allowing decision splits that maximize information gain on exactly those cues. Ensembles also benefit from variance reduction through averaging. There is an interpretability dividend as well: unlike black-box Transformers, a model driven by counts of religion-specific and offensive terms offers transparent decision signals that domain experts and policymakers can inspect and validate, a crucial property in a sensitive domain like religious hostility.

The study also probes how well the system generalizes. In cross-dataset evaluation on an external Bangla religious hate-speech dataset, the framework maintained strong performance, achieving 94.17 percent accuracy and a 95.87 percent F1-score for aggression detection, and 95.71 percent accuracy for target religion classification. A small human analysis found the classifier correctly identified the targeted religion in 87 percent of cases where religion names were absent or obfuscated, relying on contextual cues such as the word for temple. LIME explainability visualizations confirmed that predictions rest on meaningful, aggression-relevant words rather than spurious correlations.

The error analysis is candid about limitations. The model falsely flagged a sentence condemning hatred toward Buddhists as aggressive, because offensive and religion-related words co-occurred even though the statement’s polarity was negative, revealing weak handling of negation. Conversely, a stereotyping comment about Hindus that contained no explicit offensive vocabulary slipped through undetected. The authors suggest that lightweight negation-scope features or hybrid architectures combining handcrafted cues with contextual embeddings could reduce such errors. They also acknowledge that BRAC may not capture the full diversity of Bangla dialects and evolving slang, that the lexicons require expert curation, and that multimodal content such as emojis, images, and code-mixed text remains outside the current scope.

The broader significance of the work extends beyond Bangladesh. It demonstrates that in morphologically rich, low-resource languages, carefully engineered domain knowledge can rival or exceed the raw power of large pretrained models, especially when labeled data is limited and interpretability matters. By releasing BRAC and establishing the first comprehensive benchmark spanning machine learning, deep learning, and Transformer approaches for religious aggression detection, the researchers have given computational social scientists both a resource and a methodological roadmap. Their findings point toward hybrid sparse-dense systems that combine the statistical precision of lexicon-augmented ensembles with the contextual sensitivity of Transformers, a combination that could make digital spaces safer for the world’s multi-religious communities before the next viral post turns deadly.

Subject of Research: Automated detection of religious aggression and target religion in Bangla social media text using handcrafted features and ensemble learning

Article Title: Handcrafted features and ensemble learning for religious aggression detection in Bangla social media

Article References: Hossain, R., Banu, A., Mowla, A. I. G., Rana, M. M., & Hossain, A. (2026). Handcrafted features and ensemble learning for religious aggression detection in Bangla social media. Discover Artificial Intelligence, 6(1), Article 1321. https://doi.org/10.1007/s44163-026-02215-x

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02215-x

Keywords: religious aggression detection, Bangla natural language processing, social media mining, ensemble learning, XGBoost, Random Forest, TF-IDF, BanglaBERT, hate speech detection, low-resource languages, BRAC dataset, explainable AI

Cite Scienmag News

Teresa Odom. (October 1, 2026). Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media. Scienmag. https://scienmag.com/old-school-machine-learning-outsmarts-transformers-in-detecting-religious-hate-on-bangla-social-media/

Teresa Odom. "Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media." Scienmag, 1 October 2026, https://scienmag.com/old-school-machine-learning-outsmarts-transformers-in-detecting-religious-hate-on-bangla-social-media/. Accessed 1 October 2026.

Teresa Odom. "Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media." Scienmag. October 1, 2026. https://scienmag.com/old-school-machine-learning-outsmarts-transformers-in-detecting-religious-hate-on-bangla-social-media/

Tags: Bangla natural language processingBanglaBERTBRAC datasetchallenges of automated moderation in low-resource languagesdataset creation for religious hostilitydevelopment of NLP tools for Bangla languageefficacy of old-school machine learning techniquesensemble learningensemble tree-based classifiers for hate speechexplainable AIhate speech detectionimpact of online inflammatory content on real-world violencelimitations of transformer models in religious hate detectionlinguistic and cultural considerations in hate speech detectionlow-resource languagesmanual annotation of social media commentsRandom Forestreligious aggression detectionReligious hate speech detection in Bangla social mediasocial media miningsocial media violence in BangladeshTF-IDFtraditional machine learning versus transformer modelsXGBoost
Share26Tweet16
Previous Post

Fear of Missing Out Fuels Gaming Addiction Through Impulsivity, Study Finds

Next Post

Kitchen Chemistry: Lemon Juice Powers Greener Route to Drug-Like Molecules

Related Posts

Sulfur-Tweaked Catalyst Splits Water and Destroys Antibiotics With One Material
Technology and Engineering

Sulfur-Tweaked Catalyst Splits Water and Destroys Antibiotics With One Material

October 1, 2026
AI Learns to Orchestrate Robot Arms Inside Chip Factories
Technology and Engineering

AI Learns to Orchestrate Robot Arms Inside Chip Factories

October 1, 2026
Histone readers MLLT1 and MLLT3 concentrate AID to confer locus specificity
Medicine

Histone readers MLLT1 and MLLT3 concentrate AID to confer locus specificity

October 1, 2026
Broadband Light Fingerprinting Promises Sharper Chip Overlay Metrology
Technology and Engineering

Broadband Light Fingerprinting Promises Sharper Chip Overlay Metrology

October 1, 2026
Robotic Hip and Knee Replacements Show No Clear Advantage Over Conventional Surgery in Landmark UK Study
Technology and Engineering

Robotic Hip and Knee Replacements Show No Clear Advantage Over Conventional Surgery in Landmark UK Study

October 1, 2026
Sticky Gel Carrying Supercharged Stem Cell Vesicles Repairs Burned Esophagus in Rats
Technology and Engineering

Sticky Gel Carrying Supercharged Stem Cell Vesicles Repairs Burned Esophagus in Rats

October 1, 2026
Next Post
Kitchen Chemistry: Lemon Juice Powers Greener Route to Drug-Like Molecules

Kitchen Chemistry: Lemon Juice Powers Greener Route to Drug-Like Molecules

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • New Family-Trio Genetic Method Exposes Hidden Bias in Causal Gene Studies
  • Kitchen Chemistry: Lemon Juice Powers Greener Route to Drug-Like Molecules
  • Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media
  • Fear of Missing Out Fuels Gaming Addiction Through Impulsivity, Study Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading