<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>efficacy of old-school machine learning techniques &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/efficacy-of-old-school-machine-learning-techniques/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 11:20:56 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>efficacy of old-school machine learning techniques &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Old-School Machine Learning Outsmarts Transformers in Detecting Religious Hate on Bangla Social Media</title>
		<link>https://scienmag.com/old-school-machine-learning-outsmarts-transformers-in-detecting-religious-hate-on-bangla-social-media/</link>
		
		<dc:creator><![CDATA[Teresa Odom]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 11:20:56 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Bangla natural language processing]]></category>
		<category><![CDATA[BanglaBERT]]></category>
		<category><![CDATA[BRAC dataset]]></category>
		<category><![CDATA[challenges of automated moderation in low-resource languages]]></category>
		<category><![CDATA[dataset creation for religious hostility]]></category>
		<category><![CDATA[development of NLP tools for Bangla language]]></category>
		<category><![CDATA[efficacy of old-school machine learning techniques]]></category>
		<category><![CDATA[ensemble learning]]></category>
		<category><![CDATA[ensemble tree-based classifiers for hate speech]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[hate speech detection]]></category>
		<category><![CDATA[impact of online inflammatory content on real-world violence]]></category>
		<category><![CDATA[limitations of transformer models in religious hate detection]]></category>
		<category><![CDATA[linguistic and cultural considerations in hate speech detection]]></category>
		<category><![CDATA[low-resource languages]]></category>
		<category><![CDATA[manual annotation of social media comments]]></category>
		<category><![CDATA[Random Forest]]></category>
		<category><![CDATA[religious aggression detection]]></category>
		<category><![CDATA[Religious hate speech detection in Bangla social media]]></category>
		<category><![CDATA[social media mining]]></category>
		<category><![CDATA[social media violence in Bangladesh]]></category>
		<category><![CDATA[TF-IDF]]></category>
		<category><![CDATA[traditional machine learning versus transformer models]]></category>
		<category><![CDATA[XGBoost]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222318</guid>

					<description><![CDATA[Researchers in Bangladesh have released a 20,000-comment dataset for detecting religious aggression in Bangla and shown that a lexicon-boosted XGBoost and Random Forest ensemble can outperform Transformer models on the task.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers in Bangladesh has built one of the most ambitious tools yet for detecting religious hostility online, and the results upend a common assumption in modern artificial intelligence. In a study published in Discover Artificial Intelligence, Riad Hossain of East Delta University and colleagues from Chittagong University of Engineering and Technology introduce the Bangla Religious Aggression Comments dataset, or BRAC, a corpus of 20,000 manually annotated social media comments. Alongside the dataset, they present a detection system that pairs old-fashioned statistical features with an ensemble of tree-based classifiers, and it beats sophisticated Transformer models on the central task of spotting religious aggression.</p>
<p>The motivation is starkly practical. Bangladesh has repeatedly witnessed how inflammatory online content can ignite real-world violence. In October 2021, viral Facebook posts alleging a Quran desecration preceded attacks on Hindu homes and temples during Durga Puja celebrations. Five years earlier, the Nasirnagar violence followed an allegedly defamatory Facebook post and left Hindu neighborhoods devastated. In a country where religion is deeply woven into cultural identity and where more than 230 million people speak Bangla worldwide, the absence of effective automated moderation in that language represents a dangerous gap. Moderation systems trained on English-centric data routinely miss context-dependent insults, religious innuendos, and culturally grounded metaphors of hate.</p>
<p>What makes BRAC distinctive is its two-tier annotation scheme. Each comment carries a binary label marking it as aggressive or non-aggressive, and aggressive comments are further tagged with the specific religion being targeted: Muslim, Hindu, Christian, or Buddhist. The researchers argue that this fine-grained view matters because different communities face different forms of online hostility. Comments attacking Hindus during religious festivals often invoke derogatory stereotypes about idol worship, while aggression toward Muslims may center on religious attire or practices. Knowing which group is under attack allows platforms and policymakers to track rising hostility toward particular communities and potentially intervene before violence erupts.</p>
<p>Building the dataset demanded unusual rigor. The team collected comments from Facebook and YouTube, drawing aggressive material from threads with hostile religious discussions and non-aggressive material from news posts and respectful discussions of festivals, interfaith dialogue, and cultural events. Only threads with at least ten user reactions were included, and no keyword pre-filtering was applied, preserving the natural distribution of online discourse. Five trained annotators from diverse academic and residential backgrounds labeled the data under detailed guidelines. Agreement was measured with Cohen&#8217;s Kappa, reaching 0.91 for the binary aggression task and 0.84 for the harder religion-target task, figures the authors describe as high reliability. Roughly 500 comments lacking a clear religious target, along with duplicates and irrelevant entries, were discarded to produce the final corpus, which is perfectly balanced between aggressive and non-aggressive samples.</p>
<p>The technical heart of the study lies in its hybrid feature design. The researchers constructed two manually curated lexicons drawn exclusively from the training data: a religion lexicon of 1,273 root terms covering religion names, communities, practices, and institutions, and an offensive lexicon of 1,376 root terms representing abusive and derogatory expressions. These were expanded with common Bangla social media variants, including inflected forms, informal spellings, and repeated-character emphasis. Each comment was then represented by TF-IDF vectors, which capture the statistical importance of words, concatenated with counts of religion-specific and offensive terms. The authors illustrate the effect with a comment translating to Muslims must be destroyed: the words for Muslims and destroy may carry modest TF-IDF scores on their own, but the handcrafted flags boost their weight dramatically, making the aggressive intent unmistakable to the classifier.</p>
<p>After evaluating every pairwise combination of candidate learners, including logistic regression, support vector machines, decision trees, random forests, and XGBoost, the team settled on a soft-voting ensemble of XGBoost and Random Forest as its proposed model. The results were striking. For binary aggression detection, the ensemble achieved the highest accuracy of any model tested, 96.73 percent, while BanglaBERT, a Transformer pretrained specifically for Bangla, attained the best precision at 95.55 percent and the best F1-score at 96.06 percent. For the finer task of identifying the targeted religion, the two paradigms were nearly tied: the ensemble edged ahead in accuracy at 94.67 percent, while BanglaBERT led in F1-score at 94.81 percent. A two-proportion z-test confirmed that the ensemble&#8217;s accuracy gains over strong published baselines were statistically significant for aggression detection.</p>
<p>The authors offer a mathematical explanation for why sparse lexical features can outperform dense neural embeddings on moderately sized datasets. Transformer models encode text through attention-weighted averaging of token embeddings, a process that can dilute the contribution of rare but highly discriminative words such as explicit slurs or religion markers. Tree-based ensembles, by contrast, operate in a sparse feature space where each nonzero dimension corresponds to a linguistically meaningful signal, allowing decision splits that maximize information gain on exactly those cues. Ensembles also benefit from variance reduction through averaging. There is an interpretability dividend as well: unlike black-box Transformers, a model driven by counts of religion-specific and offensive terms offers transparent decision signals that domain experts and policymakers can inspect and validate, a crucial property in a sensitive domain like religious hostility.</p>
<p>The study also probes how well the system generalizes. In cross-dataset evaluation on an external Bangla religious hate-speech dataset, the framework maintained strong performance, achieving 94.17 percent accuracy and a 95.87 percent F1-score for aggression detection, and 95.71 percent accuracy for target religion classification. A small human analysis found the classifier correctly identified the targeted religion in 87 percent of cases where religion names were absent or obfuscated, relying on contextual cues such as the word for temple. LIME explainability visualizations confirmed that predictions rest on meaningful, aggression-relevant words rather than spurious correlations.</p>
<p>The error analysis is candid about limitations. The model falsely flagged a sentence condemning hatred toward Buddhists as aggressive, because offensive and religion-related words co-occurred even though the statement&#8217;s polarity was negative, revealing weak handling of negation. Conversely, a stereotyping comment about Hindus that contained no explicit offensive vocabulary slipped through undetected. The authors suggest that lightweight negation-scope features or hybrid architectures combining handcrafted cues with contextual embeddings could reduce such errors. They also acknowledge that BRAC may not capture the full diversity of Bangla dialects and evolving slang, that the lexicons require expert curation, and that multimodal content such as emojis, images, and code-mixed text remains outside the current scope.</p>
<p>The broader significance of the work extends beyond Bangladesh. It demonstrates that in morphologically rich, low-resource languages, carefully engineered domain knowledge can rival or exceed the raw power of large pretrained models, especially when labeled data is limited and interpretability matters. By releasing BRAC and establishing the first comprehensive benchmark spanning machine learning, deep learning, and Transformer approaches for religious aggression detection, the researchers have given computational social scientists both a resource and a methodological roadmap. Their findings point toward hybrid sparse-dense systems that combine the statistical precision of lexicon-augmented ensembles with the contextual sensitivity of Transformers, a combination that could make digital spaces safer for the world&#8217;s multi-religious communities before the next viral post turns deadly.</p>
<p><strong>Subject of Research:</strong> Automated detection of religious aggression and target religion in Bangla social media text using handcrafted features and ensemble learning</p>
<p><strong>Article Title:</strong> Handcrafted features and ensemble learning for religious aggression detection in Bangla social media</p>
<p><strong>Article References:</strong> Hossain, R., Banu, A., Mowla, A. I. G., Rana, M. M., &amp; Hossain, A. (2026). Handcrafted features and ensemble learning for religious aggression detection in Bangla social media. <em>Discover Artificial Intelligence, 6</em>(1), Article 1321. <a href="https://doi.org/10.1007/s44163-026-02215-x" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02215-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02215-x" rel="noopener noreferrer">10.1007/s44163-026-02215-x</a></p>
<p><strong>Keywords:</strong> religious aggression detection, Bangla natural language processing, social media mining, ensemble learning, XGBoost, Random Forest, TF-IDF, BanglaBERT, hate speech detection, low-resource languages, BRAC dataset, explainable AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222318</post-id>	</item>
	</channel>
</rss>
