<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Multilingual fake news detection &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/multilingual-fake-news-detection/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 10:57:02 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Multilingual fake news detection &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories</title>
		<link>https://scienmag.com/attention-based-ai-reads-news-word-by-word-to-catch-multilingual-fake-stories/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 10:57:02 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI accuracy in fake news classification]]></category>
		<category><![CDATA[AI-based multilingual news verification]]></category>
		<category><![CDATA[automated fake news detection in Indian languages]]></category>
		<category><![CDATA[Bi-LSTM]]></category>
		<category><![CDATA[challenges of misinformation in diverse languages]]></category>
		<category><![CDATA[COVID-19 misinformation in regional languages]]></category>
		<category><![CDATA[cross-lingual misinformation identification]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[fake news detection]]></category>
		<category><![CDATA[fake news impact on public health]]></category>
		<category><![CDATA[FastText embeddings]]></category>
		<category><![CDATA[hierarchical attention network]]></category>
		<category><![CDATA[Hierarchical Attention Network for misinformation]]></category>
		<category><![CDATA[Hindi language]]></category>
		<category><![CDATA[low-resource languages]]></category>
		<category><![CDATA[Marathi language]]></category>
		<category><![CDATA[misinformation]]></category>
		<category><![CDATA[multilingual AI news analysis]]></category>
		<category><![CDATA[Multilingual fake news detection]]></category>
		<category><![CDATA[multilingual NLP]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[scalable solutions for multilingual misinformation]]></category>
		<category><![CDATA[social media misinformation in Hindi and Marathi]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=253369</guid>

					<description><![CDATA[Researchers have built a hierarchical attention network that detects fake news in English, Hindi, and Marathi, reaching 98.9 percent accuracy and introducing a new Marathi fake news dataset.]]></description>
										<content:encoded><![CDATA[<p>Fake news does not respect borders, and it certainly does not respect languages. While social media platforms have invested heavily in automated misinformation detection for English content, the vast majority of the world&#8217;s languages remain poorly served by these systems. A new study published in Discover Artificial Intelligence by Sushama Nandgaonkar and Sunil Mane of COEP Technological University in Pune tackles this gap head-on, presenting a Hierarchical Attention Network that detects false news across English, Hindi, and Marathi with remarkable accuracy, achieving 98.9 percent accuracy and a 98.8 percent F1-score on a combined multilingual test set.</p>
<p>The stakes are particularly high in India, which has the second-largest population of internet users worldwide. According to figures cited in the study, India had more than 375 million social media users in 2019, rising to 518 million in 2020, with projections suggesting the number could reach 1.5 billion by 2040. Affordable mobile data has driven this explosive growth, and platforms such as Facebook, Twitter, and WhatsApp now carry content in dozens of languages. Misinformation that circulates in regional languages can provoke local sentiments and influence health decisions, as demonstrated during the COVID-19 pandemic, when false claims about remedies and vaccines spread faster than official corrections could keep up.</p>
<p>The researchers distinguish carefully between related concepts that are often conflated. Disinformation refers to deliberately created and spread falsehoods, rumors are unverified pieces of information circulated with intent to mislead, and misinformation is false content shared without necessarily intending to deceive. The motivations behind such content range from spreading religious hatred to gaining political advantage or disseminating false health information. Because fact-checking websites themselves can carry biases, and because human verification cannot scale to billions of posts, automated detection systems have become an essential line of defense, and the study argues that these systems must work across the languages people actually use.</p>
<p>At the heart of the new work is a two-level attention architecture that mirrors how humans read documents. The model first processes each sentence word by word using a Bidirectional Long Short-Term Memory network, which reads text in both forward and backward directions to capture context from either side of every word. A custom attention layer then assigns dynamic weights to individual words, learning during training which terms matter most for distinguishing fake from genuine content. The weighted word representations are combined into sentence vectors, and a second Bi-LSTM layer with its own attention mechanism models relationships between sentences, producing a document-level representation that feeds into a final sigmoid classifier.</p>
<p>This hierarchical design is not merely an engineering convenience. Fake news articles often contain misleading lexical patterns, emotionally polarized expressions, and contextual inconsistencies distributed across multiple sentences rather than concentrated in any single phrase. By attending to both words and sentences, the model can capture local semantic cues and document-level structure simultaneously. The attention weights also improve interpretability, allowing researchers to see which parts of an article drove a classification decision, a significant advantage over opaque black-box approaches.</p>
<p>A crucial contribution of the study is linguistic rather than architectural. Because no publicly available fake news dataset existed for Marathi, the team built one from scratch, collecting 1,957 news articles from Marathi outlets including Loksatta, Lokmat, Maharashtra Times, and Zee News, of which 707 were labeled fake and 1,250 true. Labels were assigned based on source credibility, fact-checking reports, and consistency across multiple platforms, with manual review of every article. For Hindi, the researchers combined two existing datasets from Kaggle and GitHub, while English experiments used the widely adopted ISOT dataset. To address severe class imbalance in the low-resource languages, the team applied translation-based augmentation using the IndicTrans neural translation framework, generating additional training samples while keeping the test set untouched to prevent data leakage. The final augmented corpus contained 45,386 English, 10,481 Hindi, and 5,000 Marathi samples.</p>
<p>The choice of word embeddings proved important. FastText, developed by Facebook&#8217;s AI Research lab, represents each word as a bag of character n-grams, allowing it to capture morphological information and generate vectors for out-of-vocabulary words from their sub-word components. This property is especially valuable for morphologically rich languages like Hindi and Marathi, where word forms vary extensively. The researchers compared FastText against GloVe and random initialization under identical hyperparameters, finding that pretrained embeddings substantially improved performance, with FastText also converging fastest at 383.06 seconds of training time compared with 501.22 seconds for GloVe.</p>
<p>The experimental results were striking. The HAN model achieved 99.42 percent accuracy on English news, 99.04 percent on Hindi, and 80.90 percent on Marathi, outperforming traditional machine learning baselines including Logistic Regression, Linear Support Vector Machines, and Random Forest, as well as deep learning models such as CNN and standalone Bi-LSTM. Against multilingual transformers, the picture was more nuanced: mBERT reached 98.72 percent overall accuracy and DeBERTa-v3-base 98.66 percent, slightly below the proposed model&#8217;s 98.90 percent, but DeBERTa&#8217;s performance collapsed to 74.42 percent on Marathi, exposing the sensitivity of large pretrained transformers to data imbalance. Notably, the instruction-tuned LLaMA 3.1 8B model, evaluated through prompting rather than fine-tuning, managed only 67.98 percent accuracy in zero-shot settings and 84.76 percent with six examples in context, with 1,638 failed predictions out of 10,965 test instances.</p>
<p>Ablation experiments confirmed that both attention levels contribute meaningfully. A Bi-LSTM-only model achieved an F1-score of 97.43 percent, rising to 97.82 percent with word-level attention alone and 98.17 percent with sentence-level attention alone, while the complete hierarchical architecture reached 98.72 percent. Sentence-level attention proved more influential than word-level attention in isolation, suggesting that contextual dependencies across sentences play a decisive role in identifying deceptive content. Paired t-tests across multiple random seeds showed all improvements over baselines were statistically significant, with p-values below 0.01.</p>
<p>The study is candid about its limitations. Marathi&#8217;s error rate of 19.10 percent, compared with just 0.58 percent for English and 0.96 percent for Hindi, reflects the challenge of low-resource detection, and qualitative analysis showed that the worst failures involved fake articles written in a style so similar to legitimate news that the model assigned incorrect predictions with high confidence. The authors note that generalization to other Indic languages remains unexplored and that multimodal signals such as images and metadata are not yet incorporated. Future work will expand the dataset to additional Indian regional languages and integrate textual and visual information using transformer-based Indic language models. For now, the research demonstrates that carefully designed attention mechanisms, paired with sub-word embeddings and thoughtful data augmentation, can bring state-of-the-art misinformation detection to languages that have long been left behind.</p>
<p><strong>Subject of Research:</strong> Multilingual fake news detection using hierarchical attention networks for English, Hindi, and Marathi news articles</p>
<p><strong>Article Title:</strong> Hierarchical attention network for multilingual fake news detection</p>
<p><strong>Article References:</strong> Nandgaonkar, S., &amp; Mane, S. (2026). Hierarchical attention network for multilingual fake news detection. <em>Discover Artificial Intelligence, 6</em>(1), Article 1417. <a href="https://doi.org/10.1007/s44163-026-02425-3" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02425-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02425-3" rel="noopener noreferrer">10.1007/s44163-026-02425-3</a></p>
<p><strong>Keywords:</strong> fake news detection, hierarchical attention network, multilingual NLP, Bi-LSTM, FastText embeddings, Marathi language, Hindi language, misinformation, deep learning, low-resource languages, data augmentation, natural language processing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">253369</post-id>	</item>
	</channel>
</rss>
