<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Fake news detection in India &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/fake-news-detection-in-india/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 10:22:50 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Fake news detection in India &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Custom Fine-Tuned Transformer Outperforms GPT-4 at Spotting Indian Fake News</title>
		<link>https://scienmag.com/custom-fine-tuned-transformer-outperforms-gpt-4-at-spotting-indian-fake-news/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 10:22:50 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[attention mechanisms]]></category>
		<category><![CDATA[BERT]]></category>
		<category><![CDATA[challenges of automated fact-checking in India]]></category>
		<category><![CDATA[customized transformer models for multilingual content]]></category>
		<category><![CDATA[fake news detection]]></category>
		<category><![CDATA[Fake news detection in India]]></category>
		<category><![CDATA[fine-tuned NLP models for Indian social media]]></category>
		<category><![CDATA[fine-tuning]]></category>
		<category><![CDATA[GPT-4]]></category>
		<category><![CDATA[India]]></category>
		<category><![CDATA[Indian digital media fake news detection techniques]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in regional misinformation detection]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning vs deep learning in misinformation]]></category>
		<category><![CDATA[misinformation]]></category>
		<category><![CDATA[multilingual fake news identification]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[outperforming GPT-4 in fake news detection]]></category>
		<category><![CDATA[semantic pattern recognition in Indian news]]></category>
		<category><![CDATA[text classification]]></category>
		<category><![CDATA[transformer language models for misinformation]]></category>
		<category><![CDATA[transformer models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=227131</guid>

					<description><![CDATA[Researchers in India have built a custom fine-tuned BERT-based framework that outperforms classical machine learning, deep learning baselines, and GPT-3.5 and GPT-4 at detecting fake news across a benchmark dataset of 71,898 Indian news articles.]]></description>
										<content:encoded><![CDATA[<p>Fake news has become one of the most corrosive byproducts of the social media era, eroding public trust, distorting democratic debate, and overwhelming fact-checkers who cannot keep pace with the sheer volume of fabricated content circulating online. In India, where hundreds of millions of users consume news through messaging apps and social platforms in multiple languages, the problem is particularly acute. Now, a team of researchers from Yenepoya Institute of Technology, KIIT Deemed to be University, DRIEMS University, and Nitte University reports in the International Journal of Machine Learning and Cybernetics that a carefully customized and fine-tuned transformer language model can outperform both classical machine learning pipelines and some of the world&#8217;s most powerful large language models at detecting fabricated Indian news articles.</p>
<p>The study, led by Manish Prajapati and supervised by Santos Kumar Baliarsingh, addresses a persistent weakness in automated misinformation detection: most existing systems were built on English-language corpora drawn from Western news ecosystems, and they struggle to capture the contextual and semantic patterns that characterize deceptive content in Indian digital media. Traditional machine learning approaches such as logistic regression, Naïve Bayes, and random forests rely on shallow statistical features, while recurrent architectures like LSTM and GRU networks, though better at modeling sequences, still fall short when it comes to understanding the subtle linguistic cues that separate genuine reporting from fabrication.</p>
<p>At the heart of the new framework is a customized Bidirectional Encoder Representations from Transformers, or BERT, architecture. Transformers process text through self-attention mechanisms, which allow the model to weigh the relationships between every word in a document simultaneously rather than reading it in a fixed order. This bidirectional attention means the model can detect, for example, when a sensational headline is contradicted by the body of an article, or when emotionally charged language appears in contexts where neutral reporting would be expected. The researchers applied domain-adaptive fine-tuning, adjusting the pretrained model&#8217;s internal representations to the specific vocabulary, phrasing, and rhetorical conventions of Indian news content, which differs markedly from the text distributions the original model was trained on.</p>
<p>A crucial ingredient in the work is data. The team assembled a large-scale benchmark dataset of 71,898 Indian news articles by combining multiple publicly available sources of both real and fake news, spanning politics, governance, social issues, and misinformation-related content. Dataset quality is often the bottleneck in misinformation research, since labels are frequently noisy or inconsistent. To tackle this, the researchers built an annotation pipeline that combines large language model assistance with human verification. An LLM performs an initial pass over the articles, flagging likely labels, and human annotators then verify the results, a hybrid approach designed to combine the speed of automation with the judgment of expert reviewers.</p>
<p>Before training, the corpus underwent advanced preprocessing, including text normalization, tokenization, and the generation of contextual embeddings. These steps convert raw news text into numerical representations that preserve meaning and context, giving the model a cleaner and more informative signal to learn from. The data was then split using a stratified train-validation-test strategy, which ensures that the proportion of real and fake articles remains consistent across each subset, a methodological safeguard that prevents the model from being evaluated on an unrepresentative sample.</p>
<p>The benchmarking exercise was unusually comprehensive. The custom BERT framework was pitted against classical machine learning models, deep learning baselines including LSTM and GRU networks, pretrained BERT variants, and prompt-based classification using GPT-3.5 and GPT-4. The results were striking: the fine-tuned framework achieved superior performance across accuracy, precision, recall, F1-score, and ROC-AUC, the standard metrics for classification quality. Perhaps most notably, it beat the general-purpose large language models despite their enormous scale, a finding that reinforces a growing consensus in the field: a smaller model fine-tuned on domain-specific data can outperform a much larger model that relies only on prompting.</p>
<p>The reasons are technical as much as practical. Prompt-based classification asks a general-purpose model to apply knowledge it acquired during broad pretraining, without ever updating its weights for the task at hand. Fine-tuning, by contrast, adjusts the model&#8217;s parameters directly on labeled examples from the target domain, teaching it the idiosyncratic markers of deception in Indian news, from particular narrative structures to characteristic vocabulary shifts. The custom framework also maintains computational efficiency and scalability, meaning it can plausibly be deployed for real-time monitoring rather than confined to laboratory experiments, an important consideration when misinformation spreads within minutes of publication.</p>
<p>Beyond raw accuracy, the researchers invested in interpretability, one of the most pressing concerns in applying artificial intelligence to content moderation. Using attention visualization, they examined which parts of a news article the model focused on when making its decision. The analysis revealed that the model latches onto misinformation-related linguistic patterns and contextual dependencies, effectively learning to highlight the textual signals most associated with fabrication. This kind of transparency matters because a detector that simply outputs a verdict without explanation is difficult for fact-checkers, platforms, or regulators to trust or audit. By exposing its reasoning traces, the framework becomes a tool that human reviewers can interrogate rather than a black box they must take on faith.</p>
<p>The implications extend well beyond India. Misinformation researchers have long noted that detection systems trained on one linguistic or cultural context transfer poorly to another, and the new study offers a template for building locally adapted detectors: assemble a large, domain-relevant, carefully verified dataset, fine-tune a transformer architecture on it, and validate against a broad range of baselines including frontier LLMs. The authors suggest the framework could support real-time misinformation monitoring and fact-checking applications in Indian digital media environments, where the speed and scale of viral falsehoods routinely outstrip human capacity. The study received no external funding and was conducted independently by the research team.</p>
<p>Challenges remain. Fake news evolves constantly as bad actors adapt their tactics, and any deployed detector will need continual retraining to keep pace. The dataset, while large, reflects the sources available at the time of construction, and the authors note that the underlying data is available from the corresponding author upon reasonable request. Still, the work marks a meaningful step toward practical, scalable misinformation defense: evidence that thoughtfully customized transformer models, paired with rigorous data curation and human oversight, can deliver both the accuracy and the efficiency needed to fight fake news where it spreads fastest.</p>
<p><strong>Subject of Research:</strong> Fine-tuned transformer language models for detecting fake news in Indian digital media</p>
<p><strong>Article Title:</strong> A Deployed Custom Fine-tuned Transformer Language Model Framework for Indian Fake News Detection</p>
<p><strong>Article References:</strong> Prajapati, M., Baliarsingh, S. K., Sahoo, S. S., Das, S., Revankar, P. K., &amp; Pinto, J. P. (2026). A Deployed Custom Fine-tuned Transformer Language Model Framework for Indian Fake News Detection. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 469. <a href="https://doi.org/10.1007/s13042-026-03289-w" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03289-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03289-w" rel="noopener noreferrer">10.1007/s13042-026-03289-w</a></p>
<p><strong>Keywords:</strong> fake news detection, transformer models, BERT, large language models, GPT-4, natural language processing, misinformation, text classification, fine-tuning, India, machine learning, attention mechanisms</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">227131</post-id>	</item>
	</channel>
</rss>
