Friday, October 9, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories

October 9, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 4 mins read
0
Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories

Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Fake news does not respect borders, and it certainly does not respect languages. While social media platforms have invested heavily in automated misinformation detection for English content, the vast majority of the world’s languages remain poorly served by these systems. A new study published in Discover Artificial Intelligence by Sushama Nandgaonkar and Sunil Mane of COEP Technological University in Pune tackles this gap head-on, presenting a Hierarchical Attention Network that detects false news across English, Hindi, and Marathi with remarkable accuracy, achieving 98.9 percent accuracy and a 98.8 percent F1-score on a combined multilingual test set.

The stakes are particularly high in India, which has the second-largest population of internet users worldwide. According to figures cited in the study, India had more than 375 million social media users in 2019, rising to 518 million in 2020, with projections suggesting the number could reach 1.5 billion by 2040. Affordable mobile data has driven this explosive growth, and platforms such as Facebook, Twitter, and WhatsApp now carry content in dozens of languages. Misinformation that circulates in regional languages can provoke local sentiments and influence health decisions, as demonstrated during the COVID-19 pandemic, when false claims about remedies and vaccines spread faster than official corrections could keep up.

The researchers distinguish carefully between related concepts that are often conflated. Disinformation refers to deliberately created and spread falsehoods, rumors are unverified pieces of information circulated with intent to mislead, and misinformation is false content shared without necessarily intending to deceive. The motivations behind such content range from spreading religious hatred to gaining political advantage or disseminating false health information. Because fact-checking websites themselves can carry biases, and because human verification cannot scale to billions of posts, automated detection systems have become an essential line of defense, and the study argues that these systems must work across the languages people actually use.

At the heart of the new work is a two-level attention architecture that mirrors how humans read documents. The model first processes each sentence word by word using a Bidirectional Long Short-Term Memory network, which reads text in both forward and backward directions to capture context from either side of every word. A custom attention layer then assigns dynamic weights to individual words, learning during training which terms matter most for distinguishing fake from genuine content. The weighted word representations are combined into sentence vectors, and a second Bi-LSTM layer with its own attention mechanism models relationships between sentences, producing a document-level representation that feeds into a final sigmoid classifier.

This hierarchical design is not merely an engineering convenience. Fake news articles often contain misleading lexical patterns, emotionally polarized expressions, and contextual inconsistencies distributed across multiple sentences rather than concentrated in any single phrase. By attending to both words and sentences, the model can capture local semantic cues and document-level structure simultaneously. The attention weights also improve interpretability, allowing researchers to see which parts of an article drove a classification decision, a significant advantage over opaque black-box approaches.

A crucial contribution of the study is linguistic rather than architectural. Because no publicly available fake news dataset existed for Marathi, the team built one from scratch, collecting 1,957 news articles from Marathi outlets including Loksatta, Lokmat, Maharashtra Times, and Zee News, of which 707 were labeled fake and 1,250 true. Labels were assigned based on source credibility, fact-checking reports, and consistency across multiple platforms, with manual review of every article. For Hindi, the researchers combined two existing datasets from Kaggle and GitHub, while English experiments used the widely adopted ISOT dataset. To address severe class imbalance in the low-resource languages, the team applied translation-based augmentation using the IndicTrans neural translation framework, generating additional training samples while keeping the test set untouched to prevent data leakage. The final augmented corpus contained 45,386 English, 10,481 Hindi, and 5,000 Marathi samples.

The choice of word embeddings proved important. FastText, developed by Facebook’s AI Research lab, represents each word as a bag of character n-grams, allowing it to capture morphological information and generate vectors for out-of-vocabulary words from their sub-word components. This property is especially valuable for morphologically rich languages like Hindi and Marathi, where word forms vary extensively. The researchers compared FastText against GloVe and random initialization under identical hyperparameters, finding that pretrained embeddings substantially improved performance, with FastText also converging fastest at 383.06 seconds of training time compared with 501.22 seconds for GloVe.

The experimental results were striking. The HAN model achieved 99.42 percent accuracy on English news, 99.04 percent on Hindi, and 80.90 percent on Marathi, outperforming traditional machine learning baselines including Logistic Regression, Linear Support Vector Machines, and Random Forest, as well as deep learning models such as CNN and standalone Bi-LSTM. Against multilingual transformers, the picture was more nuanced: mBERT reached 98.72 percent overall accuracy and DeBERTa-v3-base 98.66 percent, slightly below the proposed model’s 98.90 percent, but DeBERTa’s performance collapsed to 74.42 percent on Marathi, exposing the sensitivity of large pretrained transformers to data imbalance. Notably, the instruction-tuned LLaMA 3.1 8B model, evaluated through prompting rather than fine-tuning, managed only 67.98 percent accuracy in zero-shot settings and 84.76 percent with six examples in context, with 1,638 failed predictions out of 10,965 test instances.

Ablation experiments confirmed that both attention levels contribute meaningfully. A Bi-LSTM-only model achieved an F1-score of 97.43 percent, rising to 97.82 percent with word-level attention alone and 98.17 percent with sentence-level attention alone, while the complete hierarchical architecture reached 98.72 percent. Sentence-level attention proved more influential than word-level attention in isolation, suggesting that contextual dependencies across sentences play a decisive role in identifying deceptive content. Paired t-tests across multiple random seeds showed all improvements over baselines were statistically significant, with p-values below 0.01.

The study is candid about its limitations. Marathi’s error rate of 19.10 percent, compared with just 0.58 percent for English and 0.96 percent for Hindi, reflects the challenge of low-resource detection, and qualitative analysis showed that the worst failures involved fake articles written in a style so similar to legitimate news that the model assigned incorrect predictions with high confidence. The authors note that generalization to other Indic languages remains unexplored and that multimodal signals such as images and metadata are not yet incorporated. Future work will expand the dataset to additional Indian regional languages and integrate textual and visual information using transformer-based Indic language models. For now, the research demonstrates that carefully designed attention mechanisms, paired with sub-word embeddings and thoughtful data augmentation, can bring state-of-the-art misinformation detection to languages that have long been left behind.

Subject of Research: Multilingual fake news detection using hierarchical attention networks for English, Hindi, and Marathi news articles

Article Title: Hierarchical attention network for multilingual fake news detection

Article References: Nandgaonkar, S., & Mane, S. (2026). Hierarchical attention network for multilingual fake news detection. Discover Artificial Intelligence, 6(1), Article 1417. https://doi.org/10.1007/s44163-026-02425-3

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02425-3

Keywords: fake news detection, hierarchical attention network, multilingual NLP, Bi-LSTM, FastText embeddings, Marathi language, Hindi language, misinformation, deep learning, low-resource languages, data augmentation, natural language processing

Cite Scienmag News

Denise Maddox. (October 9, 2026). Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories. Scienmag. https://scienmag.com/attention-based-ai-reads-news-word-by-word-to-catch-multilingual-fake-stories/

Denise Maddox. "Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories." Scienmag, 9 October 2026, https://scienmag.com/attention-based-ai-reads-news-word-by-word-to-catch-multilingual-fake-stories/. Accessed 9 October 2026.

Denise Maddox. "Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories." Scienmag. October 9, 2026. https://scienmag.com/attention-based-ai-reads-news-word-by-word-to-catch-multilingual-fake-stories/

Tags: AI accuracy in fake news classificationAI-based multilingual news verificationautomated fake news detection in Indian languagesBi-LSTMchallenges of misinformation in diverse languagesCOVID-19 misinformation in regional languagescross-lingual misinformation identificationdata augmentationdeep learningfake news detectionfake news impact on public healthFastText embeddingshierarchical attention networkHierarchical Attention Network for misinformationHindi languagelow-resource languagesMarathi languagemisinformationmultilingual AI news analysisMultilingual fake news detectionmultilingual NLPnatural language processingscalable solutions for multilingual misinformationsocial media misinformation in Hindi and Marathi
Share26Tweet16
Previous Post

Toxic Traces in the Spud Belt: Arsenic and Lead Map the Soils of Italy’s Sila Massif

Next Post

Two Faces of One Bacterium: Study Maps Who Falls Ill with Fusobacterium and Why

Related Posts

AI-Boosted Digital Shadows Slash Wind Turbine Fatigue Prediction Errors
Climate

AI-Boosted Digital Shadows Slash Wind Turbine Fatigue Prediction Errors

October 9, 2026
AI Learns to Judge Art: Multimodal Model Scores Drawing Composition Like an Expert
Technology and Engineering

AI Learns to Judge Art: Multimodal Model Scores Drawing Composition Like an Expert

October 9, 2026
Shredded Tires Meet Cement-Free Concrete: Particle Size Decides Everything
Technology and Engineering

Shredded Tires Meet Cement-Free Concrete: Particle Size Decides Everything

October 9, 2026
Beach Plant Extract Yields Copper Ferrite Nanoparticles for Supercapacitor Electrodes
Technology and Engineering

Beach Plant Extract Yields Copper Ferrite Nanoparticles for Supercapacitor Electrodes

October 9, 2026
Feedback Control Reveals Islands of Order Hidden in Quantum Chaos
Technology and Engineering

Feedback Control Reveals Islands of Order Hidden in Quantum Chaos

October 9, 2026
Bimodal Cavity and Low-Noise Amplifier Double the Sensitivity of Q-Band Pulse EPR
Chemistry

Bimodal Cavity and Low-Noise Amplifier Double the Sensitivity of Q-Band Pulse EPR

October 9, 2026
Next Post
Two Faces of One Bacterium: Study Maps Who Falls Ill with Fusobacterium and Why

Two Faces of One Bacterium: Study Maps Who Falls Ill with Fusobacterium and Why

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • AI-Boosted Digital Shadows Slash Wind Turbine Fatigue Prediction Errors
  • AI Learns to Judge Art: Multimodal Model Scores Drawing Composition Like an Expert
  • Two Faces of One Bacterium: Study Maps Who Falls Ill with Fusobacterium and Why
  • Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading