Wednesday, September 30, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets

September 30, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets

Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets

Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Fake news has become one of the most corrosive byproducts of the connected world. As internet access has expanded, social media platforms have given false rumors the ability to travel faster than fact-checkers can respond, deceiving readers and, in some cases, causing real social harm. Traditional fact-checking, while valuable, is slow and labor-intensive; by the time a human verifier has completed an assessment, a fabricated story may already have spread across networks. A new study published in Multimedia Tools and Applications by Mayank Kumar Jain, Dinesh Gopalani, Yogesh Kumar Meena of Malaviya National Institute of Technology Jaipur, and Nishant Jain of Madhav Institute of Technology and Science Gwalior, addresses this gap with an automatic, content-based fake news detection framework built around a structured divide-and-conquer strategy.

The central premise of the research is that a detector should not depend on external social context or post-propagation metadata, such as how many times an article has been shared or who published it. Those signals are often unavailable in the critical early minutes of a story’s life, precisely when detection matters most. Instead, the proposed framework relies entirely on the text of the news article itself, combining two complementary families of features: linguistic features that capture the stylistic fingerprints of deceptive writing, and word vector representations that encode the semantic meaning of the words used. This content-only approach makes the system applicable to any news source, regardless of platform or publisher data availability.

The divide-and-conquer design systematically separates the feature engineering problem into manageable components. The researchers extracted more than eighty linguistic features from the text fields of news articles, drawing on techniques associated with tools such as Linguistic Inquiry and Word Count. These features quantify measurable properties of writing style, including word counts, sentence complexity, pronoun usage, emotional tone, and other lexical and syntactic cues that previous research has linked to deceptive or sensationalist content. In parallel, the framework generates word vector features using two well-established embedding techniques: the continuous bag-of-words model and the skip-gram model, both of which learn dense numerical representations of words based on the contexts in which they appear.

Once computed, the linguistic feature vector, which lives in an eighty-dimensional space, is concatenated with either a CBOW embedding vector or a skip-gram embedding vector, each one hundred dimensions long. The result is a hybrid feature vector of one hundred eighty dimensions that fuses stylistic and semantic information into a single representation suitable for machine learning. This structured separation, dividing the representation task into linguistic and vector components before conquering the classification task, allows the framework to remain modular: either embedding variant can be swapped in without redesigning the pipeline, and the contribution of each feature family can be evaluated independently.

To test whether the approach generalizes rather than merely memorizes one dataset’s quirks, the team conducted a rigorous multi-dataset evaluation across three widely used benchmarks: the Kaggle fake news dataset, a combined McIntire and PolitiFact corpus, and the Reuter dataset. Ten different machine learning classifiers were trained and assessed, including Support Vector Machine, Random Forest, Gradient Boosting, Extra Trees, Logistic Regression, Naive Bayes, and K-Nearest Neighbors, alongside deep learning architectures such as LSTM and CNN-based models. Every configuration was evaluated with five-fold cross-validation, a procedure that splits the data into five partitions and rotates which partition serves as the test set, ensuring that reported performance reflects consistency across corpus sizes and domains rather than a single lucky split.

The results were striking. The Support Vector Machine achieved the highest overall performance, attaining an F1-score of 98.76 percent on the Reuter dataset and 98.34 percent on the Kaggle dataset, while reaching 92.11 percent on the McIntire and PolitiFact corpus. The F1-score, which harmonically combines precision and recall, is a demanding metric: it penalizes both false alarms, where genuine news is wrongly flagged, and misses, where fabricated stories slip through. On the McIntire and PolitiFact dataset, Gradient Boosting performed best, achieving an F1-score of 92.21 percent. The fact that the top-performing classifier changes across datasets underscores why the authors emphasize multi-dataset evaluation as an essential discipline for the field.

Beyond raw accuracy, the study tackled a practical engineering concern: dimensionality. High-dimensional feature vectors increase computational cost and can encourage overfitting, where a model learns dataset-specific noise instead of generalizable patterns. The researchers applied feature selection techniques, guided in part by correlation analysis using the Pearson correlation coefficient to identify redundant or uninformative dimensions. The outcome was a reduction in dimensionality of between 25 and 53 percent without any degradation in performance, meaning the optimized feature vectors retained their discriminative power while becoming substantially leaner. For deployment scenarios where detectors must process high volumes of incoming articles in real time, this efficiency gain is far from trivial.

The work builds on a growing body of research into content-based fake news detection. Earlier efforts have explored hybrid CNN-RNN architectures, attention-enabled neural models such as AENET, transformer-based approaches for COVID-19 misinformation, and frameworks like WELFake that combine word embeddings with linguistic features. The authors’ own prior contributions, including the Confake content-based feature system and a hybrid CNN-BiLSTM model with feature selection, laid groundwork for the present framework. What distinguishes the new study is its systematic integration strategy and its insistence on benchmarking across multiple corpora with a large panel of classifiers, addressing a common weakness in the literature where methods are validated on a single dataset and reported scores cannot be compared fairly.

The stakes of this research are illustrated by real-world incidents the authors cite, including a viral hoax claiming India would be nuked and had dropped 750 bombs in Pakistan, which prompted the government to block twenty-two YouTube channels, and a fabricated message about free laptops being distributed under a government scheme that circulated widely before being debunked by fact-checkers. In both cases, the damage was done before human verification could catch up. A content-based detector that can flag suspicious articles within seconds of publication, without waiting for shares, comments, or publisher metadata to accumulate, could shorten that window dramatically and blunt the initial spread of a false story.

The researchers have made their data and code publicly available, supporting reproducibility and allowing other teams to build on the framework. Limitations remain, as they do in all content-based approaches: sophisticated disinformation campaigns can adapt their writing style to evade stylistic detection, and the benchmark datasets, while diverse, may not capture the full range of languages, topics, and deception techniques encountered in the wild. Nevertheless, the study demonstrates that a carefully structured divide-and-conquer pipeline, uniting over eighty linguistic cues with word vector embeddings and disciplined cross-dataset validation, can push detection performance above 98 percent on standard benchmarks. As misinformation tactics evolve, frameworks of this kind, grounded in measurable textual evidence rather than platform-specific signals, offer a scalable and adaptable line of defense for news consumers and platforms alike.

Subject of Research: Content-based automatic fake news detection using linguistic features and word vector embeddings with machine learning classifiers

Article Title: Implementation of a structured divide and conquer approach-based fake news detector for news sources

Article References: Jain, M. K., Gopalani, D., Meena, Y. K., & Jain, N. (2026). Implementation of a structured divide and conquer approach-based fake news detector for news sources. Multimedia Tools and Applications, 85(10), Article 787. https://doi.org/10.1007/s11042-026-21945-9

Image Credits: AI Generated

DOI: 10.1007/s11042-026-21945-9

Keywords: fake news detection, machine learning, linguistic features, word embeddings, support vector machine, natural language processing, social media, cross-validation, feature selection, misinformation, divide and conquer, text classification

Cite Scienmag News

Denise Maddox. (September 30, 2026). Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets. Scienmag. https://scienmag.com/divide-and-conquer-new-fake-news-detector-hits-98-accuracy-across-datasets/

Denise Maddox. "Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets." Scienmag, 30 September 2026, https://scienmag.com/divide-and-conquer-new-fake-news-detector-hits-98-accuracy-across-datasets/. Accessed 30 September 2026.

Denise Maddox. "Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets." Scienmag. September 30, 2026. https://scienmag.com/divide-and-conquer-new-fake-news-detector-hits-98-accuracy-across-datasets/

Tags: automated fact-checkingcontent-based fake news classifierscross-validationdivide and conquerdivide-and-conquer strategy in misinformation detectionearly-stage fake news identificationfake news detectionfeature selectionhigh-accuracy fake news detection datasetslinguistic feature extraction in journalism verificationlinguistic featuresMachine learningmachine learning models for media authenticitymisinformationmultimedia content analysis for fake newsnatural language processingnatural language processing for fake newsreal-time fake news detection frameworkssocial mediasocial media misinformation mitigationsupport vector machinetext classificationword embeddings
Share26Tweet16
Previous Post

Energy Ratings Already Shape Your Mortgage, and Climate Risk Is Next

Next Post

Where Fat Settles May Shape Polycystic Ovary Syndrome Risk, Global Study Finds

Related Posts

Aluminum Waste Red Mud Could Replace Clay in Low-Carbon Cement
Technology and Engineering

Aluminum Waste Red Mud Could Replace Clay in Low-Carbon Cement

September 30, 2026
AI Cracks the Atomic Secrets of Cement’s Most Elusive Ingredient
Technology and Engineering

AI Cracks the Atomic Secrets of Cement’s Most Elusive Ingredient

September 30, 2026
AI Health Coaches Get Personal: Massive Review Maps How Chatbots Could Tailor Care to You
Technology and Engineering

AI Health Coaches Get Personal: Massive Review Maps How Chatbots Could Tailor Care to You

September 30, 2026
AI Reads Brainwaves to Diagnose Depression, Review of 69 Studies Finds
Technology and Engineering

AI Reads Brainwaves to Diagnose Depression, Review of 69 Studies Finds

September 30, 2026
Blockchain Tokens Could Pay Citizens to Green Their Cities, Simulation Finds
Technology and Engineering

Blockchain Tokens Could Pay Citizens to Green Their Cities, Simulation Finds

September 30, 2026
AI Chatbots Redesign a Plant Molecule to Out-Bind a Gout Drug
Technology and Engineering

AI Chatbots Redesign a Plant Molecule to Out-Bind a Gout Drug

September 30, 2026
Next Post
Where Fat Settles May Shape Polycystic Ovary Syndrome Risk, Global Study Finds

Where Fat Settles May Shape Polycystic Ovary Syndrome Risk, Global Study Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Aluminum Waste Red Mud Could Replace Clay in Low-Carbon Cement
  • How Cells Build Their Protein Shredder: New Rules of Proteasome Biogenesis
  • How Malaria Parasites Punch Through Cells to Spark Protective Immunity
  • Most Older Veterans Land in Lower-Quality Nursing Homes Than Those Nearby

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading