Fake news has become one of the most corrosive byproducts of the connected world. As internet access has expanded, social media platforms have given false rumors the ability to travel faster than fact-checkers can respond, deceiving readers and, in some cases, causing real social harm. Traditional fact-checking, while valuable, is slow and labor-intensive; by the time a human verifier has completed an assessment, a fabricated story may already have spread across networks. A new study published in Multimedia Tools and Applications by Mayank Kumar Jain, Dinesh Gopalani, Yogesh Kumar Meena of Malaviya National Institute of Technology Jaipur, and Nishant Jain of Madhav Institute of Technology and Science Gwalior, addresses this gap with an automatic, content-based fake news detection framework built around a structured divide-and-conquer strategy.
The central premise of the research is that a detector should not depend on external social context or post-propagation metadata, such as how many times an article has been shared or who published it. Those signals are often unavailable in the critical early minutes of a story’s life, precisely when detection matters most. Instead, the proposed framework relies entirely on the text of the news article itself, combining two complementary families of features: linguistic features that capture the stylistic fingerprints of deceptive writing, and word vector representations that encode the semantic meaning of the words used. This content-only approach makes the system applicable to any news source, regardless of platform or publisher data availability.
The divide-and-conquer design systematically separates the feature engineering problem into manageable components. The researchers extracted more than eighty linguistic features from the text fields of news articles, drawing on techniques associated with tools such as Linguistic Inquiry and Word Count. These features quantify measurable properties of writing style, including word counts, sentence complexity, pronoun usage, emotional tone, and other lexical and syntactic cues that previous research has linked to deceptive or sensationalist content. In parallel, the framework generates word vector features using two well-established embedding techniques: the continuous bag-of-words model and the skip-gram model, both of which learn dense numerical representations of words based on the contexts in which they appear.
Once computed, the linguistic feature vector, which lives in an eighty-dimensional space, is concatenated with either a CBOW embedding vector or a skip-gram embedding vector, each one hundred dimensions long. The result is a hybrid feature vector of one hundred eighty dimensions that fuses stylistic and semantic information into a single representation suitable for machine learning. This structured separation, dividing the representation task into linguistic and vector components before conquering the classification task, allows the framework to remain modular: either embedding variant can be swapped in without redesigning the pipeline, and the contribution of each feature family can be evaluated independently.
To test whether the approach generalizes rather than merely memorizes one dataset’s quirks, the team conducted a rigorous multi-dataset evaluation across three widely used benchmarks: the Kaggle fake news dataset, a combined McIntire and PolitiFact corpus, and the Reuter dataset. Ten different machine learning classifiers were trained and assessed, including Support Vector Machine, Random Forest, Gradient Boosting, Extra Trees, Logistic Regression, Naive Bayes, and K-Nearest Neighbors, alongside deep learning architectures such as LSTM and CNN-based models. Every configuration was evaluated with five-fold cross-validation, a procedure that splits the data into five partitions and rotates which partition serves as the test set, ensuring that reported performance reflects consistency across corpus sizes and domains rather than a single lucky split.
The results were striking. The Support Vector Machine achieved the highest overall performance, attaining an F1-score of 98.76 percent on the Reuter dataset and 98.34 percent on the Kaggle dataset, while reaching 92.11 percent on the McIntire and PolitiFact corpus. The F1-score, which harmonically combines precision and recall, is a demanding metric: it penalizes both false alarms, where genuine news is wrongly flagged, and misses, where fabricated stories slip through. On the McIntire and PolitiFact dataset, Gradient Boosting performed best, achieving an F1-score of 92.21 percent. The fact that the top-performing classifier changes across datasets underscores why the authors emphasize multi-dataset evaluation as an essential discipline for the field.
Beyond raw accuracy, the study tackled a practical engineering concern: dimensionality. High-dimensional feature vectors increase computational cost and can encourage overfitting, where a model learns dataset-specific noise instead of generalizable patterns. The researchers applied feature selection techniques, guided in part by correlation analysis using the Pearson correlation coefficient to identify redundant or uninformative dimensions. The outcome was a reduction in dimensionality of between 25 and 53 percent without any degradation in performance, meaning the optimized feature vectors retained their discriminative power while becoming substantially leaner. For deployment scenarios where detectors must process high volumes of incoming articles in real time, this efficiency gain is far from trivial.
The work builds on a growing body of research into content-based fake news detection. Earlier efforts have explored hybrid CNN-RNN architectures, attention-enabled neural models such as AENET, transformer-based approaches for COVID-19 misinformation, and frameworks like WELFake that combine word embeddings with linguistic features. The authors’ own prior contributions, including the Confake content-based feature system and a hybrid CNN-BiLSTM model with feature selection, laid groundwork for the present framework. What distinguishes the new study is its systematic integration strategy and its insistence on benchmarking across multiple corpora with a large panel of classifiers, addressing a common weakness in the literature where methods are validated on a single dataset and reported scores cannot be compared fairly.
The stakes of this research are illustrated by real-world incidents the authors cite, including a viral hoax claiming India would be nuked and had dropped 750 bombs in Pakistan, which prompted the government to block twenty-two YouTube channels, and a fabricated message about free laptops being distributed under a government scheme that circulated widely before being debunked by fact-checkers. In both cases, the damage was done before human verification could catch up. A content-based detector that can flag suspicious articles within seconds of publication, without waiting for shares, comments, or publisher metadata to accumulate, could shorten that window dramatically and blunt the initial spread of a false story.
The researchers have made their data and code publicly available, supporting reproducibility and allowing other teams to build on the framework. Limitations remain, as they do in all content-based approaches: sophisticated disinformation campaigns can adapt their writing style to evade stylistic detection, and the benchmark datasets, while diverse, may not capture the full range of languages, topics, and deception techniques encountered in the wild. Nevertheless, the study demonstrates that a carefully structured divide-and-conquer pipeline, uniting over eighty linguistic cues with word vector embeddings and disciplined cross-dataset validation, can push detection performance above 98 percent on standard benchmarks. As misinformation tactics evolve, frameworks of this kind, grounded in measurable textual evidence rather than platform-specific signals, offer a scalable and adaptable line of defense for news consumers and platforms alike.
Subject of Research: Content-based automatic fake news detection using linguistic features and word vector embeddings with machine learning classifiers
Article Title: Implementation of a structured divide and conquer approach-based fake news detector for news sources
Article References: Jain, M. K., Gopalani, D., Meena, Y. K., & Jain, N. (2026). Implementation of a structured divide and conquer approach-based fake news detector for news sources. Multimedia Tools and Applications, 85(10), Article 787. https://doi.org/10.1007/s11042-026-21945-9
Image Credits: AI Generated
DOI: 10.1007/s11042-026-21945-9
Keywords: fake news detection, machine learning, linguistic features, word embeddings, support vector machine, natural language processing, social media, cross-validation, feature selection, misinformation, divide and conquer, text classification
Cite Scienmag News
Denise Maddox. (September 30, 2026). Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets. Scienmag. https://scienmag.com/divide-and-conquer-new-fake-news-detector-hits-98-accuracy-across-datasets/
Denise Maddox. "Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets." Scienmag, 30 September 2026, https://scienmag.com/divide-and-conquer-new-fake-news-detector-hits-98-accuracy-across-datasets/. Accessed 30 September 2026.
Denise Maddox. "Divide and Conquer: New Fake News Detector Hits 98% Accuracy Across Datasets." Scienmag. September 30, 2026. https://scienmag.com/divide-and-conquer-new-fake-news-detector-hits-98-accuracy-across-datasets/

