Saturday, September 12, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

New Fuzzy Machine Learning Model Tackles Imbalanced Data With Striking Accuracy

September 12, 2026
in Technology and Engineering
Teresa Odom
By Teresa Odom Scienmag Editorial Profile - Machine Learning
Reading Time: 5 mins read
0
New Fuzzy Machine Learning Model Tackles Imbalanced Data With Striking Accuracy

New Fuzzy Machine Learning Model Tackles Imbalanced Data With Striking Accuracy

New Fuzzy Machine Learning Model Tackles Imbalanced Data With Striking Accuracy

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Imbalanced data is one of the most stubborn problems in modern machine learning. Whenever a dataset contains far more examples of one class than another, standard classifiers tend to favor the majority class, quietly ignoring the rare but often critical cases. In medical diagnosis, fraud detection and fault monitoring, those rare cases are frequently the ones that matter most. A new study published in Knowledge and Information Systems by researchers at the Indian Institute of Technology Roorkee proposes a fresh attack on this problem: a class probability and polycentric fuzzy least squares twin support vector machine, abbreviated CPLTSVM, that consistently outperforms existing state-of-the-art methods across a broad suite of benchmark tests.

The work, led by Nikita Grewal and Yash Arora together with S. K. Gupta and Sanjeev Kumar, builds on the twin support vector machine, a widely recognized classification model first introduced as a faster, leaner cousin of the classical support vector machine. Rather than searching for a single separating hyperplane, the twin version constructs two nonparallel hyperplanes, one for each class, and assigns each point to whichever hyperplane it sits closer to. This structural difference typically makes training faster than the classical formulation, since each optimization problem involves the samples of only one class at a time. But the twin model has well-documented weaknesses: it is sensitive to noise and outliers, and when class sizes are unequal it can produce systematically biased decision boundaries that favor the dominant class.

The Roorkee team’s central insight is that not all samples deserve equal trust. In real datasets, some points sit comfortably in the heart of their class while others linger near the boundary or belong to scattered secondary clusters, and a few are outright errors. To capture this heterogeneity, CPLTSVM assigns each training sample a fuzzy membership weight derived from its proximity to the nearest class centers. Instead of assuming every class forms a single compact blob, the model adopts a polycentric view: each class may have several centers, reflecting the multi-cluster structure that naturally arises in high-dimensional data. Samples close to any center of their own class receive high membership values, meaning they exert strong influence on the learned hyperplane, while distant or structurally isolated points receive lower weights and therefore contribute less to the final decision rule. This mechanism embeds global structural information about the dataset directly into the training objective, something conventional support vector formulations ignore entirely.

The fuzzy weighting scheme draws on a long intellectual lineage. Fuzzy sets, introduced by Lotfi Zadeh in 1965, allow data points to belong to categories to intermediate degrees rather than in a binary fashion, and researchers have spent decades folding such gradations into kernel-based classifiers. Earlier fuzzy twin support vector machines have used information entropy, affinity measures and intuitionistic fuzzy memberships to dampen the effect of noisy points. What distinguishes the new approach is the combination of two complementary signals. The polycentric membership captures how representative a point is of its class structure, while a separate class probability estimate captures how confidently a point belongs to its assigned class at all.

That probability estimate comes from the k-nearest-neighbor technique, one of the oldest and most intuitive tools in pattern recognition. For each training sample, the model examines its closest neighbors and computes the fraction that share its label, yielding a local estimate of class membership probability. Points surrounded by neighbors of the same class earn high probability scores; points in mixed neighborhoods earn low ones. When these probabilities are folded into the fuzzy membership function, noisy samples and borderline outliers are automatically down-weighted, reducing their power to drag the decision boundary away from the true underlying structure. The result is a classifier that is robust by construction rather than by post-hoc correction.

Computational efficiency was a second design priority. Standard twin support vector machines solve two quadratic programming problems, which can be expensive on large datasets. By integrating the fuzzy membership weights into the least squares twin support vector machine formulation, which replaces the inequality constraints of the original with equality constraints and a squared loss, the authors convert training into the solution of systems of linear equations. This dramatically lowers the computational cost while preserving the benefits of the fuzzy weighting scheme, making the approach practical for the kinds of large, messy datasets where imbalance problems are most acute.

To evaluate the framework, the team ran experiments on twenty benchmark imbalanced datasets, testing both linear and nonlinear kernel versions of CPLTSVM against several state-of-the-art competitors. Performance was measured with three metrics suited to imbalanced settings: the area under the receiver operating characteristic curve, known as AUC, the geometric mean of class sensitivities, known as G-mean, and the F-measure, which balances precision and recall. The numbers are striking. The model achieved an average AUC of 89.71 percent with the linear kernel and 94.87 percent with the nonlinear kernel, and it consistently surpassed the existing techniques across the benchmark suite. The authors supplemented raw scores with comprehensive statistical analyses to confirm that the observed gains were significant rather than artifacts of dataset selection.

The researchers also demonstrated the model’s practical value on a task with real clinical stakes: breast cancer diagnosis from histopathological images using the BreakHis dataset. On this medical imaging benchmark, CPLTSVM delivered superior performance compared with baseline methods, suggesting the technique could support real-world classification pipelines where false negatives carry severe consequences. Because histopathological datasets are naturally imbalanced and noisy, with subtle visual differences between benign and malignant tissue, this application exercises exactly the weaknesses the new membership scheme was designed to address. The authors have made the source code publicly available on GitHub, allowing other researchers to reproduce the results and extend the framework.

The significance of the work extends beyond a single benchmark victory. Machine learning systems are increasingly deployed in settings where rare events dominate the consequences: detecting defective components in manufacturing lines, flagging fraudulent transactions, spotting early signals of disease, and identifying anomalies in wind turbine gearboxes, a domain where twin support vector machines have previously found application. Every one of these settings shares the same structural pathology that the Roorkee team targeted, and a classifier that resists both noise and imbalance without requiring delicate resampling or costly retraining could simplify deployment considerably. Compared with popular alternatives such as the SMOTE synthetic oversampling technique or cost-sensitive learning frameworks, which modify the data or the loss rather than the classifier’s view of structure, the fuzzy membership approach works within the learning algorithm itself, adjusting how much each sample matters based on where it sits in the geometry of the data.

There are, as with any method, avenues for future work. The polycentric membership depends on identifying class centers, and the k-nearest-neighbor probability estimation introduces a hyperparameter in the choice of k, questions of sensitivity that practitioners will want to probe. The evaluation here focuses on binary classification, the native territory of twin support vector machines, and extending the framework to multicategory problems, as earlier fuzzy proximal approaches have done, would broaden its reach. Still, the combination of strong empirical results, computational efficiency and demonstrated medical utility marks CPLTSVM as a notable advance in the ongoing effort to make classifiers trustworthy when the world refuses to hand us balanced data. As algorithms increasingly inform decisions about health, security and safety, techniques that give quiet, underrepresented samples their due weight are not merely academic refinements; they are prerequisites for reliable artificial intelligence.

Subject of Research: A fuzzy least squares twin support vector machine using class probability and polycentric memberships to improve classification of imbalanced datasets

Article Title: Class probability and polycentric fuzzy least squares twin support vector machine for imbalanced data

Article References: Grewal, N., Arora, Y., Gupta, S. K., & Kumar, S. (2026). Class probability and polycentric fuzzy least squares twin support vector machine for imbalanced data. Knowledge and Information Systems, 68(1), Article 259. https://doi.org/10.1007/s10115-026-02862-7

Image Credits: AI Generated

DOI: 10.1007/s10115-026-02862-7

Keywords: imbalanced datasets, support vector machine, twin support vector machine, fuzzy membership, polycentric membership, class probability, least squares, machine learning, medical diagnosis, breast cancer, BreakHis dataset, classification

Cite Scienmag News

Teresa Odom. (September 12, 2026). New Fuzzy Machine Learning Model Tackles Imbalanced Data With Striking Accuracy. Scienmag. https://scienmag.com/new-fuzzy-machine-learning-model-tackles-imbalanced-data-with-striking-accuracy/

Teresa Odom. "New Fuzzy Machine Learning Model Tackles Imbalanced Data With Striking Accuracy." Scienmag, 12 September 2026, https://scienmag.com/new-fuzzy-machine-learning-model-tackles-imbalanced-data-with-striking-accuracy/. Accessed 12 September 2026.

Teresa Odom. "New Fuzzy Machine Learning Model Tackles Imbalanced Data With Striking Accuracy." Scienmag. September 12, 2026. https://scienmag.com/new-fuzzy-machine-learning-model-tackles-imbalanced-data-with-striking-accuracy/

Tags: addressing minority class in imbalanced datasetsbenchmark testing of classification algorithmsBreakHis datasetbreast cancerclass probabilityclass probability models in machine learningclassificationCPLTSVM for handling rare event detectionfault monitoring with imbalanced datafuzzy membershipfuzzy twin support vector machineImbalanced data classification in machine learningimbalanced datasetsimprovements in twin support vector machine modelsleast squaresMachine learningmachine learning techniques for critical case detectionmedical diagnosismedical diagnosis and fraud detection applicationspolycentric fuzzy modelspolycentric membershipsupport vector machinesupport vector machine advancementstwin support vector machine
Share26Tweet16
Previous Post

Stress Rewires Money Choices Even Among High Earners, Study Finds

Next Post

Black Holes, Gravitational Waves and the Fate of Spacetime Take Center Stage at Vatican Conference

Related Posts

AI Learns to Read Fuzzy Bone Scans and Write Radiology Reports
Technology and Engineering

AI Learns to Read Fuzzy Bone Scans and Write Radiology Reports

September 12, 2026
AI-Powered Multimodal Sensors Learn to Untangle the World’s Overlapping Signals
Technology and Engineering

AI-Powered Multimodal Sensors Learn to Untangle the World’s Overlapping Signals

September 12, 2026
Flat Gradients Make Fake Users Deadlier: Smarter Attacks Expose Recommender Vulnerabilities
Technology and Engineering

Flat Gradients Make Fake Users Deadlier: Smarter Attacks Expose Recommender Vulnerabilities

September 12, 2026
Potassium Nickel Hydride Emerges as a Room-Temperature Hydrogen Storage Contender
Technology and Engineering

Potassium Nickel Hydride Emerges as a Room-Temperature Hydrogen Storage Contender

September 12, 2026
Two Decades of Wikipedia Research Reveal a Fractured Field Shaped by Big Data and AI
Technology and Engineering

Two Decades of Wikipedia Research Reveal a Fractured Field Shaped by Big Data and AI

September 12, 2026
Quantum Geometry and Teleportation Bound Together in a Two-Spin System
Technology and Engineering

Quantum Geometry and Teleportation Bound Together in a Two-Spin System

September 12, 2026
Next Post
Black Holes, Gravitational Waves and the Fate of Spacetime Take Center Stage at Vatican Conference

Black Holes, Gravitational Waves and the Fate of Spacetime Take Center Stage at Vatican Conference

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Black Holes, Gravitational Waves and the Fate of Spacetime Take Center Stage at Vatican Conference
  • New Fuzzy Machine Learning Model Tackles Imbalanced Data With Striking Accuracy
  • Stress Rewires Money Choices Even Among High Earners, Study Finds
  • Seven-Year Study Reveals How Loneliness and Anxiety Quietly Erode Happiness in Working Adults

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading