Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels

October 1, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels

Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels

Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Tabular machine learning quietly runs the modern world. Models trained on rows and columns of data decide whether a credit card transaction is fraudulent, whether a patient’s health indicators suggest diabetes, and whether a network connection is an intrusion. But a new study from researchers at Soongsil University in Seoul, published in Applied Intelligence, highlights a disturbing weakness in these systems: an attacker does not need to touch a single input feature to sabotage them. By simply flipping the labels attached to training samples, an adversary can distort the decision boundaries that these high-stakes models learn, and the corruption can be nearly invisible to conventional validation checks.

The research team, led by Jinhyeok Jang and corresponding author Daeseon Choi, frames the problem with unusual clarity. Label flipping is a form of data poisoning in which the observed class of a training example is changed while the underlying features remain intact. Because the features themselves look perfectly normal, simple feature-level screening catches nothing. Worse, verifying whether a label is truly correct in domains like finance or healthcare often requires scarce expert knowledge, so corrupted labels can survive unnoticed all the way into production. As organizations increasingly rely on crowdsourced annotation, outsourced labeling, and third-party data integration to cut costs, the attack surface only grows.

What makes the new work particularly timely is its focus on how modern label-flipping attacks have evolved. Early attacks flipped labels at random, scattering noise across the dataset. Recent strategies are far more surgical. Decision-boundary attacks use surrogate models to identify low-margin samples sitting perilously close to the classifier’s decision frontier, while optimization-based methods such as ALFA and its ALFA-Tilt variant solve constrained optimization problems to select the flip set that maximally degrades the learner. The authors’ visualizations show that these targeted attacks produce localized poisoning patterns, with flipped samples clustering together in feature space rather than appearing as isolated outliers. That clustering defeats density-based and global outlier detectors, because the poisoned points masquerade as a legitimate local group.

The team’s answer is a defense built on a counterintuitive idea: learn what the data looks like without trusting any labels at all. Their method, called BYOL-Union, first trains a self-supervised encoder using BYOL, or Bootstrap Your Own Latent, a technique originally developed for images that learns representations by predicting augmented views of the same instance without any negative pairs or labels. Applied to tabular data, the encoder absorbs the intrinsic geometric structure of the feature space, untouched by whatever corruption may lurk in the observed labels. Only after this label-free representation learning is complete does the system consult the labels, and then solely to measure how much each sample disagrees with its neighbors.

The disagreement scoring is where the method earns its name. For each sample, the system finds its k nearest neighbors in the embedding space and counts how many of them carry a different observed label. A genuinely flipped sample tends to sit among neighbors whose true labels differ from its own forged one, producing a high disagreement score. Crucially, the method does not rely on a single fixed view of the neighborhood. Just as image-based self-supervised learning uses crops and flips to see one object from multiple angles, the framework generates a second view by adding Gaussian noise, with a standard deviation of 0.1, to continuous features only, leaving categorical features untouched to avoid inventing invalid categories. Each sample is then probed in both the original and the perturbed embedding spaces, and suspicious sets from the two views are merged by a union rule.

That union rule is deliberately recall-oriented. A sample flagged as suspicious in either view enters the final suspicious set, maximizing the chance of catching poisoned labels that are exposed in at least one neighborhood perspective, at the cost of some additional false positives. The authors also designed for realistic deployment, where the true poisoning ratio is unknown: instead of assuming an oracle budget, the practical version uses z-score thresholding at 1.0 with any-view exceedance, which their ablations show performs close to the oracle reference while remaining entirely oracle-free.

The evaluation is unusually broad. Six public tabular benchmarks spanning network intrusion detection, finance, and healthcare, including NSL-KDD, UNSW-NB15, Bank Marketing, Credit Card Fraud, BRFSS 2015 Diabetes Health Indicators, and the Diabetes 130-US Hospitals dataset, were poisoned at ratios of 10, 20, and 30 percent under random, decision-boundary, and optimization-based attacks. Four heterogeneous target models, spanning support vector machines, deep neural networks, FT-Transformers, and XGBoost, were then trained on sanitized data. The results show a consistent pattern: under decision-boundary flipping, the original feature-space kNN detector achieved a recall of only 0.601, while the BYOL-based and SCARF-based variants reached 0.773 and 0.792 respectively, and paired Wilcoxon tests confirmed that the SSL-based detectors significantly improved both recall and F1 over Curie, LS-SVM, and original-space kNN baselines.

The study is equally candid about limits. Under ALFA-Tilt, the strongest optimization-based attack tested in a small-scale setting, detection became markedly harder for every detector, and the robust-training defense FLORAL achieved the smallest downstream accuracy gap, though it cannot identify which specific samples are poisoned and is tied to particular model architectures. The authors stress that improved detection does not always translate directly into recovered test accuracy, because removing suspicious samples changes the training set itself and can thin out supervision for minority classes. Detection quality, they argue, should be judged on its own terms, with downstream recovery treated as a complementary, setting-dependent benefit rather than the sole criterion.

Perhaps the most striking result concerns adaptive adversaries. The team constructed defense-aware attacks in which the attacker knows the sanitization mechanism and selects flips that remain locally plausible in the detector’s own neighborhood space. Against a raw feature-space detector, such an adaptive attack cut recall from 0.668 to 0.435. Yet the SSL-based union detectors held firm: BYOL-Union maintained a recall of 0.760 even when the attacker targeted its own representation space, and SCARF-Union reached 0.800 under the corresponding adaptive attack. The layered evidence from learned representation spaces, it appears, is not fully dismantled by an adversary who optimizes against any single neighborhood view.

The broader lesson resonates beyond this one defense. As machine learning systems are deployed in domains where a single misclassified transaction or missed intrusion carries real consequences, the integrity of training labels deserves the same security attention as model architecture. The Soongsil team’s framework, which will see code released on request, offers a practical, model-agnostic preprocessing step: it flags suspicious samples before any downstream classifier is fit, works alongside tree ensembles and neural networks alike, and leaves room for future extensions toward label correction, sample reweighting, and robust retraining with uncertainty. In a field where attackers increasingly aim at the data rather than the model, defenses that learn to see the data on its own terms may prove essential.

Subject of Research: Label-flipping attack detection and data sanitization in tabular machine learning using multi-view self-supervised representations

Article Title: Multi-view self-supervised learning for label-flipping robustness in tabular data: a comparative study

Article References: Multi-view self-supervised learning for label-flipping robustness in tabular data: a comparative study. (n.d.). https://doi.org/10.1007/s10489-026-07452-2

Image Credits: AI Generated

DOI: 10.1007/s10489-026-07452-2

Keywords: label flipping, data poisoning, self-supervised learning, BYOL, tabular data, data sanitization, k-nearest neighbors, adversarial machine learning, SCARF, network intrusion detection, machine learning security, Applied Intelligence

Cite Scienmag News

Denise Maddox. (October 1, 2026). Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels. Scienmag. https://scienmag.com/self-supervised-neighborhood-probing-shields-tabular-ai-from-poisoned-labels/

Denise Maddox. "Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels." Scienmag, 1 October 2026, https://scienmag.com/self-supervised-neighborhood-probing-shields-tabular-ai-from-poisoned-labels/. Accessed 1 October 2026.

Denise Maddox. "Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels." Scienmag. October 1, 2026. https://scienmag.com/self-supervised-neighborhood-probing-shields-tabular-ai-from-poisoned-labels/

Tags: adversarial data poisoning in machine learningadversarial machine learningApplied IntelligenceBYOLdata poisoningdata sanitizationdata validation challenges in healthcaredefenses against label-based attacksfinancial fraud detection vulnerabilitieshigh-stakes decision-making securityimportance of label integrity in AIk-nearest neighborslabel flippinglabel flipping attack in supervised learningmachine learning securitynetwork intrusion detectionnoise and corruption in training datanovel techniques for poisoning resistancepoisoned label detection in tabular datarobustness of tabular AI modelsSCARFself-supervised learningSelf-supervised neighborhood probingtabular data
Share26Tweet16
Previous Post

Nature, Yards and Playgrounds: Where Preschoolers Spend Less Time on Screens

Next Post

Spiders’ Bodies Shift Their Elemental Makeup Dramatically Across a Single Season

Related Posts

Quantum Shield for the Cloud: Hybrid Cryptography Takes Aim at DDoS Attacks
Technology and Engineering

Quantum Shield for the Cloud: Hybrid Cryptography Takes Aim at DDoS Attacks

October 1, 2026
Frictional Interfaces Emit Strange Non-Local Waves That Defy Classical Rupture Mechanics
Technology and Engineering

Frictional Interfaces Emit Strange Non-Local Waves That Defy Classical Rupture Mechanics

October 1, 2026
Trustworthy AI Has a Toolkit Problem, Landmark Analysis of 938 Tools Reveals
Technology and Engineering

Trustworthy AI Has a Toolkit Problem, Landmark Analysis of 938 Tools Reveals

October 1, 2026
Volcanic Rock Fibers Are Poised to Reshape the Future of Thermoplastic Composites
Technology and Engineering

Volcanic Rock Fibers Are Poised to Reshape the Future of Thermoplastic Composites

October 1, 2026
Hybrid A* and Dynamic Window Method Steers Robots Past Obstacles
Technology and Engineering

Hybrid A* and Dynamic Window Method Steers Robots Past Obstacles

October 1, 2026
Plastic Fluff That Eats Plastic: Recycled Polymer Filters Snare Microplastics and Then Get a Second Job
Technology and Engineering

Plastic Fluff That Eats Plastic: Recycled Polymer Filters Snare Microplastics and Then Get a Second Job

October 1, 2026
Next Post
Spiders’ Bodies Shift Their Elemental Makeup Dramatically Across a Single Season

Spiders' Bodies Shift Their Elemental Makeup Dramatically Across a Single Season

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Common Cosmetic Preservative Methylparaben Linked to Breast Cancer Mechanisms in Landmark Computational Study
  • Spiders’ Bodies Shift Their Elemental Makeup Dramatically Across a Single Season
  • Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels
  • Nature, Yards and Playgrounds: Where Preschoolers Spend Less Time on Screens

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading