<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>financial fraud detection vulnerabilities &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/financial-fraud-detection-vulnerabilities/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 09:31:10 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>financial fraud detection vulnerabilities &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels</title>
		<link>https://scienmag.com/self-supervised-neighborhood-probing-shields-tabular-ai-from-poisoned-labels/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 09:31:10 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial data poisoning in machine learning]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[BYOL]]></category>
		<category><![CDATA[data poisoning]]></category>
		<category><![CDATA[data sanitization]]></category>
		<category><![CDATA[data validation challenges in healthcare]]></category>
		<category><![CDATA[defenses against label-based attacks]]></category>
		<category><![CDATA[financial fraud detection vulnerabilities]]></category>
		<category><![CDATA[high-stakes decision-making security]]></category>
		<category><![CDATA[importance of label integrity in AI]]></category>
		<category><![CDATA[k-nearest neighbors]]></category>
		<category><![CDATA[label flipping]]></category>
		<category><![CDATA[label flipping attack in supervised learning]]></category>
		<category><![CDATA[machine learning security]]></category>
		<category><![CDATA[network intrusion detection]]></category>
		<category><![CDATA[noise and corruption in training data]]></category>
		<category><![CDATA[novel techniques for poisoning resistance]]></category>
		<category><![CDATA[poisoned label detection in tabular data]]></category>
		<category><![CDATA[robustness of tabular AI models]]></category>
		<category><![CDATA[SCARF]]></category>
		<category><![CDATA[self-supervised learning]]></category>
		<category><![CDATA[Self-supervised neighborhood probing]]></category>
		<category><![CDATA[tabular data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221734</guid>

					<description><![CDATA[Researchers at Soongsil University have developed a self-supervised, multi-view defense that detects label-flipping poisoning attacks in tabular machine learning by probing neighborhood label disagreement in label-free embedding spaces.]]></description>
										<content:encoded><![CDATA[<p>Tabular machine learning quietly runs the modern world. Models trained on rows and columns of data decide whether a credit card transaction is fraudulent, whether a patient&#8217;s health indicators suggest diabetes, and whether a network connection is an intrusion. But a new study from researchers at Soongsil University in Seoul, published in Applied Intelligence, highlights a disturbing weakness in these systems: an attacker does not need to touch a single input feature to sabotage them. By simply flipping the labels attached to training samples, an adversary can distort the decision boundaries that these high-stakes models learn, and the corruption can be nearly invisible to conventional validation checks.</p>
<p>The research team, led by Jinhyeok Jang and corresponding author Daeseon Choi, frames the problem with unusual clarity. Label flipping is a form of data poisoning in which the observed class of a training example is changed while the underlying features remain intact. Because the features themselves look perfectly normal, simple feature-level screening catches nothing. Worse, verifying whether a label is truly correct in domains like finance or healthcare often requires scarce expert knowledge, so corrupted labels can survive unnoticed all the way into production. As organizations increasingly rely on crowdsourced annotation, outsourced labeling, and third-party data integration to cut costs, the attack surface only grows.</p>
<p>What makes the new work particularly timely is its focus on how modern label-flipping attacks have evolved. Early attacks flipped labels at random, scattering noise across the dataset. Recent strategies are far more surgical. Decision-boundary attacks use surrogate models to identify low-margin samples sitting perilously close to the classifier&#8217;s decision frontier, while optimization-based methods such as ALFA and its ALFA-Tilt variant solve constrained optimization problems to select the flip set that maximally degrades the learner. The authors&#8217; visualizations show that these targeted attacks produce localized poisoning patterns, with flipped samples clustering together in feature space rather than appearing as isolated outliers. That clustering defeats density-based and global outlier detectors, because the poisoned points masquerade as a legitimate local group.</p>
<p>The team&#8217;s answer is a defense built on a counterintuitive idea: learn what the data looks like without trusting any labels at all. Their method, called BYOL-Union, first trains a self-supervised encoder using BYOL, or Bootstrap Your Own Latent, a technique originally developed for images that learns representations by predicting augmented views of the same instance without any negative pairs or labels. Applied to tabular data, the encoder absorbs the intrinsic geometric structure of the feature space, untouched by whatever corruption may lurk in the observed labels. Only after this label-free representation learning is complete does the system consult the labels, and then solely to measure how much each sample disagrees with its neighbors.</p>
<p>The disagreement scoring is where the method earns its name. For each sample, the system finds its k nearest neighbors in the embedding space and counts how many of them carry a different observed label. A genuinely flipped sample tends to sit among neighbors whose true labels differ from its own forged one, producing a high disagreement score. Crucially, the method does not rely on a single fixed view of the neighborhood. Just as image-based self-supervised learning uses crops and flips to see one object from multiple angles, the framework generates a second view by adding Gaussian noise, with a standard deviation of 0.1, to continuous features only, leaving categorical features untouched to avoid inventing invalid categories. Each sample is then probed in both the original and the perturbed embedding spaces, and suspicious sets from the two views are merged by a union rule.</p>
<p>That union rule is deliberately recall-oriented. A sample flagged as suspicious in either view enters the final suspicious set, maximizing the chance of catching poisoned labels that are exposed in at least one neighborhood perspective, at the cost of some additional false positives. The authors also designed for realistic deployment, where the true poisoning ratio is unknown: instead of assuming an oracle budget, the practical version uses z-score thresholding at 1.0 with any-view exceedance, which their ablations show performs close to the oracle reference while remaining entirely oracle-free.</p>
<p>The evaluation is unusually broad. Six public tabular benchmarks spanning network intrusion detection, finance, and healthcare, including NSL-KDD, UNSW-NB15, Bank Marketing, Credit Card Fraud, BRFSS 2015 Diabetes Health Indicators, and the Diabetes 130-US Hospitals dataset, were poisoned at ratios of 10, 20, and 30 percent under random, decision-boundary, and optimization-based attacks. Four heterogeneous target models, spanning support vector machines, deep neural networks, FT-Transformers, and XGBoost, were then trained on sanitized data. The results show a consistent pattern: under decision-boundary flipping, the original feature-space kNN detector achieved a recall of only 0.601, while the BYOL-based and SCARF-based variants reached 0.773 and 0.792 respectively, and paired Wilcoxon tests confirmed that the SSL-based detectors significantly improved both recall and F1 over Curie, LS-SVM, and original-space kNN baselines.</p>
<p>The study is equally candid about limits. Under ALFA-Tilt, the strongest optimization-based attack tested in a small-scale setting, detection became markedly harder for every detector, and the robust-training defense FLORAL achieved the smallest downstream accuracy gap, though it cannot identify which specific samples are poisoned and is tied to particular model architectures. The authors stress that improved detection does not always translate directly into recovered test accuracy, because removing suspicious samples changes the training set itself and can thin out supervision for minority classes. Detection quality, they argue, should be judged on its own terms, with downstream recovery treated as a complementary, setting-dependent benefit rather than the sole criterion.</p>
<p>Perhaps the most striking result concerns adaptive adversaries. The team constructed defense-aware attacks in which the attacker knows the sanitization mechanism and selects flips that remain locally plausible in the detector&#8217;s own neighborhood space. Against a raw feature-space detector, such an adaptive attack cut recall from 0.668 to 0.435. Yet the SSL-based union detectors held firm: BYOL-Union maintained a recall of 0.760 even when the attacker targeted its own representation space, and SCARF-Union reached 0.800 under the corresponding adaptive attack. The layered evidence from learned representation spaces, it appears, is not fully dismantled by an adversary who optimizes against any single neighborhood view.</p>
<p>The broader lesson resonates beyond this one defense. As machine learning systems are deployed in domains where a single misclassified transaction or missed intrusion carries real consequences, the integrity of training labels deserves the same security attention as model architecture. The Soongsil team&#8217;s framework, which will see code released on request, offers a practical, model-agnostic preprocessing step: it flags suspicious samples before any downstream classifier is fit, works alongside tree ensembles and neural networks alike, and leaves room for future extensions toward label correction, sample reweighting, and robust retraining with uncertainty. In a field where attackers increasingly aim at the data rather than the model, defenses that learn to see the data on its own terms may prove essential.</p>
<p><strong>Subject of Research:</strong> Label-flipping attack detection and data sanitization in tabular machine learning using multi-view self-supervised representations</p>
<p><strong>Article Title:</strong> Multi-view self-supervised learning for label-flipping robustness in tabular data: a comparative study</p>
<p><strong>Article References:</strong> Multi-view self-supervised learning for label-flipping robustness in tabular data: a comparative study. (n.d.). <a href="https://doi.org/10.1007/s10489-026-07452-2" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07452-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07452-2" rel="noopener noreferrer">10.1007/s10489-026-07452-2</a></p>
<p><strong>Keywords:</strong> label flipping, data poisoning, self-supervised learning, BYOL, tabular data, data sanitization, k-nearest neighbors, adversarial machine learning, SCARF, network intrusion detection, machine learning security, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221734</post-id>	</item>
	</channel>
</rss>
