<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>adversarial machine learning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/adversarial-machine-learning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 22:42:51 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>adversarial machine learning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI Fingerprinting Method Separates Stolen Models From Innocent Lookalikes</title>
		<link>https://scienmag.com/new-ai-fingerprinting-method-separates-stolen-models-from-innocent-lookalikes/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 22:42:51 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[AI fingerprinting techniques]]></category>
		<category><![CDATA[AI model fingerprinting]]></category>
		<category><![CDATA[AI model security and integrity]]></category>
		<category><![CDATA[AI model watermarking methods]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[black-box verification]]></category>
		<category><![CDATA[CIFAR-10]]></category>
		<category><![CDATA[decision difference regions]]></category>
		<category><![CDATA[deep learning model theft detection]]></category>
		<category><![CDATA[deep neural networks]]></category>
		<category><![CDATA[differentiating stolen AI models from lookalikes]]></category>
		<category><![CDATA[intellectual property protection]]></category>
		<category><![CDATA[knowledge distillation]]></category>
		<category><![CDATA[model attribution and provenance]]></category>
		<category><![CDATA[model copycat behavior identification]]></category>
		<category><![CDATA[model copyright enforcement]]></category>
		<category><![CDATA[model fingerprinting]]></category>
		<category><![CDATA[model ownership verification]]></category>
		<category><![CDATA[model stealing]]></category>
		<category><![CDATA[neural network copy detection]]></category>
		<category><![CDATA[neural network ownership verification]]></category>
		<category><![CDATA[protecting trained neural networks]]></category>
		<category><![CDATA[watermarking]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=245705</guid>

					<description><![CDATA[Researchers in China have developed a deep neural network fingerprinting method that explicitly separates stolen model copies from independently trained lookalikes using decision difference regions.]]></description>
										<content:encoded><![CDATA[<p>Deep neural networks have become some of the most valuable assets in modern technology, representing months of engineering effort, expensive compute budgets, and carefully curated training data. Yet unlike software binaries, which can be protected by conventional copyright enforcement, a trained model is notoriously difficult to police. Anyone with access to a prediction API can, in principle, distill its behavior into a copycat network, fine-tune a leaked checkpoint, or otherwise reuse the model without permission. A research team at Dalian University of Technology, working with a colleague at the Chinese Academy of Sciences, has now published a new approach to this problem in the journal Applied Intelligence, one that reframes how ownership evidence for neural networks is constructed in the first place.</p>
<p>The study, led by Tianxiang Luo with corresponding author Bo Wang, addresses a subtle but dangerous weakness in existing model fingerprinting techniques. Fingerprints are sets of test inputs whose predicted labels are expected to be preserved when a model is copied, extracted, or post-processed. If a suspect model answers the fingerprint queries the same way the owner&#8217;s model does, that is treated as evidence of theft. The catch is that independently trained models, built from scratch on similar data, can also agree on many of these inputs simply because they learned the same task. When that happens, an innocent model developer can be falsely accused of stealing someone else&#8217;s work, a risk the authors describe as leaving owners exposed to false-accusation liability.</p>
<p>Previous methods largely treated this problem with a heuristic: they searched for inputs near the source model&#8217;s decision boundary, on the assumption that only models derived from the source would inherit those boundary quirks. The new work upgrades that heuristic into an explicit optimization objective. Instead of asking whether fingerprint inputs transfer to related models, the researchers formulate fingerprint construction as a search for decision difference regions, or DDRs. These are localized areas of the input space where the source model&#8217;s decision differs sharply from the decisions of independently trained benign models. In other words, the fingerprint is designed from the outset to separate stolen copies from innocent lookalikes, rather than hoping that boundary transferability will do the job on its own.</p>
<p>Formally, the team casts DDR identification as a search problem over the input space. The goal is to find regions where the source model and a pool of independently trained models disagree in their predictions, because such disagreement zones are precisely where a copied model, which inherits the source&#8217;s internal decision structure, is most likely to side with the source rather than with the benign pool. A suspect model that matches the source&#8217;s answers inside these regions therefore provides much stronger evidence of derivation than agreement on ordinary test data. The formulation turns fingerprinting from a passive property of chosen inputs into an active, targeted construction process with a measurable separation criterion.</p>
<p>Solving this search problem efficiently is the second contribution. Exhaustively probing a high-dimensional input space to find regions where a source model diverges from a pool of benign models would be prohibitively expensive, so the authors propose two non-invasive variants that trade construction cost against verification confidence. The Fast variant performs low-cost screening, quickly identifying candidate decision difference regions without heavy computation, making it suitable for scenarios where owners need a cheap first pass or must fingerprint many models. The Selected variant applies a more demanding offline filtering stage to the candidates, retaining only the regions that satisfy stricter separation criteria, and is intended for higher-confidence verification where the consequences of a wrong accusation are severe.</p>
<p>This two-tier design gives model owners something earlier fingerprinting schemes generally lacked: a tunable dial between cost and confidence. A company deploying hundreds of models could run the Fast variant routinely to maintain baseline protection, then invoke the Selected variant when a serious dispute actually arises. Because both variants are non-invasive, they require no modification of the model itself, no embedded watermarks, and no retraining. That distinguishes the approach from watermarking techniques, which plant hidden signals inside the network&#8217;s weights or behavior and can be damaged or removed when a thief fine-tunes, prunes, or compresses the stolen model.</p>
<p>The evaluation is notably broad for this area of research. The team tested the fingerprinting scheme on CIFAR-10 and the German Traffic Sign Recognition Benchmark, using CIFAR-100 and Tiny-ImageNet as stress tests to probe how the method behaves on harder and more diverse classification tasks. The experiments spanned multiple source architectures, including the kinds of residual and convolutional networks commonly deployed in practice, and confronted the fingerprints with hard negatives, meaning independently trained models that are deliberately similar to the source, as well as with extraction attacks such as knowledge distillation and knockoff-style model stealing. This combination matters because a fingerprint that only survives easy attacks or only avoids easy false positives is of little practical value.</p>
<p>According to the published results, the DDR-based fingerprints supported ownership verification across these settings while the authors were careful to identify the limits of what their method can prove. That honesty is itself significant. Much of the literature on model protection reports headline robustness numbers without clearly delineating where the evidence would fail, which is exactly the kind of overclaiming that produces false accusations in real disputes. By giving practitioners ownership evidence with an explicit account of where the separation holds and where it does not, the Dalian team is pushing the field toward a standard more familiar from forensic science: an expert should be able to state not just a conclusion but the boundary conditions of that conclusion.</p>
<p>The stakes extend well beyond academic benchmarks. As foundation models and specialized commercial networks are increasingly offered through APIs, model extraction has become a genuine industrial espionage channel, and courts and platforms are only beginning to grapple with how ownership of a statistical artifact should be established. A false accusation against an independent developer is not a harmless error; it can trigger takedowns, contract disputes, and reputational damage. Conversely, a fingerprint that fails against a determined thief leaves the actual owner without recourse. Methods that explicitly optimize for separating these two cases, and that quantify their own reliability, are a step toward making model ownership verification evidence that could actually survive adversarial scrutiny.</p>
<p>The researchers have also released their implementation publicly, including experiment scripts, configuration files, and the model-pool specifications used in the reported experiments, hosted on GitHub under the Dalian University of Technology AI lab. That reproducibility package, combined with the use of standard public benchmarks, means other groups can independently test the DDR approach, probe its failure modes, and build on the two-variant cost-confidence trade-off. The work was supported by the National Natural Science Foundation of China and the Youth Innovation Promotion Association of the Chinese Academy of Sciences. As AI models continue to function as trade secrets in everything from medical imaging to autonomous driving, techniques like decision difference region fingerprinting suggest a future in which the question of who owns a neural network can be answered not by assertion, but by measurement, with a clearly stated margin of error.</p>
<p><strong>Subject of Research:</strong> Deep neural network model ownership verification via decision difference region fingerprinting</p>
<p><strong>Article Title:</strong> Efficient and robust DNN model fingerprint with decision difference regions identification</p>
<p><strong>Article References:</strong> Luo, T., Yang, Z., Dai, X., Wang, B., &amp; Wang, W. (2026). Efficient and robust DNN model fingerprint with decision difference regions identification. <em>Applied Intelligence, 56</em>(14), Article 405. <a href="https://doi.org/10.1007/s10489-026-07440-6" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07440-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07440-6" rel="noopener noreferrer">10.1007/s10489-026-07440-6</a></p>
<p><strong>Keywords:</strong> deep neural networks, model fingerprinting, intellectual property protection, model ownership verification, black-box verification, decision difference regions, model stealing, watermarking, CIFAR-10, knowledge distillation, adversarial machine learning, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">245705</post-id>	</item>
		<item>
		<title>Hidden Fingerprints: New Method Uncovers Backdoor Targets in Compromised Neural Networks</title>
		<link>https://scienmag.com/hidden-fingerprints-new-method-uncovers-backdoor-targets-in-compromised-neural-networks/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 12:55:00 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[autonomous vehicle neural network safety]]></category>
		<category><![CDATA[backdoor attack vulnerabilities in neural networks]]></category>
		<category><![CDATA[backdoor attacks]]></category>
		<category><![CDATA[backdoor defenses]]></category>
		<category><![CDATA[class transfer pattern in poisoned neural networks]]></category>
		<category><![CDATA[combating backdoor attacks in facial recognition]]></category>
		<category><![CDATA[data poisoning]]></category>
		<category><![CDATA[deep neural networks]]></category>
		<category><![CDATA[detecting adversarial triggers in AI models]]></category>
		<category><![CDATA[hidden fingerprint analysis in deep learning]]></category>
		<category><![CDATA[K-ARM]]></category>
		<category><![CDATA[label noise]]></category>
		<category><![CDATA[methods for uncovering backdoor targets]]></category>
		<category><![CDATA[model security]]></category>
		<category><![CDATA[Neural Cleanse]]></category>
		<category><![CDATA[neural network backdoor detection]]></category>
		<category><![CDATA[neural network security and integrity]]></category>
		<category><![CDATA[statistical fingerprinting for model security]]></category>
		<category><![CDATA[training pipeline security in deep learning]]></category>
		<category><![CDATA[transition matrix]]></category>
		<category><![CDATA[trigger inversion]]></category>
		<category><![CDATA[watermark and trigger detection in AI systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=238076</guid>

					<description><![CDATA[Researchers have developed a transition matrix-based framework that identifies the target classes of backdoor attacks in neural networks and sharply improves existing defenses.]]></description>
										<content:encoded><![CDATA[<p>Deep neural networks now power everything from facial recognition systems to autonomous vehicles, but a quiet and dangerous threat has been lurking in their training pipelines. Known as backdoor attacks, these assaults embed invisible triggers into a model during training, so that the network behaves perfectly on ordinary inputs but flips to an attacker-chosen output the moment a specific pattern appears. A stop sign with a tiny sticker, a face with a subtle perturbation, a document with a nearly imperceptible watermark — any of these can hijack the model&#8217;s decision. A new study published in Applied Intelligence by Wenbin Jiang, Shikang Xie, Jiqiang Liu, Nan Jiang and Jian Wang, researchers at Beijing Jiaotong University and Beijing University of Technology, offers a fresh way to fight back by reading the statistical fingerprints that such attacks inevitably leave behind in a model&#8217;s predictions.</p>
<p>The core insight of the work is deceptively simple: attackers have intentions, and those intentions impose structure. When a backdoor is active, poisoned samples are systematically pushed from their true classes toward a single target class chosen by the attacker. This creates what the authors call a class transfer pattern — a measurable, recurring flow of predictions from many source classes into one destination class. While most existing defenses concentrate either on discovering the trigger itself or on unlearning the backdoor behavior, they often struggle to characterize which class the attacker actually targeted. The new framework flips the problem around: instead of hunting for the trigger first, it estimates the pattern of class-to-class prediction transfers and uses that estimate to pinpoint the target class, even when no prior knowledge of the attack exists.</p>
<p>Technically, the method builds on an idea borrowed from the study of label noise. In noisy-label learning, researchers model the probability that a true label is flipped into an incorrect one using a transition matrix — a table whose entries describe how likely each class is to be misread as every other class. The authors adapt this machinery to the backdoor setting. Because poisoned samples are a small minority of the training data, the attack-induced transfers appear as large off-diagonal entries in an otherwise diagonal-dominant matrix: a well-trained model correctly classifies most samples, so the diagonal entries, representing correct predictions, dominate, while systematic misclassifications caused by poisoning stand out as anomalous rows and columns. Estimating this matrix therefore becomes a way of exposing the attack&#8217;s geometry.</p>
<p>The estimation itself proceeds in two steps and is driven by an improved information-theoretic objective. The researchers draw on the determinant of the joint matrix between predictions and labels, a loss function known in the label-noise literature as DMI, which penalizes degenerate transition estimates. In an appendix, the team provides a formal proof bounding the gradient of this loss: using the matrix differential identity for the log-determinant and Weyl&#8217;s inequality on singular values, they show that the per-sample gradient norm is bounded by the inverse of the smallest singular value of the joint matrix. This matters because it formalizes when the estimation is stable — namely, when the diagonal clean-prediction component remains stronger than the off-diagonal attack-induced component, a condition expected to hold whenever the poisoning ratio is moderate and the model retains high clean accuracy.</p>
<p>Once the transition matrix has been estimated, the identified target class becomes a powerful lever for existing defenses. Trigger-inversion methods, such as the widely used Neural Cleanse and the K-ARM optimization approach, work by reverse-engineering the smallest perturbation that can force any input to be classified as a suspect target. The trouble is that these methods must search across every possible target class, reconstructing candidate triggers for each one — a computationally expensive process prone to false positives. By narrowing the search space to the class flagged by the transition matrix, the new framework eliminates unnecessary trigger reconstruction while preserving mitigation effectiveness. The defense knows where to look before it starts looking.</p>
<p>The experimental results are striking. Across multiple datasets and attack types, the enhanced Neural Cleanse variant achieved an average attack success rate of just 2.17 percent, while the enhanced K-ARM variant drove the figure down to 1.46 percent. The strongest baseline, by comparison, averaged an attack success rate of 3.68 percent. Crucially, these gains did not come at the cost of normal performance: the defended models maintained higher clean accuracy than the baselines, meaning legitimate users would notice no degradation in everyday behavior. The evaluation covered standard image benchmarks including CIFAR-10, CIFAR-100 and Tiny ImageNet, all publicly available datasets, and spanned a range of attack designs from classic patch-based triggers to more modern imperceptible variants.</p>
<p>The breadth of attacks considered reflects the evolving threat landscape. The study&#8217;s reference list traces the lineage of backdoor research from BadNets, the 2017 work that first identified vulnerabilities in the machine learning supply chain, through targeted data-poisoning attacks, Trojaning attacks, reflection-based natural backdoors, warping-based WaNet triggers and label-consistent attacks that try to evade detection by aligning poisoned samples with their labels. Each generation of attacks has grown stealthier, and frequency-domain analyses have shown that triggers can hide in spectral regions humans barely perceive. A defense that does not depend on the specific form of the trigger — but instead on the invariant statistical consequence of any targeted attack — is inherently more robust to this arms race.</p>
<p>That invariance is precisely what makes the approach compelling. Whether an attacker uses a visible patch, a transparent overlay, a warping field or a frequency-domain perturbation, the end goal is the same: route inputs from diverse source classes into one chosen target. The transition matrix captures this routing behavior directly. The authors&#8217; theoretical analysis reinforces the point by decomposing the joint matrix into a diagonal component, representing the model&#8217;s legitimate functionality on the majority of samples, and an off-diagonal component, where large entries signal systematic misclassification toward the attacker&#8217;s target. Under the reasonable assumption that no exact linear dependencies exist among classes, the matrix is invertible, and its spectral properties govern the reliability of the estimate.</p>
<p>The practical implications extend well beyond computer vision benchmarks. Backdoor attacks are a supply-chain problem: models downloaded from public repositories, trained by third-party contractors or fine-tuned on crowdsourced data can all carry hidden triggers. Scanning such models before deployment is becoming a standard security practice, and methods that make scanning faster and more accurate have immediate value. By telling trigger-inversion defenses which class to investigate, the transition-matrix framework could reduce the computational cost of auditing large models, an increasingly important consideration as networks grow to billions of parameters. The same reasoning may eventually extend to language models, where backdoor token unlearning has emerged as a parallel research frontier.</p>
<p>Limitations remain, as with any defense. The framework&#8217;s stability guarantee rests on the assumption that correct predictions dominate misclassifications, which could weaken under extreme poisoning ratios or attacks that spread their effect across multiple targets. The authors acknowledge that the approach assumes moderate poisoning and high clean accuracy — conditions that describe most realistic attacks but not necessarily the most aggressive adversarial scenarios. Still, the work represents a meaningful conceptual shift: treating the attack target not as a mystery to be brute-forced but as a statistical signature to be estimated. As deep learning systems take on higher-stakes roles in medicine, infrastructure and transportation, defenses that exploit the inherent structure of the attacker&#8217;s intent — rather than the specifics of any single trigger — may prove to be the durable line of protection the field has been searching for.</p>
<p><strong>Subject of Research:</strong> Backdoor attack defense in deep neural networks using class transfer matrix estimation</p>
<p><strong>Article Title:</strong> Class transfer estimation in neural networks for backdoor target identification and mitigation</p>
<p><strong>Article References:</strong> Class transfer estimation in neural networks for backdoor target identification and mitigation. (n.d.). <a href="https://doi.org/10.1007/s10489-026-07470-0" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07470-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07470-0" rel="noopener noreferrer">10.1007/s10489-026-07470-0</a></p>
<p><strong>Keywords:</strong> deep neural networks, backdoor attacks, backdoor defenses, transition matrix, trigger inversion, Neural Cleanse, K-ARM, label noise, adversarial machine learning, model security, data poisoning, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">238076</post-id>	</item>
		<item>
		<title>Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels</title>
		<link>https://scienmag.com/self-supervised-neighborhood-probing-shields-tabular-ai-from-poisoned-labels/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 09:31:10 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial data poisoning in machine learning]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[BYOL]]></category>
		<category><![CDATA[data poisoning]]></category>
		<category><![CDATA[data sanitization]]></category>
		<category><![CDATA[data validation challenges in healthcare]]></category>
		<category><![CDATA[defenses against label-based attacks]]></category>
		<category><![CDATA[financial fraud detection vulnerabilities]]></category>
		<category><![CDATA[high-stakes decision-making security]]></category>
		<category><![CDATA[importance of label integrity in AI]]></category>
		<category><![CDATA[k-nearest neighbors]]></category>
		<category><![CDATA[label flipping]]></category>
		<category><![CDATA[label flipping attack in supervised learning]]></category>
		<category><![CDATA[machine learning security]]></category>
		<category><![CDATA[network intrusion detection]]></category>
		<category><![CDATA[noise and corruption in training data]]></category>
		<category><![CDATA[novel techniques for poisoning resistance]]></category>
		<category><![CDATA[poisoned label detection in tabular data]]></category>
		<category><![CDATA[robustness of tabular AI models]]></category>
		<category><![CDATA[SCARF]]></category>
		<category><![CDATA[self-supervised learning]]></category>
		<category><![CDATA[Self-supervised neighborhood probing]]></category>
		<category><![CDATA[tabular data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221734</guid>

					<description><![CDATA[Researchers at Soongsil University have developed a self-supervised, multi-view defense that detects label-flipping poisoning attacks in tabular machine learning by probing neighborhood label disagreement in label-free embedding spaces.]]></description>
										<content:encoded><![CDATA[<p>Tabular machine learning quietly runs the modern world. Models trained on rows and columns of data decide whether a credit card transaction is fraudulent, whether a patient&#8217;s health indicators suggest diabetes, and whether a network connection is an intrusion. But a new study from researchers at Soongsil University in Seoul, published in Applied Intelligence, highlights a disturbing weakness in these systems: an attacker does not need to touch a single input feature to sabotage them. By simply flipping the labels attached to training samples, an adversary can distort the decision boundaries that these high-stakes models learn, and the corruption can be nearly invisible to conventional validation checks.</p>
<p>The research team, led by Jinhyeok Jang and corresponding author Daeseon Choi, frames the problem with unusual clarity. Label flipping is a form of data poisoning in which the observed class of a training example is changed while the underlying features remain intact. Because the features themselves look perfectly normal, simple feature-level screening catches nothing. Worse, verifying whether a label is truly correct in domains like finance or healthcare often requires scarce expert knowledge, so corrupted labels can survive unnoticed all the way into production. As organizations increasingly rely on crowdsourced annotation, outsourced labeling, and third-party data integration to cut costs, the attack surface only grows.</p>
<p>What makes the new work particularly timely is its focus on how modern label-flipping attacks have evolved. Early attacks flipped labels at random, scattering noise across the dataset. Recent strategies are far more surgical. Decision-boundary attacks use surrogate models to identify low-margin samples sitting perilously close to the classifier&#8217;s decision frontier, while optimization-based methods such as ALFA and its ALFA-Tilt variant solve constrained optimization problems to select the flip set that maximally degrades the learner. The authors&#8217; visualizations show that these targeted attacks produce localized poisoning patterns, with flipped samples clustering together in feature space rather than appearing as isolated outliers. That clustering defeats density-based and global outlier detectors, because the poisoned points masquerade as a legitimate local group.</p>
<p>The team&#8217;s answer is a defense built on a counterintuitive idea: learn what the data looks like without trusting any labels at all. Their method, called BYOL-Union, first trains a self-supervised encoder using BYOL, or Bootstrap Your Own Latent, a technique originally developed for images that learns representations by predicting augmented views of the same instance without any negative pairs or labels. Applied to tabular data, the encoder absorbs the intrinsic geometric structure of the feature space, untouched by whatever corruption may lurk in the observed labels. Only after this label-free representation learning is complete does the system consult the labels, and then solely to measure how much each sample disagrees with its neighbors.</p>
<p>The disagreement scoring is where the method earns its name. For each sample, the system finds its k nearest neighbors in the embedding space and counts how many of them carry a different observed label. A genuinely flipped sample tends to sit among neighbors whose true labels differ from its own forged one, producing a high disagreement score. Crucially, the method does not rely on a single fixed view of the neighborhood. Just as image-based self-supervised learning uses crops and flips to see one object from multiple angles, the framework generates a second view by adding Gaussian noise, with a standard deviation of 0.1, to continuous features only, leaving categorical features untouched to avoid inventing invalid categories. Each sample is then probed in both the original and the perturbed embedding spaces, and suspicious sets from the two views are merged by a union rule.</p>
<p>That union rule is deliberately recall-oriented. A sample flagged as suspicious in either view enters the final suspicious set, maximizing the chance of catching poisoned labels that are exposed in at least one neighborhood perspective, at the cost of some additional false positives. The authors also designed for realistic deployment, where the true poisoning ratio is unknown: instead of assuming an oracle budget, the practical version uses z-score thresholding at 1.0 with any-view exceedance, which their ablations show performs close to the oracle reference while remaining entirely oracle-free.</p>
<p>The evaluation is unusually broad. Six public tabular benchmarks spanning network intrusion detection, finance, and healthcare, including NSL-KDD, UNSW-NB15, Bank Marketing, Credit Card Fraud, BRFSS 2015 Diabetes Health Indicators, and the Diabetes 130-US Hospitals dataset, were poisoned at ratios of 10, 20, and 30 percent under random, decision-boundary, and optimization-based attacks. Four heterogeneous target models, spanning support vector machines, deep neural networks, FT-Transformers, and XGBoost, were then trained on sanitized data. The results show a consistent pattern: under decision-boundary flipping, the original feature-space kNN detector achieved a recall of only 0.601, while the BYOL-based and SCARF-based variants reached 0.773 and 0.792 respectively, and paired Wilcoxon tests confirmed that the SSL-based detectors significantly improved both recall and F1 over Curie, LS-SVM, and original-space kNN baselines.</p>
<p>The study is equally candid about limits. Under ALFA-Tilt, the strongest optimization-based attack tested in a small-scale setting, detection became markedly harder for every detector, and the robust-training defense FLORAL achieved the smallest downstream accuracy gap, though it cannot identify which specific samples are poisoned and is tied to particular model architectures. The authors stress that improved detection does not always translate directly into recovered test accuracy, because removing suspicious samples changes the training set itself and can thin out supervision for minority classes. Detection quality, they argue, should be judged on its own terms, with downstream recovery treated as a complementary, setting-dependent benefit rather than the sole criterion.</p>
<p>Perhaps the most striking result concerns adaptive adversaries. The team constructed defense-aware attacks in which the attacker knows the sanitization mechanism and selects flips that remain locally plausible in the detector&#8217;s own neighborhood space. Against a raw feature-space detector, such an adaptive attack cut recall from 0.668 to 0.435. Yet the SSL-based union detectors held firm: BYOL-Union maintained a recall of 0.760 even when the attacker targeted its own representation space, and SCARF-Union reached 0.800 under the corresponding adaptive attack. The layered evidence from learned representation spaces, it appears, is not fully dismantled by an adversary who optimizes against any single neighborhood view.</p>
<p>The broader lesson resonates beyond this one defense. As machine learning systems are deployed in domains where a single misclassified transaction or missed intrusion carries real consequences, the integrity of training labels deserves the same security attention as model architecture. The Soongsil team&#8217;s framework, which will see code released on request, offers a practical, model-agnostic preprocessing step: it flags suspicious samples before any downstream classifier is fit, works alongside tree ensembles and neural networks alike, and leaves room for future extensions toward label correction, sample reweighting, and robust retraining with uncertainty. In a field where attackers increasingly aim at the data rather than the model, defenses that learn to see the data on its own terms may prove essential.</p>
<p><strong>Subject of Research:</strong> Label-flipping attack detection and data sanitization in tabular machine learning using multi-view self-supervised representations</p>
<p><strong>Article Title:</strong> Multi-view self-supervised learning for label-flipping robustness in tabular data: a comparative study</p>
<p><strong>Article References:</strong> Multi-view self-supervised learning for label-flipping robustness in tabular data: a comparative study. (n.d.). <a href="https://doi.org/10.1007/s10489-026-07452-2" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07452-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07452-2" rel="noopener noreferrer">10.1007/s10489-026-07452-2</a></p>
<p><strong>Keywords:</strong> label flipping, data poisoning, self-supervised learning, BYOL, tabular data, data sanitization, k-nearest neighbors, adversarial machine learning, SCARF, network intrusion detection, machine learning security, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221734</post-id>	</item>
		<item>
		<title>Hidden Switches: New Attack Plants Undetectable Backdoors in Vision Transformers</title>
		<link>https://scienmag.com/hidden-switches-new-attack-plants-undetectable-backdoors-in-vision-transformers/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 23:11:11 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial attacks on computer vision models]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[AI safety]]></category>
		<category><![CDATA[backdoor attack]]></category>
		<category><![CDATA[backdoor attack in machine learning]]></category>
		<category><![CDATA[backdoor detection in vision transformers]]></category>
		<category><![CDATA[covert backdoor triggers in deep learning]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[cybersecurity threats in AI]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning model robustness]]></category>
		<category><![CDATA[hidden switches in neural networks]]></category>
		<category><![CDATA[high-stakes AI application vulnerabilities]]></category>
		<category><![CDATA[image classification]]></category>
		<category><![CDATA[machine learning security]]></category>
		<category><![CDATA[medical imaging model security]]></category>
		<category><![CDATA[model integrity in autonomous systems]]></category>
		<category><![CDATA[model poisoning]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[self-attention mechanism exploitation]]></category>
		<category><![CDATA[supply chain security]]></category>
		<category><![CDATA[vision transformer security vulnerabilities]]></category>
		<category><![CDATA[Vision Transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=219986</guid>

					<description><![CDATA[Researchers have demonstrated a switchable backdoor attack that injects dual tokens through the layers of Vision Transformers, achieving up to 99 percent attack success while leaving clean accuracy intact.]]></description>
										<content:encoded><![CDATA[<p>Vision Transformers have quietly become the backbone of modern computer vision, powering everything from image classifiers and object detectors to medical imaging pipelines and autonomous driving prototypes. Their self-attention mechanism, which lets the model weigh relationships between every patch of an image, has displaced convolutional networks in many high-stakes applications. But a new study from researchers at the University of Information Technology, Ho Chi Minh City, and Vietnam National University Ho Chi Minh City suggests that the very architecture making these models so powerful may also make them dangerously easy to compromise. In a paper published in the International Journal of Machine Learning and Cybernetics, Dung Minh Do and Khang Nguyen describe a backdoor attack that achieves attack success rates of up to 99 percent while leaving the model&#8217;s performance on clean, unmodified images essentially untouched.</p>
<p>Backdoor attacks are among the most insidious threats in machine learning security. Unlike adversarial examples, which exploit a model at inference time by perturbing individual inputs, a backdoor is baked into the model itself during training or fine-tuning. A model with a backdoor behaves perfectly normally on ordinary data, matching the accuracy and reliability its developers expect. But when a specific trigger, a subtle pattern chosen by the attacker, appears in an input, the model&#8217;s behavior flips, producing whatever output the attacker has designated. For a deployed system, this means an adversary could, for example, cause a traffic sign classifier to misread a stop sign or a security screening system to wave through a prohibited item, all without any visible degradation in everyday performance that might raise alarms.</p>
<p>What makes the new attack notable is where it operates. Rather than modifying the weights of the transformer or poisoning the training labels in conventional ways, the method injects two specially designed tokens into the model&#8217;s token stream, progressively working from the shallowest layers to the deepest ones. Vision Transformers process images by splitting them into patches, embedding each patch as a token, and then passing those tokens through a stack of transformer blocks in which self-attention layers let tokens exchange information. By inserting crafted tokens at multiple depths, the attack ensures that the malicious signal is reinforced and refined as it travels through the network, rather than being diluted or overwritten by the model&#8217;s normal processing.</p>
<p>The progressive, layer-by-layer nature of the injection is central to the attack&#8217;s stealthiness. Defenses that inspect a single layer, or that look for anomalous attention patterns at one point in the network, can miss a signal that is distributed across the entire depth of the model. Because the injected tokens interact with the image tokens through the standard attention mechanism, they can steer the model&#8217;s internal representations toward the attacker&#8217;s chosen target class only when the trigger is present. On clean inputs, the tokens remain effectively inert, which is why the authors report that the method maintains competitive clean accuracy even as it delivers near-perfect attack success rates across multiple visual datasets.</p>
<p>The word switchable in the study&#8217;s title points to another dimension of the threat. The attack draws on a growing line of research into switchable backdoors against pre-trained vision transformers, in which a single compromised model can carry multiple behaviors that an attacker can toggle. This means a defender who discovers one trigger and filters it out cannot assume the model is safe; another trigger may remain dormant, waiting to be activated. The authors position their work as an analysis and exploitation of ViT vulnerabilities, and the breadth of prior work they survey, from BadNets and weight-poisoning attacks to attention hijacking and prompt-based backdoors, underscores how rapidly this attack surface has expanded as transformers have spread through computer vision.</p>
<p>The connection to prompt tuning is particularly significant for the current state of the field. Prompt-based methods, in which small learnable tokens are prepended to a model&#8217;s input or inserted into its layers, have become a popular way to adapt large pre-trained models to new tasks without expensive full fine-tuning. Visual prompt tuning and related techniques are widely used because they are efficient and effective. But the same mechanism that makes prompts useful, namely the ability to inject learned tokens that influence the model&#8217;s attention and representations, is exactly what this attack weaponizes. A malicious actor with access to a fine-tuning pipeline, or able to distribute a poisoned adapter or checkpoint, could embed a backdoor that looks indistinguishable from a legitimate prompt-based adaptation.</p>
<p>The supply chain implications are sobering. Modern machine learning practice relies heavily on pre-trained models downloaded from public hubs, fine-tuned adapters shared between teams, and third-party datasets. Each of these channels is a potential delivery mechanism for a backdoor. Earlier research has shown that weight poisoning attacks on pre-trained models can survive downstream fine-tuning, and that backdoors can be hidden in ways that evade standard inspection. The new study adds to this picture by showing that the token-based machinery of vision transformers offers attackers a particularly clean injection point, one that does not require the crude modifications of earlier attacks that made them easier to detect.</p>
<p>The authors evaluated their method on multiple visual datasets, measuring not only attack success rate and clean accuracy but also stealthiness and robustness. The headline figure, an attack success rate of up to 99 percent, is alarming enough, but the more troubling result is the combination of that success with preserved clean performance, since it means conventional accuracy-based validation would reveal nothing amiss. The paper also situates the work against existing defenses, which include fine-pruning approaches that remove rarely activated neurons, neural attention distillation that tries to erase trigger-related attention patterns, and input-level detection methods that look for inconsistencies in a model&#8217;s predictions under image transformations. Because the new attack distributes its signal across shallow and deep layers through dual tokens, many of these defenses, which were designed with convolutional networks or single-point injections in mind, face a harder problem.</p>
<p>The researchers are explicit about the warning their findings carry for real-world deployment. Vision Transformers are increasingly used in settings where a silent failure mode could have serious consequences, including medical diagnosis support, surveillance, industrial inspection, and safety-critical perception systems. A backdoored model in any of these contexts could be remotely triggered by an input crafted to contain the attacker&#8217;s pattern, and the compromise would be invisible in routine testing. The authors make their code publicly available, which serves the defensive side of the field as well: reproducible attack implementations are essential for developing and benchmarking countermeasures, and the history of adversarial machine learning shows that security research advances fastest when attacks are fully documented.</p>
<p>For the broader community, the study is a reminder that architectural progress and security progress have been badly out of step. The references in the paper trace a decade of deep learning breakthroughs, from early convolutional networks through EfficientNet and the original Vision Transformer, alongside a parallel literature of backdoor learning surveys, prompt injection analyses, and defense proposals. Yet the authors note that security threats to ViTs, particularly backdoor attacks, have not received research attention commensurate with the architecture&#8217;s adoption. Closing that gap will likely require defenses designed specifically for token-based architectures: methods that audit injected tokens, verify the provenance of fine-tuned checkpoints, test models against a family of triggers rather than a single known pattern, and treat the entire depth of the network, not just its input layer, as a potential attack surface. Until such defenses mature, the near-perfect stealth and effectiveness demonstrated by progressive dual-token injection stands as a stark warning that the models powering tomorrow&#8217;s vision systems may harbor switches that only their attackers know how to flip.</p>
<p><strong>Subject of Research:</strong> Backdoor attack vulnerabilities in Vision Transformers via progressive dual-token injection</p>
<p><strong>Article Title:</strong> Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers</p>
<p><strong>Article References:</strong> Do, D. M., &amp; Nguyen, K. (2026). Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 489. <a href="https://doi.org/10.1007/s13042-026-03326-8" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03326-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03326-8" rel="noopener noreferrer">10.1007/s13042-026-03326-8</a></p>
<p><strong>Keywords:</strong> Vision Transformers, backdoor attack, machine learning security, self-attention, prompt tuning, adversarial machine learning, model poisoning, supply chain security, deep learning, cybersecurity, image classification, AI safety</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">219986</post-id>	</item>
		<item>
		<title>Genetic Algorithms and GANs Converge: New Scientometric Map Reveals a Booming Field</title>
		<link>https://scienmag.com/genetic-algorithms-and-gans-converge-new-scientometric-map-reveals-a-booming-field/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 21:48:06 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[Bibliometric analysis]]></category>
		<category><![CDATA[collaboration networks]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[generative adversarial networks]]></category>
		<category><![CDATA[genetic algorithm]]></category>
		<category><![CDATA[hyperparameter optimization]]></category>
		<category><![CDATA[multi-objective optimization]]></category>
		<category><![CDATA[research trends]]></category>
		<category><![CDATA[science mapping]]></category>
		<category><![CDATA[scientometric review]]></category>
		<category><![CDATA[VOSviewer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205123</guid>

					<description><![CDATA[A new scientometric review maps the rapid rise of research combining genetic algorithms with generative adversarial networks, revealing explosive growth since 2020 alongside persistent gaps in evaluation and collaboration.]]></description>
										<content:encoded><![CDATA[<p>Two of artificial intelligence&#8217;s most powerful yet fundamentally different tools are quietly merging, and a new open-access scientometric review has, for the first time, mapped exactly how fast, how broadly, and how unevenly that merger is unfolding. Genetic Algorithms (GAs), the evolutionary search techniques inspired by natural selection, and Generative Adversarial Networks (GANs), the deep generative models trained through a duel between a generator and a discriminator, have been crossing paths with increasing frequency since 2017. A comprehensive review published in Discover Informatics by Basil Hanafi of Galgotias University and Mohammad Ali of Aligarh Muslim University analyzed 276 peer-reviewed publications retrieved from Scopus and Web of Science, and its findings reveal a research field that has exploded from near invisibility into a globally distributed, thematically rich, but still methodologically immature discipline.</p>
<p>The numbers tell a striking growth story. From a single indexed publication in 2017 and just two in 2018, annual output climbed to 17 papers in 2020, surged to 44 in 2021 and 51 in 2022, and peaked at 92 publications in 2024, the largest yearly figure in the dataset. The authors caution that the 13 records assigned to 2025 reflect only a partial year, since data retrieval ended on 23 January 2025. Overall, the corpus spans 204 sources and involves 878 authors with an average of 3.94 co-authors per paper, alongside 831 author keywords and 2,076 Keywords Plus terms. Average citation counts per document reached 11.52, with mean document age of just 2.46 years, hallmarks of a field that has grown so recently and so quickly that most of its literature has not yet reached citation maturity.</p>
<p>Why would anyone combine an evolutionary algorithm from the 1970s with a neural network architecture from 2014? The technical rationale, the review explains, lies in the notorious difficulty of training GANs. Adversarial models suffer from unstable convergence, hyperparameter sensitivity, generator–discriminator imbalance, and the chronic problem of mode collapse, where the generator produces only a narrow slice of possible outputs. GAs offer a gradient-free, population-based search that can explore many candidate configurations simultaneously, tolerate discontinuous and non-differentiable objective surfaces, and balance competing goals such as output quality, training stability, and computational efficiency. Where Bayesian optimization struggles with noisy, high-dimensional spaces and reinforcement-learning controllers add sequential training overhead, evolutionary search provides an alternative tuning strategy that has proved attractive for hyperparameter selection, architecture evolution, and multi-objective trade-offs.</p>
<p>The review&#8217;s literature synthesis identifies several landmark contributions that define the field&#8217;s methodological core. Alarsan and Younes&#8217;s GANGA framework applied genetic algorithms to GAN hyperparameter optimization and improved convergence on MNIST, while Wang and colleagues reframed adversarial training itself as an evolutionary process, mutating and selecting generator variants. On the reverse side of the convergence, He and colleagues introduced GMOEA, a GAN-driven multi-objective evolutionary algorithm that enhances convergence and diversity in high-dimensional search spaces, demonstrating that the pairing benefits optimizers as much as generators. In image translation, Xue and colleagues built AevoGAN, embedding evolutionary algorithms and channel attention into a CycleGAN architecture, and Konstantopoulou and colleagues developed GAGAN, using genetic operations to improve discriminator optimization and reduce mode collapse. Beyond imaging, Le Vine and colleagues combined conditional GANs with genetic algorithms to extract table structures from scanned documents, showing the hybrid&#8217;s utility in structured pattern recognition.</p>
<p>Geographically, the field is global but heavily concentrated. China dominates with 135 publications, followed by India with 47 and the United States with 30, while 47 countries contribute at least one paper. Cumulative curves show China&#8217;s steep climb from a single publication in 2018 to triple-digit totals, with India&#8217;s expansion accelerating sharply after 2021. Citation influence, however, tells a more nuanced story: China leads in total citations with 787, but Australia achieves the highest per-article average at 254 citations, and the corresponding-author analysis shows that multi-country collaboration remains modest, at roughly 5.4 percent of output. Tongji University tops institutional productivity with eight publications, ahead of MIT, Shanghai Jiao Tong University, and the University of Coimbra. Author productivity follows a classic Lotka-type distribution, with 87.24 percent of authors publishing only once, while a small recurring core sustains the field&#8217;s methodological continuity.</p>
<p>Thematically, keyword and co-occurrence analyses anchor the field in three intertwined strands: evolutionary optimization, adversarial generative modeling, and general deep-learning methodology. The most frequent descriptors are genetic algorithms (113 occurrences) and generative adversarial networks (111), trailed by deep learning (69), adversarial networks (48), and the recently emergent adversarial machine learning, which did not appear before 2024. Source analysis reveals moderate concentration consistent with Bradford&#8217;s law: IEEE Access leads with nine papers and Lecture Notes in Computer Science with eight, while a long tail of outlets contributes one or two documents each. The most globally cited record is a 2020 Renewable and Sustainable Energy Reviews paper on photovoltaic power forecasting with 759 citations, followed by influential works in de novo drug design, neuroscience, topology optimization, network intrusion detection, and urban design, evidence that the field&#8217;s visibility spans energy, security, biomedicine, and structural engineering alike.</p>
<p>The review is careful not to overstate the case for genetic algorithms. The authors emphasize that GA is not inherently superior to gradient-based tuning, Bayesian optimization, or reinforcement-learning controllers, and that its contribution is context-sensitive methodological support rather than a universal solution. Evolutionary gains often come at significant computational cost, and much of the supporting evidence rests on small-scale benchmarks, particularly classic 2D image datasets, limiting generalizability. Evaluation practice also remains inconsistent: studies variously report Fréchet Inception Distance, Learned Perceptual Image Patch Similarity, Inception Score, Peak Signal-to-Noise Ratio, and Structural Similarity Index depending on task and dataset, with no common framework for comparing quality, robustness, and efficiency across domains. This evaluation inconsistency, the authors argue, is one of the field&#8217;s most persistent structural weaknesses.</p>
<p>Looking forward, the scientometric evidence points to six priority directions. Scalability tops the list, since evolutionary enhancements frequently inflate computational load on already expensive models. Training stability and diversity preservation through adaptive evolutionary schemes remain unsettled. Standardized benchmarking frameworks are needed to substantiate cross-domain claims. The breadth of applications, from forecasting and cybersecurity to biomedical diagnosis and urban design, has grown faster than empirical maturity, demanding rigorous domain-specific validation with larger and more varied datasets. Interpretability and trustworthiness become critical as these systems enter healthcare and security contexts where outputs influence consequential decisions. Finally, the field&#8217;s low levels of international and interdisciplinary collaboration suggest that broader cooperation could accelerate progress where optimization, generative modeling, and domain science must intersect.</p>
<p>The larger significance of the study lies in its demonstration that scientometric mapping can discipline a fast-moving AI subfield. Rather than relying on anecdotal impressions of a hot topic, the review quantifies where GA–GAN research is concentrated, which themes are genuinely central, and where the literature is thin. Its portrait is one of a field with unambiguous methodological promise and genuine cross-domain reach, but one still held back by computational expense, optimization instability, uneven evaluation, and fragmented collaboration. Whether the GA–GAN convergence matures into a foundational hybrid methodology or remains a niche toolset, the new map gives researchers, funders, and practitioners a precise picture of the terrain they are entering, and a clear signal of where the next advances are most likely to come from.</p>
<p><strong>Subject of Research:</strong> Scientometric analysis of the research convergence between genetic algorithms and generative adversarial networks</p>
<p><strong>Article Title:</strong> Quantifying the research convergence of optimized generative adversarial networks with genetic algorithm using scientometric review analysis</p>
<p><strong>Article References:</strong> Hanafi, B., &amp; Ali, M. (2026). Quantifying the research convergence of optimized generative adversarial networks with genetic algorithm using scientometric review analysis. <em>Discover Informatics, 1</em>(1), Article 2. <a href="https://doi.org/10.1007/s44564-026-00002-5" rel="noopener noreferrer">https://doi.org/10.1007/s44564-026-00002-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44564-026-00002-5" rel="noopener noreferrer">10.1007/s44564-026-00002-5</a></p>
<p><strong>Keywords:</strong> genetic algorithm, generative adversarial networks, scientometric review, bibliometric analysis, science mapping, hyperparameter optimization, multi-objective optimization, adversarial machine learning, deep learning, research trends, collaboration networks, VOSviewer</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205123</post-id>	</item>
		<item>
		<title>Hackers Can Hijack Graph AI With Just a Handful of Poisoned Samples</title>
		<link>https://scienmag.com/hackers-can-hijack-graph-ai-with-just-a-handful-of-poisoned-samples/</link>
		
		<dc:creator><![CDATA[Hailey Crawford]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:18:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial attacks on graph-based systems]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[backdoor attacks]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[cybersecurity risks in graph neural networks]]></category>
		<category><![CDATA[data poisoning]]></category>
		<category><![CDATA[data-level prompt injection]]></category>
		<category><![CDATA[defenses against prompt injection attacks]]></category>
		<category><![CDATA[efficient graph learning vulnerabilities]]></category>
		<category><![CDATA[Few-shot learning]]></category>
		<category><![CDATA[frozen encoder]]></category>
		<category><![CDATA[GPIA]]></category>
		<category><![CDATA[graph AI model hijacking]]></category>
		<category><![CDATA[Graph neural network security]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[graph prompt learning]]></category>
		<category><![CDATA[graph security]]></category>
		<category><![CDATA[integrity of recommendation systems]]></category>
		<category><![CDATA[poisoning attacks on AI models]]></category>
		<category><![CDATA[prompt injection attack]]></category>
		<category><![CDATA[prompt injection vulnerabilities]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[scientific data mining security]]></category>
		<category><![CDATA[threat of model manipulation in AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202204</guid>

					<description><![CDATA[Researchers have unveiled a data-level prompt injection attack that hijacks graph prompt learning systems by injecting a tiny fraction of malicious samples into downstream training data, achieving over 95 percent attack success while leaving the pretrained model untouched.]]></description>
										<content:encoded><![CDATA[<p>Graph neural networks have quietly become the workhorses of some of the most security-sensitive corners of modern computing. Banks use them to spot fraudulent transaction networks, e-commerce platforms rely on them to keep recommendation systems honest, and researchers mine scientific knowledge from vast webs of linked data. Because these models are expensive to train, a new efficiency trick has swept through the field: graph prompt learning, a technique that keeps a large pretrained graph encoder frozen and adapts it to new tasks by tuning only a tiny set of prompt parameters. It is fast, cheap and remarkably effective. But according to a new study published in the journal Cybersecurity, that very efficiency may be hiding a dangerous weakness, one that lets an attacker hijack a model&#8217;s behavior without ever touching the model itself.</p>
<p>Researchers at Henan University of Science and Technology, led by Mengying Yuan and Zhiyong Zhang, describe a new class of threat they call a data-level prompt injection attack. Unlike conventional backdoor attacks against graph neural networks, which typically require poisoning the pretraining pipeline or tampering with model weights, the new attack operates entirely at the downstream adaptation stage. The attacker simply slips a small number of carefully crafted malicious graphs into the labeled dataset that a downstream user collects to tune their prompts. The pretrained encoder stays pristine, the training algorithm stays untouched, and yet the learned prompts quietly absorb an attacker-specified rule that lies dormant until the right structural pattern appears in an input graph.</p>
<p>The distinction matters because of how graph prompt learning actually works in practice. In this paradigm, prompts are not natural-language instructions of the kind familiar from large language models. Instead, they are learnable graph structures or small parameter vectors that modulate the representations produced by a frozen encoder. Because the encoder is fixed, all of the task-specific knowledge a downstream user adds ends up concentrated in those few prompt parameters. The authors argue that this concentration creates a new and previously underexplored attack surface: whatever patterns appear repeatedly in the downstream training data get amplified directly into the prompts, with no encoder capacity available to absorb or dilute the signal.</p>
<p>The attack itself, which the team calls Graph Prompt Injection Attack, or GPIA, unfolds in three stages. First, the attacker designs a compact prompt-conditioning subgraph, a small structural motif optimized offline against the frozen encoder so that graphs carrying it drift toward a chosen target representation while otherwise staying close to their clean semantics. The optimization balances two goals: pulling the injected graph&#8217;s embedding toward the target class and preserving its original behavior, controlled by a trade-off parameter. Second, the attacker attaches this subgraph to training graphs at a deliberately chosen anchor point, favoring low-degree nodes so the perturbation stays localized and hard to spot in the overall topology. Third, every injected sample is labeled with the same target label, ensuring the prompt learner receives a consistent supervision signal that ties the conditioning pattern to the attacker&#8217;s chosen prediction.</p>
<p>There is a subtle twist that makes the attack especially stealthy. Rather than arbitrarily assigning a target label, the attacker first feeds the standalone conditioning subgraph through the frozen encoder along with the initialized prompt and observes what the model naturally predicts for it. That natural prediction becomes the target label. By aligning the injected pattern with the model&#8217;s own latent classification, the attacker minimizes semantic inconsistency in the representation space, making the poisoned samples look plausible rather than anomalous. The result is a trigger that blends into the data manifold instead of standing out as an outlier.</p>
<p>Why does such a small manipulation work so well? The authors offer a mathematical explanation rooted in optimization dynamics. When a fraction of training samples is poisoned, the gradient of the loss is a weighted mix of clean and injected contributions. Clean graphs are semantically diverse, so their gradients point in many different directions and largely cancel each other out. Injected samples, by contrast, all share the same conditioning pattern and the same objective, so their gradients are highly aligned and accumulate across training steps. Meanwhile, prompt learning confines optimization to a low-dimensional parameter space, far smaller than the full model, which makes those aligned gradients easier to reinforce. The frozen encoder compounds the problem: because no encoder parameters change during adaptation, malicious signals cannot be redistributed across the network and instead act repeatedly on the prompts alone. Even a tiny poisoned fraction can therefore exert a disproportionate influence on what the prompts learn.</p>
<p>The experimental evidence is striking. Testing on five standard benchmarks, including the citation networks Cora, CiteSeer and PubMed and the e-commerce co-purchase graphs Amazon-Computers and Amazon-Photo, the researchers evaluated GPIA against adapted versions of established graph backdoor attacks such as GCBA, UGBA and CrossBA, using a frozen GAT encoder and three representative prompt frameworks: GraphPrompt, ProG and ProG-Meta. With just five percent of training samples replaced by malicious graphs, GPIA achieved attack success rates consistently above 95 percent on most datasets, while baseline attacks sometimes fell below 80 percent. Crucially, accuracy on clean inputs barely moved, degrading by no more than about three percentage points, whereas several baselines caused accuracy drops exceeding 40 percent. Small standard deviations across five independent runs confirmed the attack is stable and reproducible.</p>
<p>The attack also refuses to stay confined to its training conditions. Under feature-level distribution shifts, including additive Gaussian noise and feature scaling, attack success declined only modestly, with no abrupt collapse. In cross-dataset experiments, prompts trained on one citation network transferred their malicious behavior to another, and in cross-domain tests the conditioning pattern carried over from citation graphs to e-commerce graphs despite radically different semantics and feature spaces. Swapping the GAT encoder for a GCN left the results essentially unchanged, indicating that the vulnerability stems from the prompt adaptation mechanism itself rather than any particular architecture. An ablation study showed that the target label consistency constraint was the single most important ingredient, followed by structural optimization of the conditioning subgraph, while anchor selection played a supporting role in stability and concealment.</p>
<p>Perhaps most concerning is how little it takes. When the researchers varied the injection ratio, they found that a mere one percent of poisoned samples was enough to push attack success rates above 80 percent across all four datasets tested, with near-perfect success at five percent and saturation beyond that point. The team also probed the attack against two representative poisoning defenses, Spectral Signatures and Confident Learning, which filter out the most suspicious samples before retraining. Both defenses provided only limited mitigation, with attack success remaining high after filtering. An analysis of embedding distributions showed why: malicious samples stay close to the clean data manifold and overlap heavily with benign representations, leaving no conspicuous outliers for detectors to flag.</p>
<p>The findings carry an urgent message for anyone deploying graph prompt learning in production. Because downstream training data in real-world settings is often assembled from public repositories, crowdsourced annotations and third-party platforms, attackers have realistic entry points that require no access to the model, the pretraining pipeline or the training procedure. The authors suggest several defensive directions, including rigorous inspection of training graphs for abnormally repeated structural motifs, robust prompt learning mechanisms that use structural perturbations or stochastic masking to prevent any fixed subgraph from coupling too tightly with prompt behavior, and data sanitization that down-weights samples containing rare or artificially repeated subgraph structures. They also note open questions: attack effectiveness is expected to decline somewhat under heavier supervision, and clean-label variants, where attackers cannot control the labels of injected samples, remain an important challenge for future research. What is already clear, however, is that the efficiency that makes graph prompts so attractive also makes them exquisitely sensitive to the data they learn from, and securing that data is no longer optional.</p>
<p><strong>Subject of Research:</strong> A novel data-level prompt injection attack, GPIA, that exploits the sensitivity of graph prompt learning to small amounts of poisoned downstream training data.</p>
<p><strong>Article Title:</strong> A novel data-level prompt injection attack against graph prompt learning</p>
<p><strong>Article References:</strong> Yuan, M., Zhang, Z., Quan, G., Pan, J., &amp; Fu, Y. (2026). A novel data-level prompt injection attack against graph prompt learning. <em>Cybersecurity, 9</em>(1), Article 220. <a href="https://doi.org/10.1186/s42400-026-00650-y" rel="noopener noreferrer">https://doi.org/10.1186/s42400-026-00650-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s42400-026-00650-y" rel="noopener noreferrer">10.1186/s42400-026-00650-y</a></p>
<p><strong>Keywords:</strong> graph prompt learning, graph neural networks, prompt injection attack, data poisoning, backdoor attacks, cybersecurity, adversarial machine learning, GPIA, frozen encoder, few-shot learning, graph security, prompt tuning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202204</post-id>	</item>
	</channel>
</rss>
