<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>prompt tuning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/prompt-tuning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 02:16:26 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>prompt tuning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Spot Factory Defects From Just Four Example Images</title>
		<link>https://scienmag.com/ai-learns-to-spot-factory-defects-from-just-four-example-images/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 02:16:26 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI for manufacturing defect recognition]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[automated quality control]]></category>
		<category><![CDATA[CLIP]]></category>
		<category><![CDATA[CLIP model applications]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[computer vision in production lines]]></category>
		<category><![CDATA[defect localization]]></category>
		<category><![CDATA[factory defect detection]]></category>
		<category><![CDATA[Few-shot learning]]></category>
		<category><![CDATA[industrial inspection]]></category>
		<category><![CDATA[innovative approaches to defect detection]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in quality assurance]]></category>
		<category><![CDATA[manufacturing]]></category>
		<category><![CDATA[minimal training data for visual inspection]]></category>
		<category><![CDATA[MVTec AD]]></category>
		<category><![CDATA[prompt engineering]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[reducing reliance on labeled defect images]]></category>
		<category><![CDATA[VisA]]></category>
		<category><![CDATA[vision-language model]]></category>
		<category><![CDATA[vision-language models in manufacturing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233062</guid>

					<description><![CDATA[Researchers at Yonsei University have developed a prompt-tuning method that lets the vision-language model CLIP detect and localize industrial defects from as few as four example images, achieving 95.6 percent AUROC on MVTec-AD with a 92 percent inference speed improvement over prior methods.]]></description>
										<content:encoded><![CDATA[<p>Few ideas in modern manufacturing are as deceptively simple, and as stubbornly difficult, as teaching a machine to recognize when something looks wrong. A production line may turn out thousands of nearly identical parts, and among them, almost imperceptibly, a scratched surface, a bent pin, a missing screw. Human inspectors catch many of these flaws, but fatigue and repetition erode their accuracy. Automated systems catch many more, yet the most accurate of them have traditionally demanded something factories rarely have: enormous collections of labeled defective images for every single product type. A new study from Yonsei University in Seoul, published in Applied Intelligence, describes a way to sidestep that requirement almost entirely, letting a vision-language model learn to detect and localize defects from as few as four example images per category.</p>
<p>The research, conducted by Jiwoo Choi and Chang Ouk Kim of Yonsei University&#8217;s Department of Industrial Engineering, builds on one of the most consequential developments in computer vision of the past several years: the rise of vision-language models, or VLMs. These models, chief among them CLIP, developed by OpenAI researchers, are trained on vast quantities of image-text pairs harvested from the web. Rather than learning only to classify images into fixed categories, they learn a shared space in which pictures and sentences live side by side. That structure gives them a remarkable ability to generalize. Show CLIP an image of an object it has never formally been taught to recognize, describe that object in plain English, and the model can often match the two without any task-specific training at all, a capability known as zero-shot transfer.</p>
<p>For anomaly detection, this generality is exactly what factories need. Traditional industrial anomaly detectors typically rely on convolutional neural networks pretrained on ImageNet, extracting features and flagging anything that deviates from a learned notion of normality. These systems work well when they can be trained on hundreds or thousands of images of a specific product, but they adapt poorly when the product line changes. A VLM-based approach promises something different: the ability to describe what a defect looks like in language, for instance, a photo of a flawless bottle cap paired with the phrase a photo of a damaged bottle cap, and let the model&#8217;s learned alignment between vision and language do the rest. The catch is that this alignment depends on how the description is written.</p>
<p>That dependency is the problem of prompt engineering. In CLIP-style systems, class labels are not fed to the model as bare words; they are wrapped in natural-language templates, such as a photo of a [category], which the text encoder converts into embeddings that guide the image encoder&#8217;s interpretation. Writing these templates by hand is something of a dark art. Small changes in wording can shift performance measurably, and crafting prompts that work for a specific industrial domain, with its particular vocabulary of defects and materials, requires knowledge that quality engineers may not have. Worse, hand-written prompts are static. A factory floor is a dynamic environment: lighting changes, new defect modes appear, product designs are revised. A prompt tuned for yesterday&#8217;s conditions may quietly degrade tomorrow.</p>
<p>Choi and Kim&#8217;s answer is to stop writing prompts by hand and instead let the model learn them. Their method, a few-shot anomaly prompt tuning approach, treats the prompt itself as a set of trainable parameters. Rather than fixed words, the prompt contains continuous embedding vectors that are optimized through gradient descent on a small handful of labeled examples, in this case as few as four images per category. This idea descends from a broader line of research on prompt tuning for vision-language models, which showed that optimizing soft prompts in the model&#8217;s embedding space can match or exceed the performance of painstakingly engineered text prompts. Applied to anomaly detection, the technique lets the model discover, from a tiny dataset, exactly how the language side of the system should describe normal and defective states for the product at hand.</p>
<p>But learning from four examples creates its own hazard, one familiar to anyone who has trained machine learning models on scarce data: overconfidence. When an anomaly map is generated from so little supervision, individual regions of an image can receive spuriously high anomaly scores, producing false positives, alarms raised on perfectly good parts. In a manufacturing setting, false positives are not merely an annoyance. Every false alarm triggers inspection, rework, or a halted line, and a detector that cries wolf too often will be ignored or switched off. The Yonsei team addresses this with a second contribution, a technique they call top-k anomaly score ensemble, or TASE. Instead of trusting a single model&#8217;s peak anomaly score, the method aggregates the top-scoring regions across multiple predictions, an ensembling strategy that smooths out the idiosyncratic errors any single pass can produce and tempers the model&#8217;s tendency to overreact to borderline cases.</p>
<p>The results, reported on the two most widely used benchmarks in industrial anomaly detection, are striking. On MVTec-AD, a comprehensive real-world dataset of manufacturing defects ranging from scratched metal to contaminated tiles, the proposed model achieves an area under the receiver operating characteristic curve, or AUROC, of 95.6 percent under a 4-shot setting, meaning it saw only four example images per product category before being tested. On VisA, a harder and more varied dataset of visual anomalies, it reaches 88.5 percent. AUROC is a standard measure of a detector&#8217;s ability to distinguish defective from normal samples across all possible decision thresholds, and scores in this range, achieved with so little training data, place the method among the competitive few-shot approaches in the field.</p>
<p>Efficiency is the study&#8217;s other headline claim. Adapting large pretrained models to new tasks is often computationally punishing, and prior prompt-based anomaly detection methods have carried significant inference-time costs. The proposed approach delivers a 92 percent improvement in inference speed over a prior method, a difference that matters enormously in practice. A production line inspects parts continuously, and a detector that runs slowly either inspects a sample of the output or forces expensive hardware investments. A detector that runs quickly can, in principle, examine every part in real time. The speed gain, combined with the minimal data requirement, points toward a deployment scenario in which a factory can adapt the system to a new product in minutes rather than days, using little more than a smartphone photo of a good part and a few defective ones.</p>
<p>The broader significance of the work lies in what it says about how specialized AI is being built. The old paradigm, collect a massive labeled dataset, train a bespoke model, deploy it, is giving way to a new one: take a general-purpose foundation model trained on the open internet, and nudge it toward a narrow domain with a handful of examples and a small number of trainable parameters. Prompt tuning sits at the heart of this shift because it avoids updating the enormous backbone of the model itself. Only the lightweight prompt vectors change, which keeps training cheap, reduces the risk of catastrophic forgetting of the model&#8217;s general knowledge, and makes adaptation feasible on modest hardware. The Yonsei study demonstrates that this recipe extends beyond classification and captioning into the quality-control trenches of manufacturing.</p>
<p>There are, of course, limits that the authors&#8217; benchmarks do not erase. Four shots is a small but not negligible amount of supervision, and someone must still identify and label the defective examples. The reported AUROC figures, while strong, leave room for errors on subtle or novel defect types that differ from anything in the few-shot examples. And like all CLIP-based systems, the method inherits whatever blind spots and biases the underlying foundation model acquired during its internet-scale training. Still, the trajectory is clear. As vision-language models grow more capable, techniques like learned prompts and score ensembling are turning them from impressive generalists into practical specialists, ones that can walk onto a factory floor, look at four pictures, and start catching the defects that slip past human eyes. The work was supported by the National Research Foundation of Korea, and its code of results joins a fast-growing literature, including WinCLIP and related methods, that is rapidly redefining what is possible when inspection data is scarce and time is short.</p>
<p><strong>Subject of Research:</strong> Few-shot visual anomaly detection and localization in manufacturing using prompt tuning of the CLIP vision-language model</p>
<p><strong>Article Title:</strong> Efficient few-shot visual anomaly detection and localization via prompt tuning</p>
<p><strong>Article References:</strong> Choi, J., &amp; Kim, C. O. (2026). Efficient few-shot visual anomaly detection and localization via prompt tuning. <em>Applied Intelligence, 56</em>(15), Article 450. <a href="https://doi.org/10.1007/s10489-026-07497-3" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07497-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07497-3" rel="noopener noreferrer">10.1007/s10489-026-07497-3</a></p>
<p><strong>Keywords:</strong> anomaly detection, prompt tuning, prompt engineering, vision-language model, CLIP, few-shot learning, industrial inspection, manufacturing, computer vision, MVTec-AD, VisA, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233062</post-id>	</item>
		<item>
		<title>Hidden Switches: New Attack Plants Undetectable Backdoors in Vision Transformers</title>
		<link>https://scienmag.com/hidden-switches-new-attack-plants-undetectable-backdoors-in-vision-transformers/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 23:11:11 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial attacks on computer vision models]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[AI safety]]></category>
		<category><![CDATA[backdoor attack]]></category>
		<category><![CDATA[backdoor attack in machine learning]]></category>
		<category><![CDATA[backdoor detection in vision transformers]]></category>
		<category><![CDATA[covert backdoor triggers in deep learning]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[cybersecurity threats in AI]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning model robustness]]></category>
		<category><![CDATA[hidden switches in neural networks]]></category>
		<category><![CDATA[high-stakes AI application vulnerabilities]]></category>
		<category><![CDATA[image classification]]></category>
		<category><![CDATA[machine learning security]]></category>
		<category><![CDATA[medical imaging model security]]></category>
		<category><![CDATA[model integrity in autonomous systems]]></category>
		<category><![CDATA[model poisoning]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[self-attention mechanism exploitation]]></category>
		<category><![CDATA[supply chain security]]></category>
		<category><![CDATA[vision transformer security vulnerabilities]]></category>
		<category><![CDATA[Vision Transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=219986</guid>

					<description><![CDATA[Researchers have demonstrated a switchable backdoor attack that injects dual tokens through the layers of Vision Transformers, achieving up to 99 percent attack success while leaving clean accuracy intact.]]></description>
										<content:encoded><![CDATA[<p>Vision Transformers have quietly become the backbone of modern computer vision, powering everything from image classifiers and object detectors to medical imaging pipelines and autonomous driving prototypes. Their self-attention mechanism, which lets the model weigh relationships between every patch of an image, has displaced convolutional networks in many high-stakes applications. But a new study from researchers at the University of Information Technology, Ho Chi Minh City, and Vietnam National University Ho Chi Minh City suggests that the very architecture making these models so powerful may also make them dangerously easy to compromise. In a paper published in the International Journal of Machine Learning and Cybernetics, Dung Minh Do and Khang Nguyen describe a backdoor attack that achieves attack success rates of up to 99 percent while leaving the model&#8217;s performance on clean, unmodified images essentially untouched.</p>
<p>Backdoor attacks are among the most insidious threats in machine learning security. Unlike adversarial examples, which exploit a model at inference time by perturbing individual inputs, a backdoor is baked into the model itself during training or fine-tuning. A model with a backdoor behaves perfectly normally on ordinary data, matching the accuracy and reliability its developers expect. But when a specific trigger, a subtle pattern chosen by the attacker, appears in an input, the model&#8217;s behavior flips, producing whatever output the attacker has designated. For a deployed system, this means an adversary could, for example, cause a traffic sign classifier to misread a stop sign or a security screening system to wave through a prohibited item, all without any visible degradation in everyday performance that might raise alarms.</p>
<p>What makes the new attack notable is where it operates. Rather than modifying the weights of the transformer or poisoning the training labels in conventional ways, the method injects two specially designed tokens into the model&#8217;s token stream, progressively working from the shallowest layers to the deepest ones. Vision Transformers process images by splitting them into patches, embedding each patch as a token, and then passing those tokens through a stack of transformer blocks in which self-attention layers let tokens exchange information. By inserting crafted tokens at multiple depths, the attack ensures that the malicious signal is reinforced and refined as it travels through the network, rather than being diluted or overwritten by the model&#8217;s normal processing.</p>
<p>The progressive, layer-by-layer nature of the injection is central to the attack&#8217;s stealthiness. Defenses that inspect a single layer, or that look for anomalous attention patterns at one point in the network, can miss a signal that is distributed across the entire depth of the model. Because the injected tokens interact with the image tokens through the standard attention mechanism, they can steer the model&#8217;s internal representations toward the attacker&#8217;s chosen target class only when the trigger is present. On clean inputs, the tokens remain effectively inert, which is why the authors report that the method maintains competitive clean accuracy even as it delivers near-perfect attack success rates across multiple visual datasets.</p>
<p>The word switchable in the study&#8217;s title points to another dimension of the threat. The attack draws on a growing line of research into switchable backdoors against pre-trained vision transformers, in which a single compromised model can carry multiple behaviors that an attacker can toggle. This means a defender who discovers one trigger and filters it out cannot assume the model is safe; another trigger may remain dormant, waiting to be activated. The authors position their work as an analysis and exploitation of ViT vulnerabilities, and the breadth of prior work they survey, from BadNets and weight-poisoning attacks to attention hijacking and prompt-based backdoors, underscores how rapidly this attack surface has expanded as transformers have spread through computer vision.</p>
<p>The connection to prompt tuning is particularly significant for the current state of the field. Prompt-based methods, in which small learnable tokens are prepended to a model&#8217;s input or inserted into its layers, have become a popular way to adapt large pre-trained models to new tasks without expensive full fine-tuning. Visual prompt tuning and related techniques are widely used because they are efficient and effective. But the same mechanism that makes prompts useful, namely the ability to inject learned tokens that influence the model&#8217;s attention and representations, is exactly what this attack weaponizes. A malicious actor with access to a fine-tuning pipeline, or able to distribute a poisoned adapter or checkpoint, could embed a backdoor that looks indistinguishable from a legitimate prompt-based adaptation.</p>
<p>The supply chain implications are sobering. Modern machine learning practice relies heavily on pre-trained models downloaded from public hubs, fine-tuned adapters shared between teams, and third-party datasets. Each of these channels is a potential delivery mechanism for a backdoor. Earlier research has shown that weight poisoning attacks on pre-trained models can survive downstream fine-tuning, and that backdoors can be hidden in ways that evade standard inspection. The new study adds to this picture by showing that the token-based machinery of vision transformers offers attackers a particularly clean injection point, one that does not require the crude modifications of earlier attacks that made them easier to detect.</p>
<p>The authors evaluated their method on multiple visual datasets, measuring not only attack success rate and clean accuracy but also stealthiness and robustness. The headline figure, an attack success rate of up to 99 percent, is alarming enough, but the more troubling result is the combination of that success with preserved clean performance, since it means conventional accuracy-based validation would reveal nothing amiss. The paper also situates the work against existing defenses, which include fine-pruning approaches that remove rarely activated neurons, neural attention distillation that tries to erase trigger-related attention patterns, and input-level detection methods that look for inconsistencies in a model&#8217;s predictions under image transformations. Because the new attack distributes its signal across shallow and deep layers through dual tokens, many of these defenses, which were designed with convolutional networks or single-point injections in mind, face a harder problem.</p>
<p>The researchers are explicit about the warning their findings carry for real-world deployment. Vision Transformers are increasingly used in settings where a silent failure mode could have serious consequences, including medical diagnosis support, surveillance, industrial inspection, and safety-critical perception systems. A backdoored model in any of these contexts could be remotely triggered by an input crafted to contain the attacker&#8217;s pattern, and the compromise would be invisible in routine testing. The authors make their code publicly available, which serves the defensive side of the field as well: reproducible attack implementations are essential for developing and benchmarking countermeasures, and the history of adversarial machine learning shows that security research advances fastest when attacks are fully documented.</p>
<p>For the broader community, the study is a reminder that architectural progress and security progress have been badly out of step. The references in the paper trace a decade of deep learning breakthroughs, from early convolutional networks through EfficientNet and the original Vision Transformer, alongside a parallel literature of backdoor learning surveys, prompt injection analyses, and defense proposals. Yet the authors note that security threats to ViTs, particularly backdoor attacks, have not received research attention commensurate with the architecture&#8217;s adoption. Closing that gap will likely require defenses designed specifically for token-based architectures: methods that audit injected tokens, verify the provenance of fine-tuned checkpoints, test models against a family of triggers rather than a single known pattern, and treat the entire depth of the network, not just its input layer, as a potential attack surface. Until such defenses mature, the near-perfect stealth and effectiveness demonstrated by progressive dual-token injection stands as a stark warning that the models powering tomorrow&#8217;s vision systems may harbor switches that only their attackers know how to flip.</p>
<p><strong>Subject of Research:</strong> Backdoor attack vulnerabilities in Vision Transformers via progressive dual-token injection</p>
<p><strong>Article Title:</strong> Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers</p>
<p><strong>Article References:</strong> Do, D. M., &amp; Nguyen, K. (2026). Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 489. <a href="https://doi.org/10.1007/s13042-026-03326-8" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03326-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03326-8" rel="noopener noreferrer">10.1007/s13042-026-03326-8</a></p>
<p><strong>Keywords:</strong> Vision Transformers, backdoor attack, machine learning security, self-attention, prompt tuning, adversarial machine learning, model poisoning, supply chain security, deep learning, cybersecurity, image classification, AI safety</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">219986</post-id>	</item>
		<item>
		<title>AI Learns to Distrust Its Own Illusions: CLIP Helps Models Adapt Without Source Data</title>
		<link>https://scienmag.com/ai-learns-to-distrust-its-own-illusions-clip-helps-models-adapt-without-source-data/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 22:50:04 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[CLIP]]></category>
		<category><![CDATA[confirmation bias]]></category>
		<category><![CDATA[distribution alignment]]></category>
		<category><![CDATA[domain shift]]></category>
		<category><![CDATA[knowledge distillation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[pseudo-source domain]]></category>
		<category><![CDATA[source-free domain adaptation]]></category>
		<category><![CDATA[spurious correlations]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[vision-language models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205183</guid>

					<description><![CDATA[Researchers have developed SRRI, a source-free domain adaptation method that uses CLIP as an external verifier to de-bias pseudo-source domains and outperform state-of-the-art approaches.]]></description>
										<content:encoded><![CDATA[<p>Machine learning models are often trained in one setting and deployed in another, and that transition is rarely seamless. A classifier trained on studio photographs of cars may stumble when shown cars in rain, at night, or through a security camera. Domain adaptation is the branch of machine learning that tackles this problem, and its most demanding variant, source-free domain adaptation, adds a harsh constraint: once the model leaves the training environment, the original labeled data is gone, often for reasons of privacy, storage, or proprietary restriction. All the model can carry with it is what it learned. A new study published in the journal Machine Learning by Qing Tian, Yongjiang Liu, Keyang Cheng, Weihua Ou, Jianping Gou and colleagues confronts a subtle but pervasive failure mode in this setting, one the researchers describe with an evocative word borrowed from human cognition: illusions.</p>
<p>The core difficulty lies in how a source model, cut off from its training data, tries to make sense of an unlabeled target domain. A popular family of methods constructs a pseudo-source domain, a synthetic stand-in for the lost source data, by selecting target samples that the source model believes look source-like. Training on this pseudo-source then reduces the statistical gap between domains. The catch, as the new paper argues, is that this construction leans entirely on the source model itself, and the source model is precisely the entity most likely to be deceived. During training, it absorbed not only the true signals that define each class, such as the shape and structure of an object, but also spurious correlations linking class identity to domain-specific features like background, lighting, or environment. When asked to identify pseudo-source samples, it may confirm its own biases, mistaking samples that share superficial environmental cues with the source domain for genuinely representative ones.</p>
<p>The researchers identify two intertwined challenges that undermine this process. The first is confirmation bias: once the model commits to a belief about a sample, the subsequent training on that sample reinforces the belief, whether or not it was correct. This is the machine analogue of a person who reads only news that agrees with their views. The second is domain shift, the distributional mismatch between source and target that makes the source model&#8217;s judgments unreliable in the first place. Together, these effects can contaminate the pseudo-source domain with mislabeled or unrepresentative samples, and the errors compound as adaptation proceeds. The team&#8217;s answer is a method they call Staying Rational and Resisting Illusions, or SRRI, which refuses to let the source model grade its own homework.</p>
<p>The key innovation is the introduction of an external referee: CLIP, the Contrastive Language-Image Pre-training model developed by OpenAI researchers, which learned to align images and text by training on hundreds of millions of image-text pairs from the web. Because CLIP&#8217;s knowledge comes from a vastly broader data distribution than the narrow source domain, it does not share the source model&#8217;s spurious correlations. SRRI uses CLIP as a source of independent evidence when deciding which target samples deserve a place in the pseudo-source domain. In effect, when the source model says a sample looks familiar, SRRI asks CLIP for a second opinion before accepting the claim.</p>
<p>Technically, the method proceeds in several coordinated stages. First, SRRI employs knowledge distillation, a technique in which a teacher network&#8217;s outputs guide a student network, to help the source model disentangle class-discriminative causal features from domain-specific spurious features. The goal is to teach the model which aspects of an image actually cause its label, such as the geometry of an object, and which merely co-occur with it, such as the typical backdrop of the source photographs. This disentanglement weakens the illusions at their root, making the model&#8217;s own judgments less confounded before any pseudo-source construction begins.</p>
<p>Distillation alone, however, cannot be trusted blindly, because in some adaptation tasks the distillation process itself performs poorly, propagating errors rather than correcting them. To guard against this, the authors design a CLIP-guided dual-model validation and class balancing strategy. Every candidate pseudo-source sample must pass inspection by both the distilled source model and CLIP, and the two models&#8217; assessments are combined to filter out unreliable examples. Class balancing ensures that the retained samples cover all categories with reasonable richness, preventing the pseudo-source domain from being dominated by easy or overrepresented classes. This dual gatekeeping is what allows the method to remain robust even when one of its components falters on a given task.</p>
<p>The third pillar of SRRI is a dynamic pseudo-source domain optimization mechanism. Rather than freezing the pseudo-source once it is built, the method continuously fine-tunes the task-specific prompts of CLIP during adaptation. Prompt tuning adjusts the short text descriptions that CLIP uses to interpret images, sharpening its sensitivity to the specific categories of the target task. As these prompts improve, CLIP&#8217;s judgments on hard samples become more accurate, which in turn corrects residual bias in the target model. At the same time, the pseudo-source domain is periodically reconstructed and refined, discarding samples that no longer pass validation and admitting better ones as the models evolve. The result is a self-correcting loop in which the reference data improves alongside the adapting model.</p>
<p>With a trustworthy pseudo-source domain in place, SRRI applies robust supervised learning to train the target model on the pseudo-source samples, while simultaneously performing distribution alignment between the pseudo-source and the true target data. This alignment ensures that the model does not overfit to artifacts of the pseudo-source construction and that its decision boundaries remain well matched to the actual deployment distribution. The combination of reliable pseudo-labels, balanced classes, and distributional consistency addresses both of the fundamental challenges the authors set out to solve: confirmation bias is curbed by external validation, and domain shift is absorbed by the alignment procedure.</p>
<p>Extensive experiments reported in the paper show that SRRI outperforms state-of-the-art source-free domain adaptation methods across standard benchmarks. The improvements are attributed not to any single trick but to the architecture of trust the method builds: an independent verifier, a disentangled representation, a balanced and evolving reference set, and a training objective that keeps the target model anchored to reality. All datasets used in the study are publicly available, which should make the approach straightforward for other groups to reproduce and extend. The work was supported by the National Natural Science Foundation of China and several regional research programs, and the authors report no competing financial interests beyond these funding sources.</p>
<p>The broader significance of this research extends beyond a single benchmark. As artificial intelligence systems are increasingly deployed in hospitals, vehicles, and surveillance networks where raw training data cannot be shared, source-free adaptation will become a standard requirement rather than a niche concern. The lesson of SRRI is that a model adapting in the wild should not rely solely on its own inherited judgments, because those judgments may encode illusions about what really defines a category. By recruiting a vision-language foundation model as an external rational check, and by continuously refining both the verifier and the verified, the researchers offer a template for building machine learning systems that stay rational under pressure, resisting the very biases they were born with. In an era when AI is often criticized for confidently repeating its mistakes, a method explicitly designed to resist its own illusions is a welcome step toward more trustworthy machine intelligence.</p>
<p><strong>Subject of Research:</strong> De-biasing source-free domain adaptation using a CLIP-verified pseudo-source domain</p>
<p><strong>Article Title:</strong> Stay Rational, Resist Illusions: De-biasing Source-Free Domain Adaptation with CLIP-Verified Pseudo-Source Domain</p>
<p><strong>Article References:</strong> Tian, Q., Liu, Y., Cheng, K., Ou, W., &amp; Gou, J. (2026). Stay Rational, Resist Illusions: De-biasing Source-Free Domain Adaptation with CLIP-Verified Pseudo-Source Domain. <em>Machine Learning, 115</em>(10), Article 222. <a href="https://doi.org/10.1007/s10994-026-07163-2" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07163-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07163-2" rel="noopener noreferrer">10.1007/s10994-026-07163-2</a></p>
<p><strong>Keywords:</strong> source-free domain adaptation, CLIP, pseudo-source domain, knowledge distillation, confirmation bias, domain shift, spurious correlations, prompt tuning, distribution alignment, machine learning, vision-language models, transfer learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205183</post-id>	</item>
		<item>
		<title>Hackers Can Hijack Graph AI With Just a Handful of Poisoned Samples</title>
		<link>https://scienmag.com/hackers-can-hijack-graph-ai-with-just-a-handful-of-poisoned-samples/</link>
		
		<dc:creator><![CDATA[Hailey Crawford]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:18:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial attacks on graph-based systems]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[backdoor attacks]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[cybersecurity risks in graph neural networks]]></category>
		<category><![CDATA[data poisoning]]></category>
		<category><![CDATA[data-level prompt injection]]></category>
		<category><![CDATA[defenses against prompt injection attacks]]></category>
		<category><![CDATA[efficient graph learning vulnerabilities]]></category>
		<category><![CDATA[Few-shot learning]]></category>
		<category><![CDATA[frozen encoder]]></category>
		<category><![CDATA[GPIA]]></category>
		<category><![CDATA[graph AI model hijacking]]></category>
		<category><![CDATA[Graph neural network security]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[graph prompt learning]]></category>
		<category><![CDATA[graph security]]></category>
		<category><![CDATA[integrity of recommendation systems]]></category>
		<category><![CDATA[poisoning attacks on AI models]]></category>
		<category><![CDATA[prompt injection attack]]></category>
		<category><![CDATA[prompt injection vulnerabilities]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[scientific data mining security]]></category>
		<category><![CDATA[threat of model manipulation in AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202204</guid>

					<description><![CDATA[Researchers have unveiled a data-level prompt injection attack that hijacks graph prompt learning systems by injecting a tiny fraction of malicious samples into downstream training data, achieving over 95 percent attack success while leaving the pretrained model untouched.]]></description>
										<content:encoded><![CDATA[<p>Graph neural networks have quietly become the workhorses of some of the most security-sensitive corners of modern computing. Banks use them to spot fraudulent transaction networks, e-commerce platforms rely on them to keep recommendation systems honest, and researchers mine scientific knowledge from vast webs of linked data. Because these models are expensive to train, a new efficiency trick has swept through the field: graph prompt learning, a technique that keeps a large pretrained graph encoder frozen and adapts it to new tasks by tuning only a tiny set of prompt parameters. It is fast, cheap and remarkably effective. But according to a new study published in the journal Cybersecurity, that very efficiency may be hiding a dangerous weakness, one that lets an attacker hijack a model&#8217;s behavior without ever touching the model itself.</p>
<p>Researchers at Henan University of Science and Technology, led by Mengying Yuan and Zhiyong Zhang, describe a new class of threat they call a data-level prompt injection attack. Unlike conventional backdoor attacks against graph neural networks, which typically require poisoning the pretraining pipeline or tampering with model weights, the new attack operates entirely at the downstream adaptation stage. The attacker simply slips a small number of carefully crafted malicious graphs into the labeled dataset that a downstream user collects to tune their prompts. The pretrained encoder stays pristine, the training algorithm stays untouched, and yet the learned prompts quietly absorb an attacker-specified rule that lies dormant until the right structural pattern appears in an input graph.</p>
<p>The distinction matters because of how graph prompt learning actually works in practice. In this paradigm, prompts are not natural-language instructions of the kind familiar from large language models. Instead, they are learnable graph structures or small parameter vectors that modulate the representations produced by a frozen encoder. Because the encoder is fixed, all of the task-specific knowledge a downstream user adds ends up concentrated in those few prompt parameters. The authors argue that this concentration creates a new and previously underexplored attack surface: whatever patterns appear repeatedly in the downstream training data get amplified directly into the prompts, with no encoder capacity available to absorb or dilute the signal.</p>
<p>The attack itself, which the team calls Graph Prompt Injection Attack, or GPIA, unfolds in three stages. First, the attacker designs a compact prompt-conditioning subgraph, a small structural motif optimized offline against the frozen encoder so that graphs carrying it drift toward a chosen target representation while otherwise staying close to their clean semantics. The optimization balances two goals: pulling the injected graph&#8217;s embedding toward the target class and preserving its original behavior, controlled by a trade-off parameter. Second, the attacker attaches this subgraph to training graphs at a deliberately chosen anchor point, favoring low-degree nodes so the perturbation stays localized and hard to spot in the overall topology. Third, every injected sample is labeled with the same target label, ensuring the prompt learner receives a consistent supervision signal that ties the conditioning pattern to the attacker&#8217;s chosen prediction.</p>
<p>There is a subtle twist that makes the attack especially stealthy. Rather than arbitrarily assigning a target label, the attacker first feeds the standalone conditioning subgraph through the frozen encoder along with the initialized prompt and observes what the model naturally predicts for it. That natural prediction becomes the target label. By aligning the injected pattern with the model&#8217;s own latent classification, the attacker minimizes semantic inconsistency in the representation space, making the poisoned samples look plausible rather than anomalous. The result is a trigger that blends into the data manifold instead of standing out as an outlier.</p>
<p>Why does such a small manipulation work so well? The authors offer a mathematical explanation rooted in optimization dynamics. When a fraction of training samples is poisoned, the gradient of the loss is a weighted mix of clean and injected contributions. Clean graphs are semantically diverse, so their gradients point in many different directions and largely cancel each other out. Injected samples, by contrast, all share the same conditioning pattern and the same objective, so their gradients are highly aligned and accumulate across training steps. Meanwhile, prompt learning confines optimization to a low-dimensional parameter space, far smaller than the full model, which makes those aligned gradients easier to reinforce. The frozen encoder compounds the problem: because no encoder parameters change during adaptation, malicious signals cannot be redistributed across the network and instead act repeatedly on the prompts alone. Even a tiny poisoned fraction can therefore exert a disproportionate influence on what the prompts learn.</p>
<p>The experimental evidence is striking. Testing on five standard benchmarks, including the citation networks Cora, CiteSeer and PubMed and the e-commerce co-purchase graphs Amazon-Computers and Amazon-Photo, the researchers evaluated GPIA against adapted versions of established graph backdoor attacks such as GCBA, UGBA and CrossBA, using a frozen GAT encoder and three representative prompt frameworks: GraphPrompt, ProG and ProG-Meta. With just five percent of training samples replaced by malicious graphs, GPIA achieved attack success rates consistently above 95 percent on most datasets, while baseline attacks sometimes fell below 80 percent. Crucially, accuracy on clean inputs barely moved, degrading by no more than about three percentage points, whereas several baselines caused accuracy drops exceeding 40 percent. Small standard deviations across five independent runs confirmed the attack is stable and reproducible.</p>
<p>The attack also refuses to stay confined to its training conditions. Under feature-level distribution shifts, including additive Gaussian noise and feature scaling, attack success declined only modestly, with no abrupt collapse. In cross-dataset experiments, prompts trained on one citation network transferred their malicious behavior to another, and in cross-domain tests the conditioning pattern carried over from citation graphs to e-commerce graphs despite radically different semantics and feature spaces. Swapping the GAT encoder for a GCN left the results essentially unchanged, indicating that the vulnerability stems from the prompt adaptation mechanism itself rather than any particular architecture. An ablation study showed that the target label consistency constraint was the single most important ingredient, followed by structural optimization of the conditioning subgraph, while anchor selection played a supporting role in stability and concealment.</p>
<p>Perhaps most concerning is how little it takes. When the researchers varied the injection ratio, they found that a mere one percent of poisoned samples was enough to push attack success rates above 80 percent across all four datasets tested, with near-perfect success at five percent and saturation beyond that point. The team also probed the attack against two representative poisoning defenses, Spectral Signatures and Confident Learning, which filter out the most suspicious samples before retraining. Both defenses provided only limited mitigation, with attack success remaining high after filtering. An analysis of embedding distributions showed why: malicious samples stay close to the clean data manifold and overlap heavily with benign representations, leaving no conspicuous outliers for detectors to flag.</p>
<p>The findings carry an urgent message for anyone deploying graph prompt learning in production. Because downstream training data in real-world settings is often assembled from public repositories, crowdsourced annotations and third-party platforms, attackers have realistic entry points that require no access to the model, the pretraining pipeline or the training procedure. The authors suggest several defensive directions, including rigorous inspection of training graphs for abnormally repeated structural motifs, robust prompt learning mechanisms that use structural perturbations or stochastic masking to prevent any fixed subgraph from coupling too tightly with prompt behavior, and data sanitization that down-weights samples containing rare or artificially repeated subgraph structures. They also note open questions: attack effectiveness is expected to decline somewhat under heavier supervision, and clean-label variants, where attackers cannot control the labels of injected samples, remain an important challenge for future research. What is already clear, however, is that the efficiency that makes graph prompts so attractive also makes them exquisitely sensitive to the data they learn from, and securing that data is no longer optional.</p>
<p><strong>Subject of Research:</strong> A novel data-level prompt injection attack, GPIA, that exploits the sensitivity of graph prompt learning to small amounts of poisoned downstream training data.</p>
<p><strong>Article Title:</strong> A novel data-level prompt injection attack against graph prompt learning</p>
<p><strong>Article References:</strong> Yuan, M., Zhang, Z., Quan, G., Pan, J., &amp; Fu, Y. (2026). A novel data-level prompt injection attack against graph prompt learning. <em>Cybersecurity, 9</em>(1), Article 220. <a href="https://doi.org/10.1186/s42400-026-00650-y" rel="noopener noreferrer">https://doi.org/10.1186/s42400-026-00650-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s42400-026-00650-y" rel="noopener noreferrer">10.1186/s42400-026-00650-y</a></p>
<p><strong>Keywords:</strong> graph prompt learning, graph neural networks, prompt injection attack, data poisoning, backdoor attacks, cybersecurity, adversarial machine learning, GPIA, frozen encoder, few-shot learning, graph security, prompt tuning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202204</post-id>	</item>
	</channel>
</rss>
