<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>minimal training data for visual inspection &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/minimal-training-data-for-visual-inspection/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 02:16:26 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>minimal training data for visual inspection &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Spot Factory Defects From Just Four Example Images</title>
		<link>https://scienmag.com/ai-learns-to-spot-factory-defects-from-just-four-example-images/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 02:16:26 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI for manufacturing defect recognition]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[automated quality control]]></category>
		<category><![CDATA[CLIP]]></category>
		<category><![CDATA[CLIP model applications]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[computer vision in production lines]]></category>
		<category><![CDATA[defect localization]]></category>
		<category><![CDATA[factory defect detection]]></category>
		<category><![CDATA[Few-shot learning]]></category>
		<category><![CDATA[industrial inspection]]></category>
		<category><![CDATA[innovative approaches to defect detection]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in quality assurance]]></category>
		<category><![CDATA[manufacturing]]></category>
		<category><![CDATA[minimal training data for visual inspection]]></category>
		<category><![CDATA[MVTec AD]]></category>
		<category><![CDATA[prompt engineering]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[reducing reliance on labeled defect images]]></category>
		<category><![CDATA[VisA]]></category>
		<category><![CDATA[vision-language model]]></category>
		<category><![CDATA[vision-language models in manufacturing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233062</guid>

					<description><![CDATA[Researchers at Yonsei University have developed a prompt-tuning method that lets the vision-language model CLIP detect and localize industrial defects from as few as four example images, achieving 95.6 percent AUROC on MVTec-AD with a 92 percent inference speed improvement over prior methods.]]></description>
										<content:encoded><![CDATA[<p>Few ideas in modern manufacturing are as deceptively simple, and as stubbornly difficult, as teaching a machine to recognize when something looks wrong. A production line may turn out thousands of nearly identical parts, and among them, almost imperceptibly, a scratched surface, a bent pin, a missing screw. Human inspectors catch many of these flaws, but fatigue and repetition erode their accuracy. Automated systems catch many more, yet the most accurate of them have traditionally demanded something factories rarely have: enormous collections of labeled defective images for every single product type. A new study from Yonsei University in Seoul, published in Applied Intelligence, describes a way to sidestep that requirement almost entirely, letting a vision-language model learn to detect and localize defects from as few as four example images per category.</p>
<p>The research, conducted by Jiwoo Choi and Chang Ouk Kim of Yonsei University&#8217;s Department of Industrial Engineering, builds on one of the most consequential developments in computer vision of the past several years: the rise of vision-language models, or VLMs. These models, chief among them CLIP, developed by OpenAI researchers, are trained on vast quantities of image-text pairs harvested from the web. Rather than learning only to classify images into fixed categories, they learn a shared space in which pictures and sentences live side by side. That structure gives them a remarkable ability to generalize. Show CLIP an image of an object it has never formally been taught to recognize, describe that object in plain English, and the model can often match the two without any task-specific training at all, a capability known as zero-shot transfer.</p>
<p>For anomaly detection, this generality is exactly what factories need. Traditional industrial anomaly detectors typically rely on convolutional neural networks pretrained on ImageNet, extracting features and flagging anything that deviates from a learned notion of normality. These systems work well when they can be trained on hundreds or thousands of images of a specific product, but they adapt poorly when the product line changes. A VLM-based approach promises something different: the ability to describe what a defect looks like in language, for instance, a photo of a flawless bottle cap paired with the phrase a photo of a damaged bottle cap, and let the model&#8217;s learned alignment between vision and language do the rest. The catch is that this alignment depends on how the description is written.</p>
<p>That dependency is the problem of prompt engineering. In CLIP-style systems, class labels are not fed to the model as bare words; they are wrapped in natural-language templates, such as a photo of a [category], which the text encoder converts into embeddings that guide the image encoder&#8217;s interpretation. Writing these templates by hand is something of a dark art. Small changes in wording can shift performance measurably, and crafting prompts that work for a specific industrial domain, with its particular vocabulary of defects and materials, requires knowledge that quality engineers may not have. Worse, hand-written prompts are static. A factory floor is a dynamic environment: lighting changes, new defect modes appear, product designs are revised. A prompt tuned for yesterday&#8217;s conditions may quietly degrade tomorrow.</p>
<p>Choi and Kim&#8217;s answer is to stop writing prompts by hand and instead let the model learn them. Their method, a few-shot anomaly prompt tuning approach, treats the prompt itself as a set of trainable parameters. Rather than fixed words, the prompt contains continuous embedding vectors that are optimized through gradient descent on a small handful of labeled examples, in this case as few as four images per category. This idea descends from a broader line of research on prompt tuning for vision-language models, which showed that optimizing soft prompts in the model&#8217;s embedding space can match or exceed the performance of painstakingly engineered text prompts. Applied to anomaly detection, the technique lets the model discover, from a tiny dataset, exactly how the language side of the system should describe normal and defective states for the product at hand.</p>
<p>But learning from four examples creates its own hazard, one familiar to anyone who has trained machine learning models on scarce data: overconfidence. When an anomaly map is generated from so little supervision, individual regions of an image can receive spuriously high anomaly scores, producing false positives, alarms raised on perfectly good parts. In a manufacturing setting, false positives are not merely an annoyance. Every false alarm triggers inspection, rework, or a halted line, and a detector that cries wolf too often will be ignored or switched off. The Yonsei team addresses this with a second contribution, a technique they call top-k anomaly score ensemble, or TASE. Instead of trusting a single model&#8217;s peak anomaly score, the method aggregates the top-scoring regions across multiple predictions, an ensembling strategy that smooths out the idiosyncratic errors any single pass can produce and tempers the model&#8217;s tendency to overreact to borderline cases.</p>
<p>The results, reported on the two most widely used benchmarks in industrial anomaly detection, are striking. On MVTec-AD, a comprehensive real-world dataset of manufacturing defects ranging from scratched metal to contaminated tiles, the proposed model achieves an area under the receiver operating characteristic curve, or AUROC, of 95.6 percent under a 4-shot setting, meaning it saw only four example images per product category before being tested. On VisA, a harder and more varied dataset of visual anomalies, it reaches 88.5 percent. AUROC is a standard measure of a detector&#8217;s ability to distinguish defective from normal samples across all possible decision thresholds, and scores in this range, achieved with so little training data, place the method among the competitive few-shot approaches in the field.</p>
<p>Efficiency is the study&#8217;s other headline claim. Adapting large pretrained models to new tasks is often computationally punishing, and prior prompt-based anomaly detection methods have carried significant inference-time costs. The proposed approach delivers a 92 percent improvement in inference speed over a prior method, a difference that matters enormously in practice. A production line inspects parts continuously, and a detector that runs slowly either inspects a sample of the output or forces expensive hardware investments. A detector that runs quickly can, in principle, examine every part in real time. The speed gain, combined with the minimal data requirement, points toward a deployment scenario in which a factory can adapt the system to a new product in minutes rather than days, using little more than a smartphone photo of a good part and a few defective ones.</p>
<p>The broader significance of the work lies in what it says about how specialized AI is being built. The old paradigm, collect a massive labeled dataset, train a bespoke model, deploy it, is giving way to a new one: take a general-purpose foundation model trained on the open internet, and nudge it toward a narrow domain with a handful of examples and a small number of trainable parameters. Prompt tuning sits at the heart of this shift because it avoids updating the enormous backbone of the model itself. Only the lightweight prompt vectors change, which keeps training cheap, reduces the risk of catastrophic forgetting of the model&#8217;s general knowledge, and makes adaptation feasible on modest hardware. The Yonsei study demonstrates that this recipe extends beyond classification and captioning into the quality-control trenches of manufacturing.</p>
<p>There are, of course, limits that the authors&#8217; benchmarks do not erase. Four shots is a small but not negligible amount of supervision, and someone must still identify and label the defective examples. The reported AUROC figures, while strong, leave room for errors on subtle or novel defect types that differ from anything in the few-shot examples. And like all CLIP-based systems, the method inherits whatever blind spots and biases the underlying foundation model acquired during its internet-scale training. Still, the trajectory is clear. As vision-language models grow more capable, techniques like learned prompts and score ensembling are turning them from impressive generalists into practical specialists, ones that can walk onto a factory floor, look at four pictures, and start catching the defects that slip past human eyes. The work was supported by the National Research Foundation of Korea, and its code of results joins a fast-growing literature, including WinCLIP and related methods, that is rapidly redefining what is possible when inspection data is scarce and time is short.</p>
<p><strong>Subject of Research:</strong> Few-shot visual anomaly detection and localization in manufacturing using prompt tuning of the CLIP vision-language model</p>
<p><strong>Article Title:</strong> Efficient few-shot visual anomaly detection and localization via prompt tuning</p>
<p><strong>Article References:</strong> Choi, J., &amp; Kim, C. O. (2026). Efficient few-shot visual anomaly detection and localization via prompt tuning. <em>Applied Intelligence, 56</em>(15), Article 450. <a href="https://doi.org/10.1007/s10489-026-07497-3" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07497-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07497-3" rel="noopener noreferrer">10.1007/s10489-026-07497-3</a></p>
<p><strong>Keywords:</strong> anomaly detection, prompt tuning, prompt engineering, vision-language model, CLIP, few-shot learning, industrial inspection, manufacturing, computer vision, MVTec-AD, VisA, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233062</post-id>	</item>
	</channel>
</rss>
