<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>domain shift &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/domain-shift/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 08:30:30 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>domain shift &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Diagnose Grapevine Diseases Like a Plant Pathologist</title>
		<link>https://scienmag.com/ai-learns-to-diagnose-grapevine-diseases-like-a-plant-pathologist/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 08:30:30 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI systems for foliar disease detection]]></category>
		<category><![CDATA[AI-powered grapevine disease identification]]></category>
		<category><![CDATA[artificial intelligence in plant pathology]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[cross-attention]]></category>
		<category><![CDATA[domain shift]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[explainable AI for agriculture]]></category>
		<category><![CDATA[grapevine disease]]></category>
		<category><![CDATA[grapevine disease detection]]></category>
		<category><![CDATA[hybrid framework for plant disease diagnosis]]></category>
		<category><![CDATA[improving accuracy in plant disease diagnosis]]></category>
		<category><![CDATA[LLaVA]]></category>
		<category><![CDATA[overcoming CNN limitations in plant health]]></category>
		<category><![CDATA[plant disease classification using deep learning]]></category>
		<category><![CDATA[plant pathology]]></category>
		<category><![CDATA[PlantVillage]]></category>
		<category><![CDATA[precision agriculture]]></category>
		<category><![CDATA[QLoRA]]></category>
		<category><![CDATA[real-time vineyard disease monitoring]]></category>
		<category><![CDATA[saliency auto-crop]]></category>
		<category><![CDATA[transparent AI models for farmers]]></category>
		<category><![CDATA[vision-language models]]></category>
		<category><![CDATA[vision-language models for crop health]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=246834</guid>

					<description><![CDATA[Researchers in Vietnam have built a hybrid vision-language AI that diagnoses grapevine diseases with 98.7 percent accuracy while explaining its reasoning and resisting the background overfitting that cripples conventional CNNs in real fields.]]></description>
										<content:encoded><![CDATA[<p>Grapevines are among the most economically valuable crops on Earth, and they are under constant assault from foliar diseases such as Black Rot, ESCA, and Leaf Blight. For generations, detecting these threats has depended on the trained eye of agricultural experts, a process that is slow, expensive, and vulnerable to human subjectivity. Now, a pair of researchers from the Data Science Laboratory at the Industrial University of Ho Chi Minh City in Vietnam has unveiled an artificial intelligence system that not only diagnoses grapevine diseases with remarkable accuracy but also explains its reasoning in plain language. The study, published in Discover Artificial Intelligence, describes a Global–Local Hybrid Framework built on a vision-language model that mimics how a human pathologist actually works: first surveying the whole leaf, then zooming in on the lesion with a magnifying glass.</p>
<p>The team, Lang Hoang Son and Bui Thanh Hung, chose to move beyond the convolutional neural networks (CNNs) that have dominated plant disease classification for years. CNNs such as ResNet and EfficientNet can achieve near-perfect scores on curated laboratory datasets, but they suffer from two fundamental weaknesses. They extract raw pixel-level features without high-level semantic reasoning, and they operate as inscrutable black boxes, offering farmers a class label with no justification. Worse, researchers have documented the so-called Clever Hans phenomenon, in which CNNs quietly learn to recognize background artifacts—soil, lighting, uniform laboratory backdrops—rather than the biological morphology of the disease itself. When such models confront real-world field images, their performance can collapse.</p>
<p>Vision-language models (VLMs) such as LLaVA, which combine a CLIP image encoder with a large language model backbone, promised a way out by generating human-readable diagnostic rationales. But they brought their own problem: the attention sink phenomenon. When fed complex real-world photographs, these models squander computational attention on noisy background elements—rocks, shadows, variable illumination—instead of focusing on the microscopic necrotic lesions that actually matter. The Vietnamese team&#8217;s solution was architectural: give the model two views of every leaf at once, one global and one local, so that attention is anchored to genuine pathology.</p>
<p>The local view is produced by an unsupervised Saliency Auto-Crop algorithm that requires no bounding box annotations, a major practical advantage because manual annotation is prohibitively expensive at agricultural scale. The algorithm first converts the input image into the CIE-LAB color space, decoupling luminance from chrominance. This matters because the a* channel is especially sensitive to the green-red spectrum, allowing necrotic brown lesions to stand out against green leaf tissue regardless of lighting conditions. A weighted saliency score is computed across the image, then Otsu&#8217;s adaptive thresholding and mathematical morphological operations filter out background noise. Finally, image moments are used to calculate the centroid of the isolated lesion, and a 128-by-128-pixel patch is cropped around it. The researchers empirically tested the patch size: crops of 64 pixels risk truncating diagnostic transition zones such as chlorotic halos, while 224-pixel crops capture so much healthy tissue that the patch becomes redundant with the global view.</p>
<p>These two images—a 256-pixel global view and the 128-pixel lesion patch—are encoded by a pre-trained CLIP ViT-L/14 encoder into 576 visual tokens each, projected into the language model&#8217;s embedding space, and concatenated with a carefully engineered text prompt that explicitly tells the model which image is which. The prompt instructs the system to identify exactly one of four categories: Black Rot, ESCA, Healthy, or Leaf Blight. Fine-tuning the 7-billion-parameter LLaVA-1.5 model conventionally would demand enormous computational resources, so the team applied 4-bit QLoRA, a quantized low-rank adaptation technique. The pre-trained weights are frozen and quantized into a 4-bit NormalFloat format, and lightweight trainable adapter matrices are injected into the decoder layers. Only 7.68 percent of the total parameters were updated, allowing the entire training run to complete on free cloud GPUs—an NVIDIA Tesla T4 pair with 32 gigabytes of memory—while the model acquired specialized plant pathology knowledge.</p>
<p>The results on the PlantVillage grapevine dataset, comprising 9,027 images split into training, validation, and blind test sets, were striking. The hybrid framework achieved an overall accuracy of 98.7 percent on the blind test, with Leaf Blight recognized perfectly at 100 percent, Healthy leaves at 99.8 percent, ESCA at 98.1 percent, and Black Rot at 97.0 percent—the last slightly hampered by the visual resemblance between early black rot lesions and leaf blight. An ablation study confirmed that both image streams are indispensable: the global image alone reached 97.8 percent, the local patch alone just 90.3 percent, and the un-tuned zero-shot model a dismal 18.6 percent. The QLoRA rank mattered too, with accuracy climbing from 62.1 percent at rank 8 to 98.0 percent at rank 32 before the team settled on rank 64 for maximum stability.</p>
<p>But the most consequential experiment was the one designed to expose the Clever Hans effect. The team evaluated all models in a strictly zero-shot manner on PlantDoc, a dataset of grapevine images captured in uncontrolled field conditions with complex backgrounds and volatile lighting. Advanced CNNs that had scored near 100 percent in the laboratory plummeted: EfficientNet-B4 fell from 99.8 percent to 54.9 percent, a catastrophic domain gap of 44.9 percentage points. The hybrid VLM framework, by contrast, retained 80.5 percent accuracy overall, with healthy leaf recognition peaking at 98.6 percent, though Black Rot degraded to 60.9 percent. The researchers attribute this resilience to the saliency-cropped local patches acting as semantic anchors that prevent the model from being derailed by environmental noise. The framework also scored 98.55 percent on the external AgroMind benchmark and a perfect 100 percent on AgroCoT, both laboratory-style datasets with subtle domain shifts.</p>
<p>Interpretability was engineered into the system rather than bolted on afterward. Instead of relying on post-hoc explanation tools like Grad-CAM, which merely visualize local gradients, the team extracted cross-attention matrices directly from the 31st and final decoder layer of the language model—the layer of highest semantic abstraction, immediately preceding the output head. A forward hook captures the mean attention weights that text tokens assign to visual tokens, which are then upsampled by bilinear interpolation and overlaid on the original image as heatmaps. Layer-wise analysis showed that shallow layers produce sparse, texture-focused activations, while deeper layers concentrate attention on pathological regions. Quantitatively, the model allocated roughly 73 percent of its attention to the global image and 27 percent to the local patch, a ratio that held steady whether the leaf was diseased or healthy—a Welch&#8217;s t-test found no significant difference between the two conditions, with a negligible effect size—indicating a stable, balanced diagnostic strategy rather than arbitrary attention spikes.</p>
<p>The authors are candid about the limitations. Inference is slow: the fine-tuned LLaVA requires 2,231 milliseconds per image, or 0.448 frames per second, and consumes 149 joules per image on a Tesla T4, making real-time deployment on drones or mobile devices impractical for now. Black Rot&#8217;s scattered, multi-spot lesions sometimes defeat the single-patch cropping strategy, motivating future multi-scale, multi-crop attention mechanisms. And the residual 17-point domain gap on field imagery shows that laboratory overfitting has not been fully eliminated. The team&#8217;s roadmap includes knowledge distillation to compress the model&#8217;s reasoning into lighter architectures, deeper quantization and pruning for on-device diagnostics, and evolving the static CIE-LAB extraction module into an autonomous visual agent that can propose and refine its own lesion bounding boxes through iterative reasoning.</p>
<p>Even with those caveats, the study marks a meaningful shift in agricultural AI. Where previous systems chased marginal accuracy gains on sterile benchmarks, this framework trades a sliver of in-domain performance for transparency, natural-language reasoning, and genuine cross-domain robustness. For vineyard managers, the difference is tangible: instead of an opaque percentage, they receive a diagnosis, a rationale, and a heatmap showing exactly which lesion triggered the decision. As vision-language models continue to mature, the grapevine may become the proving ground for a new generation of AI systems that agriculture can actually trust.</p>
<p><strong>Subject of Research:</strong> Interpretable vision-language AI for grapevine disease diagnosis</p>
<p><strong>Article Title:</strong> A global-local hybrid vision-language framework for interpretable grapevine disease diagnosis</p>
<p><strong>Article References:</strong> A global-local hybrid vision-language framework for interpretable grapevine disease diagnosis. (n.d.). <a href="https://doi.org/10.1007/s44163-026-02463-x" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02463-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02463-x" rel="noopener noreferrer">10.1007/s44163-026-02463-x</a></p>
<p><strong>Keywords:</strong> vision-language models, grapevine disease, plant pathology, explainable AI, LLaVA, QLoRA, saliency auto-crop, precision agriculture, convolutional neural networks, domain shift, cross-attention, PlantVillage</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">246834</post-id>	</item>
		<item>
		<title>Frequency-Decoupled AI Cracks the Cross-Center Problem in Nuclei Segmentation</title>
		<link>https://scienmag.com/frequency-decoupled-ai-cracks-the-cross-center-problem-in-nuclei-segmentation/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 03:56:01 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Aggregated Jaccard Index]]></category>
		<category><![CDATA[AI cross-hospital generalization]]></category>
		<category><![CDATA[automated tumor grading]]></category>
		<category><![CDATA[computational pathology]]></category>
		<category><![CDATA[cross-center generalization]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[digital pathology]]></category>
		<category><![CDATA[domain shift]]></category>
		<category><![CDATA[FDS-HoVerNet framework]]></category>
		<category><![CDATA[frequency decoupling]]></category>
		<category><![CDATA[frequency domain analysis]]></category>
		<category><![CDATA[frequency-decoupled neural networks]]></category>
		<category><![CDATA[H&E staining]]></category>
		<category><![CDATA[histopathology]]></category>
		<category><![CDATA[histopathology image analysis]]></category>
		<category><![CDATA[HoVer-Net]]></category>
		<category><![CDATA[medical image segmentation challenges]]></category>
		<category><![CDATA[nuclei segmentation]]></category>
		<category><![CDATA[precision oncology]]></category>
		<category><![CDATA[robust deep learning models for pathology]]></category>
		<category><![CDATA[transfer learning in medical imaging]]></category>
		<category><![CDATA[tumor-infiltrating lymphocytes quantification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233374</guid>

					<description><![CDATA[Researchers have developed a frequency-decoupled AI framework that nearly doubles cross-center nuclei segmentation accuracy in histopathology without requiring stain normalization or target-domain annotations.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has promised to transform pathology, but a stubborn obstacle keeps getting in the way: an algorithm trained to recognize cells in slides from one hospital often stumbles when shown slides from another. A team of researchers in China now reports a solution that could make digital pathology tools dramatically more portable, and their results suggest the fix lies in an unexpected place—the frequency domain of the images themselves. Writing in the Journal of Translational Medicine, the group led by Yanyun Liu, Xiangyu Liu, and Shouping Zhu of Xidian University describes a framework called FDS-HoVerNet that, without ever seeing labeled data from a new hospital, nearly doubles the accuracy of automated nuclei segmentation compared with conventional approaches.</p>
<p>The stakes are higher than they might first appear. In quantitative histopathology, the automated identification and outlining of individual cell nuclei is the foundation on which nearly everything else is built. Tumor grading, estimates of proliferative activity, counts of tumor-infiltrating lymphocytes, and many biomarkers used in precision oncology all depend on reliably separating one nucleus from another in crowded tissue images. If the segmentation step fails, every downstream measurement inherits that failure. Yet the deep learning models that perform this task are notoriously brittle when moved between institutions, a problem that has slowed the deployment of computational pathology in real clinical and research settings.</p>
<p>The reason for this brittleness is a phenomenon known as domain shift. Pathology slides are routinely stained with hematoxylin and eosin, but the exact appearance of that stain varies with the protocols and reagents used at each laboratory. Scanners from different manufacturers apply different color responses, compression schemes, and sharpness characteristics. Even the magnification at which a slide is digitized and the resulting image resolution can differ substantially between centers. To a convolutional neural network, these variations can look like meaningful signals, so a model trained at one center learns features that are partly tied to that center&#8217;s specific staining palette and scanner fingerprint rather than to the biology of the cells themselves. When the model encounters a new visual style, its performance can collapse.</p>
<p>Existing remedies each carry significant costs. Stain normalization attempts to recolor images to a standard palette, but it requires calibration and can distort subtle morphological details. Unsupervised domain adaptation techniques need access to unlabeled images from the target institution, which raises practical and privacy hurdles when patient data cannot leave a hospital. The simplest fix—having expert pathologists annotate slides from every new center—is expensive, slow, and often infeasible at scale. The research team set out to build a framework that would sidestep all three: no stain normalization, no target-domain data, and no additional annotations.</p>
<p>Their central insight is that the domain-specific nuisance information and the biologically meaningful information occupy different regions of the image&#8217;s frequency spectrum. High-frequency components of a histology image capture rapid intensity changes—fine texture, staining noise, scanner artifacts—while low-frequency components encode the slower, smoother variations that correspond to genuine nuclear structure: the shape, size, and spatial arrangement of cell nuclei. Rather than trying to suppress staining artifacts in the pixel domain, where structure and style are hopelessly intertwined, the framework explicitly separates the image into frequency bands and processes them differently. A structure texture disentanglement mechanism reduces the influence of domain-specific staining interference in the high-frequency pathway while preserving the low-frequency morphological cues that pathologists actually rely on.</p>
<p>The second challenge the team tackled is scale. Nuclei appear at different sizes depending on magnification and resolution, and models with a single fixed receptive field tend to overfit to the scale of their training data. The framework therefore incorporates a multi-receptive-field aggregation module, which gathers information across multiple spatial extents simultaneously so that the network can recognize nuclei whether they appear large and crisp or small and compressed. Together, these two components—frequency decoupling and scale-consistent morphology processing—form a robustness-oriented architecture built on top of the well-known HoVer-Net segmentation design, which predicts a nucleus probability map, horizontal and vertical displacement fields that point each pixel toward its nucleus center, and a per-pixel type prediction that are combined to yield individual segmented nuclei.</p>
<p>To test the approach under genuinely adversarial conditions, the researchers trained the model exclusively on the CoNSeP dataset, a single-center collection of colorectal histology images, and then evaluated it on six independent external datasets spanning diverse tumor types, staining protocols, and acquisition settings. This is a demanding protocol: the model never saw any image from the evaluation centers during training, and it received no fine-tuning or adaptation on them. Performance was measured with the Aggregated Jaccard Index, a standard metric for instance segmentation that penalizes both missed nuclei and poorly delineated boundaries.</p>
<p>The results were striking. On challenging multicenter benchmarks such as Kumar and MoNuSeg, the baseline HoVer-Net model achieved an Aggregated Jaccard Index of just 0.279, reflecting the severe degradation that domain shift inflicts on conventional training. FDS-HoVerNet raised that figure to 0.550, a relative improvement of roughly 97 percent. Perhaps most remarkably, the framework approached the performance of models that had been trained with annotations from the target domain—a comparison that underscores how much of the cross-center performance gap can be closed through architectural design alone, without collecting a single new labeled image. Qualitative analyses confirmed that the framework preserved critical nuclear structural characteristics even under substantial staining variability, keeping boundaries between adjacent nuclei intact where baseline models merged or fragmented them.</p>
<p>The implications extend well beyond a single benchmark. Because the framework requires no stain-specific calibration, no site-specific adaptation, and no target-domain supervision, it offers a practical path toward interoperable AI systems in multicenter cancer research. Hospitals and biobanks could deploy a single segmentation model across their archives without the annotation campaigns that currently gate such projects, and privacy constraints that prevent sharing patient slides between institutions become far less limiting. For precision oncology, where consistent quantification of tumor morphology across centers is essential for building large, comparable cohorts, the ability to generalize from one center&#8217;s data to many others&#8217; could accelerate biomarker discovery and validation considerably.</p>
<p>Caveats remain, as they do in any early-stage methodological advance. The framework was validated on hematoxylin and eosin-stained tissue, and its behavior on other stain types, rarer tumor morphologies, or extreme scanning conditions will require further study. The authors also note that the work was shared early as a citable, peer-reviewed accepted manuscript subject to final editorial processing. Still, the core demonstration stands: by separating what a model needs to see from what it should ignore, at the level of image frequency content, the researchers have turned one of computational pathology&#8217;s most persistent failure modes into a tractable engineering problem. If the approach holds up in broader deployment, the dream of pathology AI that travels as easily as the slides it analyzes moves a significant step closer to reality.</p>
<p><strong>Subject of Research:</strong> A frequency-decoupled deep learning framework for robust cross-center nuclei instance segmentation in histopathology images</p>
<p><strong>Article Title:</strong> A frequency-decoupled framework for robust cross-center nuclei instance segmentation in histopathology</p>
<p><strong>Article References:</strong> Liu, Y., Liu, X., Shi, Y., Chen, X., He, L., Ren, S., Yao, N., Song, J., &amp; Zhu, S. (2026). A frequency-decoupled framework for robust cross-center nuclei instance segmentation in histopathology. <em>Journal of Translational Medicine</em>. <a href="https://doi.org/10.1186/s12967-026-08984-4" rel="noopener noreferrer">https://doi.org/10.1186/s12967-026-08984-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12967-026-08984-4" rel="noopener noreferrer">10.1186/s12967-026-08984-4</a></p>
<p><strong>Keywords:</strong> nuclei segmentation, histopathology, domain shift, frequency decoupling, computational pathology, precision oncology, deep learning, H&amp;E staining, cross-center generalization, HoVer-Net, Aggregated Jaccard Index, digital pathology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233374</post-id>	</item>
		<item>
		<title>AI Learns to Distrust Its Own Illusions: CLIP Helps Models Adapt Without Source Data</title>
		<link>https://scienmag.com/ai-learns-to-distrust-its-own-illusions-clip-helps-models-adapt-without-source-data/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 22:50:04 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[CLIP]]></category>
		<category><![CDATA[confirmation bias]]></category>
		<category><![CDATA[distribution alignment]]></category>
		<category><![CDATA[domain shift]]></category>
		<category><![CDATA[knowledge distillation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[pseudo-source domain]]></category>
		<category><![CDATA[source-free domain adaptation]]></category>
		<category><![CDATA[spurious correlations]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[vision-language models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205183</guid>

					<description><![CDATA[Researchers have developed SRRI, a source-free domain adaptation method that uses CLIP as an external verifier to de-bias pseudo-source domains and outperform state-of-the-art approaches.]]></description>
										<content:encoded><![CDATA[<p>Machine learning models are often trained in one setting and deployed in another, and that transition is rarely seamless. A classifier trained on studio photographs of cars may stumble when shown cars in rain, at night, or through a security camera. Domain adaptation is the branch of machine learning that tackles this problem, and its most demanding variant, source-free domain adaptation, adds a harsh constraint: once the model leaves the training environment, the original labeled data is gone, often for reasons of privacy, storage, or proprietary restriction. All the model can carry with it is what it learned. A new study published in the journal Machine Learning by Qing Tian, Yongjiang Liu, Keyang Cheng, Weihua Ou, Jianping Gou and colleagues confronts a subtle but pervasive failure mode in this setting, one the researchers describe with an evocative word borrowed from human cognition: illusions.</p>
<p>The core difficulty lies in how a source model, cut off from its training data, tries to make sense of an unlabeled target domain. A popular family of methods constructs a pseudo-source domain, a synthetic stand-in for the lost source data, by selecting target samples that the source model believes look source-like. Training on this pseudo-source then reduces the statistical gap between domains. The catch, as the new paper argues, is that this construction leans entirely on the source model itself, and the source model is precisely the entity most likely to be deceived. During training, it absorbed not only the true signals that define each class, such as the shape and structure of an object, but also spurious correlations linking class identity to domain-specific features like background, lighting, or environment. When asked to identify pseudo-source samples, it may confirm its own biases, mistaking samples that share superficial environmental cues with the source domain for genuinely representative ones.</p>
<p>The researchers identify two intertwined challenges that undermine this process. The first is confirmation bias: once the model commits to a belief about a sample, the subsequent training on that sample reinforces the belief, whether or not it was correct. This is the machine analogue of a person who reads only news that agrees with their views. The second is domain shift, the distributional mismatch between source and target that makes the source model&#8217;s judgments unreliable in the first place. Together, these effects can contaminate the pseudo-source domain with mislabeled or unrepresentative samples, and the errors compound as adaptation proceeds. The team&#8217;s answer is a method they call Staying Rational and Resisting Illusions, or SRRI, which refuses to let the source model grade its own homework.</p>
<p>The key innovation is the introduction of an external referee: CLIP, the Contrastive Language-Image Pre-training model developed by OpenAI researchers, which learned to align images and text by training on hundreds of millions of image-text pairs from the web. Because CLIP&#8217;s knowledge comes from a vastly broader data distribution than the narrow source domain, it does not share the source model&#8217;s spurious correlations. SRRI uses CLIP as a source of independent evidence when deciding which target samples deserve a place in the pseudo-source domain. In effect, when the source model says a sample looks familiar, SRRI asks CLIP for a second opinion before accepting the claim.</p>
<p>Technically, the method proceeds in several coordinated stages. First, SRRI employs knowledge distillation, a technique in which a teacher network&#8217;s outputs guide a student network, to help the source model disentangle class-discriminative causal features from domain-specific spurious features. The goal is to teach the model which aspects of an image actually cause its label, such as the geometry of an object, and which merely co-occur with it, such as the typical backdrop of the source photographs. This disentanglement weakens the illusions at their root, making the model&#8217;s own judgments less confounded before any pseudo-source construction begins.</p>
<p>Distillation alone, however, cannot be trusted blindly, because in some adaptation tasks the distillation process itself performs poorly, propagating errors rather than correcting them. To guard against this, the authors design a CLIP-guided dual-model validation and class balancing strategy. Every candidate pseudo-source sample must pass inspection by both the distilled source model and CLIP, and the two models&#8217; assessments are combined to filter out unreliable examples. Class balancing ensures that the retained samples cover all categories with reasonable richness, preventing the pseudo-source domain from being dominated by easy or overrepresented classes. This dual gatekeeping is what allows the method to remain robust even when one of its components falters on a given task.</p>
<p>The third pillar of SRRI is a dynamic pseudo-source domain optimization mechanism. Rather than freezing the pseudo-source once it is built, the method continuously fine-tunes the task-specific prompts of CLIP during adaptation. Prompt tuning adjusts the short text descriptions that CLIP uses to interpret images, sharpening its sensitivity to the specific categories of the target task. As these prompts improve, CLIP&#8217;s judgments on hard samples become more accurate, which in turn corrects residual bias in the target model. At the same time, the pseudo-source domain is periodically reconstructed and refined, discarding samples that no longer pass validation and admitting better ones as the models evolve. The result is a self-correcting loop in which the reference data improves alongside the adapting model.</p>
<p>With a trustworthy pseudo-source domain in place, SRRI applies robust supervised learning to train the target model on the pseudo-source samples, while simultaneously performing distribution alignment between the pseudo-source and the true target data. This alignment ensures that the model does not overfit to artifacts of the pseudo-source construction and that its decision boundaries remain well matched to the actual deployment distribution. The combination of reliable pseudo-labels, balanced classes, and distributional consistency addresses both of the fundamental challenges the authors set out to solve: confirmation bias is curbed by external validation, and domain shift is absorbed by the alignment procedure.</p>
<p>Extensive experiments reported in the paper show that SRRI outperforms state-of-the-art source-free domain adaptation methods across standard benchmarks. The improvements are attributed not to any single trick but to the architecture of trust the method builds: an independent verifier, a disentangled representation, a balanced and evolving reference set, and a training objective that keeps the target model anchored to reality. All datasets used in the study are publicly available, which should make the approach straightforward for other groups to reproduce and extend. The work was supported by the National Natural Science Foundation of China and several regional research programs, and the authors report no competing financial interests beyond these funding sources.</p>
<p>The broader significance of this research extends beyond a single benchmark. As artificial intelligence systems are increasingly deployed in hospitals, vehicles, and surveillance networks where raw training data cannot be shared, source-free adaptation will become a standard requirement rather than a niche concern. The lesson of SRRI is that a model adapting in the wild should not rely solely on its own inherited judgments, because those judgments may encode illusions about what really defines a category. By recruiting a vision-language foundation model as an external rational check, and by continuously refining both the verifier and the verified, the researchers offer a template for building machine learning systems that stay rational under pressure, resisting the very biases they were born with. In an era when AI is often criticized for confidently repeating its mistakes, a method explicitly designed to resist its own illusions is a welcome step toward more trustworthy machine intelligence.</p>
<p><strong>Subject of Research:</strong> De-biasing source-free domain adaptation using a CLIP-verified pseudo-source domain</p>
<p><strong>Article Title:</strong> Stay Rational, Resist Illusions: De-biasing Source-Free Domain Adaptation with CLIP-Verified Pseudo-Source Domain</p>
<p><strong>Article References:</strong> Tian, Q., Liu, Y., Cheng, K., Ou, W., &amp; Gou, J. (2026). Stay Rational, Resist Illusions: De-biasing Source-Free Domain Adaptation with CLIP-Verified Pseudo-Source Domain. <em>Machine Learning, 115</em>(10), Article 222. <a href="https://doi.org/10.1007/s10994-026-07163-2" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07163-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07163-2" rel="noopener noreferrer">10.1007/s10994-026-07163-2</a></p>
<p><strong>Keywords:</strong> source-free domain adaptation, CLIP, pseudo-source domain, knowledge distillation, confirmation bias, domain shift, spurious correlations, prompt tuning, distribution alignment, machine learning, vision-language models, transfer learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205183</post-id>	</item>
		<item>
		<title>Prompt Robustness and Fine-Tuning Tested in Open-Vocabulary Object Detection Showdown</title>
		<link>https://scienmag.com/prompt-robustness-and-fine-tuning-tested-in-open-vocabulary-object-detection-showdown/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 01:58:51 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[domain shift]]></category>
		<category><![CDATA[fine-tuning]]></category>
		<category><![CDATA[inference latency]]></category>
		<category><![CDATA[mAP]]></category>
		<category><![CDATA[object recognition]]></category>
		<category><![CDATA[open-vocabulary detection]]></category>
		<category><![CDATA[prompt sensitivity]]></category>
		<category><![CDATA[vision-language models]]></category>
		<category><![CDATA[YOLO-World]]></category>
		<category><![CDATA[YOLOE]]></category>
		<category><![CDATA[zero-shot object detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204992</guid>

					<description><![CDATA[A new comparative study finds that YOLO-World v2 offers the best balance of accuracy, speed, and prompt robustness among open-vocabulary object detectors, while YOLOE loses its open-vocabulary behavior after fine-tuning.]]></description>
										<content:encoded><![CDATA[<p>Open-vocabulary object detection has quietly become one of the most consequential ideas in modern computer vision. Instead of being locked to a fixed list of categories learned during training, these models can recognize objects described in plain language—a text prompt such as &#8220;traffic sign&#8221; or &#8220;kitchen appliance&#8221; is enough to make them find and localize instances they were never explicitly trained on. The promise is enormous: robots that understand novel instructions, surveillance systems that adapt to new threats, and annotation pipelines that label images without human effort. Yet a new systematic study from Selcuk University in Konya, Türkiye, suggests that the field&#8217;s enthusiasm for headline accuracy numbers has obscured a more complicated reality, one in which the choice of words, the cost of inference, and the fate of unseen classes after fine-tuning can matter as much as raw performance.</p>
<p>The research, published in Multimedia Tools and Applications by Melisa Alara Ozuberk and Ilkay Cinar, delivers one of the first head-to-head evaluations of three leading real-time open-vocabulary detectors: YOLO-World, its successor YOLO-World v2, and YOLOE. Rather than benchmarking accuracy alone, the authors designed their experiments around three scenarios that mirror how these systems are actually deployed: zero-shot inference on entirely new datasets, sensitivity to variations in the textual prompts that steer detection, and fine-tuning on domain-specific data followed by tests of whether open-vocabulary generalization survives. The evaluation spans three datasets with deliberately different characteristics—the classic VOC2012 segmentation subset, the HomeObjects-3K indoor detection dataset, and the demanding KITTI autonomous driving benchmark.</p>
<p>The zero-shot results reveal a striking dependence on domain. The highest performance was achieved on HomeObjects-3K, where YOLO-World v2 reached 0.443 mAP@0.5:0.95, a metric that rewards both accurate localization and correct classification across a range of overlap thresholds. KITTI, by contrast, produced the weakest results across all three models, a consequence of domain shift: the driving imagery, with its unusual viewpoints, small distant objects, and harsh lighting conditions, differs substantially from the data distributions these models encountered during pre-training. On the VOC dataset, YOLOE claimed the highest zero-shot accuracy at 0.310 mAP@0.5:0.95, outperforming both YOLO-World variants on that benchmark, although this advantage came with a caveat that emerged clearly in the timing analysis.</p>
<p>That caveat is speed. YOLO-World v2 proved to offer the best overall balance between accuracy and throughput, sustaining between 20 and 30 frames per second—comfortably real-time for many applications. YOLOE, despite its stronger zero-shot accuracy on VOC, exhibited lower inference speed in some configurations, a trade-off that could prove decisive in latency-sensitive settings such as autonomous navigation or live video analytics. The YOLO-World family also benefited from an embedding cache mechanism, which pre-computes text embeddings for the prompt vocabulary and reuses them across frames. Latency analysis with increasing prompt counts showed that this design keeps inference efficient even as the number of textual categories grows, an architectural advantage that becomes more valuable the richer the vocabulary deployed in production.</p>
<p>Perhaps the most practically important finding concerns prompt robustness—the question of how much detection quality degrades when the words fed to the model change. The authors constructed four categories of prompts for each dataset: base prompts, attribute prompts that add descriptive modifiers, longer descriptive prompts, and noisy prompts containing degraded or perturbed language. The results showed measurable performance drops under noisy prompt conditions for all models, confirming that open-vocabulary detectors are not immune to the fragility of language interfaces that has been documented across the broader vision-language literature. However, YOLO-World v2 maintained better stability across these variations than its competitors, suggesting that its training recipe or text-encoding pathway confers a degree of resilience that practitioners should weigh when deploying systems in the hands of non-expert users who cannot be relied upon to craft optimal prompts.</p>
<p>Fine-tuning delivered the expected gains but also exposed an uncomfortable truth about what adaptation costs. After fine-tuning on each dataset, mAP scores improved for all three models, demonstrating that standard transfer learning techniques remain effective when open-vocabulary detectors are specialized to a target domain. But the fine-tuned models showed zero performance on some unseen categories—classes that were never part of the fine-tuning data. This is precisely the failure mode that open-vocabulary detection is supposed to prevent, and its appearance after adaptation indicates that the boundary between open and closed vocabulary is thinner than the field often assumes.</p>
<p>The divergence between the model families was especially pronounced here. The YOLO-World family retained some of its open-vocabulary generalization ability after fine-tuning, continuing to respond to textual prompts for categories outside the training set. YOLOE, under the fine-tuning protocol adopted in the study, exhibited closed-set-like behavior: its predictions remained insensitive to the evaluated prompt variations, effectively behaving as if the text interface had been switched off and the model had reverted to a conventional fixed-category detector. For teams choosing between these architectures, the implication is significant—fine-tuning YOLOE may buy accuracy on known classes at the price of the very flexibility that motivated choosing an open-vocabulary model in the first place.</p>
<p>The study&#8217;s methodology reflects a growing recognition that evaluation practices in this field have been too narrow. The authors note that existing research has focused mainly on accuracy metrics while prompt robustness, unseen class generalization, and computational costs are rarely assessed together. By combining confusion-matrix-based analysis, standard detection metrics such as mAP at multiple intersection-over-union thresholds, and latency profiling under varying prompt counts, the work offers a template for more honest benchmarking. The datasets themselves are all publicly available—VOC2012 from the PASCAL repository, KITTI from the KITTI Vision Benchmark Suite, and HomeObjects-3K through its original repository—making the evaluation pipeline reproducible by other groups.</p>
<p>The broader context makes these findings timely. Open-vocabulary detection builds on a lineage that runs from the original YOLO real-time detector through open-set recognition and open-world detection to caption-supervised methods and the CLIP-style vision-language models that supply the text-image alignment these detectors depend on. YOLO-World, introduced in 2024, brought this capability to real-time speeds, and YOLOE pushed the concept further with its &#8220;see anything&#8221; design. Applications documented in the literature now span automatic image annotation, number plate recognition, wildlife monitoring, medical imaging, underwater fish counting, robotic navigation, and anomaly detection in surveillance—domains where the ability to name new categories without retraining is transformative.</p>
<p>For practitioners, the study&#8217;s bottom line is that model selection should be a multi-dimensional decision. Accuracy, inference speed, prompt robustness, and unseen class generalization form a set of trade-offs that no single model dominates. YOLO-World v2 emerges as the most balanced option, combining competitive accuracy, real-time throughput, an efficient embedding cache, and the strongest prompt stability. YOLOE offers the best zero-shot accuracy in some settings but pays in speed and, critically, appears to surrender its open-vocabulary character when fine-tuned. As these systems move from research demos into safety-relevant deployments—self-driving perception, medical triage, industrial inspection—the lesson of this comparative study is that the questions worth asking about a detector extend well beyond its leaderboard score, reaching into how it behaves when the words change, the domain shifts, and the training data runs out.</p>
<p><strong>Subject of Research:</strong> Comparative evaluation of prompt robustness, fine-tuning, and generalization in open-vocabulary object detection models</p>
<p><strong>Article Title:</strong> Prompt robustness, fine-tuning, and Generalization in open-vocabulary object detection: a comparative study of YOLO-World, YOLO-World v2 and YOLOE</p>
<p><strong>Article References:</strong> Ozuberk, M. A., &amp; Cinar, I. (2026). Prompt robustness, fine-tuning, and Generalization in open-vocabulary object detection: a comparative study of YOLO-World, YOLO-World v2 and YOLOE. <em>Multimedia Tools and Applications, 85</em>(10), Article 767. <a href="https://doi.org/10.1007/s11042-026-21928-w" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21928-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21928-w" rel="noopener noreferrer">10.1007/s11042-026-21928-w</a></p>
<p><strong>Keywords:</strong> open-vocabulary detection, YOLO-World, YOLOE, zero-shot object detection, prompt sensitivity, fine-tuning, computer vision, mAP, inference latency, domain shift, vision-language models, object recognition</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204992</post-id>	</item>
	</channel>
</rss>
