<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>mAP &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/map/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 04:29:27 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>mAP &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Smarter AI Spots Hidden Dental Problems in Panoramic X-Rays</title>
		<link>https://scienmag.com/smarter-ai-spots-hidden-dental-problems-in-panoramic-x-rays/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 04:29:27 +0000</pubDate>
				<category><![CDATA[Science News]]></category>
		<category><![CDATA[advanced AI frameworks for dental imaging]]></category>
		<category><![CDATA[AI-based dental lesion detection]]></category>
		<category><![CDATA[artificial intelligence in dentistry]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[computer-aided dental diagnosis tools]]></category>
		<category><![CDATA[computer-aided diagnosis]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for dental radiology]]></category>
		<category><![CDATA[dental imaging challenges]]></category>
		<category><![CDATA[dental panoramic radiographs]]></category>
		<category><![CDATA[detection of small dental lesions]]></category>
		<category><![CDATA[feature pyramid]]></category>
		<category><![CDATA[imaging of impacted teeth and root anomalies]]></category>
		<category><![CDATA[impact of AI on dental diagnosis]]></category>
		<category><![CDATA[mAP]]></category>
		<category><![CDATA[Medical Imaging]]></category>
		<category><![CDATA[object detection]]></category>
		<category><![CDATA[panoramic X-ray analysis]]></category>
		<category><![CDATA[real-time dental diagnostics]]></category>
		<category><![CDATA[real-time inference]]></category>
		<category><![CDATA[recall]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[YOLOv8n]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=251833</guid>

					<description><![CDATA[A new AI framework called RCTE improves the detection of small lesions and complex dental structures in panoramic radiographs while keeping real-time speed.]]></description>
										<content:encoded><![CDATA[<p>Dental panoramic radiographs are among the most widely used imaging tools in dentistry, offering a sweeping view of the teeth, jaws, and surrounding structures in a single exposure. Yet reading them well is notoriously difficult. Lesions can be tiny, anatomical boundaries often blur into one another, and complex structures such as impacted teeth or root anomalies can hide in plain sight. A research team led by Jiayi Peng and colleagues now reports a new artificial intelligence framework, called RCTE, that tackles these challenges head-on, delivering measurably better detection of dental structures and lesions while still running fast enough for real-time clinical use.</p>
<p>The study, published in PLOS One, addresses a long-standing bottleneck in computer-aided dental diagnosis. Modern object detectors excel at finding large, well-defined objects in natural images, but dental radiographs are a different beast. A small periapical lesion may occupy only a handful of pixels, its edges dissolving into the surrounding bone. Teeth overlap in the projection, implants mimic natural roots, and the same anatomical feature can appear at wildly different scales depending on the patient&#8217;s anatomy and positioning. Detectors trained on generic imagery tend to miss these subtle targets, and missed detections in a clinical setting can mean overlooked pathology.</p>
<p>The researchers built their framework on YOLOv8n, a lightweight member of the widely used You Only Look Once family of detectors. YOLO models are prized for speed, processing an entire image in a single pass through a neural network, which makes them attractive for clinical workflows where results must appear in seconds. The baseline YOLOv8n, however, was not designed with the peculiar demands of dental radiography in mind. The team&#8217;s contribution lies in three carefully engineered modules that retrofit the detector for exactly those demands, each targeting a distinct failure mode of the original architecture.</p>
<p>The first component is a Cross-Scale Channel Transformer, or CSCT, module. In feature pyramid architectures like YOLO&#8217;s, the network extracts features at multiple resolutions, conventionally labeled P3, P4, and P5, with P3 capturing fine, high-resolution detail and P5 capturing coarse, high-level context. Small lesions live mostly in the P3 layer, but interpreting them correctly requires context from the broader scene, such as the position of neighboring teeth or the outline of the jaw. The CSCT module lets these three levels talk to each other, using a transformer-style attention mechanism to exchange channel-wise information across scales. The result is that a faint, ambiguous patch of bone loss can be evaluated in light of its anatomical surroundings rather than in isolation.</p>
<p>The second component, a Re-parameterized Feature Pyramid Fusion structure built on a RepNCSPELAN4 block, tackles the problem of aggregating multi-scale features efficiently. Re-parameterization is a clever training trick: the network uses a richer, multi-branch structure while learning, which gives it more expressive capacity to blend features from different levels, and then the branches are mathematically collapsed into a single streamlined path for inference. The model effectively gets the accuracy benefits of a heavier architecture at the computational cost of a light one. For dental detection, where both subtle multi-scale patterns and clinical speed matter, this trade-off is precisely the right one.</p>
<p>The third innovation, a Multi-Scale EMA mechanism, refines features just before they reach the detection head. EMA stands for Efficient Multi-scale Attention, a mechanism that recalibrates features by emphasizing the spatial positions and channels most relevant to the target while suppressing background noise. By applying this recalibration across multiple scales, MS-EMA ensures that the final detection decisions rest on the clearest possible representation of each candidate object. In panoramic radiographs, where anatomical clutter is the norm rather than the exception, this last-stage cleanup proved especially valuable for reducing missed detections of complex dental structures.</p>
<p>To test the framework, the team used a publicly available dataset of dental panoramic radiographs covering eleven categories of dental structures and lesions, ranging from individual teeth and restorations to pathological findings. Performance was measured with standard object detection metrics: precision, which tracks how many detections are actually correct; recall, which tracks how many true targets are found; the F1-score, which balances the two; and mean average precision at several intersection-over-union thresholds, including mAP50, mAP75, and the stricter mAP50-95 average. These thresholds reward not just finding an object but drawing a tight, accurate box around it, which matters when a bounding box may later guide a clinician&#8217;s eye.</p>
<p>The results showed consistent gains over the baseline. Compared with the original YOLOv8n, RCTE improved mAP50 by 2.55 percentage points, mAP75 by 3.88 points, mAP50-95 by 2.40 points, and recall by 5.10 points. The recall improvement is arguably the headline number, because recall measures detection completeness, and in a screening context a missed lesion is usually more costly than a false alarm. The especially large gain at the stricter mAP75 threshold also indicates that the framework localizes targets more precisely, drawing boundaries that hug the true extent of each structure. Head-to-head comparisons against other YOLO-based detectors confirmed that RCTE offered better detection completeness and localization accuracy than its competitors, and crucially, the model retained real-time inference capability, preserving the speed advantage that makes YOLO-family detectors clinically practical.</p>
<p>What makes this work notable beyond its benchmark numbers is the way it illustrates a broader trend in medical imaging AI. Rather than inventing detection from scratch, the researchers identified the specific ways a general-purpose detector fails on dental images, such as weak multi-scale interaction, shallow feature fusion, and noisy final features, and designed targeted, composable fixes for each. This modular philosophy means the individual components, the cross-scale transformer, the re-parameterized fusion structure, and the multi-scale attention recalibration, could plausibly be ported to other radiographic domains with similar pathologies of scale and contrast, from cephalometric analysis to general skeletal imaging. The framework thus serves as both a practical tool and a design template.</p>
<p>The authors position RCTE as a potential solution for computer-aided dental image analysis, and the clinical implications are straightforward. A detector that finds more lesions, localizes them more tightly, and answers in real time could serve as a second pair of eyes for dentists, flagging findings that might otherwise slip through a busy clinic&#8217;s workflow and helping standardize the quality of radiographic screening across practitioners. The work also underscores the value of public datasets and rigorous multi-metric evaluation in pushing medical AI forward. As dental practices increasingly adopt digital imaging pipelines, frameworks like RCTE suggest a future in which the humble panoramic X-ray, read with the aid of a fast and attentive neural network, catches more disease earlier and with fewer oversights.</p>
<p><strong>Subject of Research:</strong> Multi-class object detection in dental panoramic radiographs using an enhanced YOLOv8n-based deep learning framework</p>
<p><strong>Article Title:</strong> RCTE: A multi-class object detection framework for dental panoramic radiographs</p>
<p><strong>Article References:</strong> Peng, J., Liu, J., Shen, Y., Yin, M., Liu, C., Zhou, J., Liu, J., Zhang, R., &amp; Hong, Q. (2026). RCTE: A multi-class object detection framework for dental panoramic radiographs. <em>PLOS One, 21</em>(10), e0359630. <a href="https://doi.org/10.1371/journal.pone.0359630" rel="noopener noreferrer">https://doi.org/10.1371/journal.pone.0359630</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pone.0359630" rel="noopener noreferrer">10.1371/journal.pone.0359630</a></p>
<p><strong>Keywords:</strong> dental panoramic radiographs, object detection, YOLOv8n, deep learning, computer-aided diagnosis, feature pyramid, attention mechanism, transformer, medical imaging, recall, mAP, real-time inference</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">251833</post-id>	</item>
		<item>
		<title>Prompt Robustness and Fine-Tuning Tested in Open-Vocabulary Object Detection Showdown</title>
		<link>https://scienmag.com/prompt-robustness-and-fine-tuning-tested-in-open-vocabulary-object-detection-showdown/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 01:58:51 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[domain shift]]></category>
		<category><![CDATA[fine-tuning]]></category>
		<category><![CDATA[inference latency]]></category>
		<category><![CDATA[mAP]]></category>
		<category><![CDATA[object recognition]]></category>
		<category><![CDATA[open-vocabulary detection]]></category>
		<category><![CDATA[prompt sensitivity]]></category>
		<category><![CDATA[vision-language models]]></category>
		<category><![CDATA[YOLO-World]]></category>
		<category><![CDATA[YOLOE]]></category>
		<category><![CDATA[zero-shot object detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204992</guid>

					<description><![CDATA[A new comparative study finds that YOLO-World v2 offers the best balance of accuracy, speed, and prompt robustness among open-vocabulary object detectors, while YOLOE loses its open-vocabulary behavior after fine-tuning.]]></description>
										<content:encoded><![CDATA[<p>Open-vocabulary object detection has quietly become one of the most consequential ideas in modern computer vision. Instead of being locked to a fixed list of categories learned during training, these models can recognize objects described in plain language—a text prompt such as &#8220;traffic sign&#8221; or &#8220;kitchen appliance&#8221; is enough to make them find and localize instances they were never explicitly trained on. The promise is enormous: robots that understand novel instructions, surveillance systems that adapt to new threats, and annotation pipelines that label images without human effort. Yet a new systematic study from Selcuk University in Konya, Türkiye, suggests that the field&#8217;s enthusiasm for headline accuracy numbers has obscured a more complicated reality, one in which the choice of words, the cost of inference, and the fate of unseen classes after fine-tuning can matter as much as raw performance.</p>
<p>The research, published in Multimedia Tools and Applications by Melisa Alara Ozuberk and Ilkay Cinar, delivers one of the first head-to-head evaluations of three leading real-time open-vocabulary detectors: YOLO-World, its successor YOLO-World v2, and YOLOE. Rather than benchmarking accuracy alone, the authors designed their experiments around three scenarios that mirror how these systems are actually deployed: zero-shot inference on entirely new datasets, sensitivity to variations in the textual prompts that steer detection, and fine-tuning on domain-specific data followed by tests of whether open-vocabulary generalization survives. The evaluation spans three datasets with deliberately different characteristics—the classic VOC2012 segmentation subset, the HomeObjects-3K indoor detection dataset, and the demanding KITTI autonomous driving benchmark.</p>
<p>The zero-shot results reveal a striking dependence on domain. The highest performance was achieved on HomeObjects-3K, where YOLO-World v2 reached 0.443 mAP@0.5:0.95, a metric that rewards both accurate localization and correct classification across a range of overlap thresholds. KITTI, by contrast, produced the weakest results across all three models, a consequence of domain shift: the driving imagery, with its unusual viewpoints, small distant objects, and harsh lighting conditions, differs substantially from the data distributions these models encountered during pre-training. On the VOC dataset, YOLOE claimed the highest zero-shot accuracy at 0.310 mAP@0.5:0.95, outperforming both YOLO-World variants on that benchmark, although this advantage came with a caveat that emerged clearly in the timing analysis.</p>
<p>That caveat is speed. YOLO-World v2 proved to offer the best overall balance between accuracy and throughput, sustaining between 20 and 30 frames per second—comfortably real-time for many applications. YOLOE, despite its stronger zero-shot accuracy on VOC, exhibited lower inference speed in some configurations, a trade-off that could prove decisive in latency-sensitive settings such as autonomous navigation or live video analytics. The YOLO-World family also benefited from an embedding cache mechanism, which pre-computes text embeddings for the prompt vocabulary and reuses them across frames. Latency analysis with increasing prompt counts showed that this design keeps inference efficient even as the number of textual categories grows, an architectural advantage that becomes more valuable the richer the vocabulary deployed in production.</p>
<p>Perhaps the most practically important finding concerns prompt robustness—the question of how much detection quality degrades when the words fed to the model change. The authors constructed four categories of prompts for each dataset: base prompts, attribute prompts that add descriptive modifiers, longer descriptive prompts, and noisy prompts containing degraded or perturbed language. The results showed measurable performance drops under noisy prompt conditions for all models, confirming that open-vocabulary detectors are not immune to the fragility of language interfaces that has been documented across the broader vision-language literature. However, YOLO-World v2 maintained better stability across these variations than its competitors, suggesting that its training recipe or text-encoding pathway confers a degree of resilience that practitioners should weigh when deploying systems in the hands of non-expert users who cannot be relied upon to craft optimal prompts.</p>
<p>Fine-tuning delivered the expected gains but also exposed an uncomfortable truth about what adaptation costs. After fine-tuning on each dataset, mAP scores improved for all three models, demonstrating that standard transfer learning techniques remain effective when open-vocabulary detectors are specialized to a target domain. But the fine-tuned models showed zero performance on some unseen categories—classes that were never part of the fine-tuning data. This is precisely the failure mode that open-vocabulary detection is supposed to prevent, and its appearance after adaptation indicates that the boundary between open and closed vocabulary is thinner than the field often assumes.</p>
<p>The divergence between the model families was especially pronounced here. The YOLO-World family retained some of its open-vocabulary generalization ability after fine-tuning, continuing to respond to textual prompts for categories outside the training set. YOLOE, under the fine-tuning protocol adopted in the study, exhibited closed-set-like behavior: its predictions remained insensitive to the evaluated prompt variations, effectively behaving as if the text interface had been switched off and the model had reverted to a conventional fixed-category detector. For teams choosing between these architectures, the implication is significant—fine-tuning YOLOE may buy accuracy on known classes at the price of the very flexibility that motivated choosing an open-vocabulary model in the first place.</p>
<p>The study&#8217;s methodology reflects a growing recognition that evaluation practices in this field have been too narrow. The authors note that existing research has focused mainly on accuracy metrics while prompt robustness, unseen class generalization, and computational costs are rarely assessed together. By combining confusion-matrix-based analysis, standard detection metrics such as mAP at multiple intersection-over-union thresholds, and latency profiling under varying prompt counts, the work offers a template for more honest benchmarking. The datasets themselves are all publicly available—VOC2012 from the PASCAL repository, KITTI from the KITTI Vision Benchmark Suite, and HomeObjects-3K through its original repository—making the evaluation pipeline reproducible by other groups.</p>
<p>The broader context makes these findings timely. Open-vocabulary detection builds on a lineage that runs from the original YOLO real-time detector through open-set recognition and open-world detection to caption-supervised methods and the CLIP-style vision-language models that supply the text-image alignment these detectors depend on. YOLO-World, introduced in 2024, brought this capability to real-time speeds, and YOLOE pushed the concept further with its &#8220;see anything&#8221; design. Applications documented in the literature now span automatic image annotation, number plate recognition, wildlife monitoring, medical imaging, underwater fish counting, robotic navigation, and anomaly detection in surveillance—domains where the ability to name new categories without retraining is transformative.</p>
<p>For practitioners, the study&#8217;s bottom line is that model selection should be a multi-dimensional decision. Accuracy, inference speed, prompt robustness, and unseen class generalization form a set of trade-offs that no single model dominates. YOLO-World v2 emerges as the most balanced option, combining competitive accuracy, real-time throughput, an efficient embedding cache, and the strongest prompt stability. YOLOE offers the best zero-shot accuracy in some settings but pays in speed and, critically, appears to surrender its open-vocabulary character when fine-tuned. As these systems move from research demos into safety-relevant deployments—self-driving perception, medical triage, industrial inspection—the lesson of this comparative study is that the questions worth asking about a detector extend well beyond its leaderboard score, reaching into how it behaves when the words change, the domain shifts, and the training data runs out.</p>
<p><strong>Subject of Research:</strong> Comparative evaluation of prompt robustness, fine-tuning, and generalization in open-vocabulary object detection models</p>
<p><strong>Article Title:</strong> Prompt robustness, fine-tuning, and Generalization in open-vocabulary object detection: a comparative study of YOLO-World, YOLO-World v2 and YOLOE</p>
<p><strong>Article References:</strong> Ozuberk, M. A., &amp; Cinar, I. (2026). Prompt robustness, fine-tuning, and Generalization in open-vocabulary object detection: a comparative study of YOLO-World, YOLO-World v2 and YOLOE. <em>Multimedia Tools and Applications, 85</em>(10), Article 767. <a href="https://doi.org/10.1007/s11042-026-21928-w" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21928-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21928-w" rel="noopener noreferrer">10.1007/s11042-026-21928-w</a></p>
<p><strong>Keywords:</strong> open-vocabulary detection, YOLO-World, YOLOE, zero-shot object detection, prompt sensitivity, fine-tuning, computer vision, mAP, inference latency, domain shift, vision-language models, object recognition</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204992</post-id>	</item>
	</channel>
</rss>
