<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>multi-loss optimization for image retrieval &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/multi-loss-optimization-for-image-retrieval/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 09:23:22 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>multi-loss optimization for image retrieval &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy</title>
		<link>https://scienmag.com/hybrid-cnn-vit-model-with-triple-loss-boosts-image-search-accuracy/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 09:23:22 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-powered image similarity matching]]></category>
		<category><![CDATA[CIFAR-100]]></category>
		<category><![CDATA[content-based image retrieval]]></category>
		<category><![CDATA[content-based image retrieval advancements]]></category>
		<category><![CDATA[contrastive loss]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[Convolutional Neural Networks and Vision Transformers]]></category>
		<category><![CDATA[cross-entropy loss]]></category>
		<category><![CDATA[deep learning in image search]]></category>
		<category><![CDATA[deep metric learning]]></category>
		<category><![CDATA[feature embedding]]></category>
		<category><![CDATA[fine-grained retrieval]]></category>
		<category><![CDATA[fusion of CNN and ViT in computer vision]]></category>
		<category><![CDATA[HCV-MLO]]></category>
		<category><![CDATA[hybrid architecture]]></category>
		<category><![CDATA[hybrid CNN-vision transformer architecture]]></category>
		<category><![CDATA[improving image search accuracy with hybrid models]]></category>
		<category><![CDATA[innovative deep learning frameworks in image retrieval]]></category>
		<category><![CDATA[machine learning models for visual search]]></category>
		<category><![CDATA[multi-loss optimization for image retrieval]]></category>
		<category><![CDATA[semantic image similarity detection]]></category>
		<category><![CDATA[triple loss function image retrieval]]></category>
		<category><![CDATA[triplet loss]]></category>
		<category><![CDATA[Vision Transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=234462</guid>

					<description><![CDATA[Researchers have combined convolutional neural networks and vision transformers with a three-loss training strategy to significantly improve content-based image retrieval performance.]]></description>
										<content:encoded><![CDATA[<p>Image search is quietly becoming one of the most demanding tests of artificial intelligence. When a user submits a photo and asks a system to find visually similar images across a database of millions, the machine must do far more than match colors or shapes. It must understand what makes two pictures semantically alike, a task that has pushed researchers to combine the two most powerful architectures in modern computer vision: convolutional neural networks and vision transformers. A new study published in the International Journal of Machine Learning and Cybernetics shows that fusing these architectures, and then training the result with three different loss functions at once, can substantially sharpen the quality of image retrieval.</p>
<p>The research, led by Hanh Nguyen Thi of Thuyloi University and Hanoi Architectural University together with colleagues at CMC University, the Posts and Telecommunications Institute of Technology, the Vietnam Academy of Science and Technology, and Hanoi University of Science and Technology, introduces a framework called HCV-MLO, short for Hybrid CNN-ViT with Multi-Loss Optimization. The work extends a conference paper the team presented at MAPR 2025, and it arrives at a moment when content-based image retrieval, known in the field as CBIR, is under growing pressure from ever-larger image collections and increasingly subtle queries.</p>
<p>To understand why the hybrid approach matters, it helps to look at what each architecture does well. Convolutional neural networks, the workhorses of deep learning since the breakthrough ImageNet results of 2012, scan images with filters that detect local patterns: edges, textures, corners, and small motifs that build up into recognizable objects. They are exceptionally good at capturing these fine-grained visual signatures. But convolutions look through a narrow window at each step, so a CNN can struggle to relate a detail in the top-left corner of an image to a feature in the bottom-right, especially when that relationship spans a large distance.</p>
<p>Vision transformers, introduced by Dosovitskiy and colleagues in 2021 with the memorable paper title declaring that an image is worth sixteen by sixteen words, take a different route. They chop an image into patches, treat each patch like a word in a sentence, and use self-attention to let every patch weigh its relationship to every other patch. This gives ViTs a natural talent for modeling long-range dependencies and global context, the kind of scene-level understanding that tells a retrieval system a dog on a beach belongs with other outdoor animal scenes even if the local textures differ. The catch is that transformers, on their own, can be less sensitive to the fine local detail that CNNs capture so effortlessly.</p>
<p>The Vietnamese team&#8217;s insight was that these weaknesses are complementary, and that a retrieval system should not have to choose. HCV-MLO runs parallel CNN and ViT branches over the same input image and merges their feature representations into a single embedding, a compact numerical vector that encodes the image&#8217;s semantic content. In a retrieval system, similarity between images is computed as distance between these vectors, so the entire game is to learn embeddings in which images of the same category cluster tightly while images of different categories spread apart. The hybrid design gives the embedding both the local texture sensitivity of the convolutional branch and the global relational awareness of the transformer branch.</p>
<p>But architecture alone, the researchers argue, is only half the story. The second pillar of HCV-MLO is its multi-loss optimization strategy, which trains the network by simultaneously minimizing three distinct loss functions: triplet loss, contrastive loss, and cross-entropy loss. Each of these objectives shapes the embedding space in a different way. Triplet loss, made famous by the FaceNet face recognition system in 2015, pulls an anchor image closer to a positive example of the same class while pushing it away from a negative example, directly enforcing the relative ordering that retrieval depends on. Contrastive loss works on pairs of images, rewarding the network when similar images map to nearby points and dissimilar images map to distant ones. Cross-entropy loss, the standard objective for classification, anchors the embedding to clear category boundaries and provides a stable supervisory signal.</p>
<p>To find out what each loss actually contributes, the team built three single-loss variants of their hybrid model: CNN-ViT-CE using cross-entropy alone, CNN-ViT-Contrastive using contrastive loss alone, and CNN-ViT-Triplet using triplet loss alone. This ablation-style analysis, more comprehensive than what appeared in their earlier conference version, allowed them to isolate the effect of each objective on the learned representations. The experiments covered three widely used benchmark datasets that together span a demanding range of retrieval challenges: CIFAR-100, with its one hundred diverse object categories in small, low-resolution images; CUB-200-2011, a fine-grained dataset of two hundred bird species where distinguishing one sparrow from another requires exquisite attention to subtle detail; and Stanford Dogs, another fine-grained benchmark covering one hundred twenty dog breeds.</p>
<p>The results were consistent across all three datasets. The full HCV-MLO model outperformed both single-network baselines, meaning pure CNN or pure ViT systems, and the single-loss variants, meaning the hybrid architecture trained with only one of the three objectives. The headline figure comes from CIFAR-100, where HCV-MLO achieved a mean average precision at rank one hundred, or mAP@100, of 81.99, a clear improvement over recent competitive methods in the field. Mean average precision is a strict metric for retrieval because it rewards systems that place all the relevant results near the top of the ranked list, not merely somewhere in the first hundred. The fact that the gains held up on fine-grained datasets like birds and dogs, where inter-class differences are tiny, suggests that the combination of local and global features with multi-loss training produces embeddings that are genuinely more discriminative, not just better tuned to one benchmark.</p>
<p>The study&#8217;s broader message resonates with a trend running through recent deep metric learning research, a field that has produced a rich catalog of loss functions including proxy-based objectives, multi-similarity loss, ranked list loss, and circle loss. Rather than betting on any single recipe, HCV-MLO demonstrates that stacking complementary objectives on top of a complementary architecture yields a practical and effective solution for embedding quality. The authors also note that the datasets used in the study, CIFAR-100, CUB-200-2011, and Stanford Dogs, are all publicly available, and that no new datasets were generated, making the work straightforward for other groups to reproduce and compare against.</p>
<p>For everyday applications, the implications are tangible. Better retrieval embeddings mean medical image archives where a clinician can find prior cases resembling a new scan, e-commerce platforms where a shopper&#8217;s camera snapshot surfaces the right product, and digital libraries where a rough sketch or photo unlocks visually related material. As image collections continue to balloon, the hybrid CNN-ViT strategy with multi-loss optimization offers a blueprint for systems that see both the trees and the forest, capturing the fine texture of an individual leaf while never losing sight of the shape of the whole canopy. The research was carried out without specific grant funding, and the authors report no competing financial interests, leaving the door open for the broader community to build on their hybrid, multi-loss recipe.</p>
<p><strong>Subject of Research:</strong> A hybrid CNN-ViT feature embedding framework with multi-loss optimization for content-based image retrieval</p>
<p><strong>Article Title:</strong> Hybrid CNN-ViT feature embedding for image retrieval with multi-loss optimization</p>
<p><strong>Article References:</strong> Thi, H. N., Huu, Q. N., Thuy, Q. D. T., Le, N. T. T., Van, T. N., &amp; Huu, H. N. (2026). Hybrid CNN-ViT feature embedding for image retrieval with multi-loss optimization. <em>International Journal of Machine Learning and Cybernetics, 17</em>(9), Article 446. <a href="https://doi.org/10.1007/s13042-026-03271-6" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03271-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03271-6" rel="noopener noreferrer">10.1007/s13042-026-03271-6</a></p>
<p><strong>Keywords:</strong> content-based image retrieval, convolutional neural networks, vision transformers, hybrid architecture, triplet loss, contrastive loss, cross-entropy loss, deep metric learning, feature embedding, CIFAR-100, fine-grained retrieval, HCV-MLO</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">234462</post-id>	</item>
	</channel>
</rss>
