Sunday, October 4, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy

October 4, 2026
in Technology and Engineering
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 5 mins read
0
Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy

Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy

Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Image search is quietly becoming one of the most demanding tests of artificial intelligence. When a user submits a photo and asks a system to find visually similar images across a database of millions, the machine must do far more than match colors or shapes. It must understand what makes two pictures semantically alike, a task that has pushed researchers to combine the two most powerful architectures in modern computer vision: convolutional neural networks and vision transformers. A new study published in the International Journal of Machine Learning and Cybernetics shows that fusing these architectures, and then training the result with three different loss functions at once, can substantially sharpen the quality of image retrieval.

The research, led by Hanh Nguyen Thi of Thuyloi University and Hanoi Architectural University together with colleagues at CMC University, the Posts and Telecommunications Institute of Technology, the Vietnam Academy of Science and Technology, and Hanoi University of Science and Technology, introduces a framework called HCV-MLO, short for Hybrid CNN-ViT with Multi-Loss Optimization. The work extends a conference paper the team presented at MAPR 2025, and it arrives at a moment when content-based image retrieval, known in the field as CBIR, is under growing pressure from ever-larger image collections and increasingly subtle queries.

To understand why the hybrid approach matters, it helps to look at what each architecture does well. Convolutional neural networks, the workhorses of deep learning since the breakthrough ImageNet results of 2012, scan images with filters that detect local patterns: edges, textures, corners, and small motifs that build up into recognizable objects. They are exceptionally good at capturing these fine-grained visual signatures. But convolutions look through a narrow window at each step, so a CNN can struggle to relate a detail in the top-left corner of an image to a feature in the bottom-right, especially when that relationship spans a large distance.

Vision transformers, introduced by Dosovitskiy and colleagues in 2021 with the memorable paper title declaring that an image is worth sixteen by sixteen words, take a different route. They chop an image into patches, treat each patch like a word in a sentence, and use self-attention to let every patch weigh its relationship to every other patch. This gives ViTs a natural talent for modeling long-range dependencies and global context, the kind of scene-level understanding that tells a retrieval system a dog on a beach belongs with other outdoor animal scenes even if the local textures differ. The catch is that transformers, on their own, can be less sensitive to the fine local detail that CNNs capture so effortlessly.

The Vietnamese team’s insight was that these weaknesses are complementary, and that a retrieval system should not have to choose. HCV-MLO runs parallel CNN and ViT branches over the same input image and merges their feature representations into a single embedding, a compact numerical vector that encodes the image’s semantic content. In a retrieval system, similarity between images is computed as distance between these vectors, so the entire game is to learn embeddings in which images of the same category cluster tightly while images of different categories spread apart. The hybrid design gives the embedding both the local texture sensitivity of the convolutional branch and the global relational awareness of the transformer branch.

But architecture alone, the researchers argue, is only half the story. The second pillar of HCV-MLO is its multi-loss optimization strategy, which trains the network by simultaneously minimizing three distinct loss functions: triplet loss, contrastive loss, and cross-entropy loss. Each of these objectives shapes the embedding space in a different way. Triplet loss, made famous by the FaceNet face recognition system in 2015, pulls an anchor image closer to a positive example of the same class while pushing it away from a negative example, directly enforcing the relative ordering that retrieval depends on. Contrastive loss works on pairs of images, rewarding the network when similar images map to nearby points and dissimilar images map to distant ones. Cross-entropy loss, the standard objective for classification, anchors the embedding to clear category boundaries and provides a stable supervisory signal.

To find out what each loss actually contributes, the team built three single-loss variants of their hybrid model: CNN-ViT-CE using cross-entropy alone, CNN-ViT-Contrastive using contrastive loss alone, and CNN-ViT-Triplet using triplet loss alone. This ablation-style analysis, more comprehensive than what appeared in their earlier conference version, allowed them to isolate the effect of each objective on the learned representations. The experiments covered three widely used benchmark datasets that together span a demanding range of retrieval challenges: CIFAR-100, with its one hundred diverse object categories in small, low-resolution images; CUB-200-2011, a fine-grained dataset of two hundred bird species where distinguishing one sparrow from another requires exquisite attention to subtle detail; and Stanford Dogs, another fine-grained benchmark covering one hundred twenty dog breeds.

The results were consistent across all three datasets. The full HCV-MLO model outperformed both single-network baselines, meaning pure CNN or pure ViT systems, and the single-loss variants, meaning the hybrid architecture trained with only one of the three objectives. The headline figure comes from CIFAR-100, where HCV-MLO achieved a mean average precision at rank one hundred, or mAP@100, of 81.99, a clear improvement over recent competitive methods in the field. Mean average precision is a strict metric for retrieval because it rewards systems that place all the relevant results near the top of the ranked list, not merely somewhere in the first hundred. The fact that the gains held up on fine-grained datasets like birds and dogs, where inter-class differences are tiny, suggests that the combination of local and global features with multi-loss training produces embeddings that are genuinely more discriminative, not just better tuned to one benchmark.

The study’s broader message resonates with a trend running through recent deep metric learning research, a field that has produced a rich catalog of loss functions including proxy-based objectives, multi-similarity loss, ranked list loss, and circle loss. Rather than betting on any single recipe, HCV-MLO demonstrates that stacking complementary objectives on top of a complementary architecture yields a practical and effective solution for embedding quality. The authors also note that the datasets used in the study, CIFAR-100, CUB-200-2011, and Stanford Dogs, are all publicly available, and that no new datasets were generated, making the work straightforward for other groups to reproduce and compare against.

For everyday applications, the implications are tangible. Better retrieval embeddings mean medical image archives where a clinician can find prior cases resembling a new scan, e-commerce platforms where a shopper’s camera snapshot surfaces the right product, and digital libraries where a rough sketch or photo unlocks visually related material. As image collections continue to balloon, the hybrid CNN-ViT strategy with multi-loss optimization offers a blueprint for systems that see both the trees and the forest, capturing the fine texture of an individual leaf while never losing sight of the shape of the whole canopy. The research was carried out without specific grant funding, and the authors report no competing financial interests, leaving the door open for the broader community to build on their hybrid, multi-loss recipe.

Subject of Research: A hybrid CNN-ViT feature embedding framework with multi-loss optimization for content-based image retrieval

Article Title: Hybrid CNN-ViT feature embedding for image retrieval with multi-loss optimization

Article References: Thi, H. N., Huu, Q. N., Thuy, Q. D. T., Le, N. T. T., Van, T. N., & Huu, H. N. (2026). Hybrid CNN-ViT feature embedding for image retrieval with multi-loss optimization. International Journal of Machine Learning and Cybernetics, 17(9), Article 446. https://doi.org/10.1007/s13042-026-03271-6

Image Credits: AI Generated

DOI: 10.1007/s13042-026-03271-6

Keywords: content-based image retrieval, convolutional neural networks, vision transformers, hybrid architecture, triplet loss, contrastive loss, cross-entropy loss, deep metric learning, feature embedding, CIFAR-100, fine-grained retrieval, HCV-MLO

Cite Scienmag News

Blake Davidson. (October 4, 2026). Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy. Scienmag. https://scienmag.com/hybrid-cnn-vit-model-with-triple-loss-boosts-image-search-accuracy/

Blake Davidson. "Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy." Scienmag, 4 October 2026, https://scienmag.com/hybrid-cnn-vit-model-with-triple-loss-boosts-image-search-accuracy/. Accessed 4 October 2026.

Blake Davidson. "Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy." Scienmag. October 4, 2026. https://scienmag.com/hybrid-cnn-vit-model-with-triple-loss-boosts-image-search-accuracy/

Tags: AI-powered image similarity matchingCIFAR-100content-based image retrievalcontent-based image retrieval advancementscontrastive lossconvolutional neural networksConvolutional Neural Networks and Vision Transformerscross-entropy lossdeep learning in image searchdeep metric learningfeature embeddingfine-grained retrievalfusion of CNN and ViT in computer visionHCV-MLOhybrid architecturehybrid CNN-vision transformer architectureimproving image search accuracy with hybrid modelsinnovative deep learning frameworks in image retrievalmachine learning models for visual searchmulti-loss optimization for image retrievalsemantic image similarity detectiontriple loss function image retrievaltriplet lossVision Transformers
Share26Tweet16
Previous Post

New AI Transformer Learns to See Interior Design Styles the Way Humans Do

Next Post

Where Patriarchy Runs Deepest, Women Face Nearly Double the Odds of Partner Violence, India-Wide Study Finds

Related Posts

New AI Transformer Learns to See Interior Design Styles the Way Humans Do
Technology and Engineering

New AI Transformer Learns to See Interior Design Styles the Way Humans Do

October 4, 2026
New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data
Technology and Engineering

New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data

October 4, 2026
BatPose Turns Two Cameras and a Laptop Into a Markerless 3D Motion Capture Lab
Technology and Engineering

BatPose Turns Two Cameras and a Laptop Into a Markerless 3D Motion Capture Lab

October 4, 2026
Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men
Technology and Engineering

Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men

October 4, 2026
New Model Captures Choked Gas Blasts and Wall Heat in Pressurized Vessel Discharge
Technology and Engineering

New Model Captures Choked Gas Blasts and Wall Heat in Pressurized Vessel Discharge

October 4, 2026
AI Readiness Linked to Higher Happiness Across Asia-Pacific Economies, Study Finds
Technology and Engineering

AI Readiness Linked to Higher Happiness Across Asia-Pacific Economies, Study Finds

October 4, 2026
Next Post
Where Patriarchy Runs Deepest, Women Face Nearly Double the Odds of Partner Violence, India-Wide Study Finds

Where Patriarchy Runs Deepest, Women Face Nearly Double the Odds of Partner Violence, India-Wide Study Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Clam Immunity Decoded: Mannose Receptor RpMR1 Shields Manila Clams From Deadly Vibrio Infection
  • Childhood Pneumococcal Vaccine Fails to Curb Adult Pneumonia Burden in South Korea, Decade-Long Database Study Finds
  • Where Patriarchy Runs Deepest, Women Face Nearly Double the Odds of Partner Violence, India-Wide Study Finds
  • Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,149 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading