Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI

October 1, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI

From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI

From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Artificial intelligence has quietly crossed a threshold. The most capable systems no longer see the world through a single sense: they read text, interpret images, parse audio, and watch video, all within one unified model. A new open-access review published in Discover Informatics by Azhar A. Hadi and K. P. Supreethi of Jawaharlal Nehru Technological University Hyderabad offers the most systematic map yet of this transformation, dissecting thirteen pretrained multimodal deep learning models released between 2020 and 2025 and tracing how they are reshaping fields from radiology to wildlife monitoring.

The review, which screened more than 400 candidate papers down to 120 rigorously assessed studies, arrives at a moment when the field is accelerating faster than most surveys can track. Earlier reviews, the authors argue, tended to focus narrowly on single domains such as medicine or language processing, leaving researchers without a comparative view of the architectures, datasets, and training objectives that define the current generation of models. By covering eleven application domains, including healthcare, education, autonomous transportation, e-commerce, and industrial manufacturing, the paper aims to give both newcomers and specialists a coherent picture of what these systems can actually do.

At the technical heart of the review lies a family of architectures built on the transformer, the attention-based design that revolutionized natural language processing before conquering vision. The Vision Transformer, or ViT, broke with convolutional tradition by slicing images into fixed 16-by-16-pixel patches, flattening them into vectors, and processing them exactly like words in a sentence. This simple reframing allowed vision models to inherit the scaling behavior of language models, and it became the visual backbone for much of what followed. CLIP, developed by OpenAI, went further by training an image encoder and a text encoder jointly on 400 million image-text pairs scraped from the internet, learning to pull matching pairs together in a shared latent space while pushing mismatched ones apart. That contrastive objective gave machines a rudimentary form of grounded understanding, and CLIP now serves as the visual front end for many multimodal large language models.

Subsequent generations refined this recipe in strikingly different directions. BLIP introduced a flexible encoder-decoder design and a bootstrapping mechanism called CapFilt, which generates synthetic captions and filters noisy training pairs to improve data quality. Its successor, BLIP-2, took a radically efficient path: rather than retraining everything, it froze both a pretrained image encoder and a large language model, connecting them through a lightweight Querying Transformer that translates visual features into a form the language model can consume. DeepMind’s Flamingo achieved few-shot multimodal learning by threading visual inputs into a frozen language model through a Perceiver Resampler and gated cross-attention layers, allowing it to answer visual questions from just a handful of examples. Google’s PaLI unified multilingual text generation via mT5 with high-capacity ViT encoders, while LLaVA pioneered visual instruction tuning, using GPT-4 to synthesize training dialogues and then aligning a CLIP vision encoder with a Vicuna language decoder.

The review also charts the frontier models that have captured public attention. GPT-4V extended OpenAI’s flagship language model to accept images, achieving human-level performance on many professional benchmarks while remaining weaker than humans on certain real-world tasks. GPT-4o, released in May 2024, went further still, processing text, images, audio, and video in a single network and becoming the first large language model to perform real-time emotion recognition from video. Florence-2, built on a Dual Attention Vision Transformer and trained on the FLD-5B dataset with 5.4 billion annotations across 126 million images, handles captioning, detection, segmentation, and grounding through one prompt-based sequence-to-sequence framework. SigLIP-2 replaced the conventional softmax with a sigmoid loss, decoupling batch size from training efficacy and supporting batches of up to one million examples while maintaining strong performance in more than 100 languages.

Efficiency has emerged as the defining engineering challenge, and the review documents how the newest models answer it. PaliGemma 2 pairs a SigLIP vision encoder with Gemma language models in sizes from 3 billion to 28 billion parameters, trained in three stages that progressively raise image resolution. Gemma 3 interleaves a single global attention layer with every five local layers using a restricted 1024-token sliding window, taming the memory demands of its 128,000-token context. Llama 4 adopts a sparse Mixture-of-Experts design in which a routing mechanism activates only a fraction of expert networks per token, expanding effective capacity while keeping inference cheap, and stretches context to roughly 10 million tokens for document-scale reasoning. These techniques, alongside parameter-efficient fine-tuning methods such as Low-Rank Adaptation, form what the authors call a crucial pathway to deployable multimodal AI.

The application evidence assembled in the review is striking in its breadth. In healthcare, fine-tuned combinations of vision transformers and language models reach up to 98.5 percent accuracy on lung disease diagnosis, while the PaliGemma-CXR system interprets tuberculosis chest X-rays with 90.32 percent accuracy. ChatIOS, which couples 3D point-cloud encoders with GPT-4V, achieved 93 percent intersection-over-union in automatic tooth segmentation from intraoral scans. Yet the same section catalogues sobering failures: GPT-4V identified anatomy correctly in 87.1 percent of radiology cases but pathology in only 35.2 percent, and diagnostic hallucination rates exceeding 40 percent were reported in some evaluations. In transportation, a fine-tuned PaliGemma model read license plates with 97.66 percent character accuracy, while GPT-4o managed 67 percent on pedestrian behavior prediction. In manufacturing, a CLIP-based defect classifier hit 99.9 percent AUROC with 6.6-millisecond inference, fast enough for real-time production lines.

Against these successes, the review is candid about the field’s structural weaknesses. Data remains the first bottleneck: many studies rely on small, single-center, or single-vendor datasets, annotations from lone experts, and corpora so noisy that models learn spurious correlations, the authors’ memorable example being polar bears on ice. Computational cost is the second, with models scaling to 120 billion parameters and demanding specialized hardware that restricts access for smaller research groups. Reliability is the third: models hallucinate plausible but incorrect findings, struggle with composite figures and micro-expressions, misclassify small pedestrians at low pixel density, and remain vulnerable to adversarial attacks such as projected gradient descent. Ethical concerns compound the technical ones, since multimodal datasets encode societal biases that models can perpetuate in sensitive domains like healthcare and education, and training on medical images raises privacy stakes that demand techniques such as federated learning and differential privacy.

The authors close with a research agenda that reads as a diagnosis of the field’s growing pains. They call for multi-center, human-annotated datasets that pair imaging with genetic and clinical data; for encoder refinements that can handle long audio and raw high-resolution images directly; for larger but memory-efficient transformers; and for explainability to be treated as a core evaluation requirement rather than an afterthought. Techniques like model distillation, which compresses large multimodal systems into deployable smaller versions, and Mixture-of-Experts scaling are highlighted as the most promising routes to accessible systems. What emerges from the 120 studies is a field in transition: the architectural foundations are largely settled, the benchmarks are impressive, but trust, efficiency, and generalization to the messy diversity of real-world data remain the unfinished work. For researchers deciding where to invest the next five years, this review functions as both a map and a warning.

Subject of Research: Pretrained multimodal deep learning models, their architectures, applications, and future research directions

Article Title: A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions

Article References: Hadi, A. A., & Supreethi, K. P. (2026). A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions. Discover Informatics, 1(1), Article 21. https://doi.org/10.1007/s44564-026-00022-1

Image Credits: AI Generated

DOI: 10.1007/s44564-026-00022-1

Keywords: multimodal deep learning, CLIP, vision transformers, GPT-4o, BLIP, Flamingo, PaliGemma, Llama 4, vision-language models, healthcare AI, Mixture-of-Experts, model hallucination

Cite Scienmag News

Denise Maddox. (October 1, 2026). From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI. Scienmag. https://scienmag.com/from-clip-to-llama-4-a-sweeping-review-maps-the-rise-of-pretrained-multimodal-ai/

Denise Maddox. "From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI." Scienmag, 1 October 2026, https://scienmag.com/from-clip-to-llama-4-a-sweeping-review-maps-the-rise-of-pretrained-multimodal-ai/. Accessed 1 October 2026.

Denise Maddox. "From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI." Scienmag. October 1, 2026. https://scienmag.com/from-clip-to-llama-4-a-sweeping-review-maps-the-rise-of-pretrained-multimodal-ai/

Tags: advancements from CLIP to Llama 4AI applications in wildlife monitoringAI in autonomous transportationAI in industrial manufacturingBLIPCLIPcross-modal data interpretationdeep learning models in healthcareFlamingoGPT-4ohealthcare AIinterdisciplinary AI model analysisLlama 4Mixture of Expertsmodel hallucinationmultimodal deep learningmultimodal model architecturesmultimodal training datasets and objectivesopen-access AI research reviewPaliGemmaPretrained multimodal AIsystematic review of multimodal modelsVision Transformersvision-language models
Share26Tweet16
Previous Post

Bowel Dysfunction After Rectal Cancer Surgery Splits Into Three Distinct Symptom Domains, Study Finds

Next Post

Carbon Monoxide Breath Test Outperforms Blood Test in Detecting Newborn Hemolysis, Commentary Argues

Related Posts

Liquid Metals Could Finally End the Trade-Off Between Stretchy Circuits and Real Performance
Technology and Engineering

Liquid Metals Could Finally End the Trade-Off Between Stretchy Circuits and Real Performance

October 1, 2026
Robots Learn When You Feel Unsafe: New Framework Tunes Speed and Distance in Real Time
Technology and Engineering

Robots Learn When You Feel Unsafe: New Framework Tunes Speed and Distance in Real Time

October 1, 2026
AI Alone Won’t Fix Perovskite Solar Cells, Landmark Review Warns
Technology and Engineering

AI Alone Won’t Fix Perovskite Solar Cells, Landmark Review Warns

October 1, 2026
Painkillers in Preterm Infants: Do Acetaminophen and NSAIDs Raise the Risk of Chronic Lung Disease?
Technology and Engineering

Painkillers in Preterm Infants: Do Acetaminophen and NSAIDs Raise the Risk of Chronic Lung Disease?

October 1, 2026
Coal Waste Gets a Second Life as Cement Replacement in Self-Compacting Concrete
Technology and Engineering

Coal Waste Gets a Second Life as Cement Replacement in Self-Compacting Concrete

October 1, 2026
Ensemble AI Nearly Perfectly Spots Cause and Effect in Sentences
Technology and Engineering

Ensemble AI Nearly Perfectly Spots Cause and Effect in Sentences

October 1, 2026
Next Post
Carbon Monoxide Breath Test Outperforms Blood Test in Detecting Newborn Hemolysis, Commentary Argues

Carbon Monoxide Breath Test Outperforms Blood Test in Detecting Newborn Hemolysis, Commentary Argues

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Carbon Monoxide Breath Test Outperforms Blood Test in Detecting Newborn Hemolysis, Commentary Argues
  • From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI
  • Bowel Dysfunction After Rectal Cancer Surgery Splits Into Three Distinct Symptom Domains, Study Finds
  • Liquid Metals Could Finally End the Trade-Off Between Stretchy Circuits and Real Performance

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading