<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>multimodal AI models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/multimodal-ai-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 29 Jan 2026 05:53:19 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>multimodal AI models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Next-Token Prediction Powers Large Multimodal Models</title>
		<link>https://scienmag.com/next-token-prediction-powers-large-multimodal-models/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 29 Jan 2026 05:53:19 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[autoregressive prediction capabilities]]></category>
		<category><![CDATA[byte-pair encoding]]></category>
		<category><![CDATA[Emu3 architecture]]></category>
		<category><![CDATA[high-resolution visual content compression]]></category>
		<category><![CDATA[merging textual and visual inputs]]></category>
		<category><![CDATA[multimodal AI models]]></category>
		<category><![CDATA[next-token prediction]]></category>
		<category><![CDATA[Qwen tokenizer architecture]]></category>
		<category><![CDATA[SBER-MoVQGAN technology]]></category>
		<category><![CDATA[temporal residual layers in AI]]></category>
		<category><![CDATA[unified tokenizer system]]></category>
		<category><![CDATA[vector quantization visual tokenizer]]></category>
		<guid isPermaLink="false">https://scienmag.com/next-token-prediction-powers-large-multimodal-models/</guid>

					<description><![CDATA[In the realm of artificial intelligence, a groundbreaking advance is reshaping how machines comprehend and generate interconnected sensory data. Researchers have unveiled Emu3, a next-generation multimodal model, designed to bridge language, vision, and action seamlessly through next-token prediction. This architecture heralds a paradigm shift, merging textual and visual inputs into one cohesive framework that promises [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In the realm of artificial intelligence, a groundbreaking advance is reshaping how machines comprehend and generate interconnected sensory data. Researchers have unveiled Emu3, a next-generation multimodal model, designed to bridge language, vision, and action seamlessly through next-token prediction. This architecture heralds a paradigm shift, merging textual and visual inputs into one cohesive framework that promises to redefine machine learning’s boundaries.</p>
<p>At the heart of Emu3 lies a unified tokenizer system that translates words, images, and videos into a shared discrete token space. The system uses byte-pair encoding (BPE) for textual data, while a vector quantization (VQ)-based visual tokenizer compresses immense high-resolution visual content into manageable token sequences. This shared vocabulary enables the model to treat textual and visual information equivalently, enhancing its autoregressive prediction capabilities without loss of modality-specific nuances.</p>
<p>The textual tokenizer is built on the Qwen tokenizer architecture, leveraging byte-level BPE and encompassing over 151,000 regular text tokens supplemented by 211 special tokens reserved for template control. In parallel, the visual tokenizer operates atop SBER-MoVQGAN technology, encoding 512&#215;512 pixel images or equivalent video clips into 4,096 discrete tokens from a vast codebook containing 32,768 entries. Innovations such as temporal residual layers with 3D convolution kernels allow sophisticated temporal and spatial downsampling, enabling the model to process diverse video resolutions and durations efficiently.</p>
<p>Emu3’s core structure employs a decoder-only Transformer design, containing approximately 8.5 billion parameters distributed across 32 layers. Using RMSNorm for normalization and advanced attention mechanisms such as GQA, combined with SwiGLU activation and rotary positional embeddings, the architecture is specifically optimized to integrate vision and language representations into one harmonized stream. The shared multimodal vocabulary facilitates consistent interpretation of input across different sensory domains, creating a robust foundation for complex multimodal reasoning.</p>
<p>To rigorously evaluate model design choices, Emu3 was pitted against leading architectures, including diffusion models and encoder-plus-large-language-model (LLM) composites. When compared with a diffusion transformer trained on massive datasets, Emu3’s next-token prediction model demonstrated superior convergence speed, challenging assumptions that diffusion paradigms inherently excel in visual generation. Similarly, in head-to-head tests against late-fusion LLaVA-style models—both with and without pretraining—Emu3 matched or exceeded performance, confirming that decoder-only models trained from scratch can rival hybrid architectures without reliance on pretrained visual encoders.</p>
<p>Training such a versatile system demanded innovative data curation and scheduling. Emu3 was pretrained from scratch on a carefully constructed mixture of language, image, and video data curated for quality and diversity. Resizing of visual inputs adhered to consistent aspect ratios and spatial scales near 512&#215;512 pixels, ensuring uniformity during tokenization. A dedicated curriculum spanning three stages was utilized: initial rapid convergence with longer sequences and no dropout, followed by a stability-focused phase introducing dropout regularization, and finally an expansion to ultra-long sequences accommodating video data alongside text.</p>
<p>Post-training strategies further refined Emu3’s capabilities. For text-to-image (T2I) generation, the model underwent quality-focused fine-tuning using human preference scores aggregated from multiple evaluative metrics, enabling it to produce sharper and more aesthetically appealing high-resolution images. Additionally, Direct Preference Optimization (DPO) was deployed to align outputs more closely with human taste, involving annotator-guided ranking of generated images and iterative fine-tuning against chosen preferences, balancing fidelity and diversity.</p>
<p>Extending beyond still images, Emu3 was scaled to generate coherent and temporally consistent video sequences. Video fine-tuning incorporated stringent quality and motion filters on curated five-second clips, with sequence lengths exceeding 130,000 tokens. The model’s video outputs were quantitatively assessed across 16 key dimensions, including semantic fidelity and subject-background coherence, demonstrating landmark progress in controllable and believable video synthesis.</p>
<p>Vision–language understanding tasks also showcased Emu3’s versatility. A two-stage post-training regimen first integrated large-scale image-text pair data with masked vision token losses to focus on text prediction, followed by instruction tuning on millions of question-answer pairs relating to visual inputs. This multi-modal fine-tuning enhanced performance on image understanding benchmarks, highlighting the model’s capacity for reasoning and dialogue grounded in visual context.</p>
<p>Innovatively, Emu3 supports interleaved image-text generation, where structured text instructions are augmented with inline illustrative images within a single coherent output stream. Fine-tuning for this complex format reinforces the model’s potential for producing explanatory content combining modalities naturally—foundational for applications demanding rich, multimodal communication such as educational tools and interactive storytelling.</p>
<p>Going even further, the model has been adapted for vision–language–action tasks relevant to robotics. By fine-tuning on the CALVIN benchmark, which simulates long-horizon, language-conditioned robot manipulation, Emu3 ingests visual observations and discrete action tokens alternately, predicting sequences of perception and actions accurately. This integration underscores the model’s capacity to serve as a unified controller for agents requiring continuous interpretation and decision-making across sensory and motor domains.</p>
<p>Though real-world deployment on physical robots remains a future goal, Emu3’s ability to model complex interleaved perception–action sequences without bespoke modules signals a paradigm shift. Its autoregressive formulation naturally supports conditioning on arbitrarily long histories and enables recovery from partial or noisy inputs, key capabilities for robust, real-time robotic operation under uncertain sensory conditions.</p>
<p>Emu3’s development marks a milestone in multimodal AI research, demonstrating that next-token prediction frameworks can unify language, vision, and actions at scale. It challenges prior dominance of diffusion and hybrid encoder-LLM architectures, proving that single-stream Transformers trained from scratch can deliver superior efficiency and versatility. With their advances in tokenization, model design, and fine-tuning strategies, the creators set a new standard for building large-scale, unified multimodal models with broad implications across AI-assisted creativity, robotics, and interactive intelligence.</p>
<p>As the frontier of multimodal AI continues to expand, Emu3 exemplifies how integrative architectures enable more fluid and contextual machine understanding of the world. The ability to process and generate intertwined streams of text, imagery, video, and actions without switching modalities paves the path toward truly generalized artificial intelligence systems that learn, adapt, and interact seamlessly with humans and their environments.</p>
<p>Subject of Research: Multimodal learning architectures combining language, vision, and action modalities through next-token prediction.</p>
<p>Article Title: Multimodal learning with next-token prediction for large multimodal models.</p>
<p>Article References:<br />
Wang, X., Cui, Y., Wang, J. et al. Multimodal learning with next-token prediction for large multimodal models. Nature (2026). https://doi.org/10.1038/s41586-025-10041-x</p>
<p>Image Credits: AI Generated</p>
<p>DOI: https://doi.org/10.1038/s41586-025-10041-x</p>
<p>Keywords: Emu3, multimodal models, next-token prediction, unified tokenizer, vector quantization, multimodal Transformer, vision-language-action, text-to-image, text-to-video, robotic manipulation, decoder-only architecture, direct preference optimization</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">132309</post-id>	</item>
		<item>
		<title>Insilico Medicine Unveils Nach01 Foundation Model on AWS Marketplace to Accelerate Advances in Generative Chemistry</title>
		<link>https://scienmag.com/insilico-medicine-unveils-nach01-foundation-model-on-aws-marketplace-to-accelerate-advances-in-generative-chemistry/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 10 Jun 2025 20:05:52 +0000</pubDate>
				<category><![CDATA[Chemistry]]></category>
		<category><![CDATA[artificial intelligence in pharmaceuticals]]></category>
		<category><![CDATA[AWS Marketplace biotechnology]]></category>
		<category><![CDATA[cloud-based drug design]]></category>
		<category><![CDATA[drug discovery acceleration]]></category>
		<category><![CDATA[generative chemistry]]></category>
		<category><![CDATA[Insilico Medicine]]></category>
		<category><![CDATA[machine learning in drug research]]></category>
		<category><![CDATA[molecular prediction technology]]></category>
		<category><![CDATA[multimodal AI models]]></category>
		<category><![CDATA[Nach01 foundation model]]></category>
		<category><![CDATA[pharmaceutical research innovations]]></category>
		<category><![CDATA[retrosynthesis advancements]]></category>
		<guid isPermaLink="false">https://scienmag.com/insilico-medicine-unveils-nach01-foundation-model-on-aws-marketplace-to-accelerate-advances-in-generative-chemistry/</guid>

					<description><![CDATA[In a groundbreaking development that promises to accelerate drug discovery and pharmaceutical research, Insilico Medicine, a leading clinical-stage biotechnology company harnessing generative artificial intelligence, has announced the launch of its latest foundation model, Nach01, on Amazon Web Services (AWS). This significant release, available through the AWS Marketplace, marks a pivotal advancement in the integration of [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking development that promises to accelerate drug discovery and pharmaceutical research, Insilico Medicine, a leading clinical-stage biotechnology company harnessing generative artificial intelligence, has announced the launch of its latest foundation model, Nach01, on Amazon Web Services (AWS). This significant release, available through the AWS Marketplace, marks a pivotal advancement in the integration of advanced AI technologies within the drug design domain. By leveraging cloud infrastructure and cutting-edge machine learning techniques, Nach01 stands poised to transform how researchers and pharmaceutical companies approach molecular prediction and retrosynthesis, addressing complex biochemical challenges with unprecedented accuracy and scalability.</p>
<p>At its core, Nach01 represents a novel class of multimodal foundation models capable of processing and synthesizing both structural and spatial chemical data simultaneously. Traditional AI models in drug discovery have often been limited to either textual or structural datasets, but Nach01’s architecture integrates a large language model with spatial understanding powered by point cloud transformers. This fusion allows the system to interpret molecular information in a comprehensively multidimensional manner, enhancing predictive capabilities and facilitating tasks that span from molecular property inference to the generation of novel chemical compounds. Such versatility is essential in tackling the multifaceted nature of pharmaceutical research.</p>
<p>The development of Nach01 was conducted on Amazon SageMaker, AWS’s fully-managed machine learning platform that supports the entire ML lifecycle—from data preparation and model training to deployment and monitoring. The utilization of SageMaker has endowed Nach01 not only with the ability to scale efficiently across diverse computational resources but also with seamless integration options for researchers who wish to fine-tune or deploy models in customized drug discovery pipelines. This operational flexibility ensures that both academic labs and industry players—from burgeoning startups to established pharmaceutical giants—can rapidly adopt and implement the model in their workflows.</p>
<p>Insilico Medicine’s Pharma.AI platform underpins Nach01’s capabilities by incorporating deep generative models, reinforcement learning, and transformer architectures optimized for chemistry and biochemistry applications. These advanced methodologies allow Nach01 to extrapolate chemical behaviors and interactions from vast datasets, accelerating the identification of potential drug candidates that meet precise therapeutic profiles. Moreover, the model’s proficiency in handling 2D and 3D molecular data permits a more realistic simulation of molecular dynamics, a crucial advantage for anticipating drug efficacy and toxicity before clinical testing.</p>
<p>The significance of this announcement extends beyond the technical prowess of the model itself. By distributing Nach01 via AWS Marketplace, Insilico Medicine effectively democratizes access to state-of-the-art AI-driven drug design tools. Researchers worldwide can now obtain secure, scalable access to Nach01 through standard Python APIs or cloud-native deployment strategies, reducing barriers that traditionally impeded the use of sophisticated machine learning models in life sciences. This accessibility is expected to fuel innovation and collaboration across disciplines, opening new avenues for the discovery of treatments against a diverse array of diseases.</p>
<p>Alex Zhavoronkov, PhD, Founder and CEO of Insilico Medicine, emphasized the transformative potential of Nach01 in reshaping pharmaceutical research. He described the model as a “stepping stone on our path to pharmaceutical superIntelligence,” highlighting the ambition not only to enhance current drug development pipelines but also to lay the groundwork for AI systems capable of autonomous novel medicine discovery. Zhavoronkov’s vision underscores the critical role that AI will increasingly play in resolving the longstanding challenges of drug development, including high costs, lengthy timelines, and complex molecular interactions.</p>
<p>Jon Jones, Vice President and Global Head of Startups at AWS, expressed enthusiasm about the collaboration, noting AWS’s commitment to supporting cutting-edge biochemistry models like Nach01 globally. AWS’s role in providing robust infrastructure and a reliable marketplace facilitates faster dissemination of transformative AI solutions, thereby accelerating the translation of scientific breakthroughs into real-world medical advancements. Jones framed generative AI as a crucial lever in improving patient outcomes by expediting the creation of better disease treatments.</p>
<p>From a technical standpoint, Nach01’s design integrates a natural and chemical languages + point cloud transformer approach (NACH01-PC), allowing it to navigate and generate insights across diverse chemical modalities efficiently. This architecture supports a wide array of tasks ranging from retrosynthetic pathway generation—mapping out viable synthetic routes for complex molecules—to molecular property prediction, an indispensable tool for assessing the drug-likeness and potential success of molecular candidates. The ability to fine-tune the model on bespoke datasets ensures adaptability across various therapeutic domains, including oncology, neurodegenerative diseases, and immunology.</p>
<p>The model also supports both inference and fine-tuning through Python code or API calls, providing a familiar and accessible interface for computational chemists and AI specialists. By enabling deployment on SageMaker, users benefit from scalable compute resources optimized for heavy ML workloads, essential for handling the vast chemical search spaces typically encountered in drug development. Furthermore, securing access via AWS Marketplace ensures compliance with data governance and security protocols, which are paramount in handling sensitive biomedical information.</p>
<p>Pre-launch interest in Nach01 was notably high, reflecting the community’s anticipation of its potential impact. Its release is expected to catalyze a wave of research initiatives, especially among startups and research institutions looking to harness AI for accelerated molecule optimization and design. The strategic partnership between Insilico Medicine and AWS thus represents a critical nexus of AI innovation and cloud infrastructure, jointly addressing the pressing need to modernize pharmaceutical R&amp;D processes.</p>
<p>Insilico Medicine continues to champion AI-driven breakthroughs across multiple therapeutic areas, including cancer, fibrosis, central nervous system disorders, infectious diseases, autoimmune conditions, and aging-related ailments. The introduction of Nach01 on AWS amplifies these efforts by providing a scalable, production-ready AI tool tailored for the chemical and biological complexities inherent in drug design. Through platforms like Pharma.AI and now Nach01, Insilico is setting new benchmarks in integrating computational intelligence with biomedical science, ultimately accelerating the advent of novel therapies.</p>
<p>In summary, the launch of Nach01 foundation model on Amazon Web Services signifies a watershed moment in the intersection of AI and drug discovery. By merging sophisticated multimodal AI architectures, cloud scalability, and accessible deployment frameworks, Insilico Medicine and AWS are collectively enabling a new era in pharmaceutical innovation. This progress not only portends accelerated timelines from molecule design to drug development but also heralds the promise of AI systems that may one day autonomously generate lifesaving medicines with higher precision and speed than ever before.</p>
<hr />
<p><strong>Subject of Research</strong>: Multimodal Foundation Models for AI-driven Drug Discovery and Molecular Prediction</p>
<p><strong>Article Title</strong>: Insilico Medicine Unveils Nach01: A Multimodal AI Foundation Model for Drug Design on AWS</p>
<p><strong>News Publication Date</strong>: June 10, 2025</p>
<p><strong>Web References</strong>:</p>
<ul>
<li><a href="https://insilico.com/">https://insilico.com/</a>  </li>
<li><a href="https://pharma.ai/">https://pharma.ai/</a>  </li>
</ul>
<h4><strong>Keywords</strong></h4>
<p>Generative AI, Drug Design, Machine Learning, Biochemistry, Artificial Intelligence, Molecular Prediction, Retrosynthesis, Pharmaceutical AI, Computational Chemistry</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">52647</post-id>	</item>
	</channel>
</rss>
