<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>multimodal generation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/multimodal-generation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 17:39:24 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>multimodal generation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Discrete Diffusion Models Emerge as a Rival to Autoregressive Language Models</title>
		<link>https://scienmag.com/discrete-diffusion-models-emerge-as-a-rival-to-autoregressive-language-models/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 17:39:24 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[advances in text generation methods]]></category>
		<category><![CDATA[AI model paradigm shift]]></category>
		<category><![CDATA[autoregressive language models]]></category>
		<category><![CDATA[autoregressive models]]></category>
		<category><![CDATA[billion-parameter diffusion architectures]]></category>
		<category><![CDATA[comparison of diffusion and autoregressive models]]></category>
		<category><![CDATA[D3PM]]></category>
		<category><![CDATA[diffusion modeling in NLP]]></category>
		<category><![CDATA[discrete diffusion models]]></category>
		<category><![CDATA[DREAM]]></category>
		<category><![CDATA[exposure bias in language models]]></category>
		<category><![CDATA[generative artificial intelligence]]></category>
		<category><![CDATA[inference efficiency]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[LLaDA]]></category>
		<category><![CDATA[masked diffusion]]></category>
		<category><![CDATA[multimodal generation]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[non-sequential data generation]]></category>
		<category><![CDATA[parallel decoding]]></category>
		<category><![CDATA[sequential decoding limitations]]></category>
		<category><![CDATA[text generation]]></category>
		<category><![CDATA[token prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217590</guid>

					<description><![CDATA[A new survey charts how discrete diffusion models have evolved from theoretical curiosities into billion-parameter systems that can match or beat autoregressive language models in speed, controllability, and accuracy.]]></description>
										<content:encoded><![CDATA[<p>For nearly a decade, the story of generative artificial intelligence has been told in a single voice: the autoregressive model. From the GPT series to LLaMA, the dominant recipe has been to predict one token at a time, each word conditioned on everything that came before it. That recipe has scaled spectacularly, but it carries structural baggage. Sequential decoding is inherently slow, and the gap between how such models are trained—with clean, teacher-provided context—and how they operate at inference—conditioning on their own imperfect output—creates a well-known failure mode called exposure bias. A comprehensive new survey published in the open-access journal Vicinagearth argues that a fundamentally different paradigm, discrete diffusion modeling, has now matured to the point where it can rival, and in some benchmarks beat, similarly sized autoregressive systems.</p>
<p>The new review, authored by Yixuan Li of Xi&#8217;an Jiaotong University together with colleagues at the Institute of Artificial Intelligence at China Telecom, traces the arc of discrete diffusion models (DDMs) from early prototypes to billion-parameter architectures. The core idea is borrowed from the diffusion models that revolutionized image and video generation: instead of generating data in a single pass, learn to reverse a gradual corruption process, refining pure noise into structured output step by step. The catch for language is that text is discrete. Pixels live in a continuous space where Gaussian noise is a natural perturbation; tokens do not. Injecting continuous noise into a vocabulary of categorical symbols is meaningless, so the entire mathematical machinery of diffusion had to be rebuilt for discrete state spaces.</p>
<p>The theoretical backbone of the field is the Discrete Denoising Diffusion Probabilistic Models framework, known as D3PM, which formalizes corruption as a Markov chain over categorical states. A fixed forward process gradually degrades a clean sequence, and a parameterized reverse process learns to denoise it. The behavior of the corruption is dictated by a time-dependent transition matrix applied at each step. In the absorbing formulation, tokens either stay put or fall irreversibly into a special [MASK] state—the discrete analogue of drowning a signal in noise. In the uniform formulation, a token can be replaced by any other vocabulary item with equal probability, so the sequence drifts toward a uniform distribution. Hybrid matrices blend the two, allowing already-denoised tokens to re-enter the corruption process and thereby giving the model a mechanism for self-correction during generation.</p>
<p>More exotic transition designs push the idea further. Embedding-based matrices use pretrained token embeddings to define transition probabilities, so a word is more likely to be corrupted into a semantically similar neighbor—a structured, meaningful form of noise rather than a random scramble. Continuous-time formulations replace discrete timesteps with a transition-rate matrix whose matrix exponential has an analytic solution, sidestepping expensive numerical estimation and making sophisticated self-correction computationally feasible. Across all of these variants, the reverse process is typically trained with an x0-parameterization: a neural network predicts the distribution of the original clean sequence given a corrupted one, and the known forward dynamics then yield the required reverse transition analytically. The training objective reduces to a cross-entropy loss averaged over data, timesteps, and corruption trajectories.</p>
<p>Among these options, the absorbing state has emerged as the workhorse of high-performing systems such as LLaDA, apparently because it supplies an unambiguous denoising signal that makes learning complex linguistic dependencies easier than stochastic replacement does. Theoretical work has also begun to close a longstanding gap: while convergence guarantees for continuous diffusion are well established, discrete counterparts lagged until recent analyses derived formal bounds on the Kullback-Leibler divergence between generated and target distributions in the continuous-time setting. A parallel line of inquiry reframes noise itself through the concept of positive-incentive noise—corruption that reduces task entropy rather than merely destroying information. Variational frameworks for such noise, and extensions like Rectified Noise that inject it into the velocity field of pretrained models, have reported substantial generative gains with as little as 0.39 percent additional training parameters, hinting that informative corruption may be a key frontier.</p>
<p>Turning theory into working systems has demanded a toolbox of specialized training and inference techniques. On the training side, researchers initialize diffusion models from pretrained autoregressive or BERT-style checkpoints, sometimes following a two-stage curriculum that begins with next-token prediction before switching to the diffusion objective. Complementary masking strategies construct mutually exclusive masked sample pairs from a single data point so that every token contributes to the loss, while masking schedules and loss reweighting shape which noisy examples the model sees and how much it focuses on hard-to-predict tokens. Knowledge distillation compresses a multi-step teacher into a student that generates in far fewer steps. On the inference side, confidence-based scheduling implements an easy-first strategy: the model commits to its most confident predictions each round and defers uncertain positions to later, less noisy steps, establishing global structure before filling in details.</p>
<p>Because absorbing-state models freeze revealed tokens, remasking techniques allow generated tokens to be pushed back into the [MASK] state for revision, enabling genuine iterative self-correction. Caching mechanisms analogous to the KV-cache in autoregressive serving reuse intermediate attention computations to cut the cost of repeated forward passes, and classifier-free guidance has been adapted to steer discrete generation toward desired attributes or formats. Distillation-based fast samplers train a few-step student to mimic the score trajectory of a many-step teacher, substantially reducing inference cost while preserving quality. The payoff is a qualitative change in the latency profile: where an autoregressive model faces an O(L) sequential bottleneck for a sequence of length L, a diffusion model operates in O(T) parallel refinement steps, decoupling sequence length from the number of decoding iterations.</p>
<p>The empirical evidence is no longer anecdotal. On the Countdown benchmark, the survey reports that Dream 7B, configured with five to twenty diffusion steps, simultaneously surpasses state-of-the-art autoregressive baselines such as Qwen2.5 7B in both generation accuracy and instances processed per second. Just as significant, the number of refinement steps is a tunable dial: by modulating the computational budget devoted to denoising, users can trade quality against latency at inference time, a scaling dimension conceptually akin to the reasoning-time compute exploited by models like OpenAI o1 and DeepSeek R1. The model lineage has advanced quickly, from D3PM&#8217;s foundational framework and DiffusionBERT&#8217;s proof of concept, through reparameterized objectives and simplified masked-diffusion recipes like MDLM, to large-scale systems: LLaDA trained from scratch on a mask-predict-denoise paradigm, DREAM adapted from pretrained weights with context-adaptive noise rescheduling, and conversion approaches such as DiffuGPT and DiffuLLaMA that recycle autoregressive weights into diffusion models, dramatically lowering the training barrier.</p>
<p>The application space is widening in ways that autoregressive decoding handles awkwardly. Bidirectional context makes diffusion models natural fit for text infilling, editing, and revision, exemplified by DiffusER&#8217;s edit-based reconstruction view of generation, while sequence-to-sequence systems like DiffuSeq established viability for summarization and translation with notably diverse outputs. Diffusion-of-Thought adapts stepwise denoising to mirror a chain of thought, improving multi-step reasoning, and related work extends the paradigm to commonsense and knowledge-graph reasoning. Multimodal systems such as Muddit demonstrate that the same discrete diffusion machinery unifies vision and language generation, and ViewMask-1-to-3 maintains geometric consistency across multi-view image outputs. Even autonomous driving researchers have explored discrete diffusion for reflective vision-language-action models, exploiting gradient-free self-correction for planning.</p>
<p>None of this makes the paradigm a finished story. The survey is candid about the curse of parallel decoding: generating many tokens at once can break the conditional dependencies between them, degrading coherence unless parallelism is throttled back, which erodes the speed advantage. Iterative denoising requires multiple full forward passes per generation, and the ecosystem of optimized serving frameworks, quantization, and pruning methods that matured around autoregressive models remains largely undeveloped for diffusion. Full bidirectional attention at every step scales poorly with sequence length, and no diffusion language model has yet been pushed to the trillion-parameter frontier where the largest autoregressive systems operate. Still, the trajectory is unmistakable. By recasting text synthesis as iterative refinement in discrete space, discrete diffusion models have moved from mathematical curiosity to credible foundation for next-generation language technology—faster, more controllable, and structurally aware in ways that one-token-at-a-time decoding was never designed to be.</p>
<p><strong>Subject of Research:</strong> Discrete diffusion models for natural language processing and text generation</p>
<p><strong>Article Title:</strong> An overview of discrete diffusion models in natural language processing</p>
<p><strong>Article References:</strong> Li, Y., Yu, K., Song, S., Li, Y., &amp; He, Z. (2026). An overview of discrete diffusion models in natural language processing. <em>Vicinagearth, 3</em>(1), Article 12. <a href="https://doi.org/10.1007/s44336-026-00040-5" rel="noopener noreferrer">https://doi.org/10.1007/s44336-026-00040-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-026-00040-5" rel="noopener noreferrer">10.1007/s44336-026-00040-5</a></p>
<p><strong>Keywords:</strong> discrete diffusion models, natural language processing, large language models, autoregressive models, D3PM, LLaDA, DREAM, masked diffusion, parallel decoding, text generation, inference efficiency, multimodal generation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217590</post-id>	</item>
	</channel>
</rss>
