<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>T5 &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/t5/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 20:40:45 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>T5 &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Adaptive LoRA Ranks Help AI Models Learn New Tasks Without Forgetting Old Ones</title>
		<link>https://scienmag.com/adaptive-lora-ranks-help-ai-models-learn-new-tasks-without-forgetting-old-ones/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 20:40:45 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Adaptive LoRA ranks]]></category>
		<category><![CDATA[adaptive ranks]]></category>
		<category><![CDATA[avoiding knowledge overwriting in AI models]]></category>
		<category><![CDATA[catastrophic forgetting]]></category>
		<category><![CDATA[catastrophic forgetting mitigation in large language models]]></category>
		<category><![CDATA[continual learning]]></category>
		<category><![CDATA[dynamic rank allocation in neural network layers]]></category>
		<category><![CDATA[fine-tuning neural networks without losing prior knowledge]]></category>
		<category><![CDATA[GLUE]]></category>
		<category><![CDATA[GSM8K]]></category>
		<category><![CDATA[incremental learning in language models]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[low-rank adaptation (LoRA) methodology]]></category>
		<category><![CDATA[low-rank parameter adjustment]]></category>
		<category><![CDATA[MATH]]></category>
		<category><![CDATA[memory-efficient AI training]]></category>
		<category><![CDATA[parameter-efficient fine-tuning]]></category>
		<category><![CDATA[parameter-efficient fine-tuning techniques]]></category>
		<category><![CDATA[Penn State University AI research]]></category>
		<category><![CDATA[sequential task learning in AI]]></category>
		<category><![CDATA[stability-plasticity]]></category>
		<category><![CDATA[subspace orthogonality]]></category>
		<category><![CDATA[T5]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198420</guid>

					<description><![CDATA[Penn State researchers have developed a continual learning method that adaptively adjusts LoRA ranks layer by layer, helping large language models acquire new tasks without catastrophically forgetting old ones.]]></description>
										<content:encoded><![CDATA[<p>Large language models have become astonishingly capable learners, absorbing new skills from relatively small amounts of fine-tuning data. Yet this flexibility comes with a well-known curse: when an AI model is trained sequentially on new tasks, it tends to overwrite the knowledge it acquired earlier, a phenomenon researchers call catastrophic forgetting. Now, a pair of computer scientists at Pennsylvania State University has proposed an elegant fix that works by changing a surprisingly small detail of how models are fine-tuned — the number of low-rank parameters, or rank, allotted to each layer of the network. Their method, described in the journal Machine Learning, adaptively adjusts these ranks as new tasks arrive, allowing language models to keep learning without erasing what came before.</p>
<p>The research, conducted by Fuli Qiao and Mehrdad Mahdavi of the Department of Computer Science and Engineering at Penn State, builds on Low-Rank Adaptation, or LoRA, one of the most widely used parameter-efficient fine-tuning techniques in modern artificial intelligence. Instead of retraining the billions of weights inside a large language model, LoRA freezes the original weights and injects small pairs of low-rank matrices into each layer. Only these small matrices are updated during training, which slashes memory and computation costs by orders of magnitude. The catch, the authors note, is that conventional LoRA fixes the rank to the same value across every layer and every task, a one-size-fits-all choice that leaves a crucial question unexplored: how much capacity does each layer actually need for each new task?</p>
<p>Qiao and Mahdavi&#8217;s answer is a method that treats rank as a dynamic resource rather than a fixed hyperparameter. Their approach, which the team calls CL-Rank, relies on a subspace similarity metric to measure how orthogonal — that is, how non-overlapping — the low-rank subspaces occupied by different tasks are within a given layer. When a new task arrives, the algorithm evaluates the geometry of the learned subspaces and adaptively increases the layerwise rank for the new task where overlap threatens to interfere with previously learned representations. By steering new knowledge into directions of parameter space that are roughly orthogonal to old knowledge, the method minimizes interference while still giving each new task enough expressive capacity to generalize well.</p>
<p>The underlying intuition echoes a classical idea from neuroscience and machine learning known as the stability–plasticity dilemma. A learning system must be plastic enough to absorb new information but stable enough to retain old skills. Since the pioneering work on catastrophic interference in connectionist networks in the late 1980s and the advent of modern continual learning benchmarks, researchers have tried an arsenal of remedies: replaying stored examples, regularizing important weights, growing new network branches, or constraining gradient updates to null spaces of prior tasks. Many of these approaches are either memory-hungry, brittle to task order, or impractical for models with billions of parameters. What distinguishes the new work is that it tackles the problem entirely within the cheap, low-rank adapter framework, requiring no access to old data and no freezing of the base model.</p>
<p>Technically, the method tracks the principal subspaces spanned by the low-rank updates in each layer and computes a similarity score between the subspace of an incoming task and those of earlier tasks. If the new task&#8217;s gradient updates are projected into directions that overlap heavily with prior subspaces, the rank of the adapter in that layer is expanded, giving the optimizer room to find solutions that spare old representations. Layers whose subspaces remain naturally disjoint need no expansion, so the parameter budget is spent only where it matters. The paper&#8217;s supplementary analyses reveal that the learned rank distributions differ markedly across layers and modules — for example, the variation in ranks among the value projection modules of encoder layers is larger than among query projections, while encoder layers as a whole show more consistent rank allocations than decoder layers. This suggests that different parts of a transformer genuinely serve distinct roles, and that a uniform rank silently wastes capacity in some layers while starving others.</p>
<p>To test the idea, the researchers ran experiments on T5 and Llama-2-7b language models, using GPU servers with DeepSpeed for efficient training. For T5 experiments they employed four NVIDIA A6000 GPUs with a learning rate of 1e-3 and a batch size of 32, while the larger Llama-2-7b runs used four NVIDIA A100 GPUs at a learning rate of 1e-4. The evaluation spanned standard natural language processing continual learning benchmarks built from 15 datasets, including the classic text classification benchmark of Zhang and colleagues, the GLUE and SuperGLUE suites, and the IMDB movie review corpus, with six different task sequence orders to control for ordering effects. The team also pushed into harder territory with challenging mathematical reasoning benchmarks, GSM8K and MATH, each split into sequential tasks and tested in both orders to probe how task sequencing shapes forgetting.</p>
<p>The results show that the adaptive-rank method matches or beats strong baselines on three fronts at once: it mitigates forgetting of earlier tasks, improves accuracy on the tasks being learned, and preserves strong generalization to unseen data — all in a memory-efficient manner. On the math benchmarks, where sequential training on GSM8K and MATH typically causes steep declines in earlier-task accuracy, the method with orthogonal projection consistently outperformed the O-LoRA baseline in final average testing accuracy, forgetting less on the first task while generalizing better on the second. Notably, the study observes that fully fine-tuning a T5-large model on MATH yields only about 3.0 percent accuracy, and about 4.2 percent on GSM8K, consistent with prior work — underscoring how difficult these mathematical tasks are and how much room remains for continual learning approaches that squeeze more capability out of small adapters.</p>
<p>Why should this matter beyond the benchmark tables? Continual learning is arguably the missing ingredient between today&#8217;s static AI assistants and the adaptive, lifelong learning systems that both researchers and companies envision. Every time a deployed model needs a new skill — a new language, a new domain, a new tool — retraining it from scratch or even fully fine-tuning it is prohibitively expensive. Adapter-based continual learning offers a path to modular, on-demand skill acquisition, but only if adding skills does not corrupt existing ones. By showing that the humble LoRA rank, tuned per layer and per task, is a powerful lever for balancing stability and plasticity, the Penn State work gives practitioners a new, inexpensive knob to turn. Because the method requires no stored examples of past tasks, it also sidesteps the privacy and storage concerns that plague replay-based approaches.</p>
<p>The findings also carry a broader scientific message: the internal geography of large models is far from uniform, and adaptation schemes that respect that heterogeneity can outperform uniform ones. The observation that encoder and decoder layers, and even the query and value modules within them, prefer different rank allocations suggests that future parameter-efficient methods may benefit from treating adapters as structured, layer-aware components rather than interchangeable plug-ins. The authors have released their code publicly, and all datasets used in the experiments are openly available, making the results straightforward for other groups to verify and extend. The work was partially supported by the National Science Foundation under award EFMA-2318101, with experiments designed and conducted by Fuli Qiao under the supervision of Mehrdad Mahdavi.</p>
<p>As large language models continue to spread through science, medicine, and industry, the question of how to teach them new things without breaking old ones is shifting from academic curiosity to engineering necessity. This study does not solve continual learning outright — no single method does — but it demonstrates that a careful, geometry-aware allocation of low-rank capacity can deliver competitive anti-forgetting performance at a fraction of the cost of full retraining. In a field where progress often comes from scaling up, it is a reminder that sometimes the most powerful improvements come from scaling the right things in the right places, one layer at a time.</p>
<p><strong>Subject of Research:</strong> Adaptive layerwise LoRA ranks for continual learning in large language models</p>
<p><strong>Article Title:</strong> Learning Without Forgetting for Continual Learning in LLMs Through Adaptive LoRA Ranks</p>
<p><strong>Article References:</strong> Qiao, F., &amp; Mahdavi, M. (2026). Learning Without Forgetting for Continual Learning in LLMs Through Adaptive LoRA Ranks. <em>Machine Learning, 115</em>(9), Article 212. <a href="https://doi.org/10.1007/s10994-026-07145-4" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07145-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07145-4" rel="noopener noreferrer">10.1007/s10994-026-07145-4</a></p>
<p><strong>Keywords:</strong> continual learning, large language models, LoRA, catastrophic forgetting, parameter-efficient fine-tuning, adaptive ranks, subspace orthogonality, stability-plasticity, GLUE, GSM8K, MATH, T5</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198420</post-id>	</item>
		<item>
		<title>New AI Model Supercharges Multi-Domain Dialogue Tracking With Dynamic Knowledge Fusion</title>
		<link>https://scienmag.com/new-ai-model-supercharges-multi-domain-dialogue-tracking-with-dynamic-knowledge-fusion/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:05:35 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[dialogue state tracking]]></category>
		<category><![CDATA[dynamic knowledge fusion]]></category>
		<category><![CDATA[dynamic knowledge fusion in AI]]></category>
		<category><![CDATA[handling sprawling dialogue histories]]></category>
		<category><![CDATA[hotel and restaurant booking AI]]></category>
		<category><![CDATA[improving dialogue accuracy and generalization]]></category>
		<category><![CDATA[innovative AI frameworks for dialogue understanding]]></category>
		<category><![CDATA[knowledge-augmented dialogue]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[low-resource dialogue datasets]]></category>
		<category><![CDATA[multi-domain dialogue modeling]]></category>
		<category><![CDATA[Multi-domain dialogue state tracking]]></category>
		<category><![CDATA[MultiWOZ]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[open-access AI research in dialogue systems]]></category>
		<category><![CDATA[prompt learning]]></category>
		<category><![CDATA[RoBERTa]]></category>
		<category><![CDATA[slot selection]]></category>
		<category><![CDATA[T5]]></category>
		<category><![CDATA[task-oriented conversational AI]]></category>
		<category><![CDATA[task-oriented dialogue systems]]></category>
		<category><![CDATA[taxi-hailing dialogue systems]]></category>
		<category><![CDATA[virtual assistant conversation management]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196115</guid>

					<description><![CDATA[Researchers have developed DKF-DST, a two-stage model that uses contrastive slot selection and dynamic knowledge fusion to significantly improve multi-domain dialogue state tracking accuracy on the MultiWOZ benchmark.]]></description>
										<content:encoded><![CDATA[<p>A new artificial intelligence framework promises to make virtual assistants dramatically better at following conversations that leap between booking hotels, reserving restaurants, and hailing taxis — all within a single exchange. Researchers have unveiled a model called DKF-DST, short for Dynamic Knowledge Fusion for Multi-Domain Dialogue State Tracking, which tackles two stubborn problems that have long hampered task-oriented dialogue systems: the difficulty of modeling sprawling dialogue histories and the scarcity of high-quality annotated data. Published in the open-access journal Vicinagearth, the work demonstrates measurable gains in accuracy and generalization across the field&#8217;s most demanding benchmarks.</p>
<p>Dialogue state tracking, or DST, is the quiet engine inside every competent virtual assistant. Each time a user speaks, the system must record and update a running snapshot of that user&#8217;s goals — which domain they care about, which slots such as location, price range, or date they have filled, and which values those slots hold. When a single conversation weaves together hotels, flights, and dinner reservations, the state can balloon into dozens of interdependent slot-value pairs that must be maintained flawlessly across many turns. Any slip in this internal bookkeeping propagates directly into poor responses, frustrated users, and failed tasks.</p>
<p>The research team, drawn from Xinjiang University and the Institute of Artificial Intelligence (TeleAI) at China Telecom, argues that existing approaches to multi-domain DST suffer from a fundamental tension. One family of methods encodes schema and ontology knowledge — the structured catalogs of domains, slots, and permissible values — directly into the model, but this scales poorly as the number of domains grows. Another reformulates state tracking as a question-answering problem, querying each slot one at a time, which multiplies computational cost. A third strategy simply concatenates every slot and slot-value pair into the input, a brute-force tactic the authors warn causes &#8220;attention dilution,&#8221; drowning the model in irrelevant signals and degrading its ability to focus on what matters.</p>
<p>DKF-DST resolves this tension with a two-stage architecture that decides, dynamically and on every turn, which knowledge is actually worth bringing into the conversation. In the first stage, an encoder-only network built on RoBERTa — a robustly optimized variant of the BERT transformer pre-trained on massive web-scale text — reads the dialogue history and a list of candidate slots, then scores how relevant each slot is to the ongoing exchange. Rather than relying on lexical overlap metrics like TF-IDF or BM25, which the authors show are unreliable when a word such as &#8220;cheap&#8221; could legitimately populate either a hotel price-range slot or a restaurant price-range slot, the model learns to align representations through contrastive learning.</p>
<p>The contrastive training objective is elegantly simple. During training, the encoder minimizes a binary cross-entropy loss that pulls the vector representation of a dialogue history closer to the representations of its genuinely relevant slots — those with non-empty values in the ground-truth state — while pushing it away from irrelevant ones. Relevance is measured as the dot product between the first-token representations of the dialogue and each slot. At inference time, a hyperparameter threshold, set to 0.8 after systematic experimentation, filters the slot list down to only those the model is confident the user is actively pursuing. This precision-first filtering strategy deliberately tolerates a small number of missed slots in exchange for keeping irrelevant knowledge out of the pipeline, a trade-off the ablation experiments show pays off handsomely.</p>
<p>The second stage is where the &#8220;dynamic fusion&#8221; earns its name. The selected slots are transformed into a natural-language output template — for example, if the model has flagged the taxi-departure and taxi-destination slots, the prompt becomes &#8220;The user is looking for a taxi from [0] to [1]&#8221; — and the corresponding ontology candidates, drawn from the dataset&#8217;s schema, are appended after each masked position. This filled template, together with the complete tagged dialogue history distinguishing user utterances from system responses, is fed to T5, a large pre-trained sequence-to-sequence model that treats every natural language processing task as text-to-text transformation. T5 generates a fluent summary of the dialogue state, and the final structured state is recovered by reversing the template.</p>
<p>Because only the slots identified in stage one enter stage two, the model sidesteps the input bloat that plagues competing methods. The comparison against D3ST, a state-of-the-art description-driven baseline that incorporates all slot information indiscriminately, is particularly instructive. By pruning the input before fusion, DKF-DST shortens sequences, sharpens attention, and still outperforms D3ST even though any errors made in slot selection could theoretically propagate downstream — evidence, the authors argue, of the framework&#8217;s robustness and stability under realistic conditions.</p>
<p>The evaluation was conducted on MultiWOZ, the de facto standard benchmark for multi-domain dialogue research, containing more than ten thousand human-to-human dialogues spanning seven domains: restaurant, hotel, attraction, taxi, hospital, police, and train. Because crowdsourced annotation noise has long plagued the corpus, the team tested on corrected versions 2.1 through 2.4, each of which repaired substantial fractions of erroneous state labels. Performance was measured with Joint Goal Accuracy, a demanding metric requiring every slot in a turn&#8217;s predicted state to exactly match the reference, and Slot Accuracy, which grades individual slot predictions. Against baselines including TransformerDST, SOM-DST, TripPy, SAVN, SimpleTOD, and Seq2seq-DU, DKF-DST posted the strongest results among sequence-to-sequence multi-domain trackers.</p>
<p>Ablation studies reinforced just how much weight rests on the prompt design. Strip away the prompt entirely and the model flounders, unable to generate coherent states. Remove the output template and the model loses its behavioral guidance; remove the candidate values and it loses the constrained answer space that anchors predictions to valid ontology entries. Both components proved vital, confirming that the fusion of structured domain knowledge — injected selectively and dynamically — is the mechanism driving the gains, not merely the size of the underlying language model.</p>
<p>The implications reach well beyond benchmark leaderboards. As conversational agents move into clinical consultation platforms, government services, and enterprise customer support, the ability to track user intent reliably across domain boundaries — with limited annotated data — becomes a deployment bottleneck. By pairing contrastive slot selection with prompt-based knowledge injection, DKF-DST offers a template for building assistants that generalize across tasks without requiring exhaustive retraining for every new domain, a step toward dialogue systems that feel less like brittle scripts and more like genuinely capable interlocutors.</p>
<p><strong>Subject of Research:</strong> Multi-domain dialogue state tracking using dynamic knowledge fusion and contrastive learning for task-oriented dialogue systems</p>
<p><strong>Article Title:</strong> Multi-domain dialogue state tracking based on dynamic knowledge fusion</p>
<p><strong>Article References:</strong> Su, H., Fang, R., Jiang, L., Huang, X., &amp; Song, S. (2026). Multi-domain dialogue state tracking based on dynamic knowledge fusion. <em>Vicinagearth, 3</em>(1), Article 6. <a href="https://doi.org/10.1007/s44336-026-00037-0" rel="noopener noreferrer">https://doi.org/10.1007/s44336-026-00037-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-026-00037-0" rel="noopener noreferrer">10.1007/s44336-026-00037-0</a></p>
<p><strong>Keywords:</strong> dialogue state tracking, dynamic knowledge fusion, contrastive learning, task-oriented dialogue systems, MultiWOZ, RoBERTa, T5, natural language processing, large language models, slot selection, knowledge-augmented dialogue, prompt learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196115</post-id>	</item>
	</channel>
</rss>
