<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>continual learning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/continual-learning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 19:17:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>continual learning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Artificial Neural Networks Evolve Brain-Like Modules When Learning Multiple Tasks</title>
		<link>https://scienmag.com/artificial-neural-networks-evolve-brain-like-modules-when-learning-multiple-tasks/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 19:17:25 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[brain networks]]></category>
		<category><![CDATA[brain-inspired neural network design]]></category>
		<category><![CDATA[brain-like architecture in artificial neural networks]]></category>
		<category><![CDATA[cognitive task learning in neural networks]]></category>
		<category><![CDATA[cognitive tasks]]></category>
		<category><![CDATA[computational demands driving neural architecture]]></category>
		<category><![CDATA[connectome]]></category>
		<category><![CDATA[connectome-inspired AI architecture]]></category>
		<category><![CDATA[continual learning]]></category>
		<category><![CDATA[emergence of brain-like modules in AI]]></category>
		<category><![CDATA[evolution of modular structures in AI]]></category>
		<category><![CDATA[Human Connectome Project]]></category>
		<category><![CDATA[long-range neural connections in artificial networks]]></category>
		<category><![CDATA[lottery ticket hypothesis]]></category>
		<category><![CDATA[modularity]]></category>
		<category><![CDATA[multitask learning]]></category>
		<category><![CDATA[Nature Machine Intelligence]]></category>
		<category><![CDATA[network neuroscience]]></category>
		<category><![CDATA[neural network development for complex tasks]]></category>
		<category><![CDATA[neural network modularity]]></category>
		<category><![CDATA[neural network training without physical wiring constraints]]></category>
		<category><![CDATA[neuroscience-inspired machine learning]]></category>
		<category><![CDATA[recurrent neural networks]]></category>
		<category><![CDATA[wiring cost]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=218474</guid>

					<description><![CDATA[Recurrent neural networks trained on demanding multitask curricula spontaneously develop modular, brain-like architectures, showing that computational demands rather than wiring costs alone can drive the emergence of modularity.]]></description>
										<content:encoded><![CDATA[<p>One of the deepest puzzles in neuroscience is why the brain is built the way it is. The human connectome is not a tangle of uniformly distributed wiring; it is a mosaic of densely connected clusters, or modules, each specializing in particular functions while exchanging information through a smaller set of long-range links. For decades, the dominant explanation has been spatial and economic: the brain lives inside a skull, wiring is metabolically expensive, and evolution has therefore favored architectures that minimize connection length. A new study published in Nature Machine Intelligence challenges the sufficiency of that account, showing that the sheer computational demands of learning complex tasks can, on their own, drive the emergence of modular structure in artificial neural networks—and that the resulting architectures resemble the brain&#8217;s more closely than any purely spatial model has managed.</p>
<p>The research, led by Yuhang Wu, Shi Gu, and colleagues at Zhejiang University, the University of Electronic Science and Technology of China, New York University, and the University of Pennsylvania, including network neuroscientist Dani S. Bassett, took a deliberately controlled approach. Rather than embedding their networks in physical space and imposing wiring-cost constraints, the team trained recurrent neural networks (RNNs) on batteries of cognitive tasks of the kind long used in systems neuroscience: working memory, perceptual decision-making, context-dependent categorization, and other tasks that probe the computational repertoire of prefrontal and parietal cortex. The question was simple but profound: if you strip away all spatial and metabolic pressure, does modularity still appear when a network must learn many demanding tasks at once?</p>
<p>The answer was a resounding yes. Networks trained under multitask learning paradigms developed significantly higher modularity than networks trained on a single task, and the effect grew stronger as the task load pushed against the network&#8217;s capacity. When the number of simultaneous tasks strained what the fixed pool of units could compute, the networks responded by reorganizing their internal connectivity into functionally segregated communities. This is a striking result because nothing in the training objective rewarded modularity directly. The networks were optimized only for task performance, yet the pressure of limited capacity and diverse demands was sufficient to carve the connectivity matrix into modules—much as the pressure of diverse cognitive demands may have shaped the brain&#8217;s own architecture.</p>
<p>The study went further by comparing different training regimes. Networks trained with incremental multitask learning—in which tasks were introduced sequentially and the network had to integrate each new demand into an already functioning system—developed the highest degree of modularity of all, while also maintaining superior performance across the full task set. This detail matters because it mirrors the developmental trajectory of biological brains, which do not acquire all cognitive abilities simultaneously but build them progressively over years of experience. The finding suggests that the order and pacing of task acquisition, not merely the total computational load, shapes the topology that emerges. Modularity, in this view, is not a static design feature but an adaptive response to the sequential introduction of complex problems.</p>
<p>Technically, the team quantified modularity using established network-science measures, including community-detection methods of the kind pioneered by Leicht and Newman for directed networks, applied to the learned weight matrices of the RNNs. They tracked how modular structure unfolded over the course of training, revealing that community boundaries sharpened as learning progressed and as additional tasks accumulated. They also examined the incremental addition of connections during learning, drawing an intriguing parallel to the lottery ticket hypothesis from deep learning research—the idea that sparse, trainable subnetworks exist within larger networks and are the components that effectively carry the computational load. In the task-trained RNNs, sparse modular substructures appeared to play an analogous role, suggesting a possible computational rationale for why both artificial and biological learning systems might favor segregated, sparsely interconnected architectures.</p>
<p>Perhaps the most consequential finding came when the researchers compared their task-induced networks against biological data. Using structural connectivity data from the Human Connectome Project, covering 84 cortical areas, they evaluated how closely the artificial networks matched the brain&#8217;s own organization. The task-trained networks exhibited structural properties that more closely resembled biological brain networks than models based solely on spatial constraints such as wiring-cost minimization. In other words, functional demand—the need to compute—appears to be a stronger organizing principle for brain-like topology than physical economy alone. This does not mean spatial constraints are irrelevant; the brain is certainly shaped by the geometry of the skull and the metabolic cost of axons. But the new results demonstrate that spatial models alone cannot fully explain the functional organization of brain networks, and that computational pressure fills a substantial part of that explanatory gap.</p>
<p>The work builds on a rich lineage of research at the intersection of machine learning and neuroscience. Previous studies had shown that RNNs trained on many cognitive tasks develop mixed selectivity and shared dynamical motifs that support flexible behavior, and that spatially embedded RNNs recapitulate numerous structural and functional findings from neuroscience. Other work demonstrated that brain-like functional specialization can emerge spontaneously in deep networks trained on naturalistic tasks. The new study adds a crucial piece: a controlled computational demonstration that modularity itself—arguably the signature feature of brain network organization—can be induced purely by the functional demands of multitask learning under capacity constraints. It thereby offers a causal, mechanistic account where earlier work offered correlations or spatial explanations.</p>
<p>The implications run in both directions. For neuroscience, the study provides a testable framework: if modularity is an adaptive response to cumulative cognitive demands, then developmental changes in brain network segregation should track the acquisition of complex abilities, a hypothesis consistent with prior findings that modular segregation of structural brain networks supports the development of executive function in youth. For artificial intelligence, the results hint at design principles for more adaptable machines. Modular deep learning has become a vibrant field precisely because modular systems can learn new skills without catastrophically forgetting old ones, and this study suggests that simply structuring the training curriculum—introducing tasks incrementally under realistic capacity limits—can coax modularity into existence without hand-engineered architectural constraints. That could inform how researchers build continual-learning systems, from robotics to large multimodal models, where flexibility and stability must coexist.</p>
<p>The study also speaks to a long-standing debate about the economy of brain network organization. The brain has often been described as a compromise between wiring cost and topological value, with small-world architecture emerging from that trade-off. The new findings suggest the ledger has more entries than previously appreciated: computational value, capacity limits, and the temporal sequence of learning demands all leave structural fingerprints. Modularity may be less a consequence of saving wire and more a consequence of solving problems—a functional adaptation that spatial economy then refines rather than creates. As the authors put it in their abstract, modularization emerges as an adaptive response to the sequential introduction of complex tasks, a framing that reframes the brain&#8217;s architecture as the product of a computational curriculum written by evolution and experience.</p>
<p>The team has released its code and processed connectivity data publicly via GitHub and Zenodo, allowing other researchers to reproduce the simulations and extend the approach to new task sets, architectures, and species comparisons. Cross-species network comparisons, representational similarity analyses, and generative models of the connectome are natural next steps, and the framework could eventually be applied to clinical questions, since many neuropsychiatric conditions involve disruptions of modular brain organization. For now, the study stands as a vivid example of a growing trend in modern science: artificial neural networks are no longer just engineering tools but instruments for asking why questions about the brain—and, increasingly, they are answering with structures that look strikingly familiar. When a machine learns the way a mind must, it begins, it seems, to build itself the way a brain is built.</p>
<p><strong>Subject of Research:</strong> Emergence of task-driven modular network organization in recurrent neural networks and its alignment with human brain architecture</p>
<p><strong>Article Title:</strong> Task-structured modularity emerges in artificial networks and aligns with brain architecture</p>
<p><strong>Article References:</strong> Wu, Y., Deng, S., Du, K., Mattar, M. G., Wu, Y., Bassett, D. S., Tang, H., Pan, G., &amp; Gu, S. (2026). Task-structured modularity emerges in artificial networks and aligns with brain architecture. <em>Nature Machine Intelligence</em>. <a href="https://doi.org/10.1038/s42256-026-01306-9" rel="noopener noreferrer">https://doi.org/10.1038/s42256-026-01306-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s42256-026-01306-9" rel="noopener noreferrer">10.1038/s42256-026-01306-9</a></p>
<p><strong>Keywords:</strong> recurrent neural networks, multitask learning, modularity, brain networks, connectome, cognitive tasks, network neuroscience, Human Connectome Project, continual learning, wiring cost, lottery ticket hypothesis, Nature Machine Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">218474</post-id>	</item>
		<item>
		<title>Random Weights Beat Trained Networks in Online Learning on Graphs</title>
		<link>https://scienmag.com/random-weights-beat-trained-networks-in-online-learning-on-graphs/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 21:04:56 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Bitcoin fraud detection]]></category>
		<category><![CDATA[breakthroughs in online graph learning]]></category>
		<category><![CDATA[catastrophic forgetting]]></category>
		<category><![CDATA[challenges in online graph data processing]]></category>
		<category><![CDATA[concept drift]]></category>
		<category><![CDATA[continual learning]]></category>
		<category><![CDATA[continual learning on graph structures]]></category>
		<category><![CDATA[experience replay]]></category>
		<category><![CDATA[fixed weight models in machine learning]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[graph representation learning]]></category>
		<category><![CDATA[impact of random weights on model performance]]></category>
		<category><![CDATA[implications for model training efficiency]]></category>
		<category><![CDATA[innovative approaches to lifelong learning in AI]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[neural network weight initialization strategies]]></category>
		<category><![CDATA[node classification]]></category>
		<category><![CDATA[online continual graph learning]]></category>
		<category><![CDATA[online learning]]></category>
		<category><![CDATA[performance comparison of trained vs. untrained models]]></category>
		<category><![CDATA[random weights neural networks]]></category>
		<category><![CDATA[randomized representations]]></category>
		<category><![CDATA[streaming graph data analysis]]></category>
		<category><![CDATA[streaming linear discriminant analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=210285</guid>

					<description><![CDATA[Researchers show that a randomly initialized, never-trained graph neural network paired with a lightweight streaming classifier outperforms state-of-the-art methods in online continual graph learning, with gains of up to 30 percent and no memory buffer.]]></description>
										<content:encoded><![CDATA[<p>Machine learning researchers have long assumed that the road to better performance runs through more training, bigger models, and increasingly sophisticated learning algorithms. A new study published in the journal Machine Learning turns that assumption on its head for one of the most demanding settings in artificial intelligence: learning continuously from streams of graph data. A team led by Giovanni Donghi of the University of Padua, together with Daniele Zambon, Luca Pasa, Cesare Alippi and Nicolò Navarin, has shown that a neural network whose weights are fixed at random values, never trained at all, can outperform state-of-the-art systems on online continual graph learning tasks, with improvements of up to 30 percent in some benchmarks. The finding, published as volume 115, article 220 of the journal, suggests that in this domain the key to remembering old knowledge may not be cleverer learning but deliberate refusal to learn at all.</p>
<p>The setting the researchers tackled, known as Online Continual Graph Learning, or OCGL, is among the harshest in machine learning. Nodes of a graph, which might represent Bitcoin transactions, scientific papers, social media posts or products on a shopping site, arrive one at a time in a continuous stream. Each new node brings its own attributes and connections to previously seen nodes, and the distribution of the data can shift at any moment, for instance when new classes of objects begin to appear. The model must make accurate predictions at every instant, adapt on the fly, and retain what it learned about earlier data, all under strict limits on memory and computation. Crucially, the model sees each node only once, so there is no possibility of the repeated offline training passes that standard deep learning relies on.</p>
<p>This streaming regime makes the phenomenon known as catastrophic forgetting especially damaging. When neural networks are updated sequentially on new data, the parameter changes that help with the new information often overwrite the representations that supported old tasks, erasing previously acquired knowledge. Graph data add a second, structural source of forgetting: because graph neural networks aggregate information from neighbors, the embedding of a node changes as new nodes attach to it, so even a perfectly stable classifier faces inputs that drift over time. Existing state-of-the-art methods typically fight forgetting with replay buffers that store examples of past data, or with regularization schemes that penalize changes to important weights, and both strategies carry substantial memory and computational costs in the graph setting.</p>
<p>The new approach, by contrast, decouples representation from prediction in a strikingly simple way. The authors split the model into two parts: a feature extractor that converts each node and its sampled neighborhood into a fixed-length embedding vector, and a lightweight classifier that maps embeddings to predictions. The feature extractor is a graph neural network whose weights are randomly initialized from an appropriate distribution and then frozen forever. Because the encoder never changes, the representations it produces cannot drift, and one of the two major sources of forgetting, the drift of backbone parameters, is eliminated by construction. Only the classifier is trained, and even that can be done without gradient descent using a streaming statistical method.</p>
<p>The researchers tested two flavors of randomized encoder. The first, called UGCN, is an untrained graph convolutional network with two layers of 1024 units each, whose layer outputs are concatenated into a 2048-dimensional node embedding that captures neighborhood information at multiple resolutions. The second, Graph Random Neural Features or GRNF, draws on theory showing that random samples from a universal family of graph neural networks can approximate kernel functions on graphs, and can provably separate any two non-isomorphic graphs. Both encoders operate on sparsified two-hop neighborhoods sampled around each node, keeping the computational footprint constant even as the graph grows denser. On top of either encoder, the team placed a Streaming Linear Discriminant Analysis classifier, which maintains running class means and a shared covariance matrix updated incrementally with each new example, requiring no memory buffer of past data at all.</p>
<p>The theoretical intuition behind the method is elegant. The authors formally decompose forgetting risk into three components: structural drift arising from the evolving graph itself, which no model can remove; backbone parameter drift, which the frozen encoder eliminates entirely; and classifier parameter drift, which for the streaming classifier shrinks as more examples of each class accumulate. Stability alone is not enough, of course, since a model must also remain plastic enough to absorb new classes and shifted distributions. That is where the over-parameterized random encoders earn their keep. Results from randomized network theory guarantee that, given a sufficiently large embedding dimension, randomly initialized features are expressive enough to make the downstream classification problem linearly separable, so a simple linear readout suffices.</p>
<p>The experimental results across seven benchmarks are remarkable. On six node-classification datasets, including the citation networks CoraFull and Arxiv, the co-purchase graph Amazon Computer, the Reddit post network, the heterophilic Roman Empire Wikipedia graph, and the Elliptic Bitcoin transaction network, the combination of randomized features and the streaming classifier generally beat every alternative considered, including experience replay, A-GEM, EWC, LwF and MAS applied to a conventionally trained graph network, as well as recent graph-specific replay methods such as PDGNN, SSM and TWP. In class-incremental streams, where new classes arrive in blocks, the approach approached the upper bound of joint offline training on the complete final graph, a ceiling that no online method can normally reach. Notably, the method needs no memory buffer, whereas replay-based competitors must store a substantial fraction of past nodes tailored to graph topology.</p>
<p>The robustness of the result is underlined by several control experiments. Even with as few as 64 random features, performance on most benchmarks matched or exceeded state-of-the-art trained methods, and gains had not saturated even at 4096 features. A frozen graph network pre-trained in a supervised way on part of the stream did not outperform the purely random encoder on most datasets, indicating that the advantage stems from the stability and richness of random representations rather than from any task-specific preparation. On the heterophilic Roman Empire dataset, where standard graph convolutions smooth neighboring features together, the more expressive GRNF encoder proved clearly superior. A hybrid extractor mixing features from both encoders delivered consistently robust performance across benchmarks. In memory terms, the fixed backbone plus streaming statistics occupied less space than the replay buffers of competing methods while sitting on the efficient frontier of the accuracy-versus-memory tradeoff.</p>
<p>The practical implications reach well beyond academic benchmarks. The Elliptic experiments, conducted on real Bitcoin transaction data with genuine timestamps in a time-incremental stream, demonstrate that the method works on realistic financial fraud detection problems where new transaction patterns emerge continuously. The same properties make the approach attractive for intrusion detection in Internet of Things networks, healthcare monitoring on temporal patient graphs, and recommendation systems, all applications the authors cite as motivation. Because the classifier updates are cheap streaming statistics and the encoder requires no training whatsoever, the method is naturally suited to deployment where latency is critical and predictions must be available at any moment. The authors caution that their conclusions are specific to the online setting, that a full theoretical analysis of forgetting remains future work, and that extensions to graph-level and edge-level tasks, regression and anomaly detection are still open. But the central message is already provocative: sometimes the most effective way for an artificial intelligence to remember is to stop changing its mind.</p>
<p><strong>Subject of Research:</strong> Randomized fixed representations for mitigating catastrophic forgetting in online continual graph learning</p>
<p><strong>Article Title:</strong> The Unreasonable Effectiveness of Randomized Representations in Online Continual Graph Learning</p>
<p><strong>Article References:</strong> The Unreasonable Effectiveness of Randomized Representations in Online Continual Graph Learning. (n.d.). <a href="https://doi.org/10.1007/s10994-026-07128-5" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07128-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07128-5" rel="noopener noreferrer">10.1007/s10994-026-07128-5</a></p>
<p><strong>Keywords:</strong> continual learning, graph neural networks, catastrophic forgetting, online learning, randomized representations, streaming linear discriminant analysis, node classification, experience replay, graph representation learning, concept drift, machine learning, Bitcoin fraud detection</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">210285</post-id>	</item>
		<item>
		<title>AI Learns to Forget: New Hypernetwork Framework Enables Data-Free Unlearning</title>
		<link>https://scienmag.com/ai-learns-to-forget-new-hypernetwork-framework-enables-data-free-unlearning/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 01:24:05 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI safety]]></category>
		<category><![CDATA[catastrophic forgetting]]></category>
		<category><![CDATA[continual learning]]></category>
		<category><![CDATA[Data Privacy]]></category>
		<category><![CDATA[hypernetworks]]></category>
		<category><![CDATA[machine unlearning]]></category>
		<category><![CDATA[membership inference]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[parameter generation]]></category>
		<category><![CDATA[ResNet]]></category>
		<category><![CDATA[right to be forgotten]]></category>
		<category><![CDATA[task embeddings]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205211</guid>

					<description><![CDATA[Researchers have developed a hypernetwork-based framework that lets continually learning AI systems erase specific tasks without access to the original data, while preventing both catastrophic forgetting and the relapse of supposedly forgotten knowledge.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence systems are increasingly being asked to do something that sounds paradoxical: forget. As regulators around the world tighten data protection rules and the public grows wary of how personal information is used to train machine learning models, researchers have been racing to develop techniques that allow a trained neural network to expunge specific knowledge without the enormous expense of retraining from scratch. A new study published in the journal Machine Learning takes a significant step toward making that possible in one of the hardest settings imaginable—continual learning, where a model must keep absorbing new tasks over time while its access to old data slips away.</p>
<p>The research, led by Sayanta Adhikari, Vishnuprasadh Kumaravelu, and P. K. Srijith of the Bayesian Reasoning and Inference Lab at the Indian Institute of Technology Hyderabad, introduces a framework called UnCLe, short for a Hypernetwork Framework for Data-Free Unlearning and Continual Learning. The core insight is that machine unlearning—the deliberate removal of a task&#8217;s influence from a trained model—has almost always been designed with offline training in mind, where engineers retain full access to the original dataset. In continual learning, that assumption collapses. Data arrives task by task and is typically discarded after use, so when an unlearning request arrives, the original examples may simply no longer exist.</p>
<p>The team identified two failure modes that emerge when conventional unlearning is naively applied to continual learning environments. The first is catastrophic forgetting of retained tasks, a well-known pathology in which updating a network to remove one capability wipes out unrelated capabilities it was supposed to keep. The second is subtler and, in some ways, more troubling: catastrophic remembering, in which tasks that were supposedly unlearned resurface when the model later absorbs new information. A model that appears to have forgotten sensitive data can effectively relapse, undermining the very privacy guarantees that unlearning is meant to provide.</p>
<p>UnCLe attacks both problems by restructuring how the model&#8217;s parameters are produced in the first place. Instead of training a single monolithic network, the framework employs a hypernetwork—a network that generates the weights of another network—conditioned on compact task embeddings. Each task the system encounters is represented by its own embedding vector, and the hypernetwork maps that embedding to a full set of task-specific parameters. Learning a new task therefore means learning or refining an embedding, while the shared hypernetwork machinery remains stable across the entire sequence of operations.</p>
<p>Unlearning under this scheme becomes elegantly simple at the task level. To remove a task, the framework optimizes the hypernetwork so that, for that task&#8217;s embedding, it generates parameters that behave like noise. The generated network produces uniform, maximum-entropy outputs on the forgotten task—in other words, the model becomes maximally uncertain, exactly as if it had never seen the task at all. Crucially, this procedure does not require the original training data, which is precisely what makes it data-free. The optimization is guided by a mean squared error objective that pulls the generated parameters toward freshly sampled Gaussian noise, combined with a regularization term that anchors the hypernetwork&#8217;s outputs for all previously retained tasks, preventing collateral damage.</p>
<p>The choice of a noise-matching objective, rather than a direct norm penalty, turns out to matter a great deal. The authors show mathematically that averaging the squared distance to random Gaussian samples converges to the squared L2 norm of the parameters plus a constant, meaning the MSE objective implicitly drives the forgotten task&#8217;s parameters toward zero and its outputs toward a uniform distribution. But applying the L2 norm directly would, over repeated unlearning operations, drag the hypernetwork&#8217;s own shared weights toward zero and destabilize the whole system. By contrast, sampling a fresh noise target at each optimization step constrains the forget task only in distribution, acting as an implicit regularizer that preserves the shared representation. The researchers compared alternatives—including fixed noise targets, pure norm reduction, and simply discarding the task embedding—and found their approach achieved the best balance between erasing the target task and protecting retained performance.</p>
<p>Scaling a hypernetwork to generate all the weights of a modern convolutional backbone such as ResNet18 or ResNet50 presents its own engineering challenge, since the hypernetwork&#8217;s output layer would otherwise balloon to an impractical size. The team&#8217;s solution is chunked generation: the main network&#8217;s parameters are partitioned into roughly 200 chunks, each produced by a dedicated head of the hypernetwork conditioned on a unique chunk embedding concatenated with the task embedding. These chunk embeddings are frozen after the first task to guard against forgetting, and the final layer is split into specialized heads for weights, batch normalization parameters, and residual connection parameters, reducing redundancy and computational overhead.</p>
<p>The practical consequences of this design are striking. Because unlearning operates entirely in parameter space, its computational cost is dominated by the hypernetwork&#8217;s forward and backward passes and is essentially independent of dataset size and class count. The only term that grows over a sequence is the regularization over retained tasks, and it grows linearly in the number of tasks, not data points. In conventional replay-based unlearning adapted to continual settings, by contrast, the cost scales with a replay buffer whose size is difficult to budget in advance. The researchers also introduced an annealing strategy that shrinks the burn-in phase of each unlearning operation by ten percent per operation, exploiting forward transfer to cut unlearning time without degrading quality.</p>
<p>Empirical evaluations across sequential vision benchmarks—including Permuted-MNIST, a five-dataset suite combining MNIST, Fashion-MNIST, KMNIST, notMNIST and SVHN, CIFAR-100, and TinyImageNet—showed that UnCLe can perform long interleaved sequences of learning and unlearning requests, up to 30 operations on TinyImageNet, with minimal disruption to previously acquired knowledge. Measured against baselines including fine-tuning, retraining from scratch, and hypernetwork variants that rely on natural catastrophic forgetting, UnCLe performed on par or better across metrics, and strictly better on three of five measures with a ResNet-18 backbone. It also achieved membership inference attack accuracy closest to the ideal fifty percent, a key indicator that forgotten data is genuinely indistinguishable from never-seen data—a central concern for privacy.</p>
<p>Perhaps most importantly for real-world deployment, UnCLe prevented the relapse phenomenon that plagues conventional approaches: tasks unlearned by prior methods tend to creep back once new learning occurs, whereas UnCLe&#8217;s unlearned tasks stayed forgotten even as subsequent tasks were absorbed. The authors argue this has broad implications for responsible AI governance, from honoring the right to be forgotten under data protection law to stripping biased or harmful behaviors from deployed models without full retraining. At the same time, they caution that the very possibility of relapse under weaker methods underscores the need for robust verification mechanisms. With code released publicly, the framework offers a template for AI systems that can keep learning throughout their operational life while remaining accountable to demands that they forget on command.</p>
<p><strong>Subject of Research:</strong> A hypernetwork framework enabling data-free machine unlearning within continual learning settings</p>
<p><strong>Article Title:</strong> A Hypernetwork Framework for Data-Free Unlearning and Continual Learning</p>
<p><strong>Article References:</strong> A Hypernetwork Framework for Data-Free Unlearning and Continual Learning. (n.d.). <a href="https://doi.org/10.1007/s10994-026-07125-8" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07125-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07125-8" rel="noopener noreferrer">10.1007/s10994-026-07125-8</a></p>
<p><strong>Keywords:</strong> machine unlearning, continual learning, hypernetworks, data privacy, catastrophic forgetting, task embeddings, neural networks, AI safety, right to be forgotten, ResNet, membership inference, parameter generation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205211</post-id>	</item>
		<item>
		<title>Self-Updating AI Learns to Trade as Markets Change, Boosting Returns in New Study</title>
		<link>https://scienmag.com/self-updating-ai-learns-to-trade-as-markets-change-boosting-returns-in-new-study/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 18:53:12 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Adaptive Trading Algorithms]]></category>
		<category><![CDATA[AI-Driven Market Forecasting]]></category>
		<category><![CDATA[algorithmic trading]]></category>
		<category><![CDATA[concept drift]]></category>
		<category><![CDATA[continual learning]]></category>
		<category><![CDATA[Continual Learning in Trading]]></category>
		<category><![CDATA[deep reinforcement learning]]></category>
		<category><![CDATA[Deep Reinforcement Learning for Financial Markets]]></category>
		<category><![CDATA[Evolving Market Conditions]]></category>
		<category><![CDATA[financial forecasting]]></category>
		<category><![CDATA[financial market volatility]]></category>
		<category><![CDATA[GRU]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Market Prediction and Decision-Making]]></category>
		<category><![CDATA[Market Regime Shifts]]></category>
		<category><![CDATA[maximum drawdown]]></category>
		<category><![CDATA[proximal policy optimization]]></category>
		<category><![CDATA[Reinforcement Learning Frameworks for Trading]]></category>
		<category><![CDATA[Self-Updating AI]]></category>
		<category><![CDATA[Sharpe ratio]]></category>
		<category><![CDATA[Streaming Continual Learning]]></category>
		<category><![CDATA[streaming learning]]></category>
		<category><![CDATA[Trading Algorithm Performance Improvement]]></category>
		<category><![CDATA[trading systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201368</guid>

					<description><![CDATA[Researchers have developed a deep reinforcement learning framework that continuously adapts its market forecasts, achieving an average cumulative return of 50.09 percent across six datasets.]]></description>
										<content:encoded><![CDATA[<p>Financial markets never sit still. Regimes shift, volatility clusters arrive without warning, and the statistical relationships that a trading algorithm learned last month can quietly dissolve by the next quarter. A new study tackles exactly this fragility by introducing a deep reinforcement learning framework that keeps learning as markets evolve, and its results suggest that a trading agent equipped with a continuously updated forecasting module can substantially outperform conventional reinforcement learning systems that are trained once and left alone.</p>
<p>The research, published in the Journal of Ambient Intelligence and Humanized Computing, was conducted by Hossein Abbasimehr of Azarbaijan Shahid Madani University, Reza Paki of Politecnico di Milano, and Hamidreza Asadian Rad of Iran University of Science and Technology. Their framework, called Continual Forecasting Fusion Deep Reinforcement Learning, or CFFDRL, embeds streaming continual learning directly into the pipeline of a trading agent. The central idea is deceptively simple: instead of treating market prediction and trading decision-making as two frozen stages, the framework lets the forecasting component adapt continuously to newly generated data, so that the reinforcement learning agent always acts on a view of the market that reflects its most recent behavior.</p>
<p>Deep reinforcement learning has become one of the most actively explored approaches in algorithmic trading. In a typical setup, an agent observes the state of the market, takes actions such as buying, selling, or holding, and receives rewards tied to profit or risk-adjusted performance. Over many training episodes, the agent learns a policy that maps market states to actions. The problem, the authors note, is that these systems are usually optimized on historical data and then deployed as static models. When the underlying data-generating process changes, a phenomenon known in machine learning as concept drift, the learned policy can degrade badly. A policy tuned to a bull market may hold losing positions through a regime change; a strategy tuned to low volatility may misjudge risk when turbulence returns.</p>
<p>To combat this, the researchers turned to streaming continual learning, a branch of machine learning concerned with models that learn from an unbounded flow of data without forgetting what they already know. The specific technique at the heart of CFFDRL is Continuous Piggyback, an approach that adapts to newly generated data by learning task-specific masks over a frozen pre-trained backbone network, without modifying the original weights. Rather than retraining an entire neural network each time new data arrives, which is computationally expensive and risks erasing previously learned knowledge, the framework learns lightweight binary masks that select and reconfigure pathways through the frozen network for each new forecasting task. The result is a model that can absorb new market conditions while preserving the general structure it learned earlier.</p>
<p>The authors implemented this concept inside a gated recurrent unit, a type of recurrent neural network well suited to sequential data such as prices. The resulting module, called cPB-GRU, incrementally predicts future prices from historical OHLC data, the open, high, low, and close values that form the basic vocabulary of market analysis. Crucially, the module is continuously updated during both training and testing. This means the forecasting component does not stop learning when the evaluation phase begins; it keeps adapting as fresh market observations stream in, mirroring the way a human trader might recalibrate expectations day after day.</p>
<p>The forecasts generated by the cPB-GRU module are then concatenated with the raw OHLC data to form the observation space of the reinforcement learning agent. In other words, the trading agent does not only see what has happened in the market; it also sees a continuously refreshed estimate of what the forecasting module expects to happen next. This fusion of prediction and decision-making is what gives CFFDRL its name and its edge. The agent uses the proximal policy optimization algorithm, a widely used and stable reinforcement learning method, and benefits from observations that stay informative even as the market shifts beneath it.</p>
<p>The experimental evidence is drawn from six datasets, giving the comparison a breadth that single-asset backtests often lack. Across those datasets, CFFDRL achieved an average cumulative return of 50.09 percent, compared with 33.28 percent for a standard DRL-PPO baseline and 19.48 percent for a PPO variant paired with a static GRU forecaster. The gap is striking: the continual forecasting agent delivered roughly one and a half times the average return of the standard PPO setup and more than two and a half times that of the static forecasting configuration. The comparison with PPO-Static-GRU is particularly telling, because it isolates the contribution of continual adaptation; the only substantive difference is whether the forecasting module keeps learning from new data.</p>
<p>Profit alone is not the whole story in trading research, and the framework also performed well on standard risk metrics. CFFDRL achieved the highest average Sharpe ratio among the evaluated PPO variants, at 0.10, indicating better risk-adjusted returns, and the lowest average maximum drawdown, at 19.46 percent. Maximum drawdown measures the largest peak-to-trough decline an account experiences, and a lower value signals that the strategy avoids the deepest losses, a property investors typically prize as much as raw profitability. Taken together, the results indicate that continual forecasting improves not only how much the agent earns but how smoothly and safely it earns it.</p>
<p>The broader significance of the work lies in its marriage of two research traditions that have largely developed in parallel. Continual learning researchers have built sophisticated techniques for adapting models to data streams while preventing catastrophic forgetting, but most of that work has focused on classification tasks. Reinforcement learning researchers, meanwhile, have built increasingly powerful trading agents, but often without addressing the non-stationarity of financial data head-on. By making the forecasting module a living, evolving component of the observation space, CFFDRL offers a template for how streaming continual learning can be folded into decision-making systems that operate in environments where yesterday&#8217;s patterns are never quite today&#8217;s.</p>
<p>There are, of course, limits to what any backtest can promise. Live trading introduces transaction costs, slippage, liquidity constraints, and execution delays that no simulation fully captures, and the authors&#8217; study reports no datasets generated or analyzed beyond the reported experiments. Still, the message of the research is clear and likely to resonate across quantitative finance: in non-stationary environments, the ability to keep learning is not a luxury but a determinant of performance. As automated trading systems take on a growing share of global market activity, frameworks like CFFDRL point toward a generation of agents that treat change not as a threat to be endured but as information to be absorbed, one streamed data point at a time.</p>
<p><strong>Subject of Research:</strong> A deep reinforcement learning trading framework using streaming continual learning to adapt forecasts to evolving financial markets</p>
<p><strong>Article Title:</strong> A novel deep reinforcement learning framework with task-incremental continual forecasting for trading systems</p>
<p><strong>Article References:</strong> Abbasimehr, H., Paki, R., &amp; Asadian Rad, H. (2026). A novel deep reinforcement learning framework with task-incremental continual forecasting for trading systems. <em>Journal of Ambient Intelligence and Humanized Computing</em>. <a href="https://doi.org/10.1007/s12652-026-05132-0" rel="noopener noreferrer">https://doi.org/10.1007/s12652-026-05132-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12652-026-05132-0" rel="noopener noreferrer">10.1007/s12652-026-05132-0</a></p>
<p><strong>Keywords:</strong> deep reinforcement learning, algorithmic trading, continual learning, streaming learning, concept drift, financial forecasting, GRU, proximal policy optimization, Sharpe ratio, maximum drawdown, trading systems, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201368</post-id>	</item>
		<item>
		<title>Adaptive LoRA Ranks Help AI Models Learn New Tasks Without Forgetting Old Ones</title>
		<link>https://scienmag.com/adaptive-lora-ranks-help-ai-models-learn-new-tasks-without-forgetting-old-ones/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 20:40:45 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Adaptive LoRA ranks]]></category>
		<category><![CDATA[adaptive ranks]]></category>
		<category><![CDATA[avoiding knowledge overwriting in AI models]]></category>
		<category><![CDATA[catastrophic forgetting]]></category>
		<category><![CDATA[catastrophic forgetting mitigation in large language models]]></category>
		<category><![CDATA[continual learning]]></category>
		<category><![CDATA[dynamic rank allocation in neural network layers]]></category>
		<category><![CDATA[fine-tuning neural networks without losing prior knowledge]]></category>
		<category><![CDATA[GLUE]]></category>
		<category><![CDATA[GSM8K]]></category>
		<category><![CDATA[incremental learning in language models]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[low-rank adaptation (LoRA) methodology]]></category>
		<category><![CDATA[low-rank parameter adjustment]]></category>
		<category><![CDATA[MATH]]></category>
		<category><![CDATA[memory-efficient AI training]]></category>
		<category><![CDATA[parameter-efficient fine-tuning]]></category>
		<category><![CDATA[parameter-efficient fine-tuning techniques]]></category>
		<category><![CDATA[Penn State University AI research]]></category>
		<category><![CDATA[sequential task learning in AI]]></category>
		<category><![CDATA[stability-plasticity]]></category>
		<category><![CDATA[subspace orthogonality]]></category>
		<category><![CDATA[T5]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198420</guid>

					<description><![CDATA[Penn State researchers have developed a continual learning method that adaptively adjusts LoRA ranks layer by layer, helping large language models acquire new tasks without catastrophically forgetting old ones.]]></description>
										<content:encoded><![CDATA[<p>Large language models have become astonishingly capable learners, absorbing new skills from relatively small amounts of fine-tuning data. Yet this flexibility comes with a well-known curse: when an AI model is trained sequentially on new tasks, it tends to overwrite the knowledge it acquired earlier, a phenomenon researchers call catastrophic forgetting. Now, a pair of computer scientists at Pennsylvania State University has proposed an elegant fix that works by changing a surprisingly small detail of how models are fine-tuned — the number of low-rank parameters, or rank, allotted to each layer of the network. Their method, described in the journal Machine Learning, adaptively adjusts these ranks as new tasks arrive, allowing language models to keep learning without erasing what came before.</p>
<p>The research, conducted by Fuli Qiao and Mehrdad Mahdavi of the Department of Computer Science and Engineering at Penn State, builds on Low-Rank Adaptation, or LoRA, one of the most widely used parameter-efficient fine-tuning techniques in modern artificial intelligence. Instead of retraining the billions of weights inside a large language model, LoRA freezes the original weights and injects small pairs of low-rank matrices into each layer. Only these small matrices are updated during training, which slashes memory and computation costs by orders of magnitude. The catch, the authors note, is that conventional LoRA fixes the rank to the same value across every layer and every task, a one-size-fits-all choice that leaves a crucial question unexplored: how much capacity does each layer actually need for each new task?</p>
<p>Qiao and Mahdavi&#8217;s answer is a method that treats rank as a dynamic resource rather than a fixed hyperparameter. Their approach, which the team calls CL-Rank, relies on a subspace similarity metric to measure how orthogonal — that is, how non-overlapping — the low-rank subspaces occupied by different tasks are within a given layer. When a new task arrives, the algorithm evaluates the geometry of the learned subspaces and adaptively increases the layerwise rank for the new task where overlap threatens to interfere with previously learned representations. By steering new knowledge into directions of parameter space that are roughly orthogonal to old knowledge, the method minimizes interference while still giving each new task enough expressive capacity to generalize well.</p>
<p>The underlying intuition echoes a classical idea from neuroscience and machine learning known as the stability–plasticity dilemma. A learning system must be plastic enough to absorb new information but stable enough to retain old skills. Since the pioneering work on catastrophic interference in connectionist networks in the late 1980s and the advent of modern continual learning benchmarks, researchers have tried an arsenal of remedies: replaying stored examples, regularizing important weights, growing new network branches, or constraining gradient updates to null spaces of prior tasks. Many of these approaches are either memory-hungry, brittle to task order, or impractical for models with billions of parameters. What distinguishes the new work is that it tackles the problem entirely within the cheap, low-rank adapter framework, requiring no access to old data and no freezing of the base model.</p>
<p>Technically, the method tracks the principal subspaces spanned by the low-rank updates in each layer and computes a similarity score between the subspace of an incoming task and those of earlier tasks. If the new task&#8217;s gradient updates are projected into directions that overlap heavily with prior subspaces, the rank of the adapter in that layer is expanded, giving the optimizer room to find solutions that spare old representations. Layers whose subspaces remain naturally disjoint need no expansion, so the parameter budget is spent only where it matters. The paper&#8217;s supplementary analyses reveal that the learned rank distributions differ markedly across layers and modules — for example, the variation in ranks among the value projection modules of encoder layers is larger than among query projections, while encoder layers as a whole show more consistent rank allocations than decoder layers. This suggests that different parts of a transformer genuinely serve distinct roles, and that a uniform rank silently wastes capacity in some layers while starving others.</p>
<p>To test the idea, the researchers ran experiments on T5 and Llama-2-7b language models, using GPU servers with DeepSpeed for efficient training. For T5 experiments they employed four NVIDIA A6000 GPUs with a learning rate of 1e-3 and a batch size of 32, while the larger Llama-2-7b runs used four NVIDIA A100 GPUs at a learning rate of 1e-4. The evaluation spanned standard natural language processing continual learning benchmarks built from 15 datasets, including the classic text classification benchmark of Zhang and colleagues, the GLUE and SuperGLUE suites, and the IMDB movie review corpus, with six different task sequence orders to control for ordering effects. The team also pushed into harder territory with challenging mathematical reasoning benchmarks, GSM8K and MATH, each split into sequential tasks and tested in both orders to probe how task sequencing shapes forgetting.</p>
<p>The results show that the adaptive-rank method matches or beats strong baselines on three fronts at once: it mitigates forgetting of earlier tasks, improves accuracy on the tasks being learned, and preserves strong generalization to unseen data — all in a memory-efficient manner. On the math benchmarks, where sequential training on GSM8K and MATH typically causes steep declines in earlier-task accuracy, the method with orthogonal projection consistently outperformed the O-LoRA baseline in final average testing accuracy, forgetting less on the first task while generalizing better on the second. Notably, the study observes that fully fine-tuning a T5-large model on MATH yields only about 3.0 percent accuracy, and about 4.2 percent on GSM8K, consistent with prior work — underscoring how difficult these mathematical tasks are and how much room remains for continual learning approaches that squeeze more capability out of small adapters.</p>
<p>Why should this matter beyond the benchmark tables? Continual learning is arguably the missing ingredient between today&#8217;s static AI assistants and the adaptive, lifelong learning systems that both researchers and companies envision. Every time a deployed model needs a new skill — a new language, a new domain, a new tool — retraining it from scratch or even fully fine-tuning it is prohibitively expensive. Adapter-based continual learning offers a path to modular, on-demand skill acquisition, but only if adding skills does not corrupt existing ones. By showing that the humble LoRA rank, tuned per layer and per task, is a powerful lever for balancing stability and plasticity, the Penn State work gives practitioners a new, inexpensive knob to turn. Because the method requires no stored examples of past tasks, it also sidesteps the privacy and storage concerns that plague replay-based approaches.</p>
<p>The findings also carry a broader scientific message: the internal geography of large models is far from uniform, and adaptation schemes that respect that heterogeneity can outperform uniform ones. The observation that encoder and decoder layers, and even the query and value modules within them, prefer different rank allocations suggests that future parameter-efficient methods may benefit from treating adapters as structured, layer-aware components rather than interchangeable plug-ins. The authors have released their code publicly, and all datasets used in the experiments are openly available, making the results straightforward for other groups to verify and extend. The work was partially supported by the National Science Foundation under award EFMA-2318101, with experiments designed and conducted by Fuli Qiao under the supervision of Mehrdad Mahdavi.</p>
<p>As large language models continue to spread through science, medicine, and industry, the question of how to teach them new things without breaking old ones is shifting from academic curiosity to engineering necessity. This study does not solve continual learning outright — no single method does — but it demonstrates that a careful, geometry-aware allocation of low-rank capacity can deliver competitive anti-forgetting performance at a fraction of the cost of full retraining. In a field where progress often comes from scaling up, it is a reminder that sometimes the most powerful improvements come from scaling the right things in the right places, one layer at a time.</p>
<p><strong>Subject of Research:</strong> Adaptive layerwise LoRA ranks for continual learning in large language models</p>
<p><strong>Article Title:</strong> Learning Without Forgetting for Continual Learning in LLMs Through Adaptive LoRA Ranks</p>
<p><strong>Article References:</strong> Qiao, F., &amp; Mahdavi, M. (2026). Learning Without Forgetting for Continual Learning in LLMs Through Adaptive LoRA Ranks. <em>Machine Learning, 115</em>(9), Article 212. <a href="https://doi.org/10.1007/s10994-026-07145-4" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07145-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07145-4" rel="noopener noreferrer">10.1007/s10994-026-07145-4</a></p>
<p><strong>Keywords:</strong> continual learning, large language models, LoRA, catastrophic forgetting, parameter-efficient fine-tuning, adaptive ranks, subspace orthogonality, stability-plasticity, GLUE, GSM8K, MATH, T5</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198420</post-id>	</item>
		<item>
		<title>Calibrated Prototypes Help AI Spot New Cyberattacks Without Forgetting Old Ones</title>
		<link>https://scienmag.com/calibrated-prototypes-help-ai-spot-new-cyberattacks-without-forgetting-old-ones/</link>
		
		<dc:creator><![CDATA[Hailey Crawford]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 15:53:30 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive cybersecurity systems]]></category>
		<category><![CDATA[catastrophic forgetting]]></category>
		<category><![CDATA[catastrophic forgetting in neural networks]]></category>
		<category><![CDATA[CICIDS2017]]></category>
		<category><![CDATA[continual learning]]></category>
		<category><![CDATA[continuous learning in intrusion detection systems]]></category>
		<category><![CDATA[cyberattack detection]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[few-shot class-incremental learning]]></category>
		<category><![CDATA[few-shot learning for cyber threats]]></category>
		<category><![CDATA[incremental machine learning for cybersecurity]]></category>
		<category><![CDATA[intrusion detection]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[network security]]></category>
		<category><![CDATA[neural network stability in cybersecurity]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[new methods for detecting evolving cyberattacks]]></category>
		<category><![CDATA[overcoming knowledge loss in AI security models]]></category>
		<category><![CDATA[preserving knowledge in AI-based threat detection]]></category>
		<category><![CDATA[prototype calibration]]></category>
		<category><![CDATA[prototype calibration for intrusion detection]]></category>
		<category><![CDATA[scalable cyberattack classification techniques]]></category>
		<category><![CDATA[semantic similarity]]></category>
		<category><![CDATA[UNSW-NB15]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196043</guid>

					<description><![CDATA[Researchers in India have developed BiPC-IFS, a few-shot class-incremental learning framework that lets intrusion detection systems learn new cyberattacks from minimal samples without catastrophically forgetting previous ones.]]></description>
										<content:encoded><![CDATA[<p>Cybersecurity has long suffered from a paradox at the heart of machine learning: the models that defend our networks are often the slowest to adapt to the very threats they are meant to stop. When attackers deploy a new form of intrusion, defenders must retrain their detection systems, and in doing so they frequently erase the knowledge those systems already hold about earlier attacks. Researchers at Malaviya National Institute of Technology Jaipur in India have now unveiled a framework designed to break this cycle. Their approach, called BiPC-IFS, short for Biased Prototype Calibration For Incremental Few Shot Intrusion Detection, allows an intrusion detection system to learn brand-new attack types from only a handful of examples while preserving, rather than overwriting, what it has already learned about older threats.</p>
<p>The work, published in the journal Neural Computing and Applications by Parvati Bhurani, Satyendra Singh Chouhan and Namita Mittal, addresses one of the most stubborn problems in applied machine learning, known formally as catastrophic forgetting. First documented in the late 1980s by psychologists studying connectionist networks, the phenomenon describes what happens when a neural network trained sequentially on multiple tasks loses proficiency on earlier tasks as it absorbs new ones. In the context of network security, this is not an academic curiosity. An intrusion detection system that forgets how to recognize a denial-of-service flood because it has just been taught to spot a novel botnet signature is a system that has become a liability, not a safeguard.</p>
<p>The framework the Indian team proposes falls under an emerging learning paradigm known as few-shot class-incremental learning, or FSCIL. The idea is to structure the learning problem so that a model first learns a broad set of base classes from a fully labeled dataset, and then progressively incorporates novel classes from just a few labeled samples per class, all without revisiting the original training data. This mirrors the operational reality of cybersecurity. Organizations typically possess abundant examples of well-known attacks, but when a new exploit appears in the wild, security teams may have only a few confirmed instances of it before the next wave of probes arrives. A detection model suited to this environment must therefore extract maximum information from minimal new evidence while keeping its existing knowledge intact.</p>
<p>BiPC-IFS achieves this balance through two core components: a fixed feature extractor and a prototype calibration module. The feature extractor is trained only during the base session, on the well-populated set of established attack classes, and is then frozen for the remainder of the system&#8217;s operational life. Although this might seem restrictive, the researchers found that the frozen extractor still captures meaningful similarity relationships between the base classes and the novel classes that arrive later. Because the extractor encodes the geometry of network traffic in a stable feature space, new attack types can be located within that space even when only a handful of examples exist, simply by measuring where their feature representations fall relative to everything the model already knows.</p>
<p>The second component, prototype calibration, is where the approach earns its distinctive name. In prototype-based classification, each class is represented by a single representative vector, or prototype, typically computed as the mean of the feature vectors of its training samples. With only a few samples, these novel-class prototypes are biased, pulled away from their true class centers by sampling noise and by the tendency of a model trained on base classes to interpret everything through the lens of what it already knows. Calibration corrects this bias by adjusting the prototypes before classification. The crucial design question, the authors note, is determining how much to adjust: a calibration factor that is too high can distort the original representation of the novel class, effectively overcorrecting and making the system worse than it would have been with no calibration at all.</p>
<p>What sets BiPC-IFS apart from earlier calibration techniques is the way it computes that correction. Rather than relying solely on distances in feature space, the proposed calibrated class prototype aggregates both feature-based similarity and semantic similarity among different classes. In practical terms, this means the system considers not only how close a novel attack&#8217;s samples sit to the prototypes of known attacks in the learned feature space, but also how conceptually related the classes are. Two attack types that share characteristics, for example variants of the same malware family, can inform each other&#8217;s prototypes in a way that purely geometric calibration cannot achieve. This dual-source aggregation allows the model to draw richer inferences from the sparse evidence available in each incremental session, producing prototypes that better represent the true structure of the new classes.</p>
<p>To test whether these design choices translate into real-world performance, the researchers evaluated BiPC-IFS on two of the most widely used benchmark datasets in intrusion detection research: UNSW-NB15 and CICIDS2017. The UNSW-NB15 dataset, created at the Australian Centre for Cyber Security, combines real normal traffic with nine categories of synthesized modern attacks, including backdoors, exploits, and reconnaissance activity. CICIDS2017, produced by the Canadian Institute for Cybersecurity, captures several days of benign and attack traffic covering brute-force assaults, heartbleed exploits, botnets, denial-of-service attacks, web attacks, and infiltration attempts. Together, these benchmarks provide a demanding testbed, with realistic class distributions and attack behaviors that differ substantially across categories, exactly the conditions under which incremental learning systems tend to falter.</p>
<p>The results were striking. BiPC-IFS surpassed the baseline methods it was compared against and achieved the strongest performance metrics for novel classes across both datasets. The system recorded an average accuracy of 94.91 percent across all incremental sessions, a novel class accuracy of 74.25 percent, and a performance drop, measured as the decline in accuracy over the course of learning new classes, of just 8.67 percent. That final figure is the one that matters most to security practitioners, because it quantifies how much the system forgets as it learns. A small drop means that the model&#8217;s knowledge of old attacks remains largely intact even as it absorbs new ones, which is precisely the property that conventional retraining pipelines fail to deliver.</p>
<p>The implications extend well beyond one laboratory result. Networks today face an adversary that evolves continuously, probing for unpatched vulnerabilities and mutating attack tooling faster than human analysts can label large datasets. Systems like BiPC-IFS point toward a generation of defenses that can be updated on the fly, in operational settings, without the downtime and cost of full retraining and without the silent erosion of previously learned protections. Because the feature extractor remains frozen, the computational cost of incorporating a new attack class is minimal, and the approach avoids the need to store sensitive raw traffic data from past sessions. The researchers also note that the datasets used in the study are publicly available, which should make it straightforward for other teams to reproduce the results and build on them.</p>
<p>There remain, of course, open questions. The frozen feature extractor, though shown to capture useful base-novel similarity, was never trained to see the novel classes, and future work may explore how well this holds as threat landscapes diverge further from historical attack patterns. The authors&#8217; own framing acknowledges the delicate trade-off at the center of the method: the calibration factor must be chosen carefully, since too aggressive a correction distorts the very representations it is meant to refine. Even so, the demonstration that biased prototype calibration, informed jointly by feature and semantic similarity, can push novel-class detection to over 74 percent accuracy from just a few examples marks a meaningful advance. As artificial intelligence becomes the front line of network defense, techniques that let models learn like analysts do, quickly, from limited evidence, and without forgetting hard-won lessons, may prove indispensable.</p>
<p><strong>Subject of Research:</strong> A biased prototype calibration framework for incremental few-shot learning in network intrusion detection systems.</p>
<p><strong>Article Title:</strong> BiPC-IFS: Biased Prototype Calibration For Incremental Few Shot Intrusion Detection</p>
<p><strong>Article References:</strong> Bhurani, P., Chouhan, S. S., &amp; Mittal, N. (2026). BiPC-IFS: Biased Prototype Calibration For Incremental Few Shot Intrusion Detection. <em>Neural Computing and Applications, 38</em>(17), Article 736. <a href="https://doi.org/10.1007/s00521-026-12455-8" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12455-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12455-8" rel="noopener noreferrer">10.1007/s00521-026-12455-8</a></p>
<p><strong>Keywords:</strong> intrusion detection, few-shot class-incremental learning, catastrophic forgetting, prototype calibration, machine learning, cybersecurity, network security, semantic similarity, UNSW-NB15, CICIDS2017, neural networks, continual learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196043</post-id>	</item>
	</channel>
</rss>
