<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>parameter efficiency &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/parameter-efficiency/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 23:23:37 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>parameter efficiency &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Compact AI Network Blends Mamba and Polynomial Neurons for Medical Image Segmentation</title>
		<link>https://scienmag.com/compact-ai-network-blends-mamba-and-polynomial-neurons-for-medical-image-segmentation/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 23:23:37 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[blood vessel and lesion segmentation]]></category>
		<category><![CDATA[compact neural network architecture]]></category>
		<category><![CDATA[computational efficiency in medical imaging]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[dermoscopy]]></category>
		<category><![CDATA[global and local feature extraction]]></category>
		<category><![CDATA[hybrid neural network design]]></category>
		<category><![CDATA[long-range context propagation]]></category>
		<category><![CDATA[low-parameter deep learning models]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[Mamba state-space model]]></category>
		<category><![CDATA[medical image segmentation]]></category>
		<category><![CDATA[neural architecture]]></category>
		<category><![CDATA[parameter efficiency]]></category>
		<category><![CDATA[polynomial expansion]]></category>
		<category><![CDATA[polynomial neurons]]></category>
		<category><![CDATA[polyp segmentation]]></category>
		<category><![CDATA[retinal vessels]]></category>
		<category><![CDATA[state-space models]]></category>
		<category><![CDATA[transformer alternatives for image analysis]]></category>
		<category><![CDATA[U-Net]]></category>
		<category><![CDATA[U-shaped encoder-decoder]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224290</guid>

					<description><![CDATA[Researchers have unveiled ANNE, a compact segmentation network that pairs Mamba state-space modeling with adaptive polynomial neuron expansion to match much larger AI models across eight medical imaging benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Medical image segmentation, the task of tracing the exact outlines of lesions, polyps, and blood vessels in clinical scans, has long been dominated by a trade-off. Convolutional networks excel at capturing fine local detail but struggle to connect distant regions of an image, while transformer-based models reason globally at a steep computational price. A new architecture called ANNE, described in the journal Applied Intelligence by Abel A. Reyes-Angulo of Michigan Technological University and Sidike Paheding of Fairfield University, argues that this trade-off can be broken with a surprisingly compact design that needs only about 2.77 million trainable parameters.</p>
<p>The central insight behind ANNE is that two very different mathematical tools can do complementary jobs inside a single U-shaped encoder-decoder. In the encoder, the authors deploy Mamba, a selective state-space model that has attracted wide attention since its introduction as a linear-time alternative to transformer attention. Instead of building a quadratic attention matrix over every pair of pixels, Mamba propagates information along a flattened sequence of the feature map through input-dependent state parameters, allowing long-range context to flow through the network at a cost that scales linearly with image size. This is what lets ANNE reason about broad spatial relationships without the memory burden that typically accompanies global modeling.</p>
<p>The second ingredient, and the paper&#8217;s main methodological novelty, is a module the authors call Adaptive Progressive Expansion Neurons, or A-PEN. The idea builds on a line of prior work known as progressively expanded neuron models, which increase the nonlinear expressiveness of a network by computing polynomial terms of the input features, such as squared and cubed activations, and combining them with learned weights. In earlier designs, however, the set of expansion terms and the activation applied afterward were fixed in advance. A-PEN makes both adaptive: a bank of candidate polynomial branches is weighted by normalized gates derived from trainable logits through a softmax function, and the post-expansion nonlinearity is a parametric PReLU whose negative slope is learned during training.</p>
<p>The authors are notably careful about how these gates should be interpreted. Because each normalized gate multiplies an unconstrained trainable coefficient, the effective contribution of a polynomial branch is the product of the two, meaning different gate-coefficient combinations can produce identical outputs. The gates therefore form a coupled parameterization that modulates the relative scaling of candidate branches during optimization, not a uniquely identifiable measure of how important each polynomial order is. The team also draws a precise boundary around its use of Remez approximation: the classical minimax theory of best polynomial approximation motivates the idea that compact polynomial bases can represent complex nonlinear mappings, but ANNE never executes the Remez exchange algorithm and claims no minimax guarantee. The polynomial coefficients are simply learned end to end from the segmentation loss.</p>
<p>Architecturally, ANNE follows the familiar U-shaped template that has defined biomedical segmentation since U-Net. Convolutional stem layers preserve local spatial structure at the input, Mamba blocks then model long-range dependencies at progressively coarser resolutions, and A-PEN blocks refine nonlinear feature interactions at the bottleneck and in the decoder before upsampling. Skip connections carry fine-grained encoder information to matching decoder stages, and the decoder employs multi-scale depth-wise convolutions with kernel sizes of one, three, and five. A final projection produces a binary segmentation map, with the whole pipeline trained using a weighted combination of binary cross-entropy and soft Dice loss.</p>
<p>The evaluation is unusually transparent about how comparisons were made. The authors tested ANNE on eight public 2D datasets spanning three clinical domains: dermoscopic lesion segmentation on ISIC 2016, ISIC 2017, ISIC 2018, and PH2; colonoscopic polyp segmentation on Kvasir-SEG and CVC-ClinicDB; and retinal-vessel segmentation on DRIVE and CHASE_DB1. Crucially, the results tables distinguish controlled reproductions, where competing methods were retrained under an identical protocol with the same splits, preprocessing, loss, and five fixed random seeds, from literature-reported values, which may reflect different experimental conditions. Statistical claims are restricted to the controlled comparisons, while literature numbers are described only as context.</p>
<p>Within that framework, ANNE delivers a strong and consistent showing. On ISIC 2018 and PH2 it achieved the strongest tabulated values among the compared methods, and on ISIC 2017 it recorded the lowest 95th-percentile Hausdorff distance, a boundary-sensitive metric, even though another method reported higher overlap scores. In polyp segmentation it topped the tables on Kvasir-SEG across Dice, IoU, and boundary distance, and placed second on CVC-ClinicDB with the best boundary accuracy. For retinal vessels, where thin, branching structures test a model&#8217;s ability to preserve fine detail, ANNE achieved the best tabulated IoU and boundary distance on DRIVE and swept all three metrics on CHASE_DB1.</p>
<p>Ablation experiments illuminate why the design works. Comparing a compact U-Net-like baseline against versions augmented with only A-PEN, only Mamba, or both, the authors found that A-PEN consistently improved boundary-sensitive performance while Mamba tended to boost overlap where broad context matters. The full model achieved the best Dice score in every ablation row, supporting the claim that the two modules are complementary rather than interchangeable. A separate resolution analysis showed that accuracy improved as input size rose from 128 by 128 to 256 by 256 pixels, but saturated beyond that point even as computational cost continued to climb roughly quadratically, making 256 by 256 the identified sweet spot for lesion and polyp tasks.</p>
<p>The authors are candid about the limits of their evidence. All experiments used public 2D benchmarks, so nothing is established about 3D volumes, video, or prospective clinical workflows. The evaluation is in-distribution, with no external cohorts from different scanners, institutions, or populations, and several benchmarks are small enough that uncertainty remains despite repeated runs. Efficiency was measured on an NVIDIA A100 GPU, which characterizes relative cost but says nothing about real-time performance on edge devices or hospital hardware. The team also notes that segmentation accuracy alone is insufficient for clinical deployment, which would additionally demand uncertainty estimation, failure detection, calibration, fairness analysis, and workflow integration.</p>
<p>Even with those caveats, ANNE represents a compelling data point in the ongoing effort to shrink medical AI without shrinking its competence. By pairing a linear-time state-space encoder with learnable polynomial feature expansion, the architecture shows that competitive overlap and boundary accuracy across dermoscopy, colonoscopy, and fundus imaging can be achieved with a parameter budget far below that of transformer-heavy rivals. The code, dataset split manifests, and configuration files have been released in a public repository, and the authors point toward volumetric extensions, external-domain validation, and deployment-oriented profiling on resource-constrained hardware as the next frontiers. For clinics and research labs where compute is scarce but precision is not, that combination of transparency, efficiency, and measured performance may prove just as influential as the architecture itself.</p>
<p><strong>Subject of Research:</strong> A parameter-efficient deep learning architecture combining Mamba state-space blocks and adaptive polynomial neuron expansion for 2D medical image segmentation</p>
<p><strong>Article Title:</strong> ANNE: adaptive neuron network expansion for 2D medical image segmentation</p>
<p><strong>Article References:</strong> Reyes-Angulo, A. A., &amp; Paheding, S. (2026). ANNE: adaptive neuron network expansion for 2D medical image segmentation. <em>Applied Intelligence, 56</em>(15), Article 464. <a href="https://doi.org/10.1007/s10489-026-07503-8" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07503-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07503-8" rel="noopener noreferrer">10.1007/s10489-026-07503-8</a></p>
<p><strong>Keywords:</strong> medical image segmentation, Mamba, state-space models, polynomial expansion, neural architecture, dermoscopy, polyp segmentation, retinal vessels, parameter efficiency, deep learning, U-Net, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224290</post-id>	</item>
		<item>
		<title>Neurons Point the Way: NB-Net Puts Network Width, Not Depth, in the Spotlight</title>
		<link>https://scienmag.com/neurons-point-the-way-nb-net-puts-network-width-not-depth-in-the-spotlight/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 20:52:31 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[biological inspiration in deep learning]]></category>
		<category><![CDATA[biological principles in AI]]></category>
		<category><![CDATA[biologically-inspired architecture]]></category>
		<category><![CDATA[CIFAR-10]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning model efficiency]]></category>
		<category><![CDATA[divergent-convergent neural processing]]></category>
		<category><![CDATA[enhancing deep learning models]]></category>
		<category><![CDATA[exploring alternative neural network designs]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[grouped convolutions]]></category>
		<category><![CDATA[image classification]]></category>
		<category><![CDATA[ImageNet]]></category>
		<category><![CDATA[innovative neural network structures]]></category>
		<category><![CDATA[multi-branch architecture]]></category>
		<category><![CDATA[NB-Net]]></category>
		<category><![CDATA[network width]]></category>
		<category><![CDATA[network width vs depth]]></category>
		<category><![CDATA[Neural network architecture]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[neural signal processing]]></category>
		<category><![CDATA[neuron bundle network (NB-Net)]]></category>
		<category><![CDATA[parallel computational units]]></category>
		<category><![CDATA[parameter efficiency]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198528</guid>

					<description><![CDATA[Researchers have introduced NB-Net, a biologically inspired neural network that expands in width through parallel multi-scale branches and staged feature fusion, achieving competitive accuracy on CIFAR-10 and ImageNet with controlled parameter growth.]]></description>
										<content:encoded><![CDATA[<p>For more than a decade, the story of deep learning has been told largely in terms of depth. Each new generation of record-setting models has stacked more layers onto the previous one, and the word &#8220;deep&#8221; in deep learning has become shorthand for progress itself. A new study published in Neural Processing Letters argues that this vertical obsession has left half of the design space underexplored. Researchers Longfei Tan and Huihuang Zhao of Hengyang Normal University, together with Wei-Liang Meng of the Institute of Automation at the Chinese Academy of Sciences, have introduced the Neuron Bundle Network, or NB-Net, an architecture that takes the opposite tack: instead of stretching networks downward, it widens them outward, arranging computational units in parallel bundles whose organization is inspired by how biological neurons diverge and converge their signals.</p>
<p>The biological inspiration at the heart of NB-Net comes from a structural principle familiar to neuroscientists. In living nervous systems, a single neuron frequently fans its output out to many downstream targets, and those signals are later gathered and integrated at convergent junctions further along the pathway. This divergent-convergent pattern allows nervous systems to process multiple aspects of a stimulus simultaneously before reconciling them into a unified response. The research team asked a deceptively simple question: if biological computation relies so heavily on this broad, parallel organization rather than on arbitrarily long chains of processing, could artificial networks benefit from a similar width-first philosophy, achieving strong representation quality without resorting to very deep backbones?</p>
<p>NB-Net answers that question with a concrete architectural recipe. The building block of the network is the Neuron Bundle Layer, a multi-branch structure in which several parallel convolutional branches all receive the same input tensor and process it simultaneously. Rather than forcing every branch to use identical operations, the design employs grouped convolutions with different kernel sizes across the branches. This means that at any given stage of the network, the model is simultaneously extracting features at multiple spatial scales from the very same signal. Small kernels capture fine-grained local detail, while larger kernels perceive coarser structure, and the network as a whole maintains a rich, multi-scale view of its input that a single-pathway design of comparable cost would struggle to match.</p>
<p>Of course, running many branches in parallel creates its own challenge: at some point, the network must fuse the separate streams of features back together into a coherent representation. Early attempts at such fusion in multi-branch architectures have often relied on a single 1&#215;1 convolution that collapses all channels at once, a move that can be numerically unstable when the incoming branches are numerous or heterogeneous. NB-Net instead introduces a two-stage 1&#215;1 convolution fusion module. In the first stage, the module performs a moderate compression of the concatenated channel dimension, gently reducing the volume of information rather than merging everything in one abrupt step. In the second stage, the module completes the integration of the compressed features. The authors report that this progressive fusion improves training stability compared with a single 1&#215;1 merge, because the network is never forced to reconcile wildly different feature distributions in a single transformation.</p>
<p>Repeat the Neuron Bundle Layer and its fusion module in succession, and the result is a network that grows primarily in width rather than depth. The team subjected this design to a systematic set of ablation experiments, examining how branch configuration, residual connections, and width settings each affect performance. Residual connections, the shortcut pathways popularized by deep residual networks, were found to interact meaningfully with the multi-branch structure, helping gradients flow through the widened architecture during training. The analysis of width settings demonstrated that the framework scales in a controlled fashion: increasing the number of parallel channels raises capacity, but parameter counts grow in a disciplined, predictable manner rather than exploding as they often do when architectures are naively broadened.</p>
<p>The empirical case for NB-Net rests on two of the most widely used benchmarks in computer vision. On CIFAR-10, the standard small-image classification suite that serves as a proving ground for new architectural ideas, the network delivered competitive classification accuracy against established baselines. More significantly, on ImageNet, the large-scale benchmark whose thousand object categories have historically rewarded sheer architectural depth, NB-Net achieved accuracy on par with competing designs while keeping parameter growth under tight control. In a field where improvements are often purchased with dramatic increases in model size, the ability to hold the line on parameters while remaining competitive is itself a noteworthy result, and it supports the paper&#8217;s central claim that width-oriented design, when done systematically, deserves a seat at the table alongside depth scaling.</p>
<p>The publication arrives at a moment of genuine ferment in architectural research. As the costs of training ever-larger monolithic models attract scrutiny, many groups are revisiting efficiency-driven strategies such as grouped convolutions, multi-branch topologies, and staged feature fusion. NB-Net contributes to this conversation a coherent biological rationale that ties these elements together. The divergent-convergent organization it emulates is not merely a decorative metaphor; it dictates the concrete structure of the network, from the parallel branches of each Neuron Bundle Layer to the progressive, two-stage reconciliation of their outputs. The result is an architecture in which the biological principle and the engineering mechanism reinforce one another, a combination that has historically produced some of the most durable ideas in machine learning, from convolutional receptive fields to attention mechanisms.</p>
<p>Openness was clearly a priority for the research team. The article is published open access under a Creative Commons Attribution 4.0 license, meaning that anyone can read, reuse, and build upon the work with appropriate credit. In addition, the authors have released their implementation in a public code repository on GitHub, allowing other researchers to reproduce the reported benchmarks, stress-test the two-stage fusion module on their own tasks, or graft Neuron Bundle Layers into entirely different network families. The study was supported by funding from the National Natural Science Foundation of China, the Hunan Provincial Natural Science Foundation, and the Science and Technology Innovation Program of Hunan Province, reflecting sustained institutional investment in fundamental artificial intelligence research in China.</p>
<p>Whether NB-Net or its descendants will change mainstream practice remains to be seen, but the study makes a persuasive case that the width of a network is not a second-class design variable. By borrowing a page from the divergent-convergent wiring of biological neurons, the researchers have shown that parallel, multi-scale branches, fused with care, can rival deep stacks of layers on the benchmarks that matter, all while keeping parameter budgets in check. For a discipline searching for ways to advance without simply building bigger and taller models, the message is timely: sometimes the most productive direction is sideways.</p>
<p><strong>Subject of Research:</strong> A biologically inspired, width-oriented multi-branch neural network architecture with two-stage feature fusion</p>
<p><strong>Article Title:</strong> NB-Net: A Biologically-Inspired Framework for Neural Network Width Expansion</p>
<p><strong>Article References:</strong> Tan, L., Zhao, H., &amp; Meng, W.-L. (2026). NB-Net: A Biologically-Inspired Framework for Neural Network Width Expansion. <em>Neural Processing Letters</em>. <a href="https://doi.org/10.1007/s11063-026-11880-z" rel="noopener noreferrer">https://doi.org/10.1007/s11063-026-11880-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11063-026-11880-z" rel="noopener noreferrer">10.1007/s11063-026-11880-z</a></p>
<p><strong>Keywords:</strong> neural networks, network width, biologically-inspired architecture, NB-Net, feature fusion, grouped convolutions, deep learning, image classification, CIFAR-10, ImageNet, multi-branch architecture, parameter efficiency</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198528</post-id>	</item>
	</channel>
</rss>
