<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>polyp segmentation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/polyp-segmentation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 23:23:37 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>polyp segmentation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Compact AI Network Blends Mamba and Polynomial Neurons for Medical Image Segmentation</title>
		<link>https://scienmag.com/compact-ai-network-blends-mamba-and-polynomial-neurons-for-medical-image-segmentation/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 23:23:37 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[blood vessel and lesion segmentation]]></category>
		<category><![CDATA[compact neural network architecture]]></category>
		<category><![CDATA[computational efficiency in medical imaging]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[dermoscopy]]></category>
		<category><![CDATA[global and local feature extraction]]></category>
		<category><![CDATA[hybrid neural network design]]></category>
		<category><![CDATA[long-range context propagation]]></category>
		<category><![CDATA[low-parameter deep learning models]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[Mamba state-space model]]></category>
		<category><![CDATA[medical image segmentation]]></category>
		<category><![CDATA[neural architecture]]></category>
		<category><![CDATA[parameter efficiency]]></category>
		<category><![CDATA[polynomial expansion]]></category>
		<category><![CDATA[polynomial neurons]]></category>
		<category><![CDATA[polyp segmentation]]></category>
		<category><![CDATA[retinal vessels]]></category>
		<category><![CDATA[state-space models]]></category>
		<category><![CDATA[transformer alternatives for image analysis]]></category>
		<category><![CDATA[U-Net]]></category>
		<category><![CDATA[U-shaped encoder-decoder]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224290</guid>

					<description><![CDATA[Researchers have unveiled ANNE, a compact segmentation network that pairs Mamba state-space modeling with adaptive polynomial neuron expansion to match much larger AI models across eight medical imaging benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Medical image segmentation, the task of tracing the exact outlines of lesions, polyps, and blood vessels in clinical scans, has long been dominated by a trade-off. Convolutional networks excel at capturing fine local detail but struggle to connect distant regions of an image, while transformer-based models reason globally at a steep computational price. A new architecture called ANNE, described in the journal Applied Intelligence by Abel A. Reyes-Angulo of Michigan Technological University and Sidike Paheding of Fairfield University, argues that this trade-off can be broken with a surprisingly compact design that needs only about 2.77 million trainable parameters.</p>
<p>The central insight behind ANNE is that two very different mathematical tools can do complementary jobs inside a single U-shaped encoder-decoder. In the encoder, the authors deploy Mamba, a selective state-space model that has attracted wide attention since its introduction as a linear-time alternative to transformer attention. Instead of building a quadratic attention matrix over every pair of pixels, Mamba propagates information along a flattened sequence of the feature map through input-dependent state parameters, allowing long-range context to flow through the network at a cost that scales linearly with image size. This is what lets ANNE reason about broad spatial relationships without the memory burden that typically accompanies global modeling.</p>
<p>The second ingredient, and the paper&#8217;s main methodological novelty, is a module the authors call Adaptive Progressive Expansion Neurons, or A-PEN. The idea builds on a line of prior work known as progressively expanded neuron models, which increase the nonlinear expressiveness of a network by computing polynomial terms of the input features, such as squared and cubed activations, and combining them with learned weights. In earlier designs, however, the set of expansion terms and the activation applied afterward were fixed in advance. A-PEN makes both adaptive: a bank of candidate polynomial branches is weighted by normalized gates derived from trainable logits through a softmax function, and the post-expansion nonlinearity is a parametric PReLU whose negative slope is learned during training.</p>
<p>The authors are notably careful about how these gates should be interpreted. Because each normalized gate multiplies an unconstrained trainable coefficient, the effective contribution of a polynomial branch is the product of the two, meaning different gate-coefficient combinations can produce identical outputs. The gates therefore form a coupled parameterization that modulates the relative scaling of candidate branches during optimization, not a uniquely identifiable measure of how important each polynomial order is. The team also draws a precise boundary around its use of Remez approximation: the classical minimax theory of best polynomial approximation motivates the idea that compact polynomial bases can represent complex nonlinear mappings, but ANNE never executes the Remez exchange algorithm and claims no minimax guarantee. The polynomial coefficients are simply learned end to end from the segmentation loss.</p>
<p>Architecturally, ANNE follows the familiar U-shaped template that has defined biomedical segmentation since U-Net. Convolutional stem layers preserve local spatial structure at the input, Mamba blocks then model long-range dependencies at progressively coarser resolutions, and A-PEN blocks refine nonlinear feature interactions at the bottleneck and in the decoder before upsampling. Skip connections carry fine-grained encoder information to matching decoder stages, and the decoder employs multi-scale depth-wise convolutions with kernel sizes of one, three, and five. A final projection produces a binary segmentation map, with the whole pipeline trained using a weighted combination of binary cross-entropy and soft Dice loss.</p>
<p>The evaluation is unusually transparent about how comparisons were made. The authors tested ANNE on eight public 2D datasets spanning three clinical domains: dermoscopic lesion segmentation on ISIC 2016, ISIC 2017, ISIC 2018, and PH2; colonoscopic polyp segmentation on Kvasir-SEG and CVC-ClinicDB; and retinal-vessel segmentation on DRIVE and CHASE_DB1. Crucially, the results tables distinguish controlled reproductions, where competing methods were retrained under an identical protocol with the same splits, preprocessing, loss, and five fixed random seeds, from literature-reported values, which may reflect different experimental conditions. Statistical claims are restricted to the controlled comparisons, while literature numbers are described only as context.</p>
<p>Within that framework, ANNE delivers a strong and consistent showing. On ISIC 2018 and PH2 it achieved the strongest tabulated values among the compared methods, and on ISIC 2017 it recorded the lowest 95th-percentile Hausdorff distance, a boundary-sensitive metric, even though another method reported higher overlap scores. In polyp segmentation it topped the tables on Kvasir-SEG across Dice, IoU, and boundary distance, and placed second on CVC-ClinicDB with the best boundary accuracy. For retinal vessels, where thin, branching structures test a model&#8217;s ability to preserve fine detail, ANNE achieved the best tabulated IoU and boundary distance on DRIVE and swept all three metrics on CHASE_DB1.</p>
<p>Ablation experiments illuminate why the design works. Comparing a compact U-Net-like baseline against versions augmented with only A-PEN, only Mamba, or both, the authors found that A-PEN consistently improved boundary-sensitive performance while Mamba tended to boost overlap where broad context matters. The full model achieved the best Dice score in every ablation row, supporting the claim that the two modules are complementary rather than interchangeable. A separate resolution analysis showed that accuracy improved as input size rose from 128 by 128 to 256 by 256 pixels, but saturated beyond that point even as computational cost continued to climb roughly quadratically, making 256 by 256 the identified sweet spot for lesion and polyp tasks.</p>
<p>The authors are candid about the limits of their evidence. All experiments used public 2D benchmarks, so nothing is established about 3D volumes, video, or prospective clinical workflows. The evaluation is in-distribution, with no external cohorts from different scanners, institutions, or populations, and several benchmarks are small enough that uncertainty remains despite repeated runs. Efficiency was measured on an NVIDIA A100 GPU, which characterizes relative cost but says nothing about real-time performance on edge devices or hospital hardware. The team also notes that segmentation accuracy alone is insufficient for clinical deployment, which would additionally demand uncertainty estimation, failure detection, calibration, fairness analysis, and workflow integration.</p>
<p>Even with those caveats, ANNE represents a compelling data point in the ongoing effort to shrink medical AI without shrinking its competence. By pairing a linear-time state-space encoder with learnable polynomial feature expansion, the architecture shows that competitive overlap and boundary accuracy across dermoscopy, colonoscopy, and fundus imaging can be achieved with a parameter budget far below that of transformer-heavy rivals. The code, dataset split manifests, and configuration files have been released in a public repository, and the authors point toward volumetric extensions, external-domain validation, and deployment-oriented profiling on resource-constrained hardware as the next frontiers. For clinics and research labs where compute is scarce but precision is not, that combination of transparency, efficiency, and measured performance may prove just as influential as the architecture itself.</p>
<p><strong>Subject of Research:</strong> A parameter-efficient deep learning architecture combining Mamba state-space blocks and adaptive polynomial neuron expansion for 2D medical image segmentation</p>
<p><strong>Article Title:</strong> ANNE: adaptive neuron network expansion for 2D medical image segmentation</p>
<p><strong>Article References:</strong> Reyes-Angulo, A. A., &amp; Paheding, S. (2026). ANNE: adaptive neuron network expansion for 2D medical image segmentation. <em>Applied Intelligence, 56</em>(15), Article 464. <a href="https://doi.org/10.1007/s10489-026-07503-8" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07503-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07503-8" rel="noopener noreferrer">10.1007/s10489-026-07503-8</a></p>
<p><strong>Keywords:</strong> medical image segmentation, Mamba, state-space models, polynomial expansion, neural architecture, dermoscopy, polyp segmentation, retinal vessels, parameter efficiency, deep learning, U-Net, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224290</post-id>	</item>
		<item>
		<title>Neighborhood Attention Transformer Sharpens AI Segmentation of Polyps and Skin Lesions</title>
		<link>https://scienmag.com/neighborhood-attention-transformer-sharpens-ai-segmentation-of-polyps-and-skin-lesions/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:22:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced segmentation algorithms]]></category>
		<category><![CDATA[AI in colonoscopy]]></category>
		<category><![CDATA[AI skin lesion analysis]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[attention-based neural networks]]></category>
		<category><![CDATA[automated cancer detection]]></category>
		<category><![CDATA[colonoscopy]]></category>
		<category><![CDATA[colorectal polyp detection]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[dermoscopy]]></category>
		<category><![CDATA[dermoscopy image analysis]]></category>
		<category><![CDATA[Dice score]]></category>
		<category><![CDATA[medical image analysis benchmarks]]></category>
		<category><![CDATA[medical image segmentation]]></category>
		<category><![CDATA[multi-scale feature fusion]]></category>
		<category><![CDATA[NAHFormer]]></category>
		<category><![CDATA[neighborhood attention]]></category>
		<category><![CDATA[Neighborhood Attention Transformer]]></category>
		<category><![CDATA[polyp segmentation]]></category>
		<category><![CDATA[skin lesion segmentation]]></category>
		<category><![CDATA[tissue slide analysis]]></category>
		<category><![CDATA[Transformer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202724</guid>

					<description><![CDATA[Researchers in China have developed NAHFormer, a transformer-based segmentation framework using neighborhood attention and hierarchical feature fusion that outperforms state-of-the-art methods on colonoscopy polyp and dermoscopy skin lesion benchmarks.]]></description>
										<content:encoded><![CDATA[<p>A new artificial intelligence architecture that reads medical images the way a pathologist reads tissue slides, scanning small local neighborhoods before zooming out to grasp the whole picture, is reporting some of the strongest segmentation numbers yet achieved on two notoriously difficult clinical tasks: tracing colorectal polyps in colonoscopy video frames and outlining suspicious skin lesions in dermoscopy photographs. The system, called NAHFormer, was developed by Xuehui Yin, Haonan Li, Tianxiao Hou and Chao Tang at the School of Computer Science and Technology of Chongqing University of Posts and Telecommunications, and is described in a study published in the journal Applied Intelligence. Across seven publicly available benchmark datasets, the framework consistently outperformed state-of-the-art segmentation methods, reaching a mean Dice similarity coefficient of 0.821 on the challenging ETIS polyp dataset and 0.909 on the ISIC 2018 skin lesion dataset, two results that place it at the leading edge of what automated medical image analysis currently achieves on these benchmarks.</p>
<p>The clinical motivation behind the work is straightforward and consequential. Colorectal cancer remains one of the most common and deadly malignancies worldwide, and most of these cancers arise from polyps, small growths protruding from the lining of the colon that can be removed during routine colonoscopy before they turn malignant. Detecting and delineating those polyps accurately is therefore a direct determinant of whether a lesion is excised in time. Dermoscopy, meanwhile, is the dermatologist&#8217;s principal tool for the early detection of melanoma and other skin cancers, and the precise boundary of a lesion carries decisive diagnostic weight, informing decisions about biopsy, excision margins and follow-up. In both settings, clinicians today rely on manual or semi-manual outlining of lesions in images, a process that is slow, subjective and prone to inter-observer variability. An algorithm that could reliably draw those contours automatically would not only save expert time but could also flag subtle lesions that tired human eyes might miss.</p>
<p>Getting a computer to draw those contours, however, has proven remarkably stubborn. Polyps in endoscopic images vary enormously in morphology, size and appearance; some are flat and barely distinguishable in color from the surrounding mucosa, others sit at the edge of the frame, partially obscured by specular highlights, bubbles, or motion blur. Skin lesions present a parallel problem, with irregular, sometimes fading boundaries that blend into healthy skin. Early deep learning solutions, most famously the U-Net family of convolutional networks, achieved strong results by repeatedly downsampling and upsampling images and stitching fine local texture onto coarse semantic context. But convolutional kernels see only a small window at a time, so these networks struggle to connect distant parts of an image into a coherent global understanding, and they often produce ragged or leaking boundaries around small or low-contrast objects. Transformers, which use self-attention to let every pixel consult every other pixel, solved the long-range dependency problem but introduced a new one: the computational cost of global attention grows quadratically with the number of image tokens, making naive transformer segmentation expensive and sometimes imprecise at fine scales.</p>
<p>NAHFormer&#8217;s answer begins at the front of the network. The framework employs a pyramid-structured Mix Transformer, or MiT, encoder, a hierarchical backbone in which the image is progressively broken into larger patches through successive stages, producing a ladder of feature maps that descend in spatial resolution while ascending in semantic richness. Early stages capture delicate texture and fine edges, essential for tracing the exact rim of a polyp or lesion, while later stages encode the broader context needed to say, with confidence, that a given blob is a lesion at all. This multi-scale representation, the authors argue, is what underpins the model&#8217;s generalization capability, allowing it to cope with the extraordinary diversity of lesion appearances it encounters across different patients, cameras, imaging conditions and anatomical sites rather than overfitting to the quirks of any single dataset.</p>
<p>The first bespoke innovation sits above that encoder: a Cross-Resolution Semantic Perception, or CRSP, module. Its job is to make different layers of the pyramid talk to each other in a structured way, integrating semantic information across multiple resolutions so that the coarse layers&#8217; understanding of what a lesion is can guide the fine layers&#8217; placement of its boundary. The module performs this cross-resolution conversation through neighborhood attention, an attention mechanism in which each token attends not to the entire image but only to a small window of its nearest neighbors. That locality restriction, borrowed from the neighborhood attention transformer introduced by computer vision researchers in 2023, dramatically reduces computation compared with global self-attention while retaining most of the discriminative power, because in dense prediction tasks the pixels most relevant to resolving a given boundary are typically its immediate surroundings. The practical payoff, according to the team, is precise delineation of lesion contours and boundaries, precisely the capability where earlier transformer models have tended to stumble.</p>
<p>The second innovation, a Hierarchical Feature Fusion, or HFF, module, tackles a quieter but equally corrosive problem: redundancy. In multi-scale architectures, as features from different levels are combined, the same information is often carried forward repeatedly, and low-level detail can be swamped or diluted by what are effectively duplicate signals from coarser scales. The HFF module progressively aggregates the multi-scale feature hierarchy in a stepwise fashion, deliberately suppressing redundant information as it fuses, so that the final segmentation head receives a cleaner, more informative summary of the image. This redundancy-aware aggregation, combined with the boundary-sensitive attention of the CRSP module, is what the authors describe as the framework&#8217;s distinctive pairing of redundancy-aware and boundary-aware capabilities, a combination intended to raise accuracy without inflating the network into an unwieldy behemoth.</p>
<p>The experimental evidence spans the field&#8217;s most widely used public benchmarks. On the polyp side, the team evaluated NAHFormer on five colonoscopy datasets: Kvasir, CVC-ClinicDB, CVC-ColonDB, EndoScene and ETIS. These sets differ substantially in acquisition conditions and difficulty; CVC-ClinicDB contains relatively clean and well-lit frames, while ETIS is widely regarded as the sternest test, composed of small, poorly contrasted polyps that routinely drag down the scores of models that excel elsewhere. On the dermatology side, the framework was tested on the ISIC 2017 and ISIC 2018 skin lesion archives, curated through the International Skin Imaging Collaboration&#8217;s annual challenges for melanoma detection. The pattern of results is telling: NAHFormer did not merely top the leaderboard on the easy datasets but held its advantage on the hard ones, achieving its reported mean Dice score of 0.821 on ETIS and 0.909 on ISIC 2018, indicating that its architectural choices translate into genuine robustness rather than dataset-specific tuning.</p>
<p>The Dice coefficient, the metric at the center of these comparisons, measures the overlap between the algorithm&#8217;s predicted segmentation and the ground truth annotated by experts, ranging from zero for no overlap to one for perfect agreement, so the difference between a middling model and a strong one often comes down to whether the predicted boundary clips a sliver off the lesion or lets background tissue bleed in. Success on ETIS, where polyps are small and faintly demarcated, suggests that the neighborhood attention scheme is doing exactly what it was designed to do, pinning down delicate boundaries without the blurring that coarse feature fusion can introduce. The strong ISIC result reinforces the point from the opposite direction, since skin lesion boundaries are often biologically diffuse, fading gradually into surrounding skin, and demand a different kind of perceptual finesse. A single architecture scoring highly across both regimes hints at a level of adaptability that previous task-specific models, whether convolutional designs like PraNet or hybrid transformer convolutional systems such as TransFuse, Swin-UNet, MissFormer and H2Former, have generally had to trade off against one another.</p>
<p>The broader significance of the work lies in what it suggests about the trajectory of clinical AI. Rather than choosing between the precision of convolutional networks and the global reasoning of transformers, the Chongqing team&#8217;s design stitches the two virtues together, using a hierarchical transformer backbone, locality-restricted attention and progressive fusion to keep both computation and accuracy in a clinically usable range. All of the datasets used in the study are publicly available, which means other groups can immediately stress-test the claims and build on the architecture, and the work was supported in part by the National Natural Science Foundation of China under Grant 61701060 and by a Chongqing Graduate Scientific Research Innovation Project under Grant CYS25470. The researchers frame the system as a step toward segmentation tools that could eventually assist physicians in real time, flagging and outlining lesions as images are captured. For now, the result stands as a demonstration that rethinking where an attention mechanism should look, and what it should ignore, can push the frontier of medical image analysis forward on the benchmarks that matter.</p>
<p><strong>Subject of Research:</strong> A transformer-based deep learning framework for segmenting colonoscopic polyps and dermoscopic skin lesions in medical images.</p>
<p><strong>Article Title:</strong> NAHFormer: Neighborhood self-attention and hierarchical feature fusion transformer for image medical segmentation</p>
<p><strong>Article References:</strong> Yin, X., Li, H., Hou, T., &amp; Tang, C. (2026). NAHFormer: Neighborhood self-attention and hierarchical feature fusion transformer for image medical segmentation. <em>Applied Intelligence, 56</em>(15), Article 430. <a href="https://doi.org/10.1007/s10489-026-07474-w" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07474-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07474-w" rel="noopener noreferrer">10.1007/s10489-026-07474-w</a></p>
<p><strong>Keywords:</strong> medical image segmentation, transformer, neighborhood attention, polyp segmentation, skin lesion segmentation, colonoscopy, dermoscopy, deep learning, computer vision, multi-scale feature fusion, Dice score, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202724</post-id>	</item>
	</channel>
</rss>
