<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Mamba &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/mamba/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 23:53:08 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Mamba &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Wavelets Meet Mamba: AI Reads Shoulder X-rays to Spot Rotator Cuff Tears</title>
		<link>https://scienmag.com/wavelets-meet-mamba-ai-reads-shoulder-x-rays-to-spot-rotator-cuff-tears/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 23:53:08 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI-based shoulder injury assessment]]></category>
		<category><![CDATA[artificial intelligence in medical imaging]]></category>
		<category><![CDATA[BMC Medical Imaging]]></category>
		<category><![CDATA[computer-aided diagnosis]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for radiograph interpretation]]></category>
		<category><![CDATA[frequency domain analysis in medical AI]]></category>
		<category><![CDATA[frequency domain features]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[Mamba neural network architecture]]></category>
		<category><![CDATA[medical image classification]]></category>
		<category><![CDATA[medical imaging for musculoskeletal disorders]]></category>
		<category><![CDATA[MedMamba]]></category>
		<category><![CDATA[rapid and affordable shoulder injury screening]]></category>
		<category><![CDATA[resource-limited healthcare diagnostics]]></category>
		<category><![CDATA[rotator cuff tear]]></category>
		<category><![CDATA[rotator cuff tear detection]]></category>
		<category><![CDATA[shoulder pain]]></category>
		<category><![CDATA[shoulder X-ray analysis]]></category>
		<category><![CDATA[soft tissue injury diagnosis]]></category>
		<category><![CDATA[state-space model]]></category>
		<category><![CDATA[wavelet transform]]></category>
		<category><![CDATA[wavelet-enhanced deep learning]]></category>
		<category><![CDATA[X-ray imaging]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=245834</guid>

					<description><![CDATA[Researchers in China have developed WEM-Mamba, a wavelet-enhanced state space model that classifies rotator cuff tears from plain shoulder X-rays with high accuracy and a compact computational footprint.]]></description>
										<content:encoded><![CDATA[<p>Rotator cuff tears are among the most common culprits behind shoulder pain and lost mobility, yet the imaging workhorse used to detect them is often out of reach for many patients. Magnetic resonance imaging remains the gold standard for evaluating the soft tendons of the shoulder, but it is expensive, time-consuming, and not always available, particularly in rural or resource-limited settings. Plain shoulder X-rays, by contrast, are fast, cheap, and universally accessible, which is why they dominate initial clinical assessments. The problem is that tendons barely show up on radiographs, so identifying a tear from an X-ray alone is notoriously difficult even for experienced clinicians. A team of researchers in Urumqi, China, now reports a new artificial intelligence framework designed to close that gap, squeezing diagnostic information out of X-ray images that the human eye and conventional algorithms tend to miss.</p>
<p>The study, published in BMC Medical Imaging by a multidisciplinary group from Xinjiang Medical University and its Sixth Affiliated Hospital, introduces a model called WEM-Mamba, short for Wavelet-Enhanced MedMamba. The work sits at the intersection of two of the hottest trends in medical deep learning: state space models, the efficient new architecture family that includes Mamba, and frequency domain analysis, which extracts information about textures and edges that spatial-domain models often overlook. According to the authors, most existing approaches, whether convolutional neural networks, vision transformers, or earlier Mamba variants, concentrate almost exclusively on modeling features in the spatial domain, leaving the rich frequency content of medical images underexploited. For a task like rotator cuff classification, where subtle textural signatures in the radiograph may hint at tendon damage, that omission could mean losing exactly the evidence that matters most.</p>
<p>To understand why frequency matters, it helps to think about what an X-ray actually encodes. Radiographs are dense grids of intensity values, and while clinicians read them as pictures of bones and soft tissue shadows, mathematically they can be decomposed into components that vary at different spatial scales and orientations. Fine, rapidly changing patterns correspond to high-frequency content such as edges and fine texture; smooth, slowly varying patterns correspond to low-frequency content such as broad anatomical structures. The Haar wavelet transform, a classical mathematical tool that dates back decades, provides a principled way to split an image into these frequency bands while keeping the spatial location of each feature intact. WEM-Mamba exploits this by embedding a Haar wavelet decomposition step directly inside its core processing module, so the network can examine the image simultaneously through spatial and frequency lenses.</p>
<p>Technically, the framework builds on MedMamba, a medical imaging variant of the Mamba architecture. Mamba and its relatives are built on state space models, which process sequences, including sequences of image patches, using recurrent-style computations that scale linearly with input length rather than quadratically as attention mechanisms in transformers do. That efficiency makes state space models attractive for medical imaging, where high-resolution images and limited training data are the norm. The researchers&#8217; key innovation is the Wavelet-Enhanced SS-Conv-SSM module, or WESS. In this module, an input feature map is decomposed by the Haar wavelet transform into multiple frequency sub-bands. These sub-bands are processed and then synergistically fused with the global context modeling that the state space model provides, before an inverse wavelet transform can reconstruct the enhanced representation. The result is a hybrid design in which convolutional operations capture local detail, the state space model captures long-range dependencies across the image, and the wavelet pathway injects multi-band frequency features that neither of the other two components would extract on their own.</p>
<p>The team evaluated WEM-Mamba on a dataset of shoulder X-ray images collected retrospectively from routine clinical examinations at the Sixth Affiliated Hospital of Xinjiang Medical University, with ethics approval and anonymized patient data. They benchmarked their model against fifteen representative competitors spanning the major architecture families: convolutional neural networks, transformers, multi-layer perceptrons, and Mamba variants. WEM-Mamba came out on top, achieving an accuracy of 0.8950, an F1-score of 0.9309, a precision of 0.9078, a recall of 0.9552, and an area under the curve of 0.9116. Notably, the recall figure, which measures how many true tears the model successfully flags, is the highest of its headline metrics, a property that matters greatly in screening contexts where missed diagnoses carry the greatest cost. Equally important for real-world deployment, the model is compact: it contains 14.92 million parameters and requires only 2.04 gigafloating-point operations per inference, a footprint modest enough to run on ordinary clinical hardware rather than specialized computing clusters.</p>
<p>The performance-to-cost ratio is one of the study&#8217;s most compelling aspects. Deep learning models in medical imaging have often chased accuracy at the expense of size and speed, producing architectures too heavy for the very hospitals that need them most. By combining the linear-scaling efficiency of state space models with the lightweight Haar wavelet transform, the authors demonstrate that careful architectural design, rather than brute-force scale, can deliver strong classification performance on modest computational budgets. For clinics in underserved regions, where shoulder X-rays are already the first-line examination but specialist radiology expertise may be scarce, a compact model that augments the initial radiograph with a tear probability estimate could meaningfully accelerate referrals for MRI and definitive treatment.</p>
<p>The research also carries a broader message for the machine learning community. Wavelet-based feature extraction has a long pedigree in signal processing and image compression, but it has been somewhat eclipsed in the deep learning era by purely learned representations. The success of WEM-Mamba suggests that classical mathematical transforms and modern sequence models are not rivals but complements. The wavelet decomposition imposes a physically meaningful inductive bias, telling the network explicitly where the fine textures and edges live, while the state space model supplies the flexible global reasoning that fixed transforms cannot. In domains where training data are limited, as is nearly always the case in medicine, building such structural knowledge into the architecture may prove more data-efficient than expecting a generic network to rediscover it from scratch.</p>
<p>Caveats remain, and the authors are candid about them. The study was conducted on a single-center dataset, and all conclusions are explicitly framed as valid under the data and experimental settings of this particular investigation. The researchers state that further validation on larger-scale, multi-center, and external datasets is required before the approach can be considered robust, and that the clinical application value of the framework needs continued assessment in subsequent research. Single-center datasets can encode site-specific scanner settings, patient demographics, and imaging protocols, all of which can inflate apparent performance when a model is tested on data from the same source. Multi-center trials with external test sets are the standard next step for any diagnostic AI aspiring to clinical use, and this work is no exception.</p>
<p>Nevertheless, the study represents a thoughtful and timely contribution to computer-aided diagnosis of musculoskeletal injuries. It addresses a genuine clinical bottleneck, the gap between cheap first-line X-rays and definitive but costly MRI, with a method that is technically novel, computationally lean, and empirically competitive. The work was supported by a range of Chinese research funding programs, including the National Natural Science Foundation of China and multiple Xinjiang regional initiatives, and the article is published open access, making the full technical details available to researchers worldwide. If the promised multi-center validation bears out the single-center results, wavelet-enhanced state space models like WEM-Mamba could become a standard tool in the radiology toolbox, helping ensure that a patient&#8217;s first, cheapest image is not the last word on their shoulder pain.</p>
<p><strong>Subject of Research:</strong> A wavelet-enhanced Mamba-based deep learning framework for classifying rotator cuff tears in shoulder X-ray images</p>
<p><strong>Article Title:</strong> WEM-Mamba: a Wavelet-Enhanced MedMamba framework for rotator cuff tear classification in X-ray images</p>
<p><strong>Article References:</strong> Wang, Z., Liu, Z., Li, J., Aboduaini, Y., Abulaiti, A., Li, S., Yakufu, M., Halike, A., Xv, L., &amp; Li, G. (2026). WEM-Mamba: a Wavelet-Enhanced MedMamba framework for rotator cuff tear classification in X-ray images. <em>BMC Medical Imaging</em>. <a href="https://doi.org/10.1186/s12880-026-02740-2" rel="noopener noreferrer">https://doi.org/10.1186/s12880-026-02740-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12880-026-02740-2" rel="noopener noreferrer">10.1186/s12880-026-02740-2</a></p>
<p><strong>Keywords:</strong> rotator cuff tear, deep learning, Mamba, state space model, wavelet transform, frequency domain features, X-ray imaging, medical image classification, computer-aided diagnosis, shoulder pain, MedMamba, BMC Medical Imaging</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">245834</post-id>	</item>
		<item>
		<title>Mamba AI Reads Breast Scans Slice by Slice to Find Hidden Cancers</title>
		<link>https://scienmag.com/mamba-ai-reads-breast-scans-slice-by-slice-to-find-hidden-cancers/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 05:20:55 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[3D mammography slice analysis]]></category>
		<category><![CDATA[AI-based breast cancer diagnosis]]></category>
		<category><![CDATA[AI-powered breast cancer detection accuracy]]></category>
		<category><![CDATA[BCS-DBT dataset]]></category>
		<category><![CDATA[breast cancer]]></category>
		<category><![CDATA[breast cancer detection]]></category>
		<category><![CDATA[CenterNet]]></category>
		<category><![CDATA[computer-aided diagnosis]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for medical imaging]]></category>
		<category><![CDATA[Dense breast tissue tumor detection]]></category>
		<category><![CDATA[digital breast tomosynthesis]]></category>
		<category><![CDATA[Digital breast tomosynthesis (DBT)]]></category>
		<category><![CDATA[Early breast cancer screening technologies]]></category>
		<category><![CDATA[feature propagation]]></category>
		<category><![CDATA[Inter-slice feature propagation]]></category>
		<category><![CDATA[lesion detection]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[medical imaging AI]]></category>
		<category><![CDATA[screening cascade]]></category>
		<category><![CDATA[Selective state-space models in medical AI]]></category>
		<category><![CDATA[state-space models]]></category>
		<category><![CDATA[TomoMamba deep learning framework]]></category>
		<category><![CDATA[Volumetric breast imaging analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225862</guid>

					<description><![CDATA[Researchers have unveiled TomoMamba, a two-stage AI framework that uses selective state-space models to propagate features across DBT slices, achieving strong screening performance while candidly reporting the limits of its cancer discrimination.]]></description>
										<content:encoded><![CDATA[<p>Breast cancer remains the most frequently diagnosed malignancy among women worldwide, with roughly 2.3 million new cases and 685,000 deaths each year. The stakes of catching it early could hardly be higher: the five-year relative survival rate for localized disease exceeds 99 percent, but falls sharply once the disease advances. Digital breast tomosynthesis, or DBT, was developed to help. Unlike conventional two-dimensional mammography, DBT reconstructs the breast into a stack of thin image slices, reducing the tissue overlap that can hide tumors, especially in women with dense breast tissue. Yet that same volumetric richness has created a stubborn problem for artificial intelligence: most automated systems analyze each slice in isolation, throwing away the depth-wise continuity that makes DBT valuable in the first place.</p>
<p>A team of researchers from Qatar University and the University of Hyogo has now taken a direct swing at that problem. In a study published in Multimedia Tools and Applications, they introduce TomoMamba, a two-stage deep learning framework that, according to the authors, is the first application of selective state-space models to inter-slice feature propagation in DBT. The work arrives at a moment when the Mamba architecture, a relative newcomer to the deep learning scene, has been making waves for its ability to model long sequences with linear computational cost, in contrast to the quadratic cost of transformer attention.</p>
<p>The framework&#8217;s design mirrors the clinical workflow of breast screening. Stage 1 acts as a high-throughput triage system: a dual-branch 2.5D convolutional network processes the craniocaudal and mediolateral oblique views of each breast through a shared ResNet-18 backbone, aggregates information across slices with a learnable attention mechanism, and fuses the two view-level representations into a single verdict of normal or suspicious. On the validation partition of the public BCS-DBT dataset, this screening stage achieved a breast-level area under the receiver operating characteristic curve of 0.97, correctly filtering out 94.5 percent of normal breasts while identifying 36 of 41 suspicious ones.</p>
<p>Only breasts flagged as suspicious proceed to Stage 2, where the real architectural novelty lives. A ResNet-18 backbone with a lightweight feature pyramid network extracts features from each of the fifteen representative slices retained per volume. Then, at every spatial position in the feature map, the model extracts a fifteen-element sequence of slice-wise features and feeds it through a Mamba selective state-space model. The crucial distinction from earlier recurrent approaches is that Mamba&#8217;s state-transition matrices are input-dependent: the amount each slice contributes to the running hidden state varies with the content of that slice. A lesion that activates strongly on two or three adjacent slices therefore leaves a depth summary that carries information about where along the depth axis the signal appeared, something slice-independent methods discard entirely and full 3D convolutions capture only at substantial memory cost.</p>
<p>The Mamba output is folded back into the per-slice features through a gated residual connection, a sigmoid gate that makes per-channel keep-or-suppress decisions conditioned on the depth context. Introspection across 72 validation volumes showed the gate is channel-selective rather than slice-selective: at lesion positions a mean of 40 of 128 channels fell outside the neutral band, against 23.5 at background positions, a contrast of 1.70 times with a large effect size. Downstream, an anchor-free CenterNet-style detector predicts heatmaps, box sizes and offsets directly from the Mamba-enhanced features, and classification is performed on small feature patches cropped at the detected lesion centers rather than pooled globally. That cropping choice proved essential: a pooling-only baseline oscillated between all-cancer and all-benign predictions without ever converging on the tiny training set of 200 biopsied volumes from just 101 patients.</p>
<p>The evaluation rests on BCS-DBT, hosted on The Cancer Imaging Archive and collected at Duke University Hospital, which remains the only public DBT dataset combining biopsy-confirmed pathology labels with lesion-level bounding-box annotations. The class distribution is extreme, with normal cases outnumbering cancers by more than 100 to 1, which is precisely why the authors split the problem in two. On the validation partition, Stage 2 attained a patient-level AUC of 0.80 with a lesion detection rate of 72.0 percent. On the untouched test partition, however, classification fell to 0.62, an interval that spans chance, while detection degraded only modestly to 62.5 percent. The authors are unusually candid about this: they present lesion localization, not benign-versus-cancer discrimination, as Stage 2&#8217;s demonstrated contribution.</p>
<p>Chained together, the two stages reached a breast-level AUC of 0.93 on validation and 0.92 on the held-out test set, forwarding just 14 percent of breasts for expensive diagnostic analysis while losing two of 20 cancers at screening. The authors warn that this cascade figure reflects the screening stage&#8217;s normal-versus-suspicious separation rather than genuine cancer discrimination, and that roughly one cancer in ten would be missed if the cascade were used as a standalone filter. They position the system as a reader-support tool to be used alongside, not in place of, radiologist interpretation.</p>
<p>The controlled baseline comparisons are where the architectural argument gets its sharpest test. Under identical data, preprocessing and hyperparameters, TomoMamba was pitted against a bidirectional GRU, a transformer variant and a slice-independent baseline, all matched to within one percent on parameter count. On the validation partition, the cross-slice architectures were not separable by classification AUC, and the authors decline to rank them on that metric. But at matched false-positive rates, TomoMamba localized more lesions than every alternative on validation, and on the test partition it achieved the highest patient-level AUC and the lowest Brier score of the five configurations. Most tellingly, removing the gated cross-slice residual dropped patient-level AUC from 0.798 to 0.548, the only architectural ablation that reached statistical significance at 0.009.</p>
<p>The team also confronted a subtle methodological hazard head-on. Because radiologist annotations guided which slices were retained during preprocessing, the model could in principle have exploited the annotated slice&#8217;s position as a hidden marker. The authors tested this directly: detection rates showed no significant difference depending on where the annotated slice sat in the sequence, and under a fully annotation-independent selection procedure, patient-level classification changed little, falling from 0.798 to 0.748 on validation. Since annotations do not exist at deployment, they treat the annotation-independent figures as the deployment-relevant results. Failure analysis added further nuance: the dominant failure mode was classification rather than localization, with most misclassifications clustered within 0.03 of the decision threshold, suggesting a calibration problem rather than a detection one.</p>
<p>The limitations are stated with rare honesty. The study is single-institution, single-scanner, and underpowered for benign-versus-cancer classification with only 101 biopsied training patients; no published DBT method has yet demonstrated cross-site transfer. The much-hyped efficiency of state-space models also did not materialize at this scale: at fifteen slices per volume, Mamba&#8217;s sequential scan was actually slower than both the recurrent and attention baselines, though its linear scaling should win out on deeper acquisitions. Still, the significance of the work lies in its disciplined template for evaluating new architectures on small clinical datasets, reporting all three official partitions, quantifying seed variability, and refusing to claim more than the data can support. As larger annotated DBT cohorts emerge, the question of whether selective state-space propagation becomes the backbone of volumetric cancer screening will finally get the data it deserves.</p>
<p><strong>Subject of Research:</strong> Deep learning with selective state-space models for cross-slice analysis in digital breast tomosynthesis for breast cancer screening and diagnosis</p>
<p><strong>Article Title:</strong> TomoMamba: a two-stage model with state-space cross-slice propagation for breast cancer diagnosis in digital breast tomosynthesis</p>
<p><strong>Article References:</strong> Soliman, S., Zafari, Y., Rashed, E. A., &amp; Mabrok, M. (2026). TomoMamba: a two-stage model with state-space cross-slice propagation for breast cancer diagnosis in digital breast tomosynthesis. <em>Multimedia Tools and Applications, 85</em>(10), Article 776. <a href="https://doi.org/10.1007/s11042-026-21943-x" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21943-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21943-x" rel="noopener noreferrer">10.1007/s11042-026-21943-x</a></p>
<p><strong>Keywords:</strong> digital breast tomosynthesis, breast cancer, Mamba, state-space models, deep learning, lesion detection, computer-aided diagnosis, BCS-DBT dataset, medical imaging AI, CenterNet, screening cascade, feature propagation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225862</post-id>	</item>
		<item>
		<title>Compact AI Network Blends Mamba and Polynomial Neurons for Medical Image Segmentation</title>
		<link>https://scienmag.com/compact-ai-network-blends-mamba-and-polynomial-neurons-for-medical-image-segmentation/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 23:23:37 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[blood vessel and lesion segmentation]]></category>
		<category><![CDATA[compact neural network architecture]]></category>
		<category><![CDATA[computational efficiency in medical imaging]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[dermoscopy]]></category>
		<category><![CDATA[global and local feature extraction]]></category>
		<category><![CDATA[hybrid neural network design]]></category>
		<category><![CDATA[long-range context propagation]]></category>
		<category><![CDATA[low-parameter deep learning models]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[Mamba state-space model]]></category>
		<category><![CDATA[medical image segmentation]]></category>
		<category><![CDATA[neural architecture]]></category>
		<category><![CDATA[parameter efficiency]]></category>
		<category><![CDATA[polynomial expansion]]></category>
		<category><![CDATA[polynomial neurons]]></category>
		<category><![CDATA[polyp segmentation]]></category>
		<category><![CDATA[retinal vessels]]></category>
		<category><![CDATA[state-space models]]></category>
		<category><![CDATA[transformer alternatives for image analysis]]></category>
		<category><![CDATA[U-Net]]></category>
		<category><![CDATA[U-shaped encoder-decoder]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224290</guid>

					<description><![CDATA[Researchers have unveiled ANNE, a compact segmentation network that pairs Mamba state-space modeling with adaptive polynomial neuron expansion to match much larger AI models across eight medical imaging benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Medical image segmentation, the task of tracing the exact outlines of lesions, polyps, and blood vessels in clinical scans, has long been dominated by a trade-off. Convolutional networks excel at capturing fine local detail but struggle to connect distant regions of an image, while transformer-based models reason globally at a steep computational price. A new architecture called ANNE, described in the journal Applied Intelligence by Abel A. Reyes-Angulo of Michigan Technological University and Sidike Paheding of Fairfield University, argues that this trade-off can be broken with a surprisingly compact design that needs only about 2.77 million trainable parameters.</p>
<p>The central insight behind ANNE is that two very different mathematical tools can do complementary jobs inside a single U-shaped encoder-decoder. In the encoder, the authors deploy Mamba, a selective state-space model that has attracted wide attention since its introduction as a linear-time alternative to transformer attention. Instead of building a quadratic attention matrix over every pair of pixels, Mamba propagates information along a flattened sequence of the feature map through input-dependent state parameters, allowing long-range context to flow through the network at a cost that scales linearly with image size. This is what lets ANNE reason about broad spatial relationships without the memory burden that typically accompanies global modeling.</p>
<p>The second ingredient, and the paper&#8217;s main methodological novelty, is a module the authors call Adaptive Progressive Expansion Neurons, or A-PEN. The idea builds on a line of prior work known as progressively expanded neuron models, which increase the nonlinear expressiveness of a network by computing polynomial terms of the input features, such as squared and cubed activations, and combining them with learned weights. In earlier designs, however, the set of expansion terms and the activation applied afterward were fixed in advance. A-PEN makes both adaptive: a bank of candidate polynomial branches is weighted by normalized gates derived from trainable logits through a softmax function, and the post-expansion nonlinearity is a parametric PReLU whose negative slope is learned during training.</p>
<p>The authors are notably careful about how these gates should be interpreted. Because each normalized gate multiplies an unconstrained trainable coefficient, the effective contribution of a polynomial branch is the product of the two, meaning different gate-coefficient combinations can produce identical outputs. The gates therefore form a coupled parameterization that modulates the relative scaling of candidate branches during optimization, not a uniquely identifiable measure of how important each polynomial order is. The team also draws a precise boundary around its use of Remez approximation: the classical minimax theory of best polynomial approximation motivates the idea that compact polynomial bases can represent complex nonlinear mappings, but ANNE never executes the Remez exchange algorithm and claims no minimax guarantee. The polynomial coefficients are simply learned end to end from the segmentation loss.</p>
<p>Architecturally, ANNE follows the familiar U-shaped template that has defined biomedical segmentation since U-Net. Convolutional stem layers preserve local spatial structure at the input, Mamba blocks then model long-range dependencies at progressively coarser resolutions, and A-PEN blocks refine nonlinear feature interactions at the bottleneck and in the decoder before upsampling. Skip connections carry fine-grained encoder information to matching decoder stages, and the decoder employs multi-scale depth-wise convolutions with kernel sizes of one, three, and five. A final projection produces a binary segmentation map, with the whole pipeline trained using a weighted combination of binary cross-entropy and soft Dice loss.</p>
<p>The evaluation is unusually transparent about how comparisons were made. The authors tested ANNE on eight public 2D datasets spanning three clinical domains: dermoscopic lesion segmentation on ISIC 2016, ISIC 2017, ISIC 2018, and PH2; colonoscopic polyp segmentation on Kvasir-SEG and CVC-ClinicDB; and retinal-vessel segmentation on DRIVE and CHASE_DB1. Crucially, the results tables distinguish controlled reproductions, where competing methods were retrained under an identical protocol with the same splits, preprocessing, loss, and five fixed random seeds, from literature-reported values, which may reflect different experimental conditions. Statistical claims are restricted to the controlled comparisons, while literature numbers are described only as context.</p>
<p>Within that framework, ANNE delivers a strong and consistent showing. On ISIC 2018 and PH2 it achieved the strongest tabulated values among the compared methods, and on ISIC 2017 it recorded the lowest 95th-percentile Hausdorff distance, a boundary-sensitive metric, even though another method reported higher overlap scores. In polyp segmentation it topped the tables on Kvasir-SEG across Dice, IoU, and boundary distance, and placed second on CVC-ClinicDB with the best boundary accuracy. For retinal vessels, where thin, branching structures test a model&#8217;s ability to preserve fine detail, ANNE achieved the best tabulated IoU and boundary distance on DRIVE and swept all three metrics on CHASE_DB1.</p>
<p>Ablation experiments illuminate why the design works. Comparing a compact U-Net-like baseline against versions augmented with only A-PEN, only Mamba, or both, the authors found that A-PEN consistently improved boundary-sensitive performance while Mamba tended to boost overlap where broad context matters. The full model achieved the best Dice score in every ablation row, supporting the claim that the two modules are complementary rather than interchangeable. A separate resolution analysis showed that accuracy improved as input size rose from 128 by 128 to 256 by 256 pixels, but saturated beyond that point even as computational cost continued to climb roughly quadratically, making 256 by 256 the identified sweet spot for lesion and polyp tasks.</p>
<p>The authors are candid about the limits of their evidence. All experiments used public 2D benchmarks, so nothing is established about 3D volumes, video, or prospective clinical workflows. The evaluation is in-distribution, with no external cohorts from different scanners, institutions, or populations, and several benchmarks are small enough that uncertainty remains despite repeated runs. Efficiency was measured on an NVIDIA A100 GPU, which characterizes relative cost but says nothing about real-time performance on edge devices or hospital hardware. The team also notes that segmentation accuracy alone is insufficient for clinical deployment, which would additionally demand uncertainty estimation, failure detection, calibration, fairness analysis, and workflow integration.</p>
<p>Even with those caveats, ANNE represents a compelling data point in the ongoing effort to shrink medical AI without shrinking its competence. By pairing a linear-time state-space encoder with learnable polynomial feature expansion, the architecture shows that competitive overlap and boundary accuracy across dermoscopy, colonoscopy, and fundus imaging can be achieved with a parameter budget far below that of transformer-heavy rivals. The code, dataset split manifests, and configuration files have been released in a public repository, and the authors point toward volumetric extensions, external-domain validation, and deployment-oriented profiling on resource-constrained hardware as the next frontiers. For clinics and research labs where compute is scarce but precision is not, that combination of transparency, efficiency, and measured performance may prove just as influential as the architecture itself.</p>
<p><strong>Subject of Research:</strong> A parameter-efficient deep learning architecture combining Mamba state-space blocks and adaptive polynomial neuron expansion for 2D medical image segmentation</p>
<p><strong>Article Title:</strong> ANNE: adaptive neuron network expansion for 2D medical image segmentation</p>
<p><strong>Article References:</strong> Reyes-Angulo, A. A., &amp; Paheding, S. (2026). ANNE: adaptive neuron network expansion for 2D medical image segmentation. <em>Applied Intelligence, 56</em>(15), Article 464. <a href="https://doi.org/10.1007/s10489-026-07503-8" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07503-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07503-8" rel="noopener noreferrer">10.1007/s10489-026-07503-8</a></p>
<p><strong>Keywords:</strong> medical image segmentation, Mamba, state-space models, polynomial expansion, neural architecture, dermoscopy, polyp segmentation, retinal vessels, parameter efficiency, deep learning, U-Net, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224290</post-id>	</item>
		<item>
		<title>New Mamba AI Reads Rocks in Frequency and Space to Classify Lithology</title>
		<link>https://scienmag.com/new-mamba-ai-reads-rocks-in-frequency-and-space-to-classify-lithology/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 01:31:39 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[AI in mineralogy]]></category>
		<category><![CDATA[AI-based geological exploration tools]]></category>
		<category><![CDATA[applications of artificial intelligence in tunneling and mining]]></category>
		<category><![CDATA[automated rock type recognition]]></category>
		<category><![CDATA[CNN]]></category>
		<category><![CDATA[computational efficiency in geological AI models]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for lithology identification]]></category>
		<category><![CDATA[Earth Science Informatics]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[frequency-domain enhancement]]></category>
		<category><![CDATA[geological exploration]]></category>
		<category><![CDATA[geological image classification]]></category>
		<category><![CDATA[Grad-CAM]]></category>
		<category><![CDATA[lithology classification]]></category>
		<category><![CDATA[lithology classification in geosciences]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[mineral mineral arrangement detection]]></category>
		<category><![CDATA[rock images]]></category>
		<category><![CDATA[rock sample image analysis]]></category>
		<category><![CDATA[rock surface texture analysis]]></category>
		<category><![CDATA[SpecMamba-Net architecture]]></category>
		<category><![CDATA[state-space model]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215947</guid>

					<description><![CDATA[Researchers in China have unveiled SpecMamba-Net, a dual-branch Mamba-based AI that fuses frequency-domain texture cues with linear-complexity global state-space modeling to classify rock types with record accuracy.]]></description>
										<content:encoded><![CDATA[<p>Identifying the type of rock in a photograph, a drill core sample, or a tunnel face image may look like a task for a trained geologist with a hand lens, but it is increasingly a job for artificial intelligence. In a study published in Earth Science Informatics, a research team from Guilin University of Technology in China has introduced a deep learning architecture designed to do precisely that, and to do it with an unusual combination of accuracy and computational economy. The model, named SpecMamba-Net, was developed by Yan Zhang, Fan Zhang, Ruifeng Li, and Zongguo Zhang, and tackles a long-standing bottleneck in the automated classification of lithology: the difficulty of capturing both the fine texture of a rock surface and the large-scale arrangement of its minerals within a single, efficient framework.</p>
<p>Lithology identification is one of those deceptively simple tasks that underpins an enormous range of practical work. Geological exploration, mining, tunneling, drilling operations, and hazard assessment all depend on knowing whether the material at hand is granite, basalt, limestone, or something else entirely. Misclassification can propagate through an entire project, from inaccurate resource estimates to unstable excavation designs. Traditionally, the job has fallen to specialists who read visual cues such as grain size, color, foliation, and crystalline structure. Automating that expertise has proven attractive but technically stubborn, because rocks present image-analysis challenges that generic computer vision systems were never designed to handle.</p>
<p>The dominant tool for image classification over the past decade has been the convolutional neural network, or CNN. Convolutional layers excel at detecting local patterns, such as the edges of individual mineral grains or small repetitive textures, because each filter scans only a small neighborhood of pixels at a time. That locality is a strength for fine detail but a weakness for global structure. To understand how grains are distributed across an entire rock face, or how layered patterns stretch across the whole image, a CNN must stack many layers, each expanding its effective field of view. The result is models that can be large, computationally expensive, and still imperfect at representing long-range spatial relationships, which are exactly the relationships that distinguish many lithological classes.</p>
<p>Transformers, borrowed from natural language processing, offered one way forward by letting every part of an image attend to every other part. But that power comes at a steep price: the attention mechanism scales quadratically with the number of image patches, so the cost of processing grows rapidly as image resolution or model size increases. The newer alternative inspiring the Chinese team&#8217;s work is the state-space model, and in particular its vision-oriented incarnation known as Mamba. State-space models compress a sequence&#8217;s history into a compact hidden state and update it as they scan the input, which means they can model dependencies across an entire image with computational complexity that grows only linearly rather than quadratically. For geologists staring at high-resolution outcrop photographs, that scaling difference is not an academic nicety; it determines whether the model can run on practical hardware in the field or the lab.</p>
<p>SpecMamba-Net combines this linear-complexity global modeling with a second, complementary branch devoted to the frequency domain. The idea rests on a well-known property of image analysis: while spatial representations describe where features are, their Fourier transform describes which patterns repeat and at what scales. Different rock types produce characteristic signatures in the frequency spectrum, because texture is fundamentally a statistical statement about periodicity, granularity, and roughness. The spectral enhancement branch of SpecMamba-Net transforms feature representations into the frequency domain and amplifies the discriminative spectral components that separate one lithology from another, effectively giving the network a tuning fork for rock texture that the purely spatial branch lacks.</p>
<p>The two branches work in parallel. The state-space scanning branch sweeps across the image, building a representation of global spatial dependencies and long-range mineral arrangement patterns at linear cost. The spectral branch, meanwhile, extracts frequency-domain cues tied to texture. A gated fusion mechanism then serves as the arbiter, adaptively deciding how much weight to give spectral features versus spatial evolution features for each input. Rather than forcing the network to choose between texture and structure, the gating lets the model lean on whichever cue is more informative for a given specimen, producing what the authors describe as a more comprehensive lithology representation. This kind of adaptive fusion has become an increasingly popular strategy in remote sensing and hyperspectral imaging, and the study positions SpecMamba-Net within that broader lineage of dual-branch spectral-spatial architectures.</p>
<p>To test the design, the researchers evaluated the model on a seven-class lithology dataset, using five repeated stratified holdout splits to ensure that results were not an artifact of a single lucky train-test division. SpecMamba-Net achieved a mean test accuracy of 0.8654 with a standard deviation of 0.0155, and a macro-F1 score of 0.8667 plus or minus 0.0179, the highest mean performance among all models evaluated in the comparison. The macro-F1 figure is particularly meaningful in this context because it averages performance across classes, preventing a model from posting a flattering score by excelling on common categories while neglecting rare ones. In geology, where some lithologies are inherently overrepresented in field datasets, that class-balanced metric matters.</p>
<p>The authors were careful about statistical rigor in a way that is not always typical of deep learning papers in the geosciences. Paired comparisons at the split level showed statistically supported improvements over most baseline models, which included mainstream CNNs such as ResNet18 and a range of lightweight architectures. Notably, however, the difference from ResNet18 did not reach the conventional significance threshold, a candid acknowledgment that the classic residual network remains a formidable competitor. Equally important, Grad-CAM interpretability analysis provided qualitative evidence that SpecMamba-Net focused its attention on lithologically relevant texture and structural regions of the images rather than on background artifacts or irrelevant correlations. For a field in which scientists must trust automated judgments, that kind of visualizable attention is a genuine selling point.</p>
<p>The practical implications extend across several industries. In tunneling, rapid and reliable lithology identification from tunnel face images informs support design and safety decisions in real time. In drilling, cuttings analysis has already been explored as a target for intelligent image-based logging, and a model that balances accuracy against computational cost could bring such systems closer to deployment on rigs with limited hardware. In mining and geological hazard assessment, automated classification of field photographs could accelerate mapping campaigns that currently demand scarce expert attention. The linear scaling of the Mamba backbone makes the framework attractive wherever high-resolution imagery and constrained compute collide, which describes most field geology.</p>
<p>The work, published as volume 19, article 177 of Earth Science Informatics on 5 September 2026, also reflects a broader shift in how geoscientists think about machine learning architecture. Rather than importing the largest available vision model and fine-tuning it, the authors designed a domain-motivated structure: one branch that asks what textures repeat in the image and another that asks how mineral patterns are organized across it. The favorable balance between classification performance and computational efficiency reported in the study suggests that such purpose-built hybrids may define the next generation of geological AI tools. The datasets analyzed in the research are available from the corresponding author on reasonable request, and the work was supported by the National Natural Science Foundation of China, the Guangxi Science and Technology Program, and the Guangxi Key Laboratory of Geomechanics and Geotechnical Engineering.</p>
<p><strong>Subject of Research:</strong> Deep learning-based lithology classification using frequency-domain enhancement and visual state-space modeling</p>
<p><strong>Article Title:</strong> A novel Mamba model based on frequency domain enhancement and visual state space for high-precision lithology classification</p>
<p><strong>Article References:</strong> Zhang, Y., Zhang, F., Li, R., &amp; Zhang, Z. (2026). A novel Mamba model based on frequency domain enhancement and visual state space for high-precision lithology classification. <em>Earth Science Informatics, 19</em>(10), Article 177. <a href="https://doi.org/10.1007/s12145-026-02233-4" rel="noopener noreferrer">https://doi.org/10.1007/s12145-026-02233-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12145-026-02233-4" rel="noopener noreferrer">10.1007/s12145-026-02233-4</a></p>
<p><strong>Keywords:</strong> lithology classification, Mamba, state-space model, frequency-domain enhancement, deep learning, rock images, feature fusion, CNN, Grad-CAM, geological exploration, Earth Science Informatics, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215947</post-id>	</item>
		<item>
		<title>Mamba-Based AI Sharpens Medical Image Segmentation With Position Awareness and Dynamic Upsampling</title>
		<link>https://scienmag.com/mamba-based-ai-sharpens-medical-image-segmentation-with-position-awareness-and-dynamic-upsampling/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 02:24:44 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[clinical applications of medical image segmentation]]></category>
		<category><![CDATA[computational efficiency]]></category>
		<category><![CDATA[computational efficiency in medical image analysis]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning in medical imaging]]></category>
		<category><![CDATA[dynamic upsampling]]></category>
		<category><![CDATA[dynamic upsampling in medical image analysis]]></category>
		<category><![CDATA[high-resolution medical image processing]]></category>
		<category><![CDATA[long-range context capture in medical images]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[Mamba-based AI architecture]]></category>
		<category><![CDATA[medical image segmentation]]></category>
		<category><![CDATA[multi-organ segmentation]]></category>
		<category><![CDATA[position awareness in segmentation models]]></category>
		<category><![CDATA[REF-VMUNetUp neural network]]></category>
		<category><![CDATA[rotary position embedding]]></category>
		<category><![CDATA[state-space models]]></category>
		<category><![CDATA[Synapse dataset]]></category>
		<category><![CDATA[transformer-based architectures for medical imaging]]></category>
		<category><![CDATA[U-Net]]></category>
		<category><![CDATA[VM-UNet]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212162</guid>

					<description><![CDATA[Researchers in China have developed REF-VMUNetUp, a Mamba-based network with rotary position embeddings and dynamic upsampling that improves multi-organ segmentation accuracy across multiple medical imaging benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Medical image segmentation, the task of teaching computers to trace the exact outlines of organs and lesions inside CT scans, MRIs, and endoscopic images, sits at the heart of modern computer-aided diagnosis. Radiation oncologists use these digital contours to plan treatments, surgeons rely on them to navigate around critical structures, and clinical trials depend on them to measure whether a tumor is shrinking. Yet the deep learning models that perform this work have long been caught in a trade-off. Convolutional neural networks, the workhorses of medical imaging for the past decade, are efficient and good at detecting local textures and edges, but they struggle to capture long-range relationships, the kind of whole-organ context needed to understand where one structure ends and another begins. Transformer-based architectures solve that context problem with self-attention, but their computational cost grows quadratically with the number of image patches, making them expensive and sometimes impractical for the high-resolution volumes used in hospitals.</p>
<p>A research team at Shaanxi University of Technology in Hanzhong, China, led by corresponding author Chuanghua Yang, has now introduced a new architecture designed to escape this dilemma. In a study published in the journal Complex &amp; Intelligent Systems, the researchers present REF-VMUNetUp, a network built on the Mamba architecture, a recently popularized state-space model that processes sequences with linear rather than quadratic complexity. The work builds on VM-UNet, a well-known pure-Mamba segmentation network, but the authors argue that VM-UNet leaves important performance on the table. Their new model targets three specific weaknesses: inadequate modeling of positional information, degradation of translation invariance during upsampling, and inaccurate segmentation of organ boundaries.</p>
<p>To understand why position matters, it helps to consider what Mamba actually does. Originally developed for natural language processing, Mamba treats an image as a sequence of tokens and sweeps through them with a selective state-space mechanism, deciding dynamically which information to remember or forget as it goes. This gives it the long-range modeling power of a Transformer at a fraction of the computational cost, which is precisely why it has generated such excitement in the medical imaging community. The catch is that, in its basic form, the mechanism is largely indifferent to where each token sits in the image. Two patches with similar appearance can be treated nearly identically even if one belongs to the liver and the other to the right kidney, and that ambiguity can blur the very boundaries clinicians care about most.</p>
<p>REF-VMUNetUp addresses this with a two-dimensional rotary position embedding, which the authors call 2D-RoPE. Rotary position embeddings first gained fame in large language models, where they encode the relative distance between tokens by rotating feature vectors in the complex plane; the angle of rotation depends on position, so the attention-like scoring between two tokens naturally reflects how far apart they are. The Shaanxi team adapted this idea to two dimensions, encoding both row and column coordinates so that the network can represent relative spatial relationships across the image plane. The result, according to the study, is a marked improvement in the model&#8217;s ability to represent relative position information, which in turn sharpens its grasp of anatomical layout without adding the heavy cost of explicit positional attention.</p>
<p>The second innovation tackles a subtler problem: upsampling. Segmentation networks follow an encoder-decoder design, compressing the image into abstract feature maps and then expanding them back to full resolution to produce pixel-level labels. The expansion step, or upsampling, is traditionally done with fixed interpolation rules such as bilinear or nearest-neighbor resampling, which treat every location in the image the same way. That uniformity erodes translation invariance and produces edge blurring, because a smooth interpolation cannot know that a sharp organ boundary happens to pass through a particular neighborhood. The researchers&#8217; answer is a module they call D-UP, short for dynamic upsampling, which predicts offsets dynamically for each sampling location and reorganizes channels accordingly. In effect, the network learns where to place each output pixel, warping the resampling grid to follow contours rather than ignoring them, which mitigates the edge blurring that fixed upsampling introduces.</p>
<p>The third component operates on the skip connections, the pathways that carry fine-grained detail from the encoder directly to the decoder. In classic U-Net-style architectures, these connections are what allow the final segmentation to retain precise edges, but they also pass along low-level noise and redundant features. REF-VMUNetUp inserts a module named L-HFP, a lightweight high-frequency processing unit, into the skip pathway. Its job is to adaptively enhance high-frequency features, the rapidly varying components of the signal that correspond to edges, fine textures, and thin structures. By amplifying exactly the information the decoder needs to capture critical details such as organ boundaries and tumor margins, and suppressing what it does not, the module improves the fidelity of the final output without a large computational overhead.</p>
<p>The team evaluated the architecture on the Synapse multi-organ dataset, a widely used benchmark of abdominal CT scans in which algorithms must delineate organs including the aorta, gallbladder, kidneys, liver, pancreas, and spleen. Performance was measured with two standard metrics: the Dice similarity coefficient, or DSC, which quantifies the overlap between the predicted segmentation and the ground truth, and HD95, the 95th-percentile Hausdorff distance, which captures how far the worst boundary errors stray from the true contour. REF-VMUNetUp improved mean DSC by 1.83 percentage points and reduced HD95 by 2.09 millimeters relative to the baseline comparison, gains that are meaningful in a field where fractions of a millimeter can influence treatment planning. The reduction in HD95 is particularly telling, because that metric is dominated by the largest boundary errors, precisely the failures the dynamic upsampling and high-frequency modules were designed to fix.</p>
<p>Impressive results on a single benchmark can reflect overfitting, so the researchers also tested generalization on three additional public datasets spanning different imaging modalities and clinical tasks: ACDC, which features cardiac MRI for measuring heart function; ISIC2018, a dermoscopy dataset for skin lesion segmentation; and CVC-ClinicDB, containing colonoscopy images of polyps. The model&#8217;s consistent performance across these varied domains suggests that the benefits of position-aware Mamba scanning and dynamic upsampling are not tied to one anatomy or scanner type. The authors note that all datasets used in the research are publicly available and that they declare no conflict of interest, and the work was supported by the National Natural Science Foundation of China, the Key Research and Development Program of Shaanxi Province, and a Qinchuangyuan Scientists and Engineers team grant.</p>
<p>The broader significance of the study lies in what it says about the direction of medical AI. For years, the field has been pulled between the efficiency of convolutions and the context awareness of Transformers, and Mamba-style state-space models promise a way to have both. But as this work demonstrates, simply swapping a Transformer for a Mamba block is not enough; the architecture must also respect the geometric and structural priors of imaging, from relative position to sharp boundaries, or it will forfeit much of its potential. By embedding rotary position encoding, learned dynamic resampling, and adaptive high-frequency enhancement into a linear-complexity backbone, REF-VMUNetUp offers a template for how that can be done. If such models continue to mature, the practical payoff could be faster, cheaper, and more precise automated contouring, bringing reliable multi-organ and lesion segmentation closer to routine clinical use in hospitals that cannot afford heavyweight Transformer infrastructure.</p>
<p><strong>Subject of Research:</strong> A Mamba-based deep learning architecture for medical image segmentation</p>
<p><strong>Article Title:</strong> REF-VMUNetUp: a position-aware mamba network with dynamic upsampling for medical image segmentation</p>
<p><strong>Article References:</strong> Yang, C., Yang, H., Yuan, H., Gao, W., Li, W., Pan, Z., Li, Y., &amp; Bai, X. (2026). REF-VMUNetUp: a position-aware mamba network with dynamic upsampling for medical image segmentation. <em>Complex &amp;amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02515-w" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02515-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02515-w" rel="noopener noreferrer">10.1007/s40747-026-02515-w</a></p>
<p><strong>Keywords:</strong> medical image segmentation, Mamba, state space models, VM-UNet, rotary position embedding, dynamic upsampling, deep learning, computer vision, multi-organ segmentation, Synapse dataset, U-Net, computational efficiency</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212162</post-id>	</item>
		<item>
		<title>Wavelet-Enhanced Mamba and Graph Networks Revolutionize Time Series Anomaly Detection</title>
		<link>https://scienmag.com/wavelet-enhanced-mamba-and-graph-networks-revolutionize-time-series-anomaly-detection/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 01:46:57 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[cyberattack detection in IoT networks]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for time series]]></category>
		<category><![CDATA[Graph neural network]]></category>
		<category><![CDATA[graph neural networks for data modeling]]></category>
		<category><![CDATA[healthcare sensor data anomaly detection]]></category>
		<category><![CDATA[Internet of Things]]></category>
		<category><![CDATA[Internet of Things sensor data analysis]]></category>
		<category><![CDATA[IoT sensors]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[Mamba sequence architecture in machine learning]]></category>
		<category><![CDATA[multi-modal data fusion in anomaly detection]]></category>
		<category><![CDATA[multivariate data]]></category>
		<category><![CDATA[real-time anomaly detection in industrial systems]]></category>
		<category><![CDATA[sequence modeling]]></category>
		<category><![CDATA[smart city traffic monitoring]]></category>
		<category><![CDATA[spacecraft telemetry anomaly detection]]></category>
		<category><![CDATA[state-space models]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[time series anomaly detection]]></category>
		<category><![CDATA[wavelet analysis for signal processing]]></category>
		<category><![CDATA[wavelet transform]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=209641</guid>

					<description><![CDATA[Researchers in Beijing have unveiled a wavelet-enhanced model combining Mamba sequence modeling and graph neural networks that outperforms prevailing methods for detecting anomalies in multivariate time series from IoT systems.]]></description>
										<content:encoded><![CDATA[<p>The relentless growth of the Internet of Things has transformed ordinary environments into vast webs of interconnected sensors, each streaming measurements around the clock. From smart factories to hospital wards, from city traffic systems to orbiting spacecraft, the world now generates rivers of time series data faster than human operators can watch them. Buried within those streams are the signals that matter most: the sudden voltage spike that precedes a machine failure, the erratic heartbeat that foreshadows a cardiac event, the subtle sensor drift that hints at a cyberattack. A research team in Beijing has now introduced a machine learning model designed to catch precisely those signals, and their approach combines three of the most powerful ideas in modern artificial intelligence into a single architecture.</p>
<p>The study, led by Hanjie Xu of Zhongguancun Lab and Beihang University together with Ju Ren of Tsinghua University, Jie Liu of Beijing Information Science and Technology University, Zhenya Ma, Wang Huang of Central South University, and Jianwei Niu of Beihang University, was published in the International Journal of Machine Learning and Cybernetics. The team&#8217;s model, which fuses wavelet analysis, graph neural networks, and the newly popular Mamba sequence architecture, outperformed a wide range of established baselines in extensive benchmark testing. The work addresses one of the most persistent challenges in applied machine learning: detecting anomalies in multivariate time series, where dozens or hundreds of correlated sensor channels evolve together and a fault in one can ripple through the others in complex ways.</p>
<p>Each of the three components of the new model plays a distinct role. Wavelet representations, the first ingredient, give the network a way to see a signal at multiple resolutions simultaneously. Traditional Fourier analysis tells an engineer which frequencies exist in a signal but says little about when they occurred, a limitation known as the time frequency trade off. Wavelet transforms solve this by decomposing a signal into localized oscillations of different scales, so the model can capture both a slow drift lasting minutes and a sharp transient lasting milliseconds. By enhancing the raw inputs with these wavelet features, the architecture extracts far richer representations of each moment in the stream than models that read the raw sensor values alone.</p>
<p>The second ingredient, graph neural networks, addresses the fact that sensor data is rarely a set of independent channels. In a water treatment plant, a pressure sensor upstream is physically and causally linked to a flow sensor downstream; in a server cluster, the temperature readings of adjacent racks move in concert. Graph neural networks encode these relationships explicitly, treating each sensor as a node and the learned or measured dependencies as edges. Message passing across the graph allows the model to build a picture of how the whole system should behave as a coherent network, and to flag moments when that coherence breaks down. This topology learning capability, the authors argue, is essential for detecting anomalies that only become visible when several channels are viewed together.</p>
<p>The third component, Mamba, is one of the most talked about architectures in artificial intelligence today. Introduced by Albert Gu and Tri Dao in late 2023, Mamba is a selective state space model that processes sequences in linear time, in contrast to the quadratic cost of transformer attention. For time series workloads, where context windows can stretch across thousands of timesteps and the model must run on resource constrained edge devices, that efficiency matters enormously. Mamba&#8217;s selective mechanism lets the network decide dynamically which parts of the past to remember and which to forget, giving it a powerful ability to follow long range sequential patterns without the memory burden that attends older recurrent designs.</p>
<p>What makes the new model distinctive is the way these three ideas operate together rather than in sequence. The wavelet enhanced representations give the sequence model a frequency aware view of each channel, the graph neural network binds those views into a coherent picture of system wide behavior, and the Mamba backbone explores how the entire enriched representation evolves through time. In effect, the model asks three questions at once: how is each sensor oscillating, how are the sensors related, and how are those patterns changing? An anomaly is registered when any of these three questions produces an answer that deviates from what the model has learned to expect from normal operation.</p>
<p>The evaluation was extensive. The researchers tested their architecture against a broad collection of time series anomaly detection baselines, spanning the main families of methods that have dominated the field over the past decade. Those baselines include long short term memory networks, which have been used for everything from spacecraft telemetry to rail transit monitoring; variational autoencoders and generative adversarial networks that learn to reconstruct normal data and flag what they cannot reconstruct; transformer models such as TranAD and the Anomaly Transformer, which use attention to model dependencies across time; and earlier graph based detectors, including graph attention networks and graph deviation networks, which pioneered the topological view of multivariate data. Across this competitive landscape, the wavelet enhanced Mamba and graph network model outperformed many prevailing approaches.</p>
<p>The significance of the result extends beyond leaderboard positions. Anomaly detection sits at the foundation of industrial safety, cybersecurity, financial monitoring, and healthcare. Earlier studies in the literature the team cites show detectors deployed in industrial IoT systems, smart city traffic networks, financial markets analyzed with principal component analysis and neural networks, and clinical settings where deep learning reviews promise earlier warnings. Each of these domains has its own constraints: latency requirements that forbid heavy computation, noisy training sets, and distribution shifts that arise as machines age or seasons change. A detector that combines strong accuracy with linear time sequence processing and explicit modeling of sensor topology is therefore well matched to deployment realities, particularly at the network edge where IoT devices typically live.</p>
<p>The Mamba component deserves particular attention because it represents a broader shift in the machine learning community. Since the original Mamba paper appeared, researchers have rapidly adapted the architecture to forecasting and detection tasks, producing variants such as bidirectional Mamba for time series prediction, spatial temporal Mamba for multivariate anomaly detection, and Mamba based foundation models for general forecasting. By embedding Mamba within a graph structure and enhancing its inputs with wavelets, the Beijing team positions their model at the intersection of several active research fronts. The graph element draws on a lineage that runs from graph convolutional networks through temporal graph convolutional networks for traffic prediction to recent topological analysis methods, while the wavelet element revives a classical signal processing tool whose value modern deep learning has increasingly rediscovered.</p>
<p>The study also acknowledges the practical ecosystems in which such models must be validated. The references underpinning the work include widely used benchmarks such as the MIT BIH arrhythmia database for cardiac signals, spacecraft telemetry datasets with expert labeled anomalies, and secure water treatment testbeds used in adversarial cyberattack research. Datasets of this kind have shaped the field by providing realistic fault patterns against which methods can be calibrated, and the competitive advantage the new model demonstrated against baselines on this landscape suggests it generalizes across heterogeneous data regimes rather than being tuned to a single domain. The authors declare no conflict of interest, and the work received support reflected in the supervisory and funding roles of the senior researchers on the team.</p>
<p>Looking ahead, the research points toward anomaly detectors that are simultaneously faster, more aware of physical structure, and more sensitive to the fine grained dynamics of the systems they watch. As billions more devices come online, the volume of time series data will only grow, and the cost of missing a fault will continue to rise, whether measured in damaged equipment, lost revenue, compromised security, or human health. A model that can read a system through its frequencies, its connections, and its history at linear cost offers a template for the next generation of monitoring infrastructure. The Beijing team&#8217;s results suggest that the combination of wavelets, graphs, and selective state space models is more than a fashionable assembly of components; it is a coherent answer to the question of how machines can learn what normal looks like in an increasingly connected world, and how quickly they can notice when that normality shatters.</p>
<p><strong>Subject of Research:</strong> A wavelet-enhanced Mamba and graph neural network model for detecting anomalies in multivariate time series data from IoT systems.</p>
<p><strong>Article Title:</strong> Wavelet-enhanced Mamba and graph network-based model for time series anomaly detection</p>
<p><strong>Article References:</strong> Xu, H., Ren, J., Liu, J., Ma, Z., Huang, W., &amp; Niu, J. (2026). Wavelet-enhanced Mamba and graph network-based model for time series anomaly detection. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 475. <a href="https://doi.org/10.1007/s13042-026-03250-x" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03250-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03250-x" rel="noopener noreferrer">10.1007/s13042-026-03250-x</a></p>
<p><strong>Keywords:</strong> time series, anomaly detection, Mamba, graph neural network, wavelet transform, Internet of Things, machine learning, state space models, multivariate data, deep learning, IoT sensors, sequence modeling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">209641</post-id>	</item>
		<item>
		<title>New AI Model Reads Ship Headings From Space With Record Accuracy</title>
		<link>https://scienmag.com/new-ai-model-reads-ship-headings-from-space-with-record-accuracy/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:59:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced AI for maritime safety]]></category>
		<category><![CDATA[AI ship heading detection]]></category>
		<category><![CDATA[AI-powered maritime traffic management]]></category>
		<category><![CDATA[angle encoding]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[deep learning for satellite ship recognition]]></category>
		<category><![CDATA[DOTA-ORShip]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[high-resolution satellite ship detection]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[Mamba-OrthoNet]]></category>
		<category><![CDATA[maritime monitoring]]></category>
		<category><![CDATA[maritime vessel recognition AI]]></category>
		<category><![CDATA[object detection]]></category>
		<category><![CDATA[port and harbor vessel monitoring]]></category>
		<category><![CDATA[remote sensing]]></category>
		<category><![CDATA[remote sensing ship analysis]]></category>
		<category><![CDATA[satellite imagery ship orientation]]></category>
		<category><![CDATA[satellite-based ship direction extraction]]></category>
		<category><![CDATA[ship identification from space]]></category>
		<category><![CDATA[ship image orientation accuracy]]></category>
		<category><![CDATA[ship orientation detection]]></category>
		<category><![CDATA[state-space models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203340</guid>

					<description><![CDATA[A new hybrid AI architecture called Mamba-OrthoNet fuses global and local visual features with circular angle encoding to detect ship orientation in satellite imagery with record accuracy.]]></description>
										<content:encoded><![CDATA[<p>Every day, satellites pass silently over the world&#8217;s busiest ports and shipping lanes, photographing thousands of vessels from hundreds of kilometers above. Buried inside those images is information that port authorities, maritime traffic managers, and search-and-rescue teams urgently need: not just where ships are, but which way each one is pointing. A new artificial intelligence system called Mamba-OrthoNet, described in the journal Complex &amp; Intelligent Systems, now promises to extract that directional information with unprecedented reliability, reaching 96.3 percent mean average precision and 94.8 percent orientation accuracy on a demanding benchmark of high-resolution remote-sensing imagery.</p>
<p>The problem the researchers set out to solve is deceptively simple to state and notoriously hard to crack. Detecting a ship in a satellite image is one thing; determining its orientation, meaning the direction its bow faces relative to the image frame, is quite another. In crowded harbor scenes, vessels are moored side by side, bow to stern, sometimes nearly touching. An algorithm that looks only at the big picture, capturing the global layout of the scene, can separate adjacent ships from one another but may miss the subtle visual cues that distinguish a ship&#8217;s pointed bow from its squared-off stern. Conversely, an algorithm that zooms in on fine local details can read those bow and stern cues but loses track of the surrounding context that keeps neighboring ships from blurring together. The team behind Mamba-OrthoNet argues that neither strategy alone is sufficient, and that the key to accurate orientation detection lies in fusing the two perspectives in a principled, progressive way.</p>
<p>At the heart of the new architecture is an unusual pairing of two very different computational machinery. One branch of the network is a conventional convolutional neural network, or CNN, the workhorse of computer vision that excels at picking out local patterns such as edges, textures, and fine structural details. The other branch is built on a Mamba module, a member of the state-space model family that has recently attracted intense attention as an alternative to transformers. State-space models process sequences while maintaining a compact internal memory of everything they have seen, which allows them to capture long-range, global context with far less computational cost than attention-based transformers. By running these two branches in parallel, Mamba-OrthoNet effectively gives the network both a magnifying glass and a wide-angle lens at the same time.</p>
<p>Simply running two branches, however, does not guarantee that their outputs will combine productively. The features produced by a CNN and those produced by a state-space model live in different representational spaces, and naively adding or concatenating them can introduce redundancy and conflict. To address this, the researchers introduced a component they call Orthogonal Feature Fusion, or OFF. OFF first maps both streams of features into a shared space where they can be meaningfully compared, then decomposes the CNN-derived features to extract their detail-bearing components. By doing so, it reduces redundancy between the two streams, ensuring that the global context from the Mamba branch and the local detail from the CNN branch complement rather than duplicate each other. The result is a fusion that is compatible by construction rather than left for the network to figure out through training alone.</p>
<p>Fusion in Mamba-OrthoNet is also deliberately gradual rather than abrupt. The architecture integrates the two branches through a process the authors describe as moving from representation alignment to residual refinement. A component called progressive residual enhancement, or PRE, aggregates residuals layer by layer, iteratively polishing the fused features as information flows through the network. In practical terms, this means the model does not commit to a single moment of merging global and local information; instead, it keeps refining the combined representation at multiple stages, allowing coarse scene understanding and fine-grained detail to inform each other repeatedly. The team further strengthened the network&#8217;s ability to handle ships of wildly different scales, from small fishing boats to massive container vessels, by adopting a bidirectional fusion scheme known as FPBiFusion and introducing a normalization module, ConvNorm, built with a CSPRep structure to produce stronger multi-scale representations.</p>
<p>Perhaps the most ingenious part of the system, however, lies in how it handles angles. Predicting a ship&#8217;s heading as a single regression value sounds straightforward until one remembers that angles are circular. A ship heading at 359 degrees is pointing almost exactly the same direction as one heading at 1 degree, yet a naive regression model treats those numbers as being nearly 360 degrees apart. This periodicity, and the resulting discontinuity at the 0-degree and 360-degree boundary, is a well-known source of training instability in orientation detection, causing models to produce erratic predictions and converge poorly. Different angle encodings have been proposed over the years, but each involves trade-offs between precision and stability.</p>
<p>The Mamba-OrthoNet team&#8217;s answer is a scheme they call dual-granularity ring encoding, or DGRE. The idea is to represent the full 0-to-360-degree range as a ring, combining two levels of granularity. At the coarse level, the ring is divided into cyclic bins, giving the model a stable, wrap-around-aware sense of roughly where an angle lies. At the fine level, smoothed sub-bins sharpen the estimate, allowing the model to distinguish headings that differ by only a few degrees. Because the coarse structure already respects the circular nature of angles, the fine refinement never has to fight the 0-degree and 360-degree discontinuity, and the model can produce stable predictions across the entire angular range. Ablation studies reported in the paper confirm that DGRE, together with OFF and PRE, contributes measurably to the system&#8217;s overall performance under the benchmark setting.</p>
<p>The benchmark in question, DOTA-ORShip, is a dataset of oriented ship annotations in high-resolution remote-sensing images, a setting that captures exactly the crowded, multi-scale conditions of real ports. On this dataset, Mamba-OrthoNet achieved 96.3 percent mean average precision at the standard 50 percent overlap threshold, along with 94.8 percent orientation accuracy. Those figures place the system among the strongest reported results for ship orientation detection, and the ablation experiments, in which individual components were removed to measure their contribution, support the claim that the architecture&#8217;s gains come from the deliberate interplay of its parts rather than from raw capacity alone.</p>
<p>The practical implications extend well beyond a leaderboard number. Accurate ship orientation is a foundational input for maritime domain awareness. Traffic management systems use vessel heading to predict trajectories and prevent collisions in congested straits and anchorages. Search-and-rescue operations rely on knowing which way drifting or maneuvering vessels are pointed to model drift and plan interception routes. Port operators use orientation data to verify berthing and monitor mooring compliance. A detector that confuses bow with stern, or that destabilizes when a ship&#8217;s heading crosses the 0-degree line, can propagate errors into all of these downstream systems. By combining global context, local detail, and circularly consistent angle encoding in a single network, Mamba-OrthoNet aims to remove several of the most persistent failure modes at once.</p>
<p>The work also reflects a broader shift in the deep learning landscape. State-space models like Mamba have moved rapidly from language modeling experiments into computer vision, offering a way to capture long-range dependencies with linear computational scaling, in contrast to the quadratic cost of transformer attention. The Chengdu-based research team&#8217;s demonstration that a Mamba branch can be fused productively with a CNN branch, through orthogonal feature mapping and progressive residual refinement, offers a template that may generalize beyond ships to other oriented-object detection problems in aerial imagery, from vehicles on highways to aircraft on tarmacs. The article, published open access under a Creative Commons Attribution license, was authored by Jiandan Zhong, Lingfeng Liu, Tao Yu, Zhipeng Yang, Yingxiang Li, Yajuan Xue, and Fei Song, with affiliations spanning Chengdu University of Information Technology, Xizang University, and Xihua University. As satellites multiply and image resolutions improve, the volume of remote-sensing data demanding automated interpretation will only grow, and architectures like this one suggest that the tools for reading that data, down to the direction every vessel points, are becoming both sharper and more dependable.</p>
<p><strong>Subject of Research:</strong> Deep learning-based ship orientation detection in high-resolution remote-sensing imagery</p>
<p><strong>Article Title:</strong> Mamba-OrthoNet: unified global–local feature fusion and angular encoding for accurate ship orientation detection</p>
<p><strong>Article References:</strong> Mamba-OrthoNet: unified global–local feature fusion and angular encoding for accurate ship orientation detection. (n.d.). <a href="https://doi.org/10.1007/s40747-026-02518-7" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02518-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02518-7" rel="noopener noreferrer">10.1007/s40747-026-02518-7</a></p>
<p><strong>Keywords:</strong> Mamba-OrthoNet, ship orientation detection, remote sensing, state-space models, Mamba, convolutional neural networks, feature fusion, angle encoding, object detection, DOTA-ORShip, maritime monitoring, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203340</post-id>	</item>
	</channel>
</rss>
