<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>medical image analysis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/medical-image-analysis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 13:10:18 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>medical image analysis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Read Fuzzy Bone Scans and Write Radiology Reports</title>
		<link>https://scienmag.com/ai-learns-to-read-fuzzy-bone-scans-and-write-radiology-reports/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 13:10:18 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in cancer detection through bone scans]]></category>
		<category><![CDATA[AI-driven diagnostics in nuclear medicine]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[automated radiology report generation]]></category>
		<category><![CDATA[bone metastasis]]></category>
		<category><![CDATA[bone scintigraphy]]></category>
		<category><![CDATA[challenges in nuclear imaging resolution]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[clinical report synthesis from medical images]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[cross-modal alignment]]></category>
		<category><![CDATA[cross-modal alignment in medical AI]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for nuclear medicine]]></category>
		<category><![CDATA[fuzzy bone scan analysis]]></category>
		<category><![CDATA[low-resolution medical imaging]]></category>
		<category><![CDATA[medical image analysis]]></category>
		<category><![CDATA[Medical Imaging]]></category>
		<category><![CDATA[medical report generation]]></category>
		<category><![CDATA[neural networks for radiology]]></category>
		<category><![CDATA[nuclear medicine]]></category>
		<category><![CDATA[radiology]]></category>
		<category><![CDATA[SPECT]]></category>
		<category><![CDATA[SPECT bone scintigraphy interpretation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194723</guid>

					<description><![CDATA[Researchers have developed a deep learning framework that generates clinically faithful diagnostic reports from low-resolution SPECT bone scans by combining domain-adaptive visual encoding, fine-grained cross-modal alignment, and anatomy-guided supervision.]]></description>
										<content:encoded><![CDATA[<p>Medical images are often difficult to read, but few imaging modalities test the limits of human and machine perception quite like functional nuclear medicine. Single-photon emission computed tomography, or SPECT, bone scintigraphy produces whole-body maps of radioactive tracer uptake that reveal where cancer may have spread to the skeleton. The trade-off is resolution: compared with the crisp anatomical detail of computed tomography or magnetic resonance imaging, SPECT bone scans are notoriously blurry, noisy, and low in contrast. A team of researchers in China and Australia has now unveiled a deep learning framework that confronts this challenge head-on, generating clinically faithful diagnostic reports from low-resolution bone scintigrams by tightly aligning what the image shows with what the report must say.</p>
<p>The study, published in the journal Applied Intelligence, was led by Tao Song and Qiang Lin of Northwest Minzu University in Lanzhou, together with colleagues at the Beijing Institute of Technology, Gansu Provincial Cancer Hospital, and Charles Sturt University in Australia. Their central argument is that the two obstacles holding back automated report generation in nuclear medicine—limited visual representation capacity and weak cross-modal alignment—must be solved together rather than in isolation. To that end, the team designed a unified framework that combines three cooperating components: a domain-adaptive visual feature extractor, a fine-grained image-text alignment module, and an anatomy-guided training loss that mimics how physicians actually reason through a scan.</p>
<p>The first component, called the Domain-Adaptive Visual Feature Extractor, or DAVFE, tackles the problem that most image-recognition backbones are pretrained on natural photographs that bear little resemblance to the grainy grayscale world of nuclear medicine. The researchers began with ImageNet initialization, a standard starting point that gives the network a general vocabulary of visual edges and textures. They then refined it with contrastive pretraining on large volumes of unlabeled SPECT data, a self-supervised strategy that teaches the model which scans are similar to one another and which differ, without requiring expert annotations. Finally, an attention-based recalibration mechanism, inspired by convolutional block attention designs, reweights the extracted features so that the subtle, modality-specific signatures of tracer uptake are amplified while background noise is suppressed.</p>
<p>The second component, the Feature Interaction Alignment Module, or FIAM, addresses a subtler failure mode. In many report-generation systems, the entire image is matched against the entire text as a single global summary, which is a poor fit for medicine, where a single sentence about a focal lesion in the ribs must correspond to a specific spot in the scan. FIAM explicitly models the interaction between global narrative semantics—what the overall report is describing—and lesion-level textual cues, the shorter phrases that pinpoint abnormal findings. By letting these two levels of language talk to the corresponding levels of visual representation, the module produces an alignment that is fine-grained and clinically consistent rather than coarse and generic.</p>
<p>The third ingredient, the Anatomy-guided Progressive Granularity Loss, encodes a philosophy rather than an architecture. Experienced nuclear medicine physicians do not read a bone scan in one pass; they first survey the overall distribution of tracer uptake, note the skeleton&#8217;s general anatomy, and then zoom in on suspicious hotspots to characterize them. APGL replicates this coarse-to-fine reasoning during training by applying hierarchical supervision that spans from the global image level down to individual lesions. The network is thus rewarded not only for producing fluent prose but for grounding its descriptions in the correct anatomical locations—a distinction that generic language metrics alone cannot capture.</p>
<p>To test the framework, the team assembled a clinical dataset of 2,091 SPECT bone scintigrams curated at Gansu Provincial Cancer Hospital, with the study approved by the hospital&#8217;s ethics committee under the Declaration of Helsinki. The model was benchmarked against state-of-the-art baselines using conventional natural language generation metrics—BLEU, METEOR, and ROUGE-L—which measure how closely machine-generated text matches reference reports in terms of overlapping words and phrases. The proposed approach consistently outperformed these baselines. More importantly, the researchers introduced a hierarchical Clinical Efficacy metric specifically designed to evaluate lesion localization, assessing whether the generated reports place findings in the right parts of the skeleton. Here, too, the new framework held a clear advantage.</p>
<p>Ablation studies, in which individual components are removed one at a time, confirmed that each of the three modules contributes measurably. Stripping out the domain-adaptive extractor degraded the visual representations; removing the alignment module weakened the correspondence between image content and textual descriptions; and dropping the anatomy-guided loss eroded the model&#8217;s ability to localize findings. Human evaluation and visualization analyses reinforced the quantitative results, showing that the generated reports were judged more clinically faithful, more interpretable, and more reliable than those of competing systems. Attention maps produced by the network revealed that it tended to focus on clinically relevant regions of the scan while composing corresponding sentences, offering a window into the model&#8217;s decision process that radiologists can inspect and, ultimately, trust.</p>
<p>The significance of this work extends beyond a single imaging modality. Automated report generation has advanced rapidly in radiology more broadly, with transformers, memory networks, and contrastive learning approaches producing impressive results on chest X-rays. But nuclear medicine has lagged, precisely because the images are low-resolution and the semantic gap between fuzzy functional images and precise clinical language is wider. By demonstrating that domain-adaptive pretraining, fine-grained alignment, and anatomy-guided supervision can together close that gap, the study offers a template that could transfer to other functional imaging tasks, from cardiac SPECT to positron emission tomography. It also dovetails with the team&#8217;s earlier work on deep learning segmentation of bone metastasis lesions in SPECT scans, forming a pipeline that could one day take a raw scintigram from acquisition to a draft clinical report with minimal human intervention.</p>
<p>The researchers are careful to position the system as decision support rather than a replacement for physicians. Discrepancy and error remain persistent problems in radiology, and the goal is to reduce workload while improving the reliability and consistency of diagnostic and treatment processes. Automatic draft reports could free physicians from repetitive typing, standardize terminology across departments, and serve as a second set of eyes that flags findings for review. The authors describe the approach as a pathway toward trustworthy and intelligent diagnostic support within nuclear medicine—a phrase that captures both the ambition and the caution of the field. The dataset&#8217;s validation subset is available to researchers on request, with full public release planned for the future, and the implementation can be shared through a collaboration agreement. As artificial intelligence steadily earns a place beside the radiologist&#8217;s lightbox, this study suggests that even the blurriest images in medicine may soon speak for themselves.</p>
<p><strong>Subject of Research:</strong> Deep learning-based automatic diagnostic report generation from low-resolution SPECT bone scintigrams using cross-modal visual and textual alignment</p>
<p><strong>Article Title:</strong> Deep learning-based diagnostic report generation for low-resolution functional medical images via cross-modal visual and textual alignment</p>
<p><strong>Article References:</strong> Song, T., Lin, Q., Li, T., Zeng, X., Cao, Y., Man, Z., Liu, C., Cai, Z., &amp; Huang, X. (2026). Deep learning-based diagnostic report generation for low-resolution functional medical images via cross-modal visual and textual alignment. <em>Applied Intelligence, 56</em>(14), Article 421. <a href="https://doi.org/10.1007/s10489-026-07456-y" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07456-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07456-y" rel="noopener noreferrer">10.1007/s10489-026-07456-y</a></p>
<p><strong>Keywords:</strong> deep learning, nuclear medicine, SPECT, bone scintigraphy, medical report generation, cross-modal alignment, artificial intelligence, radiology, bone metastasis, medical imaging, contrastive learning, clinical decision support</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194723</post-id>	</item>
		<item>
		<title>New AI Method Fuses Expert Opinions to Map Lung Vessels With Calibrated Confidence</title>
		<link>https://scienmag.com/new-ai-method-fuses-expert-opinions-to-map-lung-vessels-with-calibrated-confidence/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 04:42:50 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI methods for pulmonary research]]></category>
		<category><![CDATA[AI-driven medical image segmentation]]></category>
		<category><![CDATA[automated digital histology quantification]]></category>
		<category><![CDATA[computational pathology]]></category>
		<category><![CDATA[deep learning ensembles]]></category>
		<category><![CDATA[deep learning for lung tissue analysis]]></category>
		<category><![CDATA[ensemble neural network models in pathology]]></category>
		<category><![CDATA[ensemble segmentation]]></category>
		<category><![CDATA[expert opinion fusion in medical imaging]]></category>
		<category><![CDATA[histological vessel segmentation]]></category>
		<category><![CDATA[lung vessel segmentation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for vascular remodeling]]></category>
		<category><![CDATA[medical image analysis]]></category>
		<category><![CDATA[posterior fusion]]></category>
		<category><![CDATA[pulmonary hypertension]]></category>
		<category><![CDATA[pulmonary hypertension vessel analysis]]></category>
		<category><![CDATA[reliability calibration in medical AI]]></category>
		<category><![CDATA[ReliFuse]]></category>
		<category><![CDATA[scalable pulmonary disease assessment tools]]></category>
		<category><![CDATA[segmentation]]></category>
		<category><![CDATA[uncertainty calibration]]></category>
		<category><![CDATA[vessel mapping in diseased lungs]]></category>
		<category><![CDATA[vessel remodeling]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193766</guid>

					<description><![CDATA[Researchers have developed ReliFuse, a machine learning framework that fuses cached predictions from multiple segmentation experts to segment lung vessels in histology images with calibrated reliability and state-of-the-art overlap.]]></description>
										<content:encoded><![CDATA[<p>Quantifying how blood vessels remodel in diseased lungs has long been one of the most tedious bottlenecks in pulmonary research. Pathologists studying vascular changes associated with pulmonary hypertension must trace and outline vessel after vessel under a microscope, converting stained tissue sections into precise digital measurements. The work is slow, expert-dependent and difficult to scale, yet the numbers it produces underpin how researchers judge disease severity and treatment response. A team at the University of Science, Ho Chi Minh City, working with Vietnam National University, has now introduced a machine learning framework designed to automate this labor without sacrificing the reliability that clinical quantification demands.</p>
<p>The new method, called ReliFuse, is described in the journal Machine Learning and addresses a familiar irony in modern medical image analysis. Deep neural networks have become remarkably good at segmenting anatomical structures from histological images, producing masks that can rival human annotations. However, no single network is perfect, and the errors that individual models make are often complementary: one expert model may miss a faint peripheral vessel that another catches, while the second mislabels a fold of tissue that the first correctly ignores. Rather than treating these disagreements as noise, the Vietnamese team treats them as information, formulating the segmentation task as a problem of posterior fusion, in which multiple frozen expert models pool their predictions into a single, better-calibrated output.</p>
<p>What distinguishes ReliFuse from many ensemble techniques is a striking design constraint. At the fusion stage, the framework never looks at the underlying color image at all. Instead, it operates purely on cached probability maps produced beforehand by a bank of seven independently trained segmentation experts. These probability maps encode, for every pixel, how strongly each expert believes that the pixel belongs to a vessel. The fusion head then constructs so-called ensemble-state features from this stack of opinions, describing where the experts agree, where they diverge, and how their confidence is distributed. Working in logit space rather than raw probabilities, the method pools the evidence from all experts, estimates how trustworthy each local expert opinion is, and applies corrections only where the ambiguity is genuinely high.</p>
<p>Reliability estimation is the conceptual heart of the framework. For each expert model, the researchers compute validation-anchored priors from the model&#8217;s behavior on held-out validation data, giving the fusion head a sense of each expert&#8217;s typical strengths and weaknesses before it ever sees a test case. These priors are combined into a calibrated consensus opinion that serves as the starting point for the final segmentation. Crucially, ReliFuse does not rewrite the whole map. Its residual correction branch is bounded and gated by an ambiguity field, so that confident agreement among experts is preserved unchanged while corrections are concentrated exclusively in the contested regions where experts disagree or where boundary transitions are uncertain. This consensus-preservation principle ensures that the fusion step can refine the output without corrupting regions where the ensemble is already correct.</p>
<p>Training the fusion head is itself a multi-objective undertaking. The researchers combine an overlap loss with a boundary loss, a calibration loss, a consensus-preservation loss and a sparse-correction penalty. The boundary term compares gradient magnitudes between the predicted mask and the annotation, making contour errors visible even when vessels occupy few pixels. The consensus term is deliberately asymmetric, using stop-gradient operators to prevent the model from pulling its prior toward its own output or from simply lowering the ambiguity gate to dodge penalties. The sparse penalty is applied to the gated correction actually added to the logits, discouraging the network from making dense modifications everywhere rather than surgical fixes in ambiguous places. The calibration term supervises both the pooled prior and the final posterior with a Brier-style error, keeping the system&#8217;s confidence honest.</p>
<p>On a publicly available dataset of rat lung histology images with expert-annotated vessel masks, ReliFuse achieved the highest primary overlap among all methods in a matched comparison that gave every learned fusion head the same seven-expert posterior stack. The gains over the strongest competing learned fusion heads, which include ensemble-from-multiple-annotations approaches such as D-LEMA and locally calibrated federated methods such as LC-Fed, are modest in raw Dice and IoU terms. The authors are candid about this. In a paired statistical analysis across the held-out batches, the differences against these strongest references were small and not statistically significant, and the team treats those rows as evidence about effect direction and magnitude rather than proof of broad superiority.</p>
<p>Where ReliFuse genuinely pulls ahead is in the conditions that stress fusion methods hardest. In stress tests isolating batches with high expert disagreement and high vessel content, the improvements were clearest, consistent with the framework&#8217;s design focus on ambiguity and minority evidence. The method also held its own on boundary quality: while P-MoLE recorded the best boundary F1 scores and D-LEMA led on distance metrics such as HD95, ReliFuse remained close on these contour measures, indicating that its overlap gains did not come at the cost of degraded vessel geometry. A calibration and morphology analysis showed no single method dominating every diagnostic, with LC-Fed best on calibration error and D-LEMA best on centerline overlap, but ReliFuse remained competitive across morphology measures while producing the strongest primary Dice and IoU in the matched benchmark.</p>
<p>The practical economics of the approach are part of its appeal. Because the experts run only once and their probability maps are cached, the fusion stage is dramatically cheaper than re-running full segmentation networks. In the researchers&#8217; profiling experiments, recomputing the seven raw-image experts required roughly 9,357 milliseconds per batch of four images and more than 13.4 gigabytes of peak memory, far exceeding the cost of any cached-fusion pass. ReliFuse is slower than naive averaging because it must construct its diagnostic state, estimate calibrated opinions and apply gated corrections, but its parameter count remains modest relative to the base experts, and the framework is designed for settings where multiple frozen models are already available from prior development work.</p>
<p>The study is also notable for its methodological transparency. The authors report a full sensitivity analysis of how the expert bank is constructed, showing that a diversity-aware selection of experts improved every matched fusion rule compared with simply choosing the seven highest-scoring models. Ablation studies confirmed that the ambiguity gate, calibration supervision, boundary emphasis and consensus-preservation terms each contribute to the framework&#8217;s behavior, and the team documents a failure case in which the same gate that recovers a faint vessel can enlarge a false-positive region when the posterior evidence is misleading. By reporting the complete hard-subset stress matrices and labeling exploratory statistics as such, the researchers offer a template for honest evaluation in the crowded field of medical image segmentation.</p>
<p>For the researchers who need these measurements, the implications are concrete. The dataset underlying the work, published by Sinitca and colleagues in Scientific Data in 2024, contains 609 paired microphotographs and binary masks from rat models of pulmonary hypertension, split here into 517 development images and 92 held-out test images. The ReliFuse source code is publicly available on GitHub, and the framework requires no retraining of the underlying expert networks, only the lightweight fusion head. As quantitative histology moves from hand-tracing toward automated pipelines, ReliFuse suggests a pragmatic middle path: rather than chasing ever-larger single models, laboratories can combine the complementary strengths of the models they already have, and let a calibrated arbiter decide, pixel by pixel, whose opinion to trust.</p>
<p>The intellectual lineage of this approach stretches back several decades. Stacked generalization, introduced by Wolpert in 1992, established the idea of training a secondary learner to combine the outputs of base models, and Dietterich&#8217;s foundational work on ensemble methods later explained why combining diverse classifiers so often outperforms any single member. ReliFuse adapts this classical principle to dense prediction, where every pixel rather than every sample must receive a fused verdict, and where the cost of naively rerunning large networks makes caching an attractive design choice.</p>
<p>The framework also draws on a well-developed literature concerning the overconfidence of modern neural networks. Guo and colleagues demonstrated in 2017 that contemporary deep classifiers frequently produce probabilities that are poorly calibrated with respect to true correctness, a concern that is especially acute in medical settings where downstream decisions may hinge on a confidence value. Related work by Lakshminarayanan and colleagues on deep ensembles showed that simply averaging independently trained networks yields surprisingly strong uncertainty estimates, which helps explain why the reliability priors anchored on validation behavior prove so informative in the fusion stage.</p>
<p>Vessel segmentation itself has a long algorithmic history predating deep learning. Multiscale vessel-enhancement filtering, pioneered by Frangi and colleagues in 1998, remains influential, and subsequent surveys catalogued the breadth of methods, datasets and evaluation metrics used to judge tubular structure extraction. The pulmonary histology setting adds distinctive challenges, including stained tissue texture, irregular vessel branching and a pronounced class imbalance favoring background pixels, which is why overlap measures such as Dice and IoU are typically complemented by boundary and centerline diagnostics in this domain.</p>
<p>By situating posterior fusion within these established traditions, the study connects classical ensemble theory, calibration research and vascular image analysis into a single practical pipeline for quantitative histopathology.</p>
<p><strong>Subject of Research:</strong> A reliability-calibrated machine learning framework that fuses multiple segmentation experts&#x27; probability maps to automate histological pulmonary vessel segmentation.</p>
<p><strong>Article Title:</strong> ReliFuse: Reliability-Calibrated Posterior Fusion for Histological Vessel Segmentation</p>
<p><strong>Article References:</strong> Le, T. P., Nguyen, T. N., Tran, V. L. H., Doan, T. T., Nguyen, B. T., &amp; Huynh, S. T. (2026). ReliFuse: Reliability-Calibrated Posterior Fusion for Histological Vessel Segmentation. <em>Machine Learning, 115</em>(9), Article 218. <a href="https://doi.org/10.1007/s10994-026-07154-3" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07154-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07154-3" rel="noopener noreferrer">10.1007/s10994-026-07154-3</a></p>
<p><strong>Keywords:</strong> machine learning, ReliFuse, histological vessel segmentation, posterior fusion, pulmonary hypertension, medical image analysis, deep learning ensembles, uncertainty calibration, ensemble segmentation, computational pathology, vessel remodeling, segmentation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193766</post-id>	</item>
		<item>
		<title>Fuzzy attention-based encoder-decoder improves skin lesion segmentation accuracy</title>
		<link>https://scienmag.com/fuzzy-attention-based-encoder-decoder-improves-skin-lesion-segmentation-accuracy/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Thu, 03 Sep 2026 17:24:17 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[attention mechanisms in deep learning]]></category>
		<category><![CDATA[attention mechanisms in image segmentation]]></category>
		<category><![CDATA[deep learning for dermatology]]></category>
		<category><![CDATA[edge detection in dermatology images]]></category>
		<category><![CDATA[fuzzy attention encoder-decoder]]></category>
		<category><![CDATA[fuzzy attention-based encoder-decoder]]></category>
		<category><![CDATA[fuzzy set theory in medical AI]]></category>
		<category><![CDATA[Innovative Neural Network Architectures]]></category>
		<category><![CDATA[medical image analysis]]></category>
		<category><![CDATA[melanoma detection]]></category>
		<category><![CDATA[multi-national research on skin cancer]]></category>
		<category><![CDATA[neural network for skin cancer]]></category>
		<category><![CDATA[open-access skin cancer dataset]]></category>
		<category><![CDATA[open-access skin lesion datasets]]></category>
		<category><![CDATA[probabilistic neural networks]]></category>
		<category><![CDATA[probabilistic relevance modeling]]></category>
		<category><![CDATA[skin cancer edge detection]]></category>
		<category><![CDATA[skin lesion boundary detection]]></category>
		<category><![CDATA[skin lesion segmentation]]></category>
		<category><![CDATA[uncertainty modeling in medical imaging]]></category>
		<category><![CDATA[uncertainty-based image segmentation]]></category>
		<guid isPermaLink="false">https://scienmag.com/fuzzy-attention-based-encoder-decoder-improves-skin-lesion-segmentation-accuracy/</guid>

					<description><![CDATA[Melanoma, the deadliest form of skin cancer, often presents as a subtle dark patch on the skin whose edges blur almost imperceptibly into healthy tissue. Detecting those edges automatically is one of the deceptively hard problems in medical image analysis, and a new open-access study now offers an unusually elegant answer: instead of forcing a [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Melanoma, the deadliest form of skin cancer, often presents as a subtle dark patch on the skin whose edges blur almost imperceptibly into healthy tissue. Detecting those edges automatically is one of the deceptively hard problems in medical image analysis, and a new open-access study now offers an unusually elegant answer: instead of forcing a neural network to decide pixel by pixel whether something is &#8220;lesion&#8221; or &#8220;not lesion,&#8221; the researchers behind a new architecture called FAED let the network think in shades of uncertainty — the way a dermatologist actually does.</p>
<p>The work, published in the journal Complex &amp; Intelligent Systems, comes from an international team spanning SRM Institute of Science and Technology in India, the National Institute of Technology Rourkela, China University of Mining and Technology, Innopolis University in Russia, and St. Petersburg Electrotechnical University &#8220;LETI.&#8221; The team — M. R. Indresh, Soumyajit Gayen, Dmitrii Minenkov, Dmitrii Kaplun and Ram Sarkar — describes FAED, a Fuzzy Attention-aided Encoder-Decoder architecture, which swaps out the rigid binary logic of standard attention mechanisms for a soft, probabilistic notion of relevance inspired by fuzzy set theory. The results are striking not only for their accuracy but for the architecture&#8217;s remarkable frugality: with just 2.4 million parameters and roughly 4 GFLOPs of computation, FAED posts Dice scores that put it at the top tier of contemporary segmentation models while running at inference speeds measured in milliseconds.</p>
<p>The clinical stakes of this problem are easy to underestimate. Early-stage melanoma is highly curable, but the first line of defense is visual inspection of pigmented lesions, typically through dermoscopy — the imaging of skin through a magnifying device that reveals subsurface structures. Automated segmentation of dermoscopy images, the task of drawing an accurate boundary around a lesion, underpins every downstream measurement clinicians and computer-aided diagnosis systems rely on, including the asymmetry, border irregularity and color variation criteria used in melanoma risk scoring. Yet the task is plagued by low contrast between lesion and healthy skin, hair occlusions, specular reflections, and most fundamentally, ambiguous boundaries where the lesion fades gradually into its surroundings.</p>
<p>For years, the dominant tool for this job has been U-Net, a convolutional encoder-decoder architecture in which a contracting path extracts increasingly abstract features and an expanding path reconstructs a pixel-level prediction. The critical link between the two halves is a set of skip connections that pass fine-grained spatial detail from early encoder layers directly to the decoder. Most modern variants bolt attention modules onto these skip connections: the network learns to &#8220;gate&#8221; which features to pass through. But those gates are typically binary — a feature channel or spatial position is either kept or discarded. The FAED authors argue that this all-or-nothing logic is fundamentally mismatched to the nature of skin lesions, where the transition between sick and healthy tissue is gradual, not sharp. A binary gate discards exactly the soft, intermediate evidence that defines an ambiguous boundary.</p>
<p>FAED&#8217;s central innovation is its Boundary-conditioned Soft Fuzzy Attention (BSFA) module, which replaces standard skip connections altogether. Rather than multiplying features by a learned 0-or-1 mask, BSFA evaluates feature relevance using learnable Gaussian membership functions — mathematical constructs from fuzzy logic that assign each feature a continuous degree of membership, modeled as a probability-like value between zero and one. In practice, this means the network can express that a feature is &#8220;somewhat relevant&#8221; or &#8220;mostly relevant,&#8221; preserving graded boundary information that binary attention would crush. The Gaussian membership functions are themselves learnable parameters, so the network discovers its own notions of partial relevance during training rather than having them imposed by a fixed rule.</p>
<p>The architecture adds two further refinements that the authors show are individually and jointly important. The first is an Adaptive Fuzzy Mixture-based aggregation scheme. Features extracted at different depths of the network vary enormously in scale and semantics — shallow layers carry edge textures, deep layers carry abstract shape information — and fusing them well is a persistent headache in segmentation design. The fuzzy mixture approach treats each feature source as contributing to a soft ensemble, weighting its contribution according to a learned similarity-based membership rather than simple concatenation. The second refinement is an explicit Boundary Cue, a signal fed into the attention mechanism that modulates its focus along lesion perimeters. Where the fuzzy membership decides &#8220;how relevant&#8221; a feature is, the boundary cue tells the attention &#8220;where to look,&#8221; sharpening the model&#8217;s sensitivity precisely at the lesion border where errors are most costly.</p>
<p>The authors validated FAED on the four most widely used benchmarks in the field: the ISIC2016, ISIC2017 and ISIC2018 dermoscopy datasets from the International Skin Imaging Collaboration, and the smaller PH² dataset of melanocytic lesion images. The segmentation quality was measured with the Dice score, a standard metric that quantifies the overlap between the predicted lesion mask and the ground truth, where a score of 1.0 means perfect agreement. FAED achieved a Dice score of 0.9140 on ISIC2016, 0.9135 on PH², 0.8781 on ISIC2018, and 0.8615 on ISIC2017 — competitive-to-leading figures given the architecture&#8217;s size. Notably, the ISIC2016 and PH² results hover around the 0.91 mark, a level of overlap that corresponds to clinically meaningful boundary fidelity.</p>
<p>Just as important as the headline numbers is the efficiency profile, which the team documented with careful empirical measurements on an NVIDIA Tesla T4 GPU. FAED performs inference in 10.05 milliseconds per image at batch size 1, and 6.76 milliseconds per image when batched at 8 — throughput fast enough for real-time clinical workflows. Peak GPU memory during inference is similarly modest: 505 MB at batch size 1 and 948 MB at batch size 8. For context, many state-of-the-art segmentation models rely on heavyweight transformer backbones or large convolutional stacks with parameter counts in the tens of millions, demanding memory and compute budgets that make deployment on hospital hardware, edge devices or low-resource settings difficult. FAED&#8217;s 2.4 million parameters and roughly 4 GFLOPs place it in a different class entirely, suggesting that careful architectural design — rather than brute-force scale — can carry segmentation performance a long way.</p>
<p>To verify that each component of the design actually earns its place, the researchers conducted ablation studies, the standard experimental practice of removing parts of a system one at a time and measuring the drop in performance. These studies confirmed that both the prototype-based fuzzy aggregation and the boundary-conditioned modulation of attention contribute measurably to the observed improvements. In other words, the gains are not an artifact of added capacity or incidental tuning: the soft membership modeling and the explicit boundary guidance are doing real, distinguishable work. That finding matters for the broader field, because it offers evidence that how features are fused — treating fusion as a soft similarity-based membership problem — can be as consequential as how features are extracted.</p>
<p>The philosophical shift at the heart of FAED is worth dwelling on. Classical computer vision and early deep learning systems were built on crisp logic: a pixel belongs to a class, a feature passes a gate, a decision is yes or no. Fuzzy logic, introduced decades ago as a formal way of reasoning with degrees of truth, has long been touted as a natural fit for medical imaging, where human experts themselves reason in gradients — &#8220;this border looks slightly irregular,&#8221; &#8220;this region is probably part of the lesion.&#8221; What has changed recently is that learnable fuzzy components, such as Gaussian membership functions optimized end-to-end by gradient descent, can now be embedded inside deep networks so that the fuzzy rules themselves are discovered from data. FAED is a concrete demonstration that this marriage of classical soft-computing theory and modern deep learning can outperform hard-gated alternatives on a real clinical task, without any increase in architectural complexity.</p>
<p>The implications for melanoma screening are potentially significant, particularly for parts of the world where dermatologists are scarce and mobile screening programs depend on lightweight, fast and reliable algorithms. A model that runs in a few milliseconds on an entry-level GPU, fits comfortably in under a gigabyte of memory, and still achieves over 91 percent overlap with expert-annotated boundaries on benchmark datasets is precisely the kind of tool that can be embedded into telemedicine pipelines or portable dermoscope accessories. The authors caution, as all careful researchers do, that benchmark performance is a step toward clinical deployment, not the deployment itself — prospective validation on diverse skin tones, imaging devices and real-world lesion appearances remains an essential next stage for any segmentation technology destined for the clinic.</p>
<p>The article was published open access under a Creative Commons license, making the full technical description freely available to researchers and clinicians worldwide. The study was supported by the Ministry of Economic Development of the Russian Federation. As peer-reviewed, citable research made available early for faster dissemination, it joins a growing body of work arguing that the future of medical AI lies not only in ever-larger models, but in smarter ones — systems that, like the physicians they assist, know how to say &#8220;maybe.&#8221; FAED&#8217;s fuzzy attention may be an early but compelling example of that principle turned into working code.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Deep learning–based skin lesion segmentation in dermoscopy images using fuzzy attention mechanisms</p>
<p><strong>Article Title:</strong> FAED: fuzzy attention-aided encoder-decoder architecture for skin lesion segmentation</p>
<p><strong>Article References:</strong> Indresh, M. R., Gayen, S., Minenkov, D., Kaplun, D., &amp; Sarkar, R. (2026). FAED: fuzzy attention-aided encoder-decoder architecture for skin lesion segmentation. <em>Complex &amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02482-2" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02482-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02482-2" target="_blank" rel="noopener noreferrer">10.1007/s40747-026-02482-2</a></p>
<p><strong>Keywords:</strong> Skin lesion segmentation, Dermoscopy, Fuzzy attention, Boundary-aware segmentation, U-Net model, Feature fusion, Melanoma diagnosis, Encoder-decoder architecture, Gaussian membership functions, ISIC datasets</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">186487</post-id>	</item>
		<item>
		<title>ALPaCA Adapts Llama for Pathology Context Analysis and Slide-Level Question Answering</title>
		<link>https://scienmag.com/alpaca-adapts-llama-for-pathology-context-analysis-and-slide-level-question-answering/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 21 Aug 2026 00:21:34 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI adaptation for pathology]]></category>
		<category><![CDATA[AI in pathology]]></category>
		<category><![CDATA[cellular morphology detection]]></category>
		<category><![CDATA[digital pathology]]></category>
		<category><![CDATA[large language models for medical diagnosis]]></category>
		<category><![CDATA[medical image analysis]]></category>
		<category><![CDATA[pathology context understanding]]></category>
		<category><![CDATA[slide-level question answering]]></category>
		<category><![CDATA[tissue organization recognition]]></category>
		<category><![CDATA[tumor architecture analysis]]></category>
		<category><![CDATA[visual evidence integration in medical AI]]></category>
		<category><![CDATA[whole-slide image analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/alpaca-adapts-llama-for-pathology-context-analysis-and-slide-level-question-answering/</guid>

					<description><![CDATA[Pathology has entered an era in which a single medical image can contain more information than any human can comfortably inspect at once. Whole-slide images, or WSIs, convert glass microscope slides into enormous digital files that may contain billions of pixels, revealing tumor architecture, cellular morphology, tissue organization and subtle diagnostic clues across multiple scales. [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Pathology has entered an era in which a single medical image can contain more information than any human can comfortably inspect at once. Whole-slide images, or WSIs, convert glass microscope slides into enormous digital files that may contain billions of pixels, revealing tumor architecture, cellular morphology, tissue organization and subtle diagnostic clues across multiple scales. Yet asking an artificial intelligence system a simple question about an entire slide remains extraordinarily difficult. A new study published in <em>Nature Communications</em> introduces ALPaCA, a system designed to adapt the Llama family of large language models for pathology context analysis and slide-level question answering.</p>
<p>The work by Gao, He, Su and colleagues addresses a central challenge in medical artificial intelligence: connecting visual evidence distributed across a massive pathology slide with natural-language reasoning. Conventional computer-vision models are often trained to classify small image patches or predict a diagnosis from preselected regions. That approach can be effective when the task is narrowly defined, but it struggles when a pathologist asks a broader question such as which tissue compartments are present, where abnormal structures are located, or how multiple regions contribute to an overall interpretation. ALPaCA is designed to move beyond isolated image recognition by building a structured connection between slide content and language-based analysis.</p>
<p>The difficulty begins with scale. A high-resolution WSI cannot usually be inserted directly into a language model because it is far larger than the model’s input capacity. The slide must first be divided into smaller visual regions, commonly called patches or tiles. These regions can then be processed by an image encoder that converts their visual features into numerical representations. The resulting information must be compressed, organized and presented to a language model in a way that preserves the relationships between local findings and the global structure of the specimen. If that process loses spatial context, the model may identify a feature correctly while misunderstanding its significance within the slide.</p>
<p>ALPaCA’s core idea is to adapt Llama so that it can interpret pathology-specific visual context rather than treating a slide as a collection of unrelated image fragments. In practical terms, this involves connecting visual representations extracted from pathology images with the language model’s token-based reasoning system. The model can then receive visual evidence and generate answers in natural language, potentially explaining what it observes and linking local morphology to a slide-level conclusion. This kind of design represents a shift from simple image classification toward multimodal question answering, where the system must identify relevant evidence, integrate it and formulate a response.</p>
<p>The researchers’ approach is especially important because pathology questions are rarely limited to one visual object. A pathologist may need to compare several areas, determine whether a pattern is widespread or focal, distinguish normal from abnormal tissue, or interpret the relationship between cellular details and larger anatomical structures. These tasks demand what researchers often call context-aware reasoning. A gland, nucleus or inflammatory region can have different meanings depending on where it appears, what surrounds it and how frequently it occurs. By adapting a general-purpose language model to pathology context, ALPaCA aims to make those relationships accessible through interactive questions rather than fixed diagnostic labels alone.</p>
<p>Slide-level question answering could eventually provide a more flexible interface for digital pathology. Instead of asking a model only to produce a predetermined category, users could pose targeted questions about the content of a specimen. Such systems might help retrieve relevant regions, summarize morphological patterns, compare findings across tissue compartments or support the review of complex cases. In a research or clinical workflow, a language-based interface could also make computational analysis easier for users who are not specialists in machine learning. However, the value of such a system depends on whether its answers are grounded in the actual slide rather than generated from statistical associations or plausible-sounding language.</p>
<p>That issue places interpretability and reliability at the center of the ALPaCA study. Large language models are powerful generators of text, but they can also produce confident answers that are incomplete, ambiguous or incorrect. In pathology, an unsupported statement is more than a technical error: it could influence a diagnostic decision. A useful slide-question-answering system therefore needs to connect its responses to visual evidence and ideally indicate which regions support a conclusion. Context analysis can help with this requirement by encouraging the model to reason over multiple locations, but it does not eliminate the need for expert oversight, careful validation and transparent evaluation.</p>
<p>The study also highlights a broader trend in medical AI. Rather than building a separate model for every narrowly defined task, researchers are increasingly adapting foundation models that already possess broad capabilities in language, representation learning or visual interpretation. Llama provides a language-based foundation that can be specialized with pathology data and visual inputs. The advantage of this strategy is flexibility: one adapted model may support many forms of interaction, from descriptive questions to evidence-based comparisons. The challenge is that medical specialization requires high-quality, well-annotated data and strict controls against hallucination, bias and the misuse of incomplete clinical information.</p>
<p>Pathology is particularly demanding because tissue appearance varies with organ type, staining protocol, scanner characteristics, preparation quality and disease stage. A model trained on one collection of slides may perform differently when confronted with images from another laboratory or population. It must also distinguish meaningful biological variation from technical artifacts. These concerns make external validation essential. A system that answers questions accurately on a research benchmark may still require substantial testing before it can be integrated into routine diagnostic practice. ALPaCA’s significance therefore lies not only in its immediate performance, but also in the direction it represents: pathology AI that is conversational, context-sensitive and designed to work with the full complexity of digital slides.</p>
<p>The arrival of ALPaCA signals a growing ambition for computational pathology: to create systems that do not merely recognize patterns, but participate in a structured dialogue about what those patterns mean. If further studies confirm that the approach can produce accurate, visually grounded and reproducible answers across diverse specimens, slide-level question answering could become a powerful tool for research, education and clinical decision support. It will not replace pathologists, whose expertise includes clinical history, uncertainty management and responsibility for patient care. Instead, its most valuable role may be to help experts navigate enormous quantities of visual information, focus attention on relevant regions and turn digital slides into evidence that can be examined through natural language.</p>
<p><strong>Subject of Research</strong>: ALPaCA, a pathology-focused multimodal artificial intelligence system that adapts Llama for context analysis and slide-level question answering.</p>
<p><strong>Article Title</strong>: ALPaCA: Adapting Llama for Pathology Context Analysis to enable slide-level question answering.</p>
<p><strong>Article References</strong>: Gao, Z., He, K., Su, W. <i>et al.</i> “ALPaCA: Adapting Llama for Pathology Context Analysis to enable slide-level question answering.” <i>Nature Communications</i> (2026). <a href="https://doi.org/10.1038/s41467-026-76372-z">https://doi.org/10.1038/s41467-026-76372-z</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: 10.1038/s41467-026-76372-z</p>
<p><strong>Keywords</strong>: computational pathology, digital pathology, whole-slide images, multimodal AI, large language models, Llama, pathology context analysis, slide-level question answering, medical imaging, artificial intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">180703</post-id>	</item>
	</channel>
</rss>
