<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>clinical decision support &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/clinical-decision-support/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 22:39:53 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>clinical decision support &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Privacy-Safe AI Learns When to Take ICU Patients Off the Ventilator</title>
		<link>https://scienmag.com/privacy-safe-ai-learns-when-to-take-icu-patients-off-the-ventilator/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 22:39:53 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-driven ICU ventilator weaning decision support]]></category>
		<category><![CDATA[airway management]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[conservative Q-learning]]></category>
		<category><![CDATA[cyber-physical systems for ICU patient management]]></category>
		<category><![CDATA[data fusion techniques in ICU patient]]></category>
		<category><![CDATA[diaphragm ultrasound]]></category>
		<category><![CDATA[differential privacy]]></category>
		<category><![CDATA[ethical considerations in AI for sensitive health data]]></category>
		<category><![CDATA[federated learning]]></category>
		<category><![CDATA[Internet of Things]]></category>
		<category><![CDATA[IoT sensor-based patient monitoring in ICU]]></category>
		<category><![CDATA[machine learning models for diaphragm weakness prediction]]></category>
		<category><![CDATA[mechanical ventilation weaning]]></category>
		<category><![CDATA[MIMIC-IV]]></category>
		<category><![CDATA[multimodal data fusion]]></category>
		<category><![CDATA[offline reinforcement learning]]></category>
		<category><![CDATA[otolaryngology]]></category>
		<category><![CDATA[privacy engineering in critical care AI applications]]></category>
		<category><![CDATA[privacy-preserving machine learning in critical care]]></category>
		<category><![CDATA[reducing ventilator-associated pneumonia through intelligent monitoring]]></category>
		<category><![CDATA[risk assessment for early extubation in head and neck surgery]]></category>
		<category><![CDATA[tailored ventilator weaning protocols for airway pathology]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224058</guid>

					<description><![CDATA[Researchers in Shanghai have built a privacy-preserving federated AI system that fuses bedside ventilator, ultrasound, and blood gas data with offline reinforcement learning to help clinicians decide when patients with complex airway conditions can be safely weaned from mechanical ventilation.]]></description>
										<content:encoded><![CDATA[<p>One of the most consequential decisions in intensive care is also one of the most difficult to make: when to liberate a patient from mechanical ventilation. Remove the breathing tube too early, and the patient may fail extubation, requiring emergency reintubation with all the associated risks of airway trauma, aspiration, and prolonged hospitalization. Wait too long, and the patient faces ventilator-associated pneumonia, diaphragm weakness, sedation accumulation, and rising costs. For a specific and vulnerable group of patients—those with upper airway pathology following head and neck surgery, those with laryngeal dysfunction, and those with obstructive airway conditions—the standard weaning criteria developed for general intensive care unit populations often fall short, because the underlying physiology of airway compromise does not map cleanly onto the rules derived from broader ICU cohorts. A new study published in Complex &amp; Intelligent Systems by Fangling Peng, Fei Pei, and Hong Zhou of the Department of Otolaryngology at Shidong Hospital in Shanghai proposes a technological answer that is as much about privacy engineering as it is about machine learning.</p>
<p>The researchers built a cyber-physical framework anchored in Internet of Things sensing at the bedside. The system continuously acquires and fuses four fundamentally different kinds of clinical data: ventilator waveform parameters that describe the mechanics of each delivered breath, diaphragm ultrasound imaging features that reveal whether the patient&#8217;s principal breathing muscle is strong enough to sustain independent ventilation, arterial blood gas indices that capture the adequacy of oxygenation and carbon dioxide clearance, and static patient baseline profiles such as demographics and comorbidity information. Each of these streams has its own sampling rate, dimensionality, and noise characteristics. The framework therefore employs a gated recurrent architecture—a class of neural network designed to carry relevant information across time steps while discarding noise—augmented by a cross-modal attention mechanism. In practical terms, attention allows the model, at every moment, to decide which data stream deserves the most weight: when the ventilator waveform shows rapid shallow breathing, the model may attend more strongly to respiratory mechanics; when the waveform looks stable but the diaphragm ultrasound shows thinning muscle, the imaging signal can dominate the assessment.</p>
<p>On top of this multimodal perception layer sits the decision-making core, and here the study makes a choice that reflects a hard lesson learned across the field of medical artificial intelligence. Rather than training a reinforcement learning agent by trial and error in a live clinical environment—an approach that would be ethically untenable, since an exploring algorithm might deliberately test suboptimal ventilator settings on real patients—the team used offline reinforcement learning. The specific algorithm, Conservative Q-Learning, learns a dynamic ventilation parameter adjustment policy entirely from retrospective data. Its defining trick is conservatism: the agent learns the value of actions that appear in the historical record, but it systematically penalizes the estimated value of out-of-distribution actions that the dataset never demonstrates. This prevents the classic failure mode of offline reinforcement learning, in which an agent extrapolates confidently about actions no clinician ever took and recommends them with unwarranted certainty. The result is a policy that can suggest how ventilation parameters might be adjusted over time while remaining anchored to what experienced clinicians actually did, and to the outcomes that followed.</p>
<p>Complementing the policy agent, the framework includes a discrete-time survival model that produces calibrated, uncertainty-aware estimates of each individual patient&#8217;s probability of successful weaning. This is a critical piece of clinical trustworthiness. A single point prediction—say, a 78 percent chance of weaning success—tells a clinician little about how confident the model really is. By working in discrete time steps and reporting uncertainty, the survival model tells the care team not only what the system expects but how much the system itself doubts that expectation, allowing human judgment to calibrate accordingly. The pairing of a reinforcement learning policy with an uncertainty-aware probabilistic forecast reflects a broader trend in clinical machine learning: decision support systems are increasingly designed to quantify their own limitations rather than present deceptively crisp answers.</p>
<p>The privacy architecture is where the study pushes furthest beyond conventional clinical prediction models. Hospitals are naturally reluctant to pool intensive care data, which is among the most sensitive health information that exists, and legal frameworks in many jurisdictions restrict or complicate cross-institution data sharing. The researchers therefore adopted a federated learning protocol based on FedProx, an algorithm in which each participating hospital trains the model locally on its own patients and shares only model parameter updates—never raw data—with a coordinating server. FedProx adds a proximal term that keeps each hospital&#8217;s local model from drifting too far from the global consensus, a safeguard that matters in clinical settings where different units see different patient mixes and hardware varies. On top of federation, the framework applies Gaussian differential privacy with a privacy budget of epsilon equal to 1.0, a mathematically rigorous guarantee that the contribution of any single patient&#8217;s record to the learned model is bounded and quantifiable. A privacy budget of 1.0 is a meaningfully strict setting; it trades some statistical efficiency for a strong formal assurance that no individual&#8217;s ventilator traces, imaging studies, or blood gas values can be reconstructed from the shared model updates.</p>
<p>Recognizing that no two hospitals are identical, the team also incorporated MAML-style local personalization, drawing on the Model-Agnostic Meta-Learning paradigm. MAML trains a model not to perform a single task well but to be rapidly adaptable, so that each participating otolaryngology or head and neck surgery unit can fine-tune the shared model to its own patient population with minimal local data. This addresses a persistent tension in federated medical AI: a single global model may average away the very idiosyncrasies that matter for a specialized unit, while fully independent local models sacrifice the statistical power of pooled learning. Meta-learned personalization offers a middle path—the global model learns how to learn, and each site adapts quickly.</p>
<p>The evaluation was substantial. The team trained and tested the framework on 36,181 weaning episodes drawn from two widely used critical care databases, MIMIC-IV and eICU-CRD, which together capture thousands of intensive care stays from multiple hospitals. Performance was measured with the area under the receiver operating characteristic curve, or AUROC, a standard metric of discriminative ability in which 0.5 represents chance and 1.0 represents perfect separation of outcomes. On internal validation the system achieved an AUROC of 0.893, and on external validation—testing on data the model had not been tuned against—it reached 0.871. The modest drop between the two figures is itself informative, suggesting the model generalizes rather than memorizes. Perhaps most striking for privacy-minded readers, the federated variant of the system narrowed the performance gap to fully centralized training to within 0.001 AUROC. In other words, the hospitals could, in principle, learn together across institutional boundaries with almost no measurable cost in predictive accuracy, while keeping every raw record inside its originating institution and carrying a formal differential privacy guarantee.</p>
<p>The clinical target of the work deserves emphasis. Airway liberation—deciding when a patient can safely breathe without mechanical support—is particularly fraught in otolaryngological practice. After head and neck surgery, swelling, surgical anatomy, and impaired laryngeal function can make the airway fragile in ways that routine weaning parameters, which were largely validated on cardiac and general medical ICU populations, do not capture. Diaphragm ultrasound adds a direct window into respiratory muscle readiness, ventilator waveforms expose the pattern of patient-ventilator interaction, and blood gases confirm physiological stability, but integrating these heterogeneous signals in real time has traditionally been a matter of clinical intuition. The framework described in the study formalizes that integration, turning four bedside data streams into a continuously updated, uncertainty-aware estimate of weaning readiness, alongside a policy that indicates how ventilation might be adjusted as the patient progresses.</p>
<p>It is worth noting what the study does and does not claim. The system was evaluated on retrospective databases, not deployed prospectively at the bedside, and the authors report that the research received no dedicated funding and that the authors declare no competing financial interests. The work was published open access on 4 September 2026 in Complex &amp; Intelligent Systems, a peer-reviewed journal, and carries the DOI 10.1007/s40747-026-02488-w. Retrospective validation is the necessary first step in the long path toward clinical deployment, which would require prospective trials, regulatory review, and integration with hospital information systems. But the architectural choices—offline learning that never experiments on patients, federated training that never moves raw data, differential privacy that bounds individual leakage, and uncertainty quantification that flags the model&#8217;s own doubt—are precisely the design patterns that regulators and clinicians have been demanding from medical AI.</p>
<p>The broader significance of the study lies in its demonstration that privacy and performance need not be opposing forces in clinical intelligence. For years, the assumption in health data science was that the price of protecting patient privacy was a meaningful loss of model accuracy, and that institutions would have to choose between collaborative learning and competitive performance. By combining IoT-scale multimodal sensing, conservative offline reinforcement learning, federated optimization with personalization, and formal privacy guarantees—and by showing that the federated system trails centralized training by less than a thousandth of an AUROC point—the Shanghai team has offered a concrete existence proof that this trade-off can be nearly eliminated. If subsequent prospective studies confirm the retrospective results, the framework could point the way toward a generation of hospital AI that learns from every patient it serves without ever requiring any single patient&#8217;s data to leave the ward.</p>
<p><strong>Subject of Research:</strong> A privacy-preserving federated IoT and offline reinforcement learning framework for predicting mechanical ventilation weaning in patients with upper airway pathology</p>
<p><strong>Article Title:</strong> Privacy-preserving federated IoT intelligence for multimodal airway liberation decision support via offline reinforcement learning</p>
<p><strong>Article References:</strong> Privacy-preserving federated IoT intelligence for multimodal airway liberation decision support via offline reinforcement learning. (n.d.). <a href="https://doi.org/10.1007/s40747-026-02488-w" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02488-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02488-w" rel="noopener noreferrer">10.1007/s40747-026-02488-w</a></p>
<p><strong>Keywords:</strong> mechanical ventilation weaning, federated learning, offline reinforcement learning, Internet of Things, differential privacy, clinical decision support, diaphragm ultrasound, MIMIC-IV, airway management, conservative Q-learning, multimodal data fusion, otolaryngology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224058</post-id>	</item>
		<item>
		<title>AI Learns to Read Hearing Tests Like an Audiologist, But Not Yet Like a Doctor</title>
		<link>https://scienmag.com/ai-learns-to-read-hearing-tests-like-an-audiologist-but-not-yet-like-a-doctor/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 08:10:47 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI accuracy in interpreting hearing tests]]></category>
		<category><![CDATA[AI audiology interpretation]]></category>
		<category><![CDATA[AI in hearing health assessment]]></category>
		<category><![CDATA[AI-driven audiology workflows]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[artificial intelligence in otolaryngology]]></category>
		<category><![CDATA[audiogram and tympanogram interpretation]]></category>
		<category><![CDATA[audiology]]></category>
		<category><![CDATA[automated hearing test analysis]]></category>
		<category><![CDATA[clinical data extraction from hearing tests]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[digital hearing test interpretation]]></category>
		<category><![CDATA[hearing loss]]></category>
		<category><![CDATA[Journal of Medical Systems]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[machine learning for audiology reports]]></category>
		<category><![CDATA[medical imaging AI]]></category>
		<category><![CDATA[modular AI systems in healthcare]]></category>
		<category><![CDATA[multimodal AI]]></category>
		<category><![CDATA[multimodal AI in medical diagnostics]]></category>
		<category><![CDATA[patient communication]]></category>
		<category><![CDATA[prompt engineering]]></category>
		<category><![CDATA[pure-tone audiometry]]></category>
		<category><![CDATA[tympanometry]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221274</guid>

					<description><![CDATA[A modular study of multimodal AI shows audiologist-guided prompts dramatically improve hearing-test interpretation, while no-response handling and cross-model reliability remain critical weaknesses.]]></description>
										<content:encoded><![CDATA[<p>Hearing tests produce some of medicine&#8217;s most deceptively simple images. An audiogram is a grid of symbols marking the faintest sounds a patient can detect at each frequency, in each ear, with and without masking noise. A tympanogram traces how the eardrum moves under changing pressure. Interpreting these charts requires more than reading numbers: it demands knowledge of specialty conventions, masking rules, and the subtle distinction between a threshold that was not measured and a sound so loud the patient still could not hear it. A new study published in the Journal of Medical Systems has now tested, with unusual methodological rigor, whether multimodal artificial intelligence can perform this interpretation, and where it still fails.</p>
<p>The research team, led by investigators at the University of Hong Kong and Ningbo Hospital of Integrated Traditional Chinese and Western Medicine in China, took a deliberately modular approach. Rather than asking an AI to produce an end-to-end diagnosis from an image, they split the problem into two separate modules. The first tested whether a large multimodal model could transcribe and interpret pure-tone audiometry and tympanometry images. The second tested whether AI workflows could calculate twenty-five prespecified clinical fields and draft professional and patient-facing reports in Chinese, given already-verified structured data. This separation matters because a single end-to-end score can hide whether errors come from reading the image, applying specialty rules, or communicating the result.</p>
<p>The study drew on 158 outpatient audiology encounters collected between December 2024 and March 2025, of which 155 records representing 151 unique patients and 302 ears were eligible. Reference standards were built painstakingly: two trained transcribers independently entered every threshold while masked to each other, and two audiologists with 13 and 16 years of clinical experience classified tympanogram curves, agreeing on 296 of 300 dually classified ears, a Cohen&#8217;s kappa of 0.979. Air-bone gaps, hearing-loss degrees, and loss types were derived under prespecified rules, including a local convention that a meaningful air-bone gap required at least two comparable frequencies with gaps of 15 dB or more.</p>
<p>The heart of the first module was the audiologist skill: a carefully engineered prompt, not a fine-tuned model, that encoded audiological practice. Early development errors were revealing. The model initially produced thresholds not on the standard 5-dB grid, misapplied degree boundaries, overcalled conductive components, treated insufficient bone-conduction evidence as a negative air-bone gap, and read cancelled 95-dB acoustic-reflex marks as present responses. The refined skill imposed a fixed sequence: verify the image and ear, inspect axes and legends, assign symbols, transcribe before calculating, preserve no-response entries, then derive and cross-check. When masking could not be assigned unambiguously, the prompt instructed the model to abstain rather than guess.</p>
<p>The results of the locked test were striking. In 50 development records, the skill raised hearing-loss-type agreement from 85.0 to 94.0 percent and acoustic-reflex agreement from 70.4 to 98.4 percent compared with a schema-only prompt. In the held-out evaluation of 101 independent patients using a Codex GPT5.5 agentic workflow, the system achieved 1809 of 1820 exact numeric thresholds, 99.4 percent, and 91 of 101 patients met every numeric and no-response criterion. Degree was correct in all 101 patients, hearing-loss type in 100, and tympanometry measurements in all 578 entries. Yet the Achilles&#8217; heel persisted: only 9 of 15 no-response entries were correct, and 18 of 19 air-bone-gap mismatches occurred because the model forced a negative label when the evidence supported an indeterminate one.</p>
<p>A post hoc robustness analysis using DeepSeek-V4-Flash-Vision-Exp on the same patients showed how much performance depends on the implementation. With the same locked skill, hearing-loss-type agreement rose from 37.6 to 70.3 percent, a dramatic improvement, but exact numeric-threshold agreement reached only 72.1 percent, reflex agreement hovered near 69 percent, and no-response agreement was zero. The authors are careful to note this was a descriptive cross-model comparison, not a matched foundation-model experiment, since the execution environments differed. The lesson, however, is clear: specialty guidance can substantially improve rule-dependent interpretation, but raw visual accuracy remains tied to the underlying model and workflow.</p>
<p>The second module addressed reporting. Here the AI received adjudicated structured values rather than its own image predictions, calculated twenty-five prespecified fields, and drafted separate Chinese professional and patient-facing reports. After two audiologists reviewed 60 initial cases and identified overreliance on the speech-frequency average, omission of high-frequency losses, and weak integration of history with results, the prompts were refined and locked. In the formal evaluation of 91 independent patients, all twenty-five rule-derived fields were correct in 91 of 91 Codex-workflow cases and 87 of 91 DeepSeek-workflow cases. Both audiologists rated every single report from both workflows as accurate or basically accurate; no report received a rating of clear error or potentially misleading.</p>
<p>An exploratory lay evaluation added a human dimension. Five lay raters compared pre-refinement patient-facing reports from the two workflows in 48 patients. DeepSeek reports were preferred for explanations in 44 of 48 patients and for next steps in 30, but they were also significantly longer, with a median of 392 versus 235 Chinese characters. Overall preference and perceived ease did not differ significantly. The authors emphasize that preference and readability proxies do not establish comprehension, and that no lay evaluation of the refined reports was conducted. This matters because hearing-health materials often exceed recommended reading levels, and limited health literacy can coexist with hearing loss in older adults.</p>
<p>The study&#8217;s limitations are candidly enumerated. It came from a single hospital with two devices over four months. Only 15 no-response entries existed, from just three validation patients. The two audiologists who rated the final reports were the same ones who had refined the prompts, raising the possibility of incorporation bias. Even with zero unfavorable ratings among 91 patients, the statistical upper bound on the unfavorable-report rate remains roughly 4.1 percent. Reproducibility was constrained by reliance on proprietary services: the exact Codex snapshot and sampling settings were unavailable, and provider data retention could not be excluded. No end-to-end test connected the two modules, so error propagation from image to report was never measured.</p>
<p>What the study ultimately offers is an architecture rather than a product. The authors envision a safety-conscious pipeline in which image transcription, deterministic validation and calculation, report drafting, uncertainty flags, and clinician approval remain visible, auditable handoff points. This aligns with what Chinese audiologists themselves have said in qualitative work: AI may absorb repetitive technical work, but communication, judgment, and responsibility should remain clinician-led. The findings support prospective evaluation of modular, clinician-supervised assistance, not autonomous diagnosis. The next step, the authors argue, is a silent prospective deployment that connects the modules while preserving intermediate outputs, measuring abstention, correction burden, review time, patient comprehension, and downstream clinical decisions. Until then, the audiogram-reading AI remains a promising apprentice, one that can transcribe nearly every threshold perfectly yet still needs its audiologist to teach it what silence means.</p>
<p><strong>Subject of Research:</strong> Audiologist-guided multimodal AI for interpreting pure-tone audiometry and tympanometry and generating clinical reports</p>
<p><strong>Article Title:</strong> Audiologist-Guided Multimodal AI for Pure-Tone Audiometry and Tympanometry Interpretation and Reporting</p>
<p><strong>Article References:</strong> Wu, X., Shen, X., Mo, C., Shao, S., Wang, J., &amp; Wang, S. (2026). Audiologist-Guided Multimodal AI for Pure-Tone Audiometry and Tympanometry Interpretation and Reporting. <em>Journal of Medical Systems, 50</em>(1), Article 140. <a href="https://doi.org/10.1007/s10916-026-02463-5" rel="noopener noreferrer">https://doi.org/10.1007/s10916-026-02463-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10916-026-02463-5" rel="noopener noreferrer">10.1007/s10916-026-02463-5</a></p>
<p><strong>Keywords:</strong> artificial intelligence, audiology, pure-tone audiometry, tympanometry, large language models, multimodal AI, clinical decision support, hearing loss, medical imaging AI, patient communication, prompt engineering, Journal of Medical Systems</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221274</post-id>	</item>
		<item>
		<title>Federated AI Predicts Sepsis in ICUs Without Sharing Patient Data</title>
		<link>https://scienmag.com/federated-ai-predicts-sepsis-in-icus-without-sharing-patient-data/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 18:15:35 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI-based sepsis prediction without patient data sharing]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[cross-institutional healthcare data collaboration]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for early sepsis detection]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[electronic health records for predictive analytics]]></category>
		<category><![CDATA[FedAvg]]></category>
		<category><![CDATA[federated learning]]></category>
		<category><![CDATA[federated learning in healthcare]]></category>
		<category><![CDATA[health informatics]]></category>
		<category><![CDATA[healthcare data privacy regulations and AI solutions]]></category>
		<category><![CDATA[ICU]]></category>
		<category><![CDATA[ICU patient monitoring with AI]]></category>
		<category><![CDATA[improving sepsis outcomes with federated AI]]></category>
		<category><![CDATA[innovative approaches to healthcare data governance]]></category>
		<category><![CDATA[machine learning models respecting patient privacy]]></category>
		<category><![CDATA[privacy-preserving AI]]></category>
		<category><![CDATA[privacy-preserving machine learning for ICU patients]]></category>
		<category><![CDATA[sepsis prediction]]></category>
		<category><![CDATA[SHAP explainability]]></category>
		<category><![CDATA[simulation of hospital data for AI training]]></category>
		<category><![CDATA[Transformer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217930</guid>

					<description><![CDATA[Researchers have built a federated deep learning system that predicts sepsis across simulated hospitals with 93.6 percent accuracy while keeping all patient data local.]]></description>
										<content:encoded><![CDATA[<p>Sepsis, the body&#8217;s runaway immune response to infection, remains one of the leading causes of death in intensive care units across the United States, killing patients not because clinicians lack treatments but because the window for intervention is brutally narrow. Every hour that passes without recognition of the syndrome measurably worsens survival odds, which is why hospitals have long sought computational tools that can flag deteriorating patients before the classic signs become unmistakable. Electronic health records hold the raw material for such early warnings: heart rates, laboratory values, medication records, and scores of other variables streaming in from every bedside. The obstacle has never been the data itself but the walls around it. Privacy regulations and institutional governance policies make it extraordinarily difficult for hospitals to pool patient records, which means most predictive models are trained on a single institution&#8217;s population and often fail to generalize elsewhere. A new study published in Discover Social Science and Health proposes a way around this impasse, using federated learning to train a powerful hybrid deep learning model across simulated hospitals without any patient data ever leaving its home institution.</p>
<p>The research team, led by Miad Islam of Saint Leo University together with collaborators from institutions in Bangladesh, the United Kingdom, and the United States, built a framework that combines three complementary neural network architectures into a single sepsis prediction engine. The first component is a one-dimensional convolutional neural network, or 1D-CNN, which excels at scanning sequences of clinical measurements and picking out local patterns, much as a radiologist might scan an image for telling features. Layered on top of that is a bidirectional long short-term memory network, or BiLSTM, a recurrent architecture that reads the patient&#8217;s clinical timeline in both directions, capturing how early vital-sign changes foreshadow later laboratory abnormalities and how recent developments reframe earlier ambiguity. The final piece is a Transformer attention block, the same family of architecture that powers modern language models, which allows the network to weigh the relative importance of different clinical variables and time points dynamically rather than treating every input as equally significant. Together, these components form a model capable of learning the subtle, multi-scale signatures that precede sepsis onset.</p>
<p>The privacy mechanism at the heart of the study is federated learning, a training paradigm in which the data never moves. Instead of shipping records to a central server, each participating hospital trains the model locally on its own patients. Only the resulting model parameters, essentially long lists of numerical weights, are transmitted to a coordinating server, which averages them using an algorithm known as Federated Averaging, or FedAvg. The updated global model is then sent back to each site for another round of local training. Over many such rounds, the shared model gradually absorbs the statistical diversity of all participating institutions without any single record, diagnosis, or lab value ever being exchanged. For the purposes of this study, the researchers simulated this arrangement by partitioning a MIMIC-IV-style ICU dataset across three virtual hospital clients, allowing them to evaluate how the federated approach behaves under realistic distributed conditions while working with data that posed no privacy risk.</p>
<p>The results are striking for a model trained under such constraints. The federated hybrid framework achieved an accuracy of 93.6 percent and an area under the receiver operating characteristic curve, or AUROC, of 0.959, a standard measure of a classifier&#8217;s ability to distinguish septic from non-septic patients across all possible thresholds. Precision came in at 0.871, meaning that when the model raised an alarm it was usually justified, and specificity reached 98.24 percent, indicating it rarely flagged healthy patients falsely. The F1-score, which balances precision against recall, was 0.759, a figure that reflects the inherent difficulty of detecting a condition that is rare relative to the ICU population. For comparison, the researchers also trained a centralized baseline model on all the data pooled together, and it reached an AUROC of 0.997. That gap is real, but the federated model&#8217;s performance remains firmly in the range clinicians would consider useful, and it was achieved without the data centralization that privacy law makes impractical.</p>
<p>Perhaps even more consequential for real-world deployment is the framework&#8217;s communication efficiency. In federated systems, the cost of repeatedly shipping model weights between hospitals and the coordinating server can become prohibitive, particularly for institutions with limited bandwidth. Across all federated training rounds in this study, the total communication cost was just 36.09 megabytes, a figure small enough to travel over ordinary hospital networks in seconds. That efficiency matters because it suggests the approach could scale to many more participating sites without the coordination overhead becoming a bottleneck. In distributed healthcare environments, where connectivity is often uneven and IT infrastructure varies widely between a major academic medical center and a community hospital, keeping the communication footprint light is not a luxury but a prerequisite.</p>
<p>A model that clinicians cannot understand is a model they will not trust, so the researchers turned to SHAP, or SHapley Additive exPlanations, a technique borrowed from cooperative game theory that assigns each input variable a quantified contribution to every individual prediction. The SHAP analysis identified the Sequential Organ Failure Assessment score, a composite measure of dysfunction across six organ systems, as the most influential predictor of sepsis risk, followed by lactate level, white blood cell count, creatinine, and procalcitonin. Each of these is a familiar marker to any intensivist: lactate rises when tissues are starved of oxygen, white blood cells surge or crash during systemic infection, creatinine signals kidney injury, and procalcitonin is a well-established biomarker of bacterial sepsis. The fact that the model&#8217;s attention converged on variables that already carry clinical weight is reassuring, because it suggests the network learned genuine physiology rather than exploiting statistical artifacts in the data.</p>
<p>The authors are candid about the limits of their work, and their honesty is itself instructive. The study used a MIMIC-IV-style dataset rather than live records from multiple hospitals, and the federated setup was a simulation rather than a deployment across real institutions with genuinely heterogeneous patient populations, equipment, and documentation practices. The researchers also note that they did not formally define an adversary or threat model. They did not assess whether a malicious server or a colluding client could infer patient-level information from the exchanged model updates, a known vulnerability in federated systems where gradient information can sometimes leak details about training data. Techniques such as differential privacy, secure aggregation, and homomorphic encryption exist to close these gaps, and the authors flag rigorous validation on real multi-institutional ICU records as a necessary next step before any claim of scalable, secure deployment can be made.</p>
<p>Those caveats notwithstanding, the study lands at a moment when the tension between data hunger and data privacy has become the defining challenge of clinical artificial intelligence. The most accurate models are typically those trained on the largest and most diverse datasets, yet the most diverse datasets are precisely the ones that privacy law keeps fragmented across institutions. Federated learning offers a mathematically principled compromise, and pairing it with architectures that capture both local temporal patterns and long-range dependencies, as this hybrid CNN-BiLSTM-Transformer design does, may prove to be a template for predictive medicine well beyond sepsis. Acute kidney injury, respiratory failure, and cardiac deterioration are all conditions where early, distributed, privacy-preserving prediction could change outcomes.</p>
<p>What makes this work compelling as a piece of the larger puzzle is its demonstration that privacy and performance need not be framed as a zero-sum trade. A model trained without ever seeing another hospital&#8217;s records came within a few percentage points of the centralized ideal, and it did so with a communication budget measured in tens of megabytes and with explanations that map onto established clinical intuition. The road from a three-client simulation to a network of hospitals sharing a living model is long, and it will require threat modeling, regulatory negotiation, and prospective clinical validation. But the direction of travel is clear. If the next generation of ICU decision support can be trained collectively while respecting the sovereignty of every patient record, the hours that matter most in sepsis may finally be spent on treatment rather than detection.</p>
<p><strong>Subject of Research:</strong> Privacy-preserving federated deep learning for sepsis prediction in intensive care units</p>
<p><strong>Article Title:</strong> A hybrid deep learning framework for privacy-preserving sepsis prediction in distributed ICU environments using federated learning simulation</p>
<p><strong>Article References:</strong> Islam, M., Mohiuddin, T., Rahman, M. A., Sharfuddin, M., Islam, M. S., Sunny, S. R., &amp; Lokesh, E. (2026). A hybrid deep learning framework for privacy-preserving sepsis prediction in distributed ICU environments using federated learning simulation. <em>Discover Social Science and Health</em>. <a href="https://doi.org/10.1007/s44155-026-00489-1" rel="noopener noreferrer">https://doi.org/10.1007/s44155-026-00489-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44155-026-00489-1" rel="noopener noreferrer">10.1007/s44155-026-00489-1</a></p>
<p><strong>Keywords:</strong> federated learning, sepsis prediction, deep learning, ICU, electronic health records, privacy-preserving AI, Transformer, BiLSTM, SHAP explainability, clinical decision support, health informatics, FedAvg</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217930</post-id>	</item>
		<item>
		<title>AI Model Flags Gastric Cancer in Patients With Psychological Symptoms</title>
		<link>https://scienmag.com/ai-model-flags-gastric-cancer-in-patients-with-psychological-symptoms/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 18:12:40 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[advancements in AI-based diagnostic tools for gastric cancer]]></category>
		<category><![CDATA[clinical challenges in diagnosing gastric tumors]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[distinguishing gastric cancer from benign gastric conditions]]></category>
		<category><![CDATA[early detection of gastric malignancies]]></category>
		<category><![CDATA[endoscopic surveillance guidelines for gastric precancerous conditions]]></category>
		<category><![CDATA[endoscopy]]></category>
		<category><![CDATA[gastric cancer]]></category>
		<category><![CDATA[gastric cancer detection]]></category>
		<category><![CDATA[impact of mental health on gastrointestinal disease diagnosis]]></category>
		<category><![CDATA[intestinal metaplasia]]></category>
		<category><![CDATA[intestinal metaplasia and gastric cancer]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in gastrointestinal diagnostics]]></category>
		<category><![CDATA[monocytes]]></category>
		<category><![CDATA[nested cross-validation]]></category>
		<category><![CDATA[predictive markers]]></category>
		<category><![CDATA[psychological symptoms]]></category>
		<category><![CDATA[psychological symptoms in cancer diagnosis]]></category>
		<category><![CDATA[role of AI in cancer screening]]></category>
		<category><![CDATA[serum albumin]]></category>
		<category><![CDATA[SHAP]]></category>
		<category><![CDATA[significance of psychological factors in gastric cancer]]></category>
		<category><![CDATA[XGBoost]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217894</guid>

					<description><![CDATA[Researchers in Qingdao developed an interpretable XGBoost model that distinguishes gastric cancer from intestinal metaplasia in patients with psychological symptoms using seven routine clinical features.]]></description>
										<content:encoded><![CDATA[<p>Gastric cancer remains one of the world&#8217;s most lethal malignancies, and one of its most insidious features is how quietly it can masquerade as something far less threatening. For patients who present with psychological symptoms such as anxiety or depression, the diagnostic picture becomes even murkier, because distress and dyspepsia often travel together and can obscure the warning signs of a developing tumor. A new study published in Cancer Causes &amp; Control by Shi-ran Wang, Yu-quan Mao, and Guo-jie Hu of The Affiliated Hospital of Qingdao University tackles this diagnostic gray zone head-on, using machine learning to separate two conditions that clinicians frequently struggle to tell apart at the bedside: current gastric cancer and intestinal metaplasia, a precancerous transformation of the stomach lining.</p>
<p>The stakes of this distinction are considerable. Intestinal metaplasia, or IM, is a well-recognized precursor state in the Correa cascade of gastric carcinogenesis, in which chronic inflammation drives the stomach&#8217;s normal epithelium to take on intestinal characteristics. Most patients with IM will never develop cancer, but their risk is elevated enough that international guidelines, including the 2025 MAPS III update from the European Society of Gastrointestinal Endoscopy, recommend structured endoscopic surveillance. Gastric cancer, by contrast, demands immediate oncological workup and treatment. When a patient with psychological symptoms arrives complaining of vague abdominal discomfort, weight loss, or appetite changes, deciding how urgently to pursue endoscopy and biopsy can determine whether a tumor is caught at a curable stage or discovered only after it has advanced.</p>
<p>The research team designed a retrospective, single-center, cross-sectional study that enrolled 302 patients with psychological symptoms, of whom 185 had intestinal metaplasia and 117 had gastric cancer confirmed by standard diagnostic pathways. The cohort was randomly split into a training set of 212 patients and an independently retained validation set of 90 patients, a design that guards against the optimistic bias that plagues many machine learning studies in medicine. Rather than feeding every available variable into their algorithms, the investigators first pruned the feature space using three complementary selection methods: least absolute shrinkage and selection operator regression, known as LASSO, the Boruta algorithm, and recursive feature elimination. This consensus approach converged on seven features that carried the most diagnostic signal: age, serum albumin, sex, the type of psychological symptom, total bilirubin, smoking status, and monocyte count.</p>
<p>With those seven variables in hand, the team compared seven different machine learning algorithms using nested cross-validation, a rigorous technique in which an inner loop tunes the model&#8217;s hyperparameters while an outer loop estimates how the tuned model will perform on unseen data. Nested cross-validation is widely regarded as the gold standard for honest performance estimation, yet it remains underused in clinical prediction research. Among the contenders, XGBoost, a gradient-boosted decision tree ensemble celebrated for its performance on tabular medical data, emerged as the winner with a mean outer-fold area under the receiver operating characteristic curve of 0.839, plus or minus 0.029. The AUC, a measure of how well a model ranks diseased patients above healthy ones across all thresholds, ranges from 0.5, equivalent to a coin flip, to 1.0, perfect discrimination.</p>
<p>When the final XGBoost model was unleashed on the 90-patient validation cohort it had never seen, it achieved an AUC of 0.793, with a 95 percent confidence interval spanning 0.670 to 0.895. Its overall accuracy reached 0.800, with a specificity of 0.891, meaning it correctly identified most patients with intestinal metaplasia, and a sensitivity of 0.657, meaning it caught roughly two-thirds of the actual cancers. The F1 score, which balances precision and recall, came in at 0.719, and the Brier score, a measure of calibration that penalizes both wrong predictions and misplaced confidence, was 0.167. Decision curve analysis, a method that quantifies the net clinical benefit of acting on a model&#8217;s predictions across a range of threshold probabilities, suggested that using the model could yield meaningful benefit compared with treating all patients or none.</p>
<p>What elevates this study above the crowded field of medical machine learning papers is its commitment to interpretability. Black-box models have earned well-deserved skepticism from clinicians who cannot act on predictions they cannot understand. The Qingdao team therefore applied SHAP, or Shapley Additive Explanations, a technique borrowed from cooperative game theory that assigns each feature a contribution value for every individual prediction. The SHAP analysis revealed that age and the type of psychological symptom were the leading drivers of the model&#8217;s classifications, followed by the laboratory markers. This transparency allows physicians to see not just what the model predicts but why, and to judge whether those reasons align with clinical plausibility.</p>
<p>The feature list itself tells an intriguing biological story. Serum albumin, a protein synthesized by the liver, tends to fall in the context of both malignancy-associated inflammation and malnutrition, and low albumin has repeatedly been linked to poor outcomes in gastric cancer. Total bilirubin, another hepatic marker, has been associated with the clinical characteristics of gastric cancer patients in prior hematological studies. Monocytes, the innate immune cells that circulate in blood and seed tumors as macrophages, are increasingly recognized as players in the tumor microenvironment, and an elevated monocyte count can reflect the systemic inflammatory milieu of an active cancer. Smoking status and sex are established epidemiological risk factors, while age simply reflects the cumulative probability of neoplastic progression. The prominence of psychological symptom type is perhaps the most novel finding, hinting that the pattern of a patient&#8217;s distress may carry diagnostic information that has been largely overlooked.</p>
<p>The connection between psychological symptoms and gastric pathology is not merely coincidental. Previous research has documented a high prevalence of psychological distress among gastric cancer patients, and experimental work has shown that chronic stress can accelerate gastric cancer progression, with recent studies implicating gut microbial metabolites in stress-driven tumor growth. Conversely, malnutrition and systemic inflammation associated with advanced malignancy can themselves produce depressive and anxious symptoms, creating a bidirectional loop. Patients with functional gastrointestinal disorders and precancerous conditions also experience elevated rates of anxiety and depression, which is precisely why the distinction between IM and cancer in this population is so difficult and so consequential.</p>
<p>The authors are careful, appropriately so, about the limits of their work. The model showed only moderate discrimination, and the study was retrospective and single-center, drawing on patients from one Chinese hospital. The team explicitly states that the model should not replace endoscopic or histopathological diagnosis, which remain the definitive standards, and that external multicenter validation is required before any broader clinical implementation. Psychological symptom classification in the study relied on established rating instruments, including the Hamilton scales for depression and anxiety, but symptom type as a predictor will need careful operationalization in other healthcare settings and languages.</p>
<p>Even with those caveats, the study points toward a genuinely useful clinical role. Imagine a triage tool embedded in the electronic health record that, for a patient with psychological symptoms and dyspeptic complaints, computes a probability of current gastric cancer from seven routinely collected variables. A high score could accelerate the endoscopy queue; a low score could support a more measured surveillance pathway consistent with guidelines for metaplasia. In health systems where endoscopic resources are scarce and waiting lists are long, such prioritization could translate directly into earlier cancer detection. The Qingdao model is not yet that tool, but it is a credible, honestly evaluated step along the way, and its emphasis on interpretability sets a standard that clinical machine learning studies would do well to follow. As gastric cancer continues to impose a heavy global burden, with incidence and mortality projected to remain substantial through 2035, every legitimate advance in distinguishing the benign from the malignant deserves attention.</p>
<p><strong>Subject of Research:</strong> Machine learning for distinguishing gastric cancer from intestinal metaplasia in patients with psychological symptoms</p>
<p><strong>Article Title:</strong> A machine learning model for distinguishing gastric cancer from intestinal metaplasia in patients with psychological symptoms</p>
<p><strong>Article References:</strong> A machine learning model for distinguishing gastric cancer from intestinal metaplasia in patients with psychological symptoms. (n.d.). <a href="https://doi.org/10.1007/s10552-026-02251-z" rel="noopener noreferrer">https://doi.org/10.1007/s10552-026-02251-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10552-026-02251-z" rel="noopener noreferrer">10.1007/s10552-026-02251-z</a></p>
<p><strong>Keywords:</strong> gastric cancer, intestinal metaplasia, machine learning, XGBoost, SHAP, psychological symptoms, endoscopy, predictive markers, nested cross-validation, clinical decision support, serum albumin, monocytes</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217894</post-id>	</item>
		<item>
		<title>AI Tells Doctors When It Is Unsure About Cancer Treatment Success</title>
		<link>https://scienmag.com/ai-tells-doctors-when-it-is-unsure-about-cancer-treatment-success/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Sun, 27 Sep 2026 19:51:36 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI confidence in medical predictions]]></category>
		<category><![CDATA[assessing treatment efficacy in lung cancer]]></category>
		<category><![CDATA[biomarker data integration in cancer therapy]]></category>
		<category><![CDATA[Biomarkers]]></category>
		<category><![CDATA[cancer treatment response prediction]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[conformal prediction]]></category>
		<category><![CDATA[cytokines]]></category>
		<category><![CDATA[early response indicators in cancer therapy]]></category>
		<category><![CDATA[FDG PET]]></category>
		<category><![CDATA[immune system biomarkers in cancer treatment]]></category>
		<category><![CDATA[Immunotherapy]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for cancer treatment outcomes]]></category>
		<category><![CDATA[metastatic non-small cell lung cancer]]></category>
		<category><![CDATA[oncology clinical decision support systems]]></category>
		<category><![CDATA[PD-L1]]></category>
		<category><![CDATA[personalized cancer therapy decision tools]]></category>
		<category><![CDATA[PET imaging in cancer prognosis]]></category>
		<category><![CDATA[T-cell receptor repertoire]]></category>
		<category><![CDATA[treatment response prediction]]></category>
		<category><![CDATA[tumor response prediction using AI]]></category>
		<category><![CDATA[uncertainty quantification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217087</guid>

					<description><![CDATA[University of Washington researchers have built a decision support system that combines PET imaging, immune repertoire sequencing, and blood cytokines to predict lung cancer treatment response while quantifying its own uncertainty.]]></description>
										<content:encoded><![CDATA[<p>Predicting whether a cancer patient will respond to therapy has long been one of oncology&#8217;s most consequential gambles. For people with metastatic non-small cell lung cancer, the most common and deadliest form of the disease, first-line chemoimmunotherapy can produce dramatic tumor shrinkage in some patients while leaving others with little benefit and substantial toxicity. Now, a research team at the University of Washington and Fred Hutchinson Cancer Center has built a prototype clinical decision support system that fuses three very different biological signals into a single prediction, and, crucially, tells clinicians exactly how much it trusts each answer.</p>
<p>The study, published in the Journal of Translational Medicine, drew on data from the PET-BRIGHT clinical trial, in which thirty-five patients with metastatic non-small cell lung cancer received the standard combination of carboplatin, pemetrexed, and the immunotherapy pembrolizumab. Each patient underwent a FDG-PET/CT scan and a blood draw at two moments: before treatment began, and again after just three weeks, a single cycle of therapy. From these timepoints the researchers extracted three families of biomarkers: imaging metrics that quantify how much glucose metabolically active tumor tissue consumes, measures of T-cell receptor diversity that capture the breadth of the immune system&#8217;s anti-cancer repertoire, and a panel of inflammatory cytokines circulating in the blood.</p>
<p>The choice of these three modalities reflects a growing recognition that no single biomarker can capture the complexity of immunotherapy response. FDG-PET imaging measures uptake of a radioactive glucose analog, revealing where and how aggressively tumor cells are metabolizing energy, with summary metrics such as standardized uptake value and total lesion glycolysis summarizing whole-body tumor burden. T-cell receptor sequencing reads out the diversity of immune recognition, on the theory that a richer repertoire of tumor-targeting T cells signals a stronger response to immune checkpoint blockade. Cytokines, the small signaling proteins through which immune cells communicate, offer a systemic window into inflammatory states that may either support or suppress anti-tumor immunity.</p>
<p>Against these signals, the researchers benchmarked the current clinical standard: PD-L1 tumor proportion score, an immunohistochemistry measurement that guides treatment selection in lung cancer today. The results were sobering for the status quo. PD-L1 achieved an area under the receiver operating characteristic curve of just 0.58, barely better than a coin flip at distinguishing responders from non-responders. This weak discrimination, the authors note, is precisely what motivates a multimodal, uncertainty-aware approach, because patients and clinicians currently must commit to months of therapy armed with a biomarker that carries limited predictive weight.</p>
<p>The machine learning pipeline itself was deliberately conservative, a design choice appropriate for a small cohort. Using nested leave-one-out cross-validation with automated feature selection, the team identified a single representative biomarker per modality, guarding against the overfitting that plagues models trained on dozens of features and only a handful of patients. Classification was performed with class-balanced logistic regression, an interpretable method chosen over more opaque architectures, and performance was assessed both by discrimination, how well the model separated responders from non-responders, and by calibration, whether predicted probabilities matched observed outcomes as measured by the Brier score.</p>
<p>The headline innovation is the application of conformal prediction, a statistical framework that converts a model&#8217;s raw output into prediction sets with formal guarantees. Rather than declaring that a patient will respond or not, the conformal framework targets a specified reliability level, in this case eighty percent, and returns either a confident single-class prediction or a set containing both possibilities, signaling that the model cannot confidently classify that patient. The proportion of confident singleton predictions therefore becomes a direct, interpretable measure of uncertainty at the level of the individual patient, not merely the population average that conventional accuracy metrics provide.</p>
<p>The performance results chart a coherent story about how biomarker information evolves over treatment. Before therapy began, T-cell receptor diversity was the strongest single modality, reaching an AUROC of 0.80, followed by PET imaging at 0.76, while cytokines lagged at 0.54. Combining PET and T-cell data through late fusion, in which separately trained modality-specific models are merged, pushed pre-treatment discrimination to 0.85, and raised the proportion of confident singleton predictions from 78 percent to 89 percent. After the three-week blood draws and scans were added, the hierarchy shifted: cytokines became the strongest single modality at 0.78 while T-cell metrics were attenuated at 0.66, and late fusion of PET and cytokines achieved the highest overall discrimination of 0.86. Thirteen of twenty-two model and timepoint combinations beat a permutation-based null at the conventional significance threshold, and empirical coverage landed at 78 percent and 82 percent, close to the eighty percent reliability target the conformal framework was set to honor.</p>
<p>That dynamic shift in which modality carries the most predictive weight is itself scientifically interesting. It suggests that pre-treatment immune repertoire breadth anticipates who will benefit from checkpoint blockade, but that once therapy begins, the inflammatory milieu measured in plasma becomes the more informative readout of whether the treatment is taking hold. A static biomarker panel frozen at baseline would miss this evolution entirely, and the study&#8217;s longitudinal design, with its week-three reassessment, demonstrates how a decision support system could update its confidence as new evidence accumulates over the first critical weeks of therapy.</p>
<p>The authors are appropriately careful about the limitations of what they have built. Thirty-five patients from a single center constitute a proof-of-concept feasibility study, and the findings are explicitly framed as hypothesis-generating rather than practice-changing. Small cohorts make any machine learning model vulnerable to optimistic bias, which is why the permutation testing and conservative feature selection matter so much here. Before such a system could inform real treatment decisions, it would need validation in larger, multi-institution cohorts spanning different tumor types, treatment regimens, and demographic populations, along with calibration checks to confirm that the conformal guarantees hold outside the development data.</p>
<p>Even so, the prototype points toward a future in which clinical AI systems do something more honest than issuing confident verdicts. By pairing multimodal biomarker fusion with formal uncertainty quantification, the framework identifies not only which patients are likely to respond, but also, and perhaps more valuably, those for whom no confident prediction can be made, flagging them for closer monitoring or alternative strategies. In a disease where the current standard biomarker performs barely better than chance, knowing the limits of prediction may prove as clinically important as the predictions themselves. The interface prototype developed alongside the statistical framework shows how such uncertainty could be surfaced directly to oncologists, turning a black-box risk score into a transparent, caveat-aware recommendation that clinicians can weigh against the realities of each patient&#8217;s care.</p>
<p><strong>Subject of Research:</strong> Uncertainty-aware multimodal prediction of chemoimmunotherapy response in metastatic non-small cell lung cancer</p>
<p><strong>Article Title:</strong> Toward uncertainty-aware clinical decision support for treatment response prediction in metastatic NSCLC: integrating FDG-PET, T-cell repertoire, and cytokines with conformal prediction</p>
<p><strong>Article References:</strong> Yaseen, F., Hippe, D. S., Cui, S., Fu, J., Kim, Y., Grassberger, C., Deng, L., Ye, T., Kinahan, P. E., Zeng, J., Gennari, J. H., &amp; Bowen, S. R. (2026). Toward uncertainty-aware clinical decision support for treatment response prediction in metastatic NSCLC: integrating FDG-PET, T-cell repertoire, and cytokines with conformal prediction. <em>Journal of Translational Medicine</em>. <a href="https://doi.org/10.1186/s12967-026-09030-z" rel="noopener noreferrer">https://doi.org/10.1186/s12967-026-09030-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12967-026-09030-z" rel="noopener noreferrer">10.1186/s12967-026-09030-z</a></p>
<p><strong>Keywords:</strong> metastatic non-small cell lung cancer, FDG-PET, T-cell receptor repertoire, cytokines, conformal prediction, clinical decision support, immunotherapy, PD-L1, biomarkers, machine learning, uncertainty quantification, treatment response prediction</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217087</post-id>	</item>
		<item>
		<title>AI Joins the Slide and the Chart to Predict Colorectal Cancer Risk</title>
		<link>https://scienmag.com/ai-joins-the-slide-and-the-chart-to-predict-colorectal-cancer-risk/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 22:38:51 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced machine learning techniques for personalized cancer risk prediction]]></category>
		<category><![CDATA[AI-driven colorectal cancer risk prediction]]></category>
		<category><![CDATA[artificial intelligence in oncology]]></category>
		<category><![CDATA[cancer diagnosis]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[Colorectal cancer]]></category>
		<category><![CDATA[combining imaging and clinical observations for cancer risk assessment]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[early detection of colorectal cancer using AI]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[high accuracy in cancer risk stratification]]></category>
		<category><![CDATA[histopathology]]></category>
		<category><![CDATA[integration of histopathological images and clinical data]]></category>
		<category><![CDATA[multimodal AI]]></category>
		<category><![CDATA[multimodal deep learning models for cancer diagnosis]]></category>
		<category><![CDATA[neural computing applications in medical diagnosis]]></category>
		<category><![CDATA[neural network architecture for medical imaging]]></category>
		<category><![CDATA[predictive analytics for cancer prognosis]]></category>
		<category><![CDATA[predictive medicine]]></category>
		<category><![CDATA[risk stratification]]></category>
		<category><![CDATA[SMOTE]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[VGG16]]></category>
		<category><![CDATA[VGG16 deep learning model in healthcare]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216825</guid>

					<description><![CDATA[A multimodal deep learning model combining histopathological images with clinical data achieved 97.17 percent accuracy in stratifying colorectal cancer patients into high- and low-risk groups.]]></description>
										<content:encoded><![CDATA[<p>Colorectal cancer remains one of the most common and deadliest malignancies worldwide, and the difference between a favorable outcome and a devastating one often hinges on how early the disease is caught and how accurately a patient&#8217;s risk is stratified. A new study published in Neural Computing and Applications by a team of researchers led by P. Margaret Savitha of Christ University in Bangalore, India, takes aim at exactly this problem with an artificial intelligence system that refuses to look at just one kind of data. Instead of relying solely on tissue images or solely on clinical measurements, the team built a multimodal deep learning model that fuses histopathological imagery with structured clinical observations, producing a single, unified risk verdict for each patient. The results are striking: on a held-out test set, the model achieved an overall accuracy of 97.17 percent, a sensitivity of 94.33 percent, a perfect specificity of 100 percent, and an area under the receiver operating characteristic curve of 0.9987.</p>
<p>The architecture at the heart of the study is built around a Visual Geometry Group network with sixteen layers, universally known in the deep learning community as VGG16. Developed originally for large-scale natural image recognition, VGG16 is a convolutional neural network characterized by stacks of small three-by-three convolutional filters arranged in progressively deeper blocks, each followed by pooling layers that shrink the spatial dimensions of the feature maps while increasing their depth. What VGG16 lacks in architectural novelty it makes up for in reliability and transferability: its filters, particularly when initialized with weights pre-trained on massive image corpora, capture texture, granularity, and structural patterns that translate remarkably well to biomedical imagery. In this study, the image branch of the model ingested histopathology tiles — small, high-resolution crops of stained tissue sections — and learned to extract morphological signatures associated with malignancy, such as glandular disorganization, nuclear atypia, and abnormal cellular density.</p>
<p>But tissue morphology tells only part of the story in colorectal cancer. Clinicians weighing treatment intensity and surveillance frequency also depend on systemic indicators: patient demographics, laboratory values, tumor characteristics, and other clinical observations recorded in patient charts. The second branch of the model was therefore designed to process these tabular clinical features. Numerical inputs were normalized using standard scaling techniques — the kind of min-max and z-score transformations long used to place heterogeneous clinical variables on comparable footing — so that no single measurement would dominate the learning process simply because of its units. The two branches, one visual and one tabular, were then concatenated into a joint representation, a fused feature vector in which morphological evidence and systemic evidence coexist. A sigmoid classifier sits at the end of this pipeline, outputting a probability that maps each patient into a high-risk or low-risk category.</p>
<p>The training data combined two complementary sources. On the imaging side, the researchers worked with a dataset of 5,000 histopathology image tiles, drawing on the widely used Kather texture collection of colorectal tissue patches. On the clinical side, they incorporated 235 patient records from a de-identified colorectal dataset, ensuring that the model learned from real-world clinical profiles rather than synthetic constructs. Because clinical datasets of this size are often imbalanced — with one risk class outnumbering the other — the team employed the Synthetic Minority Over-sampling Technique, or SMOTE, a well-established algorithm that generates synthetic examples of the underrepresented class by interpolating between existing minority samples in feature space. This step is critical: without it, a classifier can achieve deceptively high accuracy by simply predicting the majority class every time, while completely failing the patients who matter most.</p>
<p>The performance numbers reported on the held-out test set deserve close scrutiny. An accuracy of 97.17 percent means the model assigned the correct risk category to nearly every patient it evaluated. The sensitivity of 94.33 percent is arguably the most clinically meaningful figure: it reflects the proportion of genuinely high-risk patients the system correctly flagged, and missing such patients — false negatives — carries the gravest consequences in oncology. The specificity of 100 percent indicates that not a single low-risk patient in the test set was wrongly classified as high risk, avoiding unnecessary anxiety, invasive follow-up procedures, and overtreatment. The area under the curve of 0.9987, a measure of the model&#8217;s ability to discriminate between classes across all possible decision thresholds, sits tantalizingly close to the theoretical maximum of 1.0. Crucially, the multimodal system outperformed unimodal baselines that saw only images or only clinical data, providing empirical support for the central thesis of the work: morphology and systemic context are complementary signals, and their fusion yields a risk assessment neither could deliver alone.</p>
<p>What elevates this work beyond a leaderboard result is its attention to interpretability. Deep learning models are notoriously opaque — millions of parameters interact in ways that resist straightforward human audit — and this opacity has been a persistent barrier to clinical adoption. The researchers addressed this by building explainability and automated report generation directly into their framework. Rather than emitting an inscrutable risk score, the system translates its predictions into analysis and explainable forms, producing reports that clinicians can read, question, and act upon. This design philosophy aligns with a broader movement in medical artificial intelligence, visible across recent literature on colorectal cancer diagnosis, that treats explainable AI not as an optional add-on but as a prerequisite for trust. When a pathologist or oncologist can see why a model reached its conclusion, the technology shifts from a black-box oracle to a collaborative second opinion.</p>
<p>The study sits within a rapidly accelerating field. Recent years have witnessed deep learning diagnostic frameworks for colorectal cancer built on histopathological images, explainable deep learning systems for colon cancer diagnosis, and interpretable machine learning platforms that read pathology slides directly. Transfer learning — the practice of repurposing networks pre-trained on general imagery for specialized medical tasks — has enabled large emulated prospective studies of pathological diagnosis, while convolutional approaches combined with support vector machines have been used to predict prognosis and mutational signatures from routine hematoxylin and eosin slides. Parallel efforts have explored blood-based multiomics integration, serum glycoproteome profiling, exosomal proteomic signatures, microbiome biomarker discovery, and microRNA markers for early detection. The Bangalore team&#8217;s contribution to this landscape is the explicit marriage of the imaging pipeline with the clinical chart, a fusion strategy that mirrors how physicians actually reason, integrating what they see under the microscope with what they know about the patient as a whole.</p>
<p>The clinical implications of reliable, automated risk stratification are substantial. In current practice, risk assessment in colorectal cancer leans heavily on the TNM staging system, which classifies tumors by depth of invasion, nodal involvement, and metastatic status, supplemented by molecular markers such as microsatellite instability and histological features like lymphovascular and perineural invasion. Each of these inputs is valuable but subject to interobserver variability, and integrating them into a coherent treatment plan is a cognitively demanding task. A validated multimodal AI system could serve as a consistent, tireless adjunct — triaging patients into high-risk and low-risk groups to guide the intensity of adjuvant therapy, the frequency of surveillance colonoscopy, and the urgency of specialist referral. In settings with limited access to experienced pathologists, such systems could democratize diagnostic expertise, extending high-quality risk assessment to hospitals and clinics that lack subspecialty staffing.</p>
<p>Caution is nonetheless warranted before such tools reach the clinic. The model was trained and evaluated on a dataset of 235 patient records and 5,000 image tiles, and while the held-out test results are exceptional, external validation on independent, multi-center cohorts remains the essential next step for any diagnostic AI. The authors themselves note that the study involved no human participants and used publicly available, anonymized data, meaning that prospective clinical evaluation — ideally designed as an emulated or true trial — has yet to be performed. Dataset shift, staining variability between laboratories, scanner differences, and demographic heterogeneity can all erode performance when a model trained on one population is deployed on another, a lesson repeatedly demonstrated across the medical imaging literature. The path from a 97 percent test accuracy to a deployed clinical decision-support tool runs through regulatory review, workflow integration studies, and careful monitoring for failure modes.</p>
<p>Even with those caveats, the study offers a compelling glimpse of where cancer risk assessment is heading. The era of artificial intelligence that looks at a single data type in isolation is giving way to systems that, like experienced clinicians, synthesize evidence across modalities — the architecture of the tissue and the physiology of the patient considered together. If the near-perfect discrimination reported here can be replicated in larger, more diverse cohorts, multimodal deep learning could become a standard component of colorectal cancer care, catching high-risk patients earlier, sparing low-risk patients unnecessary intervention, and doing so with reports that doctors can actually understand. For a disease that will claim hundreds of thousands of lives this year, a model that fuses the microscope and the medical record into one coherent, explainable verdict is more than an incremental technical achievement; it is a template for how machine intelligence can be woven into the most consequential decisions in medicine.</p>
<p><strong>Subject of Research:</strong> Multimodal deep learning for colorectal cancer risk stratification using histopathological images and clinical data</p>
<p><strong>Article Title:</strong> Multimodal deep learning for colorectal cancer risk assessment</p>
<p><strong>Article References:</strong> Multimodal deep learning for colorectal cancer risk assessment. (n.d.). <a href="https://doi.org/10.1007/s00521-026-12473-6" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12473-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12473-6" rel="noopener noreferrer">10.1007/s00521-026-12473-6</a></p>
<p><strong>Keywords:</strong> colorectal cancer, deep learning, multimodal AI, VGG16, histopathology, risk stratification, clinical decision support, explainable AI, SMOTE, transfer learning, cancer diagnosis, predictive medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216825</post-id>	</item>
		<item>
		<title>Fuzzy Logic Gives Big Data Decision-Making a Flexible New Edge in Medicine</title>
		<link>https://scienmag.com/fuzzy-logic-gives-big-data-decision-making-a-flexible-new-edge-in-medicine/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 21:25:22 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive data clustering in healthcare]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[big data analytics in healthcare]]></category>
		<category><![CDATA[BRATS dataset]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[decision systems for shifting medical landscapes]]></category>
		<category><![CDATA[dynamic clustering in medicine]]></category>
		<category><![CDATA[FESO-DKMCA]]></category>
		<category><![CDATA[flexible clustering algorithms for patient data]]></category>
		<category><![CDATA[Fuzzy Ebola Search Optimized Dynamic K-means Clustering]]></category>
		<category><![CDATA[fuzzy logic]]></category>
		<category><![CDATA[fuzzy logic in medical decision-making]]></category>
		<category><![CDATA[fuzzy membership functions in medical datasets]]></category>
		<category><![CDATA[handling outdated medical information]]></category>
		<category><![CDATA[histogram equalization]]></category>
		<category><![CDATA[hybrid algorithms for clinical data]]></category>
		<category><![CDATA[improving clinical decision accuracy]]></category>
		<category><![CDATA[K-means clustering]]></category>
		<category><![CDATA[knowledge-based decision-making]]></category>
		<category><![CDATA[local binary patterns]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[medical big data analysis techniques]]></category>
		<category><![CDATA[Medical Imaging]]></category>
		<category><![CDATA[uncertainty quantification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216409</guid>

					<description><![CDATA[Researchers in India report that a fuzzy logic-driven dynamic clustering algorithm outperformed standard machine learning methods on brain tumor data, pointing toward medical decision systems that adapt as knowledge evolves.]]></description>
										<content:encoded><![CDATA[<p>Medicine moves fast, and the knowledge that clinical decisions rest on can age with alarming speed. A study published in the journal Knowledge and Information Systems tackles this problem head-on, presenting a comparative examination of knowledge-based decision-making models that weave fuzzy logic into the big data paradigm. The research team, led by Honganur Raju Manjunath of JAIN (Deemed-to-be University) in Bangalore together with colleagues from institutions across India, argues that when medical information is anchored to static or outdated sources, even the most sophisticated analytics pipeline can quietly drift toward obsolescence. Their answer is a hybrid algorithm that lets data points belong to clusters by degree rather than by hard boundary, allowing a decision system to track a shifting medical landscape instead of freezing it in time.</p>
<p>The heart of the proposal is a mouthful with a memorable acronym: the Fuzzy Ebola Search Optimized Dynamic K-means Clustering Algorithm, or FESO-DKMCA. Classical K-means clustering assigns every data point to exactly one cluster, drawing crisp walls through the data space. That rigidity becomes a liability in medical datasets, where patients, images, and biomarkers rarely fall neatly into categories. FESO-DKMCA replaces those walls with fuzzy membership functions, so each point carries a graded degree of belonging to multiple clusters simultaneously. As new data arrive and distributions shift, the memberships are continuously reassigned, and the cluster structure evolves with them. The Ebola Search optimization component then steers the clustering process toward configurations that best fit the data, giving the model a dynamic quality that static clustering schemes lack.</p>
<p>The conceptual machinery here draws on a decades-old insight from Lotfi Zadeh&#8217;s fuzzy set theory: real-world categories are matters of degree. A tumor region on a brain scan is not simply malignant or benign; a tissue texture is not simply normal or abnormal. By encoding such gradations mathematically, fuzzy membership functions let an algorithm express uncertainty rather than hide it. In the big data setting, where millions of records stream through analytics platforms, that expressiveness matters. The authors position their work within a broader literature on clinical decision support systems, smart healthcare, and the quantification of uncertainty in machine-assisted medical decisions, arguing that fuzzy approaches are particularly well suited to environments where knowledge itself is continuously revised.</p>
<p>To test the idea, the team ran a comparative analysis against standard machine learning methods on the BRATS medical dataset, a widely used benchmark of brain tumor magnetic resonance imaging scans. The experimental pipeline followed a classic three-stage design. First, the raw images were pre-processed using histogram equalization, a technique that redistributes pixel intensities across the image&#8217;s dynamic range to enhance contrast and make subtle tissue boundaries more visible to downstream analysis. Second, feature extraction was performed with local binary patterns, or LBP, a texture descriptor that encodes local spatial structure by comparing each pixel with its neighbors and recording the results as a compact binary code. These LBP features then served as the input space on which the clustering and classification methods competed.</p>
<p>The results, as reported in the paper, favor the fuzzy approach across every headline metric. FESO-DKMCA achieved an accuracy of 95 percent, precision of 90 percent, an F1 score of 89 percent, sensitivity of 96 percent, specificity of 91 percent, and an area under the receiver operating characteristic curve of 92 percent. Each number tells a distinct story. Sensitivity, the proportion of true positive cases correctly identified, is especially critical in oncology, where a missed tumor carries far graver consequences than a false alarm. The 96 percent sensitivity figure suggests the fuzzy clustering preserved fine distinctions that crisper methods blurred. Specificity, at 91 percent, indicates the model did not pay for that vigilance with an avalanche of false positives, a balance that many high-recall classifiers struggle to strike.</p>
<p>The F1 score of 89 percent, which harmonizes precision and recall into a single figure, and the AUC of 92 percent, which summarizes discriminative ability across all classification thresholds, round out a profile of consistent rather than cherry-picked performance. In comparative terms, the authors report that these figures establish the superiority of FESO-DKMCA over the conventional machine learning baselines included in the study. While the abstract does not enumerate every baseline score, the framing is clear: the dynamic, membership-driven clustering outperformed static alternatives on the same pre-processed features, isolating the fuzzy mechanism as the decisive variable in the comparison.</p>
<p>Why should fuzzy membership functions translate into better performance on medical data? The authors&#8217; argument runs through the nature of the datasets themselves. Medical data are noisy, high-dimensional, and riddled with overlapping classes. Two patients with similar imaging profiles may follow different clinical trajectories; two tissue samples with adjacent texture signatures may belong to different diagnostic categories. Hard clustering forces the algorithm to commit to a partition before it has enough evidence, and that premature commitment propagates errors downstream into whatever decision the system supports. Fuzzy clustering, by contrast, defers commitment. It keeps multiple hypotheses alive, weighted by their degree of support, and only resolves them as the evidence accumulates. In a big data environment where evidence accumulates constantly, that deferred-commitment architecture becomes a genuine advantage rather than a computational luxury.</p>
<p>The study also speaks to a problem that has haunted clinical decision support systems since their inception: knowledge decay. Reviews of the field have documented both the promise and the risks of computer-assisted diagnosis, and ethicists have warned that algorithmic decision-making in healthcare demands careful governance. The Indian team&#8217;s contribution is architectural rather than purely algorithmic. By building continuous reassignment into the clustering core, they create a system whose internal representation of knowledge is never final. When research advances render old categories obsolete, the model does not need to be rebuilt from scratch; its memberships migrate toward the new structure in the data. In principle, this makes decisions remain grounded in current knowledge, which is precisely the failure mode the authors set out to address.</p>
<p>The work sits at a busy intersection of research currents. Bibliometric analyses have charted the explosive growth of fuzzy techniques in big data applications, and surveys have catalogued the deployment of machine learning across IoT-enabled smart healthcare platforms, from infectious disease decision support to supplier selection in healthcare supply chains. Fuzzy graph structures have been applied to decision-making analysis, fuzzy systems have modeled customer loyalty and social network dynamics, and uncertainty quantification has been championed as a necessity in machine-assisted medicine. FESO-DKMCA synthesizes these threads, combining an evolutionary optimization strategy with fuzzy clustering and validating the combination on a benchmark that the medical imaging community recognizes. The authors also connect their approach to the wider big data decision-making literature spanning finance, smart cities, and board-level strategy, suggesting the fuzzy dynamic clustering pattern could generalize beyond medicine.</p>
<p>Caveats remain, as they always do. The evaluation rests on a single medical dataset, and the authors note that no new datasets were generated or analyzed during the study, relying instead on the publicly available BRATS 2018 collection. Real clinical deployment would demand validation across modalities, populations, and hospitals, along with the interpretability and ethical safeguards that regulators and bioethicists increasingly require. Still, the study offers a concrete demonstration that the mathematics of vagueness, far from being a philosophical curiosity, can outperform rigid alternatives on the messy, shifting data that modern medicine actually produces. As big data analytics matures, the lesson of FESO-DKMCA may prove durable: in domains where knowledge itself evolves, the algorithms that thrive are the ones built to change their minds.</p>
<p><strong>Subject of Research:</strong> Fuzzy logic-based knowledge-driven decision-making in big data analytics for medical applications</p>
<p><strong>Article Title:</strong> A comparative examination of knowledge-based decision-making using fuzzy logic in the big data landscape</p>
<p><strong>Article References:</strong> Manjunath, H. R., Sutaria, K., Loonkar, S., Wadhwa, B., Gupta, S. K., &amp; Dey, P. (2026). A comparative examination of knowledge-based decision-making using fuzzy logic in the big data landscape. <em>Knowledge and Information Systems, 68</em>(1), Article 266. <a href="https://doi.org/10.1007/s10115-026-02874-3" rel="noopener noreferrer">https://doi.org/10.1007/s10115-026-02874-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10115-026-02874-3" rel="noopener noreferrer">10.1007/s10115-026-02874-3</a></p>
<p><strong>Keywords:</strong> fuzzy logic, big data, machine learning, knowledge-based decision-making, FESO-DKMCA, K-means clustering, clinical decision support, BRATS dataset, local binary patterns, histogram equalization, medical imaging, uncertainty quantification</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216409</post-id>	</item>
		<item>
		<title>AI Framework Bridges the Knowledge Gap in Safer Medication Recommendations</title>
		<link>https://scienmag.com/ai-framework-bridges-the-knowledge-gap-in-safer-medication-recommendations/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 21:06:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI medication recommendation frameworks]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[bucket effect in AI systems]]></category>
		<category><![CDATA[challenges in AI medication models]]></category>
		<category><![CDATA[clinical dataset performance enhancements]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[cross-modal alignment]]></category>
		<category><![CDATA[drug information sources in AI models]]></category>
		<category><![CDATA[drug knowledge representation]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[improvements in AI-driven clinical decision support]]></category>
		<category><![CDATA[knowledge gap in healthcare]]></category>
		<category><![CDATA[knowledge graph]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[medication recommendation]]></category>
		<category><![CDATA[MIMIC-III]]></category>
		<category><![CDATA[MIMIC-IV]]></category>
		<category><![CDATA[molecular representation]]></category>
		<category><![CDATA[multi-knowledge integration in clinical AI]]></category>
		<category><![CDATA[multi-modal data in healthcare AI]]></category>
		<category><![CDATA[pharmacotherapy]]></category>
		<category><![CDATA[safety in electronic health record analysis]]></category>
		<category><![CDATA[unified knowledge space for medication suggestions]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216349</guid>

					<description><![CDATA[Researchers in China have developed MKMed, a cross-modal AI framework that aligns five types of drug knowledge to overcome the so-called bucket effect and improve the accuracy and safety of medication recommendations on clinical benchmark datasets.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence systems that recommend medications to hospital patients have long been promised as a way to help clinicians sift through complex electronic health records and arrive at safer, more effective treatment combinations. Yet a fundamental weakness has quietly undermined many of these systems, and a new study has finally given it a name: the bucket effect. Researchers at Yanshan University in China, writing in the journal Applied Intelligence, describe how the performance of medication recommendation models is often constrained not by their strongest knowledge sources but by their weakest ones, much as a barrel can only hold as much water as its shortest stave allows. Their proposed solution, a framework called MKMed for Multi-Knowledge Medication recommendation, aligns five different kinds of drug knowledge into a single unified representation space, and in doing so delivers measurable improvements over the most advanced existing systems on two of the most widely used clinical datasets in the field.</p>
<p>The core problem the team identified is deceptively simple. Modern medication recommendation models increasingly enrich their understanding of each drug by drawing on multiple complementary sources of knowledge: free-text descriptions of what a medication does, images associated with the compound, its molecular structure, its chemical properties, and its position within large biomedical knowledge graphs that map relationships between drugs, diseases, proteins and side effects. Prior studies have shown that incorporating more of this medication-related knowledge significantly improves the quality of the internal representations a model builds for each drug. But here lies the catch: not every medication is blessed with all of these knowledge types simultaneously. Some drugs may have rich textual descriptions but no available molecular imaging data. Others may be well characterized structurally yet nearly absent from curated knowledge graphs. The availability of knowledge across modalities is uneven, incomplete and, crucially, unevenly distributed in ways that vary from drug to drug.</p>
<p>The researchers went beyond merely observing this imbalance. Through a comprehensive statistical analysis of the distribution of modality coverage across medications, they quantified just how severe the bucket effect actually is, demonstrating that a substantial fraction of drugs lack one or more of the knowledge types that state-of-the-art models implicitly assume will be present. When a model encounters a drug missing a modality it was trained to exploit, the quality of that drug&#8217;s representation degrades, and with it the quality of the entire recommendation. In a clinical setting, where a recommendation engine might be asked to suggest a combination of several medications for a patient with multiple chronic conditions, a single poorly represented drug can drag down the coherence and safety of the whole prescription. The bucket effect, in other words, is not a marginal nuisance but a structural flaw baked into the data landscape of pharmacology itself.</p>
<p>MKMed attacks the problem at its architectural root. The centerpiece of the framework is a cross-modal medication encoder whose job is to take heterogeneous knowledge modalities, each with its own format, dimensionality and statistical character, and project them into a shared representation space where a drug&#8217;s identity is captured consistently regardless of which knowledge types happen to be available for it. The encoder is pre-trained using contrastive learning, a technique that has transformed representation learning across machine learning in recent years. In contrastive pre-training, the model is shown pairs or groups of examples and taught to pull together the representations of items that are semantically related while pushing apart those that are not. Applied here, contrastive learning across five complementary modalities, text, image, molecular structure, chemical properties and knowledge graph embeddings, teaches the encoder to recognize the underlying unity of a drug beneath its fragmented data trail.</p>
<p>The technical machinery behind this alignment draws on several strands of recent research. The molecular structure modality is processed using graph neural networks of the kind popularized by powerful message-passing architectures, treating each molecule as a graph of atoms and bonds. Textual descriptions are encoded with language models in the lineage of large-scale natural language pre-training, while image data is handled with vision transformer approaches that have become standard for visual recognition at scale. Knowledge graph information is embedded using translation-based techniques originally developed for multi-relational data, drawing on resources such as the Drug Repurposing Knowledge Graph and the PubChem database of chemical properties. The pre-training pipeline also incorporates cheminformatics tooling, including the RDKit library, to extract chemical descriptors. By the time the encoder has finished pre-training, each drug carries a representation that reflects whatever knowledge exists for it, and degrades gracefully, rather than catastrophically, when some of that knowledge is missing.</p>
<p>Once the cross-modal encoder has produced its unified drug representations, MKMed integrates them with patient electronic health record data to generate personalized medication recommendations. The patient side of the task is itself challenging: an EHR contains diagnoses, procedures and laboratory findings that must be synthesized into a picture of what a particular person actually needs. The framework fuses this patient context with the aligned drug representations to score candidate medications and assemble them into a recommended set. The design philosophy is that neither side of the equation should be impoverished by the other&#8217;s gaps. A patient whose record points toward a rarely documented drug should still receive a sound recommendation, because that drug&#8217;s representation has been anchored in whatever knowledge does exist and aligned with the shared space occupied by better-documented alternatives.</p>
<p>The empirical case for the approach rests on extensive experiments against state-of-the-art baseline systems on MIMIC-III and MIMIC-IV, the freely accessible critical care databases maintained through PhysioNet that have become the de facto benchmarks for clinical machine learning research. MKMed consistently outperformed the strongest baselines across multiple evaluation metrics. On MIMIC-III, the framework achieved improvements of 1.9 percent in Jaccard similarity, a metric that captures how well the recommended set of medications overlaps with what clinicians actually prescribed, and 1.3 percent in PRAUC, the area under the precision-recall curve, which is particularly informative when positive cases are rare. In a field where incremental gains of a fraction of a percent are often hard-won, consistent improvements of this magnitude across metrics and datasets represent a meaningful advance, and the gains were achieved specifically by mitigating the sparse-knowledge weakness that had constrained earlier systems.</p>
<p>The significance of this work extends beyond leaderboard numbers. Medication recommendation sits at the intersection of patient safety and clinical efficiency, and errors in this domain carry real consequences: drug-drug interactions, adverse reactions and inappropriate combinations for patients with complex comorbidities. A model whose representations are systematically weaker for poorly documented drugs risks being least reliable precisely for the patients who are hardest to treat, including those on rare or newly approved medications. By explicitly engineering robustness into sparse knowledge settings, MKMed points toward recommendation systems that are more dependable across the full breadth of the pharmacopoeia rather than only for its well-documented corners. The researchers suggest that cross-modal knowledge alignment of this kind could support more reliable and safer clinical decision-making in real-world healthcare scenarios, a claim that the benchmark results, at least, lend credible support.</p>
<p>The team has also made its work unusually accessible for follow-on research. The code associated with the study is publicly available in a GitHub repository, representative samples of the molecular multimodal pretraining data are published with instructions for acquisition, and the MIMIC-III and MIMIC-IV datasets themselves can be obtained through PhysioNet by researchers who complete the required credentialing process and data use agreements. This openness matters, because the bucket effect is unlikely to be solved by a single paper. As multimodal artificial intelligence spreads through medicine, from drug discovery to diagnostic imaging to treatment planning, the question of how to build coherent representations from incomplete, unevenly distributed knowledge will only grow more pressing. MKMed offers both a diagnosis of that problem, quantified with statistical rigor, and a working architectural answer, and it hands the research community the tools to test, extend and challenge that answer in the years ahead.</p>
<p><strong>Subject of Research:</strong> A multi-knowledge alignment framework for AI-based medication recommendation using electronic health records</p>
<p><strong>Article Title:</strong> MKMed: Multi-Knowledge alignment framework for medication recommendation</p>
<p><strong>Article References:</strong> Ma, H., Wu, G., Mu, S., Li, C., &amp; Liang, S. (2026). MKMed: Multi-Knowledge alignment framework for medication recommendation. <em>Applied Intelligence, 56</em>(15), Article 452. <a href="https://doi.org/10.1007/s10489-026-07337-4" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07337-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07337-4" rel="noopener noreferrer">10.1007/s10489-026-07337-4</a></p>
<p><strong>Keywords:</strong> medication recommendation, machine learning, electronic health records, contrastive learning, cross-modal alignment, molecular representation, knowledge graph, MIMIC-III, MIMIC-IV, clinical decision support, pharmacotherapy, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216349</post-id>	</item>
		<item>
		<title>AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds</title>
		<link>https://scienmag.com/ai-chatbots-match-human-doctors-in-recommending-laser-eye-surgery-study-finds/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 00:16:37 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI chatbots]]></category>
		<category><![CDATA[AI limitations in complex medical choices]]></category>
		<category><![CDATA[AI triage for eye procedures]]></category>
		<category><![CDATA[AI vs human ophthalmologists]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[BMC Medicine]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[Cost-effectiveness]]></category>
		<category><![CDATA[DeepSeek]]></category>
		<category><![CDATA[GPT-4o]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in healthcare]]></category>
		<category><![CDATA[laser eye surgery decision-making]]></category>
		<category><![CDATA[LASIK]]></category>
		<category><![CDATA[LASIK and SMILE surgical recommendations]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[medical AI performance comparison]]></category>
		<category><![CDATA[medical decision support AI]]></category>
		<category><![CDATA[ophthalmology]]></category>
		<category><![CDATA[ophthalmology AI diagnosis]]></category>
		<category><![CDATA[real-world AI validation in ophthalmology]]></category>
		<category><![CDATA[refractive surgery]]></category>
		<category><![CDATA[refractive surgery AI accuracy]]></category>
		<category><![CDATA[SMILE]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215569</guid>

					<description><![CDATA[A study of nearly 12,000 patients found that leading large language models matched or exceeded an intermediate ophthalmologist in recommending refractive surgery, though experts caution the AI should remain a decision-support tool.]]></description>
										<content:encoded><![CDATA[<p>A team of ophthalmologists in China has put some of the world&#8217;s most powerful artificial intelligence chatbots through one of the most demanding real-world tests yet devised for medical AI: deciding which type of refractive surgery, if any, suits a patient&#8217;s eyes. The results, published in BMC Medicine, show that five leading large language models could match or exceed the performance of an intermediate-level ophthalmologist when triaging patients for procedures such as LASIK and SMILE, reaching accuracies above 98.5 percent in some tasks. But the study also draws a careful line: the machines excel at straightforward yes-or-no judgments yet remain noticeably shakier when asked to choose among multiple surgical options, a nuance the researchers say should shape how clinics deploy these tools.</p>
<p>Refractive surgery is one of the most common elective procedures in medicine, encompassing techniques that reshape the cornea with lasers or implant a corrective lens inside the eye. Choosing the right procedure is far from trivial. Surgeons must weigh corneal thickness, refractive error, anterior chamber depth, pupil size, lifestyle factors and a catalogue of contraindications. A patient who is an ideal candidate for Small Incision Lenticule Extraction, or SMILE, might be a poor fit for transepithelial photorefractive keratectomy, known as TransPRK, and vice versa. Misjudging that calculus can lead to complications, retreatments or long-term visual problems, which is why surgical suitability decisions are traditionally reserved for trained specialists.</p>
<p>The study, led by Qi Wan, Ran Wei, Jing Tang, Ying-ping Deng and Ke Ma of the Department of Ophthalmology at West China Hospital of Sichuan University, drew on preoperative data from 11,966 consecutive patients evaluated at their institution. That sheer scale is what makes the work unusual. Most AI-in-medicine benchmarks rely on a few hundred curated cases; this one threw nearly twelve thousand messy, real-world clinical records at the algorithms. To establish ground truth, a panel of three senior refractive surgeons independently reviewed every case, assigning each a recommendation score from 0 to 100 and a suitability classification across four procedures: Femtosecond LASIK, SMILE, TransPRK and the Implantable Collamer Lens, or ICL. The panel&#8217;s agreement was excellent, with a kappa statistic exceeding 0.85, meaning the experts rarely disagreed about what the right answer was.</p>
<p>Five large language models faced the exam: DeepSeek-Chat, GLM-4.7, GPT-4o, Kimi-K2-Thinking and Qwen-Max. Each was fed the same structured, expert-mimicking prompt for every case and asked to produce the same recommendation scores and classifications the human panel had given. Alongside the machines, an intermediate-level physician, representing a mid-career ophthalmologist, independently evaluated all 11,966 cases, providing a human benchmark that is arguably more realistic than comparing AI to world-renowned professors. The study then scored everyone, human and machine alike, on a battery of standard metrics: accuracy, the area under the receiver operating characteristic curve for binary decisions, Cohen&#8217;s kappa for multi-class agreement, correlation coefficients for the numeric scores, and the root mean square and mean absolute errors measuring how far predictions strayed from expert judgment.</p>
<p>The headline result is that the top-performing language models were remarkably good at the binary question of whether a patient is suitable for a given procedure. For LASIK and SMILE, the best models achieved accuracies above 98.5 percent and areas under the curve exceeding 0.96, figures that place them in territory generally considered excellent for clinical classification. Perhaps more striking was the performance of the intermediate physician, who scored significantly lower than the leading models, particularly on complex classifications. That comparison matters because it reframes the debate about AI in medicine. The question is no longer simply whether machines can match elite specialists, but whether they can consistently outperform the average clinician making routine decisions, and on this evidence, in this narrow domain, they can.</p>
<p>The picture grows more complicated when the task shifts from two options to four. In the multi-class setting, where a model must pick the single best-suited procedure among LASIK, SMILE, TransPRK and ICL, agreement with the expert panel was more modest. Qwen-Max led this metric, achieving a Cohen&#8217;s kappa of up to 0.743, which statisticians generally label substantial agreement but which falls well short of the panel&#8217;s own internal consistency. The researchers interpret this gap candidly: the models are best understood as decision-support tools rather than autonomous decision-makers. In other words, an AI can reliably flag that a patient is a candidate for corneal refractive surgery, but a human expert should still make the final call about which procedure, especially in borderline cases where corneal topography, pupil characteristics and patient preference interact in subtle ways.</p>
<p>Beyond raw accuracy, the study examined how closely the models&#8217; numeric recommendation scores tracked the experts&#8217; 0-to-100 ratings. Here Qwen-Max and DeepSeek-Chat showed strong correlation with the panel&#8217;s scores, suggesting the models grasp not just the category of recommendation but its gradations, recognizing, say, that a patient with thin corneas and moderate myopia deserves a lower suitability score for LASIK than one with abundant corneal tissue. Regression errors captured the same story: the strongest models deviated from expert scores by margins small enough to be clinically useful, while weaker models and the intermediate physician showed wider scatter. This graded, score-like behavior matters for real deployment, because a tool that merely says yes or no discards the nuance surgeons use to counsel patients about risk.</p>
<p>One of the study&#8217;s most practically significant contributions is its analysis of cost and speed, dimensions rarely quantified in medical AI evaluations. DeepSeek-Chat emerged as the efficiency champion, combining the lowest cost per query with the fastest response times while still delivering strong accuracy. GPT-4o, by contrast, was the most expensive of the five. These economics are not trivia. In high-volume screening scenarios, where thousands of preoperative assessments must be triaged before a surgeon ever sees the patient, a model that is nearly as accurate but a fraction of the price can transform workflow economics. The authors point to resource-limited settings, including regions with few refractive surgeons, as the clearest beneficiaries, since a cheap, fast, accurate pre-screening layer could extend specialist-grade triage to populations that currently lack access.</p>
<p>Technically, the study also offers a template for how such evaluations should be run. The structured expert-mimicking prompt, the massive consecutive-patient dataset, the multi-surgeon gold standard with quantified inter-rater agreement, and the multi-dimensional metric battery together form a benchmarking blueprint that other specialties can copy. Too many published AI evaluations rest on convenience samples and single-metric reporting; this work demonstrates what a more rigorous standard looks like, including the honest acknowledgment that kappa values in multi-class tasks remain moderate. It is worth noting that the study is retrospective: the models reviewed recorded data rather than live patients, and real clinical deployment would raise additional questions about data privacy, liability, and how surgeons integrate algorithmic advice into consultations.</p>
<p>What emerges is a measured but genuinely exciting picture of where medical language models stand. On narrow, well-defined classification tasks grounded in structured clinical data, they now perform at or above the level of mid-career physicians, at pennies per case and in seconds. On the harder judgment calls that define expert practice, they still lag behind senior specialists and need human oversight. The West China Hospital team frames the technology exactly as the evidence supports: a powerful assistive layer for screening, triage and decision support, positioned to augment rather than replace the surgeon&#8217;s judgment. As these models continue to improve, and as prospective trials validate retrospective results, the routine preoperative assessment of refractive surgery candidates may become one of the first places where patients routinely benefit from an AI second opinion, whether or not they ever know it is there.</p>
<p><strong>Subject of Research:</strong> Evaluation of large language models for refractive surgery recommendation and clinical decision support</p>
<p><strong>Article Title:</strong> Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation</p>
<p><strong>Article References:</strong> Wan, Q., Wei, R., Tang, J., Deng, Y.-P., &amp; Ma, K. (2026). Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation. <em>BMC Medicine</em>. <a href="https://doi.org/10.1186/s12916-026-05262-4" rel="noopener noreferrer">https://doi.org/10.1186/s12916-026-05262-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12916-026-05262-4" rel="noopener noreferrer">10.1186/s12916-026-05262-4</a></p>
<p><strong>Keywords:</strong> large language models, refractive surgery, LASIK, SMILE, ophthalmology, artificial intelligence, clinical decision support, GPT-4o, DeepSeek, cost-effectiveness, machine learning, BMC Medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215569</post-id>	</item>
		<item>
		<title>Agentic AI in Medicine Needs Clinical Testing Harnesses, Researchers Warn</title>
		<link>https://scienmag.com/agentic-ai-in-medicine-needs-clinical-testing-harnesses-researchers-warn/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 23:15:37 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[agentic AI]]></category>
		<category><![CDATA[AI governance]]></category>
		<category><![CDATA[AI safety]]></category>
		<category><![CDATA[AI safety and reliability in clinical settings]]></category>
		<category><![CDATA[AI system scaffolding and external tool integration]]></category>
		<category><![CDATA[AI-driven medical decision-making]]></category>
		<category><![CDATA[automation bias]]></category>
		<category><![CDATA[autonomous artificial intelligence in healthcare]]></category>
		<category><![CDATA[biomedical engineering]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[clinical testing]]></category>
		<category><![CDATA[clinical validation of agentic AI systems]]></category>
		<category><![CDATA[development of clinical testing harness for AI]]></category>
		<category><![CDATA[digital health]]></category>
		<category><![CDATA[ethical considerations of autonomous decision-support]]></category>
		<category><![CDATA[future of AI-enabled autonomous patient management]]></category>
		<category><![CDATA[human-in-the-loop]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[medical AI validation]]></category>
		<category><![CDATA[multi-step AI task execution in medicine]]></category>
		<category><![CDATA[regulatory implications of AI with planning capabilities]]></category>
		<category><![CDATA[risks of autonomous AI in patient care]]></category>
		<category><![CDATA[trustworthy AI]]></category>
		<category><![CDATA[validation challenges for agentic AI in hospitals]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215220</guid>

					<description><![CDATA[Researchers argue that multi-step agentic AI systems break traditional clinical validation models and propose structured testing harnesses adapted from simulation, credentialing, and morbidity review.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence in hospitals has long followed a familiar pattern: a model produces a single output, such as a risk score or a diagnostic suggestion, and a clinician reviews it before anything happens to the patient. A new letter in the Annals of Biomedical Engineering argues that this entire evaluation framework is about to be upended. Ethan Waisberg and Joseph W. Guarnieri, affiliated with the University of Cambridge, the Blue Marble Space Institute of Science, and the Guarnieri Research Group, contend that agentic artificial intelligence systems, which can plan and execute multi-step tasks on their own, break the assumptions underlying how medicine has traditionally validated decision-support technology. Their proposed response is a concept they call the clinical testing harness, and it may determine whether autonomous AI earns a place at the bedside or remains confined to the laboratory.</p>
<p>The core of the problem lies in what the authors describe as agent scaffolding. This is the software infrastructure wrapped around a large language model that transforms passive text generation into active, multi-step execution. With scaffolding in place, an AI system does not merely answer a question. It can retrieve patient records, invoke external tools such as calculators, databases, or imaging analyzers, interpret the results of those tool calls, and then take further actions based on what it finds, all before a human clinician sees anything at all. The relevant unit of clinical risk, the authors argue, therefore shifts from a single inference to an entire trajectory of actions. A model might compose a flawless summary of a chest radiograph and still fail catastrophically if it queries the wrong data source, misconfigures a drug-dosing tool, or chains a series of individually reasonable steps into a collectively dangerous plan.</p>
<p>Waisberg and Guarnieri organize the resulting challenges into four categories, each of which exposes a weakness in current approaches to responsible AI development in medicine. The first is silent error propagation across multi-step tasks. In a single-output system, an error appears in the one thing the clinician reviews. In an agentic system, a mistake made in step two, say a misparsed laboratory value, can silently contaminate steps three through ten, and by the time a conclusion reaches human review, it may look polished and internally consistent even though it rests on corrupted intermediate work. The error is not merely hidden; it is laundered through subsequent stages of reasoning until it becomes difficult to detect.</p>
<p>The second challenge concerns oversight that reviews conclusions rather than processes. Traditional clinical decision support places a single recommendation in front of a clinician, who evaluates it against their own expertise. But agentic systems present the end product of a long chain of retrievals, tool invocations, and intermediate judgments. A clinician who signs off on the final answer has effectively approved a process they never observed. The authors point out that this compounds a well-documented human tendency: automation bias, the inclination to accept machine-generated recommendations without sufficient scrutiny, has been recognized in the medical informatics literature for over a decade. When the machine does far more work invisibly, the bias has more room to operate.</p>
<p>The third challenge is validation that does not survive changes to the model or the tooling. In conventional medical AI, a model is trained, validated on a fixed dataset, and then deployed with its behavior frozen. Regulatory pathways and clinical trust both rest on the assumption that what was tested is what runs. Agentic systems violate this assumption structurally. A developer might swap the underlying language model for a newer version, update a connected tool, or modify the scaffolding logic, and the behavior of the entire system can change in ways that invalidate prior testing. The authors note that medicine has already lived through a cautionary version of this problem with static models: the widely implemented Epic sepsis prediction model, when externally validated by independent researchers, performed far worse than its developers reported, missing the majority of sepsis cases while generating a flood of false alarms. If a frozen, single-purpose model can degrade that badly under real-world conditions, the authors imply, a self-directed system whose components constantly evolve poses a categorically harder validation problem.</p>
<p>The fourth challenge is scope that widens faster than the evidence supporting it. Agentic AI is generality by design. The same scaffolding that lets a system answer a question about diabetic retinopathy can, in principle, let it draft referral letters, order tests, adjust medication lists, or triage messages, and the temptation in deployment is to keep expanding what the agent is allowed to do. But each expansion of scope requires its own evidence base, and the pace of capability demonstration is outrunning the pace of clinical validation. The letter cites recent work including MedAgentBench, a virtual electronic health record environment for benchmarking medical LLM agents, and AgentClinic, a multimodal benchmark for tool-using clinical AI agents, alongside a validated autonomous oncology decision-making agent, as signs of a field whose technical ambitions have sprinted ahead of its safety infrastructure.</p>
<p>The authors&#8217; solution is not a new invention but an adaptation of tools medicine already trusts. They propose the clinical testing harness, a structured evaluation environment with four components. The first is a scenario library built from clinical edge cases, the rare, ambiguous, and dangerous presentations where decision support is most likely to fail and most likely to matter. The second is full-trajectory observability, meaning that evaluators can inspect not just the final output but every retrieval, tool call, and intermediate decision the agent made along the way. This directly answers the silent propagation problem: if every step is logged and examinable, an error introduced in step two can be traced before it metastasizes into a confident, polished, and wrong conclusion.</p>
<p>The third component is explicit escalation testing, which evaluates whether the agent recognizes the limits of its own competence. A safe clinical agent must not only answer questions correctly; it must know when a situation demands a human, and must reliably hand control to a clinician when uncertainty is high, the case is atypical, or the stakes are severe. The fourth component is staged evidence thresholds tied to scope of practice. Instead of a binary approval, an agent would earn capabilities incrementally: broad evidence might support letting it retrieve and summarize information, while higher-stakes actions such as influencing treatment decisions would remain gated until the system had demonstrated reliability at each preceding level. This mirrors how medicine itself grants privileges, and it prevents the scope-expansion problem by making each widening of responsibility contingent on demonstrated performance.</p>
<p>The elegance of the proposal, and arguably its feasibility, comes from its institutional ancestry. Waisberg and Guarnieri observe that medicine already possesses every ingredient of a testing harness in mature form. Simulation-based training subjects clinicians to edge cases before they touch patients. Credentialing and scope-of-practice rules tie permitted actions to demonstrated competence. Morbidity and mortality conferences systematically dissect adverse outcomes to find process failures rather than merely blaming individuals. The clinical testing harness is these institutions translated into software: scenario libraries are simulation, staged thresholds are credentialing, and full-trajectory review is the M and M conference applied to machine reasoning. The letter, published on 23 September 2026 and reviewed under Associate Editor Joel Stitzel, reports no funding and no conflicts of interest, and draws on the authors&#8217; prior work on large language models in medical imaging, concerns about ChatGPT in academia and medicine, and the problem of correlated failure in delegated AI supervision.</p>
<p>Whether the vision takes hold will depend on regulators, developers, and health systems accepting an uncomfortable premise: that the familiar benchmark scores of a language model tell clinicians almost nothing about the safety of an agent built around it. The past decade of clinical AI offers a sobering precedent, from sepsis models that faltered under external validation to the recognition that human oversight is neither automatic nor reliably vigilant. Agentic systems promise genuine gains, potentially extending expert-level reasoning into settings where specialists are scarce, but only if the unit of testing expands to match the unit of risk. A trajectory of actions demands a trajectory of evidence, and the letter&#8217;s central claim is that medicine need not invent a new safety culture to provide it. It need only point the tools it already trusts at a new kind of clinician, one that never tires, never second-guesses, and acts before anyone is watching.</p>
<p><strong>Subject of Research:</strong> Responsible development and clinical evaluation of agentic artificial intelligence systems in medicine</p>
<p><strong>Article Title:</strong> Agentic AI in Medicine: Challenges for Responsible Development and the Case for Clinical Testing Harnesses</p>
<p><strong>Article References:</strong> Waisberg, E., &amp; Guarnieri, J. W. (2026). Agentic AI in Medicine: Challenges for Responsible Development and the Case for Clinical Testing Harnesses. <em>Annals of Biomedical Engineering</em>. <a href="https://doi.org/10.1007/s10439-026-04381-6" rel="noopener noreferrer">https://doi.org/10.1007/s10439-026-04381-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10439-026-04381-6" rel="noopener noreferrer">10.1007/s10439-026-04381-6</a></p>
<p><strong>Keywords:</strong> agentic AI, clinical decision support, AI safety, large language models, clinical testing, medical AI validation, automation bias, digital health, AI governance, human-in-the-loop, trustworthy AI, biomedical engineering</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215220</post-id>	</item>
	</channel>
</rss>
