<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI language models in healthcare &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-language-models-in-healthcare/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 07 May 2026 21:48:33 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI language models in healthcare &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Study Reveals AI Language Models Encounter Challenges with Basic Hospital Data Tasks</title>
		<link>https://scienmag.com/study-reveals-ai-language-models-encounter-challenges-with-basic-hospital-data-tasks/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 07 May 2026 21:48:33 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[administrative data tasks in hospitals]]></category>
		<category><![CDATA[AI for patient load monitoring]]></category>
		<category><![CDATA[AI language models in healthcare]]></category>
		<category><![CDATA[AI performance on hospital resource allocation]]></category>
		<category><![CDATA[challenges of AI in hospital administration]]></category>
		<category><![CDATA[democratizing data access in healthcare]]></category>
		<category><![CDATA[EHR data querying by non-technical staff]]></category>
		<category><![CDATA[GPT-4o and Llama in medical data analysis]]></category>
		<category><![CDATA[healthcare operational reporting with AI]]></category>
		<category><![CDATA[large language models for electronic health records]]></category>
		<category><![CDATA[limitations of LLMs in clinical workflows]]></category>
		<category><![CDATA[natural language processing for healthcare data]]></category>
		<guid isPermaLink="false">https://scienmag.com/study-reveals-ai-language-models-encounter-challenges-with-basic-hospital-data-tasks/</guid>

					<description><![CDATA[A recent investigation into the practical capabilities of large language models (LLMs) reveals significant limitations in their use for routine administrative tasks within hospital environments. Conducted by Eyal Klang and colleagues at the Icahn School of Medicine at Mount Sinai in New York, the study critically evaluates the performance of state-of-the-art LLMs on essential number-crunching [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>A recent investigation into the practical capabilities of large language models (LLMs) reveals significant limitations in their use for routine administrative tasks within hospital environments. Conducted by Eyal Klang and colleagues at the Icahn School of Medicine at Mount Sinai in New York, the study critically evaluates the performance of state-of-the-art LLMs on essential number-crunching operations that healthcare administrators depend on daily. Published in PLOS Digital Health, these findings provide essential technical insights into the challenges facing AI implementation in clinical administrative workflows.</p>
<p>Hospitals nowadays rely heavily on electronic health records (EHRs) – structured datasets that capture patient information, resource availability, and care events. Administrators utilize this data to monitor patient loads, allocate resources, and generate operational reports. Traditionally, these tasks are performed by specialized data analysts deploying programming languages and database queries, a process often fraught with delays when rapid answers are needed for decision-making. The promise of LLMs like GPT-4o and Llama has been to democratize data access by allowing non-technical staff to query these datasets directly using natural language prompts.</p>
<p>In the study, researchers subjected nine leading LLMs to a rigorous battery of tests designed to emulate two foundational administrative functions: counting how many patients meet a specific clinical condition and filtering records based on multiple inclusion criteria simultaneously. The data itself was sourced from a substantial real-world dataset of over 50,000 emergency department visits within the Mount Sinai Health System, grounding the evaluation in practical, messy clinical data rather than synthetic or simplified examples.</p>
<p>The initial experiments employed straightforward prompting techniques, where models were simply asked direct questions such as “How many patients were admitted from this table?” Across the board, all tested LLMs demonstrated subpar accuracy, failing to provide reliable answers when handling these structured queries. This underlines a fundamental disconnect between LLM training—and their practical applicability to real-world numerical and logical operations in healthcare datasets.</p>
<p>To enhance performance, the researchers explored a chain-of-thought prompting approach. This method instructs the model to transparently reason through the problem step-by-step before arriving at the final answer, theoretically enabling more accurate and consistent outputs. However, the results were underwhelming; only modest improvements were observed on smaller tables, and as the size and complexity of the data increased, accuracy declined precipitously. For instance, even GPT-4o, the best performing model under this regime, saw accuracy plummet from approximately 95% on small datasets to below 60% when confronted with larger tables.</p>
<p>Recognizing that prompting alone may not suffice, the research shifted focus to a tool-based model execution approach. Here, LLMs were tasked with generating executable code, such as SQL or Python scripts, to process the data programmatically. This method leverages the LLM’s natural language understanding to translate queries into precise machine-readable commands, which are then run directly against the EHR data for guaranteed accuracy. Impressively, this approach substantially improved results for the most advanced models. GPT-4o and Qwen-2.5-72B demonstrated near-perfect accuracy under these conditions, successfully navigating the intricacies of complex filters and large datasets.</p>
<p>Despite these successes, not all models fared well. LLMs optimized for speed and efficiency, such as distilled variants of DeepSeek, struggled to produce usable outputs even when provided with the ability to generate and run code. Furthermore, the Llama-3.1-8B model encountered major difficulties, failing to produce functional results in the majority of assessments and being ultimately excluded from further analysis. These discrepancies highlight the diverse capabilities within the current LLM ecosystem and caution against broad assumptions regarding their utility in structured data environments.</p>
<p>The study’s findings carry critical implications for the future deployment of LLMs in healthcare administration. Benjamin Glicksberg, one of the authors, emphasized that without integrating tool-based strategies—combining LLM-generated code with actual execution—large language models remain fundamentally unsuitable for standalone use in clinical administrative settings. Clinical workflows frequently involve complex structured data requiring absolute reliability and precision, conditions under which straightforward natural language query processing by LLMs falls short.</p>
<p>Moreover, the requirement for “agentic” approaches is underscored by this work. Agentic AI involves systems that act semi-autonomously, leveraging external tools and code execution capabilities to ensure results remain consistent and verifiable. By integrating LLMs with backend code execution engines, hospitals could dramatically accelerate administrative processes while maintaining data integrity. Such hybrid solutions may bridge the gap between cutting-edge AI capabilities and the stringent accuracy demands of healthcare operations.</p>
<p>This study shines a spotlight on the often-overlooked challenges of applying AI in clinical data environments. While the hype around LLMs centers on their conversational fluency and general knowledge, the ability to perform precise numerical computations and filtered data retrieval within complex EHR systems requires a fundamentally different kind of model reliability. The researchers’ meticulous experimental design and real-world data usage offer a vital reality check for the healthcare sector’s ongoing AI ambitions.</p>
<p>Lastly, the authors note that their work did not receive any external funding, and no competing interests were declared. The open-access publication ensures that the full details, along with extensive methodological descriptions and results, remain available to researchers, clinicians, and AI developers aiming to advance safe and effective AI integration into hospital administration.</p>
<p>Overall, these findings caution healthcare providers and AI developers alike to calibrate expectations around LLMs’ current abilities in administrative contexts. They also highlight the powerful potential unlocked by hybrid human-AI systems that combine natural language understanding with robust programming and execution frameworks. As digital healthcare continues to evolve, researchers and practitioners will need to navigate these complex trade-offs to harness AI’s benefits without compromising accuracy and trustworthiness.</p>
<p><strong>Web References:</strong><br />
<a href="https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0001326">https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0001326</a></p>
<hr />
<p><strong>Subject of Research</strong>: Not applicable<br />
<strong>Article Title</strong>: Large language models are poor clinical administrators: An evaluation of structured queries in real-world electronic health records<br />
<strong>News Publication Date</strong>: 7-May-2026<br />
<strong>References</strong>: Klang E, Sorin V, Korfiatis P, Sawant AS, Freeman R, Charney AW, et al. (2026) Large language models are poor clinical administrators: An evaluation of structured queries in real-world electronic health records. PLOS Digit Health 5(5): e0001326. DOI: 10.1371/journal.pdig.0001326</p>
<h4><strong>Keywords</strong></h4>
<p>Large Language Models, Electronic Health Records, Clinical Administration, Artificial Intelligence, GPT-4o, Tool-based AI, Chain-of-Thought Prompting, Healthcare Data Analytics, AI Reliability, Code Generation, Hospital Resource Management, Clinical Workflow Automation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">157482</post-id>	</item>
		<item>
		<title>WVU Researchers Explore the Boundaries of AI in Emergency Room Diagnoses</title>
		<link>https://scienmag.com/wvu-researchers-explore-the-boundaries-of-ai-in-emergency-room-diagnoses/</link>
		
		<dc:creator><![CDATA[Reid Dalton]]></dc:creator>
		<pubDate>Tue, 20 May 2025 20:17:11 +0000</pubDate>
				<category><![CDATA[Mathematics]]></category>
		<category><![CDATA[AI in emergency room diagnostics]]></category>
		<category><![CDATA[AI language models in healthcare]]></category>
		<category><![CDATA[ChatGPT performance evaluation]]></category>
		<category><![CDATA[clinical decision-making with AI]]></category>
		<category><![CDATA[de-identified physician notes study]]></category>
		<category><![CDATA[diagnostic accuracy using AI]]></category>
		<category><![CDATA[emergency department AI applications]]></category>
		<category><![CDATA[enhancing diagnostic tools with AI]]></category>
		<category><![CDATA[limitations of AI in medical diagnoses]]></category>
		<category><![CDATA[real-world clinical data analysis]]></category>
		<category><![CDATA[symptom presentation challenges in AI]]></category>
		<category><![CDATA[WVU research on AI healthcare]]></category>
		<guid isPermaLink="false">https://scienmag.com/wvu-researchers-explore-the-boundaries-of-ai-in-emergency-room-diagnoses/</guid>

					<description><![CDATA[Artificial intelligence (AI) technologies have found a burgeoning role in modern healthcare, promising enhancements in diagnostic accuracy and clinical decision-making. Recent research from West Virginia University (WVU) propels this promise into the emergency department setting, where rapid and precise diagnosis is critical yet often challenging. WVU scientists, led by Gangqing “Michael” Hu, assistant professor at [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence (AI) technologies have found a burgeoning role in modern healthcare, promising enhancements in diagnostic accuracy and clinical decision-making. Recent research from West Virginia University (WVU) propels this promise into the emergency department setting, where rapid and precise diagnosis is critical yet often challenging. WVU scientists, led by Gangqing “Michael” Hu, assistant professor at the WVU School of Medicine, have conducted a pioneering evaluation of multiple iterations of ChatGPT, a state-of-the-art AI language model, assessing its performance in diagnosing emergency department patients based on physicians’ clinical notes. Their findings, published in <em>Scientific Reports</em>, underscore both the potential and current limitations of AI in emergency diagnostics, particularly in the context of symptom presentation.</p>
<p>The core objective of Hu’s study was to interrogate how different versions of ChatGPT handle diagnostic tasks given real-world clinical data. Using de-identified physician notes from 30 emergency department cases, the research team prompted various ChatGPT model iterations—including GPT-3.5, GPT-4, GPT-4o, and the o1 series—to generate their top three diagnostic suggestions. The study’s methodological rigor involved comparing the models&#8217; diagnostic precision and accuracy against actual clinical outcomes to draw a comprehensive performance profile. This approach provides a window into how AI tools can supplement, but not yet replace, human clinical judgment.</p>
<p>One of the profound insights emerging from this investigation is the discrepancy in AI performance between cases with classic, textbook symptoms and those with atypical or “challenging” presentations. For patients exhibiting hallmark signs of disease, ChatGPT models demonstrated promising diagnostic assistance capabilities, supporting physicians by suggesting accurate differential diagnoses. However, when confronted with complex cases lacking traditional symptomatic cues—such as pneumonia cases without accompanying fever—AI’s capacity to correctly identify diagnoses notably diminished. These failures illuminate the inherent difficulty AI models face when operating beyond their training data’s typical patterns, emphasizing the necessity for richer, more diverse datasets.</p>
<p>The researchers note that current AI diagnostic models primarily ingest unstructured text input—in this case, physicians’ notes—without access to multimodal clinical information. Consequently, ChatGPT’s diagnostic reasoning is limited by the breadth and variability of its textual training corpora and the information provided. Hu posits that enhancing future AI frameworks with additional clinical data streams—such as imaging results, laboratory findings, and comprehensive patient histories—could improve the fidelity and robustness of AI-assisted diagnoses in emergency contexts. Integration of these heterogeneous data types would transform AI from a purely linguistic interpreter to a more holistic clinical decision support system.</p>
<p>Analysis of the longitudinal performance of ChatGPT iterations reveals an interesting but cautious trajectory of improvement. While no statistically significant advance was observed when considering the inclusion of AI-generated diagnoses within the top three suggestions, the accuracy of the very top, or primary, diagnosis recommendation improved by approximately 15 to 20 percent in newer models relative to their predecessors. This subtle enhancement suggests iterative refinement in model capabilities but also highlights the persistent challenges in achieving consistently high precision necessary for clinical reliability.</p>
<p>The study underscores a key principle in the deployment of AI-assisted diagnostic tools: the indispensability of human oversight. Given the models’ current inadequate performance on complex cases, physician expertise remains essential to interpret AI outputs critically and corroborate or refute AI-generated hypotheses. This interplay forms a hybrid intelligence paradigm, wherein AI accelerates data synthesis and hypothesis generation while clinicians provide contextual judgment, ensuring that patient care remains both accurate and personalized.</p>
<p>Beyond diagnostic accuracy, Hu envisions AI modalities evolving towards greater transparency and explicability. He stresses the importance of AI systems that do not merely generate results but also reveal their reasoning pathways, enabling clinicians to understand and trust their recommendations. Such “explainable AI” is critical to fostering confidence among healthcare providers, enhancing AI’s integration into clinical workflows, and ultimately improving patient outcomes. Achieving this level of transparency will require methodological innovations in how AI models represent and communicate uncertainty and rationale.</p>
<p>Moreover, Hu’s research team explores imaginative avenues to augment diagnostic reasoning by leveraging multi-agent AI simulations. Drawing on prior work where ChatGPT-4 was deployed in role-playing scenarios—emulating specialists such as physiotherapists, psychologists, and nutritionists engaged in panel discussions—this approach aims to replicate the collaborative diagnostic processes typical in clinical environments. The proposed conversational model suggests that dynamic interactions among diverse AI agents could produce more nuanced, accurate diagnostic assessments, reflecting interdisciplinary integration akin to human medical teams.</p>
<p>Despite these promising strides, the researchers caution that current AI systems, including ChatGPT, do not qualify as certified medical devices and should not be used as standalone diagnostic solutions. In clinical settings where expanded data types, such as imaging, are incorporated, AI models must operate within secure, privacy-compliant hospital clusters as open-source platforms. Compliance with regulatory standards and patient confidentiality laws remains a non-negotiable prerequisite for AI deployment in healthcare institutions.</p>
<p>The study acknowledges support from the National Science Foundation and the National Institutes of Health, emphasizing the significance of federally-funded research in advancing AI applications in medicine. Additional contributors include postdoctoral fellow Jinge Wang, lab volunteer Kenneth Shue, and Li Liu from Arizona State University, reflecting a multidisciplinary collaboration essential to tackling complex problems at the intersection of computer science, bioinformatics, and clinical medicine.</p>
<p>Looking ahead, Hu advocates for future research to focus not only on enhancing AI’s diagnostic performance but also on its capacity to articulate reasoning in clinically meaningful ways. He suggests that improved explainability could facilitate critical emergency department decisions such as triage prioritization and treatment pathway selection, augmenting both efficiency and patient safety.</p>
<p>In summary, the pioneering evaluation of ChatGPT models in emergency diagnostics performed by WVU scientists reveals a nuanced landscape marked by AI’s emerging utility balanced against intrinsic challenges. While encouraging diagnostic accuracy for prototypical cases validates the promise of language models as assistive tools, persistent deficiencies in recognizing atypical disease presentations underscore the imperative for richer data integration, transparent reasoning, and robust human-AI collaboration. This research not only advances scientific understanding of AI capabilities at the clinical frontline but also charts a thoughtful course towards responsible integration of AI in patient-centered care.</p>
<hr />
<p><strong>Subject of Research</strong>:<br />
Evaluation of ChatGPT AI model iterations for diagnostic assistance in emergency department patients using clinical notes.</p>
<p><strong>Article Title</strong>:<br />
Preliminary evaluation of ChatGPT model iterations in emergency department diagnostics</p>
<p><strong>News Publication Date</strong>:<br />
26-Mar-2025</p>
<p><strong>Web References</strong>:  </p>
<ul>
<li><a href="https://www.wvu.edu/">https://www.wvu.edu/</a>  </li>
<li><a href="https://directory.hsc.wvu.edu/Profile/60888">https://directory.hsc.wvu.edu/Profile/60888</a>  </li>
<li><a href="https://medicine.wvu.edu/">https://medicine.wvu.edu/</a>  </li>
<li><a href="https://medicine.wvu.edu/micro/">https://medicine.wvu.edu/micro/</a>  </li>
<li><a href="https://health.wvu.edu/research-and-graduate-education/research/core-facilities/bioinformatics-core/">https://health.wvu.edu/research-and-graduate-education/research/core-facilities/bioinformatics-core/</a>  </li>
<li><a href="https://www.nature.com/articles/s41598-025-95233-1#citeas">https://www.nature.com/articles/s41598-025-95233-1#citeas</a>  </li>
<li><a href="http://dx.doi.org/10.1038/s41598-025-95233-1">http://dx.doi.org/10.1038/s41598-025-95233-1</a>  </li>
<li><a href="https://mededu.jmir.org/2024/1/e51157/">https://mededu.jmir.org/2024/1/e51157/</a></li>
</ul>
<p><strong>References</strong>:<br />
Hu, G. M., Wang, J., Shue, K., Liu, L. (2025). Preliminary evaluation of ChatGPT model iterations in emergency department diagnostics. <em>Scientific Reports</em>. DOI: 10.1038/s41598-025-95233-1</p>
<p><strong>Image Credits</strong>:<br />
WVU Photo/Greg Ellis</p>
<p><strong>Keywords</strong>:<br />
Artificial intelligence, Disease prevention, Clinical medicine, Medical tests, Artificial consciousness, Artificial neural networks, Cognitive robotics, Forward chaining, Generative AI, Genetic algorithms, Logic based AI, Adaptive systems, Cybernetics, Robotics, Computer science</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">46603</post-id>	</item>
	</channel>
</rss>
