<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>limitations of AI in medicine &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/limitations-of-ai-in-medicine/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 28 May 2026 18:15:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>limitations of AI in medicine &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Doctor GPT: AI Achieves Nearly 76% Accuracy in Answering Healthcare Queries</title>
		<link>https://scienmag.com/doctor-gpt-ai-achieves-nearly-76-accuracy-in-answering-healthcare-queries/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 28 May 2026 18:15:25 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI and patient information accuracy]]></category>
		<category><![CDATA[AI chatbot accuracy in health queries]]></category>
		<category><![CDATA[AI in healthcare]]></category>
		<category><![CDATA[AI medical advice safety]]></category>
		<category><![CDATA[consumer-focused AI health tools]]></category>
		<category><![CDATA[healthcare AI evaluation study]]></category>
		<category><![CDATA[large language models for medical advice]]></category>
		<category><![CDATA[limitations of AI in medicine]]></category>
		<category><![CDATA[Penn State Diagnose-a-thon event]]></category>
		<category><![CDATA[public engagement with health AI]]></category>
		<category><![CDATA[real-world AI healthcare applications]]></category>
		<category><![CDATA[symptom checking with AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/doctor-gpt-ai-achieves-nearly-76-accuracy-in-answering-healthcare-queries/</guid>

					<description><![CDATA[In recent years, the rise of artificial intelligence (AI) technologies, particularly large language models (LLMs), has opened new frontiers in multiple fields, including healthcare. A groundbreaking study led by researchers at Penn State has now provided a rigorous evaluation of how AI-powered chatbots respond to everyday health-related inquiries posed by the general public. The study [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In recent years, the rise of artificial intelligence (AI) technologies, particularly large language models (LLMs), has opened new frontiers in multiple fields, including healthcare. A groundbreaking study led by researchers at Penn State has now provided a rigorous evaluation of how AI-powered chatbots respond to everyday health-related inquiries posed by the general public. The study uncovers that these AI systems achieve an accuracy rate of approximately 76% when addressing routine health questions, a figure that simultaneously highlights both the promise and the perils of deploying such technologies in real-world medical contexts.</p>
<p>The research uniquely focusses on the perspective of the average internet user, a group that frequently turns to AI as a modern-day symptom checker, reminiscent of how Google was traditionally used for preliminary health information. This user-centered approach is critical because prior studies predominantly examined LLMs from expert or academic lenses, often overlooking practical consumer interactions. By focusing on typical health queries submitted by laypersons, the study offers vital insights into the effectiveness and safety of AI-based medical advice in daily life.</p>
<p>To gather authentic data reflecting real-world usage, the research team organized an innovative event known as the &#8220;Diagnose-a-thon&#8221; at Penn State. This competition attracted 34 participants spanning faculty, staff, and students across various academic levels. Participants generated a substantial dataset of 212 health-related prompts, encompassing both genuine and hypothetical conditions, crafted from patient and clinician viewpoints. They then queried four distinct state-of-the-art LLMs: ChatGPT-4o, ChatGPT-3.5, Gemini-1.5 Pro, and Llama3-8b. By allowing participants to select their preferred AI model without constraints, the study faithfully replicated the autonomous and diverse usage patterns found in natural settings.</p>
<p>An essential part of the study involved a rigorous evaluation stage where nine board-certified physicians assessed the treatments and information handed back by the LLMs. The evaluation metric was comprehensive, assessing both the clinical accuracy and the potential harm posed by the AI-generated answers, measured on a nuanced six-point scale from very low to very high. This detailed scoring system illuminated how AI diagnostic responses vary across medical specialties and contexts, a level of granularity rarely seen in previous AI investigations.</p>
<p>The findings showed a variable performance landscape across medical disciplines. Obstetrics, gynecology, and otolaryngology yielded the highest levels of correct information with minimal risks, showcasing scenarios where LLMs currently excel. Conversely, fields such as internal medicine, neurology, and dermatology demonstrated more significant challenges for AI systems, where inaccuracies and higher harm potentials were more prevalent. These results underscore an important reality: certain specialized medical domains demand more caution when leveraging AI tools, especially if these tools are employed by untrained individuals.</p>
<p>A fascinating specificity in the study revealed that prompts with a length between 60 and 250 characters tended to produce more accurate AI responses. This suggests that message framing and prompt articulation play crucial roles in steering AI models toward clinically valid outputs. Moreover, highly specialized or narrowly focused questions posed difficulties, suggesting that broad generalist models still face significant hurdles when addressing deeply technical or nuanced medical issues.</p>
<p>Beyond evaluating off-the-shelf AI models, the research team experimented with a novel augmentation approach by retraining the base LLMs using an extensive corpus of medical textbooks, clinical guidelines, and peer-reviewed literature typical of medical school curricula. The goal was to determine whether such domain-specific tuning could enhance clinical validity while reducing harmful outputs. Surprisingly, medical professionals and trainees reviewing these augmented models showed a preference for responses from the original Gemini and Llama bases over the retrained versions. No statistically significant preference was observed regarding ChatGPT’s base versus augmented models. This counterintuitive result suggests that current fine-tuning strategies may not straightforwardly translate into improved clinical communication by AI.</p>
<p>The implications of these findings are profound for the future integration of AI into healthcare delivery. As Dr. Jennifer Kraschnewski, a co-author of the study and a practicing physician, articulates, AI represents a transformative force with the potential to augment clinician capabilities rather than replace human doctors. The challenge lies in harnessing AI tools in ways that bolster medical professionals’ diagnostic processes, reduce cognitive burdens, and improve patient outcomes without exposing patients to the risks of AI errors in unsupervised contexts.</p>
<p>Crucially, the study emphasizes that despite satisfactory accuracy scores in the mid-70s percentage range, the AI models still exhibited an error rate exceeding 20%. This rate is approximately double that of human physicians and highlights the potential for AI to propagate misinformation leading to harm if used uncritically by patients themselves. Such statistical insights counsel for cautious and responsible deployment of AI technologies in healthcare, underscoring the necessity of preserving human clinical oversight.</p>
<p>The study also offers a nuanced view on AI’s evolving role: rather than supplanting the physician’s role, AI could serve as a catalyst to &#8220;upskill&#8221; clinicians by providing rapid evidence summaries, differential diagnosis suggestions, and decision support, streamlining care processes. The research community is thus encouraged to focus on developing AI systems tailored to professional use, with interfaces and interpretability tuned for clinical environments.</p>
<p>Penn State’s research ecosystem facilitated this multidisciplinary collaboration, bringing together expertise in informatics, intelligent systems, clinical medicine, and AI ethics. Their participatory research design, which mimics user autonomy and real-world interaction dynamics, sets a new methodological standard for evaluating AI systems in societally critical domains. It also expands the discourse on AI accountability and transparency by highlighting the tangible benefits and limitations observed when AI systems engage with health-related content.</p>
<p>Given the inevitable persistence of AI tools in healthcare, public education and digital literacy emerge as pivotal. The study’s co-authors advocate for initiatives that enhance consumer understanding of AI’s strengths and weaknesses in medical diagnosis. Such literacy efforts will empower users to critically appraise AI-generated advice, reducing overreliance and potential misuses.</p>
<p>In summary, this Penn State study, to be presented at the 2026 ACM Fairness, Accountability, and Transparency (FAccT) conference, offers a watershed moment in understanding how large language models intersect with everyday healthcare. Their findings resonate with a dual narrative: AI carries tremendous promise to revolutionize medical diagnostics and patient care when stewarded responsibly, but also harbors non-negligible risks, particularly if accessible without proper clinical guidance. As artificial intelligence advances, the path forward must balance innovation with prudence, ensuring these systems enhance rather than undermine the intricate art of medicine.</p>
<hr />
<p><strong>Subject of Research</strong>: Evaluation of large language models’ accuracy and safety in responding to everyday health-related queries by general users.</p>
<p><strong>Article Title</strong>: Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases</p>
<p><strong>News Publication Date</strong>: 25-Jun-2026</p>
<p><strong>Web References</strong>:<br />
<a href="http://dx.doi.org/10.48550/arXiv.2506.13805">10.48550/arXiv.2506.13805</a><br />
<a href="https://facctconference.org/2026/acceptedpapers.html">2026 ACM FAccT Conference</a></p>
<p><strong>References</strong>:<br />
The study data is derived from peer evaluations by board-certified physicians, augmented training on medical textbooks and peer-reviewed articles, and participatory crowdsourced clinical cases generated during the Diagnose-a-thon event hosted by Penn State’s Center for Socially Responsible Artificial Intelligence.</p>
<h4><strong>Keywords</strong></h4>
<p>Generative AI, Artificial Intelligence, Large Language Models, Healthcare, Medical Diagnosis, Clinical Accuracy, AI Ethics, Doctor-Patient Relationship, AI Safety, Medical Informatics, Healthcare Technology, AI in Medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">162310</post-id>	</item>
		<item>
		<title>AI vs. Clinicians: New Study Compares Diagnostic Accuracy</title>
		<link>https://scienmag.com/ai-vs-clinicians-new-study-compares-diagnostic-accuracy/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Tue, 03 Jun 2025 19:16:54 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI diagnostic accuracy]]></category>
		<category><![CDATA[AI in complex medical scenarios]]></category>
		<category><![CDATA[AI response consistency in healthcare]]></category>
		<category><![CDATA[challenges in AI healthcare solutions]]></category>
		<category><![CDATA[comparative study of AI and clinicians]]></category>
		<category><![CDATA[ethical considerations in AI healthcare]]></category>
		<category><![CDATA[experiential judgment in clinical decision-making]]></category>
		<category><![CDATA[healthcare AI study]]></category>
		<category><![CDATA[human clinicians vs AI]]></category>
		<category><![CDATA[limitations of AI in medicine]]></category>
		<category><![CDATA[patient inquiries analysis]]></category>
		<category><![CDATA[strengths of AI in medical questions]]></category>
		<guid isPermaLink="false">https://scienmag.com/ai-vs-clinicians-new-study-compares-diagnostic-accuracy/</guid>

					<description><![CDATA[In a landmark comparative study published in the Journal of Health Organization and Management, researchers from the University of Maine have embarked on a rigorous investigation to evaluate the diagnostic capabilities of artificial intelligence (AI) models against those of seasoned human clinicians when handling multifaceted and sensitive medical queries. By analyzing an extensive dataset comprising [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a landmark comparative study published in the <em>Journal of Health Organization and Management</em>, researchers from the University of Maine have embarked on a rigorous investigation to evaluate the diagnostic capabilities of artificial intelligence (AI) models against those of seasoned human clinicians when handling multifaceted and sensitive medical queries. By analyzing an extensive dataset comprising over 7,000 anonymized patient inquiries sourced from both the United States and Australia, the study offers an unprecedented look into the strengths, limitations, and ethical considerations surrounding AI-driven healthcare solutions amid escalating global health workforce challenges.</p>
<p>The research revealed that AI systems demonstrate considerable proficiency when addressing factual and procedural medical questions, aligning closely with established expert knowledge in these domains. Nevertheless, these models frequently faltered when confronted with nuanced questions requiring explanatory reasoning—commonly framed as “why” and “how” types—highlighting the persistent gap between algorithmic output and the depth of human clinical insight. This disparity underscores the current boundaries of AI’s interpretive and contextual understanding in complex healthcare scenarios, which remain fundamentally reliant on experiential judgment.</p>
<p>One of the study’s most striking findings concerned the consistency of AI responses. Within a given session, AI systems maintained stable answers, yet when the same queries were posed across multiple sessions, variations emerged. Such discrepancies raise significant concerns, especially when diagnostic accuracy and patient safety hang in the balance. This inconsistency suggests that while AI can be a valuable aid, it cannot yet replace the nuanced and adaptive thinking that human clinicians bring to evolving medical cases. These results call for ongoing refinement of AI algorithms to enhance reliability and foster trust among users.</p>
<p>The research further delves into the qualitative aspects of AI-generated responses, particularly their emotional resonance and communicative style. Unlike human clinicians, whose responses exhibited variable lengths tailored to the complexity of inquiries, AI answers were notably uniform, generally comprising between 400 and 475 words regardless of the question’s nature. Moreover, vocabulary analysis revealed AI’s tendency to employ clinical jargon without adapting its language to patient comprehension or emotional sensitivity. This mechanical delivery often lacked the empathy critical in contexts such as mental health discussions or terminal illness consultations, where human warmth and compassion fundamentally shape therapeutic rapport.</p>
<p>Experts consulting on the study emphasized that medical practice hinges on interpersonal connections unreplicable by AI. Physical presence, nuanced communication, and empathetic engagement form the cornerstone of effective healing, roles that technology cannot supplant. Kelley Strout, associate professor at UMaine’s School of Nursing, highlighted that the true transformative potential lies in synergistic integration—where AI augments clinical judgment and compassion rather than attempting to substitute human care providers. Such integration, however, mandates stringent ethical frameworks and vigilant oversight to preempt errors and unintended consequences.</p>
<p>Contextualizing the study within the broader healthcare landscape reveals a striking urgency propelled by systemic strain, particularly in the U.S. The nation grapples with acute shortages in primary and specialty care providers, exacerbating wait times, inflating costs, and disproportionately affecting rural populations. Projections paint a sobering picture: nonmetropolitan areas alone are expected to face a 42% shortfall in primary care physicians by 2037, intensifying existing healthcare disparities. In parallel, the aging population—projected to increase by more than 50% among those aged 65 and older between 2022 and 2026—further amplifies the demand for effective and accessible health services.</p>
<p>Against this backdrop, AI emerges as a potential ally in alleviating some pressure points. The technology could offer round-the-clock virtual assistance, triaging capabilities, and augment patient-provider communication via portals and remote platforms. Yet, researchers caution that the rapid rollout of AI tools, absent comprehensive regulatory guardrails and ethical safeguards, risks eroding care quality and may exacerbate societal inequities, especially if AI systems are trained on limited, non-representative datasets. Ensuring inclusivity in AI development is paramount to avoiding the reinforcement of existing healthcare biases.</p>
<p>A critical lesson drawn from prior technological adoptions—such as the widespread implementation of electronic health records (EHR)—resonates through the study. Despite the promise of EHRs to streamline workflows and improve outcomes, many systems were originally designed around billing imperatives, not clinical efficacy or user experience, resulting in provider dissatisfaction and compromised patient engagement. The study’s experts urge that AI developers heed these mistakes by centering patient outcomes and provider workflows in system design, thus fostering tools that genuinely enhance care delivery rather than provoke frustration or disengagement.</p>
<p>Moreover, the study highlights the pressing need for addressing accountability and patient privacy in the context of AI’s increasing role in clinical decision-making. Ethical concerns loom large, demanding thoughtful policies tailored to the regulatory and cultural environments of implementation locales. Transparency surrounding AI decision processes and mechanisms for error reporting and correction will be essential for widespread acceptance. Without these foundational pillars, AI’s role risks becoming a source of ambiguity and mistrust rather than clarity.</p>
<p>Despite the challenges, the study supports a growing consensus: AI technologies hold immense potential to optimize healthcare by augmenting rather than replacing human providers. By efficiently sifting through vast datasets and highlighting patterns, AI can expedite diagnosis and recommendation processes in ways previously unattainable. However, the emotional intelligence that human clinicians bring, coupled with their ethical judgment, remains irreplaceable in delivering patient-centered care. Balancing these dimensions stands as the new frontier in digital health evolution.</p>
<p>Future research directions outlined by the study underscore the importance of advancing AI’s interpretive capabilities while concurrently managing ethical risks. Tailoring AI tools to diverse healthcare systems—accounting for differences in regulation, culture, and infrastructure—is critical to ensuring equitable and effective deployment. Such adaptations will enable AI to function as a truly supportive asset, enhancing the humanity and efficiency of medical practice without undermining the clinician-patient relationship.</p>
<p>Technological progress in artificial intelligence is poised to reshape healthcare delivery profoundly. Yet, as C. Matt Graham, author of the study, poignantly states, “Technology should enhance the humanity of medicine, not diminish it.” The crux of innovation lies in designing AI systems to serve as complementary extensions of the clinician’s expertise—intelligent assistants enabling more informed, compassionate, and timely care—rather than autonomous decision-makers disconnected from the nuances that define human-centered medicine.</p>
<p>As healthcare systems worldwide wrestle with growing logistical complexities, workforce shortages, and expanding patient needs, the integration of AI offers both formidable opportunities and daunting challenges. This University of Maine study represents a pivotal stepping stone toward clarifying AI’s role, charting a course that maximizes benefits while safeguarding core humanistic values intrinsic to the art and science of healing.</p>
<hr />
<p><strong>Subject of Research</strong>: Comparative analysis of AI and human clinician diagnostic performance on complex medical queries across different healthcare systems</p>
<p><strong>Article Title</strong>: Artificial intelligence vs human clinicians: a comparative analysis of complex medical query handling across the USA and Australia</p>
<p><strong>News Publication Date</strong>: 27-May-2025</p>
<p><strong>Web References</strong>:</p>
<ul>
<li>Journal article: <a href="https://www.emerald.com/insight/content/doi/10.1108/jhom-02-2025-0100/full/html"><a href="https://www.emerald.com/insight/content/doi/10.1108/jhom-02-2025-0100/full/html">https://www.emerald.com/insight/content/doi/10.1108/jhom-02-2025-0100/full/html</a></a>  </li>
<li>DOI: <a href="http://dx.doi.org/10.1108/JHOM-02-2025-0100"><a href="http://dx.doi.org/10.1108/JHOM-02-2025-0100">http://dx.doi.org/10.1108/JHOM-02-2025-0100</a></a>  </li>
</ul>
<p><strong>References</strong>:</p>
<ul>
<li>Health Resources and Services Administration (2024). State of the Primary Care Workforce Report. <a href="https://bhw.hrsa.gov/sites/default/files/bureau-health-workforce/state-of-the-primary-care-workforce-report-2024.pdf"><a href="https://bhw.hrsa.gov/sites/default/files/bureau-health-workforce/state-of-the-primary-care-workforce-report-2024.pdf">https://bhw.hrsa.gov/sites/default/files/bureau-health-workforce/state-of-the-primary-care-workforce-report-2024.pdf</a></a></li>
</ul>
<p><strong>Keywords</strong>: Artificial intelligence, Generative AI, Machine learning, Computer science, Technology, Health and medicine, Health care, Health care delivery, Clinical medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">50938</post-id>	</item>
	</channel>
</rss>
