<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI vs physician performance &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-vs-physician-performance/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 30 Apr 2026 18:54:39 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI vs physician performance &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Landmark Clinical Reasoning Test Shows AI Surpasses Physicians, Setting New Standard for Advanced Evaluation</title>
		<link>https://scienmag.com/landmark-clinical-reasoning-test-shows-ai-surpasses-physicians-setting-new-standard-for-advanced-evaluation/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 30 Apr 2026 18:54:39 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced medical AI evaluation]]></category>
		<category><![CDATA[AI clinical decision support systems]]></category>
		<category><![CDATA[AI diagnostic accuracy]]></category>
		<category><![CDATA[AI vs physician performance]]></category>
		<category><![CDATA[artificial intelligence in healthcare]]></category>
		<category><![CDATA[clinical reasoning AI]]></category>
		<category><![CDATA[collaborative AI medical research]]></category>
		<category><![CDATA[electronic health records complexity]]></category>
		<category><![CDATA[emergency department decision making]]></category>
		<category><![CDATA[Harvard Medical School AI study]]></category>
		<category><![CDATA[large language model diagnostics]]></category>
		<category><![CDATA[real patient chart analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/landmark-clinical-reasoning-test-shows-ai-surpasses-physicians-setting-new-standard-for-advanced-evaluation/</guid>

					<description><![CDATA[In a groundbreaking study conducted by a collaborative team of physicians and computer scientists from Harvard Medical School and Beth Israel Deaconess Medical Center, a large language model (LLM), a form of advanced artificial intelligence, has demonstrated remarkable capabilities in performing complex clinical reasoning tasks typically undertaken by human physicians. Published on April 30, 2026, [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking study conducted by a collaborative team of physicians and computer scientists from Harvard Medical School and Beth Israel Deaconess Medical Center, a large language model (LLM), a form of advanced artificial intelligence, has demonstrated remarkable capabilities in performing complex clinical reasoning tasks typically undertaken by human physicians. Published on April 30, 2026, in the prestigious journal Science, this research represents one of the most comprehensive comparisons to date between AI systems and medical doctors across a wide spectrum of diagnostic and decision-making challenges within emergency department settings.</p>
<p>The investigation centered on whether an LLM could navigate the intricacies of reviewing real, unfiltered patient charts—often fraught with incomplete, inconsistent, or ambiguous data—and effectively synthesize the information to arrive at accurate diagnoses and recommend appropriate next steps. Unlike many prior studies that rely on sanitized or idealized datasets, this research embraced the inherent complexity and &#8220;messiness&#8221; of live electronic health records (EHRs), thereby reflecting authentic clinical environments and offering a robust assessment of AI’s practical performance.</p>
<p>Employing evaluation benchmarks rooted in long-established standards for assessing physician competence—some dating back to methodologies developed in the 1950s—the researchers subjected the model to rigorous diagnostic challenges, clinical reasoning exercises, and real-time emergency department case analyses. The LLM was tested continuously at various critical junctures of patient care, from initial triage when data are sparse to admission decisions informed by more comprehensive clinical findings.</p>
<p>Remarkably, the AI model not only matched but often surpassed the diagnostic accuracy of experienced attending physicians during these early decision points. This finding was particularly striking given the traditionally unpredictable and data-scarce nature of early emergency assessments. Researchers noted that the model&#8217;s ability to operate under these conditions signaled a transformative shift in AI’s readiness to contribute meaningfully to frontline medical decision-making.</p>
<p>Co-senior author Arjun (Raj) Manrai, assistant professor of biomedical informatics at Harvard Medical School, emphasized that while the AI model eclipsed previous iterations and physician baselines across multiple clinical tasks, this accomplishment does not imply that autonomous AI-driven medical practice is imminent. Instead, he underscored the importance of conducting rigorous prospective clinical trials to systematically evaluate the impact and safety of integrating AI tools in diverse care settings before widespread adoption.</p>
<p>Peter Brodeur, MD, MA, a co-first author and clinical researcher at BIDMC, highlighted a significant implication of these findings for the future of AI evaluation metrics. Traditional assessment methodologies, such as multiple-choice tests long used to gauge medical knowledge, no longer offer sufficient resolution to differentiate the rapidly advancing capabilities of modern AI systems, which are now routinely achieving near-perfect scores. This ceiling effect necessitates innovative, contextually rich benchmarks that mirror the nuanced realities of clinical practice.</p>
<p>Furthermore, the study’s design preserved the authenticity of emergency department workflows by presenting the LLM with clinical data precisely as recorded in the EHR, unprocessed and unfiltered. Adam Rodman, MD, MPH, hospitalist and co-senior author, noted the deliberate avoidance of data smoothing techniques common in many AI trials, thereby challenging the model to contend with the full breadth of real-world clinical variability and imperfections.</p>
<p>Despite the model’s promising performance, the researchers maintain a cautious stance regarding its clinical deployment. They acknowledge that although the AI may frequently propose the correct leading diagnosis, it might also recommend additional tests or interventions that are unnecessary or potentially harmful, underscoring that human clinicians must remain integral to the diagnostic workflow to ensure patient safety and care quality.</p>
<p>Thomas Buckley, a doctoral student at Harvard’s AI in Medicine PhD program and co-first author of the study, emphasized the significance of assessing AI’s capabilities early in the diagnostic trajectory, when patient information is limited. This approach more accurately reflects real-world decision-making processes and challenges, challenging the AI to demonstrate proficiency in ambiguous and evolving clinical scenarios rather than well-defined, retrospective cases.</p>
<p>Collectively, these results herald a pivotal moment in the field of medical artificial intelligence. Rather than viewing these systems’ promising diagnostic accuracy as endpoints, the authors advocate for their evaluation through the lens of medical science’s gold standard: controlled clinical trials in authentic healthcare environments. This approach will elucidate the true benefits, limitations, and safety considerations inherent in adopting AI-assisted clinical practice.</p>
<p>The institutions spearheading this research—Harvard Medical School and Beth Israel Deaconess Medical Center—are renowned for their leadership in medical innovation, education, and research. Their combined expertise has facilitated a landmark study that not only challenges previous assumptions about AI’s clinical abilities but also sets a new benchmark for future investigations exploring how artificial intelligence can augment human judgment in medicine.</p>
<p>Looking ahead, the study propels the conversation about AI’s role in healthcare beyond theoretical performance metrics into practical, patient-centered applications. It underscores the pressing need for interdisciplinary collaboration among technologists, clinicians, ethicists, and policymakers to navigate the complex landscape of AI integration responsibly and effectively.</p>
<p>In sum, this research redefines expectations for large language models in clinical environments, proving that AI systems are now capable of reasoning and decision-making at a level that rivals seasoned physicians, particularly in the fast-paced and unpredictable context of emergency medicine. However, it equally stresses that the path forward requires prudence, comprehensive validation, and a reaffirmation of the indispensable role of human expertise in ensuring patient welfare.</p>
<hr />
<p><strong>Subject of Research</strong>: Not applicable</p>
<p><strong>Article Title</strong>: Performance of a large language model on the reasoning tasks of a physician</p>
<p><strong>News Publication Date</strong>: 30-Apr-2026</p>
<p><strong>Web References</strong>: <a href="http://dx.doi.org/10.1126/science.adz4433">10.1126/science.adz4433</a></p>
<h4><strong>Keywords</strong></h4>
<p>AI common sense knowledge, Computer science, Machine learning, Clinical medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">155775</post-id>	</item>
		<item>
		<title>AI Surpasses Physicians in Summarizing Complex Cancer Pathology Reports</title>
		<link>https://scienmag.com/ai-surpasses-physicians-in-summarizing-complex-cancer-pathology-reports/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Thu, 09 Apr 2026 18:04:24 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[AI advancements in cancer diagnostics]]></category>
		<category><![CDATA[AI in oncology pathology]]></category>
		<category><![CDATA[AI vs physician performance]]></category>
		<category><![CDATA[biomarker testing in cancer]]></category>
		<category><![CDATA[cancer pathology report summarization]]></category>
		<category><![CDATA[clinical decision support AI]]></category>
		<category><![CDATA[genetic information in cancer diagnosis]]></category>
		<category><![CDATA[histopathological data AI analysis]]></category>
		<category><![CDATA[immunohistochemical report summarization]]></category>
		<category><![CDATA[large language models in medicine]]></category>
		<category><![CDATA[lung cancer diagnostic data analysis]]></category>
		<category><![CDATA[personalized cancer treatment AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/ai-surpasses-physicians-in-summarizing-complex-cancer-pathology-reports/</guid>

					<description><![CDATA[In a remarkable advancement that merges oncology with cutting-edge artificial intelligence, researchers at Northwestern Medicine have unveiled compelling evidence pointing to the superior performance of AI models in summarizing complex cancer pathology reports. This breakthrough, detailed in a study published on April 8, 2026, in JCO Clinical Cancer Informatics, highlights the transformative potential of AI [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a remarkable advancement that merges oncology with cutting-edge artificial intelligence, researchers at Northwestern Medicine have unveiled compelling evidence pointing to the superior performance of AI models in summarizing complex cancer pathology reports. This breakthrough, detailed in a study published on April 8, 2026, in JCO Clinical Cancer Informatics, highlights the transformative potential of AI to enhance clinical practice, particularly in the nuanced and demanding field of oncology.</p>
<p>Pathology reports have long served as the cornerstone for cancer diagnosis and treatment planning. However, as biomarker testing has proliferated and patient survival rates have improved, these reports have grown increasingly voluminous and intricate. Clinicians often face the challenge of sifting through multi-institutional, longitudinal data dense with histopathological, immunohistochemical, and genetic information, all under significant time constraints. Northwestern’s latest research addresses this critical bottleneck by deploying advanced large language models (LLMs) to generate succinct, comprehensive summaries that capture essential clinical details more reliably than physicians’ own written summaries.</p>
<p>The study&#8217;s authors meticulously analyzed 94 de-identified lung cancer pathology reports, encompassing a broad spectrum of diagnostic data including microscopic tumor characteristics, protein expression profiles, and molecular genetics that inform personalized treatment decisions. The team evaluated six open-source AI language models—Meta’s Llama 3.0, 3.1, and 3.2 variants, Google’s Gemma 9B, DeepSeek-R1, and Mistral 7.2B—each engineered to interpret and synthesize complex textual clinical data without reliance on external cloud-based chatbot frameworks.</p>
<p>Following model-generated summarization, a panel of expert oncologists rigorously assessed the outputs against physician-written clinical summaries. The consensus was striking: AI-generated summaries consistently outperformed their human counterparts, particularly in accurately incorporating molecular and genetic findings crucial for therapeutic strategies. The models’ ability to standardize and elevate the completeness of these summaries marks a significant milestone in addressing informational overload in oncology.</p>
<p>“The complexity of cancer care means clinicians must integrate ever-growing volumes of data, often under intense time pressures,” explained Dr. Mohamed Abazeed, senior study author and Chair of Radiation Oncology at Northwestern University Feinberg School of Medicine. “Our findings underscore that AI doesn’t replace clinical expertise but rather serves as a potent tool to ensure no critical pathological or genomic detail is overlooked—which can be a game-changer for patient outcomes.”</p>
<p>Not all AI architectures performed equally. DeepSeek and Meta’s Llama 3.1 models emerged as the strongest performers, demonstrating superior accuracy and completeness in summarization tasks. Importantly, these models are designed for local deployment, enabling hospital IT systems to integrate AI tools while maintaining patient data privacy—an increasingly vital consideration given heightened concerns about health information security.</p>
<p>Beyond accuracy, the potential clinical impact of this technology is profound. As Dr. Yirong Liu, lead author and radiation oncology resident at McGaw Medical Center, noted, “Patients with complex cancers undergo multiple biopsies and genetic tests across time. Their pathology reports often span dozens of pages. AI-driven summaries can spotlight elusive but critical information—like actionable genetic mutations—that might otherwise be missed, thereby enhancing treatment personalization and improving survival rates.”</p>
<p>The team is currently advancing this research by developing an application powered by Llama 3.1 which will enable clinicians to upload pathology reports and instantly receive AI-generated summaries for review. Nevertheless, the researchers emphasize that before such solutions enter routine clinical practice, extensive validation and testing across broader patient cohorts and cancer types are essential to establish reliability and safety.</p>
<p>This convergence of oncology and artificial intelligence represents a broader trend toward harnessing machine learning tools to manage clinical complexity and optimize workflow efficiency. Unlike conversational chatbots that generate generalized text, these AI systems are specifically trained to digest and condense exhaustive, technical reports into actionable clinical insights, thereby relieving physicians from repetitive, time-consuming documentation tasks.</p>
<p>The implications extend beyond lung cancer, with the potential to revolutionize pathology reporting in other cancer types and chronic diseases that require integrating multifaceted diagnostic data. By ensuring higher fidelity in the transmission of critical diagnostic information, AI-enabled summaries could become an indispensable support layer, augmenting clinical judgment and facilitating more informed decision-making pathways.</p>
<p>Funding for this pioneering work came from prestigious sources, including the Canadian Institute of Health Research and Amazon Web Services’ Social Impact program, reflecting the growing recognition of AI’s pivotal role in healthcare innovation. As these technologies mature, studies like Northwestern’s provide a foundational blueprint for developing AI-driven tools that prioritize patient safety, data security, and enhanced clinical usability.</p>
<p>The Northwestern Medicine study titled “Toward Automating the Summarization of Cancer Pathology Reports Using Large Language Models to Improve Clinical Usability” signals a transformative step forward. It illuminates a future where AI not only augments human intelligence but also fundamentally reshapes how vital medical knowledge is processed, delivered, and utilized in cancer care—potentially translating to better outcomes and improved quality of life for patients worldwide.</p>
<p>Subject of Research: Automating summarization of complex cancer pathology reports using large language models to improve clinical decision-making.</p>
<p>Article Title: Toward Automating the Summarization of Cancer Pathology Reports Using Large Language Models to Improve Clinical Usability</p>
<p>News Publication Date: April 8, 2026</p>
<p>Web References: DOI 10.1200/CCI-25-00284 (JCO Clinical Cancer Informatics)</p>
<p>References: Northwestern University study, JCO Clinical Cancer Informatics, April 8, 2026</p>
<p>Image Credits: Northwestern University</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">150255</post-id>	</item>
	</channel>
</rss>
