<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>ChatGPT-4 surgical applications &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/chatgpt-4-surgical-applications/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 09:02:21 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>ChatGPT-4 surgical applications &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Summaries of Landmark Surgical Studies Pass Expert Accuracy Test</title>
		<link>https://scienmag.com/ai-summaries-of-landmark-surgical-studies-pass-expert-accuracy-test/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 09:02:21 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI accuracy in medical summaries]]></category>
		<category><![CDATA[AI surgical study summaries]]></category>
		<category><![CDATA[AI validation in clinical research]]></category>
		<category><![CDATA[AI-driven medical literature review]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[ChatGPT-4 surgical applications]]></category>
		<category><![CDATA[evidence-based surgery]]></category>
		<category><![CDATA[evidence-based surgical practice]]></category>
		<category><![CDATA[F1-score]]></category>
		<category><![CDATA[graduate medical education]]></category>
		<category><![CDATA[hallucination]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in medicine]]></category>
		<category><![CDATA[medical education and AI]]></category>
		<category><![CDATA[medical literature summarization]]></category>
		<category><![CDATA[natural language processing in healthcare]]></category>
		<category><![CDATA[Perplexity Pro]]></category>
		<category><![CDATA[precision and recall]]></category>
		<category><![CDATA[resident training]]></category>
		<category><![CDATA[surgical education]]></category>
		<category><![CDATA[surgical education challenges]]></category>
		<category><![CDATA[surgical landmark study analysis]]></category>
		<category><![CDATA[surgical trainee learning tools]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=234350</guid>

					<description><![CDATA[A new study finds that ChatGPT-4.0 and Perplexity Pro can summarize landmark surgical studies with near-expert accuracy, and surgical residents say the AI-generated summaries are a valuable supplement to their education.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has now been put to one of medicine&#8217;s most demanding tests: distilling the landmark studies that define modern surgical practice into accurate, teachable summaries. A new study from researchers at the University of Texas Southwestern Medical Center, published in Global Surgical Education, the Journal of the Association for Surgical Education, reports that two widely available large language models, ChatGPT-4.0 and Perplexity Pro, can generate one-page summaries of foundational surgical papers with accuracy scores that rival careful human work. The findings arrive at a moment when surgical trainees face an impossible reading list. PubMed alone now contains more than 40 million citations and abstracts, with over 1.5 million added each year, and the American Council for Graduate Medical Education explicitly requires that residents be able to locate, appraise and assimilate evidence from scientific studies related to their patients. The question the researchers asked was deceptively simple: can a chatbot be trusted to help?</p>
<p>The study&#8217;s design was deliberately rigorous. Rather than asking the models to find important papers on their own, the investigators asked attending surgeons from four subspecialties, emergency general surgery, trauma surgery, surgical critical care and pediatric surgery, to nominate the studies they considered essential for resident education. This expert-driven approach produced a reference set of 26 landmark studies, including six in emergency general surgery, five in trauma surgery, six in surgical critical care and nine in pediatric surgery. Because the goal was to measure how faithfully a model could condense a known paper, not how well it could identify influential literature, curating the list by hand was a critical methodological choice. It meant that every summary could later be checked line by line against a definitive source document chosen by people who teach these papers for a living.</p>
<p>Both models were then given the same standardized task. Each study was uploaded as a PDF in a fresh browser session to eliminate contamination from prior conversation history, and the models received a prompt engineered to simulate expert teaching: they were instructed to act as a subspecialty surgeon teaching residents, to produce a one-page summary covering background, study design, inclusion and exclusion criteria, treatment groups, key findings and clinical applications, and to briefly describe how the data had changed clinical practice since publication. All of the selected studies were publicly available, so no proprietary or restricted content entered the models. The standardized prompt matters more than it might appear, because prompt wording is known to shape the quality and emphasis of large language model output, and a single consistent prompt allowed fair head-to-head comparison between the two systems.</p>
<p>To score the results, the researchers borrowed evaluation machinery from machine learning itself. Each summary was independently assessed by two faculty reviewers from the relevant subspecialty, both blinded to the fact that the summaries were AI-generated and told to evaluate them as if residents had prepared them for educational purposes. Documents were randomly labeled to prevent bias based on perceived source quality. Reviewers classified each of six domains, background, methods, inclusion and exclusion criteria, treatment groups, key findings and clinical relevance, as a true positive, a false positive or a false negative. A true positive meant the AI content was accurate and present in the original paper; a false positive meant the summary contained information absent from the source, the phenomenon better known as a hallucination; a false negative meant the original paper contained essential information the model failed to include. From these labels the team calculated precision, recall and the composite F1 score for each model and domain.</p>
<p>The headline numbers were strikingly strong. ChatGPT-4.0 achieved a composite precision of 0.93, a recall of 0.96 and an F1 score of 0.95. Perplexity Pro performed comparably, with a precision of 0.96, a recall of 0.94 and an F1 score of 0.95. In the machine learning literature, F1 scores above 0.70 are generally considered acceptable and scores above 0.90 excellent, though appropriate thresholds depend on task complexity. By that yardstick, both models delivered excellent performance on a task that demands deep domain knowledge. A Mann-Whitney U test found no statistically significant differences between the two systems on any composite metric, suggesting that the choice between a general-purpose assistant and a web-connected AI search engine mattered less than one might expect for this constrained summarization task.</p>
<p>The weak spot, when the scores were broken down by domain, was instructive. Both models performed best on background, treatment groups, key findings and clinical relevance, but stumbled on inclusion and exclusion criteria, the fine-grained eligibility rules that determine exactly which patients a study applies to. ChatGPT-4.0 posted a precision of 0.93, a recall of 0.84 and an F1 of 0.89 in that domain, while Perplexity Pro achieved a perfect precision of 1.00 but a recall of just 0.81, for an F1 of 0.90. Crucially, the shortfall was driven by omission rather than invention: the models tended to drop detailed eligibility information rather than fabricate it. This reflects a known tendency of large language models to prioritize information they perceive as most salient, retaining study objectives and headline findings while discarding granular methodological detail, a pattern that prior research has linked to overgeneralization of scientific findings in AI summaries.</p>
<p>Qualitative reviewer comments reinforced the quantitative picture. Most errors were omissions, typically involving incomplete reporting of inclusion and exclusion criteria, secondary outcomes or detailed methodological descriptions. True hallucinations were uncommon and most often took the form of overstated clinical implications, with models extrapolating broader practice changes than the original manuscripts supported. Importantly, reviewers rarely identified hallucinations that would meaningfully alter the interpretation of a study&#8217;s primary findings, and they frequently noted that the summaries captured the main results accurately while providing concise, readable overviews. Interrater agreement across all reviewer pairs was 84 percent overall, ranging from 77 percent in pediatric surgery to 91 percent in surgical critical care, indicating that the expert scoring itself was reasonably consistent even if percent agreement does not account for chance concordance.</p>
<p>The educational verdict came from the residents themselves. A voluntary survey distributed to 62 general surgery residents across all postgraduate years drew 15 responses, a 24 percent response rate. The results were emphatic: every respondent rated the summaries as clear or very clear, all but one said the summaries captured key findings well or very well, 80 percent judged them accurate or very accurate, and all rated them helpful or very helpful for understanding core concepts. Every respondent agreed the summaries would be a beneficial supplement to regular study materials, and all recommended continued use of LLM-generated summaries. Thirteen of fifteen felt the one-page format struck the right balance between brevity and comprehensiveness. Trust, however, remained calibrated: only a third of residents said they would trust the summaries without additional verification, while the rest wanted review by faculty, by other residents or by themselves against the original papers, a pattern the authors read as trainees valuing the tool while still recognizing the importance of consulting primary literature.</p>
<p>The study is the first, according to its authors, to validate large language models for summarizing key surgical literature, and it extends a body of work showing that models like ChatGPT can pass standardized medical exams, generate realistic clinical vignettes and deliver feedback to trainees. The authors are candid about limitations. Summaries were generated with a single standardized prompt, leaving open whether prompt optimization could improve performance; only four subspecialties were included; accuracy was scored at the section level rather than claim by claim, which could underestimate performance; and rapidly evolving model outputs may not be reproducible over time. They also point to retrieval augmented generation as a promising fix, citing work in which a modified GPT-4 answered clinical guideline questions correctly 84 percent of the time, up from 57 percent for the base model. For now, the message is one of cautious enthusiasm: AI-generated summaries of landmark surgical studies are accurate enough to earn a place in resident education, provided that human oversight, and the original papers, remain close at hand.</p>
<p><strong>Subject of Research:</strong> Evaluating large language models for generating accurate summaries of landmark surgical studies for resident education</p>
<p><strong>Article Title:</strong> Harnessing large language models to summarize landmark surgical studies for resident education</p>
<p><strong>Article References:</strong> Pettigrew, M. F., Tyler, L. A., Gregory, A. L., Nomellini, V., Bhat, S. G., Murphy, J. T., Purcell, L. N., Butler, D., Park, C., Clark, A., &amp; Abdelfattah, K. R. (2026). Harnessing large language models to summarize landmark surgical studies for resident education. <em>Global Surgical Education &#8211; Journal of the Association for Surgical Education, 5</em>(1), Article 135. <a href="https://doi.org/10.1007/s44186-026-00537-z" rel="noopener noreferrer">https://doi.org/10.1007/s44186-026-00537-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44186-026-00537-z" rel="noopener noreferrer">10.1007/s44186-026-00537-z</a></p>
<p><strong>Keywords:</strong> artificial intelligence, large language models, ChatGPT, Perplexity Pro, surgical education, resident training, medical literature summarization, hallucination, precision and recall, F1 score, graduate medical education, evidence-based surgery</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">234350</post-id>	</item>
	</channel>
</rss>
