<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI vs human peer review in plastic surgery &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-vs-human-peer-review-in-plastic-surgery/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 03:56:02 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI vs human peer review in plastic surgery &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Manuscript Reviewer Outperforms Human Peer Reviewers in Plastic Surgery Study</title>
		<link>https://scienmag.com/ai-manuscript-reviewer-outperforms-human-peer-reviewers-in-plastic-surgery-study/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 03:56:02 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI manuscript review]]></category>
		<category><![CDATA[AI performance in medical research evaluation]]></category>
		<category><![CDATA[AI tools for journal feedback]]></category>
		<category><![CDATA[AI vs human peer review in plastic surgery]]></category>
		<category><![CDATA[AI-driven manuscript evaluation]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[ChatGPT in scholarly publishing]]></category>
		<category><![CDATA[democratizing access to scientific review]]></category>
		<category><![CDATA[equity in scientific publishing]]></category>
		<category><![CDATA[global surgery]]></category>
		<category><![CDATA[language barriers in peer review]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in academic peer review]]></category>
		<category><![CDATA[low-and-middle-income countries]]></category>
		<category><![CDATA[manuscript evaluation]]></category>
		<category><![CDATA[meta-prompting]]></category>
		<category><![CDATA[peer review]]></category>
		<category><![CDATA[plastic surgery]]></category>
		<category><![CDATA[prompt engineering for manuscript assessment]]></category>
		<category><![CDATA[research equity]]></category>
		<category><![CDATA[Review Quality Instrument]]></category>
		<category><![CDATA[scientific publishing]]></category>
		<category><![CDATA[systemic biases in peer review]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=251621</guid>

					<description><![CDATA[An iteratively optimized ChatGPT-based tool scored significantly higher than human peer reviewers on a validated measure of manuscript review quality, offering accessible pre-submission feedback for researchers worldwide.]]></description>
										<content:encoded><![CDATA[<p>An artificial intelligence tool built from a standard ChatGPT account has scored higher than human peer reviewers in a head-to-head test of manuscript evaluation quality, according to a new study published in Global Surgical Education, the journal of the Association for Surgical Education. The research, led by plastic surgeon Ashit Patel and resident Raunak Goyal at Duke University along with collaborators at four other US medical institutions, suggests that careful, iterative prompt engineering—not expensive custom systems—may be enough to give researchers structured, journal-quality feedback on their papers before submission. The finding lands at a moment when the scientific community is wrestling with both the promise and the peril of large language models in scholarly publishing.</p>
<p>The motivation behind the study is rooted in a long-documented inequity in global science. Authors from low- and middle-income countries and from regions where English is not the primary language face systemic disadvantages when preparing manuscripts for international journals. Previous studies cited by the team have shown that geographic bias shapes peer reviewer selection, that English-language dominance disadvantages researchers in places like Colombia, and that historically excluded groups continue to encounter barriers in the peer review process. Rejection and resubmission cycles impose real costs on early-career researchers everywhere, but the burden falls hardest on those without access to experienced mentors or institutional support.</p>
<p>The Duke-led team set out to test whether an accessible AI tool could help level that playing field. They designed a controlled comparison of three progressively refined AI approaches, all reachable through an ordinary ChatGPT account, meaning no specialized software, no programming expertise, and no additional cost beyond a standard subscription. The first tool was a baseline deployment of GPT-5 with no special configuration. The second was a Custom GPT augmented with a knowledge base of peer review guidance. The third and most sophisticated version layered on iterative optimization: enhanced system instructions developed through meta-prompting, a technique in which the model itself is used to help draft and refine the instructions given to it, combined with journal-specific guidelines drawn from the field of plastic surgery.</p>
<p>To evaluate these tools rigorously, the researchers assembled sixty manuscript evaluations covering fifteen plastic surgery articles, deliberately spread across five clinical studies, five clinical trials, and five review articles. Each AI-generated evaluation was compared against human peer reviews obtained from journals that practice open peer review, where reviewer reports are published alongside the paper. Two blinded raters then scored every evaluation—human and machine alike—using the Review Quality Instrument, a validated assessment tool developed in 1999 that measures review quality on a scale with a maximum score of 40. The raters did not know which reviews came from humans and which came from AI.</p>
<p>The results were striking. The iteratively optimized Custom GPT, Tool 3, achieved a mean Review Quality Instrument score of 28.5 with a standard deviation of 3.8, significantly higher than every other review type in the study. Human peer reviews averaged 25.2 plus or minus 4.9, the knowledge-base Custom GPT averaged 24.3 plus or minus 5.0, and the baseline GPT-5 averaged 24.0 plus or minus 1.8. The differences were statistically significant at p less than 0.001. In other words, the carefully engineered AI tool did not merely match human reviewers—it outperformed them on average, while the unoptimized baseline model sat slightly below human performance.</p>
<p>Digging into the subcomponents of review quality revealed where the optimized tool excelled most. Using Kendall&#8217;s W, a statistical measure of effect size for non-parametric comparisons, the team found that Tool 3 showed its largest advantages in methodological assessment, evidence provision, and interpretation evaluation, with effect sizes of 0.69, 0.59, and 0.63 respectively. These are precisely the dimensions that separate a superficial review from a genuinely useful one: whether the reviewer critically appraises the study design, whether claims are backed by specific evidence from the manuscript, and whether the reviewer interprets the findings in context rather than offering generic praise or criticism.</p>
<p>Equally notable was the tool&#8217;s consistency. Across the three manuscript types—clinical studies, clinical trials, and reviews—the optimized AI&#8217;s scores ranged only from 28.0 to 29.1, a narrow band suggesting stable performance regardless of article genre. Human reviews, by contrast, showed considerably greater variability, with mean scores ranging from 22.1 to 27.0 across manuscript types. That variability mirrors a well-known reality of peer review: quality depends heavily on the individual reviewer, their time constraints, and their familiarity with the subject matter. A tool that delivers reliably structured feedback could smooth out some of that randomness for authors seeking pre-submission guidance.</p>
<p>The study builds on a growing body of work examining whether large language models can provide useful feedback on research papers. Earlier analyses, including a large-scale empirical study published in NEJM AI, found that models like GPT-4 could identify weaknesses in papers that human reviewers sometimes missed, though concerns about hallucinated criticisms and superficiality persisted. Subsequent studies in hand surgery and transplantation research have compared AI-generated and human peer reviews with mixed results. What distinguishes the new work is its focus on accessibility and optimization: rather than asking whether AI can review manuscripts in the abstract, the team asked how much better a freely accessible tool can become through systematic instruction refinement and domain-specific knowledge integration.</p>
<p>The implications for global surgical research are considerable. A researcher in a resource-limited setting could, in principle, paste a draft manuscript into the optimized tool and receive structured, actionable feedback on methodology, evidence, and interpretation—feedback that might otherwise require a mentor at a well-funded institution. The authors emphasize that the tool requires no technical expertise and no additional cost, which addresses two of the most common barriers to AI implementation in healthcare identified in prior literature: control and cost. Because the tool is designed for pre-submission refinement rather than replacing journal peer review, it also sidesteps some of the thornier ethical questions about AI&#8217;s role in editorial decision-making, though the broader debate about generative AI in scientific publishing remains unresolved, as a 2024 Lancet commentary highlighted.</p>
<p>Cautions remain. The study evaluated reviews of plastic surgery manuscripts scored by two raters using a single validated instrument, and the human comparison reviews came from open peer review journals, which may not represent the full spectrum of review quality across all journals. The Review Quality Instrument measures the construct of a review report, not the ultimate accuracy of an AI&#8217;s scientific judgments, and large language models can still produce confident errors. The authors report no external funding and no conflicts of interest, and they have made their evaluation scores available in supplementary materials for scrutiny. Still, the central result stands as a provocative data point: with thoughtful meta-prompting and domain knowledge, a tool anyone can access from a standard ChatGPT account produced manuscript evaluations that blinded raters judged better than those written by human experts. For the millions of researchers excluded from the informal networks of academic mentorship, that could be a genuinely democratizing development.</p>
<p><strong>Subject of Research:</strong> Development and evaluation of an accessible AI-powered tool for manuscript peer review in plastic surgery research</p>
<p><strong>Article Title:</strong> Development and assessment of an accessible AI-powered tool for manuscript evaluation through iterative optimization in plastic surgery research</p>
<p><strong>Article References:</strong> Goyal, R., Liang, J., Kulkarni, M. S., Nguyen, P., Gillipelli, S., &amp; Patel, A. (2026). Development and assessment of an accessible AI-powered tool for manuscript evaluation through iterative optimization in plastic surgery research. <em>Global Surgical Education &#8211; Journal of the Association for Surgical Education, 5</em>(1), Article 188. <a href="https://doi.org/10.1007/s44186-026-00599-z" rel="noopener noreferrer">https://doi.org/10.1007/s44186-026-00599-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44186-026-00599-z" rel="noopener noreferrer">10.1007/s44186-026-00599-z</a></p>
<p><strong>Keywords:</strong> artificial intelligence, large language models, ChatGPT, peer review, manuscript evaluation, plastic surgery, global surgery, meta-prompting, Review Quality Instrument, research equity, low- and middle-income countries, scientific publishing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">251621</post-id>	</item>
	</channel>
</rss>
