Artificial intelligence is steadily moving into the most intimate corners of medical training, and one of the newest frontiers is the automated feedback that large language models now generate for surgical residents after simulation exercises. A research team at Mayo Clinic has introduced a way to ask a question that has largely gone unexamined: when an AI tells a trainee how well they performed, is that feedback actually telling the truth? The study, published in Global Surgical Education, the journal of the Association for Surgical Education, offers a psychometric framework that borrows two of the oldest ideas in educational measurement, validity and reliability, and translates them into metrics designed specifically for machine-generated narrative feedback.
The problem the researchers identified is deceptively simple. Existing validation efforts for AI in surgical assessment tend to focus on classification accuracy, meaning how reliably an algorithm can distinguish a novice from an expert. But classification is not feedback. A trainee does not benefit from being sorted into a category; they benefit from language that accurately reflects their level of performance and that responds sensibly when their performance changes. Until now, no widely adopted framework existed to measure whether the tone of AI-generated commentary appropriately mirrors how well a resident actually performed, or whether that commentary shifts consistently in the right direction when scores rise or fall.
To fill that gap, Mohamed S. Baloul, Jonathan D’Angelo and their colleagues proposed two complementary metrics. The first is Calibration, expressed as the squared correlation between the sentiment of the feedback and the resident’s actual performance scores. In psychometric terms, it functions as an analogue of validity evidence, asking whether the emotional register of the machine’s commentary matches the quality of the work being described. The second is the Feedback Robustness Index, or FRI, which measures the percentage of times the direction of the feedback correctly follows a controlled change in performance scores. This acts as an analogue of reliability, probing whether the system behaves consistently rather than drifting arbitrarily. Sentiment was quantified using two established natural language processing tools, TextBlob and VADER, providing a reproducible, algorithmic measure of how positive or negative each piece of feedback was.
The experimental design was unusually systematic. The team drew assessments from nine second-year general surgery residents per station, deliberately selected to span the full range of performance, across five standardized simulation stations: FLS Knot Tying, FLS Circle Cutting, Vascular Anastomosis, Endostitch, and Imaging Interpretation. These stations were scored with rubrics ranging from minimal, essentially time and accuracy alone, to detailed, behaviorally anchored descriptors that spell out what specific levels of skill look like. Feedback was generated by Qwen3 4B, an open-source large language model running locally, a choice that keeps sensitive trainee data on institutional hardware and makes the entire pipeline transparent and reproducible.
The crucial manipulation came next. For each assessment, the researchers perturbed the resident’s scores in five controlled ways: a global improvement, a global decline, a targeted correction of a weakness, a targeted deterioration of a strength, and a minimal change that should barely register. In total, 225 baseline-to-perturbed comparisons were generated, each producing a fresh piece of AI feedback that could be scored for sentiment and compared with what the score changes should have implied. This perturbation approach transforms a vague worry about AI quality into a measurable stress test, revealing exactly where the model’s commentary tracks reality and where it loses the plot.
The results were striking in their variation, and they point to a single dominant factor: the design of the assessment rubric. The two stations equipped with detailed, behaviorally anchored rubrics produced the strongest AI feedback. FLS Knot Tying achieved a calibration of R-squared 0.40 with a Feedback Robustness Index of 84 percent, while Vascular Anastomosis reached R-squared 0.37 and an FRI of 73 percent. At those stations, the language model’s commentary genuinely rose and fell with performance, and did so most of the time the scores changed. For trainees, that means the feedback they read bore a meaningful relationship to what they actually did with the needle driver or the anastomosis.
The picture deteriorated sharply where rubrics were thinner. Stations with structured but not behaviorally anchored rubrics showed moderate robustness yet essentially zero calibration: the Endostitch station recorded R-squared 0.00 with an FRI of 64 percent, and Imaging Interpretation posted R-squared 0.00 with an FRI of 69 percent. In plain terms, the AI could often tell when performance had improved or declined, but the overall tone of its feedback bore no relationship to the resident’s actual level. Worst of all was the station with a minimal, two-item rubric, FLS Circle Cutting, where calibration sat near zero at R-squared 0.03 and the Feedback Robustness Index of 49 percent was no better than a coin flip. Across all assessments, the mean FRI of 68 percent versus a mean R-squared of just 0.16 captured the study’s central asymmetry: AI feedback moves in the right direction more often than it matches the resident’s overall performance level.
Why would rubric detail matter so much to a language model? The answer likely lies in how these systems generate text. A behaviorally anchored rubric gives the model rich, specific descriptors of what novice, intermediate, and expert performance actually look like, providing the raw material for commentary that distinguishes a struggling resident from a strong one. A sparse rubric offering only time and accuracy leaves the model with almost nothing to anchor its language to, so it defaults to generic phrasing that cannot differentiate performance levels even when the numbers change. The researchers also situate their findings in a broader concern from the AI literature: sycophancy, the well-documented tendency of language models to tell users what they want to hear. A sycophantic feedback generator could shower praise on weak performers or fail to sharpen its criticism when scores plummet, which is precisely the failure mode the FRI is designed to detect.
The practical implications extend well beyond one simulation lab. Surgical education has long wrestled with feedback that is inconsistent, delayed, or influenced by unconscious biases, and large language models promise scalable, immediately available narrative commentary at a fraction of the cost of faculty time. But this study demonstrates that the quality of that commentary is not a fixed property of the model; it is a joint product of the model and the assessment instrument feeding it. A program that invests in a capable language model but pairs it with a skeletal rubric may be deploying a system whose praise and criticism are essentially uncorrelated with trainee performance, potentially misleading learners at the most formative stage of their development.
The authors’ recommendation is therefore twofold. Programs implementing AI-generated feedback should evaluate those systems with calibration and robustness metrics, alongside conventional expert review, before putting them in front of trainees. And they should ensure that the assessment rubrics underlying those systems contain detailed performance criteria, because the framework developed here shows that rubric design is the variable most tightly associated with whether AI feedback is appropriate and consistent. As institutions increasingly hand narrative evaluation to machines, this study offers something the field has lacked: a rigorous, replicable way to audit what the machine is actually saying about a surgeon’s hands before anyone trusts it enough to learn from it.
Subject of Research: Psychometric evaluation of AI-generated feedback quality in surgical simulation assessment
Article Title: A psychometric framework for AI feedback systems in surgical simulation assessments
Article References: Baloul, M. S., Giri, O., Mehta, A., Cui, D., & D’Angelo, J. (2026). A psychometric framework for AI feedback systems in surgical simulation assessments. Global Surgical Education – Journal of the Association for Surgical Education, 5(1), Article 183. https://doi.org/10.1007/s44186-026-00588-2
Image Credits: AI Generated
DOI: 10.1007/s44186-026-00588-2
Keywords: artificial intelligence, large language models, surgical education, psychometrics, simulation assessment, feedback calibration, Feedback Robustness Index, sentiment analysis, rubric design, medical training, Qwen3, surgical residents
Cite Scienmag News
Courtney Benton. (September 30, 2026). New Psychometric Tests Reveal When AI Feedback on Surgeons Fails. Scienmag. https://scienmag.com/new-psychometric-tests-reveal-when-ai-feedback-on-surgeons-fails/
Courtney Benton. "New Psychometric Tests Reveal When AI Feedback on Surgeons Fails." Scienmag, 30 September 2026, https://scienmag.com/new-psychometric-tests-reveal-when-ai-feedback-on-surgeons-fails/. Accessed 30 September 2026.
Courtney Benton. "New Psychometric Tests Reveal When AI Feedback on Surgeons Fails." Scienmag. September 30, 2026. https://scienmag.com/new-psychometric-tests-reveal-when-ai-feedback-on-surgeons-fails/

