Wednesday, September 30, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Social Science

New Psychometric Tests Reveal When AI Feedback on Surgeons Fails

September 30, 2026
in Social Science
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 5 mins read
0
New Psychometric Tests Reveal When AI Feedback on Surgeons Fails

New Psychometric Tests Reveal When AI Feedback on Surgeons Fails

New Psychometric Tests Reveal When AI Feedback on Surgeons Fails

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Artificial intelligence is steadily moving into the most intimate corners of medical training, and one of the newest frontiers is the automated feedback that large language models now generate for surgical residents after simulation exercises. A research team at Mayo Clinic has introduced a way to ask a question that has largely gone unexamined: when an AI tells a trainee how well they performed, is that feedback actually telling the truth? The study, published in Global Surgical Education, the journal of the Association for Surgical Education, offers a psychometric framework that borrows two of the oldest ideas in educational measurement, validity and reliability, and translates them into metrics designed specifically for machine-generated narrative feedback.

The problem the researchers identified is deceptively simple. Existing validation efforts for AI in surgical assessment tend to focus on classification accuracy, meaning how reliably an algorithm can distinguish a novice from an expert. But classification is not feedback. A trainee does not benefit from being sorted into a category; they benefit from language that accurately reflects their level of performance and that responds sensibly when their performance changes. Until now, no widely adopted framework existed to measure whether the tone of AI-generated commentary appropriately mirrors how well a resident actually performed, or whether that commentary shifts consistently in the right direction when scores rise or fall.

To fill that gap, Mohamed S. Baloul, Jonathan D’Angelo and their colleagues proposed two complementary metrics. The first is Calibration, expressed as the squared correlation between the sentiment of the feedback and the resident’s actual performance scores. In psychometric terms, it functions as an analogue of validity evidence, asking whether the emotional register of the machine’s commentary matches the quality of the work being described. The second is the Feedback Robustness Index, or FRI, which measures the percentage of times the direction of the feedback correctly follows a controlled change in performance scores. This acts as an analogue of reliability, probing whether the system behaves consistently rather than drifting arbitrarily. Sentiment was quantified using two established natural language processing tools, TextBlob and VADER, providing a reproducible, algorithmic measure of how positive or negative each piece of feedback was.

The experimental design was unusually systematic. The team drew assessments from nine second-year general surgery residents per station, deliberately selected to span the full range of performance, across five standardized simulation stations: FLS Knot Tying, FLS Circle Cutting, Vascular Anastomosis, Endostitch, and Imaging Interpretation. These stations were scored with rubrics ranging from minimal, essentially time and accuracy alone, to detailed, behaviorally anchored descriptors that spell out what specific levels of skill look like. Feedback was generated by Qwen3 4B, an open-source large language model running locally, a choice that keeps sensitive trainee data on institutional hardware and makes the entire pipeline transparent and reproducible.

The crucial manipulation came next. For each assessment, the researchers perturbed the resident’s scores in five controlled ways: a global improvement, a global decline, a targeted correction of a weakness, a targeted deterioration of a strength, and a minimal change that should barely register. In total, 225 baseline-to-perturbed comparisons were generated, each producing a fresh piece of AI feedback that could be scored for sentiment and compared with what the score changes should have implied. This perturbation approach transforms a vague worry about AI quality into a measurable stress test, revealing exactly where the model’s commentary tracks reality and where it loses the plot.

The results were striking in their variation, and they point to a single dominant factor: the design of the assessment rubric. The two stations equipped with detailed, behaviorally anchored rubrics produced the strongest AI feedback. FLS Knot Tying achieved a calibration of R-squared 0.40 with a Feedback Robustness Index of 84 percent, while Vascular Anastomosis reached R-squared 0.37 and an FRI of 73 percent. At those stations, the language model’s commentary genuinely rose and fell with performance, and did so most of the time the scores changed. For trainees, that means the feedback they read bore a meaningful relationship to what they actually did with the needle driver or the anastomosis.

The picture deteriorated sharply where rubrics were thinner. Stations with structured but not behaviorally anchored rubrics showed moderate robustness yet essentially zero calibration: the Endostitch station recorded R-squared 0.00 with an FRI of 64 percent, and Imaging Interpretation posted R-squared 0.00 with an FRI of 69 percent. In plain terms, the AI could often tell when performance had improved or declined, but the overall tone of its feedback bore no relationship to the resident’s actual level. Worst of all was the station with a minimal, two-item rubric, FLS Circle Cutting, where calibration sat near zero at R-squared 0.03 and the Feedback Robustness Index of 49 percent was no better than a coin flip. Across all assessments, the mean FRI of 68 percent versus a mean R-squared of just 0.16 captured the study’s central asymmetry: AI feedback moves in the right direction more often than it matches the resident’s overall performance level.

Why would rubric detail matter so much to a language model? The answer likely lies in how these systems generate text. A behaviorally anchored rubric gives the model rich, specific descriptors of what novice, intermediate, and expert performance actually look like, providing the raw material for commentary that distinguishes a struggling resident from a strong one. A sparse rubric offering only time and accuracy leaves the model with almost nothing to anchor its language to, so it defaults to generic phrasing that cannot differentiate performance levels even when the numbers change. The researchers also situate their findings in a broader concern from the AI literature: sycophancy, the well-documented tendency of language models to tell users what they want to hear. A sycophantic feedback generator could shower praise on weak performers or fail to sharpen its criticism when scores plummet, which is precisely the failure mode the FRI is designed to detect.

The practical implications extend well beyond one simulation lab. Surgical education has long wrestled with feedback that is inconsistent, delayed, or influenced by unconscious biases, and large language models promise scalable, immediately available narrative commentary at a fraction of the cost of faculty time. But this study demonstrates that the quality of that commentary is not a fixed property of the model; it is a joint product of the model and the assessment instrument feeding it. A program that invests in a capable language model but pairs it with a skeletal rubric may be deploying a system whose praise and criticism are essentially uncorrelated with trainee performance, potentially misleading learners at the most formative stage of their development.

The authors’ recommendation is therefore twofold. Programs implementing AI-generated feedback should evaluate those systems with calibration and robustness metrics, alongside conventional expert review, before putting them in front of trainees. And they should ensure that the assessment rubrics underlying those systems contain detailed performance criteria, because the framework developed here shows that rubric design is the variable most tightly associated with whether AI feedback is appropriate and consistent. As institutions increasingly hand narrative evaluation to machines, this study offers something the field has lacked: a rigorous, replicable way to audit what the machine is actually saying about a surgeon’s hands before anyone trusts it enough to learn from it.

Subject of Research: Psychometric evaluation of AI-generated feedback quality in surgical simulation assessment

Article Title: A psychometric framework for AI feedback systems in surgical simulation assessments

Article References: Baloul, M. S., Giri, O., Mehta, A., Cui, D., & D’Angelo, J. (2026). A psychometric framework for AI feedback systems in surgical simulation assessments. Global Surgical Education – Journal of the Association for Surgical Education, 5(1), Article 183. https://doi.org/10.1007/s44186-026-00588-2

Image Credits: AI Generated

DOI: 10.1007/s44186-026-00588-2

Keywords: artificial intelligence, large language models, surgical education, psychometrics, simulation assessment, feedback calibration, Feedback Robustness Index, sentiment analysis, rubric design, medical training, Qwen3, surgical residents

Cite Scienmag News

Courtney Benton. (September 30, 2026). New Psychometric Tests Reveal When AI Feedback on Surgeons Fails. Scienmag. https://scienmag.com/new-psychometric-tests-reveal-when-ai-feedback-on-surgeons-fails/

Courtney Benton. "New Psychometric Tests Reveal When AI Feedback on Surgeons Fails." Scienmag, 30 September 2026, https://scienmag.com/new-psychometric-tests-reveal-when-ai-feedback-on-surgeons-fails/. Accessed 30 September 2026.

Courtney Benton. "New Psychometric Tests Reveal When AI Feedback on Surgeons Fails." Scienmag. September 30, 2026. https://scienmag.com/new-psychometric-tests-reveal-when-ai-feedback-on-surgeons-fails/

Tags: AI language models for surgical resident evaluationAI-based surgical feedback validationArtificial Intelligencechallenges in AI-based surgical assessmenteffectiveness of AI feedback in surgical educationevaluation of machine-generated surgical performance reportsfeedback calibrationFeedback Robustness Indeximpact of AI feedback tone on surgical trainee learninglarge language modelslimitations of AI in medical skill assessmentmeasuring authenticity of AI-generated surgical feedbackmedical trainingpsychometric framework for AI assessmentpsychometricsQwen3reliability and validity of AI in medical trainingrubric designsentiment analysissimulation assessmentsurgical educationsurgical residentssurgical simulation AI feedback accuracyvalidation of AI tools for surgical skill improvement
Share26Tweet16
Previous Post

Terraces That Heal: Conservation Structures Rebuild Soil and Water in Ethiopia’s Highlands

Next Post

Zinc Oxide Quantum Dots Take a Step Toward Spin-Based Quantum Computing

Related Posts

Satellite images reveal Ghana’s Garden City is losing its green soul
Social Science

Satellite images reveal Ghana’s Garden City is losing its green soul

September 30, 2026
Why China Is Betting on Gradual Reform, Not Revolution, in the Global Order
Social Science

Why China Is Betting on Gradual Reform, Not Revolution, in the Global Order

September 30, 2026
Happy Teachers, Motivated Students: Giant Meta-Analysis Confirms the Classroom Link
Social Science

Happy Teachers, Motivated Students: Giant Meta-Analysis Confirms the Classroom Link

September 30, 2026
Researchers Name the Infrastructure Firms Powering AI Deepfake Pornography Sites
Social Science

Researchers Name the Infrastructure Firms Powering AI Deepfake Pornography Sites

September 30, 2026
A Statistical Rethink Suggests Judges May Not Be ‘Liberated’ the Way Criminologists Thought
Social Science

A Statistical Rethink Suggests Judges May Not Be ‘Liberated’ the Way Criminologists Thought

September 30, 2026
New Digital Literacy Scale Puts Computational Thinking and Lifelong Learning to the Test
Social Science

New Digital Literacy Scale Puts Computational Thinking and Lifelong Learning to the Test

September 30, 2026
Next Post
Zinc Oxide Quantum Dots Take a Step Toward Spin-Based Quantum Computing

Zinc Oxide Quantum Dots Take a Step Toward Spin-Based Quantum Computing

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Zinc Oxide Quantum Dots Take a Step Toward Spin-Based Quantum Computing
  • New Psychometric Tests Reveal When AI Feedback on Surgeons Fails
  • Terraces That Heal: Conservation Structures Rebuild Soil and Water in Ethiopia’s Highlands
  • Machine Learning Cracks the Optical Code of Lithium Bromide Solutions

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading