In the operating rooms and clinics of modern teaching hospitals, the education of a young surgeon increasingly depends on a deceptively small artifact: a short block of narrative text attached to a workplace assessment. A new study from the University of Florida College of Medicine and collaborating institutions asks a question that residency programs everywhere have quietly wondered about — does a faculty member who submits more of these assessments actually write better feedback? The answer, published in Global Surgical Education, the journal of the Association for Surgical Education, appears to be yes, and the finding carries practical consequences for how surgical programs train and support their teaching faculty.
The research team, led by Michel S. Kabbash, examined Entrustable Professional Activity, or EPA, micro-assessments submitted to general surgery residents over the 2023 to 2025 academic years. EPAs are a cornerstone of competency-based medical education, a reform movement that shifted training away from time-based benchmarks and toward demonstrated ability to perform trusted professional tasks. Rather than judging a resident on a single high-stakes examination, EPA micro-assessments capture frequent, low-stakes observations of real clinical work — in this study, the management of three bread-and-butter surgical conditions: gallbladder disease, appendicitis, and inguinal hernia. Each micro-assessment includes a narrative comment in which the supervising faculty surgeon explains what the resident did well, what needs work, and how to improve.
The central hypothesis was straightforward: faculty who file more EPA evaluations might simply be more engaged in assessment overall, and that engagement might translate into higher-quality written feedback. To test it, the researchers needed an objective way to score feedback quality, which is notoriously difficult to quantify. They turned to the QuAL rubric — the Quality of Assessment for Learning instrument — which rates short workplace-based comments on a zero-to-five-point scale across three dimensions: evidence, meaning whether the comment cites specific observed behaviors; suggestion, meaning whether it offers actionable guidance; and connection, meaning whether it links the observation to the suggestion in a coherent way. The QuAL score was developed specifically to rate the kind of brief, free-text comments that dominate modern assessment platforms, giving the study a validated yardstick rather than a subjective impression of quality.
Two independent raters scored all of the narrative feedback, and the agreement between them was strong, with an interrater reliability of 0.85 — a figure that lends considerable statistical credibility to the quality ratings. In total, the team analyzed 661 EPA micro-assessments submitted by 26 faculty members. The median QuAL score across all comments was 3 on the six-point scale, and the median comment length was 60 words. Those two numbers turned out to be tightly linked: word count correlated positively with QuAL score, with a Spearman correlation coefficient of 0.81 and a p-value below 0.001. In plain terms, longer comments were substantially more likely to contain concrete evidence, useful suggestions, and a logical bridge between the two.
The volume question at the heart of the study also resolved in favor of the prolific evaluators. Faculty who submitted fewer than 25 evaluations during the study window produced feedback with a median QuAL score of 2.88, while those who submitted 25 or more reached a median of 3.25, a statistically significant difference with a p-value of 0.034. Faculty experience in the study cohort ranged widely, from one to 37 years in practice, and that variable mattered too. To disentangle the effects of volume, length, and experience — which naturally cluster together in busy senior surgeons — the researchers used regression modeling based on generalized estimating equations, a statistical approach suited to clustered data in which one faculty member contributes many assessments. The results held up: the number of evaluations submitted, the word count of the narrative feedback, and the number of years a faculty member had been in practice were each independently associated with higher-quality feedback, all with p-values below 0.05.
Why would frequent assessors write better comments? The study itself points toward targets for faculty development rather than offering a definitive psychological explanation, but the broader medical education literature suggests plausible mechanisms. Assessment is a skill, and like operative technique it improves with deliberate, repeated practice. A faculty member who completes dozens of EPA evaluations each year is repeatedly forced to articulate what competent performance looks like, to distinguish acceptable from excellent work, and to translate clinical observations into educational language. That repetition may build a kind of assessment fluency. Experience in practice likely compounds the effect: a surgeon 20 years into a career has seen a wider spectrum of resident performance and can draw on a richer mental library of specific examples, corrective strategies, and developmental trajectories when writing a comment.
The strong correlation between word count and quality deserves particular attention, because it cuts both ways. On one hand, it suggests a simple lever for improvement: programs that encourage faculty to write more substantive comments — moving beyond a terse sentence toward the 60-word-plus territory where evidence, suggestion, and connection can actually fit — may see measurable gains in feedback quality. On the other hand, length alone is not the goal. A rambling 200-word comment with no actionable suggestion would score poorly on the QuAL rubric regardless of its heft. The correlation indicates that in real-world practice, the faculty who write longer comments tend to be the ones including the specific ingredients that make feedback useful, not merely padding their text. Brevity, in the context of EPA micro-assessments, appears to come at a genuine educational cost.
The findings arrive at a moment when competency-based medical education is expanding globally and EPA-based assessment is becoming the default architecture for postgraduate training in many specialties. That transition has created an enormous appetite for faculty assessment data, and programs often measure engagement by counting submissions. This study adds an important nuance to that accounting: quantity and quality are not independent, and the faculty members who contribute the most data points may also be supplying the most educationally valuable ones. For program directors, that reframes low-volume assessors not merely as a data gap but as a quality gap — their residents may be receiving thinner, less actionable guidance precisely where more feedback structure is needed.
The practical implications for faculty development are concrete. The authors explicitly frame their results as identifying targets for improving EPA narrative feedback, and the data suggest where to aim: faculty early in their careers and faculty with low submission volumes are the groups most likely to produce lower-scoring comments. Interventions could include rater training workshops that model high-quality narrative comments, prompts embedded in assessment platforms that encourage evaluators to cite specific observed behaviors and pair them with concrete next steps, and departmental norms that treat assessment volume as a professional expectation rather than an optional courtesy. Previous work in the medical education literature has shown that faculty development programs can improve the quality of written evaluations, and this study provides the volume, length, and experience variables that such programs can monitor as markers of progress.
There are, as with any single-institution study, limits to how far the conclusions travel. The analysis covered 26 faculty members at one academic department, focused on three common surgical conditions, and relied on a six-point quality rubric that, while validated, cannot capture every dimension of what makes a comment transformative for a learner — including the trust and relationship between teacher and trainee that researchers call the educational alliance. The data underlying the study are not publicly available, and the observational design cannot prove that writing more evaluations causes better feedback, however strongly the regression suggests the association. Still, the core message is hard to ignore: in the ecosystem of surgical training, the surgeons who show up most often as assessors, write the most substantive comments, and carry the deepest wells of clinical experience are the ones producing feedback that best equips residents to grow. For a profession built on apprenticeship, that is a finding worth acting on.
Subject of Research: Quality of faculty narrative feedback in Entrustable Professional Activity assessments of general surgery residents
Article Title: Does more mean better? Evaluating faculty-to-resident feedback in general surgery
Article References: Kabbash, M. S., Fieber, J. H., Shaw, C. M., Cochran, A. A., Sarosi, G. A., & Falcone, J. L. (2026). Does more mean better? Evaluating faculty-to-resident feedback in general surgery. Global Surgical Education – Journal of the Association for Surgical Education, 5(1), Article 154. https://doi.org/10.1007/s44186-026-00563-x
Image Credits: AI Generated
DOI: 10.1007/s44186-026-00563-x
Keywords: surgical education, entrustable professional activities, feedback quality, resident assessment, competency-based medical education, general surgery, faculty development, QuAL rubric, narrative feedback, micro-assessment, medical education, workplace-based assessment
Cite Scienmag News
Courtney Benton. (October 1, 2026). More Feedback May Mean Better Feedback: Study Scores Surgeons’ Comments to Residents. Scienmag. https://scienmag.com/more-feedback-may-mean-better-feedback-study-scores-surgeons-comments-to-residents/
Courtney Benton. "More Feedback May Mean Better Feedback: Study Scores Surgeons’ Comments to Residents." Scienmag, 1 October 2026, https://scienmag.com/more-feedback-may-mean-better-feedback-study-scores-surgeons-comments-to-residents/. Accessed 1 October 2026.
Courtney Benton. "More Feedback May Mean Better Feedback: Study Scores Surgeons’ Comments to Residents." Scienmag. October 1, 2026. https://scienmag.com/more-feedback-may-mean-better-feedback-study-scores-surgeons-comments-to-residents/

