Every surgeon remembers the moment they stopped being watched and started being trusted. A new pilot study suggests that moment may be audible in the operating room itself, written into the very pattern of words exchanged between an attending surgeon and a trainee. Researchers who recorded 25 real general surgery operations found that the language flowing between teacher and learner changes in measurable, predictable ways as a resident gains operative independence, raising the possibility that a microphone could one day do what subjective end-of-case evaluations have long struggled to do: capture surgical readiness as it actually happens.
The study, led by Katharine E. Caldwell of the Medical University of South Carolina with colleagues at Washington University in Saint Louis and Stanford University, and published in Global Surgical Education, the Journal of the Association for Surgical Education, set out to solve a stubborn problem in surgical training. Modern programs are shifting toward competency-based frameworks built around Entrustable Professional Activities, or EPAs, which are designed to judge whether a trainee is ready for independent practice. Yet EPA ratings still depend on retrospective, written assessments completed after the operation ends. Those evaluations can suffer from recall bias, incomplete case capture, and documented racial and sex-based evaluator bias, and they may miss the subtle, minute-to-minute dynamics that define how a trainee actually performs under pressure.
The researchers’ hypothesis was elegantly simple: the operating room conversation is itself a data stream. When a trainee is struggling, the attending speaks more, directs more, and takes over more often. When a trainee is ready to operate with only indirect supervision, the conversational balance flips. To test this, the team equipped attending surgeon and trainee dyads with lapel microphones during general surgery operations across three divisions: minimally invasive surgery, colorectal surgery, and surgical oncology. Recordings began immediately after the surgical time out and ended at skin closure, protecting patient privacy while capturing the full operative dialogue. Procedures included cholecystectomies, ventral and inguinal hernia repairs, small bowel and gastric resections, and colectomies, performed open, laparoscopically, and robotically, with a mean operative duration of 130.6 minutes.
From those recordings, the team generated an enormous corpus: 34,392 individual utterances. To make sense of it, they developed a framework of 38 unique operative linguistic codes using a modified grounded theory approach. Codes captured technical instruction, instrument requests, feedback, takeovers, off-target talking, and shared mental modelling, the practice of verbally indicating important anatomy, planning the next steps, or voicing uncertainty. Some codes were split by direction or valence, so feedback could be positive or negative, and takeovers could be classified as completion, demonstration, or safety events. Utterances deemed ordinary conversational filler, such as simple acknowledgements or clarifying questions, were excluded from analysis.
Artificial intelligence entered the workflow as an accelerant, not an arbiter. The researchers initialized a large language model, ChatGPT-4o, with the codebook and anchor examples, then ran a supervised training phase across 3,000 transcript lines in ten-line increments, with a human researcher reviewing and correcting each categorization. In a final layer of quality control, a human investigator reviewed 100 percent of the AI-generated codes, correcting errors that clustered around boundary cases, such as distinguishing off-target chatter with a third party from genuine explanation, or deciding whether a brief utterance like good or okay was feedback or mere agreement. Every final code in the study was assigned by a human, and the authors are explicit that AI-only coding remains unvalidated for this purpose and would require far more data and improved model accuracy before fully automated analysis becomes feasible.
After each operation, the attending rated the trainee on an EPA-based entrustability scale ranging from Level 1, limited participation, to Level 4, practice-ready. For analysis, cases were divided into lower-autonomy cases, where the trainee needed direct supervision, and higher-autonomy cases, where the trainee operated under indirect supervision or was deemed practice-ready. The linguistic contrasts between the two groups were striking. More autonomous trainees generated 28.4 percent of the words spoken during a case, compared with just 10.8 percent for trainees under direct supervision, a difference that held even after normalizing for case length and individual speaking rate. Talk, in other words, tracks trust.
The content of trainee speech shifted as sharply as its volume. Higher-autonomy trainees initiated 40.7 percent of instrument requests versus 17.8 percent in the lower-autonomy group, and they delivered 21.8 percent of technical instruction directed at the attending surgeon, compared with a mere 1.4 percent among less independent learners. They also led a dramatically larger share of shared mental modelling, 54.5 percent versus 16.8 percent, meaning they were the ones calling out anatomy, proposing the next operative step, and articulating uncertainty. Takeover events, in which the attending steps in to complete, demonstrate, or secure a critical maneuver, fell from an average of 7.6 per case in the direct supervision group to 0.6 per case among more autonomous trainees. In the lower-autonomy group, the majority of takeovers were completions, the attending finishing what the trainee could not, whereas among indirectly supervised trainees, takeovers more often took the form of demonstration.
The attending surgeons’ language told the complementary story. When operating with highly trusted trainees, attendings engaged in significantly more off-target talking, 31.7 percent of utterances versus 8.7 percent, conversation unconnected to the immediate operative task, a behavioral signature of reduced need for continuous coaching. In lower-autonomy cases, such chatter was largely confined to the opening and closing of the case, vanishing during critical operative portions when every word carried weight. Total feedback volume was similar across groups, at 2.1 versus 1.5 percent of utterances, but the valence shifted decisively: attendings delivered 63.8 percent of feedback as positive in higher-autonomy cases, compared with just 24.0 percent in lower-autonomy ones. Technical feedback dominated overall, accounting for 85.3 percent of all feedback given.
The authors are careful to frame this as a pilot with real limitations. It was a single-center study of 25 operations at a large Midwestern academic medical center; all the attending surgeons were male, nearly half the trainees were fellows, and all procedures were common general surgery operations with complex cases deliberately excluded. The team could not analyze how race or gender shaped communication patterns, despite prior evidence that these factors influence feedback dynamics in surgical teams, and they did not control for familiarity between attending and trainee, which is known to affect team performance and entrustability. The Hawthorne effect looms as well: participants knew they were being recorded and may have altered their speech, prompting the group to investigate less invasive black box style recording technologies for future work.
Even so, the implications are considerable. If operative dialogue can be captured and coded at scale, every case could yield an objective, behavior-based supplement to EPA ratings, giving trainees individualized performance data and giving faculty a mirror for their own teaching styles, potentially transforming faculty development alongside trainee assessment. The researchers plan multicenter validation, integration with existing competency frameworks, and studies linking linguistic markers to real-time operative performance metrics. For now, the study’s most provocative message is conceptual: surgical autonomy is not just something evaluators imagine after the fact, but something audible in real time, one utterance at a time. The operating room, it turns out, has been telling us who is ready all along.
Subject of Research: Using live operative audio recordings and linguistic analysis to evaluate surgical resident autonomy
Article Title: How we talk and teach in the operating room: using live operative recordings to evaluate resident autonomy
Article References: Caldwell, K. E., Beneville, B. T., Bennett, J., Jama, M. A., Fox, C., Ferzoco, M., Lewis, L., Tong, J., & Awad, M. M. (2026). How we talk and teach in the operating room: using live operative recordings to evaluate resident autonomy. Global Surgical Education – Journal of the Association for Surgical Education, 5(1), Article 179. https://doi.org/10.1007/s44186-026-00584-6
Image Credits: AI Generated
DOI: 10.1007/s44186-026-00584-6
Keywords: surgical education, resident autonomy, operating room communication, Entrustable Professional Activities, operative linguistics, surgical training, feedback, shared mental modelling, AI transcription, competency-based assessment, general surgery, entrustability
Cite Scienmag News
Courtney Benton. (September 20, 2026). Surgeons’ Words in the Operating Room Reveal When Residents Are Ready to Fly Solo. Scienmag. https://scienmag.com/surgeons-words-in-the-operating-room-reveal-when-residents-are-ready-to-fly-solo/
Courtney Benton. "Surgeons’ Words in the Operating Room Reveal When Residents Are Ready to Fly Solo." Scienmag, 20 September 2026, https://scienmag.com/surgeons-words-in-the-operating-room-reveal-when-residents-are-ready-to-fly-solo/. Accessed 20 September 2026.
Courtney Benton. "Surgeons’ Words in the Operating Room Reveal When Residents Are Ready to Fly Solo." Scienmag. September 20, 2026. https://scienmag.com/surgeons-words-in-the-operating-room-reveal-when-residents-are-ready-to-fly-solo/

