<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>resident autonomy &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/resident-autonomy/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 23:39:29 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>resident autonomy &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>How Surgeons Decide to Trust Residents at the Robotic Console</title>
		<link>https://scienmag.com/how-surgeons-decide-to-trust-residents-at-the-robotic-console/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 23:39:29 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[barriers to resident autonomy in robotic procedures]]></category>
		<category><![CDATA[competency-based assessment]]></category>
		<category><![CDATA[development of robotic portfolios for residents]]></category>
		<category><![CDATA[dual console]]></category>
		<category><![CDATA[entrustment]]></category>
		<category><![CDATA[faculty perspectives on robotic surgery]]></category>
		<category><![CDATA[general surgery residency]]></category>
		<category><![CDATA[impact of robotic technology on surgical training]]></category>
		<category><![CDATA[operative autonomy]]></category>
		<category><![CDATA[patient outcomes in robotic surgery]]></category>
		<category><![CDATA[qualitative research]]></category>
		<category><![CDATA[qualitative study in surgical education]]></category>
		<category><![CDATA[resident autonomy]]></category>
		<category><![CDATA[resident skill assessment in robotic surgery]]></category>
		<category><![CDATA[robotic console decision-making]]></category>
		<category><![CDATA[Robotic surgery]]></category>
		<category><![CDATA[robotic surgery training]]></category>
		<category><![CDATA[surgeon trust in residents]]></category>
		<category><![CDATA[surgical education]]></category>
		<category><![CDATA[surgical education and mentorship]]></category>
		<category><![CDATA[surgical subspecialties and robotic procedures]]></category>
		<category><![CDATA[surgical training]]></category>
		<category><![CDATA[telestration]]></category>
		<category><![CDATA[thematic analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224338</guid>

					<description><![CDATA[A qualitative study of thirteen robotic surgery faculty reveals that attendings weigh a pre-test probability of resident readiness, that a documented robotic portfolio raises their willingness to entrust trainees, and that direct observation at the console remains essential.]]></description>
										<content:encoded><![CDATA[<p>Robotic surgery has quietly transformed the modern operating room, but it has also created an uncomfortable paradox for surgical education. The very machines that give patients smaller incisions and faster recoveries can leave trainees watching from the periphery, their hands hovering over bedside instruments rather than resting on the console controls. A new qualitative study published in Global Surgical Education, the journal of the Association for Surgical Education, takes a close look at why attending surgeons hesitate to hand over the robotic console to general surgery residents, and whether a simple documentation tool called a robotic portfolio might change the calculus of trust.</p>
<p>The research team, led by Laura Washburn of the University of Pittsburgh Medical Center together with colleagues at the University of Pittsburgh School of Medicine and The Ohio State University Wexner Medical Center, interviewed thirteen faculty surgeons who regularly operate robotically alongside general surgery trainees. Using purposive sampling, the investigators deliberately recruited attending surgeons across different subspecialties and career stages, ensuring that the findings would not simply reflect the habits of one narrow surgical niche. The most common robotic case types represented in the faculty&#8217;s practices were hernia repairs, followed by colorectal and bariatric procedures, a mix that mirrors the bread-and-butter robotic workload at many academic centers.</p>
<p>The methodological approach was deliberately open-ended. Semi-structured interviews allowed the faculty to describe in their own words how they decide, case by case, whether a resident is ready to take the console. Transcripts were analyzed verbatim using inductive thematic analysis, the widely used qualitative technique associated with Braun and Clarke, in which themes emerge from the data rather than being imposed by a pre-existing framework. This matters because entrustment is not a purely technical judgment; it is a human decision shaped by experience, context, and intuition, and only an open-ended method can surface those subtleties.</p>
<p>What emerged was a striking metaphor borrowed from clinical medicine: entrustment as a kind of pre-test probability. Just as a clinician estimates the likelihood of disease before ordering a diagnostic test, attending surgeons described forming an internal estimate of a resident&#8217;s readiness before ever stepping into the operating room. Factors feeding that estimate include the trainee&#8217;s prior operative experiences, their reputation among colleagues, their demonstrated initiative, and the complexity and risk profile of the planned case. The higher this pre-test probability, the more willing the attending is to grant autonomy from the first port placement to the final closure.</p>
<p>But the pre-test probability is only the opening bid. The study found that faculty consistently insist on validating their initial estimate through direct observation before fully entrusting a resident with the robotic console. No matter how impressive a resident&#8217;s documented experience looks on paper, attendings want to see the trainee&#8217;s hands move, watch how they handle tissue, and gauge their responses to unexpected findings in real time. This validation step is a safeguard rooted in patient safety, and it explains why objective credentials alone rarely unlock autonomy; they open the door, but only demonstrated performance walks the trainee through it.</p>
<p>It is precisely at this juncture that the robotic portfolio enters the story. The researchers developed a sample resident portfolio containing objective details about the trainee&#8217;s robotic experiences, including case volumes, console time, and procedural roles. When faculty reviewed the portfolio during the interviews, they described it as genuinely informative, and several reported that it raised their pre-test probability of prospective entrustment. In practical terms, the portfolio functions like a well-documented history in clinical reasoning: it does not replace the diagnostic test of direct observation, but it meaningfully shifts the starting point, making attendings more inclined to plan for resident console time rather than default to supervision.</p>
<p>The robotic platform itself emerged as a double-edged sword in the entrustment equation. On the promoting side, faculty highlighted the dual console, which lets attending and trainee sit at linked stations and swap control of the instruments instantly, and telestration, the technique of drawing on the video feed to guide the trainee&#8217;s next move. These features create teaching moments that are difficult or impossible in traditional laparoscopy, where the attending and trainee share a single camera and awkward instrument exchanges. The dual console in particular lowers the perceived risk of granting autonomy, because the attending can reclaim control within a fraction of a second if the dissection strays into danger.</p>
<p>On the challenging side, the study participants emphasized that residents must adapt to the robotic console itself, learning to interpret tactile visual cues in an environment where their hands never touch tissue. The robotic system translates hand movements into scaled instrument motions and filters out natural tremor, but it also strips away the haptic feedback that open and laparoscopic surgeons rely on. Trainees must learn to read tissue resistance through visual deformation, instrument interaction, and subtle changes in the operative field. Faculty described this perceptual recalibration as a genuine learning curve, one that complicates simple judgments about how much console experience should translate into entrustment.</p>
<p>The implications reach well beyond the two academic centers studied. Robotic surgery training has documented barriers to resident console time, and prior work has shown that trainees often struggle with the transition from bedside assistant to console surgeon. The findings suggest that competency-based, rather than time-based, progression is achievable if programs give faculty better information about what each resident has actually done. A portfolio that aggregates console hours, completed procedures, simulator performance, and bedside roles could standardize the conversation between attending and trainee before the case begins, replacing vague impressions with shared data. At the same time, the study is a caution against treating any document as a substitute for observation; faculty in this research remained clear that validation at the console is non-negotiable.</p>
<p>For surgical educators, the roadmap is now more concrete. Programs adopting robotic portfolios should pair them with structured opportunities for direct observation, deliberate use of dual-console teaching, and telestration-based coaching tailored to the perceptual demands of the robotic environment. For residents, the message is that documented experience builds the attending&#8217;s pre-test probability, but earning full entrustment still requires demonstrating skill when the attending is watching. And for patients, the reassurance is that the system of graduated autonomy, with all its built-in skepticism, remains firmly anchored in safety. As robotic platforms proliferate across hospitals worldwide, understanding the psychology of surgical trust may prove as important as the technology itself, and this study offers one of the clearest portraits yet of how that trust is built, tested, and ultimately granted at the console.</p>
<p><strong>Subject of Research:</strong> Faculty entrustment of general surgery residents in robotic surgery and the use of a resident robotic portfolio to support prospective entrustment decisions</p>
<p><strong>Article Title:</strong> Entrustment of general surgery residents in robotic surgery and utility of a robotic portfolio to promote prospective entrustment</p>
<p><strong>Article References:</strong> Entrustment of general surgery residents in robotic surgery and utility of a robotic portfolio to promote prospective entrustment. (n.d.). <a href="https://doi.org/10.1007/s44186-026-00555-x" rel="noopener noreferrer">https://doi.org/10.1007/s44186-026-00555-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44186-026-00555-x" rel="noopener noreferrer">10.1007/s44186-026-00555-x</a></p>
<p><strong>Keywords:</strong> robotic surgery, surgical education, resident autonomy, entrustment, general surgery residency, competency-based assessment, dual console, telestration, qualitative research, thematic analysis, surgical training, operative autonomy</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224338</post-id>	</item>
		<item>
		<title>AI Turns Operating Room Talk Into Actionable Surgical Feedback</title>
		<link>https://scienmag.com/ai-turns-operating-room-talk-into-actionable-surgical-feedback/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 16:37:49 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI-based analysis of surgical conversations]]></category>
		<category><![CDATA[AI-driven surgical training feedback]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[artificial intelligence in surgical education]]></category>
		<category><![CDATA[capturing and analyzing operating room dialogue]]></category>
		<category><![CDATA[competency-based education]]></category>
		<category><![CDATA[GPT-4o]]></category>
		<category><![CDATA[human oversight in AI medical systems]]></category>
		<category><![CDATA[human-in-the-loop]]></category>
		<category><![CDATA[improving surgical training through AI]]></category>
		<category><![CDATA[intraoperative communication analysis]]></category>
		<category><![CDATA[intraoperative teaching]]></category>
		<category><![CDATA[language models for surgical feedback]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[operative feedback]]></category>
		<category><![CDATA[qualitative coding]]></category>
		<category><![CDATA[real-time operative skill assessment]]></category>
		<category><![CDATA[resident autonomy]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[structured surgical coaching with AI]]></category>
		<category><![CDATA[surgeon-trainee communication enhancement]]></category>
		<category><![CDATA[surgical education]]></category>
		<category><![CDATA[surgical performance improvement tools]]></category>
		<category><![CDATA[surgical training]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=206875</guid>

					<description><![CDATA[Researchers at Washington University developed an AI-assisted, human-verified coding pipeline that reduced the time needed to analyze intraoperative teaching dialogue by more than 80 percent while achieving 98 percent accuracy, turning transient operating room conversation into structured, actionable feedback for surgical trainees and faculty.]]></description>
										<content:encoded><![CDATA[<p>Every conversation between a surgeon and a trainee during an operation carries educational weight, yet almost none of it survives the moment. Once the drapes come down, the teaching that happened over the hum of the surgical field exists only in memory, scattered across dozens of cases and delivered in fragments between critical steps of a procedure. A new proof-of-concept study from Washington University in St. Louis suggests that this ephemeral dialogue can now be captured, analyzed, and transformed into structured feedback at a scale that was previously impossible, using an artificial intelligence system built on a large language model and carefully constrained by human oversight.</p>
<p>The research, published in Global Surgical Education, the Journal of the Association for Surgical Education, tackles a long-standing problem in surgical training. Intraoperative feedback is widely recognized as essential for skill acquisition, performance improvement, and progression toward independent practice, but trainees consistently report receiving less of it than they need. Surveys have shown stark gaps between what faculty believe they are teaching and what residents perceive they are learning. In one national survey, only 18 percent of residents reported that faculty helped identify personal operative goals before surgery, and only 37 percent reported discussion of areas for improvement. The operating room is a cognitively demanding environment where feedback is delivered informally and in real time, leaving little opportunity for systematic documentation or later reflection.</p>
<p>The verbal exchange between attending surgeons and trainees during operations represents a rich, largely untapped data source for understanding how teaching actually happens. Qualitative research could unlock this resource, but traditional qualitative methods are notoriously labor-intensive, requiring hours of manual coding and thematic analysis for every hour of recorded dialogue. That bottleneck has kept operative dialogue largely out of educational research and out of practical feedback systems. The Washington University team, led by Blake T. Beneville, Katharine E. Caldwell, and Michael M. Awad, set out to determine whether artificial intelligence could break that bottleneck without sacrificing the rigor that makes qualitative analysis meaningful.</p>
<p>The study team audio-recorded dialogue from 25 general and colorectal surgical operations performed between January and March 2025, using lapel microphones worn in the operating room. The procedures, including cholecystectomies, ventral and inguinal hernia repairs, and colectomies across open, laparoscopic, and robotic approaches, were chosen because they allowed trainees to perform a significant portion of the operation independently. Eight attending surgeons participated, along with colorectal fellows, minimally invasive surgery fellows, and general surgery residents across all five postgraduate year levels. Recordings were transcribed automatically, then manually reviewed by the research team to correct errors and remove all identifying information, including names and protected health details. The dialogue itself was left untouched, preserving slang, disfluencies, and the natural texture of operative conversation.</p>
<p>Central to the method was a 38-category thematic codebook developed through a modified grounded theory approach. The research team began with deductive codes focused on technical direction, essential team communication, and supervision, then added inductive codes from open coding of five pilot recordings to capture phenomena not anticipated in advance. The final codebook categorized every utterance into functions such as instruction, explanation, coaching, feedback, questioning, shared mental model communication, clarification, and care coordination, along with contextual elements like acknowledgments, jokes, off-target banter, instrument requests, and moments when the attending took over the operation. Feedback utterances were further characterized by sentiment and constructiveness, allowing the analysis to distinguish constructive coaching from simple praise or criticism.</p>
<p>For the automated analysis, the team built a custom GPT based on the GPT-4o architecture, enhanced with retrieval-augmented generation, a technique in which the model draws on reference documents rather than relying solely on pre-trained knowledge. In this case, the finalized codebook and the manually coded pilot transcripts were stored as internal references, constraining the model to apply only the study-specific codes and reducing the risk of hallucinated categories. The underlying model was not fine-tuned or retrained; instead, it was configured through explicit instructions and iteratively calibrated against investigator-coded transcripts until it reliably assigned a single code to each sentence-level line of transcript text.</p>
<p>The workflow was deliberately human-in-the-loop. Transcripts were fed to the AI assistant in batches of 100 lines, the maximum the interface could reliably process at the time. The model assigned codes to each line, and investigators reviewed and corrected the output in real time, adding missing codes and fixing incorrect ones before submitting the next batch. Senior researchers then adjudicated the final assignments. This end-to-end process, spanning audio capture, de-identification, transcript reformatting, AI coding, real-time correction, and senior review, constituted the complete pipeline tested in the study.</p>
<p>The efficiency gains were dramatic. The AI-assisted workflow required roughly 10 minutes per 100 transcript lines, compared with more than 60 minutes for manual coding, meaning the automated process took less than 20 percent of the manual effort. Across the full dataset of 28,799 transcript lines, the team estimated a time savings of 239 hours. The overall error rate for the AI-assisted workflow was just 2.0 percent, approximately 98 percent accuracy, and after initial calibration, batch-level agreement between AI-generated and investigator-confirmed codes reached 95 percent or higher, meaning reviewers needed to intervene on fewer than 5 percent of lines. Most errors clustered in boundary cases: distinguishing brief affirmations like good or okay from genuine feedback, separating interjections from off-target conversation, and occasionally misattributing teaching directed at a third party such as a medical student. These errors were caught and corrected during review and did not propagate into the final dataset.</p>
<p>Reliability was assessed by having the AI independently code the same five transcript segments twice in separate runs, spanning 2,245 lines and five distinct attending-trainee dyads with different teaching styles. The two runs agreed on 79 percent of lines, corresponding to substantial agreement by Cohen&#8217;s kappa of 0.75. Importantly, the researchers note, this measures the reproducibility of a non-deterministic model rather than its accuracy against a human standard; human verification remains the basis for validity. Agreement was highest for explicit utterance types and lowest for codes requiring contextual interpretation, a pattern that points toward a hybrid model in which AI performs high-throughput first-pass coding while investigators focus their attention on ambiguous segments.</p>
<p>The implications extend well beyond research efficiency. The authors envision a pipeline that could generate rapid post-case feedback summaries for trainees, highlighting key teaching moments with verbatim quotes, counting the feedback statements received, describing coaching strategies, and flagging missed teaching opportunities. Faculty could receive longitudinal teaching profiles showing their questioning frequency, feedback valence, and debriefing consistency, making visible how they calibrate autonomy over time, a central element of operative entrustment that is otherwise difficult to appraise. Programs could monitor learning environment features across services and case types, and future systems could map dialogue evidence to Entrustable Professional Activities, linking operative communication to competency-based decisions about supervision and readiness for practice. The team is already working on an end-to-end application that would handle secure audio capture, transcription, AI-assisted coding, human verification, and report generation automatically.</p>
<p>The authors are careful to frame the work as a proof of concept rather than a validated, deployable system. The evaluation was limited to general and colorectal surgery at a single institution, and other specialties would require locally calibrated codebooks and validation. Transcript-based coding captures only verbal communication, missing nonverbal cues, tone, and gestures that shape educational meaning. The workflow relied on a proprietary, closed-source model whose outputs may vary across runs, and the researchers emphasize that any deployed system would need periodic re-validation as underlying models change. They also stress that AI outputs should be treated as decision support, monitored for bias, and never as a replacement for human educational judgment. Still, the central finding stands: with an AI-supported, human-verified workflow, the fleeting conversation of the operating room can become a durable, analyzable, and actionable record, one that could strengthen trainee coaching, support faculty development, and move surgical education toward a more intentional, data-informed future.</p>
<p><strong>Subject of Research:</strong> An AI-assisted qualitative coding pipeline for analyzing intraoperative attending-trainee teaching dialogue in surgery</p>
<p><strong>Article Title:</strong> From OR dialogue to actionable feedback: a scalable method for analyzing intraoperative teaching</p>
<p><strong>Article References:</strong> Beneville, B. T., Lewis, L., Ferzoco, M. J., Caldwell, K. E., Bennett, J., Jama, M. A., Fox, C., Tong, J., &amp; Awad, M. M. (2026). From OR dialogue to actionable feedback: a scalable method for analyzing intraoperative teaching. <em>Global Surgical Education &#8211; Journal of the Association for Surgical Education, 5</em>(1), Article 162. <a href="https://doi.org/10.1007/s44186-026-00570-y" rel="noopener noreferrer">https://doi.org/10.1007/s44186-026-00570-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44186-026-00570-y" rel="noopener noreferrer">10.1007/s44186-026-00570-y</a></p>
<p><strong>Keywords:</strong> intraoperative teaching, surgical education, artificial intelligence, large language models, qualitative coding, operative feedback, surgical training, GPT-4o, retrieval-augmented generation, competency-based education, resident autonomy, human-in-the-loop</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">206875</post-id>	</item>
		<item>
		<title>Surgeons&#8217; Words in the Operating Room Reveal When Residents Are Ready to Fly Solo</title>
		<link>https://scienmag.com/surgeons-words-in-the-operating-room-reveal-when-residents-are-ready-to-fly-solo/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:14:02 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI transcription]]></category>
		<category><![CDATA[competency-based assessment]]></category>
		<category><![CDATA[competency-based surgical education]]></category>
		<category><![CDATA[entrustability]]></category>
		<category><![CDATA[Entrustable Professional Activities]]></category>
		<category><![CDATA[entrustment decision-making in surgery]]></category>
		<category><![CDATA[feedback]]></category>
		<category><![CDATA[general surgery]]></category>
		<category><![CDATA[language analysis in operating rooms]]></category>
		<category><![CDATA[microphone-based surgical skill evaluation]]></category>
		<category><![CDATA[operating room communication]]></category>
		<category><![CDATA[operative independence indicators]]></category>
		<category><![CDATA[operative linguistics]]></category>
		<category><![CDATA[predictive analysis of surgical performance]]></category>
		<category><![CDATA[real-time surgical readiness measurement]]></category>
		<category><![CDATA[resident autonomy]]></category>
		<category><![CDATA[resident autonomy in surgery]]></category>
		<category><![CDATA[shared mental modelling]]></category>
		<category><![CDATA[surgeon-trainee communication patterns]]></category>
		<category><![CDATA[surgical case evaluation methods]]></category>
		<category><![CDATA[surgical education]]></category>
		<category><![CDATA[surgical education research]]></category>
		<category><![CDATA[surgical training]]></category>
		<category><![CDATA[Surgical training assessment]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202055</guid>

					<description><![CDATA[Researchers recorded 25 real operations and found that measurable speech patterns between attending surgeons and trainees closely track operative autonomy and entrustability.]]></description>
										<content:encoded><![CDATA[<p>Every surgeon remembers the moment they stopped being watched and started being trusted. A new pilot study suggests that moment may be audible in the operating room itself, written into the very pattern of words exchanged between an attending surgeon and a trainee. Researchers who recorded 25 real general surgery operations found that the language flowing between teacher and learner changes in measurable, predictable ways as a resident gains operative independence, raising the possibility that a microphone could one day do what subjective end-of-case evaluations have long struggled to do: capture surgical readiness as it actually happens.</p>
<p>The study, led by Katharine E. Caldwell of the Medical University of South Carolina with colleagues at Washington University in Saint Louis and Stanford University, and published in Global Surgical Education, the Journal of the Association for Surgical Education, set out to solve a stubborn problem in surgical training. Modern programs are shifting toward competency-based frameworks built around Entrustable Professional Activities, or EPAs, which are designed to judge whether a trainee is ready for independent practice. Yet EPA ratings still depend on retrospective, written assessments completed after the operation ends. Those evaluations can suffer from recall bias, incomplete case capture, and documented racial and sex-based evaluator bias, and they may miss the subtle, minute-to-minute dynamics that define how a trainee actually performs under pressure.</p>
<p>The researchers&#8217; hypothesis was elegantly simple: the operating room conversation is itself a data stream. When a trainee is struggling, the attending speaks more, directs more, and takes over more often. When a trainee is ready to operate with only indirect supervision, the conversational balance flips. To test this, the team equipped attending surgeon and trainee dyads with lapel microphones during general surgery operations across three divisions: minimally invasive surgery, colorectal surgery, and surgical oncology. Recordings began immediately after the surgical time out and ended at skin closure, protecting patient privacy while capturing the full operative dialogue. Procedures included cholecystectomies, ventral and inguinal hernia repairs, small bowel and gastric resections, and colectomies, performed open, laparoscopically, and robotically, with a mean operative duration of 130.6 minutes.</p>
<p>From those recordings, the team generated an enormous corpus: 34,392 individual utterances. To make sense of it, they developed a framework of 38 unique operative linguistic codes using a modified grounded theory approach. Codes captured technical instruction, instrument requests, feedback, takeovers, off-target talking, and shared mental modelling, the practice of verbally indicating important anatomy, planning the next steps, or voicing uncertainty. Some codes were split by direction or valence, so feedback could be positive or negative, and takeovers could be classified as completion, demonstration, or safety events. Utterances deemed ordinary conversational filler, such as simple acknowledgements or clarifying questions, were excluded from analysis.</p>
<p>Artificial intelligence entered the workflow as an accelerant, not an arbiter. The researchers initialized a large language model, ChatGPT-4o, with the codebook and anchor examples, then ran a supervised training phase across 3,000 transcript lines in ten-line increments, with a human researcher reviewing and correcting each categorization. In a final layer of quality control, a human investigator reviewed 100 percent of the AI-generated codes, correcting errors that clustered around boundary cases, such as distinguishing off-target chatter with a third party from genuine explanation, or deciding whether a brief utterance like good or okay was feedback or mere agreement. Every final code in the study was assigned by a human, and the authors are explicit that AI-only coding remains unvalidated for this purpose and would require far more data and improved model accuracy before fully automated analysis becomes feasible.</p>
<p>After each operation, the attending rated the trainee on an EPA-based entrustability scale ranging from Level 1, limited participation, to Level 4, practice-ready. For analysis, cases were divided into lower-autonomy cases, where the trainee needed direct supervision, and higher-autonomy cases, where the trainee operated under indirect supervision or was deemed practice-ready. The linguistic contrasts between the two groups were striking. More autonomous trainees generated 28.4 percent of the words spoken during a case, compared with just 10.8 percent for trainees under direct supervision, a difference that held even after normalizing for case length and individual speaking rate. Talk, in other words, tracks trust.</p>
<p>The content of trainee speech shifted as sharply as its volume. Higher-autonomy trainees initiated 40.7 percent of instrument requests versus 17.8 percent in the lower-autonomy group, and they delivered 21.8 percent of technical instruction directed at the attending surgeon, compared with a mere 1.4 percent among less independent learners. They also led a dramatically larger share of shared mental modelling, 54.5 percent versus 16.8 percent, meaning they were the ones calling out anatomy, proposing the next operative step, and articulating uncertainty. Takeover events, in which the attending steps in to complete, demonstrate, or secure a critical maneuver, fell from an average of 7.6 per case in the direct supervision group to 0.6 per case among more autonomous trainees. In the lower-autonomy group, the majority of takeovers were completions, the attending finishing what the trainee could not, whereas among indirectly supervised trainees, takeovers more often took the form of demonstration.</p>
<p>The attending surgeons&#8217; language told the complementary story. When operating with highly trusted trainees, attendings engaged in significantly more off-target talking, 31.7 percent of utterances versus 8.7 percent, conversation unconnected to the immediate operative task, a behavioral signature of reduced need for continuous coaching. In lower-autonomy cases, such chatter was largely confined to the opening and closing of the case, vanishing during critical operative portions when every word carried weight. Total feedback volume was similar across groups, at 2.1 versus 1.5 percent of utterances, but the valence shifted decisively: attendings delivered 63.8 percent of feedback as positive in higher-autonomy cases, compared with just 24.0 percent in lower-autonomy ones. Technical feedback dominated overall, accounting for 85.3 percent of all feedback given.</p>
<p>The authors are careful to frame this as a pilot with real limitations. It was a single-center study of 25 operations at a large Midwestern academic medical center; all the attending surgeons were male, nearly half the trainees were fellows, and all procedures were common general surgery operations with complex cases deliberately excluded. The team could not analyze how race or gender shaped communication patterns, despite prior evidence that these factors influence feedback dynamics in surgical teams, and they did not control for familiarity between attending and trainee, which is known to affect team performance and entrustability. The Hawthorne effect looms as well: participants knew they were being recorded and may have altered their speech, prompting the group to investigate less invasive black box style recording technologies for future work.</p>
<p>Even so, the implications are considerable. If operative dialogue can be captured and coded at scale, every case could yield an objective, behavior-based supplement to EPA ratings, giving trainees individualized performance data and giving faculty a mirror for their own teaching styles, potentially transforming faculty development alongside trainee assessment. The researchers plan multicenter validation, integration with existing competency frameworks, and studies linking linguistic markers to real-time operative performance metrics. For now, the study&#8217;s most provocative message is conceptual: surgical autonomy is not just something evaluators imagine after the fact, but something audible in real time, one utterance at a time. The operating room, it turns out, has been telling us who is ready all along.</p>
<p><strong>Subject of Research:</strong> Using live operative audio recordings and linguistic analysis to evaluate surgical resident autonomy</p>
<p><strong>Article Title:</strong> How we talk and teach in the operating room: using live operative recordings to evaluate resident autonomy</p>
<p><strong>Article References:</strong> Caldwell, K. E., Beneville, B. T., Bennett, J., Jama, M. A., Fox, C., Ferzoco, M., Lewis, L., Tong, J., &amp; Awad, M. M. (2026). How we talk and teach in the operating room: using live operative recordings to evaluate resident autonomy. <em>Global Surgical Education &#8211; Journal of the Association for Surgical Education, 5</em>(1), Article 179. <a href="https://doi.org/10.1007/s44186-026-00584-6" rel="noopener noreferrer">https://doi.org/10.1007/s44186-026-00584-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44186-026-00584-6" rel="noopener noreferrer">10.1007/s44186-026-00584-6</a></p>
<p><strong>Keywords:</strong> surgical education, resident autonomy, operating room communication, Entrustable Professional Activities, operative linguistics, surgical training, feedback, shared mental modelling, AI transcription, competency-based assessment, general surgery, entrustability</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202055</post-id>	</item>
	</channel>
</rss>
