Tuesday, September 22, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Social Science

AI Turns Operating Room Talk Into Actionable Surgical Feedback

September 22, 2026
in Social Science
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 5 mins read
0
AI Turns Operating Room Talk Into Actionable Surgical Feedback

AI Turns Operating Room Talk Into Actionable Surgical Feedback

AI Turns Operating Room Talk Into Actionable Surgical Feedback

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every conversation between a surgeon and a trainee during an operation carries educational weight, yet almost none of it survives the moment. Once the drapes come down, the teaching that happened over the hum of the surgical field exists only in memory, scattered across dozens of cases and delivered in fragments between critical steps of a procedure. A new proof-of-concept study from Washington University in St. Louis suggests that this ephemeral dialogue can now be captured, analyzed, and transformed into structured feedback at a scale that was previously impossible, using an artificial intelligence system built on a large language model and carefully constrained by human oversight.

The research, published in Global Surgical Education, the Journal of the Association for Surgical Education, tackles a long-standing problem in surgical training. Intraoperative feedback is widely recognized as essential for skill acquisition, performance improvement, and progression toward independent practice, but trainees consistently report receiving less of it than they need. Surveys have shown stark gaps between what faculty believe they are teaching and what residents perceive they are learning. In one national survey, only 18 percent of residents reported that faculty helped identify personal operative goals before surgery, and only 37 percent reported discussion of areas for improvement. The operating room is a cognitively demanding environment where feedback is delivered informally and in real time, leaving little opportunity for systematic documentation or later reflection.

The verbal exchange between attending surgeons and trainees during operations represents a rich, largely untapped data source for understanding how teaching actually happens. Qualitative research could unlock this resource, but traditional qualitative methods are notoriously labor-intensive, requiring hours of manual coding and thematic analysis for every hour of recorded dialogue. That bottleneck has kept operative dialogue largely out of educational research and out of practical feedback systems. The Washington University team, led by Blake T. Beneville, Katharine E. Caldwell, and Michael M. Awad, set out to determine whether artificial intelligence could break that bottleneck without sacrificing the rigor that makes qualitative analysis meaningful.

The study team audio-recorded dialogue from 25 general and colorectal surgical operations performed between January and March 2025, using lapel microphones worn in the operating room. The procedures, including cholecystectomies, ventral and inguinal hernia repairs, and colectomies across open, laparoscopic, and robotic approaches, were chosen because they allowed trainees to perform a significant portion of the operation independently. Eight attending surgeons participated, along with colorectal fellows, minimally invasive surgery fellows, and general surgery residents across all five postgraduate year levels. Recordings were transcribed automatically, then manually reviewed by the research team to correct errors and remove all identifying information, including names and protected health details. The dialogue itself was left untouched, preserving slang, disfluencies, and the natural texture of operative conversation.

Central to the method was a 38-category thematic codebook developed through a modified grounded theory approach. The research team began with deductive codes focused on technical direction, essential team communication, and supervision, then added inductive codes from open coding of five pilot recordings to capture phenomena not anticipated in advance. The final codebook categorized every utterance into functions such as instruction, explanation, coaching, feedback, questioning, shared mental model communication, clarification, and care coordination, along with contextual elements like acknowledgments, jokes, off-target banter, instrument requests, and moments when the attending took over the operation. Feedback utterances were further characterized by sentiment and constructiveness, allowing the analysis to distinguish constructive coaching from simple praise or criticism.

For the automated analysis, the team built a custom GPT based on the GPT-4o architecture, enhanced with retrieval-augmented generation, a technique in which the model draws on reference documents rather than relying solely on pre-trained knowledge. In this case, the finalized codebook and the manually coded pilot transcripts were stored as internal references, constraining the model to apply only the study-specific codes and reducing the risk of hallucinated categories. The underlying model was not fine-tuned or retrained; instead, it was configured through explicit instructions and iteratively calibrated against investigator-coded transcripts until it reliably assigned a single code to each sentence-level line of transcript text.

The workflow was deliberately human-in-the-loop. Transcripts were fed to the AI assistant in batches of 100 lines, the maximum the interface could reliably process at the time. The model assigned codes to each line, and investigators reviewed and corrected the output in real time, adding missing codes and fixing incorrect ones before submitting the next batch. Senior researchers then adjudicated the final assignments. This end-to-end process, spanning audio capture, de-identification, transcript reformatting, AI coding, real-time correction, and senior review, constituted the complete pipeline tested in the study.

The efficiency gains were dramatic. The AI-assisted workflow required roughly 10 minutes per 100 transcript lines, compared with more than 60 minutes for manual coding, meaning the automated process took less than 20 percent of the manual effort. Across the full dataset of 28,799 transcript lines, the team estimated a time savings of 239 hours. The overall error rate for the AI-assisted workflow was just 2.0 percent, approximately 98 percent accuracy, and after initial calibration, batch-level agreement between AI-generated and investigator-confirmed codes reached 95 percent or higher, meaning reviewers needed to intervene on fewer than 5 percent of lines. Most errors clustered in boundary cases: distinguishing brief affirmations like good or okay from genuine feedback, separating interjections from off-target conversation, and occasionally misattributing teaching directed at a third party such as a medical student. These errors were caught and corrected during review and did not propagate into the final dataset.

Reliability was assessed by having the AI independently code the same five transcript segments twice in separate runs, spanning 2,245 lines and five distinct attending-trainee dyads with different teaching styles. The two runs agreed on 79 percent of lines, corresponding to substantial agreement by Cohen’s kappa of 0.75. Importantly, the researchers note, this measures the reproducibility of a non-deterministic model rather than its accuracy against a human standard; human verification remains the basis for validity. Agreement was highest for explicit utterance types and lowest for codes requiring contextual interpretation, a pattern that points toward a hybrid model in which AI performs high-throughput first-pass coding while investigators focus their attention on ambiguous segments.

The implications extend well beyond research efficiency. The authors envision a pipeline that could generate rapid post-case feedback summaries for trainees, highlighting key teaching moments with verbatim quotes, counting the feedback statements received, describing coaching strategies, and flagging missed teaching opportunities. Faculty could receive longitudinal teaching profiles showing their questioning frequency, feedback valence, and debriefing consistency, making visible how they calibrate autonomy over time, a central element of operative entrustment that is otherwise difficult to appraise. Programs could monitor learning environment features across services and case types, and future systems could map dialogue evidence to Entrustable Professional Activities, linking operative communication to competency-based decisions about supervision and readiness for practice. The team is already working on an end-to-end application that would handle secure audio capture, transcription, AI-assisted coding, human verification, and report generation automatically.

The authors are careful to frame the work as a proof of concept rather than a validated, deployable system. The evaluation was limited to general and colorectal surgery at a single institution, and other specialties would require locally calibrated codebooks and validation. Transcript-based coding captures only verbal communication, missing nonverbal cues, tone, and gestures that shape educational meaning. The workflow relied on a proprietary, closed-source model whose outputs may vary across runs, and the researchers emphasize that any deployed system would need periodic re-validation as underlying models change. They also stress that AI outputs should be treated as decision support, monitored for bias, and never as a replacement for human educational judgment. Still, the central finding stands: with an AI-supported, human-verified workflow, the fleeting conversation of the operating room can become a durable, analyzable, and actionable record, one that could strengthen trainee coaching, support faculty development, and move surgical education toward a more intentional, data-informed future.

Subject of Research: An AI-assisted qualitative coding pipeline for analyzing intraoperative attending-trainee teaching dialogue in surgery

Article Title: From OR dialogue to actionable feedback: a scalable method for analyzing intraoperative teaching

Article References: Beneville, B. T., Lewis, L., Ferzoco, M. J., Caldwell, K. E., Bennett, J., Jama, M. A., Fox, C., Tong, J., & Awad, M. M. (2026). From OR dialogue to actionable feedback: a scalable method for analyzing intraoperative teaching. Global Surgical Education – Journal of the Association for Surgical Education, 5(1), Article 162. https://doi.org/10.1007/s44186-026-00570-y

Image Credits: AI Generated

DOI: 10.1007/s44186-026-00570-y

Keywords: intraoperative teaching, surgical education, artificial intelligence, large language models, qualitative coding, operative feedback, surgical training, GPT-4o, retrieval-augmented generation, competency-based education, resident autonomy, human-in-the-loop

Cite Scienmag News

Courtney Benton. (September 22, 2026). AI Turns Operating Room Talk Into Actionable Surgical Feedback. Scienmag. https://scienmag.com/ai-turns-operating-room-talk-into-actionable-surgical-feedback/

Courtney Benton. "AI Turns Operating Room Talk Into Actionable Surgical Feedback." Scienmag, 22 September 2026, https://scienmag.com/ai-turns-operating-room-talk-into-actionable-surgical-feedback/. Accessed 22 September 2026.

Courtney Benton. "AI Turns Operating Room Talk Into Actionable Surgical Feedback." Scienmag. September 22, 2026. https://scienmag.com/ai-turns-operating-room-talk-into-actionable-surgical-feedback/

Tags: AI-based analysis of surgical conversationsAI-driven surgical training feedbackArtificial Intelligenceartificial intelligence in surgical educationcapturing and analyzing operating room dialoguecompetency-based educationGPT-4ohuman oversight in AI medical systemshuman-in-the-loopimproving surgical training through AIintraoperative communication analysisintraoperative teachinglanguage models for surgical feedbacklarge language modelsoperative feedbackqualitative codingreal-time operative skill assessmentresident autonomyretrieval-augmented generationstructured surgical coaching with AIsurgeon-trainee communication enhancementsurgical educationsurgical performance improvement toolssurgical training
Share26Tweet16
Previous Post

New Atomic-Scale Method Promises Longer-Lasting Batteries

Next Post

PKR inhibitor imoxin fails to ease diaphragm disease in male mdx mice

Related Posts

Slums Are Heating Up Fast: Satellites Reveal Nairobi, Kampala and Dar es Salaam’s Invisible Heat Crisis
Social Science

Slums Are Heating Up Fast: Satellites Reveal Nairobi, Kampala and Dar es Salaam’s Invisible Heat Crisis

September 22, 2026
Cameras Cut Crime, But Only Where Neighborhoods Let Them Work
Social Science

Cameras Cut Crime, But Only Where Neighborhoods Let Them Work

September 22, 2026
New 3C Model Aims to Fix How Teachers Are Trained to Teach Coding
Social Science

New 3C Model Aims to Fix How Teachers Are Trained to Teach Coding

September 22, 2026
New Index Reveals How Similar EU Nations Really Are on Gender Equality
Social Science

New Index Reveals How Similar EU Nations Really Are on Gender Equality

September 22, 2026
Artificial Intelligence Is Rewriting How Schools Teach, Test and Govern
Social Science

Artificial Intelligence Is Rewriting How Schools Teach, Test and Govern

September 22, 2026
Reporter game therapy technique quiets sensory brain responses in autistic youth
Social Science

Reporter game therapy technique quiets sensory brain responses in autistic youth

September 22, 2026
Next Post
PKR inhibitor imoxin fails to ease diaphragm disease in male mdx mice

PKR inhibitor imoxin fails to ease diaphragm disease in male mdx mice

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • PKR inhibitor imoxin fails to ease diaphragm disease in male mdx mice
  • AI Turns Operating Room Talk Into Actionable Surgical Feedback
  • New Atomic-Scale Method Promises Longer-Lasting Batteries
  • Why Young Adults Skip Health Apps: A New Model Reveals What Makes eHealth Stick

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading