<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>entrustability &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/entrustability/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 09:35:24 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>entrustability &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores</title>
		<link>https://scienmag.com/new-scoring-algorithm-turns-routine-surgical-ratings-into-fairer-resident-entrustability-scores/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 09:35:24 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[assessment noise reduction]]></category>
		<category><![CDATA[competency-based medical education]]></category>
		<category><![CDATA[entrustability]]></category>
		<category><![CDATA[Entrustable Professional Activities]]></category>
		<category><![CDATA[entrustable professional activities (EPAs)]]></category>
		<category><![CDATA[entrustment scoring]]></category>
		<category><![CDATA[general surgery residency]]></category>
		<category><![CDATA[intraoperative assessment]]></category>
		<category><![CDATA[Multilevel modeling]]></category>
		<category><![CDATA[programmatic assessment]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[rater stringency]]></category>
		<category><![CDATA[resident trustworthiness scoring]]></category>
		<category><![CDATA[scoring algorithm]]></category>
		<category><![CDATA[statistical analysis of surgical skills]]></category>
		<category><![CDATA[surgical competence benchmarking]]></category>
		<category><![CDATA[surgical education]]></category>
		<category><![CDATA[surgical education measurement]]></category>
		<category><![CDATA[surgical residency evaluation]]></category>
		<category><![CDATA[surgical resident assessment]]></category>
		<category><![CDATA[surgical training quality]]></category>
		<category><![CDATA[workplace-based assessment]]></category>
		<category><![CDATA[workplace-based assessments]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221774</guid>

					<description><![CDATA[Researchers have developed a multilevel modeling algorithm that converts routine intraoperative EPA assessments into statistically adjusted entrustability scores for general surgery residents, accounting for procedural difficulty, repetition, and rater stringency.]]></description>
										<content:encoded><![CDATA[<p>Every time a surgical attending hands a resident the scalpel and steps back, an invisible judgment is being made: how much can this trainee be trusted to do this operation, right now, on this patient? That judgment—called entrustment—has long been the currency of surgical education, yet it has remained stubbornly difficult to measure in a way that is fair, rigorous, and useful. A new study from researchers at Massachusetts General Brigham and the University of Illinois, published in Global Surgical Education, the Journal of the Association for Surgical Education, now offers a mathematically grounded way to convert the everyday stream of workplace-based assessments into a single, statistically refined entrustability score for each resident. The work, led by Dr. Dandan Chen and colleagues, tackles one of the most persistent problems in competency-based medical education: raw assessment data are noisy, and the noise can obscure the signal of genuine trainee growth.</p>
<p>The study&#8217;s foundation is the entrustable professional activity, or EPA, a framework that has rapidly reshaped how surgical residency programs think about assessment. Rather than rating abstract competencies such as professionalism or medical knowledge in isolation, EPAs describe the concrete tasks a surgeon must be able to perform—like performing an laparoscopic cholecystectomy or managing an acutely ill patient—and ask supervising attendings to rate the level of supervision a resident actually required. In 2024, a national pilot study published in Annals of Surgery demonstrated that general surgery programs across the United States could implement EPA-based assessment at scale, and subsequent work has validated EPAs in national samples of programs. But implementation is only half the battle. Once programs collect thousands of microassessments, they face a harder question: what do all those numbers actually mean?</p>
<p>That question matters because intraoperative assessments are confounded by factors that have nothing to do with the resident&#8217;s ability. A resident who performs a straightforward hernia repair for the tenth time may earn a high entrustment rating, while the same resident struggling through a complex, first-time pancreatic resection may look far weaker than they truly are. The attending doing the rating adds another layer of variability: some raters are systematically generous, others harsh, and some vary widely from case to case. A simple average of ratings therefore blends together trainee skill, procedural difficulty, repetition, and rater temperament into a single number that is difficult to interpret. The research team set out to disentangle these threads using a statistical approach known as multilevel modeling, a technique with deep roots in educational measurement and generalizability theory.</p>
<p>The dataset behind the study is substantial. The researchers analyzed 949 intraoperative EPA microassessments collected during the 2024–2025 academic year at a large academic general surgery residency program. The assessments spanned 57 residents across all five postgraduate training years, rated by 72 different attending surgeons, and covered 14 distinct intraoperative EPA categories. This volume and diversity of data is exactly what makes the modeling approach feasible: with enough ratings per resident, per procedure type, and per rater, the model can statistically separate the components of each score. The study used de-identified assessment data and was reviewed by the local Institutional Review Board, which determined it to be exempt from full ethical review.</p>
<p>The core of the algorithm is a multilevel model with random intercepts for residents, for EPA categories, and for attending raters, combined with a fixed effect for procedural repetition. In practical terms, the model treats each individual entrustment rating as the sum of several influences: the resident&#8217;s underlying ability, the inherent difficulty of the procedure being performed, the stringency of the particular attending, and the learning that comes simply from having done the operation before. By estimating each of these components simultaneously, the model produces a composite entrustability score for every resident that is adjusted for the circumstances under which their ratings were earned. Procedural repetition emerged as a meaningful factor, with a positive association to entrustability (a coefficient of 0.028, statistically significant at p &lt; 0.001)—confirming quantitatively what every surgeon knows intuitively: doing an operation more times makes you better at it.</p>
<p>The psychometric evaluation of the resulting scores is where the study becomes particularly compelling for measurement scientists and program directors alike. The composite scores were normally distributed, with a mean of 2.67 on the entrustment scale and a range spanning 1.29 to 3.77. Normality matters because many downstream statistical procedures and competency benchmarks assume a roughly bell-shaped distribution, and skewed or clumped score distributions can distort decisions about promotion or remediation. Perhaps most reassuringly, the modeled scores correlated at r = 0.98 with the raw average ratings, indicating that the algorithm refines rather than replaces the intuitive signal that attendings already provide. The model does not invent a new construct; it sharpens an existing one by stripping away measurement noise.</p>
<p>Validity evidence came from a classic known-groups analysis. If the scoring algorithm captures something real about surgical development, senior residents should score higher than junior residents—and they did, decisively. Senior residents earned a mean composite score of 3.32 compared with 2.33 for junior residents, a difference that was highly statistically significant (p &lt; 0.001). This gradient across training levels supports the construct validity of the measure: the algorithm tracks the developmental trajectory that surgical educators expect to see as residents accumulate operative experience. The analysis also documented substantial variability in procedure difficulties and in attending rater stringency, confirming that the confounders the model adjusts for are not hypothetical—they are present and measurable in real program data.</p>
<p>For residency programs, the implications are practical. Competency-based, time-variable training has been proposed as a way to allow trainees to progress at their own pace, but regulatory structures and assessment systems have struggled to keep pace with that vision. A scoring algorithm that can synthesize local EPA data into defensible resident-level estimates gives program directors and clinical competency committees a tool for making promotion decisions that is more rigorous than eyeballing averages. It also provides earlier warning signals: a resident whose adjusted entrustability trajectory lags behind peers can be identified and supported before problems compound. The approach aligns with the Standards for Educational and Psychological Testing, the joint framework from the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education that governs how validity evidence should be accumulated for educational measures.</p>
<p>The study also speaks to a broader movement in surgical education toward large-scale workplace assessment. Systems such as SIMPL, developed to implement operative performance assessments at scale, have shown that mobile microassessment tools can generate the volume of data needed for meaningful statistical modeling. What has often been missing is the analytical layer—methods that transform raw ratings into trustworthy summaries. The multilevel approach described in this study draws on hierarchical linear modeling traditions dating back to the foundational work of Bryk and Raudenbush, and on generalizability theory frameworks articulated by measurement specialists such as Bloch and Norman. By applying these established psychometric tools to the specific problem of intraoperative entrustment, the researchers bridge a gap between assessment collection and assessment interpretation.</p>
<p>Limitations remain, as the authors acknowledge through the structure of their analysis. The data come from a single large academic program in the New England region, and the underlying data are not publicly available due to institutional restrictions protecting learner confidentiality. Whether the algorithm&#8217;s parameters—particularly the estimates of procedure difficulty and rater stringency—generalize to other programs with different case mixes and faculty cultures will require multi-institutional replication. The strong correlation with raw means also suggests that for many purposes, simple averages may already capture most of the signal; the model&#8217;s advantage lies in edge cases, where unusual case mixes or idiosyncratic raters could otherwise distort a resident&#8217;s apparent performance. Still, as EPA-based assessment becomes embedded in surgical training nationwide, this study offers a template for turning the daily stream of supervisory judgments into scores that programs can trust—scores that measure the resident, not the moment.</p>
<p><strong>Subject of Research:</strong> Development of a multilevel scoring algorithm for intraoperative resident entrustability using EPA assessment data</p>
<p><strong>Article Title:</strong> Developing a scoring algorithm for intraoperative entrustability among general surgery residents using local EPA data</p>
<p><strong>Article References:</strong> Chen, D., Mckinley, S., Thomas, J., Witt, E., Phitayakorn, R., Greer, J., Moses, J., &amp; Smink, D. (2026). Developing a scoring algorithm for intraoperative entrustability among general surgery residents using local EPA data. <em>Global Surgical Education &#8211; Journal of the Association for Surgical Education, 5</em>(1), Article 153. <a href="https://doi.org/10.1007/s44186-026-00550-2" rel="noopener noreferrer">https://doi.org/10.1007/s44186-026-00550-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44186-026-00550-2" rel="noopener noreferrer">10.1007/s44186-026-00550-2</a></p>
<p><strong>Keywords:</strong> entrustable professional activities, surgical education, general surgery residency, multilevel modeling, entrustability, workplace-based assessment, psychometrics, competency-based medical education, intraoperative assessment, scoring algorithm, rater stringency, programmatic assessment</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221774</post-id>	</item>
		<item>
		<title>Surgeons&#8217; Words in the Operating Room Reveal When Residents Are Ready to Fly Solo</title>
		<link>https://scienmag.com/surgeons-words-in-the-operating-room-reveal-when-residents-are-ready-to-fly-solo/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:14:02 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI transcription]]></category>
		<category><![CDATA[competency-based assessment]]></category>
		<category><![CDATA[competency-based surgical education]]></category>
		<category><![CDATA[entrustability]]></category>
		<category><![CDATA[Entrustable Professional Activities]]></category>
		<category><![CDATA[entrustment decision-making in surgery]]></category>
		<category><![CDATA[feedback]]></category>
		<category><![CDATA[general surgery]]></category>
		<category><![CDATA[language analysis in operating rooms]]></category>
		<category><![CDATA[microphone-based surgical skill evaluation]]></category>
		<category><![CDATA[operating room communication]]></category>
		<category><![CDATA[operative independence indicators]]></category>
		<category><![CDATA[operative linguistics]]></category>
		<category><![CDATA[predictive analysis of surgical performance]]></category>
		<category><![CDATA[real-time surgical readiness measurement]]></category>
		<category><![CDATA[resident autonomy]]></category>
		<category><![CDATA[resident autonomy in surgery]]></category>
		<category><![CDATA[shared mental modelling]]></category>
		<category><![CDATA[surgeon-trainee communication patterns]]></category>
		<category><![CDATA[surgical case evaluation methods]]></category>
		<category><![CDATA[surgical education]]></category>
		<category><![CDATA[surgical education research]]></category>
		<category><![CDATA[surgical training]]></category>
		<category><![CDATA[Surgical training assessment]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202055</guid>

					<description><![CDATA[Researchers recorded 25 real operations and found that measurable speech patterns between attending surgeons and trainees closely track operative autonomy and entrustability.]]></description>
										<content:encoded><![CDATA[<p>Every surgeon remembers the moment they stopped being watched and started being trusted. A new pilot study suggests that moment may be audible in the operating room itself, written into the very pattern of words exchanged between an attending surgeon and a trainee. Researchers who recorded 25 real general surgery operations found that the language flowing between teacher and learner changes in measurable, predictable ways as a resident gains operative independence, raising the possibility that a microphone could one day do what subjective end-of-case evaluations have long struggled to do: capture surgical readiness as it actually happens.</p>
<p>The study, led by Katharine E. Caldwell of the Medical University of South Carolina with colleagues at Washington University in Saint Louis and Stanford University, and published in Global Surgical Education, the Journal of the Association for Surgical Education, set out to solve a stubborn problem in surgical training. Modern programs are shifting toward competency-based frameworks built around Entrustable Professional Activities, or EPAs, which are designed to judge whether a trainee is ready for independent practice. Yet EPA ratings still depend on retrospective, written assessments completed after the operation ends. Those evaluations can suffer from recall bias, incomplete case capture, and documented racial and sex-based evaluator bias, and they may miss the subtle, minute-to-minute dynamics that define how a trainee actually performs under pressure.</p>
<p>The researchers&#8217; hypothesis was elegantly simple: the operating room conversation is itself a data stream. When a trainee is struggling, the attending speaks more, directs more, and takes over more often. When a trainee is ready to operate with only indirect supervision, the conversational balance flips. To test this, the team equipped attending surgeon and trainee dyads with lapel microphones during general surgery operations across three divisions: minimally invasive surgery, colorectal surgery, and surgical oncology. Recordings began immediately after the surgical time out and ended at skin closure, protecting patient privacy while capturing the full operative dialogue. Procedures included cholecystectomies, ventral and inguinal hernia repairs, small bowel and gastric resections, and colectomies, performed open, laparoscopically, and robotically, with a mean operative duration of 130.6 minutes.</p>
<p>From those recordings, the team generated an enormous corpus: 34,392 individual utterances. To make sense of it, they developed a framework of 38 unique operative linguistic codes using a modified grounded theory approach. Codes captured technical instruction, instrument requests, feedback, takeovers, off-target talking, and shared mental modelling, the practice of verbally indicating important anatomy, planning the next steps, or voicing uncertainty. Some codes were split by direction or valence, so feedback could be positive or negative, and takeovers could be classified as completion, demonstration, or safety events. Utterances deemed ordinary conversational filler, such as simple acknowledgements or clarifying questions, were excluded from analysis.</p>
<p>Artificial intelligence entered the workflow as an accelerant, not an arbiter. The researchers initialized a large language model, ChatGPT-4o, with the codebook and anchor examples, then ran a supervised training phase across 3,000 transcript lines in ten-line increments, with a human researcher reviewing and correcting each categorization. In a final layer of quality control, a human investigator reviewed 100 percent of the AI-generated codes, correcting errors that clustered around boundary cases, such as distinguishing off-target chatter with a third party from genuine explanation, or deciding whether a brief utterance like good or okay was feedback or mere agreement. Every final code in the study was assigned by a human, and the authors are explicit that AI-only coding remains unvalidated for this purpose and would require far more data and improved model accuracy before fully automated analysis becomes feasible.</p>
<p>After each operation, the attending rated the trainee on an EPA-based entrustability scale ranging from Level 1, limited participation, to Level 4, practice-ready. For analysis, cases were divided into lower-autonomy cases, where the trainee needed direct supervision, and higher-autonomy cases, where the trainee operated under indirect supervision or was deemed practice-ready. The linguistic contrasts between the two groups were striking. More autonomous trainees generated 28.4 percent of the words spoken during a case, compared with just 10.8 percent for trainees under direct supervision, a difference that held even after normalizing for case length and individual speaking rate. Talk, in other words, tracks trust.</p>
<p>The content of trainee speech shifted as sharply as its volume. Higher-autonomy trainees initiated 40.7 percent of instrument requests versus 17.8 percent in the lower-autonomy group, and they delivered 21.8 percent of technical instruction directed at the attending surgeon, compared with a mere 1.4 percent among less independent learners. They also led a dramatically larger share of shared mental modelling, 54.5 percent versus 16.8 percent, meaning they were the ones calling out anatomy, proposing the next operative step, and articulating uncertainty. Takeover events, in which the attending steps in to complete, demonstrate, or secure a critical maneuver, fell from an average of 7.6 per case in the direct supervision group to 0.6 per case among more autonomous trainees. In the lower-autonomy group, the majority of takeovers were completions, the attending finishing what the trainee could not, whereas among indirectly supervised trainees, takeovers more often took the form of demonstration.</p>
<p>The attending surgeons&#8217; language told the complementary story. When operating with highly trusted trainees, attendings engaged in significantly more off-target talking, 31.7 percent of utterances versus 8.7 percent, conversation unconnected to the immediate operative task, a behavioral signature of reduced need for continuous coaching. In lower-autonomy cases, such chatter was largely confined to the opening and closing of the case, vanishing during critical operative portions when every word carried weight. Total feedback volume was similar across groups, at 2.1 versus 1.5 percent of utterances, but the valence shifted decisively: attendings delivered 63.8 percent of feedback as positive in higher-autonomy cases, compared with just 24.0 percent in lower-autonomy ones. Technical feedback dominated overall, accounting for 85.3 percent of all feedback given.</p>
<p>The authors are careful to frame this as a pilot with real limitations. It was a single-center study of 25 operations at a large Midwestern academic medical center; all the attending surgeons were male, nearly half the trainees were fellows, and all procedures were common general surgery operations with complex cases deliberately excluded. The team could not analyze how race or gender shaped communication patterns, despite prior evidence that these factors influence feedback dynamics in surgical teams, and they did not control for familiarity between attending and trainee, which is known to affect team performance and entrustability. The Hawthorne effect looms as well: participants knew they were being recorded and may have altered their speech, prompting the group to investigate less invasive black box style recording technologies for future work.</p>
<p>Even so, the implications are considerable. If operative dialogue can be captured and coded at scale, every case could yield an objective, behavior-based supplement to EPA ratings, giving trainees individualized performance data and giving faculty a mirror for their own teaching styles, potentially transforming faculty development alongside trainee assessment. The researchers plan multicenter validation, integration with existing competency frameworks, and studies linking linguistic markers to real-time operative performance metrics. For now, the study&#8217;s most provocative message is conceptual: surgical autonomy is not just something evaluators imagine after the fact, but something audible in real time, one utterance at a time. The operating room, it turns out, has been telling us who is ready all along.</p>
<p><strong>Subject of Research:</strong> Using live operative audio recordings and linguistic analysis to evaluate surgical resident autonomy</p>
<p><strong>Article Title:</strong> How we talk and teach in the operating room: using live operative recordings to evaluate resident autonomy</p>
<p><strong>Article References:</strong> Caldwell, K. E., Beneville, B. T., Bennett, J., Jama, M. A., Fox, C., Ferzoco, M., Lewis, L., Tong, J., &amp; Awad, M. M. (2026). How we talk and teach in the operating room: using live operative recordings to evaluate resident autonomy. <em>Global Surgical Education &#8211; Journal of the Association for Surgical Education, 5</em>(1), Article 179. <a href="https://doi.org/10.1007/s44186-026-00584-6" rel="noopener noreferrer">https://doi.org/10.1007/s44186-026-00584-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44186-026-00584-6" rel="noopener noreferrer">10.1007/s44186-026-00584-6</a></p>
<p><strong>Keywords:</strong> surgical education, resident autonomy, operating room communication, Entrustable Professional Activities, operative linguistics, surgical training, feedback, shared mental modelling, AI transcription, competency-based assessment, general surgery, entrustability</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202055</post-id>	</item>
	</channel>
</rss>
