Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Social Science

New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores

October 1, 2026
in Social Science
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 5 mins read
0
New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores

New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores

New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every time a surgical attending hands a resident the scalpel and steps back, an invisible judgment is being made: how much can this trainee be trusted to do this operation, right now, on this patient? That judgment—called entrustment—has long been the currency of surgical education, yet it has remained stubbornly difficult to measure in a way that is fair, rigorous, and useful. A new study from researchers at Massachusetts General Brigham and the University of Illinois, published in Global Surgical Education, the Journal of the Association for Surgical Education, now offers a mathematically grounded way to convert the everyday stream of workplace-based assessments into a single, statistically refined entrustability score for each resident. The work, led by Dr. Dandan Chen and colleagues, tackles one of the most persistent problems in competency-based medical education: raw assessment data are noisy, and the noise can obscure the signal of genuine trainee growth.

The study’s foundation is the entrustable professional activity, or EPA, a framework that has rapidly reshaped how surgical residency programs think about assessment. Rather than rating abstract competencies such as professionalism or medical knowledge in isolation, EPAs describe the concrete tasks a surgeon must be able to perform—like performing an laparoscopic cholecystectomy or managing an acutely ill patient—and ask supervising attendings to rate the level of supervision a resident actually required. In 2024, a national pilot study published in Annals of Surgery demonstrated that general surgery programs across the United States could implement EPA-based assessment at scale, and subsequent work has validated EPAs in national samples of programs. But implementation is only half the battle. Once programs collect thousands of microassessments, they face a harder question: what do all those numbers actually mean?

That question matters because intraoperative assessments are confounded by factors that have nothing to do with the resident’s ability. A resident who performs a straightforward hernia repair for the tenth time may earn a high entrustment rating, while the same resident struggling through a complex, first-time pancreatic resection may look far weaker than they truly are. The attending doing the rating adds another layer of variability: some raters are systematically generous, others harsh, and some vary widely from case to case. A simple average of ratings therefore blends together trainee skill, procedural difficulty, repetition, and rater temperament into a single number that is difficult to interpret. The research team set out to disentangle these threads using a statistical approach known as multilevel modeling, a technique with deep roots in educational measurement and generalizability theory.

The dataset behind the study is substantial. The researchers analyzed 949 intraoperative EPA microassessments collected during the 2024–2025 academic year at a large academic general surgery residency program. The assessments spanned 57 residents across all five postgraduate training years, rated by 72 different attending surgeons, and covered 14 distinct intraoperative EPA categories. This volume and diversity of data is exactly what makes the modeling approach feasible: with enough ratings per resident, per procedure type, and per rater, the model can statistically separate the components of each score. The study used de-identified assessment data and was reviewed by the local Institutional Review Board, which determined it to be exempt from full ethical review.

The core of the algorithm is a multilevel model with random intercepts for residents, for EPA categories, and for attending raters, combined with a fixed effect for procedural repetition. In practical terms, the model treats each individual entrustment rating as the sum of several influences: the resident’s underlying ability, the inherent difficulty of the procedure being performed, the stringency of the particular attending, and the learning that comes simply from having done the operation before. By estimating each of these components simultaneously, the model produces a composite entrustability score for every resident that is adjusted for the circumstances under which their ratings were earned. Procedural repetition emerged as a meaningful factor, with a positive association to entrustability (a coefficient of 0.028, statistically significant at p < 0.001)—confirming quantitatively what every surgeon knows intuitively: doing an operation more times makes you better at it.

The psychometric evaluation of the resulting scores is where the study becomes particularly compelling for measurement scientists and program directors alike. The composite scores were normally distributed, with a mean of 2.67 on the entrustment scale and a range spanning 1.29 to 3.77. Normality matters because many downstream statistical procedures and competency benchmarks assume a roughly bell-shaped distribution, and skewed or clumped score distributions can distort decisions about promotion or remediation. Perhaps most reassuringly, the modeled scores correlated at r = 0.98 with the raw average ratings, indicating that the algorithm refines rather than replaces the intuitive signal that attendings already provide. The model does not invent a new construct; it sharpens an existing one by stripping away measurement noise.

Validity evidence came from a classic known-groups analysis. If the scoring algorithm captures something real about surgical development, senior residents should score higher than junior residents—and they did, decisively. Senior residents earned a mean composite score of 3.32 compared with 2.33 for junior residents, a difference that was highly statistically significant (p < 0.001). This gradient across training levels supports the construct validity of the measure: the algorithm tracks the developmental trajectory that surgical educators expect to see as residents accumulate operative experience. The analysis also documented substantial variability in procedure difficulties and in attending rater stringency, confirming that the confounders the model adjusts for are not hypothetical—they are present and measurable in real program data.

For residency programs, the implications are practical. Competency-based, time-variable training has been proposed as a way to allow trainees to progress at their own pace, but regulatory structures and assessment systems have struggled to keep pace with that vision. A scoring algorithm that can synthesize local EPA data into defensible resident-level estimates gives program directors and clinical competency committees a tool for making promotion decisions that is more rigorous than eyeballing averages. It also provides earlier warning signals: a resident whose adjusted entrustability trajectory lags behind peers can be identified and supported before problems compound. The approach aligns with the Standards for Educational and Psychological Testing, the joint framework from the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education that governs how validity evidence should be accumulated for educational measures.

The study also speaks to a broader movement in surgical education toward large-scale workplace assessment. Systems such as SIMPL, developed to implement operative performance assessments at scale, have shown that mobile microassessment tools can generate the volume of data needed for meaningful statistical modeling. What has often been missing is the analytical layer—methods that transform raw ratings into trustworthy summaries. The multilevel approach described in this study draws on hierarchical linear modeling traditions dating back to the foundational work of Bryk and Raudenbush, and on generalizability theory frameworks articulated by measurement specialists such as Bloch and Norman. By applying these established psychometric tools to the specific problem of intraoperative entrustment, the researchers bridge a gap between assessment collection and assessment interpretation.

Limitations remain, as the authors acknowledge through the structure of their analysis. The data come from a single large academic program in the New England region, and the underlying data are not publicly available due to institutional restrictions protecting learner confidentiality. Whether the algorithm’s parameters—particularly the estimates of procedure difficulty and rater stringency—generalize to other programs with different case mixes and faculty cultures will require multi-institutional replication. The strong correlation with raw means also suggests that for many purposes, simple averages may already capture most of the signal; the model’s advantage lies in edge cases, where unusual case mixes or idiosyncratic raters could otherwise distort a resident’s apparent performance. Still, as EPA-based assessment becomes embedded in surgical training nationwide, this study offers a template for turning the daily stream of supervisory judgments into scores that programs can trust—scores that measure the resident, not the moment.

Subject of Research: Development of a multilevel scoring algorithm for intraoperative resident entrustability using EPA assessment data

Article Title: Developing a scoring algorithm for intraoperative entrustability among general surgery residents using local EPA data

Article References: Chen, D., Mckinley, S., Thomas, J., Witt, E., Phitayakorn, R., Greer, J., Moses, J., & Smink, D. (2026). Developing a scoring algorithm for intraoperative entrustability among general surgery residents using local EPA data. Global Surgical Education – Journal of the Association for Surgical Education, 5(1), Article 153. https://doi.org/10.1007/s44186-026-00550-2

Image Credits: AI Generated

DOI: 10.1007/s44186-026-00550-2

Keywords: entrustable professional activities, surgical education, general surgery residency, multilevel modeling, entrustability, workplace-based assessment, psychometrics, competency-based medical education, intraoperative assessment, scoring algorithm, rater stringency, programmatic assessment

Cite Scienmag News

Courtney Benton. (October 1, 2026). New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores. Scienmag. https://scienmag.com/new-scoring-algorithm-turns-routine-surgical-ratings-into-fairer-resident-entrustability-scores/

Courtney Benton. "New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores." Scienmag, 1 October 2026, https://scienmag.com/new-scoring-algorithm-turns-routine-surgical-ratings-into-fairer-resident-entrustability-scores/. Accessed 1 October 2026.

Courtney Benton. "New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores." Scienmag. October 1, 2026. https://scienmag.com/new-scoring-algorithm-turns-routine-surgical-ratings-into-fairer-resident-entrustability-scores/

Tags: assessment noise reductioncompetency-based medical educationentrustabilityEntrustable Professional Activitiesentrustable professional activities (EPAs)entrustment scoringgeneral surgery residencyintraoperative assessmentMultilevel modelingprogrammatic assessmentpsychometricsrater stringencyresident trustworthiness scoringscoring algorithmstatistical analysis of surgical skillssurgical competence benchmarkingsurgical educationsurgical education measurementsurgical residency evaluationsurgical resident assessmentsurgical training qualityworkplace-based assessmentworkplace-based assessments
Share26Tweet16
Previous Post

Impulsivity Intensifies Substance Use Link After Sexual Abuse, Global Survey Finds

Next Post

Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive

Related Posts

Impulsivity Intensifies Substance Use Link After Sexual Abuse, Global Survey Finds
Social Science

Impulsivity Intensifies Substance Use Link After Sexual Abuse, Global Survey Finds

October 1, 2026
How Big Must a City Park Be to Cool Its Surroundings? Satellite Data Reveal a Surprising Ceiling
Social Science

How Big Must a City Park Be to Cool Its Surroundings? Satellite Data Reveal a Surprising Ceiling

October 1, 2026
Universities Rethink the Curriculum as Generative AI Reshapes Higher Education
Social Science

Universities Rethink the Curriculum as Generative AI Reshapes Higher Education

October 1, 2026
Two Hurdles Stand Between Nigerian Farmers and Agricultural Insurance
Social Science

Two Hurdles Stand Between Nigerian Farmers and Agricultural Insurance

October 1, 2026
Strategic Autonomy: How Trumpism Is Forcing Europe to Reinvent Itself
Social Science

Strategic Autonomy: How Trumpism Is Forcing Europe to Reinvent Itself

October 1, 2026
India’s Chronic Disease Burden Hits Women Harder, National Survey Analysis Reveals
Social Science

India’s Chronic Disease Burden Hits Women Harder, National Survey Analysis Reveals

October 1, 2026
Next Post
Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive

Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Game Theory Reveals When Shale Gas Firms Should Think Long Term
  • Proteomic Aging Clocks Expose the Tangled Truth Behind Alcohol’s Healthy Drinking Paradox
  • Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive
  • New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading