Ever since the COVID-19 pandemic forced medical schools and residency programs to move their high-stakes examinations online, educators have wrestled with a deceptively simple question: is a virtual clinical exam really equivalent to an in-person one? A new study from researchers at New York University Grossman School of Medicine, published in the Journal of General Internal Medicine, offers the most granular answer yet — and it comes with a statistical twist. By applying item response theory, a psychometric technique borrowed from educational testing, the team found that while overall communication scores looked reassuringly similar across virtual and in-person formats, at least one specific communication behavior behaved very differently depending on the modality. The finding challenges the assumption that a checklist rating of “well done” carries the same meaning in a video call as it does across a hospital bedside, and it hands assessment designers a powerful new tool for auditing exam comparability at the level of individual behaviors rather than aggregate scores.
The study focused on the Objective Structured Clinical Examination, or OSCE, a cornerstone of medical training in which trainees rotate through stations and interact with standardized patients — trained actors portraying realistic clinical scenarios. In this case, 126 first-year internal medicine residents at NYU completed six communication-focused OSCE cases between 2019 and 2023, with 82 assessed in person and 54 assessed virtually. In each case, the standardized patients rated resident performance across three core communication domains: information gathering, relationship development, and patient education. Crucially, the ratings used a behaviorally anchored scale — “not done,” “partially done,” or “well done” — meaning each rating level corresponded to observable clinical behaviors rather than a vague global impression. This design choice mattered enormously for the analysis that followed.
Most studies comparing virtual and in-person OSCEs have stopped at overall performance scores, concluding broadly that the two formats produce similar results. The NYU team, led by Christine P. Beltran and Colleen Gillespie, argued that such aggregate comparisons can mask important differences. Two exams might yield identical average scores while individual checklist items behave differently across formats — for example, if a particular behavior is easier to demonstrate, or easier to notice, over video. To detect these hidden discrepancies, the researchers turned to the graded response model, a form of item response theory that estimates what they call item-level thresholds. Each threshold represents the amount of underlying communication proficiency a resident needs to move from one rating category to the next — from “not done” to “partially done,” or from “partially done” to “well done.” If a threshold is lower in one modality, that means residents need less actual skill to earn the same rating there.
The technical machinery is worth unpacking, because it represents a meaningful methodological advance for medical education assessment. In a graded response model, every checklist item is characterized by a discrimination parameter, describing how well the item distinguishes between residents of different ability levels, and a set of threshold parameters, marking the points on the latent proficiency continuum where the probability of achieving each successive rating crosses 50 percent. By fitting the model separately to virtual and in-person data and then comparing threshold estimates with Wald tests — a standard statistical procedure for testing whether estimated parameters differ significantly — the researchers could ask, item by item, whether “partially done” or “well done” meant the same thing in both settings. This stands in contrast to traditional approaches that compare mean scores, which assume that identical numbers imply identical underlying constructs.
The headline results were, in most respects, reassuring. Most residents demonstrated sufficient skill to receive “partially done” or “well done” ratings on the majority of communication items, regardless of format, and overall communication performance appeared broadly similar across modalities. For program directors worried that the pandemic-era pivot to virtual assessment diluted their exams, the message is largely encouraging: the global picture of resident communication competence looks comparable whether the encounter happens in an exam room or on a screen. But the item-level analysis told a subtler story. One specific behavior — “using words the patient understood and explaining jargon” — showed a statistically significant difference between modalities. Residents required a lower level of underlying proficiency to receive a “partially done” rating on that item in virtual encounters compared with in-person ones, with a Wald test statistic of z = 3.92 and an adjusted p-value of 0.001.
Why would explaining medical jargon in plain language be easier to credit over video? The authors are careful to lay out several non-exclusive explanations. It may reflect genuine modality-related variation in the skill itself: residents might consciously simplify their language when they cannot rely on physical props, printed materials, or the full repertoire of in-person cues, making plain speech more prevalent on screen. Alternatively, the difference may lie in rater scoring — standardized patients evaluating a video encounter may attend differently to language, or may be more generous when communication is constrained by technology. Both mechanisms could produce the same statistical signature, and the study design cannot fully disentangle them. What the finding does establish, however, is that at least one checklist item does not function equivalently across formats, which means raw scores from virtual and in-person exams are not perfectly interchangeable, even when their averages look the same.
The broader context makes the study timely. Telehealth has evolved from a pandemic stopgap into a permanent fixture of American medicine, and organizations such as the Association of American Medical Colleges have published telehealth competency frameworks urging training programs to teach and assess virtual care skills deliberately. Previous studies of virtual OSCEs — spanning physical medicine and rehabilitation residencies, pediatric high-stakes exams, dental education, and systematic reviews of implementation across the health professions — have generally reported acceptable feasibility and comparable overall performance, but few have probed whether individual rating categories mean the same thing across formats. Earlier psychometric work had already applied item response theory to OSCE data to explore rater effects and item discrimination, and researchers had documented differences in how satisfied patients feel about physician communication during telemedicine visits. The NYU study ties these threads together, using threshold analysis as a diagnostic for format comparability.
For the assessment community, the practical implications are concrete. Programs running mixed-modality OSCEs — increasingly common as telehealth training becomes standard — can use graded response modeling as a routine quality-assurance step, flagging items whose thresholds drift between formats before those items contaminate pass-fail decisions. The one flagged behavior in this study, avoiding jargon, is itself a natural candidate for targeted curriculum revision: if residents earn partial credit for plain language more easily online, virtual encounters may need explicit prompts or rubric anchors that distinguish genuine skill from modality-driven simplification. The authors also note that their sample, drawn from a single internal medicine program across six cases, limits generalizability, and that the modest virtual cohort of 54 residents constrains statistical power. The project, funded through the AAMC Competency-Based Education in Telehealth Challenge Grant, was reviewed by NYU’s institutional process and classified as educational quality improvement using fully anonymized, de-identified data, with the informed consent requirement waived accordingly. The work was previously presented as a poster at the 2025 Society of General Internal Medicine Annual Conference.
Perhaps the most durable contribution of the study is conceptual: it reframes what “equivalence” means in clinical skills assessment. Averages alone, the authors argue, can lull educators into false confidence, because two formats can converge on the same totals while individual behaviors are weighted differently by raters or expressed differently by trainees. Threshold analysis offers a way to interrogate whether an observed rating reflects the same level of underlying communication proficiency across settings — essentially asking not just whether residents score the same, but whether the same score certifies the same competence. As telemedicine cements its place in routine care, ensuring that the exams certifying tomorrow’s physicians measure what they claim, in every modality, is no longer a technical nicety. It is a patient-safety question, and this study provides a template for answering it one checklist item at a time.
Subject of Research: Item response theory analysis of internal medicine residents' OSCE communication skill ratings in virtual versus in-person examinations
Article Title: When “Well-Done” Is Not the Same: Item Response Theory Analysis of Medicine Residents’ OSCE Communication Ratings Across Modalities
Article References: Beltran, C. P., Nallamaddi, S., Wilhite, J. A., Hardowar, K., Hanley, K., Altshuler, L., Zabar, S. R., & Gillespie, C. (2026). When “Well-Done” Is Not the Same: Item Response Theory Analysis of Medicine Residents’ OSCE Communication Ratings Across Modalities. Journal of General Internal Medicine. https://doi.org/10.1007/s11606-026-10798-5
Image Credits: AI Generated
DOI: 10.1007/s11606-026-10798-5
Keywords: OSCE, item response theory, communication skills, medical residents, telehealth, virtual assessment, standardized patients, internal medicine, graded response model, medical education, assessment comparability, patient communication
Cite Scienmag News
Ophelia Keating. (September 23, 2026). Same Score, Different Skill? Rethinking Virtual Doctor Communication Exams. Scienmag. https://scienmag.com/same-score-different-skill-rethinking-virtual-doctor-communication-exams/
Ophelia Keating. "Same Score, Different Skill? Rethinking Virtual Doctor Communication Exams." Scienmag, 23 September 2026, https://scienmag.com/same-score-different-skill-rethinking-virtual-doctor-communication-exams/. Accessed 23 September 2026.
Ophelia Keating. "Same Score, Different Skill? Rethinking Virtual Doctor Communication Exams." Scienmag. September 23, 2026. https://scienmag.com/same-score-different-skill-rethinking-virtual-doctor-communication-exams/

