When a trauma patient arrives at an emergency department, the difference between a smooth resuscitation and a chaotic one often has little to do with technical skill. It hinges on leadership, communication, cooperation, and the ability of a team to anticipate what comes next. These are the non-technical skills that researchers have spent decades trying to measure, and one of the most trusted instruments for doing so is the Team Emergency Assessment Measure, known simply as TEAM. Now a research team working across institutions in China and Malaysia has taken a crucial step: they have translated, culturally adapted, and rigorously tested a Simplified Chinese version of the questionnaire, opening the door for its use in the world’s largest Chinese-speaking clinical workforce.
The study, published in BMC Medical Education, addresses a gap that has persisted despite the TEAM tool’s translation into multiple languages. The original instrument evaluates eleven observable behaviors during emergency team performance, covering domains such as leadership and team cooperation, along with global ratings of team performance and the observer’s overall impression. Until now, no Simplified Chinese version had been formally validated, which meant that emergency departments in mainland China lacked a linguistically and culturally equivalent instrument for structured observation and feedback during training in trauma and resuscitation scenarios.
The translation process followed established cross-cultural adaptation guidelines, a pipeline designed to ensure that a questionnaire means the same thing in a new language as it did in the original. The researchers began with independent forward translations of the scale into Simplified Chinese, then synthesized those versions into a single draft. That draft was back-translated into English by translators blinded to the original instrument, allowing the team to detect any drift in meaning. A multidisciplinary panel of authors, doctors, nurses, and English-language specialists then reviewed the versions side by side, reconciling discrepancies in wording until the Chinese text captured the intent of every item.
Crucially, the team did not stop at linguistic equivalence. They collaborated directly with the original authors of the TEAM questionnaire to maintain conceptual equivalence, and they assessed cultural relevance through Delphi rounds and cognitive interviews with domain experts. Delphi methods involve iterative rounds of structured expert consultation, in which panelists rate and comment on each item until consensus emerges. Cognitive interviews probe whether respondents understand an item the way its authors intended, surfacing ambiguities that a literal translation would never reveal. Together, these steps ensured that the Chinese wording resonated with the clinical realities and communication norms of Chinese emergency teams rather than simply mirroring English phrasing.
The psychometric evaluation centered on video-based assessment. Two trained raters independently scored video recordings of 46 emergency team scenarios involving severe trauma or resuscitation, producing 46 pairs of ratings and 92 individual ratings in total. Using video rather than live observation allowed both raters to see identical performances, isolating the agreement between raters from the noise of real-time perspective differences. The study received ethics approval from the Ethics Committee of the Second Affiliated Hospital of Zhejiang University School of Medicine, and written informed consent was obtained from participating healthcare professionals and from patients or their legally authorized representatives for the research use of recordings.
The results provide strong preliminary support for the new version. Expert assessment of content validity yielded a scale-level content validity index average of 0.98, remarkably close to the perfect score of 1.0, indicating near-unanimous expert agreement that the eleven behavioral items remain relevant and clear in Chinese. The universal agreement index, which is stricter because it requires every expert to rate an item as relevant, stood at 0.82, still comfortably above commonly accepted thresholds. These figures suggest that the adaptation preserved the substance of the original instrument while making it culturally natural for Chinese clinicians.
Internal consistency, which measures how coherently the items hang together as a single scale, was good, with a Cronbach’s alpha of 0.896. This statistic ranges from zero to one, and values above 0.9 are often considered excellent, though excessively high values can signal redundant items. Item-total correlations, which indicate how strongly each item relates to the overall score, ranged from 0.547 to 0.764, meaning every behavioral item contributed meaningfully to the measurement of the underlying construct. No item behaved as an outlier or dragged down the scale’s coherence, an encouraging sign that the translated items measure a unified dimension of team performance.
Inter-rater reliability, arguably the most important property for an observational rating tool, was also strong. The average-measure intraclass correlation coefficient for the total score across the eleven behavioral items was 0.902, with a 95 percent confidence interval of 0.836 to 0.938. An ICC of that magnitude indicates that two independent trained observers watching the same resuscitation would arrive at very similar assessments of team performance, which is exactly what a tool intended for feedback and training requires. At the level of individual items, average-measure ICCs ranged from 0.553 to 0.849, a more varied picture suggesting that some specific behaviors, such as those that are harder to observe or more context-dependent, may be somewhat more challenging to rate consistently than the scale as a whole.
The analysis of floor and ceiling effects revealed a pattern worth noting. No item showed a floor effect, meaning raters did not cluster scores at the bottom of the scale even for weaker team performances. However, six items showed ceiling effects ranging from 15.22 percent to 26.09 percent, indicating that a meaningful share of the observed teams received the maximum rating on those behaviors. Ceiling effects can compress the range of scores at the top end and reduce the instrument’s sensitivity to differences among high-performing teams. The authors suggest that future research should examine the structural validity and responsiveness of the Simplified Chinese version, as well as its performance across other clinical settings, to determine whether the ceiling pattern persists and how best to address it.
The practical implications extend well beyond psychometrics. Structured assessment tools like TEAM are the backbone of simulation-based training in emergency medicine, giving instructors an objective framework for debriefing after mock codes and trauma drills. A validated Simplified Chinese version means that Chinese emergency departments, simulation centers, and nursing schools can now deliver feedback grounded in an instrument whose measurement properties have been demonstrated in their own language and cultural context. It also enables Chinese teams to participate in international research comparing team performance across countries, since scores from the adapted version can be interpreted against the same construct as the original. The researchers, led by Yukun Zhang of the Second Affiliated Hospital of Zhejiang University School of Medicine and Universiti Malaya, together with LiYoong Tang, MeiChan Chong, Shanshan Li, and Yuwei Wang, frame the work as a foundation rather than a finish line. The evidence they report covers content validity, internal consistency, and inter-rater reliability in trauma and resuscitation scenarios; broader claims about responsiveness to training interventions and generalizability to other acute care settings await further study. For now, the study stands as a careful example of how a measurement instrument crosses a linguistic border without losing its meaning, and as a practical gift to the clinicians who train the teams that respond when every second counts.
Subject of Research: Cross-cultural adaptation and psychometric validation of a Simplified Chinese version of the Team Emergency Assessment Measure for emergency team performance
Article Title: Cross-cultural adaptation, translation, and psychometric evaluation of the simplified Chinese version of the Team Emergency Assessment Measure (TEAM)
Article References: Zhang, Y., Tang, L., Chong, M., Li, S., & Wang, Y. (2026). Cross-cultural adaptation, translation, and psychometric evaluation of the simplified Chinese version of the Team Emergency Assessment Measure (TEAM). BMC Medical Education. https://doi.org/10.1186/s12909-026-10529-8
Image Credits: AI Generated
DOI: 10.1186/s12909-026-10529-8
Keywords: TEAM questionnaire, non-technical skills, cross-cultural adaptation, psychometrics, emergency medicine, teamwork, resuscitation, trauma, inter-rater reliability, content validity, medical education, translation
Cite Scienmag News
Courtney Benton. (October 5, 2026). Chinese Version of Emergency Teamwork Rating Tool Proves Reliable in Trauma and Resuscitation Scenarios. Scienmag. https://scienmag.com/chinese-version-of-emergency-teamwork-rating-tool-proves-reliable-in-trauma-and-resuscitation-scenarios/
Courtney Benton. "Chinese Version of Emergency Teamwork Rating Tool Proves Reliable in Trauma and Resuscitation Scenarios." Scienmag, 5 October 2026, https://scienmag.com/chinese-version-of-emergency-teamwork-rating-tool-proves-reliable-in-trauma-and-resuscitation-scenarios/. Accessed 5 October 2026.
Courtney Benton. "Chinese Version of Emergency Teamwork Rating Tool Proves Reliable in Trauma and Resuscitation Scenarios." Scienmag. October 5, 2026. https://scienmag.com/chinese-version-of-emergency-teamwork-rating-tool-proves-reliable-in-trauma-and-resuscitation-scenarios/

