When future teachers evaluate one another’s practice lessons, a curious gap opens between what they score and what they say. A new study of 702 peer evaluations from a Japanese teacher-preparation program finds that pre-service teachers’ written comments do become more corrective when their numerical ratings are harsher—but only in a general sense. The comments rarely name, let alone critique, the specific teaching dimensions the evaluator actually rated low. The research, published in Discover Education by Shota Shirasaka of Fukuoka Institute of Technology and colleagues Takahisa Imagawa and Shuichi Enokida of Kyushu Institute of Technology, offers one of the most transparent looks yet at how the numbers and the words in peer feedback drift apart.
The setting was microteaching, a staple of teacher education in which trainees deliver short lessons to their peers and then swap roles as evaluators. Each of 27 third-year undergraduates—specializing in subjects from industrial technology to mathematics and science—delivered a three-minute lesson modeled on a short homeroom activity. Every participant then rated all 26 classmates and wrote a short free-text comment in Japanese about each performance. After excluding self-evaluations, the researchers had 702 paired evaluations, each combining nine five-point ratings (covering voice clarity, speaking speed, facial expression, gaze, gesture, board work, questioning, and content clarity, plus an overall score) with a comment averaging just 41.6 characters.
A crucial design feature shaped the entire analysis: the ratings were collected under a forced distribution. For each item, every evaluator had to assign the top score of 5 to exactly two peers, a 4 to seven peers, a 3 to nine, a 2 to seven, and a 1 to two—a fixed 2–7–9–7–2 shape. This means a rating of 1 or 2 does not indicate objectively poor teaching; it indicates that the evaluator placed that peer in the bottom of their own ranking. The researchers treated the ratings as within-evaluator relative ranks, which gave them an internal benchmark: if rating and commenting were tightly linked, corrective language should cluster precisely where the evaluator’s own low ratings fell.
To measure the comments, the team deliberately avoided large language models and human coding, both of which introduce reproducibility problems. Instead, they built a frozen, rule-based Japanese lexicon and scanned every comment with the morphological analyzer MeCab for five explicit markers: generic praise, concrete instructional referent, learner-oriented mention, explicit feed-forward (improvement-suggestion grammar such as ‘it would be better to’), and explicit criticism. Every reported label can be exactly reproduced from the published dictionaries. As a stress test, the researchers also tried classifying comments by their similarity to Sentence-BERT category prototypes; that probe agreed poorly with the rule-based labels, reinforcing their choice of deterministic rules for such short, praise-heavy text.
The descriptive results confirmed a pattern long noted in the peer-feedback literature: novices praise. Generic praise appeared in 67.1 percent of comments. Explicit feed-forward language was detected in only 14.4 percent of comments under the frozen lexicon and 13.0 percent under a stricter ablation, while explicit criticism appeared in just 6.4 percent and 4.6 percent respectively. The similarity of the frozen and strict rates suggests the scarcity of corrective language is robust rather than an artifact of permissive matching, and an audit of all positive cases found no obvious false positives. Concrete instructional referents ranged from 53.6 to 71.4 percent depending on dictionary breadth, and learner-oriented mentions appeared in 21.4 percent of comments.
The first key finding was a genuine, graded coupling between rating severity and corrective wording. Among the 574 comments accompanying at least one low rating (a 1 or 2 on the eight teaching items), 22.1 percent contained an explicit feed-forward or critical marker, compared with only 10.9 percent of the 128 comments where all items were rated 3 to 5—a difference of +0.118 that a within-evaluator permutation test put at p = 0.008. The corrective-marker rate rose steadily with both the number of low-rated items and the severity of the lowest rating (Spearman ρ = +0.205 and −0.204, each p = 0.0002). In plain terms: the harsher an evaluator’s overall placement of a peer, the more likely their comment carried corrective language.
But the second finding dismantled any hope that this corrective language was targeted. Of the 491 comments that explicitly mentioned at least one of the eight predefined teaching dimensions, only 180 mentioned a dimension the same evaluator had rated low—a 36.7 percent match rate that actually fell slightly below the 39.6 percent chance baseline generated by re-pairing comments with the evaluator’s own low-rated dimension sets (one-sided p = 0.958). When a pre-service teacher named a teaching dimension, it was no more likely than chance to be one they had scored low. Per-item breakdowns told the same story: the difference in same-dimension mention rates between low- and high-rated items was near zero or negative on every item except speaking speed.
Stranger still, even the mentions that did land on a rated-low dimension were usually framed positively. Among those 180 comments, the clause surrounding the mention carried an explicit praise marker more often than a correction marker—55.0 percent versus 26.7 percent at the clause level, and 75.6 percent versus 34.4 percent across the whole comment. The researchers interpret this as ‘loose coupling’: the numerical and textual channels of the same evaluation share a global structure, in which overall severity nudges corrective wording, but the text adds no dimension-specific information beyond the ratings themselves. Notably, a companion analysis of the same cohort found the numerical ratings to be highly halo-structured—a single general impression rather than independent dimension judgments—so weak textual alignment is compatible with how the ratings behave, not paradoxical.
The study has clear limits, which the authors state plainly. It is exploratory, based on a single cohort, with short Japanese comments; the markers capture only explicit surface wording, so implicit or indirectly phrased feedback falls outside the measured construct, and the scarcity rates are lower bounds. Comment length mattered too: the corrective-marker rate climbed from about 8 percent in the shortest length quartile to about 36 percent in the longest, suggesting some of the scarcity reflects brevity rather than disposition alone—though the dimension-targeting failure was not explained by length. Timing also differed: final ratings could be adjusted after re-viewing recorded sessions, while comments were typically written immediately, and no edit histories were recorded.
The practical implications reach both teacher education and the growing field of automated feedback support. For trainers, the message is that ranking peers and articulating usable, dimension-specific suggestions are distinct skills that may need to be scaffolded separately rather than assumed to develop together. For developers of feedback tools, the finding is a caution: systems cannot assume that a low rating will come with detected corrective or dimension-relevant text to build on. The authors suggest future work could test whether prompts or scaffolds help learners convert a noticed weakness into explicit wording, apply the same pipeline to rubric-anchored rating formats, or benchmark large language model classifiers against human codes. For now, the study stands as a precise, reproducible demonstration that what novice evaluators write and what they score are two related but separable acts—one that any effort to automate or improve peer feedback must reckon with.
Subject of Research: The correspondence between numerical ratings and written comment content in pre-service teacher microteaching peer evaluation
Article Title: Peer comment wording in pre-service teacher microteaching tracks rating severity but not the dimensions rated low
Article References: Shirasaka, S., Imagawa, T., & Enokida, S. (2026). Peer comment wording in pre-service teacher microteaching tracks rating severity but not the dimensions rated low. Discover Education, 5(1), Article 1141. https://doi.org/10.1007/s44217-026-02238-7
Image Credits: AI Generated
DOI: 10.1007/s44217-026-02238-7
Keywords: peer feedback, microteaching, pre-service teachers, teacher education, lexicon-based text analysis, natural language processing, forced distribution ratings, feed-forward, evaluative judgement, feedback literacy, permutation test, halo effect
Cite Scienmag News
Courtney Benton. (October 7, 2026). When Student Teachers Rate Peers Harshly, Their Written Comments Stay Vague. Scienmag. https://scienmag.com/when-student-teachers-rate-peers-harshly-their-written-comments-stay-vague/
Courtney Benton. "When Student Teachers Rate Peers Harshly, Their Written Comments Stay Vague." Scienmag, 7 October 2026, https://scienmag.com/when-student-teachers-rate-peers-harshly-their-written-comments-stay-vague/. Accessed 7 October 2026.
Courtney Benton. "When Student Teachers Rate Peers Harshly, Their Written Comments Stay Vague." Scienmag. October 7, 2026. https://scienmag.com/when-student-teachers-rate-peers-harshly-their-written-comments-stay-vague/

