When artificial intelligence systems are asked to grade the same spoken English exam that experienced human examiners grade, the results are anything but uniform. A new study published in Discover Education has found that the much-discussed gap between AI and human scoring in high-stakes language assessment is not a property of artificial intelligence in general, but of specific systems and, crucially, of how they receive the test itself. The findings suggest that the future of test preparation lies not in choosing between human and machine judgment, but in deliberately splitting the work between them.
Researchers led by Nur Aeni of Universitas Negeri Makassar in Indonesia recorded the IELTS-style speaking performances of 30 English-as-a-foreign-language university students, each completing the full three-part interview format: an introduction and interview, an individual long turn, and a two-way discussion. The identical recordings were then scored by two experienced English-speaking lecturers and by three commercially available AI platforms: ChatGPT, Otter AI, and Claude AI. Each system received a standardized prompt instructing it to apply the official IELTS Speaking Band Descriptors across the four rated competencies of fluency and coherence, lexical resource, grammatical range and accuracy, and pronunciation. No system was fine-tuned or custom-trained for the task.
When the three AI scores were averaged into a single composite, the picture looked familiar to anyone following the debate over automated assessment. Human raters assigned significantly higher scores overall, with a mean difference of 0.522 band points, a result that held up at the p < .001 level with a large effect size. The correlation between the human and AI composites was moderate and positive (r = .497, p = .005), meaning the two methods agreed broadly on the ranking of candidates but shared only about a quarter of their scoring variance. Twenty-eight of the thirty students received lower scores from the machines than from the humans, and the gap widened for the strongest speakers, with automated scores rarely exceeding 7.0 while human ratings frequently approached 8.0.
But the composite figure concealed the study’s most striking discovery. When the researchers disaggregated the results system by system, ChatGPT’s scores turned out to be statistically indistinguishable from those of one of the two human raters, with an effect size near zero. Otter AI and, especially, Claude AI diverged substantially from both humans, with Claude AI differing from the more lenient rater by nearly three standard deviations. A repeated-measures ANOVA across all five raters confirmed the pattern, and pairwise correlations told the same story: ChatGPT correlated significantly with both human raters, while Claude AI correlated with neither and barely overlapped with Otter AI. Treating AI scoring as a single undifferentiated category, the authors argue, obscures more than it reveals.
What explains the divergence? The researchers propose a deceptively simple answer: input modality. ChatGPT received the audio recordings directly through its voice input function, giving it access to the same acoustic information humans use, including intonation, stress placement, pausing, and articulation quality. Claude AI could not accept audio at the time of data collection, so recordings were first transcribed by a separate application, and Otter AI, a transcription platform by design, also scored from text. A system working from a transcript cannot genuinely hear hesitation, filler sounds, or pronunciation errors; at best it can infer them from transcription artifacts. The conservative scores of the transcript-based systems may therefore reflect an information deficit rather than stricter standards, a claim the authors frame as testable in future replications that score the same performances through all three input modalities in parallel.
Alternative explanations remain plausible, and the study is candid about them. The three systems differ in architecture and purpose: ChatGPT and Claude AI are general-purpose large language models capable of holistic rubric reasoning, while Otter AI’s scoring is a secondary feature built on top of transcription. Its score variability, roughly double that of the other two systems, is consistent with a shallower judgment process. Cultural and linguistic bias also looms over the results. None of the systems was calibrated for Indonesian English varieties, and prior research has shown that automated speech evaluators can penalize accents underrepresented in their training data, potentially disadvantaging test-takers from specific linguistic backgrounds. The authors stress, however, that their comparative design cannot isolate accent exposure as the causal mechanism.
The human side of the equation was no paragon of consistency either. The two human raters differed significantly from each other, averaging 6.03 versus 6.47 band points, with an intraclass correlation of only moderate strength. This transparency matters: part of the human-AI gap reflects genuine variability among human judges, and the human composite cannot be treated as an infallible gold standard. A more conservative non-parametric check complicated the picture further, as the Spearman rank correlation between human and AI composite scores was weak and non-significant, suggesting the two methods agree on the existence of a mean-level gap but less certainly on the relative ranking of individual candidates. Notably, IELTS speaking inter-rater reliability typically falls in the 0.7 to 0.8 range under standard conditions, and the two lecturers in this study were experienced teachers rather than certified IELTS examiners.
From these patterns the researchers build their central argument: the complementarity of human and machine judgment is an opportunity, not a problem. AI demonstrates superior consistency in measuring technical accuracy, the rule-governed features such as pronunciation metrics, grammatical structures, and speech rate that survive quantification. Humans excel at assessing communicative effectiveness, pragmatic intent, and the interactional contingencies of a live conversation. The study therefore proposes a hybrid pedagogical framework in which AI handles low-stakes practice and instantaneous technical feedback while teachers devote their time to high-value instruction on pragmatic and interactive speaking skills. In one proposed cycle, students perform a task on an AI platform, refine their delivery based on immediate feedback, and only then submit the performance to a teacher for holistic evaluation. Teachers can also mine aggregated AI data to diagnose class-wide weaknesses, for instance discovering that a large share of students score low on grammatical range in a particular task type.
The practical guidance for educators is pointed. Audio-input capability should be treated as a threshold requirement, not a convenience feature, when the goal includes pronunciation feedback, because systems that cannot hear the candidate cannot fairly score the criteria that depend on hearing. Transcript-based tools may still add value for lexical and grammatical feedback, where the relevant information survives transcription, but they should not be trusted for holistic proficiency judgments. Tool selection should also be purposeful and plural: a conservatively scoring system might serve as a demanding benchmark for technical accuracy, while a system closer to human norms might better simulate test conditions and build confidence. No single current platform replicates the full spectrum of human judgment.
The authors are equally clear about what their study does not establish. None of the systems was calibrated against certified IELTS rating standards, so the score differences may partly reflect differences in what each system was measuring rather than in accuracy on a shared scale. The sample of 30 Indonesian learners was not power-analyzed, and the hybrid framework itself remains a direction for future validation rather than an empirically tested model. Still, the core recommendation is well supported by the data: evaluations of AI-assisted assessment should always report results separately for each system, because pooling masks exactly the kind of system-specific divergence this study documented. The debate over whether AI or humans grade better, the authors conclude, is for the classroom a red herring. The consequential question is how to deploy each where it excels, letting machines process the data and deliver instant objective feedback while human teachers do what machines cannot: inspire, motivate, and teach the nuanced art of human communication.
Subject of Research: Comparison of human and AI scoring of IELTS speaking tests to inform hybrid pedagogy
Article Title: Aligning artificial intelligence and human judgment in IELTS speaking assessment to develop a hybrid pedagogical framework for test preparation
Article References: Aeni, N., Syawaluddin, A., & Strid, J. E. (2026). Aligning artificial intelligence and human judgment in IELTS speaking assessment to develop a hybrid pedagogical framework for test preparation. Discover Education, 5(1), Article 1008. https://doi.org/10.1007/s44217-026-02185-3
Image Credits: AI Generated
DOI: 10.1007/s44217-026-02185-3
Keywords: IELTS, artificial intelligence, language assessment, automated scoring, ChatGPT, Claude AI, Otter AI, speech recognition, EFL learners, hybrid pedagogy, formative assessment, rater reliability
Cite Scienmag News
Courtney Benton. (October 9, 2026). AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other. Scienmag. https://scienmag.com/ai-examiners-score-ielts-speaking-tests-differently-than-humans-and-differently-than-each-other/
Courtney Benton. "AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other." Scienmag, 9 October 2026, https://scienmag.com/ai-examiners-score-ielts-speaking-tests-differently-than-humans-and-differently-than-each-other/. Accessed 9 October 2026.
Courtney Benton. "AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other." Scienmag. October 9, 2026. https://scienmag.com/ai-examiners-score-ielts-speaking-tests-differently-than-humans-and-differently-than-each-other/

