Friday, October 9, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Science Education

AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other

October 9, 2026
in Science Education
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 5 mins read
0
AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other

AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

When artificial intelligence systems are asked to grade the same spoken English exam that experienced human examiners grade, the results are anything but uniform. A new study published in Discover Education has found that the much-discussed gap between AI and human scoring in high-stakes language assessment is not a property of artificial intelligence in general, but of specific systems and, crucially, of how they receive the test itself. The findings suggest that the future of test preparation lies not in choosing between human and machine judgment, but in deliberately splitting the work between them.

Researchers led by Nur Aeni of Universitas Negeri Makassar in Indonesia recorded the IELTS-style speaking performances of 30 English-as-a-foreign-language university students, each completing the full three-part interview format: an introduction and interview, an individual long turn, and a two-way discussion. The identical recordings were then scored by two experienced English-speaking lecturers and by three commercially available AI platforms: ChatGPT, Otter AI, and Claude AI. Each system received a standardized prompt instructing it to apply the official IELTS Speaking Band Descriptors across the four rated competencies of fluency and coherence, lexical resource, grammatical range and accuracy, and pronunciation. No system was fine-tuned or custom-trained for the task.

When the three AI scores were averaged into a single composite, the picture looked familiar to anyone following the debate over automated assessment. Human raters assigned significantly higher scores overall, with a mean difference of 0.522 band points, a result that held up at the p < .001 level with a large effect size. The correlation between the human and AI composites was moderate and positive (r = .497, p = .005), meaning the two methods agreed broadly on the ranking of candidates but shared only about a quarter of their scoring variance. Twenty-eight of the thirty students received lower scores from the machines than from the humans, and the gap widened for the strongest speakers, with automated scores rarely exceeding 7.0 while human ratings frequently approached 8.0.

But the composite figure concealed the study’s most striking discovery. When the researchers disaggregated the results system by system, ChatGPT’s scores turned out to be statistically indistinguishable from those of one of the two human raters, with an effect size near zero. Otter AI and, especially, Claude AI diverged substantially from both humans, with Claude AI differing from the more lenient rater by nearly three standard deviations. A repeated-measures ANOVA across all five raters confirmed the pattern, and pairwise correlations told the same story: ChatGPT correlated significantly with both human raters, while Claude AI correlated with neither and barely overlapped with Otter AI. Treating AI scoring as a single undifferentiated category, the authors argue, obscures more than it reveals.

What explains the divergence? The researchers propose a deceptively simple answer: input modality. ChatGPT received the audio recordings directly through its voice input function, giving it access to the same acoustic information humans use, including intonation, stress placement, pausing, and articulation quality. Claude AI could not accept audio at the time of data collection, so recordings were first transcribed by a separate application, and Otter AI, a transcription platform by design, also scored from text. A system working from a transcript cannot genuinely hear hesitation, filler sounds, or pronunciation errors; at best it can infer them from transcription artifacts. The conservative scores of the transcript-based systems may therefore reflect an information deficit rather than stricter standards, a claim the authors frame as testable in future replications that score the same performances through all three input modalities in parallel.

Alternative explanations remain plausible, and the study is candid about them. The three systems differ in architecture and purpose: ChatGPT and Claude AI are general-purpose large language models capable of holistic rubric reasoning, while Otter AI’s scoring is a secondary feature built on top of transcription. Its score variability, roughly double that of the other two systems, is consistent with a shallower judgment process. Cultural and linguistic bias also looms over the results. None of the systems was calibrated for Indonesian English varieties, and prior research has shown that automated speech evaluators can penalize accents underrepresented in their training data, potentially disadvantaging test-takers from specific linguistic backgrounds. The authors stress, however, that their comparative design cannot isolate accent exposure as the causal mechanism.

The human side of the equation was no paragon of consistency either. The two human raters differed significantly from each other, averaging 6.03 versus 6.47 band points, with an intraclass correlation of only moderate strength. This transparency matters: part of the human-AI gap reflects genuine variability among human judges, and the human composite cannot be treated as an infallible gold standard. A more conservative non-parametric check complicated the picture further, as the Spearman rank correlation between human and AI composite scores was weak and non-significant, suggesting the two methods agree on the existence of a mean-level gap but less certainly on the relative ranking of individual candidates. Notably, IELTS speaking inter-rater reliability typically falls in the 0.7 to 0.8 range under standard conditions, and the two lecturers in this study were experienced teachers rather than certified IELTS examiners.

From these patterns the researchers build their central argument: the complementarity of human and machine judgment is an opportunity, not a problem. AI demonstrates superior consistency in measuring technical accuracy, the rule-governed features such as pronunciation metrics, grammatical structures, and speech rate that survive quantification. Humans excel at assessing communicative effectiveness, pragmatic intent, and the interactional contingencies of a live conversation. The study therefore proposes a hybrid pedagogical framework in which AI handles low-stakes practice and instantaneous technical feedback while teachers devote their time to high-value instruction on pragmatic and interactive speaking skills. In one proposed cycle, students perform a task on an AI platform, refine their delivery based on immediate feedback, and only then submit the performance to a teacher for holistic evaluation. Teachers can also mine aggregated AI data to diagnose class-wide weaknesses, for instance discovering that a large share of students score low on grammatical range in a particular task type.

The practical guidance for educators is pointed. Audio-input capability should be treated as a threshold requirement, not a convenience feature, when the goal includes pronunciation feedback, because systems that cannot hear the candidate cannot fairly score the criteria that depend on hearing. Transcript-based tools may still add value for lexical and grammatical feedback, where the relevant information survives transcription, but they should not be trusted for holistic proficiency judgments. Tool selection should also be purposeful and plural: a conservatively scoring system might serve as a demanding benchmark for technical accuracy, while a system closer to human norms might better simulate test conditions and build confidence. No single current platform replicates the full spectrum of human judgment.

The authors are equally clear about what their study does not establish. None of the systems was calibrated against certified IELTS rating standards, so the score differences may partly reflect differences in what each system was measuring rather than in accuracy on a shared scale. The sample of 30 Indonesian learners was not power-analyzed, and the hybrid framework itself remains a direction for future validation rather than an empirically tested model. Still, the core recommendation is well supported by the data: evaluations of AI-assisted assessment should always report results separately for each system, because pooling masks exactly the kind of system-specific divergence this study documented. The debate over whether AI or humans grade better, the authors conclude, is for the classroom a red herring. The consequential question is how to deploy each where it excels, letting machines process the data and deliver instant objective feedback while human teachers do what machines cannot: inspire, motivate, and teach the nuanced art of human communication.

Subject of Research: Comparison of human and AI scoring of IELTS speaking tests to inform hybrid pedagogy

Article Title: Aligning artificial intelligence and human judgment in IELTS speaking assessment to develop a hybrid pedagogical framework for test preparation

Article References: Aeni, N., Syawaluddin, A., & Strid, J. E. (2026). Aligning artificial intelligence and human judgment in IELTS speaking assessment to develop a hybrid pedagogical framework for test preparation. Discover Education, 5(1), Article 1008. https://doi.org/10.1007/s44217-026-02185-3

Image Credits: AI Generated

DOI: 10.1007/s44217-026-02185-3

Keywords: IELTS, artificial intelligence, language assessment, automated scoring, ChatGPT, Claude AI, Otter AI, speech recognition, EFL learners, hybrid pedagogy, formative assessment, rater reliability

Cite Scienmag News

Courtney Benton. (October 9, 2026). AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other. Scienmag. https://scienmag.com/ai-examiners-score-ielts-speaking-tests-differently-than-humans-and-differently-than-each-other/

Courtney Benton. "AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other." Scienmag, 9 October 2026, https://scienmag.com/ai-examiners-score-ielts-speaking-tests-differently-than-humans-and-differently-than-each-other/. Accessed 9 October 2026.

Courtney Benton. "AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other." Scienmag. October 9, 2026. https://scienmag.com/ai-examiners-score-ielts-speaking-tests-differently-than-humans-and-differently-than-each-other/

Tags: AI system reliability in language assessmentAI-based IELTS speaking test scoringand Claude AI in language testingArtificial Intelligenceautomated scoringchallenges in automated spoken English evaluationChatGPTClaude AIcomparison of ChatGPTdifferences between human and AI examinersEFL learnersformative assessmentfuture of language assessment with AI integrationhybrid pedagogyIELTSimpact of test administration on AI scoringimplications of AI and human hybrid gradinginfluence of test input quality on AI scoring accuracylanguage assessmentlimitations of AI in evaluating pronunciation and fluencyOtter AIrater reliabilityscoring consistency across AI platformsspeech recognitionstandardization of AI scoring
Share26Tweet16
Previous Post

Massive New Database Catalogs China’s 5,142 Reservoirs in Unprecedented Detail

Next Post

Nuclear Fallout Fingerprints Reveal a Shifting, Warming Arctic Gateway

Related Posts

Three-Way Data Detective Work Exposes Hidden CT Curriculum Gap Behind Licensure Slump
Science Education

Three-Way Data Detective Work Exposes Hidden CT Curriculum Gap Behind Licensure Slump

October 9, 2026
How Victorian geologists cracked Cornwall’s mysterious slab of ancient ocean crust
Science Education

How Victorian geologists cracked Cornwall’s mysterious slab of ancient ocean crust

October 9, 2026
Supermarket Sweep for the Green Deal: Festival Toolkit Puts Critical Raw Materials in Shoppers’ Baskets
Earth Science

Supermarket Sweep for the Green Deal: Festival Toolkit Puts Critical Raw Materials in Shoppers’ Baskets

October 9, 2026
Wartime Ukraine shows why education must count as scientific infrastructure
Science Education

Wartime Ukraine shows why education must count as scientific infrastructure

October 9, 2026
Training Frontline Health Workers for Migrant Mothers and Children in Central America
Science Education

Training Frontline Health Workers for Migrant Mothers and Children in Central America

October 9, 2026
After a Virtual Trip Through the Solar System, Students Rate Their Digital Skills
Science Education

After a Virtual Trip Through the Solar System, Students Rate Their Digital Skills

October 9, 2026
Next Post
Nuclear Fallout Fingerprints Reveal a Shifting, Warming Arctic Gateway

Nuclear Fallout Fingerprints Reveal a Shifting, Warming Arctic Gateway

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Nuclear Fallout Fingerprints Reveal a Shifting, Warming Arctic Gateway
  • AI Examiners Score IELTS Speaking Tests Differently Than Humans — And Differently Than Each Other
  • Massive New Database Catalogs China’s 5,142 Reservoirs in Unprecedented Detail
  • Why UK spinal cord injury rehab is a missed window for exercise research

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading