In a modest computer laboratory at a public girls’ school in Sahiwal, Pakistan, fifty ninth-grade students recently took part in an experiment that may reshape how English is taught in classrooms where speaking practice is scarce. Over twelve weeks, half of the students practiced spoken English with ChatGPT, talking to the AI through voice and text, answering its follow-up questions, and receiving instant feedback on their responses. The other half continued with conventional instruction. When researchers compared the two groups afterward, the difference was striking: the ChatGPT group had climbed from the lowest rungs of the international proficiency ladder to intermediate territory, while their classmates had barely moved. The study, published in Discover Education, offers some of the most concrete experimental evidence yet that a general-purpose chatbot, when embedded in a carefully designed curriculum, can transform oral language skills.
The stakes are high in a country where spoken English opens doors to social, academic, and professional opportunity, yet where proficiency remains stubbornly low. Pakistan ranked 67th among 116 countries in the Education First English Proficiency Report, and earlier research by the same team found that 88 percent of secondary school students performed at only the most basic speaking levels, classified as A1 or A2 on the Common European Framework of Reference (CEFR). Just 12 percent reached the pre-intermediate B1 level, and not a single student achieved advanced proficiency. The reason, the researchers argue, is structural: Pakistani public schools emphasize rote memorization and written exercises, leaving students with years of English instruction but almost no meaningful oral practice.
The study was grounded in Vygotsky’s Sociocultural Theory, which holds that learners advance through interaction with a More Knowledgeable Other within their Zone of Proximal Development, the space between what a learner can do alone and what they can achieve with guidance. In this framing, ChatGPT served as a scaffold: an endlessly patient conversational partner that adjusted its questions to the student’s level, prompted elaboration, and corrected errors in real time. After each individual practice session, students joined collaborative classroom activities such as role plays, pair work, and group discussions, allowing them to consolidate what the AI practice had unlocked.
The intervention itself was built with the ADDIE instructional design model and aligned with the Grade 9 English syllabus. Each week began with a five-minute teacher briefing, followed by sixty minutes of ChatGPT-assisted speaking practice in the school’s IT laboratory, and closed with twenty-five minutes of collaborative learning. Students used standardized prompts to steer the conversation. In one documented exchange on the topic of child labor, a student first asked ChatGPT to act as an English tutor and provide a word bank of useful vocabulary, then asked factual questions about the topic, and finally instructed the AI to act as a speaking partner, asking questions one at a time and giving brief feedback after each answer. Crucially, students kept the same conversation thread throughout the twelve weeks, so the AI retained the history of prior interactions and could build on them.
Measuring the outcome required a rigorous assessment framework. The researchers developed a Spoken English Proficiency Test modeled loosely on the IELTS structure, comprising a warm-up conversation, a four-to-five-minute monologue, and a five-to-six-minute dialogue. Each task was scored against CEFR descriptors across five dimensions: range, the breadth of vocabulary and expression; accuracy, grammatical correctness; fluency, the smoothness of delivery; interaction, the ability to engage a conversational partner; and coherence, the logical organization of speech. Each component was scored from 1 to 6, corresponding to CEFR levels A1 through C2, for a maximum total of 60 points. The instrument showed exceptional internal consistency, with a Cronbach’s alpha of 0.964.
In a novel twist, the assessment itself was conducted by ChatGPT. The researchers uploaded the CEFR scale, the marking scheme, and the test materials, then prompted the AI to act as a CEFR English-speaking examiner. Students spoke directly to ChatGPT in Voice Mode while the researcher supervised, and the AI generated real-time transcripts and scored the speaking components. Human language experts, blinded to group allocation, oversaw quality control and independently scored the interaction component of the monologue task, since ChatGPT could not evaluate non-verbal engagement such as facial expressions and gestures. This human-AI collaborative approach aimed to combine standardized, consistent scoring with expert judgment.
The results were unambiguous. Before the intervention, the two groups were statistically equivalent, with all students clustered in the basic A1-A2 range. After twelve weeks, the experimental group’s mean overall score reached 36.56, compared with 17.60 for the control group, a difference the Mann-Whitney U test confirmed as highly significant, with U = 1.50, z = -6.08, and p < .001. The effect size of r = .86 is unusually large for an educational intervention. Perhaps more telling is the raw gain: the ChatGPT group improved by an average of 20.44 points, while the control group gained just 1.04. On the CEFR scale, experimental students progressed from Basic User to Independent User levels, with one learner reaching C1, the threshold of Proficient User.
The improvements held across every measured dimension and both task types. In monologue performance, where students spoke independently on topics such as their dream job, the experimental group significantly outperformed the control group on range, accuracy, fluency, interaction, and coherence, all at p < .001. The same pattern emerged in dialogue tasks, where students conversed on everyday topics such as weekend plans. This across-the-board pattern matters because previous studies of ChatGPT in language education tended to focus on motivation, confidence, or overall proficiency; this study is among the first to isolate and document gains in the specific qualitative dimensions that define spoken competence under an internationally recognized framework.
The findings align with a growing body of international evidence. Studies in other contexts have reported that interactions with ChatGPT improve oral speaking performance, that consistent engagement with AI-assisted speaking activities produces greater gains than sporadic use, and that twelve-week chatbot interventions help learners outperform traditionally taught peers. Researchers have also documented improvements in vocabulary use, grammatical accuracy, and fluency from repeated AI-supported practice. Yet the literature carries caveats: systematic reviews warn about ethical concerns, uneven quality of AI-generated feedback, and the fact that chatbots cannot fully replicate human interaction. Effectiveness, the evidence suggests, depends less on the technology itself than on pedagogical integration, teacher guidance, and instructional design, all of which were central to the Pakistani module.
The authors are careful about the limits of their claims. The sample comprised fifty female students from a single urban public school, a homogeneous group in age, socio-economic background, and educational history, which may have reduced variability and inflated effect sizes. The twelve-week duration cannot speak to long-term retention, and the researchers recommend delayed post-tests, larger and more diverse samples, and mixed-methods designs to capture how the learning actually happens. The theoretical interpretation, that ChatGPT provided scaffolding within students’ zones of proximal development, remains a plausible explanation rather than a directly demonstrated mechanism. Still, the practical implications are hard to ignore. For classrooms where one teacher faces dozens of students and individual speaking practice is a logistical impossibility, an AI conversational partner offers a scalable supplement, and the study suggests that policymakers weighing digital infrastructure and teacher training for AI-supported language education now have experimental evidence to inform those decisions.
Subject of Research: Effectiveness of a ChatGPT-assisted learning module for improving English-speaking proficiency among secondary school students
Article Title: Integration of ChatGPT to enhance English speaking skills among secondary school students in Pakistan
Article References: Nageen, S., Khawaja, H. U. H., Sarwar, M., Moin, M., Imran, Z., & Jabeen, S. (2026). Integration of ChatGPT to enhance English speaking skills among secondary school students in Pakistan. Discover Education, 5(1), Article 1137. https://doi.org/10.1007/s44217-026-02278-z
Image Credits: AI Generated
DOI: 10.1007/s44217-026-02278-z
Keywords: ChatGPT, artificial intelligence, English speaking proficiency, CEFR, Pakistan, secondary education, language learning, experimental study, educational technology, Vygotsky, scaffolding, randomized controlled trial
Cite Scienmag News
Courtney Benton. (October 7, 2026). ChatGPT Tutoring Lifts Pakistani Students’ English Speaking by Two CEFR Levels. Scienmag. https://scienmag.com/chatgpt-tutoring-lifts-pakistani-students-english-speaking-by-two-cefr-levels/
Courtney Benton. "ChatGPT Tutoring Lifts Pakistani Students’ English Speaking by Two CEFR Levels." Scienmag, 7 October 2026, https://scienmag.com/chatgpt-tutoring-lifts-pakistani-students-english-speaking-by-two-cefr-levels/. Accessed 7 October 2026.
Courtney Benton. "ChatGPT Tutoring Lifts Pakistani Students’ English Speaking by Two CEFR Levels." Scienmag. October 7, 2026. https://scienmag.com/chatgpt-tutoring-lifts-pakistani-students-english-speaking-by-two-cefr-levels/

