Artificial intelligence chatbots have swept through classrooms with promises of personalized tutoring, but rigorous evidence about what they actually teach remains scarce. A new quasi-experimental study from Nigeria offers one of the clearest tests yet, and its results are strikingly two-sided. Researchers at the University of Nigeria, Nsukka, working with colleagues at the National University of Lesotho and the University of Pretoria, found that undergraduates who learned Java programming with Google Gemini embedded in a problem-based learning framework made dramatically larger gains in coding skill than peers taught with instructional videos. Yet the same AI-supported students showed no significant advantage in critical thinking or problem-solving, and the authors suggest that over-reliance on the chatbot may even have blunted deeper cognitive engagement.
The study, published in the Journal of New Approaches in Educational Research, recruited 62 second-year computer science education students, 42 women and 20 men aged 16 to 25, from two public universities in southeastern Nigeria. All had completed prerequisite programming courses and were enrolled in a Java-based course mandated by Nigeria’s National Universities Commission. Because intact classes rather than individuals were assigned to conditions, the researchers used a non-equivalent control group design, splitting participants into an experimental group of 25 students taught with AI-assisted problem-based learning and a control group of 37 taught with video-supported problem-based learning. Students were further classified by academic ability, using the 50th percentile of an eight-item reasoning test, and by age, creating a 2 by 2 by 2 factorial structure across eight distinct conditions.
The six-week intervention was carefully choreographed so that both groups covered identical curricular content while differing only in instructional support. After a joint orientation and baseline testing in week one, the AI group began using Google Gemini as a real-time coding assistant. In week two they tackled open-ended, real-world problems in small teams of roughly seven, prompting the AI to brainstorm, generate, and refine solution algorithms, then comparing its suggestions against their own. The video group, by contrast, watched expert-led tutorials on algorithmic design and worked through similar exercises under instructor supervision, relying on peer discussion rather than dynamic feedback.
The divergence deepened as the tasks grew harder. In week three, focused on program structuring and debugging, the AI students wrote Java programs and received immediate, AI-generated explanations of syntax and logic errors, with Gemini suggesting refactored code options that students iteratively corrected. The video group analyzed prerecorded debugging demonstrations and practiced replicating and fixing similar errors in the lab. Week four brought testing and simulation: the experimental group ran logic-based test cases on their scripts and evaluated output consistency with AI-guided feedback, while control students collaboratively tested sample programs and discussed modifications with peers and instructors. In week five, each AI-supported team built a functional Java application, such as a calculator, grading system, or login interface, with Gemini suggesting logic structures, flagging inefficiencies, and offering interface tips. Control students developed comparable applications but received instructor feedback only after submission.
When the researchers analyzed post-test scores with multivariate analysis of covariance, controlling for pre-test performance, the headline result was unmistakable. The instructional approach explained 43.4 percent of the multivariate variance across the combined outcomes, and univariate tests pinpointed computer programming skill as the driver: group membership produced a large effect on programming scores, F(1, 51) = 26.051, p < 0.001, with a partial eta squared of 0.338. Adjusted means told the same story. The AI-assisted group averaged 170.93 on the programming assessment against 145.37 for the video group, a Bonferroni-adjusted difference of 25.56 points with a 95 percent confidence interval running from 15.51 to 35.61. For a cohort of novices wrestling with one of the most notoriously difficult introductory subjects, that gap represents a substantial acceleration in skill acquisition.
The cognitive outcomes, however, refused to follow. On the critical thinking scale, the video group actually posted a marginally higher adjusted mean, 26.37 versus 24.16, though the difference fell short of significance at p = 0.107. Problem-solving scores were similarly indistinguishable, at 25.90 versus 24.94, p = 0.502. The authors interpret this asymmetry cautiously but pointedly, attributing the AI group’s weaker cognitive development to students’ dependence on the chatbot during instruction, reduced human interaction, and a focus on efficiency over learning. When Gemini could generate working code from a well-crafted prompt, students sometimes watched the program run in the Java IDE without mastering the underlying principles, a pattern the researchers describe as producing code rather than cultivating the ability to debug and reason through problems.
The study also revealed a significant interaction between teaching strategy and academic ability, Pillai’s Trace = 0.148, p = 0.048, indicating that the intervention’s effectiveness depended on who was using it. Estimated marginal means showed high-ability students outperforming low-ability peers across all three outcomes, with the widest margin in problem-solving, 26.40 versus 24.44. The researchers connect this to mathematical foundation and adaptability: students with stronger logical reasoning absorbed the AI-supported instruction more effectively, echoing broader findings that academic ability correlates with skill acquisition. Age, by contrast, produced no significant main effects or interactions, though descriptive patterns hinted that students over 20 scored slightly higher on programming and problem-solving, possibly reflecting cognitive readiness for abstract tasks, while students 20 or younger showed higher critical thinking scores, perhaps benefiting more from interactive technologies.
The theoretical framing helps explain why the AI worked as a coding scaffold but not as a thinking coach. Drawing on Dreyfus’s five-stage model of skill acquisition, the researchers structured tasks to move students from algorithm development through structured programming to full application building, mirroring the progression from novice to competent programmer. They also invoked the concepts of absorptive capacity and collective intelligence, arguing that AI-supported groups engage in an iterative loop of feedback, exploration, and knowledge construction in which the chatbot functions as a cognitive collaborator. That loop evidently accelerates procedural mastery, where immediate error correction and personalized hints reduce cognitive load. But the same immediacy may short-circuit the productive struggle, argumentation, and reflection that critical thinking measures were designed to capture, consistent with recent survey evidence that generative AI can reduce cognitive effort among knowledge workers.
The measurement apparatus lends the findings credibility. Programming skills were assessed with the 40-item Computer Programming Skills Assessment Rating Scale, spanning algorithm development, structured programming, the full program lifecycle, and software application development, with a Cronbach’s alpha of 0.91. The seven-item critical thinking scale and six-item problem-solving scale, both validated through Lawshe content validity procedures and pilot testing, achieved alphas of 0.81 and 0.87 respectively, while the academic ability test reached a KR-20 of 0.79. The team tested normality with both Shapiro-Wilk and Kolmogorov-Smirnov procedures, verified homogeneity of error variances with Levene’s test, and, because Box’s M indicated unequal covariance matrices, relied on Pillai’s Trace, the multivariate statistic most robust to that violation. Missing data under five percent were handled with expectation-maximization imputation.
The authors are candid about limitations: the quasi-experimental design without random assignment, the six-week window that may have been too short for higher-order cognitive gains to emerge, reliance on some self-report instruments, and a context specific to Java at two Nigerian universities. Still, the practical implications are concrete. They argue that AI tools like Gemini can serve as real-time scaffolds that close competence gaps for novice programmers, particularly in resource-limited settings, and they call for teacher training in digital pedagogies, national policy frameworks promoting AI literacy and equitable access, and investment in digital infrastructure. For educators everywhere wrestling with whether chatbots belong in the computer science classroom, the message is nuanced: AI-assisted problem-based learning can teach students to code faster than video alone, but cultivating independent thinkers still demands deliberate design, sustained human interaction, and more time than a single semester affords.
Subject of Research: Effect of AI-assisted problem-based learning on programming skills, critical thinking, and problem-solving in undergraduate Java education
Article Title: Fostering programming skill and critical thinking through AI-assisted PBL integration
Article References: Omeh, C. B., Ayanwale, M. A., Mnguni, L. E., & Olelewe, C. J. (2025). Fostering programming skill and critical thinking through AI-assisted PBL integration. Journal of New Approaches in Educational Research, 14(1), Article 22. https://doi.org/10.1007/s44322-025-00041-0
Image Credits: AI Generated
DOI: 10.1007/s44322-025-00041-0
Keywords: artificial intelligence in education, Google Gemini, problem-based learning, Java programming, critical thinking, problem-solving skills, quasi-experimental study, Nigeria, higher education, programming instruction, generative AI, computer science education
Cite Scienmag News
Courtney Benton. (October 1, 2026). AI Tutor Boosts Java Coding Skills but Not Critical Thinking, Study Finds. Scienmag. https://scienmag.com/ai-tutor-boosts-java-coding-skills-but-not-critical-thinking-study-finds/
Courtney Benton. "AI Tutor Boosts Java Coding Skills but Not Critical Thinking, Study Finds." Scienmag, 1 October 2026, https://scienmag.com/ai-tutor-boosts-java-coding-skills-but-not-critical-thinking-study-finds/. Accessed 1 October 2026.
Courtney Benton. "AI Tutor Boosts Java Coding Skills but Not Critical Thinking, Study Finds." Scienmag. October 1, 2026. https://scienmag.com/ai-tutor-boosts-java-coding-skills-but-not-critical-thinking-study-finds/

