Every teacher knows the moment: a student solves one problem correctly, stumbles on the next, and within a handful of answers an experienced educator forms a mental map of what that learner actually understands. Replicating that intuition in software has been the goal of knowledge tracing, a field of educational data science that models a student’s mastery of underlying skills from their exercise records and predicts how they will perform on future questions. For a decade, deep learning has driven remarkable gains in this task, with recurrent networks, attention mechanisms, memory networks and graph-based architectures all pushing predictive accuracy upward. Yet a new study argues that these systems, for all their statistical power, have drifted away from the very scenario they were meant to serve: real classrooms, where teachers must judge students from limited evidence and then explain their reasoning in words.
That gap is the starting point for a research article published on 22 September 2025 in Frontiers of Digital Education, a Springer journal, by Haoxuan Li of Beihang University, Jifan Yu of Tsinghua University and colleagues. The team reformulates knowledge tracing as a new task they call explainable few-shot knowledge tracing, and proposes a cognition-guided framework built on large language models that can track a student’s knowledge state from only a few exercise records while producing natural language explanations of its conclusions. Across three widely used benchmark datasets, the authors report that large language models perform comparably to, or better than, competitive deep knowledge tracing methods, despite operating under conditions that would cripple conventional models.
To appreciate why this matters, it helps to understand how traditional knowledge tracing works. The field traces its lineage to Bayesian knowledge tracing, introduced by Albert Corbett and John Anderson in 1994, which treats each skill as a binary state, learned or unlearned, and updates the probability of mastery each time a student answers a question. From 2015 onward, deep knowledge tracing replaced these hand-built probabilistic updates with recurrent neural networks that learn hidden representations of student ability directly from long sequences of responses. Later refinements added self-attention, graph neural networks that exploit relationships between exercises, and memory networks with dynamic key-value stores. These models are powerful, but they share two structural dependencies: they need extensive interaction data per student to converge, and their output is a bare number, a predicted probability of a correct answer, with no account of why.
Both dependencies clash with teaching practice. A teacher assessing a student in a tutoring session sees perhaps a dozen attempts, not thousands, and must nonetheless form a judgment. That judgment is then communicated as feedback: the student confuses the distributive property with the associative property, or consistently misapplies a sign rule when moving terms across an equation. The Beihang and Tsinghua authors describe this mismatch as current methods falling into the cracks between laboratory benchmarks and real-world pedagogy. Their response is to redefine the task itself: instead of predicting a numerical score from abundant data, the model should infer a student’s cognitive state from sparse evidence and articulate that inference in language a teacher or student can read.
Large language models are, in a technical sense, an unexpected but fitting instrument for this reformulation. Models in the lineage of GLM and LLaMA are pretrained on vast text corpora and exhibit emergent abilities in reasoning and generation, meaning they can perform tasks they were never explicitly trained for. Crucially, they carry prior knowledge about academic subjects themselves: what a quadratic equation is, what concept a fraction problem tests, which misconceptions typically arise. A conventional deep knowledge tracing model sees only anonymized question identifiers and correctness bits; a language model sees the actual content of the exercise and can reason about the cognitive skill it probes. This allows the framework to lean on semantic understanding rather than sheer volume of interaction history, which is precisely what the few-shot setting demands.
The architecture the researchers propose is cognition-guided, a design choice that distinguishes it from simply prompting a chatbot with a transcript. The framework structures the model’s reasoning around cognitive dimensions of learning, guiding the language model to decompose a student’s performance into mastery of specific knowledge components rather than emitting an undifferentiated guess. The student’s few exercise records are rendered into a structured prompt, the model reasons over which underlying skills each question engages, and it then generates both a prediction of future performance and a natural language explanation tracing the evidence for its judgment. In effect, the explanation is not bolted on after the fact, as with post-hoc interpretability techniques applied to neural networks, but emerges from the same reasoning chain that produces the prediction.
The empirical evaluation is where the claim becomes concrete. The authors tested their framework on three widely used knowledge tracing datasets, benchmarking against competitive deep learning baselines drawn from the field’s standard toolkit, including models assembled through the PYKT benchmarking library. Under few-shot conditions, where each student contributes only a small number of records, the language model-based approach achieved results comparable to or superior than the deep baselines, which typically require far more data to reach their reported performance. The significance is twofold. Practically, it suggests that meaningful student modeling may be possible in cold-start scenarios, new courses, new platforms, or individual tutoring, where deep models have historically floundered. Scientifically, it demonstrates that the semantic knowledge embedded in pretrained language models can substitute, at least in part, for the statistical signal that large datasets normally provide.
The study does not present itself as a finished solution. The authors explicitly discuss potential directions and call for future improvements, acknowledging that the field is at an early stage of understanding how language models should be adapted to educational measurement. Open questions loom large. Language models can hallucinate, producing confident but wrong explanations, and in an educational context an incorrect explanation of a student’s misconception could misdirect instruction. The cost of running large models at scale across millions of learners is nontrivial compared with compact neural networks. And the psychometric tradition, from item response theory to the Rasch model, has spent a century building rigorous measurement theory that these new generative approaches have not yet absorbed. The authors’ framing, grounded in the established literature on educational assessment, suggests they see their work as a bridge between that tradition and modern generative AI rather than a replacement for it.
Even so, the implications ripple outward. Explainability is becoming a regulatory and ethical requirement for AI systems that make consequential decisions about people, and few decisions are more consequential than those shaping a student’s education. A model that can say, in plain language, that a student has mastered linear equations but struggles with word problems gives teachers something actionable; a probability output gives them nothing. The work also joins a broader wave of research applying large language models to education, from computerized adaptive testing to automated feedback on mathematics responses, indicating that the generative AI era may reshape educational assessment as profoundly as the deep learning era did. If the trend holds, the next generation of tutoring systems may not merely predict what a student will get right, but explain, in a teacher’s own vocabulary, what the student knows, what they are missing, and what to do next. That, the authors suggest, is the standard real teaching has always demanded, and one that machine learning is only now beginning to meet.
Subject of Research: Using large language models for explainable few-shot knowledge tracing in educational assessment
Article Title: Explainable Few-Shot Knowledge Tracing
Article References: Explainable Few-Shot Knowledge Tracing. (n.d.). https://doi.org/10.1007/s44366-025-0071-x
Image Credits: AI Generated
DOI: 10.1007/s44366-025-0071-x
Keywords: knowledge tracing, large language models, explainability, educational assessment, few-shot learning, student modeling, deep learning, intelligent tutoring systems, natural language explanations, educational data mining, cognitive modeling, personalized education
Cite Scienmag News
Courtney Benton. (September 30, 2026). AI Learns to Read Students’ Minds from Just a Handful of Answers. Scienmag. https://scienmag.com/ai-learns-to-read-students-minds-from-just-a-handful-of-answers/
Courtney Benton. "AI Learns to Read Students’ Minds from Just a Handful of Answers." Scienmag, 30 September 2026, https://scienmag.com/ai-learns-to-read-students-minds-from-just-a-handful-of-answers/. Accessed 30 September 2026.
Courtney Benton. "AI Learns to Read Students’ Minds from Just a Handful of Answers." Scienmag. September 30, 2026. https://scienmag.com/ai-learns-to-read-students-minds-from-just-a-handful-of-answers/








