Every online learning platform faces an awkward moment when it launches in a new subject area. The sophisticated algorithms that track which students have mastered which concepts are, in most cases, statistical machines that learn from history. They need thousands of recorded answers, correct and incorrect, before they can say anything meaningful about a learner’s knowledge. In a brand-new domain, that history simply does not exist, and the diagnostic engine sits idle. Researchers at Anhui University in Hefei, China, have now shown that large language models may be able to fill this gap almost entirely from scratch, diagnosing student abilities in a subject they have never seen interaction data for.
The team, led by Haiping Ma and Xingyi Zhang, frames the problem as zero-shot cross-domain cognitive diagnosis, abbreviated ZCCD. Cognitive diagnosis is the task of estimating a student’s mastery of fine-grained knowledge concepts, such as specific skills within mathematics or language learning, from their response logs. Classical diagnostic models, including neural approaches developed over the past several years, excel when they can train on abundant interaction records within a single domain. But when a platform expands into a new subject, there are no logs to learn from. The researchers call this the diagnostic system cold-start problem, and it has become one of the central bottlenecks in intelligent education as online learning expands rapidly across subjects and institutions.
Previous attempts at cold-start diagnosis have tried to engineer around the missing data. Some approaches transfer embeddings of knowledge concepts across domains using concept graphs, while others rely on a small batch of early students in the target domain to bootstrap the model. These methods help, but they still require some structural alignment between domains or some minimal target-domain data. The Anhui team asked a more radical question: could a general-purpose large language model, with no fine-tuning and no target-domain response logs at all, act as the bridge between domains? Their answer is a paradigm they call large language model-guided cognitive state transfer, or LCST.
The core insight of LCST is to recast cognitive diagnosis as a natural language task. Instead of representing a student’s knowledge state as a vector of latent parameters learned by a neural network, the method expresses it in words. The system prompts a large language model with descriptions of the knowledge concepts in the source domain, the student’s recorded performance on exercises covering those concepts, and descriptions of the concepts in the target domain. The model is then asked to reason about which target-domain abilities the student is likely to have mastered, given the profile of strengths and weaknesses observed in the source domain. In effect, the language model serves as an intermediary that interprets a learner’s cognitive state in one subject and projects it into another.
What makes this plausible is the way large language models encode relationships between concepts. Because these models are trained on vast corpora of educational and general text, they carry implicit knowledge about how skills relate to one another, for example that proficiency in algebraic manipulation tends to accompany proficiency in solving linear equations, or that grammatical understanding underpins reading comprehension. The researchers exploit this by having the model analyze the relationships between knowledge concepts in both domains and use those relationships to guide the transfer of mastery estimates. The approach draws on techniques such as chain-of-thought prompting, which encourages models to reason step by step rather than jumping to conclusions, and prompt engineering, which shapes the input format so the model can best apply its internal knowledge to the diagnostic task.
The team evaluated LCST on real-world datasets, including educational data spanning multiple subject domains, and compared it against existing cold-start and transfer-learning baselines for cognitive diagnosis. The results showed that the language-model-guided approach significantly improved diagnostic performance in the target domain compared with methods that lacked access to such semantic reasoning. Notably, the model achieved this without any prior interaction data from the target domain, which is precisely the condition under which conventional diagnostic models fail completely. The experiments used several prominent language models, including open-weight families such as Llama and Gemma as well as more capable proprietary systems, suggesting that the paradigm is not tied to a single proprietary model but reflects a general capability of modern language models.
The technical evaluation relied on standard metrics from the diagnostic modeling literature, including measures of how well predicted mastery patterns align with actual student performance, evaluated with area-under-the-curve style criteria commonly used to assess binary classification quality. The authors also situate their work within a rapidly growing body of research showing that large language models can act as zero-shot reasoners across many tasks, from ranking items in recommender systems to tracking dialogue states, without task-specific training. The cognitive diagnosis result extends this pattern to a domain where the underlying data, student response logs, are sparse, noisy, and structurally different from the text corpora the models were trained on.
The implications for education technology could be substantial. A platform that adds a new course area today typically needs weeks or months of usage before its recommendation and assessment engines become useful. If a language model can bootstrap reasonable diagnostic estimates immediately, platforms could deliver personalized exercise recommendations and ability assessments from day one, improving early learner engagement and reducing dropout during the vulnerable initial period. The approach also hints at a broader role for language models in education: rather than serving only as content generators or chat tutors, they may function as inferential engines that reason about learner cognition, a role the authors describe as language models acting as educational experts.
There are, of course, important caveats. The method depends on the quality and cultural coverage of the language model’s internal knowledge, and prior work has documented biases in how these models treat different topics and regions. Diagnostic estimates produced without any target-domain data will inevitably be less precise than those refined by actual interaction logs, so the most realistic deployment may treat zero-shot transfer as a starting point that is progressively corrected as real data accumulates. Interpretability is another consideration: because the model reasons in natural language, its justifications can in principle be inspected by educators, but those explanations must be validated rather than taken at face value. The Anhui team’s work, published in Frontiers of Digital Education and supported by the National Natural Science Foundation of China, nonetheless marks a striking demonstration that the semantic knowledge inside large language models can substitute, at least partially, for the statistical history that diagnostic systems have always required. As intelligent education systems spread to new subjects, languages, and regions, the ability to diagnose learners without waiting for data may prove to be one of the most consequential applications of language models in the classroom.
Subject of Research: Zero-shot cross-domain cognitive diagnosis of student knowledge using large language models
Article Title: Large Language Models Are Zero-Shot Cross-Domain Diagnosticians in Cognitive Diagnosis
Article References: Large Language Models Are Zero-Shot Cross-Domain Diagnosticians in Cognitive Diagnosis. (n.d.). https://doi.org/10.1007/s44366-025-0054-y
Image Credits: AI Generated
DOI: 10.1007/s44366-025-0054-y
Keywords: cognitive diagnosis, large language models, zero-shot learning, cold-start problem, intelligent education, prompt engineering, knowledge tracing, educational data mining, cross-domain transfer, AI in education, personalized learning, intelligent tutoring systems
Cite Scienmag News
Courtney Benton. (October 3, 2026). AI Tutors Without Training Wheels: Language Models Diagnose New Subjects Zero-Shot. Scienmag. https://scienmag.com/ai-tutors-without-training-wheels-language-models-diagnose-new-subjects-zero-shot/
Courtney Benton. "AI Tutors Without Training Wheels: Language Models Diagnose New Subjects Zero-Shot." Scienmag, 3 October 2026, https://scienmag.com/ai-tutors-without-training-wheels-language-models-diagnose-new-subjects-zero-shot/. Accessed 3 October 2026.
Courtney Benton. "AI Tutors Without Training Wheels: Language Models Diagnose New Subjects Zero-Shot." Scienmag. October 3, 2026. https://scienmag.com/ai-tutors-without-training-wheels-language-models-diagnose-new-subjects-zero-shot/

