Sunday, October 4, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Science Education

AI Grades the Reflections: Can Language Models Code What Students Really Think?

October 4, 2026
in Science Education
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 5 mins read
0
AI Grades the Reflections: Can Language Models Code What Students Really Think?

AI Grades the Reflections: Can Language Models Code What Students Really Think?

AI Grades the Reflections: Can Language Models Code What Students Really Think?

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

When a graduate course in systems analysis and design handed its students over to roughly thirty custom-built artificial intelligence tutors for an entire semester, the instructors knew the experiment would reveal something about how students learn with generative AI. What they did not anticipate was that the same technology would then be turned on the students’ own words, in a methodological test that asks one of the most pressing questions in modern educational research: can large language models be trusted to analyze what students say about AI itself? A new exploratory study published in Discover Education by Viktors J. Muiznieks, Billie Anderson, and Tyler Price of Southern New Hampshire University offers one of the most candid answers yet, and the verdict is a carefully bounded yes.

The course at the heart of the study was no ordinary lecture series. Across a sixteen-week semester, twenty graduate students in a master’s-level information technology program worked with approximately thirty course-specific large language model applications, built through OpenAI’s custom GPT interface. Each application ran pre-scripted prompt sequences that cast the AI in structured instructional roles: simulator, role-player, Socratic dialogue partner, mentor, teammate, or chatbot. Students used these agents to conduct systems analysis, design user interfaces, model business cases, run multi-phase negotiation scenarios, and explore ethical considerations. Crucially, students accessed the tools through shared URLs without being able to view or modify the underlying prompt architecture, which meant the instructional design, not the students’ prompting skill, determined how the AI behaved in each activity.

At the end of the semester, all twenty students completed an anonymous, Institutional Review Board-approved survey containing nine Likert-scale items and eight open-ended questions. The quantitative results were strikingly positive. Most responses to the engagement and motivation items fell into the Strongly Agree category, and fifteen students strongly agreed that generative AI helped them understand complex concepts. Most students also agreed or strongly agreed that the tools supported independent exploration and helped them work more efficiently. The most emphatic result of all concerned the future: eighteen of nineteen respondents selected Definitely Yes when asked whether they would continue using generative AI tools in future courses or projects.

But the researchers are careful, almost insistently so, about what those numbers do not mean. The survey was developed for this exploratory course evaluation and was never psychometrically validated as a scale. It was administered once, after a highly scaffolded, technology-intensive course, with no baseline measures, no comparison condition, and no objective assessments. Self-reported learning gains, the literature reminds us, often track satisfaction and motivation more closely than actual cognitive achievement. The study therefore treats its findings as evidence about students’ perceived experiences, not as proof that the AI-supported course improved learning, critical thinking, retention, or transfer. Responses to the critical-thinking item were notably more varied than the rest, a hesitation that echoes recent meta-analytic work finding no statistically significant average effect of generative AI on metacognition even when other outcomes improve.

The genuinely novel contribution lies in the second half of the study, where the researchers staged a quiet contest between human judgment and machine judgment. Two human researchers independently read all the open-ended responses and, through iterative discussion, built a shared thematic codebook for each of the eight questions. Then two contemporary large language models, GPT-4o and Claude 3.7 Sonnet, were given the same codebook and asked to apply it to the same responses in the same binary format, with no example-coded cases and no corrective feedback. Agreement was quantified using Cohen’s kappa for each coder pair and Fleiss’ kappa across all four coders, benchmarks interpreted through the classic Landis and Koch thresholds.

The results revealed a clear hierarchy of reliability. Human-human agreement ranged from moderate to almost perfect, with kappa values between 0.52 and 0.84. Average human-LLM agreement was lower and more variable, spanning 0.36 to 0.62, and agreement between the two models themselves ranged from 0.30 to 0.70. Across all eight questions, Claude 3.7 Sonnet aligned more closely with the human coders than GPT-4o did when the two human-model kappas were averaged, though the authors warn that this is a snapshot of two proprietary models under one prompt design and a single coding run, not a general ranking of model quality. Four-rater Fleiss’ kappa ranged from 0.42 to 0.63, peaking on the most concrete prompts.

The pattern of where agreement flourished and where it collapsed is perhaps the most scientifically interesting finding. The strongest four-rater agreement, kappa of 0.63, occurred for a concrete retrospective question asking students to describe specific instances where AI helped them understand a topic or solve a problem. Responses to such questions contain observable task cues: coding, debugging, project work, interface design. The weakest agreement, kappa of 0.42, occurred on the two prospective prompts, one asking students to propose desired future features and another asking them to imagine a strategy for a difficult hypothetical assignment. These questions invite multiple reasonable levels of abstraction, from technical feature to learning process to emotional outcome, and the models fragmented broad human themes into narrower subcodes or missed human-identified themes altogether.

Some discrepancies were more than statistical noise. In one question about challenges encountered, both models treated statements reporting no challenge as a substantive theme, while the human coders recognized them as the absence of a challenge code, effectively assigning meaning to non-responses. The authors also point to codebook maturity as a confounding factor: the codebook was developed from the same small set of twenty students’ responses, and human-human kappa was only moderate for four of the eight questions, meaning some theme boundaries were unstable even for trained researchers. Brief open-ended responses compound the difficulty, because meaning is often implicit, compressed, and socially situated. A student’s mention of working faster might signal efficiency, reduced anxiety, or creeping dependence, and disentangling those possibilities requires exactly the contextual judgment that current models lack.

The study frames its instructional findings through the hybrid intelligent feedback framework, which defines effective practice as a pedagogically designed combination of complementary human and artificial cognition. In the course, the LLM applications supplied immediate examples, stepwise prompts, role-play, and iterative suggestions, while the instructor set disciplinary goals, ethical boundaries, and verification expectations, and students remained responsible for evaluating and using what the AI produced. The authors apply this lens retrospectively and with deliberate restraint, distinguishing the pedagogical level of hybridity, where humans and AI shared instructional responsibility, from the methodological level, where humans and models shared a coding workflow. Conflating the two, they argue, would blur claims about classroom experience with claims about research reliability.

The practical upshot is a bounded division of labor rather than a replacement scenario. The researchers conclude that large language models can assist qualitative researchers with screening, discrepancy detection, and code suggestions, but that human researchers remain necessary for contextual interpretation, validation, and adjudication. They recommend that researchers establish mature codebooks with explicit definitions, positive and negative examples, and rules for non-substantive responses before any model comparison; document the provider, model version, prompts, and number of runs for auditability; and treat LLM coding as a reproducibility and sensitivity problem requiring multiple independent runs rather than a single definitive judgment. For educators, the study suggests a sequence of bounded purpose, structured interaction, comparison against disciplinary criteria, and closing reflection on what was accepted, revised, or rejected. In an era when students overwhelmingly intend to keep using these tools, the study’s most durable message may be that the technology works best when it is designed into the course, verified by the learner, and never allowed to have the last word on what learning means.

Subject of Research: Using large language models to thematically code graduate students' reflections on generative AI in coursework

Article Title: Using large language models to analyze student reflections on generative AI in graduate coursework

Article References: Muiznieks, V. J., Anderson, B., & Price, T. (2026). Using large language models to analyze student reflections on generative AI in graduate coursework. Discover Education, 5(1), Article 1065. https://doi.org/10.1007/s44217-026-02231-0

Image Credits: AI Generated

DOI: 10.1007/s44217-026-02231-0

Keywords: generative AI, large language models, graduate education, thematic coding, inter-rater reliability, qualitative research, student reflections, hybrid intelligent feedback, GPT-4o, Claude 3.7 Sonnet, instructional design, AI literacy

Cite Scienmag News

Courtney Benton. (October 4, 2026). AI Grades the Reflections: Can Language Models Code What Students Really Think? Scienmag. https://scienmag.com/ai-grades-the-reflections-can-language-models-code-what-students-really-think/

Courtney Benton. "AI Grades the Reflections: Can Language Models Code What Students Really Think?" Scienmag, 4 October 2026, https://scienmag.com/ai-grades-the-reflections-can-language-models-code-what-students-really-think/. Accessed 4 October 2026.

Courtney Benton. "AI Grades the Reflections: Can Language Models Code What Students Really Think?" Scienmag. October 4, 2026. https://scienmag.com/ai-grades-the-reflections-can-language-models-code-what-students-really-think/

Tags: AI education assessmentAI literacyAI role-playing in learningAI tutors in graduate coursesAI-based student reflections gradingAI-driven evaluation of student workanalyzing student opinions with AIClaude 3.7 Sonnetethical implications of AI in educationgenerative AIgenerative AI in systems analysisGPT applications for educational researchGPT-4ograduate educationhybrid intelligent feedbackimpact of AI on graduate teaching methodsinstructional designinter-rater reliabilitylarge language modelslarge language models in student feedback analysisqualitative researchstudent reflectionsthematic codingtrustworthiness of AI in education
Share26Tweet16
Previous Post

Horned Melon Extract Shows Liver Risk in Mice After Prolonged Use

Next Post

From Coal Tar to OLEDs: The Named Reactions That Build Carbazole

Related Posts

Refugee Women on Lesbos Redefine Sexual and Reproductive Health as a Matter of Justice
Science Education

Refugee Women on Lesbos Redefine Sexual and Reproductive Health as a Matter of Justice

October 4, 2026
HKU Faculty of Education Climbs to Second Place Worldwide in 2026 ShanghaiRanking Subject League Table
Science Education

HKU Faculty of Education Climbs to Second Place Worldwide in 2026 ShanghaiRanking Subject League Table

October 4, 2026
Soft-Embalmed Cadavers Win Over Surgeons-in-Training in Malaysian Study
Science Education

Soft-Embalmed Cadavers Win Over Surgeons-in-Training in Malaysian Study

October 3, 2026
Anxious Words: How Learners’ Speech Reveals Fear in Online and In-Person English Classes
Science Education

Anxious Words: How Learners’ Speech Reveals Fear in Online and In-Person English Classes

October 3, 2026
AI Genomics Pioneer from Hong Kong Wins 2026 APEC ASPIRE Prize
Science Education

AI Genomics Pioneer from Hong Kong Wins 2026 APEC ASPIRE Prize

October 3, 2026
What Really Drives Health Inequity in the Americas? A Landmark Review Maps the Engines of Power
Science Education

What Really Drives Health Inequity in the Americas? A Landmark Review Maps the Engines of Power

October 3, 2026
Next Post
From Coal Tar to OLEDs: The Named Reactions That Build Carbazole

From Coal Tar to OLEDs: The Named Reactions That Build Carbazole

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • From Coal Tar to OLEDs: The Named Reactions That Build Carbazole
  • AI Grades the Reflections: Can Language Models Code What Students Really Think?
  • Horned Melon Extract Shows Liver Risk in Mice After Prolonged Use
  • Why Some People Approve of Lynching: Moral Reasoning Holds the Key

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading