<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI-driven evaluation of student work &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-driven-evaluation-of-student-work/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 01:36:13 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI-driven evaluation of student work &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Grades the Reflections: Can Language Models Code What Students Really Think?</title>
		<link>https://scienmag.com/ai-grades-the-reflections-can-language-models-code-what-students-really-think/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 01:36:13 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[AI education assessment]]></category>
		<category><![CDATA[AI literacy]]></category>
		<category><![CDATA[AI role-playing in learning]]></category>
		<category><![CDATA[AI tutors in graduate courses]]></category>
		<category><![CDATA[AI-based student reflections grading]]></category>
		<category><![CDATA[AI-driven evaluation of student work]]></category>
		<category><![CDATA[analyzing student opinions with AI]]></category>
		<category><![CDATA[Claude 3.7 Sonnet]]></category>
		<category><![CDATA[ethical implications of AI in education]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[generative AI in systems analysis]]></category>
		<category><![CDATA[GPT applications for educational research]]></category>
		<category><![CDATA[GPT-4o]]></category>
		<category><![CDATA[graduate education]]></category>
		<category><![CDATA[hybrid intelligent feedback]]></category>
		<category><![CDATA[impact of AI on graduate teaching methods]]></category>
		<category><![CDATA[instructional design]]></category>
		<category><![CDATA[inter-rater reliability]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in student feedback analysis]]></category>
		<category><![CDATA[qualitative research]]></category>
		<category><![CDATA[student reflections]]></category>
		<category><![CDATA[thematic coding]]></category>
		<category><![CDATA[trustworthiness of AI in education]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=232894</guid>

					<description><![CDATA[A new study finds that while graduate students reported highly positive experiences with course-specific AI tutors, large language models asked to code the students' own reflections agreed with human researchers far less reliably, especially on reflective and hypothetical prompts.]]></description>
										<content:encoded><![CDATA[<p>When a graduate course in systems analysis and design handed its students over to roughly thirty custom-built artificial intelligence tutors for an entire semester, the instructors knew the experiment would reveal something about how students learn with generative AI. What they did not anticipate was that the same technology would then be turned on the students&#8217; own words, in a methodological test that asks one of the most pressing questions in modern educational research: can large language models be trusted to analyze what students say about AI itself? A new exploratory study published in Discover Education by Viktors J. Muiznieks, Billie Anderson, and Tyler Price of Southern New Hampshire University offers one of the most candid answers yet, and the verdict is a carefully bounded yes.</p>
<p>The course at the heart of the study was no ordinary lecture series. Across a sixteen-week semester, twenty graduate students in a master&#8217;s-level information technology program worked with approximately thirty course-specific large language model applications, built through OpenAI&#8217;s custom GPT interface. Each application ran pre-scripted prompt sequences that cast the AI in structured instructional roles: simulator, role-player, Socratic dialogue partner, mentor, teammate, or chatbot. Students used these agents to conduct systems analysis, design user interfaces, model business cases, run multi-phase negotiation scenarios, and explore ethical considerations. Crucially, students accessed the tools through shared URLs without being able to view or modify the underlying prompt architecture, which meant the instructional design, not the students&#8217; prompting skill, determined how the AI behaved in each activity.</p>
<p>At the end of the semester, all twenty students completed an anonymous, Institutional Review Board-approved survey containing nine Likert-scale items and eight open-ended questions. The quantitative results were strikingly positive. Most responses to the engagement and motivation items fell into the Strongly Agree category, and fifteen students strongly agreed that generative AI helped them understand complex concepts. Most students also agreed or strongly agreed that the tools supported independent exploration and helped them work more efficiently. The most emphatic result of all concerned the future: eighteen of nineteen respondents selected Definitely Yes when asked whether they would continue using generative AI tools in future courses or projects.</p>
<p>But the researchers are careful, almost insistently so, about what those numbers do not mean. The survey was developed for this exploratory course evaluation and was never psychometrically validated as a scale. It was administered once, after a highly scaffolded, technology-intensive course, with no baseline measures, no comparison condition, and no objective assessments. Self-reported learning gains, the literature reminds us, often track satisfaction and motivation more closely than actual cognitive achievement. The study therefore treats its findings as evidence about students&#8217; perceived experiences, not as proof that the AI-supported course improved learning, critical thinking, retention, or transfer. Responses to the critical-thinking item were notably more varied than the rest, a hesitation that echoes recent meta-analytic work finding no statistically significant average effect of generative AI on metacognition even when other outcomes improve.</p>
<p>The genuinely novel contribution lies in the second half of the study, where the researchers staged a quiet contest between human judgment and machine judgment. Two human researchers independently read all the open-ended responses and, through iterative discussion, built a shared thematic codebook for each of the eight questions. Then two contemporary large language models, GPT-4o and Claude 3.7 Sonnet, were given the same codebook and asked to apply it to the same responses in the same binary format, with no example-coded cases and no corrective feedback. Agreement was quantified using Cohen&#8217;s kappa for each coder pair and Fleiss&#8217; kappa across all four coders, benchmarks interpreted through the classic Landis and Koch thresholds.</p>
<p>The results revealed a clear hierarchy of reliability. Human-human agreement ranged from moderate to almost perfect, with kappa values between 0.52 and 0.84. Average human-LLM agreement was lower and more variable, spanning 0.36 to 0.62, and agreement between the two models themselves ranged from 0.30 to 0.70. Across all eight questions, Claude 3.7 Sonnet aligned more closely with the human coders than GPT-4o did when the two human-model kappas were averaged, though the authors warn that this is a snapshot of two proprietary models under one prompt design and a single coding run, not a general ranking of model quality. Four-rater Fleiss&#8217; kappa ranged from 0.42 to 0.63, peaking on the most concrete prompts.</p>
<p>The pattern of where agreement flourished and where it collapsed is perhaps the most scientifically interesting finding. The strongest four-rater agreement, kappa of 0.63, occurred for a concrete retrospective question asking students to describe specific instances where AI helped them understand a topic or solve a problem. Responses to such questions contain observable task cues: coding, debugging, project work, interface design. The weakest agreement, kappa of 0.42, occurred on the two prospective prompts, one asking students to propose desired future features and another asking them to imagine a strategy for a difficult hypothetical assignment. These questions invite multiple reasonable levels of abstraction, from technical feature to learning process to emotional outcome, and the models fragmented broad human themes into narrower subcodes or missed human-identified themes altogether.</p>
<p>Some discrepancies were more than statistical noise. In one question about challenges encountered, both models treated statements reporting no challenge as a substantive theme, while the human coders recognized them as the absence of a challenge code, effectively assigning meaning to non-responses. The authors also point to codebook maturity as a confounding factor: the codebook was developed from the same small set of twenty students&#8217; responses, and human-human kappa was only moderate for four of the eight questions, meaning some theme boundaries were unstable even for trained researchers. Brief open-ended responses compound the difficulty, because meaning is often implicit, compressed, and socially situated. A student&#8217;s mention of working faster might signal efficiency, reduced anxiety, or creeping dependence, and disentangling those possibilities requires exactly the contextual judgment that current models lack.</p>
<p>The study frames its instructional findings through the hybrid intelligent feedback framework, which defines effective practice as a pedagogically designed combination of complementary human and artificial cognition. In the course, the LLM applications supplied immediate examples, stepwise prompts, role-play, and iterative suggestions, while the instructor set disciplinary goals, ethical boundaries, and verification expectations, and students remained responsible for evaluating and using what the AI produced. The authors apply this lens retrospectively and with deliberate restraint, distinguishing the pedagogical level of hybridity, where humans and AI shared instructional responsibility, from the methodological level, where humans and models shared a coding workflow. Conflating the two, they argue, would blur claims about classroom experience with claims about research reliability.</p>
<p>The practical upshot is a bounded division of labor rather than a replacement scenario. The researchers conclude that large language models can assist qualitative researchers with screening, discrepancy detection, and code suggestions, but that human researchers remain necessary for contextual interpretation, validation, and adjudication. They recommend that researchers establish mature codebooks with explicit definitions, positive and negative examples, and rules for non-substantive responses before any model comparison; document the provider, model version, prompts, and number of runs for auditability; and treat LLM coding as a reproducibility and sensitivity problem requiring multiple independent runs rather than a single definitive judgment. For educators, the study suggests a sequence of bounded purpose, structured interaction, comparison against disciplinary criteria, and closing reflection on what was accepted, revised, or rejected. In an era when students overwhelmingly intend to keep using these tools, the study&#8217;s most durable message may be that the technology works best when it is designed into the course, verified by the learner, and never allowed to have the last word on what learning means.</p>
<p><strong>Subject of Research:</strong> Using large language models to thematically code graduate students&#x27; reflections on generative AI in coursework</p>
<p><strong>Article Title:</strong> Using large language models to analyze student reflections on generative AI in graduate coursework</p>
<p><strong>Article References:</strong> Muiznieks, V. J., Anderson, B., &amp; Price, T. (2026). Using large language models to analyze student reflections on generative AI in graduate coursework. <em>Discover Education, 5</em>(1), Article 1065. <a href="https://doi.org/10.1007/s44217-026-02231-0" rel="noopener noreferrer">https://doi.org/10.1007/s44217-026-02231-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44217-026-02231-0" rel="noopener noreferrer">10.1007/s44217-026-02231-0</a></p>
<p><strong>Keywords:</strong> generative AI, large language models, graduate education, thematic coding, inter-rater reliability, qualitative research, student reflections, hybrid intelligent feedback, GPT-4o, Claude 3.7 Sonnet, instructional design, AI literacy</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">232894</post-id>	</item>
	</channel>
</rss>
