A single midterm exam at Brown University has become the flashpoint in a battle that is quietly dismantling a century of thinking about how schools measure learning. In spring 2026, a professor suspected that his take-home midterm had been answered largely by ChatGPT: the 77 students in the class averaged an astonishing 96 percent, and many answers reproduced the chatbot’s idiosyncratic reasoning almost verbatim. When the professor replaced the final exam with an in-person version, eighteen students dropped the course, the average collapsed to below 50 percent, and nineteen students failed. The episode, recounted in an editorial by Milo Koretsky of Tufts University and colleagues in the International Journal of STEM Education, is not merely a story about cheating. It is evidence that the basic inference at the heart of every exam, that the submitted work reveals what a student actually knows, has broken down across large parts of higher education.
The editorial, written with Meixia Ding of Temple University, Thomas Chiu of the Chinese University of Hong Kong, Jonas Hallström of Linköping University, and Yeping Li of Texas A&M University, argues that the response to generative artificial intelligence must be framed as a problem of validity rather than a problem of policing. Assessment researchers have long described measurement in education through the assessment triangle, a framework from the National Research Council with three vertices: cognition, the learning goals defining what students should know and be able to do; observation, the tasks that elicit evidence of that learning; and interpretation, the inferences drawn from the collected evidence. Generative AI destabilizes every vertex simultaneously, and the authors used the triangle as a lens for a scoping review spanning the Scopus, Web of Science, and ERIC databases.
The technical heart of the problem sits on the observation vertex. Generative AI disrupts the assumption that a polished product is credible evidence of student capability, because a chatbot such as ChatGPT or Claude can often produce correct solutions to narrowly defined problems. Empirical surveys confirm that educators and students view the validity threat as most acute for essays, take-home exams, and computer programming assignments, formats where a coherent response can be generated with little visible trace of independent thinking. Interestingly, the threat runs in both directions. Restrictive anti-cheating controls, such as converting every flexible assignment into a high-stakes proctored exam, can reduce accessibility, authenticity, and alignment with professional practice, thereby weakening validity from the opposite side. What matters is not AI use in the abstract but how a specific use interacts with a specific assessment to provide or curtail evidence of the targeted cognition.
Instructors are already redesigning their practice in response. The emerging strategy is to embed evidence of learning within the solution process itself rather than in a final artifact. Techniques include annotated drafts, in-class checkpoints, and reflective accounts of decision-making, all of which make student thinking more visible and reveal how learners frame problems, respond to feedback, revise ideas, and exercise judgment. These process-based approaches have revived formats long considered impractical, most notably the oral examination. Across STEM fields, instructors have drawn on learning sciences research to improve the reliability and fairness of oral exams: sharing rubrics in advance, recording sessions for calibration, offering rehearsal opportunities, and providing exemplar recordings. Contrary to intuition, several studies report that oral exams can require less total time than written tests, and efficiency improves when students are assessed in groups or sessions are distributed across faculty.
Authentic assessment, tasks that mirror the real work of STEM professionals, offers a second redesign pathway. Instead of single-correct-answer questions that a chatbot reproduces effortlessly, authentic tasks demand decision-making, construction of evidence-based explanations, troubleshooting, open-ended design, and interpretation of noisy data while wrestling with ambiguity and tradeoffs. Some versions position students’ problem-solving in interaction with others, eliciting the discipline-specific language and negotiation norms of professional practice. Yet the authors caution that authenticity can still be gamed unless the design also captures process, performance, and interaction data. The deeper fix, they argue, is cultural: students must come to see their work as building usable knowledge aligned with their own career goals, rather than viewing assessment as something to be policed.
A further complication is that AI literacy is emerging as a professional capability in its own right, requiring assessments to distinguish three separable constructs: foundational disciplinary competence, knowledge about AI, and sound judgment when using AI-supported tools. A student may understand AI concepts without using AI responsibly, or operate an AI tool effectively without grasping its limitations. Since AI is now routinely available in professional STEM practice, banning it from school assessments may produce evidence that is secure but misaligned with the capabilities students will need in their careers. Unrestricted use, however, could short-circuit the development of foundational knowledge. The proposed solution is a coherent assessment system mixing multiple forms: some tasks establishing independent capability, others evaluating the ability to use, critique, and verify AI-supported work, including portfolios that combine AI-free tasks, tasks permitting specified AI support, and tasks requiring students to improve AI output.
Policy is scrambling to keep pace, and the review finds that institutional responses have developed quickly and unevenly, producing interpretive chaos in which identical behavior is legitimate in one course and misconduct in another. Students commonly use chatbots for brainstorming, editing, translation, and partial drafting without viewing it as an integrity violation, because acceptable use depends on unstated assumptions about originality, effort, and ownership. Disclosure requirements exist on paper, but fear of academic consequences, ambiguous instructions, and inconsistent enforcement discourage honest declaration. Detection technology offers no rescue: AI detectors are unreliable and appear to disproportionately disadvantage certain student populations. The authors advocate task-level definitions of acceptable use, supported by concrete examples, coupled to a positive classroom culture in which students understand why their work matters and see instructors as partners in their success rather than adversaries armed with surveillance software.
The second emerging theme flips the perspective: AI is increasingly part of the assessment infrastructure itself, operating on the interpretation vertex of the triangle. Automated scoring has progressed from keyword-matching and feature-based systems to transformer-based large language models capable of locating student responses within theoretically grounded models of developing understanding. Studies now demonstrate that AI can assess disciplinary knowledge expressed in students’ own words, opening the possibility of rich constructed-response tasks at scales impossible for human graders, particularly in large introductory university courses. Yet the authors stress that quality depends on alignment with learning goals and that reliability cannot be assumed, especially for the higher-order thinking STEM professions demand. Concerns about bias, transparency, consistency, and transfer across tasks persist, and there is a visible shift from summative automated scoring toward formative feedback. A disturbing counter-current also appears in the literature: some students are bypassing instructor guidance entirely, depending on GenAI to learn STEM topics, a pattern likely to intensify as AI companies market aggressively to students.
The third theme concerns the fidelity of AI-generated assessment materials themselves. Generative AI is now used to create classroom dialogues, teaching cases, laboratory scenarios, and simulated student work, and researchers have begun evaluating these artifacts across dimensions including content, linguistic, cognitive, behavioral, structural, and pedagogical fidelity, with psychological fidelity, the credibility of simulated emotion, identified as an additional frontier. Current evidence reveals substantial limitations: simplified interaction structures, repetitive behaviors, unrealistic student errors, and insufficient responsiveness to individual learners. AI-generated tasks may also contain inaccurate content and impose cognitive demands inappropriate for the intended learners. The editorial concludes that such artifacts must pass through human-in-the-loop evaluation before use, and that the field’s ultimate question is deceptively simple: what evidence is needed to justify claims about student learning, and what role should AI play in producing it? With cognition, observation, and interpretation all shifting at once, the authors argue that wholesale reconsideration, not incremental patching, is the only viable response.
Subject of Research: The impact of generative AI on assessment design, validity, and policy in STEM education
Article Title: The shifting landscape of assessment in STEM education in the age of generative AI
Article References: Koretsky, M. D., Ding, M., Chiu, T. K. F., Hallström, J., & Li, Y. (2026). The shifting landscape of assessment in STEM education in the age of generative AI. International Journal of STEM Education, 13(1), Article 59. https://doi.org/10.1186/s40594-026-00648-5
Image Credits: AI Generated
DOI: 10.1186/s40594-026-00648-5
Keywords: generative AI, STEM education, assessment, validity, academic integrity, ChatGPT, authentic assessment, oral exams, AI literacy, automated scoring, assessment redesign, educational policy
Cite Scienmag News
Courtney Benton. (September 30, 2026). When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis. Scienmag. https://scienmag.com/when-chatgpt-aces-the-exam-stem-assessment-faces-a-validity-crisis/
Courtney Benton. "When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis." Scienmag, 30 September 2026, https://scienmag.com/when-chatgpt-aces-the-exam-stem-assessment-faces-a-validity-crisis/. Accessed 30 September 2026.
Courtney Benton. "When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis." Scienmag. September 30, 2026. https://scienmag.com/when-chatgpt-aces-the-exam-stem-assessment-faces-a-validity-crisis/

