Wednesday, September 30, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Science Education

When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis

September 30, 2026
in Science Education
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 5 mins read
0
When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis

When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis

When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

A single midterm exam at Brown University has become the flashpoint in a battle that is quietly dismantling a century of thinking about how schools measure learning. In spring 2026, a professor suspected that his take-home midterm had been answered largely by ChatGPT: the 77 students in the class averaged an astonishing 96 percent, and many answers reproduced the chatbot’s idiosyncratic reasoning almost verbatim. When the professor replaced the final exam with an in-person version, eighteen students dropped the course, the average collapsed to below 50 percent, and nineteen students failed. The episode, recounted in an editorial by Milo Koretsky of Tufts University and colleagues in the International Journal of STEM Education, is not merely a story about cheating. It is evidence that the basic inference at the heart of every exam, that the submitted work reveals what a student actually knows, has broken down across large parts of higher education.

The editorial, written with Meixia Ding of Temple University, Thomas Chiu of the Chinese University of Hong Kong, Jonas Hallström of Linköping University, and Yeping Li of Texas A&M University, argues that the response to generative artificial intelligence must be framed as a problem of validity rather than a problem of policing. Assessment researchers have long described measurement in education through the assessment triangle, a framework from the National Research Council with three vertices: cognition, the learning goals defining what students should know and be able to do; observation, the tasks that elicit evidence of that learning; and interpretation, the inferences drawn from the collected evidence. Generative AI destabilizes every vertex simultaneously, and the authors used the triangle as a lens for a scoping review spanning the Scopus, Web of Science, and ERIC databases.

The technical heart of the problem sits on the observation vertex. Generative AI disrupts the assumption that a polished product is credible evidence of student capability, because a chatbot such as ChatGPT or Claude can often produce correct solutions to narrowly defined problems. Empirical surveys confirm that educators and students view the validity threat as most acute for essays, take-home exams, and computer programming assignments, formats where a coherent response can be generated with little visible trace of independent thinking. Interestingly, the threat runs in both directions. Restrictive anti-cheating controls, such as converting every flexible assignment into a high-stakes proctored exam, can reduce accessibility, authenticity, and alignment with professional practice, thereby weakening validity from the opposite side. What matters is not AI use in the abstract but how a specific use interacts with a specific assessment to provide or curtail evidence of the targeted cognition.

Instructors are already redesigning their practice in response. The emerging strategy is to embed evidence of learning within the solution process itself rather than in a final artifact. Techniques include annotated drafts, in-class checkpoints, and reflective accounts of decision-making, all of which make student thinking more visible and reveal how learners frame problems, respond to feedback, revise ideas, and exercise judgment. These process-based approaches have revived formats long considered impractical, most notably the oral examination. Across STEM fields, instructors have drawn on learning sciences research to improve the reliability and fairness of oral exams: sharing rubrics in advance, recording sessions for calibration, offering rehearsal opportunities, and providing exemplar recordings. Contrary to intuition, several studies report that oral exams can require less total time than written tests, and efficiency improves when students are assessed in groups or sessions are distributed across faculty.

Authentic assessment, tasks that mirror the real work of STEM professionals, offers a second redesign pathway. Instead of single-correct-answer questions that a chatbot reproduces effortlessly, authentic tasks demand decision-making, construction of evidence-based explanations, troubleshooting, open-ended design, and interpretation of noisy data while wrestling with ambiguity and tradeoffs. Some versions position students’ problem-solving in interaction with others, eliciting the discipline-specific language and negotiation norms of professional practice. Yet the authors caution that authenticity can still be gamed unless the design also captures process, performance, and interaction data. The deeper fix, they argue, is cultural: students must come to see their work as building usable knowledge aligned with their own career goals, rather than viewing assessment as something to be policed.

A further complication is that AI literacy is emerging as a professional capability in its own right, requiring assessments to distinguish three separable constructs: foundational disciplinary competence, knowledge about AI, and sound judgment when using AI-supported tools. A student may understand AI concepts without using AI responsibly, or operate an AI tool effectively without grasping its limitations. Since AI is now routinely available in professional STEM practice, banning it from school assessments may produce evidence that is secure but misaligned with the capabilities students will need in their careers. Unrestricted use, however, could short-circuit the development of foundational knowledge. The proposed solution is a coherent assessment system mixing multiple forms: some tasks establishing independent capability, others evaluating the ability to use, critique, and verify AI-supported work, including portfolios that combine AI-free tasks, tasks permitting specified AI support, and tasks requiring students to improve AI output.

Policy is scrambling to keep pace, and the review finds that institutional responses have developed quickly and unevenly, producing interpretive chaos in which identical behavior is legitimate in one course and misconduct in another. Students commonly use chatbots for brainstorming, editing, translation, and partial drafting without viewing it as an integrity violation, because acceptable use depends on unstated assumptions about originality, effort, and ownership. Disclosure requirements exist on paper, but fear of academic consequences, ambiguous instructions, and inconsistent enforcement discourage honest declaration. Detection technology offers no rescue: AI detectors are unreliable and appear to disproportionately disadvantage certain student populations. The authors advocate task-level definitions of acceptable use, supported by concrete examples, coupled to a positive classroom culture in which students understand why their work matters and see instructors as partners in their success rather than adversaries armed with surveillance software.

The second emerging theme flips the perspective: AI is increasingly part of the assessment infrastructure itself, operating on the interpretation vertex of the triangle. Automated scoring has progressed from keyword-matching and feature-based systems to transformer-based large language models capable of locating student responses within theoretically grounded models of developing understanding. Studies now demonstrate that AI can assess disciplinary knowledge expressed in students’ own words, opening the possibility of rich constructed-response tasks at scales impossible for human graders, particularly in large introductory university courses. Yet the authors stress that quality depends on alignment with learning goals and that reliability cannot be assumed, especially for the higher-order thinking STEM professions demand. Concerns about bias, transparency, consistency, and transfer across tasks persist, and there is a visible shift from summative automated scoring toward formative feedback. A disturbing counter-current also appears in the literature: some students are bypassing instructor guidance entirely, depending on GenAI to learn STEM topics, a pattern likely to intensify as AI companies market aggressively to students.

The third theme concerns the fidelity of AI-generated assessment materials themselves. Generative AI is now used to create classroom dialogues, teaching cases, laboratory scenarios, and simulated student work, and researchers have begun evaluating these artifacts across dimensions including content, linguistic, cognitive, behavioral, structural, and pedagogical fidelity, with psychological fidelity, the credibility of simulated emotion, identified as an additional frontier. Current evidence reveals substantial limitations: simplified interaction structures, repetitive behaviors, unrealistic student errors, and insufficient responsiveness to individual learners. AI-generated tasks may also contain inaccurate content and impose cognitive demands inappropriate for the intended learners. The editorial concludes that such artifacts must pass through human-in-the-loop evaluation before use, and that the field’s ultimate question is deceptively simple: what evidence is needed to justify claims about student learning, and what role should AI play in producing it? With cognition, observation, and interpretation all shifting at once, the authors argue that wholesale reconsideration, not incremental patching, is the only viable response.

Subject of Research: The impact of generative AI on assessment design, validity, and policy in STEM education

Article Title: The shifting landscape of assessment in STEM education in the age of generative AI

Article References: Koretsky, M. D., Ding, M., Chiu, T. K. F., Hallström, J., & Li, Y. (2026). The shifting landscape of assessment in STEM education in the age of generative AI. International Journal of STEM Education, 13(1), Article 59. https://doi.org/10.1186/s40594-026-00648-5

Image Credits: AI Generated

DOI: 10.1186/s40594-026-00648-5

Keywords: generative AI, STEM education, assessment, validity, academic integrity, ChatGPT, authentic assessment, oral exams, AI literacy, automated scoring, assessment redesign, educational policy

Cite Scienmag News

Courtney Benton. (September 30, 2026). When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis. Scienmag. https://scienmag.com/when-chatgpt-aces-the-exam-stem-assessment-faces-a-validity-crisis/

Courtney Benton. "When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis." Scienmag, 30 September 2026, https://scienmag.com/when-chatgpt-aces-the-exam-stem-assessment-faces-a-validity-crisis/. Accessed 30 September 2026.

Courtney Benton. "When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis." Scienmag. September 30, 2026. https://scienmag.com/when-chatgpt-aces-the-exam-stem-assessment-faces-a-validity-crisis/

Tags: academic integrityAI literacyAI-generated exam responsesassessmentassessment redesignauthentic assessmentautomated scoringchallenges of detecting AI-assisted cheatingChatGPTChatGPT and academic integrityeducational policyeducational policy response to AI-generated workevolving assessment strategies for AI integrationfuture of fair and reliable student assessmentsgenerative AIimpact of AI on higher education evaluationimplications of AI for measuring student understandinglimitations of traditional testing methodsoral examsreassessing exam effectiveness in the AI eraSTEM educationvalidityvalidity concerns in online and take-home examsvalidity crisis in STEM assessments
Share26Tweet16
Previous Post

Tensor-Network Solver Gets a Major Overhaul for Hard Optimization Problems

Next Post

Chip-Sized OCT Scanner Brings 3D Industrial Inspection Into Tight Spaces

Related Posts

New Model Argues VR Learning Depends on Experience, Not Just Headsets
Science Education

New Model Argues VR Learning Depends on Experience, Not Just Headsets

September 30, 2026
Ghana’s Community Health Training Grounds Could Become a Launchpad for Team-Based Care
Science Education

Ghana’s Community Health Training Grounds Could Become a Launchpad for Team-Based Care

September 30, 2026
Wealth Divide Leaves Nigeria’s Poorest Children Most Vulnerable to Missed Vaccines
Science Education

Wealth Divide Leaves Nigeria’s Poorest Children Most Vulnerable to Missed Vaccines

September 30, 2026
Menopause Training in Family Medicine Residencies Remains Uneven, National Survey Finds
Science Education

Menopause Training in Family Medicine Residencies Remains Uneven, National Survey Finds

September 30, 2026
AI Arrived Before the Evidence: Scientists Map How Universities Can Change Under Uncertainty
Science Education

AI Arrived Before the Evidence: Scientists Map How Universities Can Change Under Uncertainty

September 30, 2026
Repeating a Grade Leaves Spanish Teens Feeling Less Confident, Connected and Curious
Science Education

Repeating a Grade Leaves Spanish Teens Feeling Less Confident, Connected and Curious

September 30, 2026
Next Post
Chip-Sized OCT Scanner Brings 3D Industrial Inspection Into Tight Spaces

Chip-Sized OCT Scanner Brings 3D Industrial Inspection Into Tight Spaces

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Long Waits and High Deaths: Kenya’s Congenital Heart Disease Care Gap Laid Bare
  • Radiation-Proof Batteries Could Power the Hunt for Ghost Particles
  • Chip-Sized OCT Scanner Brings 3D Industrial Inspection Into Tight Spaces
  • When ChatGPT Aces the Exam: STEM Assessment Faces a Validity Crisis

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading