Generative artificial intelligence has been pitched as everything from a tireless tutor to a replacement grader, but a new classroom experiment suggests the technology is best understood as something far more modest: a competent peer. In a study published in Ecology and Evolution, an instructor at St. Mary’s University in San Antonio, Texas, put ChatGPT-5 head-to-head with herself and with student reviewers across three undergraduate biology courses, scoring the same drafts with the same rubrics. The results offer one of the most granular looks yet at how well large language models can approximate expert judgment in real science writing assignments, and at how students actually feel about feedback when they know a machine produced part of it.
The study, conducted during the Fall 2025 semester, spanned three courses of increasing disciplinary complexity: a non-majors food and nutrition course with 45 students, an introduction to bioinformatics course with 17 majors, and an upper-level genes, genomes, and genomics course with 27 majors. Two writing-to-learn assignment genres were tested. Students in the first two courses designed infographics meant to communicate a scientific concept to a general audience, while students in the genomics course wrote evidence-based explanatory articles on emerging ethical issues in genomic science. Every assignment followed a scaffolded, draft-based process, with rubrics distributed from the outset so that all evaluators, human and machine alike, were working from identical criteria.
The workflow was carefully engineered to keep the three feedback streams independent. Each draft simultaneously received formative feedback from the instructor, from two classmates whose scores were averaged into a single composite, and from ChatGPT-5 operating through the OpenAI web interface with memory features disabled, ensuring the model retained nothing between evaluations. The AI received the assignment instructions, the analytic rubric, and a standardized prompt directing it to complete the rubric and provide criterion-specific written comments, with particular attention to sources and citations. No rater saw any other rater’s scores. Only after final grades were posted were feedback sources revealed to students, immediately before they completed an anonymous perception survey.
To measure how closely the three sources aligned, the study used intraclass correlation coefficients, a statistical tool that quantifies agreement among raters on the same set of items. Values below 0.50 were interpreted as poor, 0.50 to 0.75 as moderate, 0.75 to 0.90 as good, and above 0.90 as excellent. The headline finding was that overall agreement among AI, instructor, and peer evaluations was good for the non-majors infographic assignment, with an ICC of 0.79, and good for the upper-level explanatory article, at 0.80, but only moderate for the majors’ bioinformatics infographic, at 0.54. In other words, the machine and the humans largely agreed, but the strength of that consensus depended heavily on what kind of assignment was on the table.
That modality effect became even clearer when the rubric categories were examined individually. For the text-based academic articles, agreement was strongest on evidence and reasoning and broader impact, both hovering around 0.75, while sources and citations lagged at 0.36. For the bioinformatics infographics, agreement collapsed entirely on organization and storytelling and on design and visuals, both hitting an ICC of 0.00, even as sources and citations reached a solid 0.79. The pattern suggests that large language models align most reliably with instructor judgment when assessing written scientific argumentation, and least reliably when the task demands interpretation of visual design, representational choices, and the integration of graphics with scientific content, a domain where multimodal evaluation imposes interpretive demands that current models handle unevenly.
The study also tracked how AI-instructor agreement shifted as students revised their work, and here the two assignment types diverged in opposite directions. For the infographics, agreement between AI and instructor scores declined modestly from draft to final submission, dropping from 0.66 to 0.62 in the non-majors course and from 0.54 to 0.29 in the bioinformatics course. For the academic articles, the opposite occurred: overall agreement rose from 0.72 to 0.80, with notable gains in scientific accuracy, evidence and reasoning, argumentation and organization, and broader impact. One plausible reading is that as students revised text-based arguments using the combined feedback, their work converged on standards that both the instructor and the model recognized, while visual revisions moved in directions the AI judged differently from the human expert.
Student perceptions added a second, equally important layer of evidence. Across both infographic courses, students rated instructor feedback significantly higher than peer and AI feedback on nearly every measure: identifying weaknesses, providing specific and actionable suggestions, improving organization and scientific accuracy, and, most strikingly, trust, which showed the largest effect size in the study with a Kendall’s W of 0.44. Peer and AI feedback were rated similarly on most dimensions, effectively positioning the chatbot as a substitute classmate rather than a surrogate professor. For the upper-level article assignment, differences among the three sources shrank, with significant gaps remaining only for identifying weaknesses, specificity, scientific accuracy, and trust. Notably, students reported comparable effort applying feedback regardless of its source, meaning AI suggestions did not add cognitive burden to the revision process.
The survey also probed why students sometimes ignored the machine. Among the 49 respondents who reported rejecting at least one AI suggestion, the most common reason, cited by 55 percent, was that the AI feedback conflicted with feedback from another source, followed by concerns that it was inaccurate or unclear, each cited by 37 percent, and that it did not match the rubric, cited by 29 percent. Yet the overall verdict on the AI was cautiously positive: 39 percent of students said AI feedback improved the quality of their final product and another 41 percent said it may have. More broadly, 80 percent reported that their ability to communicate science improved through the project, and 84 percent said their understanding of biological content deepened, suggesting that a multi-source feedback workflow including AI did not dilute perceived learning gains.
The author is careful to frame these findings as evidence about one implementation rather than a verdict on any particular model, and the limitations are real. The study took place at a single Hispanic-serving institution with modest enrollments, sample sizes for the reliability analyses ranged from 15 to 38 assignments, and confidence intervals around some ICC estimates were wide. Assignment type, course context, and student experience varied together and could not be disentangled. The perception survey was administered only after feedback sources were disclosed, so students’ ratings of usefulness and trust likely reflect the authority of the instructor’s name as much as the intrinsic quality of the comments. Peer reviewers, moreover, received no formal training, which may have depressed their scores relative to what trained peer review could achieve.
Even with those caveats, the practical message for educators is reasonably clear. Structured implementation matters: the AI in this study operated within assignment-specific prompts, analytic rubrics, and instructor oversight, not as an unrestricted generative tool, and that discipline appears to be what made its feedback usable. The findings support treating large language models as stage-specific scaffolds that help students identify baseline weaknesses, clarify expectations, and engage in early revision cycles, thereby freeing instructor time for the conceptual depth, disciplinary nuance, and rhetorical judgment that remain stubbornly human territory. Students themselves seem to have internalized this hierarchy, trusting the machine about as much as a classmate and the professor far more. As AI systems evolve, the open questions are shifting from whether models can grade to when and how their feedback should enter the workflow, and future work comparing multiple models, blinding feedback sources, and spanning more institutions will be needed to map that territory. For now, the evidence suggests the smartest role for ChatGPT in the science classroom is neither ghost-grader nor gimmick, but a third voice in the room, one worth hearing and wise to verify.
Subject of Research: Comparison of AI, peer, and instructor feedback on undergraduate biology writing assignments
Article Title: Evaluating Artificial Intelligence, Peer, and Instructor Feedback Across Scientific Communication Assignments in Undergraduate Biology
Article References: Boies, L. (2026). Evaluating Artificial Intelligence, Peer, and Instructor Feedback Across Scientific Communication Assignments in Undergraduate Biology. Ecology and Evolution, 16(9), Article e74365. https://doi.org/10.1002/ece3.74365
Image Credits: AI Generated
DOI: 10.1002/ece3.74365
Keywords: artificial intelligence, ChatGPT, formative feedback, science communication, undergraduate biology, peer review, writing-to-learn, inter-rater reliability, STEM education, large language models, assessment rubrics, student perceptions
Cite Scienmag News
Drew Townsend. (September 24, 2026). ChatGPT Graded Biology Homework Like a Peer, But Professors Still Rule the Classroom. Scienmag. https://scienmag.com/chatgpt-graded-biology-homework-like-a-peer-but-professors-still-rule-the-classroom/
Drew Townsend. "ChatGPT Graded Biology Homework Like a Peer, But Professors Still Rule the Classroom." Scienmag, 24 September 2026, https://scienmag.com/chatgpt-graded-biology-homework-like-a-peer-but-professors-still-rule-the-classroom/. Accessed 24 September 2026.
Drew Townsend. "ChatGPT Graded Biology Homework Like a Peer, But Professors Still Rule the Classroom." Scienmag. September 24, 2026. https://scienmag.com/chatgpt-graded-biology-homework-like-a-peer-but-professors-still-rule-the-classroom/

