Friday, October 2, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Medicine

AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds

October 2, 2026
in Medicine
Ophelia Keating
By Ophelia Keating Scienmag Editorial Profile - Health Services Research
Reading Time: 5 mins read
0
AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds

AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds

AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Artificial intelligence is quietly creeping into one of science’s most guarded rituals: peer review. Editors short on time and reviewers exhausted by endless requests have already begun experimenting with large language models to help evaluate manuscripts, and surveys suggest the practice is more widespread than journals officially acknowledge. But handing scientific judgment to machines trained on human text raises an uncomfortable question: will these systems inherit the biases that have long plagued human peer review, from gender disparities to geographic favoritism? A new experimental study published in the Journal of General Internal Medicine offers one of the most controlled answers yet, and the result is surprising in its cleanliness.

Paul Sebo of the University of Geneva and Ting Wang of Emporia State University set out to test whether two leading AI models, OpenAI’s ChatGPT and Anthropic’s Claude, would score identical scientific abstracts differently depending on the name and country attached to them. They also wanted to know something equally important for any future editorial application: whether the models give the same score to the same work when asked twice. The stakes are considerable. If AI reviewers are both unbiased and reproducible, they could become powerful tools for triaging the flood of submissions that modern journals face. If they are neither, their adoption could quietly distort who gets published and whose research shapes medicine.

The experimental design was deliberately austere. The researchers randomly selected ten general internal medicine journals from the Journal Citation Reports, each with an impact factor of at least 1.5, and pulled five original research abstracts per journal from the Web of Science, for a total of fifty abstracts published between 2023 and 2026. Each abstract was then evaluated under four fictional author identities: an American woman named Rachel Smith, an American man named John Smith, a woman from Côte d’Ivoire named Fatoumata Diallo, and a man from Côte d’Ivoire named Amadou Diallo. Using fictional identities rather than real researchers allowed the team to manipulate exactly one variable, the identity cue, while keeping everything else constant and avoiding the ethical tangle of attaching published work to real people.

Every abstract-identity combination was scored twice by each model, with each evaluation conducted in a fresh chat session to prevent any memory of previous judgments from contaminating the results. The prompt was identical every time: act as a scientific reviewer, and rate the abstract on three dimensions using a 0 to 10 scale, namely overall scientific quality, novelty, and likelihood of acceptance at a high-quality conference or journal. The models were instructed to output only three numbers, no explanations. That discipline produced 800 evaluations in total, 400 per model, all collected between April 15 and April 30, 2026, using the standard user interfaces with default settings.

The headline finding is striking in its symmetry. For ChatGPT, median quality and novelty scores were identical across all four identities, with only minimal variation in acceptance scores that followed no consistent gender or geographic pattern. For Claude, all three scores were identical across identities. Multivariable ordinal logistic regression, adjusted for journal, found no overall association between author identity and any score, with one partial exception: a global test for Claude’s acceptance scores hinted at a possible association, but no individual comparison reached statistical significance. The authors urge caution even about the isolated differences that did appear, such as slightly lower novelty and acceptance scores that ChatGPT gave to abstracts attributed to the American man, because the overall tests were not significant and many comparisons were performed on a modest number of unique abstracts.

Reproducibility, the second pillar of the study, was equally encouraging. When the same abstract was scored twice under the same identity, percent agreement exceeded 0.98 for both models across all three outcomes. Fleiss’ kappa, a statistic that corrects for agreement expected by chance and weights larger disagreements more heavily, ranged from 0.88 to 0.89 for ChatGPT, indicating excellent consistency, and from 0.74 to 0.80 for Claude, indicating substantial agreement. ChatGPT was thus the steadier grader, though both models were far more consistent with themselves than human reviewers typically are with each other, a well-documented weakness of traditional peer review.

The scores were not blind to everything, however. Abstracts from higher-impact journals received systematically higher ratings across all three dimensions, with each point of impact factor associated with roughly 35 to 40 percent higher odds of a better score. Interestingly, the journals’ names and impact factors were never shown to the models, so this pattern cannot reflect prestige bias directly. It more likely reflects genuine differences in the writing quality and study characteristics of abstracts published in more selective venues, though the relationship was not perfectly linear, and one mid-tier journal, Annals of Family Medicine, consistently ranked low despite its respectable impact factor of 5.1. The models also differed from each other: Claude assigned significantly lower quality and acceptance scores than ChatGPT overall, suggesting that different AI systems bring meaningfully different evaluative standards to the same text.

The authors are careful, almost insistently so, about what these results do not prove. Consistency is not validity. The fact that a model gives the same score twice says nothing about whether that score is accurate, insightful, or useful for editorial decisions, and the study included no human reference standard against which to calibrate the numbers. The restricted range of scores, mostly clustered between 2 and 9 with medians of 5 to 7, may itself have inflated the agreement statistics. The task was also radically simplified compared with real peer review: abstracts rather than full manuscripts, numerical scores rather than narrative critique, and a single standardized prompt rather than the messy, iterative dialogue that characterizes genuine refereeing. Subtler biases that might surface in open-ended written feedback would be invisible in this design.

There are further limits worth keeping in mind. Fifty abstracts from a single medical specialty cannot speak for all of science, and only two models, at specific versions, were tested; both are updated frequently, so the findings may not hold for future releases. The identity manipulation covered only two countries and two genders, a narrow slice of the diversity of the global research community, and the researchers could not verify whether the models actually registered the identity cues at all. An absence of score differences could mean the models are genuinely impartial, or simply that they ignored the names entirely. The study was also not preregistered, which the authors acknowledge transparently.

Still, the implications are provocative. At a moment when journals are wrestling with reviewer shortages and with evidence that human peer review itself suffers from gender, geographic, and institutional bias, the prospect of an evaluator that treats a manuscript from Geneva and one from Abidjan identically is genuinely newsworthy. The authors suggest that structured, low-stakes applications such as initial screening, detection of reporting deficiencies, and editorial triage may be the most realistic near-term uses for these tools, rather than replacing human reviewers outright. The most informative next step, they argue, would be to run LLM-based review on manuscripts for which real human review reports and editorial decisions already exist, allowing a direct head-to-head comparison of judgments and biases. Until then, the message of this study is a measured one: in the sterile laboratory of standardized abstract scoring, the machines showed no favoritism and remarkable self-consistency. Whether they can sustain that discipline in the messy, argumentative world of real scientific publishing remains an open and urgent question.

Subject of Research: Bias and reproducibility of large language models as evaluators of scientific abstracts in peer review

Article Title: Bias and Reliability of AI-Based Peer Review: A Comparative Study of ChatGPT and Claude Evaluating Scientific Abstracts

Article References: Sebo, P., & Wang, T. (2026). Bias and Reliability of AI-Based Peer Review: A Comparative Study of ChatGPT and Claude Evaluating Scientific Abstracts. Journal of General Internal Medicine. https://doi.org/10.1007/s11606-026-10806-8

Image Credits: AI Generated

DOI: 10.1007/s11606-026-10806-8

Keywords: artificial intelligence, large language models, peer review, ChatGPT, Claude, bias, reproducibility, scientific publishing, gender bias, geographic bias, abstracts, internal medicine

Cite Scienmag News

Ophelia Keating. (October 2, 2026). AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds. Scienmag. https://scienmag.com/ai-reviewers-show-no-gender-or-geographic-bias-in-abstract-scoring-study-finds/

Ophelia Keating. "AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds." Scienmag, 2 October 2026, https://scienmag.com/ai-reviewers-show-no-gender-or-geographic-bias-in-abstract-scoring-study-finds/. Accessed 2 October 2026.

Ophelia Keating. "AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds." Scienmag. October 2, 2026. https://scienmag.com/ai-reviewers-show-no-gender-or-geographic-bias-in-abstract-scoring-study-finds/

Tags: abstractsAI fairness in scienceAI peer review biasAI-assisted manuscript reviewAnthropic Claude review performanceArtificial Intelligencebiasbias-free AI scientific judgmentChatGPTChatGPT scientific abstract evaluationClaudeethical considerations in AI peer reviewgender biasgender bias in scientific evaluationgeographic biasgeographic bias in AI scoringimpact of AI on scientific publishinginternal medicinelarge language modelslarge language models in peer reviewpeer reviewreproducibilityreproducibility of AI assessmentscientific publishing
Share26Tweet16
Previous Post

When Classrooms Learn to Read Feelings: The Hidden Ethics of Emotion-Sensing AI in Schools

Next Post

Light-Polluted Nights and Heat Push Day-Active Tiger Mosquitoes Into the Dark

Related Posts

Rising Air Pollution Quietly Paused Methane’s Climb, Study Finds
Medicine

Rising Air Pollution Quietly Paused Methane’s Climb, Study Finds

October 2, 2026
Chinese Herbal Formula Faces Its Toughest Test: Landmark Trial Targets Precancerous Colon Polyps
Medicine

Chinese Herbal Formula Faces Its Toughest Test: Landmark Trial Targets Precancerous Colon Polyps

October 2, 2026
Exercise May Lower Multiple Sclerosis Risk, But Sun Exposure Holds the Key
Medicine

Exercise May Lower Multiple Sclerosis Risk, But Sun Exposure Holds the Key

October 2, 2026
COVID-19 Booster Debate: Why Heterologous Advantage May Be Overstated
Medicine

COVID-19 Booster Debate: Why Heterologous Advantage May Be Overstated

October 2, 2026
AI-Powered Reconstruction Slashes Radiation Dose in Heart Scans Without Sacrificing Image Quality
Medicine

AI-Powered Reconstruction Slashes Radiation Dose in Heart Scans Without Sacrificing Image Quality

October 2, 2026
Inverted T-waves in Athletes Are Common but Usually Benign, Massive Analysis Finds
Medicine

Inverted T-waves in Athletes Are Common but Usually Benign, Massive Analysis Finds

October 2, 2026
Next Post
Light-Polluted Nights and Heat Push Day-Active Tiger Mosquitoes Into the Dark

Light-Polluted Nights and Heat Push Day-Active Tiger Mosquitoes Into the Dark

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Turmeric’s Weakness Fixed: Sugar Nanoparticles Boost Curcumin Solubility 1,100-Fold
  • Risk-averse manufacturers may do better outsourcing remanufacturing than keeping it in-house
  • Light-Polluted Nights and Heat Push Day-Active Tiger Mosquitoes Into the Dark
  • AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading