Monday, October 5, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

ChatGPT-5 Outperforms Classic Alvarado Score in Detecting Appendicitis, Study Finds

October 5, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
ChatGPT-5 Outperforms Classic Alvarado Score in Detecting Appendicitis, Study Finds

ChatGPT-5 Outperforms Classic Alvarado Score in Detecting Appendicitis, Study Finds

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Acute appendicitis is the most common cause of the acute abdomen, and every emergency physician knows the stakes: miss it, and the inflamed appendix can rupture, flooding the abdomen with bacteria; overcall it, and a patient undergoes surgery they never needed. For decades, clinicians have leaned on the Alvarado score, a simple checklist of symptoms, signs and blood values, to tip the balance. Now a retrospective diagnostic accuracy study from Mansoura University Hospital in Egypt suggests that a large language model, ChatGPT-5, can read the same clinical data and outperform that venerable score in a striking way, particularly in its ability to rule the disease out.

The research team, led by Amr A. Elgharib and colleagues, assembled a cohort of 162 patients aged 15 and over who presented with abdominal pain and were clinically suspected of having appendicitis between March 2024 and March 2025. All of them underwent appendectomy, and the final word on their diagnosis came from histopathological examination of the removed appendix, the definitive reference standard. The cohort was deliberately enriched: roughly 70 percent of patients turned out to have confirmed appendicitis, a prevalence far higher than in a general emergency department, which shapes how the results must be interpreted.

The researchers fed ChatGPT-5 structured clinical narratives for each patient, including age, sex, symptoms such as fever, nausea, vomiting and right lower quadrant pain, physical examination findings like rebound tenderness, and laboratory values including white blood cell count and neutrophil percentage. The prompt cast the model as a general surgery physician in an emergency department and, crucially, forbade it from using any established appendicitis scoring system such as Alvarado, AIR, RIPASA or AAS. Instead, the model had to rely on its own clinical reasoning, outputting either a diagnosis of acute or non-acute appendicitis and, for positive cases, a percentage probability stratified into low, intermediate or high risk.

The study was designed in two phases following the classic machine learning convention of a 70-30 split. In Phase 1, 114 cases were presented without any prior exposure to labeled examples, simulating how a clinician might use a readily available chatbot off the shelf. After the model’s predictions were recorded, the true diagnoses were revealed, and in Phase 2 the remaining 48 cases were used to test whether this in-context learning improved performance. The prompt format was fixed across all cases, and the model’s memory was cleared before each evaluation round to test its default behavior.

The headline numbers are eye-catching. In Phase 1 without randomization, ChatGPT-5 achieved 100 percent sensitivity, meaning it caught every single case of appendicitis with zero false negatives, alongside 81.6 percent specificity, a positive predictive value of 91.6 percent, a negative predictive value of 100 percent, and overall accuracy of 93.9 percent. The Alvarado score, by contrast, managed only 67.1 percent sensitivity, missing 25 true cases, though its specificity was higher at 92.1 percent, and its overall accuracy came to 75.4 percent. Statistical comparison using McNemar’s test for paired proportions confirmed that the model’s sensitivity and accuracy advantages were highly significant, with p-values below 0.001.

But the story took a twist when the researchers examined a subtle methodological trap known as order leakage. When cases are presented in a sequence, a language model can inadvertently exploit the ordering of the data rather than genuine clinical reasoning, artificially inflating its apparent accuracy. To probe this, the team re-ran the evaluations with the case order randomized. Sensitivity barely budged, remaining at 98.7 percent in Phase 1, but specificity collapsed dramatically from 81.6 percent to 36.8 percent, a difference that was highly statistically significant. In other words, part of the model’s apparent precision in ruling appendicitis in may have been an artifact of how the cases were sequenced in the prompt.

Phase 2 told a more balanced story. After in-context learning, the non-randomized model reached 100 percent sensitivity, 93.8 percent specificity, 97.0 percent positive predictive value, 100 percent negative predictive value and 97.9 percent accuracy, statistically indistinguishable from the Alvarado score’s 96.9 percent sensitivity and 93.8 percent specificity. Agreement between the two methods, measured with Cohen’s kappa, climbed from a moderate 0.503 in Phase 1 to an almost perfect 0.952 in Phase 2. The authors note that in-context learning produced a modest, non-significant gain in specificity, while sensitivity was largely unaffected, suggesting the model’s core ability to exclude appendicitis was robust from the start.

The findings sit within a growing but cautionary literature. Previous evaluations have found that ChatGPT’s answers to surgeon-designed appendicitis management questions were clinically pertinent but inconsistent, and that ChatGPT-3.5 performed significantly worse than clinicians for some acute abdominal conditions such as cholecystitis and diverticulitis. A systematic review of 29 studies on artificial intelligence in appendicitis highlighted wide heterogeneity in inputs, validation strategies and metrics. Imaging-based machine learning models reading CT scans have achieved sensitivities around 77 percent and specificities around 86 percent. The new study’s authors also point to broader evidence that data leakage has inflated reported performance in fields ranging from Parkinson’s disease detection to brain MRI classification, reinforcing their insistence on randomization as a guard against over-optimistic estimates.

The authors are careful about what their results do and do not license. Because the cohort consisted exclusively of surgically confirmed patients with high pre-test probability, the estimates cannot simply be generalized to the unselected emergency department population in whom appendicitis must be ruled in or out, and the study was limited to a single center, a single model and a single comparator, with newer scores like AIR and RIPASA omitted because the required data were not consistently available. ChatGPT-5, they conclude, cannot yet be relied upon to confirm acute appendicitis, given the Alvarado score’s higher specificity, and the findings should be regarded as hypothesis-generating pending prospective validation.

What the study does establish is a proof of concept with real clinical texture: a general-purpose language model, guided by a carefully engineered prompt and given nothing more than the routine clinical data already sitting in a patient’s chart, can match or exceed a century-old scoring heuristic at the task that matters most for safety, excluding the disease. The authors envision ChatGPT-5 as a clinical decision-support tool to assist professionals during early assessment, particularly in resource-limited settings, always under physician supervision and never as an autonomous diagnostician. They also warn explicitly against patient self-diagnosis, noting that false reassurance from a chatbot could delay timely medical evaluation. As large language models continue their march into medicine, this Egyptian study offers both an encouraging data point and a methodological warning: before trusting an AI’s diagnostic brilliance, check whether the cases were shuffled.

Subject of Research: Diagnostic accuracy of ChatGPT-5 compared with the Alvarado score for acute appendicitis

Article Title: Diagnostic performance of ChatGPT for acute appendicitis compared with the Alvarado score: a retrospective diagnostic accuracy study

Article References: Elgharib, A. A., Ghanem, M. I., Ibrahim, R. S., Elwakeel, N., & Shemes, A. (2026). Diagnostic performance of ChatGPT for acute appendicitis compared with the Alvarado score: a retrospective diagnostic accuracy study. Discover Artificial Intelligence, 6(1), Article 1352. https://doi.org/10.1007/s44163-026-02332-7

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02332-7

Keywords: ChatGPT-5, large language models, acute appendicitis, Alvarado score, diagnostic accuracy, artificial intelligence, clinical decision support, histopathology, order leakage, in-context learning, emergency medicine, prompt engineering

Cite Scienmag News

Denise Maddox. (October 5, 2026). ChatGPT-5 Outperforms Classic Alvarado Score in Detecting Appendicitis, Study Finds. Scienmag. https://scienmag.com/chatgpt-5-outperforms-classic-alvarado-score-in-detecting-appendicitis-study-finds/

Denise Maddox. "ChatGPT-5 Outperforms Classic Alvarado Score in Detecting Appendicitis, Study Finds." Scienmag, 5 October 2026, https://scienmag.com/chatgpt-5-outperforms-classic-alvarado-score-in-detecting-appendicitis-study-finds/. Accessed 5 October 2026.

Denise Maddox. "ChatGPT-5 Outperforms Classic Alvarado Score in Detecting Appendicitis, Study Finds." Scienmag. October 5, 2026. https://scienmag.com/chatgpt-5-outperforms-classic-alvarado-score-in-detecting-appendicitis-study-finds/

Tags: acute appendicitisAI diagnostic tools for appendicitisAI in acute abdomen assessmentAI outperforming traditional scoring systemsAI-based symptom and sign analysis in emergency careAlvarado scoreappendicitis diagnosis accuracyArtificial IntelligenceChatGPT-5ChatGPT-5 clinical decision supportChatGPT-5 versus clinical scoring systemsclinical decision supportcomparison of AI and Alvarado score in appendicitis detectiondiagnostic accuracyEmergency Medicinehistopathological confirmation of appendicitishistopathologyin-context learninglarge language modelsorder leakageprompt engineeringretrospective diagnostic accuracy studyrole of AI in ruling out appendicitisuse of large language models in emergency medicine
Share26Tweet16
Previous Post

Weather and Pollution Show Little Sway Over Respiratory Pathogens in Central China

Next Post

Moroccan Oases and Atlas Mountains Hide a Surprising Wealth of Fig Tree Genetic Diversity

Related Posts

Hidden atomic distortions explain why promising lithium battery cathodes waste energy
Technology and Engineering

Hidden atomic distortions explain why promising lithium battery cathodes waste energy

October 5, 2026
Teaching Tiny Networks: New Quantization Method Pushes 1-Bit AI Toward Full-Precision Accuracy
Technology and Engineering

Teaching Tiny Networks: New Quantization Method Pushes 1-Bit AI Toward Full-Precision Accuracy

October 5, 2026
Nanopore Sequencing Spots Deadly Fungal Bloodstream Infections in Hours, Not Days
Technology and Engineering

Nanopore Sequencing Spots Deadly Fungal Bloodstream Infections in Hours, Not Days

October 5, 2026
Quantum Rivals, Delayed Data: Economists Find a Universal Stability Boundary in Quantum Duopolies
Technology and Engineering

Quantum Rivals, Delayed Data: Economists Find a Universal Stability Boundary in Quantum Duopolies

October 5, 2026
Tiny AI brain lets a $10 microcontroller remember hidden objects and grab them
Technology and Engineering

Tiny AI brain lets a $10 microcontroller remember hidden objects and grab them

October 5, 2026
Thermography-Guided Redesign of Film Heater Traces Cuts Temperature Swings by a Third
Technology and Engineering

Thermography-Guided Redesign of Film Heater Traces Cuts Temperature Swings by a Third

October 5, 2026
Next Post
Moroccan Oases and Atlas Mountains Hide a Surprising Wealth of Fig Tree Genetic Diversity

Moroccan Oases and Atlas Mountains Hide a Surprising Wealth of Fig Tree Genetic Diversity

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Who Trusts AI in Europe? Class and Government Views Shape Public Opinion
  • Satellite Record Reveals 25 Years of Dust and Haze Over Nigeria
  • Moroccan Oases and Atlas Mountains Hide a Surprising Wealth of Fig Tree Genetic Diversity
  • ChatGPT-5 Outperforms Classic Alvarado Score in Detecting Appendicitis, Study Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading