Saturday, September 26, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Medicine

AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds

September 26, 2026
in Medicine
Ophelia Keating
By Ophelia Keating Scienmag Editorial Profile - Health Services Research
Reading Time: 5 mins read
0
AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds

AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds

AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

A team of ophthalmologists in China has put some of the world’s most powerful artificial intelligence chatbots through one of the most demanding real-world tests yet devised for medical AI: deciding which type of refractive surgery, if any, suits a patient’s eyes. The results, published in BMC Medicine, show that five leading large language models could match or exceed the performance of an intermediate-level ophthalmologist when triaging patients for procedures such as LASIK and SMILE, reaching accuracies above 98.5 percent in some tasks. But the study also draws a careful line: the machines excel at straightforward yes-or-no judgments yet remain noticeably shakier when asked to choose among multiple surgical options, a nuance the researchers say should shape how clinics deploy these tools.

Refractive surgery is one of the most common elective procedures in medicine, encompassing techniques that reshape the cornea with lasers or implant a corrective lens inside the eye. Choosing the right procedure is far from trivial. Surgeons must weigh corneal thickness, refractive error, anterior chamber depth, pupil size, lifestyle factors and a catalogue of contraindications. A patient who is an ideal candidate for Small Incision Lenticule Extraction, or SMILE, might be a poor fit for transepithelial photorefractive keratectomy, known as TransPRK, and vice versa. Misjudging that calculus can lead to complications, retreatments or long-term visual problems, which is why surgical suitability decisions are traditionally reserved for trained specialists.

The study, led by Qi Wan, Ran Wei, Jing Tang, Ying-ping Deng and Ke Ma of the Department of Ophthalmology at West China Hospital of Sichuan University, drew on preoperative data from 11,966 consecutive patients evaluated at their institution. That sheer scale is what makes the work unusual. Most AI-in-medicine benchmarks rely on a few hundred curated cases; this one threw nearly twelve thousand messy, real-world clinical records at the algorithms. To establish ground truth, a panel of three senior refractive surgeons independently reviewed every case, assigning each a recommendation score from 0 to 100 and a suitability classification across four procedures: Femtosecond LASIK, SMILE, TransPRK and the Implantable Collamer Lens, or ICL. The panel’s agreement was excellent, with a kappa statistic exceeding 0.85, meaning the experts rarely disagreed about what the right answer was.

Five large language models faced the exam: DeepSeek-Chat, GLM-4.7, GPT-4o, Kimi-K2-Thinking and Qwen-Max. Each was fed the same structured, expert-mimicking prompt for every case and asked to produce the same recommendation scores and classifications the human panel had given. Alongside the machines, an intermediate-level physician, representing a mid-career ophthalmologist, independently evaluated all 11,966 cases, providing a human benchmark that is arguably more realistic than comparing AI to world-renowned professors. The study then scored everyone, human and machine alike, on a battery of standard metrics: accuracy, the area under the receiver operating characteristic curve for binary decisions, Cohen’s kappa for multi-class agreement, correlation coefficients for the numeric scores, and the root mean square and mean absolute errors measuring how far predictions strayed from expert judgment.

The headline result is that the top-performing language models were remarkably good at the binary question of whether a patient is suitable for a given procedure. For LASIK and SMILE, the best models achieved accuracies above 98.5 percent and areas under the curve exceeding 0.96, figures that place them in territory generally considered excellent for clinical classification. Perhaps more striking was the performance of the intermediate physician, who scored significantly lower than the leading models, particularly on complex classifications. That comparison matters because it reframes the debate about AI in medicine. The question is no longer simply whether machines can match elite specialists, but whether they can consistently outperform the average clinician making routine decisions, and on this evidence, in this narrow domain, they can.

The picture grows more complicated when the task shifts from two options to four. In the multi-class setting, where a model must pick the single best-suited procedure among LASIK, SMILE, TransPRK and ICL, agreement with the expert panel was more modest. Qwen-Max led this metric, achieving a Cohen’s kappa of up to 0.743, which statisticians generally label substantial agreement but which falls well short of the panel’s own internal consistency. The researchers interpret this gap candidly: the models are best understood as decision-support tools rather than autonomous decision-makers. In other words, an AI can reliably flag that a patient is a candidate for corneal refractive surgery, but a human expert should still make the final call about which procedure, especially in borderline cases where corneal topography, pupil characteristics and patient preference interact in subtle ways.

Beyond raw accuracy, the study examined how closely the models’ numeric recommendation scores tracked the experts’ 0-to-100 ratings. Here Qwen-Max and DeepSeek-Chat showed strong correlation with the panel’s scores, suggesting the models grasp not just the category of recommendation but its gradations, recognizing, say, that a patient with thin corneas and moderate myopia deserves a lower suitability score for LASIK than one with abundant corneal tissue. Regression errors captured the same story: the strongest models deviated from expert scores by margins small enough to be clinically useful, while weaker models and the intermediate physician showed wider scatter. This graded, score-like behavior matters for real deployment, because a tool that merely says yes or no discards the nuance surgeons use to counsel patients about risk.

One of the study’s most practically significant contributions is its analysis of cost and speed, dimensions rarely quantified in medical AI evaluations. DeepSeek-Chat emerged as the efficiency champion, combining the lowest cost per query with the fastest response times while still delivering strong accuracy. GPT-4o, by contrast, was the most expensive of the five. These economics are not trivia. In high-volume screening scenarios, where thousands of preoperative assessments must be triaged before a surgeon ever sees the patient, a model that is nearly as accurate but a fraction of the price can transform workflow economics. The authors point to resource-limited settings, including regions with few refractive surgeons, as the clearest beneficiaries, since a cheap, fast, accurate pre-screening layer could extend specialist-grade triage to populations that currently lack access.

Technically, the study also offers a template for how such evaluations should be run. The structured expert-mimicking prompt, the massive consecutive-patient dataset, the multi-surgeon gold standard with quantified inter-rater agreement, and the multi-dimensional metric battery together form a benchmarking blueprint that other specialties can copy. Too many published AI evaluations rest on convenience samples and single-metric reporting; this work demonstrates what a more rigorous standard looks like, including the honest acknowledgment that kappa values in multi-class tasks remain moderate. It is worth noting that the study is retrospective: the models reviewed recorded data rather than live patients, and real clinical deployment would raise additional questions about data privacy, liability, and how surgeons integrate algorithmic advice into consultations.

What emerges is a measured but genuinely exciting picture of where medical language models stand. On narrow, well-defined classification tasks grounded in structured clinical data, they now perform at or above the level of mid-career physicians, at pennies per case and in seconds. On the harder judgment calls that define expert practice, they still lag behind senior specialists and need human oversight. The West China Hospital team frames the technology exactly as the evidence supports: a powerful assistive layer for screening, triage and decision support, positioned to augment rather than replace the surgeon’s judgment. As these models continue to improve, and as prospective trials validate retrospective results, the routine preoperative assessment of refractive surgery candidates may become one of the first places where patients routinely benefit from an AI second opinion, whether or not they ever know it is there.

Subject of Research: Evaluation of large language models for refractive surgery recommendation and clinical decision support

Article Title: Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation

Article References: Wan, Q., Wei, R., Tang, J., Deng, Y.-P., & Ma, K. (2026). Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation. BMC Medicine. https://doi.org/10.1186/s12916-026-05262-4

Image Credits: AI Generated

DOI: 10.1186/s12916-026-05262-4

Keywords: large language models, refractive surgery, LASIK, SMILE, ophthalmology, artificial intelligence, clinical decision support, GPT-4o, DeepSeek, cost-effectiveness, machine learning, BMC Medicine

Cite Scienmag News

Ophelia Keating. (September 26, 2026). AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds. Scienmag. https://scienmag.com/ai-chatbots-match-human-doctors-in-recommending-laser-eye-surgery-study-finds/

Ophelia Keating. "AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds." Scienmag, 26 September 2026, https://scienmag.com/ai-chatbots-match-human-doctors-in-recommending-laser-eye-surgery-study-finds/. Accessed 26 September 2026.

Ophelia Keating. "AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds." Scienmag. September 26, 2026. https://scienmag.com/ai-chatbots-match-human-doctors-in-recommending-laser-eye-surgery-study-finds/

Tags: AI chatbotsAI limitations in complex medical choicesAI triage for eye proceduresAI vs human ophthalmologistsArtificial IntelligenceBMC Medicineclinical decision supportCost-effectivenessDeepSeekGPT-4olarge language modelslarge language models in healthcarelaser eye surgery decision-makingLASIKLASIK and SMILE surgical recommendationsMachine learningmedical AI performance comparisonmedical decision support AIophthalmologyophthalmology AI diagnosisreal-world AI validation in ophthalmologyrefractive surgeryrefractive surgery AI accuracySMILE
Share26Tweet16
Previous Post

Deep Infiltrating Endometriosis Linked to Higher Placenta and Newborn Risks in Pregnancy

Next Post

Why Self-Objectifying Women Feel Lonelier: Rivalry With Other Women Holds the Key

Related Posts

Deep Infiltrating Endometriosis Linked to Higher Placenta and Newborn Risks in Pregnancy
Medicine

Deep Infiltrating Endometriosis Linked to Higher Placenta and Newborn Risks in Pregnancy

September 26, 2026
Cold Packs, Nerve Stimulation and Radiofrequency Emerge as Drug-Free Options for Postpartum Perineal Pain
Medicine

Cold Packs, Nerve Stimulation and Radiofrequency Emerge as Drug-Free Options for Postpartum Perineal Pain

September 26, 2026
Vision Problems May Be the Earliest Warning Sign of Lewy Body Dementia, New Study Finds
Medicine

Vision Problems May Be the Earliest Warning Sign of Lewy Body Dementia, New Study Finds

September 26, 2026
When Fat Drives Sleep Apnea: Scientists Push to Define Adiposity-Attributable Disease
Medicine

When Fat Drives Sleep Apnea: Scientists Push to Define Adiposity-Attributable Disease

September 26, 2026
Vitamin-Derived Molecule Shows First Signs of Treating Mitochondrial DNA Disease in Humans
Medicine

Vitamin-Derived Molecule Shows First Signs of Treating Mitochondrial DNA Disease in Humans

September 26, 2026
Racing the Clock: Athletes’ Glucose Soars Higher in Competition Than Training
Medicine

Racing the Clock: Athletes’ Glucose Soars Higher in Competition Than Training

September 25, 2026
Next Post
Why Self-Objectifying Women Feel Lonelier: Rivalry With Other Women Holds the Key

Why Self-Objectifying Women Feel Lonelier: Rivalry With Other Women Holds the Key

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • New AI Study Reveals What Actually Matters When Mapping Tissue Architecture
  • Viral Saboteur Unmasked: How EBV Silences a Key Immune Molecule to Drive Nasopharyngeal Cancer
  • Why Self-Objectifying Women Feel Lonelier: Rivalry With Other Women Holds the Key
  • AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading