Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

AI Chatbots Fail Safety Warnings When Patients Ask About Pregabalin

October 1, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
AI Chatbots Fail Safety Warnings When Patients Ask About Pregabalin

AI Chatbots Fail Safety Warnings When Patients Ask About Pregabalin

AI Chatbots Fail Safety Warnings When Patients Ask About Pregabalin

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Millions of patients now turn to artificial intelligence chatbots with questions about their medications, from dosing schedules to side effects, often without ever speaking to a pharmacist or physician. A new study puts that habit under the microscope, and the results are a sobering mix of reassurance and alarm. Researchers from Burkina Faso, Niger, and Cameroon evaluated four of the world’s most widely used large language models, ChatGPT-4o, Microsoft Copilot, Gemini 2.0 Flash, and Claude 3.5 Sonnet, as they answered fifteen real patient questions about pregabalin, a blockbuster drug prescribed for epilepsy, neuropathic pain, generalized anxiety disorder, and fibromyalgia. The verdict: the chatbots were largely accurate, but they stumbled badly on the one thing that matters most in medication advice, warning patients when danger lurks.

The study, published in Discover Artificial Intelligence, was designed to mirror how ordinary people actually use these tools. On a single day, 30 June 2025, the team submitted all fifteen questions to each model through standard consumer web interfaces, exactly as a patient would, with no API access, no custom settings, and no prompt engineering. The questions themselves were drawn exhaustively from the official NHS page of common questions about pregabalin, a resource compiled and validated by clinical pharmacists for the general public. Each query came wrapped in a standardized instruction asking for a clear, factual, patient-appropriate answer that must include a warning if any potential danger existed. That explicit safety instruction is what makes the study’s central finding so striking.

Pregabalin was a deliberate and timely choice. The drug is a structural analogue of the neurotransmitter GABA, but it does not act on GABA receptors directly; instead it binds selectively to the alpha-2-delta-1 subunit of voltage-gated calcium channels, dampening the release of excitatory neurotransmitters. Its clinical footprint has expanded dramatically over the past decade. Global annual sales of gabapentinoids, the drug class that includes pregabalin, rose by 114.5 percent between 2012 and 2022, and by a staggering 180.9 percent in low- and middle-income countries, where specialist care is scarce and patients are especially likely to seek answers from digital platforms. Prescribing rates more than doubled over ten years, reaching 7.2 prescriptions per 100 patients in 2019, with the highest volumes among elderly patients juggling multiple medications. Off-label and recreational use have broadened the population asking questions even further.

To judge the answers, the researchers assembled a multidisciplinary panel of a rheumatologist, a neurologist, and a pharmacologist, each with at least five years of clinical experience, reflecting the three main specialties that prescribe and monitor pregabalin. The evaluators worked independently and blind to which model had produced each response, after a calibration session in which they jointly scored three practice answers to align their criteria. Six dimensions were assessed: warning compliance, accuracy, completeness, content safety, readability, and clinical relevance. The study followed the TRIPOD-LLM reporting guideline, an emerging standard for research involving large language models, and the team calculated inter-rater reliability using Fleiss’ kappa and intraclass correlation coefficients.

The headline numbers on accuracy look encouraging. ChatGPT-4o, Gemini 2.0 Flash, and Copilot each scored 0.933 for factual correctness, while Claude 3.5 Sonnet scored 0.867. Clinical relevance ranged from 0.867 to 1.000 across the four models. But the confidence intervals overlapped so extensively that no model could claim reliable superiority over any other, a point the authors emphasize repeatedly. With only fifteen questions and a single query per model, the differences fall well within the statistical noise. The researchers are candid about this: the point estimates are best read as an illustration of the evaluation framework rather than stable performance benchmarks, and the single-query design means the stochastic variability of these models, which can produce substantively different answers to identical prompts, was never characterized.

The safety findings are where the study bites. Every one of the fifteen questions was judged by panel consensus to carry a clinically relevant safety concern warranting a warning, and the prompt explicitly demanded one. Yet warning compliance scores ranged only from 0.467 for ChatGPT-4o and Gemini 2.0 Flash to 0.733 for Claude 3.5 Sonnet. In plain terms, the models ignored a direct safety instruction in roughly 27 to 53 percent of cases. The authors argue this is arguably more alarming than a simple failure to volunteer warnings, because it demonstrates that even explicit prompting cannot guarantee reliable safety communication. For a drug with dependence potential, meaningful interaction risks, and a heavy presence among polypharmacy patients, that gap between instruction and behavior is the study’s most clinically consequential result.

When the researchers dissected the errors, a clear pattern emerged. Oversimplification dominated, accounting for 56.3 percent of all recorded errors, followed by omission of essential information at 34.8 percent, with outright hallucinations making up 8.9 percent. Gemini 2.0 Flash logged the most errors at 32, while Claude 3.5 Sonnet logged the fewest at 25. The authors connect these patterns to plausible, though unproven, architectural explanations: token probability maximization may favor statistically common, simplified phrasing over clinically nuanced detail; attention-based architectures may underweight rare but critical content; and reinforcement learning from human feedback may optimize for fluency and consensus at the cost of completeness. Hallucinations, though rarest, remain the most dangerous failure mode, since a fabricated dosing recommendation or a fictitious drug interaction could translate directly into harm.

The study is notable as much for its transparency as for its findings. In an unusual move, the authors disclose that the rater-by-item scoring matrix and the original response corpus were not archived after analysis, that the rule used to collapse three raters’ scores into one per-question value was never recorded, and that formal paired statistical tests such as McNemar’s test therefore became impossible. Copilot’s characteristic inline citations may also have compromised the blinding protocol. Rather than presenting plausible reconstructions as fact, the team reports these gaps openly, framing the work as an exploratory evaluation whose primary contribution is the multidimensional framework itself and a model for honest documentation of methodological limitations in early-stage LLM research.

What should be done with these results? The authors recommend hybrid architectures that ground chatbot outputs in structured, validated pharmaceutical databases rather than relying on parametric knowledge alone, along with dedicated fine-tuning on pharmacovigilance corpora so that safety alerts are generated reliably rather than left to general instruction-following. They also propose uncertainty detection mechanisms that would signal a model’s limits and steer users toward professional consultation, and even a continuous error-monitoring system modeled on existing pharmacovigilance infrastructure. The limitations are real: one medication, fifteen questions, one query per model, English only, and no testing with actual patients. The team, based in Francophone West and Central Africa, already plans a follow-up in French, where the stakes may be even higher. For now, the message to patients is clear: chatbots can be a reasonable starting point for basic medication facts, but when it comes to the warnings that keep you safe, they still cannot be trusted to speak up.

Subject of Research: Evaluation of large language models answering patient questions about the medication pregabalin

Article Title: A structured exploratory multidisciplinary evaluation of four large language models responding to frequently asked patient questions about pregabalin

Article References: A structured exploratory multidisciplinary evaluation of four large language models responding to frequently asked patient questions about pregabalin. (n.d.). https://doi.org/10.1007/s44163-026-02394-7

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02394-7

Keywords: large language models, ChatGPT-4o, Copilot, Gemini, Claude, pregabalin, patient safety, medication information, warning compliance, hallucination, pharmacovigilance, TRIPOD-LLM

Cite Scienmag News

Denise Maddox. (October 1, 2026). AI Chatbots Fail Safety Warnings When Patients Ask About Pregabalin. Scienmag. https://scienmag.com/ai-chatbots-fail-safety-warnings-when-patients-ask-about-pregabalin/

Denise Maddox. "AI Chatbots Fail Safety Warnings When Patients Ask About Pregabalin." Scienmag, 1 October 2026, https://scienmag.com/ai-chatbots-fail-safety-warnings-when-patients-ask-about-pregabalin/. Accessed 1 October 2026.

Denise Maddox. "AI Chatbots Fail Safety Warnings When Patients Ask About Pregabalin." Scienmag. October 1, 2026. https://scienmag.com/ai-chatbots-fail-safety-warnings-when-patients-ask-about-pregabalin/

Tags: AI chatbot accuracy in medical queriesAI chatbot medication safety warningsAI chatbot risks in pharmacologyAI chatbots pregabalin side effectsAI in patient education and safetyAI-driven medical information reliabilityChatGPT-4oClaudeCopilotevaluation of ChatGPT-4 and similar modelsGeminihallucinationhealthcare AI safety concernslarge language modelslarge language models patient medication advicemedication informationmedication warning system limitationsnatural language processing in healthcarepatient safetypatient safety and AI chatbot oversightpharmacovigilancepregabalinTRIPOD-LLMwarning compliance
Share26Tweet16
Previous Post

Children See Fractions as Parts, Adults See Them as Wholes, Study Finds

Next Post

AI Meets Expert Judgment to Map Safer Waste Dumps in Indian City

Related Posts

Swarms of Drones Learn to Search Smarter With Brain-Inspired Game Theory
Technology and Engineering

Swarms of Drones Learn to Search Smarter With Brain-Inspired Game Theory

October 1, 2026
Brain Rhythms Reveal How Skin Temperature Shapes the Feeling of Body Ownership
Technology and Engineering

Brain Rhythms Reveal How Skin Temperature Shapes the Feeling of Body Ownership

October 1, 2026
Hypoxia-Programmed Macrophage Vesicles Turn the Immune System Into a Bone-Healing Engine
Technology and Engineering

Hypoxia-Programmed Macrophage Vesicles Turn the Immune System Into a Bone-Healing Engine

October 1, 2026
Seventeen Years of Classifier Chains: Landmark Review Maps the Hidden Backbone of Multi-Label AI
Technology and Engineering

Seventeen Years of Classifier Chains: Landmark Review Maps the Hidden Backbone of Multi-Label AI

October 1, 2026
Machine Learning Meets Operations Research in New Closed-Loop IoT Decision Engine
Technology and Engineering

Machine Learning Meets Operations Research in New Closed-Loop IoT Decision Engine

October 1, 2026
AI Heatmaps May Be Lying: New Test Exposes Flawed Explanations in Image-Recognition Networks
Technology and Engineering

AI Heatmaps May Be Lying: New Test Exposes Flawed Explanations in Image-Recognition Networks

October 1, 2026
Next Post
AI Meets Expert Judgment to Map Safer Waste Dumps in Indian City

AI Meets Expert Judgment to Map Safer Waste Dumps in Indian City

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • AI Meets Expert Judgment to Map Safer Waste Dumps in Indian City
  • AI Chatbots Fail Safety Warnings When Patients Ask About Pregabalin
  • Children See Fractions as Parts, Adults See Them as Wholes, Study Finds
  • Circular Economy Can Shield Against Known Shocks While Making Systems Brittle, Study Warns

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading