Wednesday, October 7, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Medicine

AI Chatbots Mostly Safe on Sudden Cardiac Death Advice, but Safety Gaps Persist Across All Six Models Tested

October 7, 2026
in Medicine
Ophelia Keating
By Ophelia Keating Scienmag Editorial Profile - Health Services Research
Reading Time: 6 mins read
0
AI Chatbots Mostly Safe on Sudden Cardiac Death Advice, but Safety Gaps Persist Across All Six Models Tested

AI Chatbots Mostly Safe on Sudden Cardiac Death Advice, but Safety Gaps Persist Across All Six Models Tested

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

When someone types a frightening question into a chatbot at two in the morning—Am I having a heart attack? Can my teenager collapse on the football field? Should I restart exercise after COVID?—the answer they receive may shape whether they call an ambulance, start chest compressions, or simply go back to sleep. A new peer-reviewed study published in BMC Public Health has put six of the world’s most widely used generative AI chatbots through one of the most demanding public-health stress tests yet devised, asking them hundreds of real questions about sudden cardiac death and scoring every answer for safety, accuracy, empathy, quality, transparency, and readability. The results are both encouraging and sobering: most responses were judged safe and clinically relevant, yet every single model produced at least some answers containing potential safety concerns, and the recurring failure modes clustered around exactly the moments when minutes matter most.

The research team, led by cardiologists at Bozhou People’s Hospital in Anhui Province, China, together with a colleague at the First Affiliated Hospital of the University of Science and Technology of China, evaluated ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, Doubao, and Perplexity. Rather than inventing test questions in a conference room, the investigators built a standardized set of 56 sudden-cardiac-death-related public consultation questions from an unusually broad evidence base: Google Trends search data, online patient forums, public question-and-answer platforms, clinical guidelines, expert consensus statements, systematic reviews, and qualitative interviews with patients and family members. The topics spanned the full landscape of lay concern about sudden cardiac death, including symptom triage, emergency response, cardiopulmonary resuscitation and automated external defibrillator use, inherited cardiac risk, myocarditis-related worries, returning to exercise after illness, screening, and decision-making about implantable cardioverter-defibrillators.

Each of the 56 questions was submitted once to each of the six chatbots between May 10 and May 16, 2026, generating 336 individual model-response items for analysis. Five senior cardiology raters, blinded to which chatbot had produced which answer, then scored every response across a battery of validated instruments: a safety assessment, an accuracy assessment, an empathy scale, the DISCERN instrument for judging the quality of written health information, the EQIP instrument for ensuring the quality of patient education materials, the JAMA benchmark criteria, and the Global Quality Scale. Readability was calculated separately using six formula-based indices. Statistical comparisons between models used Cochran’s Q test and Friedman tests, with Kendall’s W reported as the effect size, and the agreement among raters was strikingly high, with a Fleiss’ kappa of 0.856 for the safety domain and intraclass correlation coefficients ranging from 0.801 to 0.887 for the other manually assessed metrics.

The headline finding on safety was genuinely reassuring at first glance. Safe responses predominated across all six models, suggesting that the current generation of publicly accessible chatbots has largely absorbed the basic guardrails of cardiac emergency communication. But the more consequential finding was that responses containing potential safety concerns appeared in every chatbot’s output during the study window. ChatGPT Plus 5.5 thinking and Perplexity showed the highest observed proportions of safe responses, while Doubao showed the lowest. Importantly, the qualitative severity analysis found that most flagged responses contained only a single potential safety concern, and the distribution of concerns was concentrated at the low and low-to-moderate severity levels, with higher-severity problems occurring less frequently. That pattern is consistent with a technology that has improved substantially but has not yet reached the reliability threshold that safety-critical health communication demands.

The qualitative analysis of flagged responses is arguably the most clinically valuable part of the study, because it names the specific failure patterns that clinicians and developers can act on. The reviewers repeatedly identified answers that delayed emergency activation for chest pain, syncope, or post-viral symptoms—the classic red-flag presentations of sudden cardiac death where immediate escalation to emergency medical services is the single most important instruction. They found instances of over-reassurance about prevention, implying that a person’s risk had been dismissed when individualized assessment was warranted. They documented fixed return-to-exercise timelines after infection or COVID-19, offered as blanket rules rather than guidance tailored to a patient’s cardiac evaluation. Some responses gave pulse-check instructions that could delay the start of CPR, a direct contradiction of resuscitation guidance that emphasizes rapid compressions. Others delivered oversimplified advice on screening, electrolyte disturbances, or implantable cardioverter-defibrillator decisions, flattening genuinely complex clinical judgments into deceptively simple answers.

Beyond safety, the study revealed sharp and statistically significant differences between models on nearly every quality dimension. Accuracy differed significantly across the six chatbots, with a Kendall’s W of 0.864 indicating strong rank concordance among raters; ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, and Perplexity achieved similar median accuracy scores, while Doubao performed lower. Empathy also differed significantly, with a Kendall’s W of 0.531, and DeepSeek-V4-pro recorded the highest observed empathy score—a notable result, since empathetic framing is increasingly recognized as a determinant of whether anxious users actually follow health advice. Significant differences emerged for DISCERN, EQIP, JAMA benchmark, and Global Quality Scale scores as well, all at P < 0.001, confirming that the choice of chatbot materially changes the quality of the health information a member of the public receives.

DeepSeek-V4-pro emerged as the strongest overall performer on information quality within this dataset, achieving the highest DISCERN and EQIP scores and the most favorable readability profile, while Gemini 3.5 Flash showed comparable EQIP performance and Microsoft Copilot smart recorded the highest JAMA benchmark score. Yet two weaknesses were universal. Transparency remained limited across the board: chatbots rarely made clear where their information came from, how current it was, or how confident users should be in it, which matters enormously when a reader must decide whether to trust an answer about inherited cardiac risk. And every model produced responses with relatively high reading-grade requirements, meaning that the very populations most likely to need accessible cardiac health information—older adults, people with limited health literacy, non-native speakers—may struggle to extract the key safety messages from dense, jargon-laden text.

The authors are admirably explicit about the limits of their own findings. Each question was sampled only once during a single defined week in May 2026, and all prompts were submitted in English through public interfaces accessed from China. Generative models are stochastic and continuously updated, so the same question asked a month later might yield a different answer. The researchers therefore frame their results as a time-, language-, and access-specific snapshot rather than a stable ranking of the underlying models. This caveat is not a weakness of the study but a feature of honest evaluation methodology in a field where the tools change faster than the literature can assess them, and it underscores why one-off benchmark scores should never be treated as permanent verdicts on any AI system.

The practical implications are clear for three audiences at once. For members of the public, the message is that chatbots can be useful for general education about sudden cardiac death—understanding what an implantable defibrillator does, why myocarditis follows some viral infections, or what screening involves—but they must never substitute for calling emergency services or consulting a clinician when symptoms suggest a cardiac emergency. For clinicians and public health communicators, the study identifies the exact content areas where AI guidance most often goes astray: emergency escalation, CPR and AED instructions, return-to-exercise timing after infection, individualized risk interpretation, and screening decisions. For developers, the study’s recommendations read as a roadmap: prioritize evidence-linked content, build clear red-flag escalation into every relevant response, adopt plain-language communication, and submit systems to ongoing expert review rather than one-time validation.

What makes this study resonate beyond cardiology is its methodological template. By combining blinded expert scoring across validated instruments with a fine-grained qualitative analysis of safety failures, and by anchoring the question set in the genuine concerns of patients and families, the researchers have shown how AI health tools should be evaluated: not with a single accuracy number, but across the full multidimensional surface of safety, accuracy, empathy, quality, transparency, and readability. Sudden cardiac death is an unforgiving test case—outcomes are measured in minutes, and a subtly wrong answer about chest pain or pulse-checking can be fatal. On that test, the six chatbots earned a qualified pass: good enough to inform, not yet good enough to advise alone. As generative AI becomes a default first stop for health questions worldwide, studies like this one provide the evidence base that regulators, developers, and clinicians will need to decide where these tools belong in the chain of care—and where they must, without exception, hand the user straight to a human.

Subject of Research: Evaluation of generative AI chatbots for public health consultation on sudden cardiac death

Article Title: Multidimensional evaluation of generative AI chatbots for public consultation on sudden cardiac death: safety, accuracy, empathy, reliability, quality, and readability across six models

Article References: Chen, Y., Ma, H., Wang, H., Chen, D., Wang, H., & Jiang, R. (2026). Multidimensional evaluation of generative AI chatbots for public consultation on sudden cardiac death: safety, accuracy, empathy, reliability, quality, and readability across six models. BMC Public Health. https://doi.org/10.1186/s12889-026-29766-z

Image Credits: AI Generated

DOI: 10.1186/s12889-026-29766-z

Keywords: generative AI, chatbots, sudden cardiac death, public health, CPR, AED, patient safety, health communication, readability, empathy, large language models, BMC Public Health

Cite Scienmag News

Ophelia Keating. (October 7, 2026). AI Chatbots Mostly Safe on Sudden Cardiac Death Advice, but Safety Gaps Persist Across All Six Models Tested. Scienmag. https://scienmag.com/ai-chatbots-mostly-safe-on-sudden-cardiac-death-advice-but-safety-gaps-persist-across-all-six-models-tested/

Ophelia Keating. "AI Chatbots Mostly Safe on Sudden Cardiac Death Advice, but Safety Gaps Persist Across All Six Models Tested." Scienmag, 7 October 2026, https://scienmag.com/ai-chatbots-mostly-safe-on-sudden-cardiac-death-advice-but-safety-gaps-persist-across-all-six-models-tested/. Accessed 7 October 2026.

Ophelia Keating. "AI Chatbots Mostly Safe on Sudden Cardiac Death Advice, but Safety Gaps Persist Across All Six Models Tested." Scienmag. October 7, 2026. https://scienmag.com/ai-chatbots-mostly-safe-on-sudden-cardiac-death-advice-but-safety-gaps-persist-across-all-six-models-tested/

Tags: AEDAI chatbot safety in medical adviceand other AI models in healthcareBMC Public Healthchatbotsclinical accuracy of generative AI chatbotscomparison of ChatGPTCPRcritical moments in emergency medical AI responsesempathyempathy and transparency in AI medical adviceevaluation of AI models for emergency health guidanceGeminigenerative AIhealth communicationimportance of accuracylarge language modelslimitations of current AI chatbots in life-threatening situationspatient safetypotential safety concerns in AI-generated health informationPublic healthpublic health stress testing of AI chatbotsreadabilitysafety gaps in AI health responsessudden cardiac deathsudden cardiac death risk assessment
Share26Tweet16
Previous Post

Pearls on a String: Astronomers Reveal Hidden Knots in Milky Way’s Only Type Iax Supernova Remnant

Next Post

NCCN Summit Puts Cancer Prevention and Screening at the Heart of Health Policy

Related Posts

After the ICU Door Closes: One in Four HIV Patients Rehospitized or Dead Within Weeks of Cryptococcal Meningitis Discharge
Medicine

After the ICU Door Closes: One in Four HIV Patients Rehospitized or Dead Within Weeks of Cryptococcal Meningitis Discharge

October 7, 2026
Landmark Study Maps the Normal Child Heart From Birth to 18 With MRI
Medicine

Landmark Study Maps the Normal Child Heart From Birth to 18 With MRI

October 6, 2026
Mindfulness Meets Magic Mushrooms: USC Trials Psilocybin Therapy for Depression
Medicine

Mindfulness Meets Magic Mushrooms: USC Trials Psilocybin Therapy for Depression

October 6, 2026
Coupon Clipping at the Pharmacy Counter: What Manufacturer Discounts Really Do to GLP-1 Drug Spending
Medicine

Coupon Clipping at the Pharmacy Counter: What Manufacturer Discounts Really Do to GLP-1 Drug Spending

October 6, 2026
Fast Walkers May Escape the Cognitive Toll of Aging, Dementia Risk Study Finds
Medicine

Fast Walkers May Escape the Cognitive Toll of Aging, Dementia Risk Study Finds

October 6, 2026
One Short Night of Sleep Can Skew Concussion Test Scores, Study Warns
Medicine

One Short Night of Sleep Can Skew Concussion Test Scores, Study Warns

October 6, 2026
Next Post
NCCN Summit Puts Cancer Prevention and Screening at the Heart of Health Policy

NCCN Summit Puts Cancer Prevention and Screening at the Heart of Health Policy

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • NCCN Summit Puts Cancer Prevention and Screening at the Heart of Health Policy
  • AI Chatbots Mostly Safe on Sudden Cardiac Death Advice, but Safety Gaps Persist Across All Six Models Tested
  • Pearls on a String: Astronomers Reveal Hidden Knots in Milky Way’s Only Type Iax Supernova Remnant
  • Sausage Tree Seeds Lose Their Power After Just One Year in Storage, Benin Study Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading