Sunday, September 20, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Medicine

ChatGPT Matches Psychiatrist Guidelines on Teen Eating Disorders, Study Finds

September 20, 2026
in Medicine
Glenn Wilkins
By Glenn Wilkins Scienmag Editorial Profile - Clinical Psychology
Reading Time: 5 mins read
0
ChatGPT Matches Psychiatrist Guidelines on Teen Eating Disorders, Study Finds

ChatGPT Matches Psychiatrist Guidelines on Teen Eating Disorders, Study Finds

ChatGPT Matches Psychiatrist Guidelines on Teen Eating Disorders, Study Finds

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Artificial intelligence chatbots are increasingly becoming a first stop for health information, and few areas demand more careful handling than adolescent eating disorders, where misinformation can carry life-threatening consequences. A new study published in the Journal of Eating Disorders has put one of the most widely used AI systems, ChatGPT, through a rigorous clinical examination, asking whether its answers to questions about teenage anorexia nervosa, bulimia and related conditions actually match the standards set by the American Psychiatric Association. The verdict is cautiously encouraging: the chatbot’s responses showed substantial alignment with the guideline, scored well on quality assessments, and were judged reliable by a panel of specialists. Yet the researchers also found enough variability in reliability and a reading difficulty level that may exclude many families, suggesting that AI tools remain a supplement to clinical expertise rather than a substitute for it.

The research team, led by Ayşe Gül Güven of the Division of Adolescent Medicine at Ankara University Faculty of Medicine, together with colleagues from adolescent medicine, child and adolescent psychiatry, public health and clinical nutrition, conducted a cross-sectional content analysis of ChatGPT’s performance. They constructed 34 clinical questions drawn directly from the American Psychiatric Association’s 2023 eating disorders guideline, covering the full arc of care from recognition and diagnosis through nutritional rehabilitation and medical monitoring to psychological and pharmacological treatment. Each question was posed to ChatGPT, and the responses it generated became the dataset for expert evaluation. Because the study analyzed AI-generated text rather than human participants, ethical approval was not required, and the authors note that ChatGPT was used solely as the subject of evaluation, not as a tool in study design, data analysis or manuscript preparation.

To judge the responses, the researchers assembled a panel of nine independent experts spanning three professional groups: physicians specializing in adolescent medicine, child and adolescent psychiatrists, and dietitians. This multidisciplinary design was deliberate. Eating disorders sit at the intersection of pediatrics, psychiatry and nutrition science, and each discipline brings a distinct lens to questions about refeeding protocols, weight restoration targets, family-based treatment and the medical complications of starvation. Every expert rated the same AI-generated answers, allowing the team to measure not only how good the responses were but also how consistently professionals from different backgrounds agreed on that judgment.

Quality and reliability were quantified using two established instruments. The Global Quality Score, or GQS, rates health information on a scale in which higher values indicate better overall quality, taking into account accuracy, completeness and suitability for the intended audience. The modified DISCERN scale, or mDISCERN, evaluates the reliability of published health information by examining whether sources are cited, whether aims are clear, and whether the material distinguishes evidence from speculation. The mean GQS scores pointed to high quality across all three clinical domains tested: diagnosis and assessment earned 4.27 plus or minus 0.39, nutritional management and medical monitoring earned 4.25 plus or minus 0.50, and treatment and clinical interventions earned 4.37 plus or minus 0.33. The corresponding mDISCERN scores of 29.24 plus or minus 4.25, 30.99 plus or minus 3.12 and 29.75 plus or minus 4.66 indicated reasonable reliability, with the treatment domain performing particularly well.

One of the most striking findings was the consistency of evaluation across professional backgrounds. Statistical testing revealed no significant differences in GQS or mDISCERN scores between the adolescent medicine physicians, the psychiatrists and the dietitians, with all comparisons yielding p values above 0.05. In practical terms, this means that a nutrition specialist judging an AI answer about refeeding and a psychiatrist judging an AI answer about psychotherapy converged on similar assessments of quality. Such convergence strengthens confidence that the high scores were not an artifact of one discipline’s perspective but reflected genuine alignment between the chatbot’s output and guideline-based clinical content.

Agreement among raters, however, was not uniform. The researchers calculated Fleiss’ kappa coefficients, a statistical measure of inter-rater agreement that corrects for agreement expected by chance, and found values ranging from 0.28 to 0.56 across disciplines. Kappa statistics in this range span what researchers conventionally label fair to moderate agreement, indicating that while the experts broadly concurred, individual judgments about the reliability of specific responses varied. This variability is itself informative: it suggests that evaluating AI-generated clinical content is not a fully standardized exercise, and that different specialists may weight aspects such as completeness, safety caveats or source transparency differently when reading the same answer.

Readability emerged as the study’s clearest caution. Using multiple standard readability indices, the team found that ChatGPT’s responses to eating disorder questions demanded a reading level spanning from high school to college difficulty. For a topic in which patients are often adolescents and the people searching for answers are frequently worried parents, this is a meaningful barrier. Health communication research has long argued that accessible medical information should target roughly an eighth-grade reading level, far below what the chatbot produced here. An AI system can be factually accurate and still fail a family that cannot comfortably parse its sentences, and the authors highlight this gap as a key limitation of relying on such tools for patient-facing information.

The findings arrive amid explosive growth in the use of large language models for health queries. These models generate fluent, confident-sounding text by predicting likely word sequences based on vast training corpora, and their outputs can look authoritative even when they drift from evidence. That is precisely why guideline-concordance studies matter. By comparing AI answers against a curated, expert-vetted document such as the APA’s 2023 eating disorders guideline, researchers can measure whether the model’s statistical fluency translates into clinically defensible content. In this case, the answer was largely yes: the chatbot demonstrated substantial concordance with the guideline’s recommendations across diagnosis, nutritional management and treatment, a result the authors describe as evidence that AI systems may serve as complementary sources for accessing guideline-based information.

The researchers are careful, however, to frame the result within its limits. Reliability scores, while reasonable, showed variability, and the Fleiss’ kappa range indicates that even trained clinicians did not always agree on how dependable a given response was. Readability levels were high enough to raise concerns about equitable access to the information. And the study evaluated a single AI system’s responses at a single point in time; large language models are updated frequently, and their outputs can differ between sessions, meaning performance measured today may not describe the system a patient encounters tomorrow. The authors conclude that artificial intelligence tools should be used cautiously and should not replace clinical expertise or evidence-based guidelines in clinical decision-making.

For clinicians, the study offers a practical message: AI chatbots are neither a minefield to be banned nor an oracle to be trusted, but a rapidly evolving information layer whose quality can and should be audited against authoritative standards. For families navigating the frightening early days of a suspected eating disorder, the takeaway is more direct. The answers a chatbot provides may indeed reflect what psychiatric guidelines recommend, but the reading level may be challenging, the reliability may vary from question to question, and nothing generated by a language model replaces the assessment of a trained adolescent medicine physician, psychiatrist or dietitian. As AI systems become more embedded in everyday health searches, studies of this kind provide the benchmark data needed to hold them to the standards of the professions whose knowledge they summarize.

Subject of Research: Evaluation of ChatGPT response quality, reliability, and readability against APA guidelines for adolescent eating disorders

Article Title: Consistency of ChatGPT responses with the American Psychiatric Association guidelines in adolescent eating disorders: evaluation of quality, reliability, and readability

Article References: Güven, A. G., Mirioğlu, S., Kurt, B., Temeltürk, R. D., & Aycan, Z. (2026). Consistency of ChatGPT responses with the American Psychiatric Association guidelines in adolescent eating disorders: evaluation of quality, reliability, and readability. Journal of Eating Disorders. https://doi.org/10.1186/s40337-026-01749-w

Image Credits: AI Generated

DOI: 10.1186/s40337-026-01749-w

Keywords: ChatGPT, artificial intelligence, adolescent eating disorders, American Psychiatric Association guidelines, large language models, Global Quality Score, modified DISCERN, readability, anorexia nervosa, clinical decision-making, Ankara University, Journal of Eating Disorders

Cite Scienmag News

Glenn Wilkins. (September 20, 2026). ChatGPT Matches Psychiatrist Guidelines on Teen Eating Disorders, Study Finds. Scienmag. https://scienmag.com/chatgpt-matches-psychiatrist-guidelines-on-teen-eating-disorders-study-finds/

Glenn Wilkins. "ChatGPT Matches Psychiatrist Guidelines on Teen Eating Disorders, Study Finds." Scienmag, 20 September 2026, https://scienmag.com/chatgpt-matches-psychiatrist-guidelines-on-teen-eating-disorders-study-finds/. Accessed 20 September 2026.

Glenn Wilkins. "ChatGPT Matches Psychiatrist Guidelines on Teen Eating Disorders, Study Finds." Scienmag. September 20, 2026. https://scienmag.com/chatgpt-matches-psychiatrist-guidelines-on-teen-eating-disorders-study-finds/

Tags: adolescent eating disordersadolescent mental health information accuracyAI chatbot accuracy for adolescent anorexia and bulimiaAI readability and accessibility for familiesAI tool reliability in eating disorder diagnosisAI versus psychiatrist standards in eating disorder treatmentAI-driven health information for teenagersAmerican Psychiatric Association eating disorder guidelines evaluationAmerican Psychiatric Association guidelinesAnkara Universityanorexia nervosaArtificial IntelligenceChatGPTChatGPT clinical performance in mental healthChatGPT content analysis for adolescent psychiatryclinical decision-makingclinical quality assessment of ChatGPT in adolescent mentalGlobal Quality ScoreJournal of Eating Disorderslarge language modelsmodified DISCERNreadabilitysupplementing clinical expertise with AI in eating disorder careteenage eating disorder guidelines
Share26Tweet16
Previous Post

What Does the Ocean We Want Actually Look Like? Four Countries Offer Answers

Next Post

Deleting a Detox Enzyme Shields Mouse Livers From Fat but Worsens Blood Sugar

Related Posts

Deleting a Detox Enzyme Shields Mouse Livers From Fat but Worsens Blood Sugar
Medicine

Deleting a Detox Enzyme Shields Mouse Livers From Fat but Worsens Blood Sugar

September 20, 2026
Families of Children Treated for Rare Diseases Say Long-Term Care Falls Short
Medicine

Families of Children Treated for Rare Diseases Say Long-Term Care Falls Short

September 20, 2026
Four Asia-Pacific Nations, Four Paths: Why Cancer Genomics Success Hinges on Health Systems, Not Sequencers
Medicine

Four Asia-Pacific Nations, Four Paths: Why Cancer Genomics Success Hinges on Health Systems, Not Sequencers

September 20, 2026
Colonoscopy Reveals Hidden Causes of Childhood Intussusception
Medicine

Colonoscopy Reveals Hidden Causes of Childhood Intussusception

September 20, 2026
Engineered nanopore reads amino acids, sugars and RNA building blocks at once
Medicine

Engineered nanopore reads amino acids, sugars and RNA building blocks at once

September 20, 2026
Poor Nutrition Makes People Smell More Attractive to Mosquitoes, Study Finds
Medicine

Poor Nutrition Makes People Smell More Attractive to Mosquitoes, Study Finds

September 20, 2026
Next Post
Deleting a Detox Enzyme Shields Mouse Livers From Fat but Worsens Blood Sugar

Deleting a Detox Enzyme Shields Mouse Livers From Fat but Worsens Blood Sugar

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • New Softball Simulation Captures Spin and Friction of Oblique Impacts
  • Deleting a Detox Enzyme Shields Mouse Livers From Fat but Worsens Blood Sugar
  • ChatGPT Matches Psychiatrist Guidelines on Teen Eating Disorders, Study Finds
  • What Does the Ocean We Want Actually Look Like? Four Countries Offer Answers

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading