Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Science Education

AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds

October 1, 2026
in Science Education
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 5 mins read
0
AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds

AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds

AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Artificial intelligence chatbots have become ubiquitous study companions for students in medicine and dentistry, promising instant answers to virtually any exam question. But a new study from Hacettepe University in Türkiye suggests that getting the right answer is only half the story. When three leading large language models were put to the test on real periodontology exam questions, they performed remarkably similarly in terms of raw accuracy — yet differed significantly in how well they could explain their reasoning. The research, published in BMC Medical Education, reveals that the type and topic of a question shape the quality of an AI’s explanation far more than whether the answer itself is correct.

The study, conducted by Hanife Merva Parlak of the Department of Periodontology at Hacettepe University’s Faculty of Dentistry in Ankara, set out to answer a deceptively simple question: do question type and topic affect how well large language models answer periodontology questions? This matters because the debate over artificial intelligence in health science education has largely focused on accuracy scores, with less attention paid to the explanatory quality that determines whether a chatbot can actually teach rather than merely answer. A model that picks the correct option but offers a vague, inconsistent, or misleading rationale may be worse for learning than one that occasionally errs but explains itself clearly.

To investigate, Parlak assembled 134 multiple-choice questions drawn from the Dental Specialization Examination administered in Türkiye, the high-stakes test that dentists must pass to enter specialty training. The questions were fed to three widely used consumer-facing artificial intelligence systems: ChatGPT-5, Gemini 2.5 Flash, and Copilot. Each model’s responses were then assessed on two separate dimensions. The first was straightforward accuracy — did the model select the correct answer? The second was more nuanced: the quality and adequacy of the explanation accompanying each answer, judged on a Likert scale, a standard survey instrument that rates responses along a graded spectrum of quality.

The questions themselves were not uniform. As described in the study’s supplementary materials, they spanned different cognitive levels, following the classic taxonomy of educational objectives. Standard questions tested basic recall — the remembering level — such as memorized facts about periodontal disease classification. Information-based questions probed understanding, requiring the model to interpret and apply knowledge in a slightly transformed context. Case-based questions, the most demanding category, simulated clinical scenarios at the applying level, asking models to reason through patient presentations the way a periodontist would. The question pool also covered both basic science content and clinical periodontology, allowing the researcher to disentangle whether models perform differently on foundational biology versus applied clinical reasoning.

The headline finding on accuracy was, in a sense, a non-finding: the total accuracy rates of ChatGPT-5, Gemini 2.5 Flash, and Copilot were statistically similar. All three models have absorbed vast corpora of biomedical text, and by now the ability of frontier chatbots to answer multiple-choice dental questions at or near specialist level has been demonstrated repeatedly across medical specialties. What distinguished the models was something subtler. Gemini 2.5 Flash outperformed its rivals in the adequacy of its response explanations, a difference that reached statistical significance. In other words, when it came to articulating why an answer was correct, Gemini’s outputs were judged more complete and more useful than those of ChatGPT-5 or Copilot.

The pattern deepened when the analysis turned to question categories. Gemini provided more consistent explanations than the other models on basic science questions and on clinical periodontology questions alike, and it maintained that consistency across standardized and information-based question formats. Each of these differences was statistically significant. This suggests that Gemini’s advantage was not confined to one corner of the syllabus but reflected a more general capacity to produce coherent, pedagogically sound rationales across the breadth of periodontology content tested.

To move beyond simple group comparisons, the study employed regression analyses, statistical techniques that estimate how much each predictor variable contributes to an outcome while accounting for the others. The results were telling: question type and topic were both significantly related to explanation quality. Accuracy, by contrast, appeared comparatively robust to these factors. The models generally knew the right answer regardless of whether a question tested recall or clinical application, or whether it concerned the microbiology of periodontal pockets or the surgical management of mucogingival defects. But the clarity, depth, and consistency of the reasoning they offered in support of those answers fluctuated considerably depending on what was being asked and how.

This dissociation between accuracy and explanation quality carries real implications for how artificial intelligence should be integrated into health professional education. Multiple-choice examinations, including the specialty exams used to license and certify clinicians, reward the selection of a correct option. A student using a chatbot as a study aid, however, typically learns from the explanation, not the letter of the answer. If a model’s rationale is thin, inconsistent, or subtly wrong even when its final choice is right, the student may internalize flawed mental models that surface later at the chairside. Conversely, an explanation of high quality can transform a simple quiz question into a miniature tutorial, connecting the answer to underlying mechanisms, differential diagnoses, and treatment principles.

The findings also speak to the technical character of large language models. These systems generate answers by predicting likely text sequences based on patterns learned during training, rather than by consulting a verified knowledge base. Their factual accuracy on well-represented exam content can be excellent, because such content appears abundantly in textbooks, review articles, and question banks. Explanation quality, however, depends on the model’s ability to organize and verbalize reasoning in a way that is both correct and appropriately calibrated to the question’s cognitive demand — a harder task that varies more with the structure of the prompt. Case-based questions, which require integrating scattered clinical clues, appear to stress this capacity in ways that simple factual recall does not, and the regression results indicate that this stress shows up in explanation quality even when the final answer survives intact.

Parlak’s conclusion is measured. Large language models, the study finds, are promising tools for periodontology education, but they retain limitations that require improvement before they can be relied upon uncritically. The author also cautions that the findings should be interpreted within the constraints of the study design: the questions came from a single national specialty examination, were posed in Turkish, and the evaluation of explanation quality, though systematic, ultimately rests on structured human judgment rather than an objective gold standard. Consumer chatbot interfaces also evolve rapidly, meaning that any snapshot comparison reflects a particular moment in a fast-moving technological race. Still, the central message is likely to resonate well beyond periodontology. As educators worldwide weigh whether to embrace, restrict, or redesign assessment in the age of generative artificial intelligence, this study suggests the right question is not simply whether chatbots get the answer right, but whether they can explain it in a way that actually teaches. On that measure, the models are not yet interchangeable — and the format of the question matters more than anyone might have guessed.

Subject of Research: Comparative performance of large language models in answering periodontology multiple-choice questions

Article Title: Do question type and topic affect the performance of large language models in answering periodontology questions? A comparative study

Article References: Parlak, H. M. (2026). Do question type and topic affect the performance of large language models in answering periodontology questions? A comparative study. BMC Medical Education. https://doi.org/10.1186/s12909-026-10516-z

Image Credits: AI Generated

DOI: 10.1186/s12909-026-10516-z

Keywords: large language models, artificial intelligence, periodontology, dental education, ChatGPT-5, Gemini 2.5 Flash, Copilot, multiple-choice questions, medical education, explanation quality, question type, regression analysis

Cite Scienmag News

Courtney Benton. (October 1, 2026). AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds. Scienmag. https://scienmag.com/ai-chatbots-ace-periodontology-facts-but-stumble-on-explanations-study-finds/

Courtney Benton. "AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds." Scienmag, 1 October 2026, https://scienmag.com/ai-chatbots-ace-periodontology-facts-but-stumble-on-explanations-study-finds/. Accessed 1 October 2026.

Courtney Benton. "AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds." Scienmag. October 1, 2026. https://scienmag.com/ai-chatbots-ace-periodontology-facts-but-stumble-on-explanations-study-finds/

Tags: accuracy versus explanation quality in AIAI chatbots in medical and dental educationAI in health science educationAI-assisted learning in periodontologyArtificial Intelligencechallenges of AI explainability in medicineChatGPT-5Copilotdental educationeffectiveness of AI chatbots in medical examsevaluation of AI chatbots for clinical knowledgeexplanation qualityGemini 2.5 Flashimpact of question type on AI performancelarge language modelslarge language models in healthcarelimitations of AI in explaining reasoningMedical Educationmultiple-choice questionsperiodontologyquestion typeregression analysisresearch on AI explanation capabilitiesrole of AI in dental student study tools
Share26Tweet16
Previous Post

How Microbial Imbalances Shape Lung Cancer Growth and Treatment Failure

Next Post

Funnel-Shaped Rectenna Array Turns 5G Millimeter-Wave Signals Into Usable Power

Related Posts

Europe’s Endocrinologists Open New School to Tackle the Obesity Crisis
Science Education

Europe’s Endocrinologists Open New School to Tackle the Obesity Crisis

October 1, 2026
Who Slips Through the Cracks? Massive Study Maps Cervical Screening Non-Attendance in Flanders
Science Education

Who Slips Through the Cracks? Massive Study Maps Cervical Screening Non-Attendance in Flanders

October 1, 2026
Scientists Map 25 Years of Sustainability Research in Farming Education
Science Education

Scientists Map 25 Years of Sustainability Research in Farming Education

October 1, 2026
AI Grading on Campus: New Model Says Legitimacy, Not Capability, Decides Success
Science Education

AI Grading on Campus: New Model Says Legitimacy, Not Capability, Decides Success

October 1, 2026
New Nurses Rate Kuwait’s Ministry of Health Rotation Programme Highly, but Call for Better Supervision
Science Education

New Nurses Rate Kuwait’s Ministry of Health Rotation Programme Highly, but Call for Better Supervision

October 1, 2026
Physiotherapist-Led Exercise Therapy Protects Young Knees After ACL Reconstruction
Science Education

Physiotherapist-Led Exercise Therapy Protects Young Knees After ACL Reconstruction

October 1, 2026
Next Post
Funnel-Shaped Rectenna Array Turns 5G Millimeter-Wave Signals Into Usable Power

Funnel-Shaped Rectenna Array Turns 5G Millimeter-Wave Signals Into Usable Power

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Funnel-Shaped Rectenna Array Turns 5G Millimeter-Wave Signals Into Usable Power
  • AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds
  • How Microbial Imbalances Shape Lung Cancer Growth and Treatment Failure
  • Nature Prescriptions for Older Adults Show Promise, But the Evidence Is Not Yet Strong Enough

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading