<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI-assisted learning in periodontology &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-assisted-learning-in-periodontology/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 23:27:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI-assisted learning in periodontology &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Chatbots Ace Periodontology Facts but Stumble on Explanations, Study Finds</title>
		<link>https://scienmag.com/ai-chatbots-ace-periodontology-facts-but-stumble-on-explanations-study-finds/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 23:27:25 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[accuracy versus explanation quality in AI]]></category>
		<category><![CDATA[AI chatbots in medical and dental education]]></category>
		<category><![CDATA[AI in health science education]]></category>
		<category><![CDATA[AI-assisted learning in periodontology]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[challenges of AI explainability in medicine]]></category>
		<category><![CDATA[ChatGPT-5]]></category>
		<category><![CDATA[Copilot]]></category>
		<category><![CDATA[dental education]]></category>
		<category><![CDATA[effectiveness of AI chatbots in medical exams]]></category>
		<category><![CDATA[evaluation of AI chatbots for clinical knowledge]]></category>
		<category><![CDATA[explanation quality]]></category>
		<category><![CDATA[Gemini 2.5 Flash]]></category>
		<category><![CDATA[impact of question type on AI performance]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in healthcare]]></category>
		<category><![CDATA[limitations of AI in explaining reasoning]]></category>
		<category><![CDATA[Medical Education]]></category>
		<category><![CDATA[multiple-choice questions]]></category>
		<category><![CDATA[periodontology]]></category>
		<category><![CDATA[question type]]></category>
		<category><![CDATA[regression analysis]]></category>
		<category><![CDATA[research on AI explanation capabilities]]></category>
		<category><![CDATA[role of AI in dental student study tools]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224302</guid>

					<description><![CDATA[A new comparative study finds that ChatGPT-5, Gemini 2.5 Flash, and Copilot answer periodontology exam questions with similar accuracy, but question type and topic strongly influence the quality of their explanations.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence chatbots have become ubiquitous study companions for students in medicine and dentistry, promising instant answers to virtually any exam question. But a new study from Hacettepe University in Türkiye suggests that getting the right answer is only half the story. When three leading large language models were put to the test on real periodontology exam questions, they performed remarkably similarly in terms of raw accuracy — yet differed significantly in how well they could explain their reasoning. The research, published in BMC Medical Education, reveals that the type and topic of a question shape the quality of an AI&#8217;s explanation far more than whether the answer itself is correct.</p>
<p>The study, conducted by Hanife Merva Parlak of the Department of Periodontology at Hacettepe University&#8217;s Faculty of Dentistry in Ankara, set out to answer a deceptively simple question: do question type and topic affect how well large language models answer periodontology questions? This matters because the debate over artificial intelligence in health science education has largely focused on accuracy scores, with less attention paid to the explanatory quality that determines whether a chatbot can actually teach rather than merely answer. A model that picks the correct option but offers a vague, inconsistent, or misleading rationale may be worse for learning than one that occasionally errs but explains itself clearly.</p>
<p>To investigate, Parlak assembled 134 multiple-choice questions drawn from the Dental Specialization Examination administered in Türkiye, the high-stakes test that dentists must pass to enter specialty training. The questions were fed to three widely used consumer-facing artificial intelligence systems: ChatGPT-5, Gemini 2.5 Flash, and Copilot. Each model&#8217;s responses were then assessed on two separate dimensions. The first was straightforward accuracy — did the model select the correct answer? The second was more nuanced: the quality and adequacy of the explanation accompanying each answer, judged on a Likert scale, a standard survey instrument that rates responses along a graded spectrum of quality.</p>
<p>The questions themselves were not uniform. As described in the study&#8217;s supplementary materials, they spanned different cognitive levels, following the classic taxonomy of educational objectives. Standard questions tested basic recall — the remembering level — such as memorized facts about periodontal disease classification. Information-based questions probed understanding, requiring the model to interpret and apply knowledge in a slightly transformed context. Case-based questions, the most demanding category, simulated clinical scenarios at the applying level, asking models to reason through patient presentations the way a periodontist would. The question pool also covered both basic science content and clinical periodontology, allowing the researcher to disentangle whether models perform differently on foundational biology versus applied clinical reasoning.</p>
<p>The headline finding on accuracy was, in a sense, a non-finding: the total accuracy rates of ChatGPT-5, Gemini 2.5 Flash, and Copilot were statistically similar. All three models have absorbed vast corpora of biomedical text, and by now the ability of frontier chatbots to answer multiple-choice dental questions at or near specialist level has been demonstrated repeatedly across medical specialties. What distinguished the models was something subtler. Gemini 2.5 Flash outperformed its rivals in the adequacy of its response explanations, a difference that reached statistical significance. In other words, when it came to articulating why an answer was correct, Gemini&#8217;s outputs were judged more complete and more useful than those of ChatGPT-5 or Copilot.</p>
<p>The pattern deepened when the analysis turned to question categories. Gemini provided more consistent explanations than the other models on basic science questions and on clinical periodontology questions alike, and it maintained that consistency across standardized and information-based question formats. Each of these differences was statistically significant. This suggests that Gemini&#8217;s advantage was not confined to one corner of the syllabus but reflected a more general capacity to produce coherent, pedagogically sound rationales across the breadth of periodontology content tested.</p>
<p>To move beyond simple group comparisons, the study employed regression analyses, statistical techniques that estimate how much each predictor variable contributes to an outcome while accounting for the others. The results were telling: question type and topic were both significantly related to explanation quality. Accuracy, by contrast, appeared comparatively robust to these factors. The models generally knew the right answer regardless of whether a question tested recall or clinical application, or whether it concerned the microbiology of periodontal pockets or the surgical management of mucogingival defects. But the clarity, depth, and consistency of the reasoning they offered in support of those answers fluctuated considerably depending on what was being asked and how.</p>
<p>This dissociation between accuracy and explanation quality carries real implications for how artificial intelligence should be integrated into health professional education. Multiple-choice examinations, including the specialty exams used to license and certify clinicians, reward the selection of a correct option. A student using a chatbot as a study aid, however, typically learns from the explanation, not the letter of the answer. If a model&#8217;s rationale is thin, inconsistent, or subtly wrong even when its final choice is right, the student may internalize flawed mental models that surface later at the chairside. Conversely, an explanation of high quality can transform a simple quiz question into a miniature tutorial, connecting the answer to underlying mechanisms, differential diagnoses, and treatment principles.</p>
<p>The findings also speak to the technical character of large language models. These systems generate answers by predicting likely text sequences based on patterns learned during training, rather than by consulting a verified knowledge base. Their factual accuracy on well-represented exam content can be excellent, because such content appears abundantly in textbooks, review articles, and question banks. Explanation quality, however, depends on the model&#8217;s ability to organize and verbalize reasoning in a way that is both correct and appropriately calibrated to the question&#8217;s cognitive demand — a harder task that varies more with the structure of the prompt. Case-based questions, which require integrating scattered clinical clues, appear to stress this capacity in ways that simple factual recall does not, and the regression results indicate that this stress shows up in explanation quality even when the final answer survives intact.</p>
<p>Parlak&#8217;s conclusion is measured. Large language models, the study finds, are promising tools for periodontology education, but they retain limitations that require improvement before they can be relied upon uncritically. The author also cautions that the findings should be interpreted within the constraints of the study design: the questions came from a single national specialty examination, were posed in Turkish, and the evaluation of explanation quality, though systematic, ultimately rests on structured human judgment rather than an objective gold standard. Consumer chatbot interfaces also evolve rapidly, meaning that any snapshot comparison reflects a particular moment in a fast-moving technological race. Still, the central message is likely to resonate well beyond periodontology. As educators worldwide weigh whether to embrace, restrict, or redesign assessment in the age of generative artificial intelligence, this study suggests the right question is not simply whether chatbots get the answer right, but whether they can explain it in a way that actually teaches. On that measure, the models are not yet interchangeable — and the format of the question matters more than anyone might have guessed.</p>
<p><strong>Subject of Research:</strong> Comparative performance of large language models in answering periodontology multiple-choice questions</p>
<p><strong>Article Title:</strong> Do question type and topic affect the performance of large language models in answering periodontology questions? A comparative study</p>
<p><strong>Article References:</strong> Parlak, H. M. (2026). Do question type and topic affect the performance of large language models in answering periodontology questions? A comparative study. <em>BMC Medical Education</em>. <a href="https://doi.org/10.1186/s12909-026-10516-z" rel="noopener noreferrer">https://doi.org/10.1186/s12909-026-10516-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12909-026-10516-z" rel="noopener noreferrer">10.1186/s12909-026-10516-z</a></p>
<p><strong>Keywords:</strong> large language models, artificial intelligence, periodontology, dental education, ChatGPT-5, Gemini 2.5 Flash, Copilot, multiple-choice questions, medical education, explanation quality, question type, regression analysis</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224302</post-id>	</item>
	</channel>
</rss>
