<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI in prosthodontics board exam &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-in-prosthodontics-board-exam/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 00:55:57 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI in prosthodontics board exam &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Takes the Prosthodontics Board Exam: Chatbots Flirt With Resident-Level Scores</title>
		<link>https://scienmag.com/ai-takes-the-prosthodontics-board-exam-chatbots-flirt-with-resident-level-scores/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 00:55:57 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI accuracy and variability in specialized knowledge]]></category>
		<category><![CDATA[AI and human comparison in dental exams]]></category>
		<category><![CDATA[AI in prosthodontics board exam]]></category>
		<category><![CDATA[American College of Prosthodontists]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[artificial intelligence in dental specialization assessments]]></category>
		<category><![CDATA[chatbot performance in dental licensing tests]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[ChatGPT-3.5 and 4 in prosthodontics]]></category>
		<category><![CDATA[clinical knowledge testing with chatbots]]></category>
		<category><![CDATA[dental education]]></category>
		<category><![CDATA[evaluation of AI capabilities in healthcare]]></category>
		<category><![CDATA[Google Gemini]]></category>
		<category><![CDATA[hallucination]]></category>
		<category><![CDATA[healthcare education]]></category>
		<category><![CDATA[impact of AI on dental education and licensing]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in medical exams]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning models in medical certification]]></category>
		<category><![CDATA[machine reasoning in clinical dentistry]]></category>
		<category><![CDATA[medical licensing exams]]></category>
		<category><![CDATA[prosthodontics]]></category>
		<category><![CDATA[test-retest reliability]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=229883</guid>

					<description><![CDATA[A new study tested ChatGPT-3.5, ChatGPT-4, Bing, and Google Gemini on 300 questions from the 2023 and 2024 National Prosthodontic Resident Examination, finding that top models approached resident-level accuracy but showed significant inconsistency between testing sessions.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has already passed bar exams, medical licensing tests, and chess grandmasters, but a new study has now put four of the world&#8217;s most prominent chatbots through one of dentistry&#8217;s most demanding knowledge assessments: the National Prosthodontic Resident Examination, the annual test taken by dentists training to become specialists in rebuilding and replacing teeth. The results, published in the journal Heliyon, are both striking and sobering. The best-performing models scored within striking distance of the average human resident, yet their answers shifted between test sessions in ways that reveal how fragile machine reasoning remains when confronted with genuinely specialized clinical knowledge.</p>
<p>The research team, led by Marwa Shembesh and Cortino Sukotjo and spanning institutions in the United States and abroad, assembled 300 multiple-choice questions drawn from the 2023 and 2024 examinations administered by the American College of Prosthodontists. Each year&#8217;s exam contained 150 questions, written by prosthodontics program directors and fellows of the college, and the correct answers were verified against the official answer key. Four large language models were put to the test: ChatGPT-3.5 and ChatGPT-4 from OpenAI, Microsoft&#8217;s Bing Copilot, and Google&#8217;s Gemini. The investigators created fresh accounts for each model to eliminate any influence of stored conversation history, used standardized zero-shot prompts instructing each system to pick the single best answer, and set the sampling temperature to zero for API-based queries to favor deterministic output.</p>
<p>The examination itself is structured around a three-tier hierarchy of topics that mirrors the written section of the American Board of Prosthodontics certification. Tier 1 covers the core of the specialty: complete dentures, fixed and removable partial dentures, dental biomaterials, occlusion, implant prosthodontics, and esthetics. Tier 2 spans adjacent disciplines such as periodontics, pharmacology, wound healing, and temporomandibular disorders, while Tier 3 collects the broader periphery, from biostatistics and diagnostic imaging to sleep disorders, oral pathology, and maxillofacial prosthodontics. In 2023, more than 61 percent of questions fell into Tier 1, with implant prosthodontics and removable partial dentures the most heavily represented topics. In 2024, dental materials and fixed prosthodontics dominated, and Tier 1 still accounted for nearly 57 percent of the exam.</p>
<p>The headline numbers are remarkable. In the 2023 examination, ChatGPT-4 answered 72.7 percent of questions correctly on the first testing round, edging out Gemini at 67.3 percent, Bing at 62.7 percent, and ChatGPT-3.5 at 62.0 percent. The average human resident that year scored 64 percent. In other words, the most advanced chatbot of its generation outperformed the typical prosthodontics resident on the same questions, at least descriptively. A year later the picture shifted: Gemini led the first round with 66.7 percent, followed by ChatGPT-4 at 64.0 percent, Bing at 60.0 percent, and ChatGPT-3.5 at 44.7 percent, against a resident average of 61 percent. Because the human and machine scores were not compared with paired statistical tests, the authors caution that these comparisons are descriptive rather than proof of superiority.</p>
<p>Statistical analysis told a more nuanced story. In 2023, the differences among the four models never reached statistical significance at either testing point. In 2024, however, the overall difference was significant at the first time point, with ChatGPT-4, Bing, and Gemini each significantly outperforming ChatGPT-3.5 after Bonferroni correction; Gemini&#8217;s advantage over the older model amounted to 22 percentage points. By the second round two weeks later, the omnibus difference was still significant, but no individual pairwise comparison survived correction, underscoring how unstable model rankings can be across sessions and examination years.</p>
<p>Perhaps the most unsettling finding concerns consistency. When the researchers repeated the entire examination two weeks after the baseline session, three of the four models held steady in 2023, with observed agreement between sessions ranging from 78.7 to 85.3 percent and Cohen&#8217;s kappa values between 0.518 and 0.647. But in 2024, ChatGPT-4&#8217;s accuracy dropped from 64.0 percent to 51.3 percent, a statistically significant decline of 12.7 percentage points. The authors note that unobserved changes on the provider side, such as backend model updates between testing sessions, cannot be excluded as a source of this variability, since exact platform-level technical details were not retained during data collection. For educators, the message is clear: the same chatbot asked the same question two weeks apart may give a different answer, even under tightly controlled conditions.</p>
<p>The topic-tier analysis added another layer of insight. Descriptively, all four models found Tier 3 questions, the broad peripheral subjects, easiest, pooling to 74.2 percent accuracy in 2023 and 69.9 percent in 2024, while Tier 1, the heart of prosthodontics, proved hardest at 61.1 percent and 53.1 percent respectively. After accounting for repeated responses to the same questions using generalized estimating equations, the tier effect was statistically significant only in 2024, when Tier 3 questions carried roughly twice the odds of a correct answer compared with Tier 1. Within the core Tier 1 material, ChatGPT-4 and Gemini both significantly outperformed ChatGPT-3.5, with odds ratios of 1.80 and 1.92 respectively, while Bing&#8217;s advantage did not survive statistical adjustment.</p>
<p>Counterintuitively, newer was not always better. In the 2023 exam, the older ChatGPT-3.5 actually beat ChatGPT-4 on Tier 2 questions at baseline, 77.8 percent to 70.4 percent, and on Tier 3 as well. In 2024, ChatGPT-3.5&#8217;s second-round total accuracy of 52.7 percent narrowly exceeded ChatGPT-4&#8217;s 51.3 percent. The authors point to similar reversals reported elsewhere, including cases where GPT-3.5 outperformed GPT-4 on classification tasks because its simpler pattern recognition avoided overgeneralization. Model performance, they conclude, depends on the interplay of architecture, training data, and question type rather than a simple hierarchy of model generations.</p>
<p>The study is the first to evaluate large language models on the National Prosthodontic Resident Examination, filling a gap in a literature that has mostly focused on medicine, periodontology, and endodontics. Its implications reach beyond dentistry. The findings suggest that while chatbots can serve as supplementary study aids for residents reviewing implant dentistry or dental materials, they cannot yet be trusted as authoritative sources, particularly on the specialized core knowledge that defines a specialty. The authors also flag concerns about examination integrity if candidates gain access to such tools during secure testing, and they emphasize that model-generated content must always be verified against reliable sources because these systems can fabricate plausible-sounding but incorrect information, a phenomenon known as hallucination.</p>
<p>The researchers acknowledge limitations that temper the conclusions. The question bank is proprietary and accessible only to members of the American College of Prosthodontists, the dataset consisted solely of multiple-choice items that may not capture the complexity of real clinical decision-making, and the two-week testing window cannot predict how performance might drift with future model updates. Future work, they suggest, should test open-ended and case-based questions that better mirror chairside judgment, extend the evaluation across longer periods and more specialties, and probe how training strategies shape performance across different question types. For now, the study stands as a vivid snapshot of a fast-moving frontier: machines that can nearly match the specialists of tomorrow on their own exam, yet still wobble from one week to the next, and that therefore belong beside the textbook, not in place of the clinician.</p>
<p><strong>Subject of Research:</strong> Performance of large language models on the American College of Prosthodontists national resident examination</p>
<p><strong>Article Title:</strong> The performance of different large language models in the national prosthodontic resident examination in 2023 and 2024</p>
<p><strong>Article References:</strong> Shembesh, M., Koseoglu, M., Fang, Q., Gheisarifar, M., Kattadiyil, M. T., Yuan, J. C.-C., Barao, V. A., &amp; Sukotjo, C. (2026). The performance of different large language models in the national prosthodontic resident examination in 2023 and 2024. <em>Heliyon, 12</em>(15), Article e45465. <a href="https://doi.org/10.1016/j.heliyon.2026.e45465" rel="noopener noreferrer">https://doi.org/10.1016/j.heliyon.2026.e45465</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.heliyon.2026.e45465" rel="noopener noreferrer">10.1016/j.heliyon.2026.e45465</a></p>
<p><strong>Keywords:</strong> artificial intelligence, large language models, ChatGPT, Google Gemini, prosthodontics, dental education, medical licensing exams, test-retest reliability, hallucination, American College of Prosthodontists, machine learning, healthcare education</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">229883</post-id>	</item>
	</channel>
</rss>
