Artificial intelligence has already written essays, passed bar exams, and drafted computer code, but can it plan the reconstruction of a patient’s broken-down dentition? A new study from Prince Sattam bin Abdulaziz University in Saudi Arabia set out to answer that question with unusual rigor, pitting three of the world’s most prominent large language models against one another in a head-to-head test of fixed prosthodontic treatment planning. The results, published in BMC Medical Education, are a fascinating mix of promise and caution: one model edged ahead on completeness and adherence to clinical references, another shone brightest on patient safety, and yet the overall verdict was far less decisive than the raw numbers might suggest.
The research team, led by Heba Wageh Abozaed and colleagues, constructed thirty standardized clinical scenarios representing the kinds of cases a prosthodontist encounters in daily practice. Each scenario was fed to three AI systems: Qwen 3.7 Plus, developed by Alibaba; Claude Sonnet 5, from Anthropic; and Mistral Medium 3.5, the European contender from Mistral AI. Every model generated a treatment plan for every scenario, producing ninety plans in total. This standardized design matters enormously. Earlier studies of AI in dentistry often used ad hoc prompts and single-model assessments, making comparisons nearly impossible. By holding the clinical material constant and varying only the model, the researchers isolated exactly what they wanted to measure: the quality of machine-generated clinical reasoning.
Judging the plans fell to two blinded prosthodontic evaluators who worked independently, neither knowing which model had produced which plan. Each plan was scored on a seven-domain rubric using a five-point Likert scale, covering diagnostic accuracy, treatment selection, treatment sequencing, completeness, adherence to references, patient safety, and overall clinical appropriateness. That structure is worth pausing on, because treatment planning in fixed prosthodontics is not a single decision but a chain of them. A clinician must first interpret the diagnostic information correctly, then choose among crowns, fixed partial dentures, implants, or other restorative pathways, then order those interventions sensibly, all while protecting the patient from harm and grounding every choice in accepted evidence. A model that writes fluent prose but skips a critical step, or recommends an irreversible procedure without adequate justification, would lose points precisely where it matters most.
The statistical machinery behind the comparison was appropriately serious for non-normal, paired ordinal data. The Friedman test, a non-parametric equivalent of repeated-measures analysis of variance, assessed whether the three models differed overall, and it did so emphatically: a chi-square statistic of 35.05 with two degrees of freedom, yielding a p-value below 0.001, and a Kendall’s W of 0.584 indicating a substantial consistency in how the models were ordered across the thirty scenarios. Pairwise contrasts followed using Wilcoxon signed-rank tests with Holm adjustment, a correction that controls the inflated risk of false positives when multiple comparisons are run on the same data. Every pairwise difference survived that correction, meaning the gaps between the models were statistically robust rather than artifacts of chance.
On the raw numbers, Qwen 3.7 Plus took the crown. Its mean total score of 33.82, with a standard deviation of 1.74, topped Claude Sonnet 5’s 32.52 (standard deviation 2.30) and left Mistral Medium 3.5 trailing at 29.68 (standard deviation 2.66). Because the rubric spans seven domains on a five-point scale, the theoretical maximum total is 35, which makes Qwen’s average strikingly close to perfection and highlights just how compressed the top of the scale had become. Within the individual domains, the picture grew more nuanced. Qwen scored highest on completeness and on adherence to references, suggesting it was the most thorough and the most evidence-anchored of the three. Claude, however, achieved the highest patient safety scores, a domain that many clinicians would argue outweighs almost any other. Mistral lagged across all seven domains, a uniform rather than selective weakness.
Then comes the twist that transforms this from a straightforward leaderboard into a genuinely important methodological caution. The two evaluators did not agree with each other very well. The single-measure intraclass correlation coefficient came in at just 0.378, well below conventional thresholds for acceptable reliability, and even the average-measure ICC, which reflects the reliability of the mean of both raters, reached only 0.549. Quadratic-weighted Cohen’s kappa told a similarly sobering story. In plain terms, the two expert judges, both blinded and both using the same rubric, saw the same ninety plans and scored them in ways that diverged substantially.
A sensitivity analysis made the consequences explicit. When each evaluator’s scores were analyzed separately, the model ranking flipped. Evaluator 1 placed Claude marginally above Qwen, while Evaluator 2 ranked Qwen clearly ahead of Claude. The researchers traced this discrepancy to a pronounced ceiling effect in Evaluator 1’s ratings: that evaluator assigned the maximum score to the majority of Qwen and Claude cases, compressing the scale so severely that meaningful differences between the two leading models simply could not register. Ceiling effects are a well-known hazard in performance assessment, and here they were strong enough to overturn the headline finding of the primary analysis. The authors are admirably direct about the implication: these findings do not support universal superiority of any single model.
What, then, can responsibly be concluded? Mistral’s last-place finish across every domain appears to be the most stable result, consistent under both evaluators and across the sensitivity checks. Qwen’s strengths in completeness and reference adherence, and Claude’s edge in patient safety, are real observations from the primary analysis, but the fragility of the inter-rater agreement means any claim that one model is definitively better than the other should be treated as provisional. The study also stops short of testing whether AI-assisted planning actually improves outcomes for patients or learning outcomes for students; educational effectiveness was not directly assessed. What the study does establish is that the leading models can produce treatment plans that experts, on average, judge to be clinically appropriate across a demanding seven-domain rubric, and that the differences among them, while statistically detectable, are modest and measurement-dependent.
The educational implications may prove to be the most durable legacy of this work. The authors suggest that large language models could serve as adjunctive tools for case-based learning and for critical appraisal of treatment-planning decisions, and that framing fits the evidence well. A dental student who generates a plan with an AI system and then interrogates it, asking why a particular sequence was chosen, which alternatives were rejected, and whether the safety considerations were adequately addressed, is engaging in exactly the reflective reasoning that prosthodontic education tries to cultivate. The models become sparring partners rather than oracles. That use case also sidesteps the most dangerous failure mode, which would be a clinician or trainee accepting a machine-generated plan uncritically, particularly in a field where errors can mean irreversible removal of tooth structure.
For the broader conversation about AI in medicine, the study offers a template worth copying. Standardized scenarios, blinded expert evaluation, a multi-domain rubric, non-parametric statistics with multiplicity correction, and, crucially, a sensitivity analysis that exposed how much the conclusions depended on the raters themselves. The uncomfortable but valuable lesson is that evaluating AI is itself a measurement problem, and a rubric that allows one expert to hand out perfect scores to most cases will tell you more about the rubric than about the machine. As these models continue to improve at astonishing speed, the bottleneck in studies like this one may shift from the capability of the AI to the reliability of the humans judging it. Fixing that, perhaps through larger panels of evaluators, refined anchor descriptions for each score point, or harder scenarios that push plans away from the ceiling, will be essential before anyone can say with confidence which artificial intelligence plans a crown best. For now, the honest answer is that the top contenders are close, the measurement tools are shaky, and the most sensible role for AI in the prosthodontic workflow today is as a well-informed colleague whose plans are always worth questioning.
Subject of Research: Comparative evaluation of large language models for fixed prosthodontic treatment planning
Article Title: Comparative evaluation of three large language models for fixed prosthodontic treatment planning: a standardized scenario-based study
Article References: Abozaed, H. W., Alshenaiber, R., Albader, S., Algohar, A., Murayshed, M. S., & Alokla, M. (2026). Comparative evaluation of three large language models for fixed prosthodontic treatment planning: a standardized scenario-based study. BMC Medical Education. https://doi.org/10.1186/s12909-026-10560-9
Image Credits: AI Generated
DOI: 10.1186/s12909-026-10560-9
Keywords: large language models, artificial intelligence, prosthodontics, treatment planning, dental education, Qwen, Claude, Mistral, clinical decision-making, inter-rater reliability, BMC Medical Education, AI in dentistry
Cite Scienmag News
Courtney Benton. (October 9, 2026). AI Dentists Face Off: Three Chatbots Tested on Real Treatment Plans. Scienmag. https://scienmag.com/ai-dentists-face-off-three-chatbots-tested-on-real-treatment-plans/
Courtney Benton. "AI Dentists Face Off: Three Chatbots Tested on Real Treatment Plans." Scienmag, 9 October 2026, https://scienmag.com/ai-dentists-face-off-three-chatbots-tested-on-real-treatment-plans/. Accessed 9 October 2026.
Courtney Benton. "AI Dentists Face Off: Three Chatbots Tested on Real Treatment Plans." Scienmag. October 9, 2026. https://scienmag.com/ai-dentists-face-off-three-chatbots-tested-on-real-treatment-plans/

