Every dental student knows the ritual: a block of white soap, a scalpel, and hours of careful scraping to transform an unremarkable rectangle into a convincing replica of a maxillary left permanent central incisor. The exercise, a staple of preclinical dental morphology education, is designed to train the eye and the hand before students ever touch a real tooth. But grading those carvings has always been a labor-intensive, imperfect process. Human examiners, however experienced, disagree with one another, and the sheer volume of specimens to assess places a heavy burden on faculty time. A new study from researchers at Biruni University in Istanbul, published in BMC Medical Education, asks whether a deep learning system can shoulder some of that load, and its answer is a carefully qualified yes.
The research team, led by prosthodontists Goknur Ozturk and Ozlem Kara together with computer engineering colleagues, built and internally validated an artificial intelligence pipeline that scores soap carvings criterion by criterion, mimicking the way educators actually evaluate student work. Rather than issuing a single holistic grade, the system was trained to assess ten distinct morphological features of the carved tooth, each rated on a scale from zero to ten. That criterion-specific design matters, because the pedagogical value of assessment lies not in a number but in the feedback it generates: a student needs to know whether the incisal edge is misshapen or the cervical line is misplaced, not merely that the carving earned a six.
To create the dataset, the team collected 280 soap-carved replicas of the FDI tooth 21, the maxillary left permanent central incisor, and photographed each specimen from five standardized views, yielding 1,400 images. Three educators then independently scored every carving on all ten criteria. Before any machine learning took place, the researchers quantified how consistently the three human raters judged the work. Using intraclass correlation coefficients, they found average-measure reliability ranging from 0.849 to 0.994 across the ten criteria, figures that indicate strong consensus when multiple expert judgments are pooled. The mean of the three educators’ scores for each criterion became the reference standard against which the artificial intelligence would be measured, an approach that smooths out individual idiosyncrasies and approximates the collective judgment of an examination board.
The machine learning architecture at the heart of the study is ResNet50, a convolutional neural network originally developed for image recognition and now widely repurposed across medicine. The researchers adapted it for multi-output regression, meaning a single model simultaneously predicts ten numeric scores, one per morphological criterion, from the set of five photographs of each carving. Training employed two-stage transfer learning, a technique in which the network first learns general visual features from a large external image dataset and is then fine-tuned on the soap carving images themselves. To make the most of a modest dataset and guard against the accident of a favorable or unfavorable data split, the team used five-fold cross-validation repeated across three random seeds, training ensembles of models whose predictions were aggregated for the final assessment.
The specimens were divided into a development set of 224 carvings used for training and an internal hold-out test set of 56 carvings that the models never saw during learning. Performance on that untouched test set was evaluated with a battery of statistical measures, each reported with bootstrap-derived 95 percent confidence intervals. The ensemble achieved a macro-averaged mean absolute error of 1.204 points on the ten-point scale, with a confidence interval running from 1.059 to 1.374, and a macro-averaged root mean squared error of 1.584. In practical terms, the system’s typical deviation from the averaged human reference was roughly one and a quarter points per criterion, a level of accuracy that could meaningfully flag a carving as strong or weak even if it cannot yet replicate a fine-grained examiner judgment.
Beneath those averages, however, lay substantial variation from criterion to criterion, and the authors are candid about it. The best-predicted feature, criterion 10, showed a mean absolute error of just 0.709, suggesting the network could track educator consensus closely on that aspect of tooth morphology. The worst, criterion 9, had a mean absolute error of 2.029, more than double. Spearman rank correlations between machine predictions and human reference scores ranged from 0.297 to 0.560 across criteria, while absolute-agreement intraclass correlation coefficients between the artificial intelligence and the multi-rater reference spanned only 0.309 to 0.429. The macro-averaged coefficient of determination, or R-squared, came in at 0.223, indicating that the model captured a modest fraction of the variance in human scores. Bland-Altman analysis, the standard tool for assessing agreement between two measurement methods, revealed small mean biases but limits of agreement wide enough to rule out treating the machine as a drop-in replacement for a human grader.
Why does performance vary so much across criteria? The study does not offer a definitive answer, but the pattern is consistent with what is known about both human and machine vision. Some morphological features, such as overall proportions or the presence of well-defined developmental grooves, produce strong visual signals that a convolutional network can readily detect. Others may depend on subtle contours, fine surface texture, or shading cues that are difficult to capture even in standardized photographs, and where even the three human raters may have found it harder to converge. The researchers also examined Gradient-weighted Class Activation Mapping, or Grad-CAM, on eight hold-out specimens to visualize which image regions drove the model’s predictions. The resulting heat maps showed heterogeneous activation patterns, meaning the network did not consistently focus on the same anatomical zones across specimens, a qualitative hint that its internal reasoning had not fully locked onto the features educators consider decisive.
The authors are explicit that these results do not support autonomous summative grading, the high-stakes use case in which an algorithm would assign official examination marks. Interchangeability with educator assessment, they conclude, is not demonstrated by the agreement statistics. What the findings do support is a more modest and arguably more realistic role: human-supervised formative assessment. In that scenario, the system could provide students with rapid, criterion-specific feedback during practice sessions, long before an instructor reviews the work, freeing faculty time for teaching rather than tallying and giving learners a faster feedback loop precisely when they are still developing their skills. The educator remains in the loop as the arbiter of record, with the machine serving as a screening and feedback instrument.
The study’s limitations are as instructive as its results. It was a single-center investigation using a single tooth type, the maxillary left central incisor, and a single assessment medium, standardized photographs rather than the physical specimens an examiner can turn in hand and inspect under varied light. Internal validation on a hold-out set, however rigorously conducted with cross-validation and multiple seeds, cannot substitute for external validation on data from other institutions, other raters, and other carving exercises. The authors state plainly that prospective external validation is required before any implementation in a live curriculum. That caution is a welcome corrective to the breathless tone that often surrounds educational artificial intelligence, and it reflects a maturing research culture in which feasibility studies are expected to define their own boundaries.
Still, the trajectory is notable. Soap carving is only the most traditional of preclinical assessments; the same photographic, multi-rater, criterion-specific framework could in principle extend to wax-ups, digital tooth preparations, and other visually assessed psychomotor tasks across the health professions. The Biruni study, which received no external funding and was approved by the university’s non-interventional research ethics committee, provides an honest baseline: a deep learning system can approach the averaged judgment of expert educators on some criteria, falls short on others, and is not yet fit to grade alone. For students sweating over a bar of soap at two in the morning, the near future may hold an algorithmic teaching assistant that points out a flawed mamelon or an over-carved lingual fossa within seconds, while the final verdict on their craftsmanship stays firmly in human hands.
Subject of Research: Deep learning assessment of preclinical dental soap carvings
Article Title: Development and internal validation of a deep learning system for criterion-specific assessment of preclinical dental soap carvings: a multi-rater study
Article References: Ozturk, G., Kara, O., Purmut, H. E., Birdal, M. B., & Sezer, A. (2026). Development and internal validation of a deep learning system for criterion-specific assessment of preclinical dental soap carvings: a multi-rater study. BMC Medical Education. https://doi.org/10.1186/s12909-026-10554-7
Image Credits: AI Generated
DOI: 10.1186/s12909-026-10554-7
Keywords: artificial intelligence, deep learning, dental education, preclinical assessment, soap carving, automated scoring, ResNet50, inter-rater reliability, medical education, computer vision, formative assessment, Bland-Altman analysis
Cite Scienmag News
Blake Davidson. (October 8, 2026). AI Learns to Grade Dental Students’ Soap Carvings, But Not Yet Like a Human Examiner. Scienmag. https://scienmag.com/ai-learns-to-grade-dental-students-soap-carvings-but-not-yet-like-a-human-examiner/
Blake Davidson. "AI Learns to Grade Dental Students’ Soap Carvings, But Not Yet Like a Human Examiner." Scienmag, 8 October 2026, https://scienmag.com/ai-learns-to-grade-dental-students-soap-carvings-but-not-yet-like-a-human-examiner/. Accessed 8 October 2026.
Blake Davidson. "AI Learns to Grade Dental Students’ Soap Carvings, But Not Yet Like a Human Examiner." Scienmag. October 8, 2026. https://scienmag.com/ai-learns-to-grade-dental-students-soap-carvings-but-not-yet-like-a-human-examiner/

