For decades, the success of a facelift has been judged largely the way art critics judge a painting: by eye, by opinion, and by consensus that is often anything but consistent. A surgeon looks at before-and-after photographs, the patient looks in the mirror, and sometimes an expert panel weighs in. Each observer brings a different standard, and the result is a field where outcomes are described in adjectives rather than data. A new pilot study published in BMC Plastic and Reconstructive Surgery now suggests that artificial intelligence could change that, by generating reproducible numerical scores for facial rejuvenation surgery directly from standard clinical images.
The research, led by an international team of surgeons and researchers from institutions including the University Hospital Regensburg, Charité–Universitätsmedizin Berlin, and Cedars-Sinai Medical Center, tested a fully automated pipeline for evaluating what surgeons call combined facial aesthetic surgery, or CFAS. This term describes operations in which two or more facial rejuvenation procedures are performed in a single operative session, for example a deep-plane facelift performed together with eyelid surgery and a temporal lift. Such combined procedures are common in elective facial plastic surgery, yet their results have traditionally been measured with subjective tools that suffer from high variability between different raters.
At the heart of the study is a machine learning algorithm known as CAARISMA ARMM, the Aesthetic Research and Rating Metrics Model developed by ICA Aesthetic Navigation in Frankfurt am Main, Germany. The researchers paired this algorithm with the Vectra H2 system from Canfield Scientific, a widely used clinical camera platform that captures a three-dimensional stereophotogrammetric reconstruction of the face. From that 3D capture, a standardized two-dimensional frontal projection is rendered and fed into the algorithm, which analyzes 1,080 variables across 17 facial landmarks spanning the periorbital, perioral, malar, jawline, and cervical regions. The specific landmark coordinates and feature weightings remain proprietary to the developer, but the system has previously been applied to objectify outcomes in face transplantation and facial palsy reanimation surgery.
The output consists of three scores, each ranging from 0 to 100: a Facial Youthfulness Index (FYI), a Facial Aesthetic Index (FAI), and a Skin Quality Index (SQI). To test the pipeline, the team retrospectively analyzed ten female patients treated at a single private aesthetic surgery practice in Paris, France. All patients underwent deep-plane facelift and facial lipofilling, with additional procedures layered on in various combinations: upper blepharoplasty in seven patients, lower blepharoplasty in six, temporal lift in five, and lip lift in four. The mean age of the cohort was 73.2 years. Standardized frontal images taken before surgery and three months after surgery were processed through the automated system, with no manual scoring step involved.
The results were statistically significant across all three indices. The Facial Youthfulness Index rose from a mean of 40.8 to 42.1, a relative gain of 2.1 percent with a p-value of 0.016. The Facial Aesthetic Index climbed from 73.5 to 83.3, a relative improvement of 12.3 percent with a p-value of 0.008. The Skin Quality Index showed the largest and most consistent change, rising from 58.5 to 68.4, a relative gain of 14.3 percent with a p-value of 0.001. Among the secondary, component-level measures, the biggest relative improvements appeared in skin texture parameters: fine relief improved by 56.2 percent, roughness by 15.6 percent, and rough relief by 13.0 percent, all with p-values of 0.006. Wrinkle scores improved most in the crow’s feet region, with a relative gain of 10.4 percent, followed by the infraorbital region at 6.6 percent.
One striking technical feature of the pipeline is its determinism. Because the algorithm produces the same output for the same input, two investigators who independently re-processed the identical standardized images achieved perfect agreement. The authors are careful to note that this reflects computational reproducibility rather than conventional interrater reliability, since no subjective judgment was involved at any point. In other words, the system eliminates one of the most stubborn problems in aesthetic surgery research, the disagreement between human raters, but it does not by itself prove that the numbers it produces measure anything a human would recognize as improvement.
That caveat runs throughout the paper, and the authors are unusually candid about it. The study is explicitly labeled a pilot feasibility study, and the team stresses that the numerical gains should not be interpreted as evidence of clinically perceptible improvement. The minimal clinically important difference for each index has not been established, meaning no one yet knows how large a score change must be before a patient or a blinded evaluator would actually notice a difference. The three-month follow-up interval is also short relative to the biology of deep-plane facelift surgery, since residual swelling, tissue induration, and scar maturation can still be resolving at twelve weeks. A later assessment at nine to twelve months would better capture consolidated outcomes.
The cohort itself was small and strikingly homogeneous: ten Caucasian French women from a single practice, retrospectively selected. That homogeneity limits generalizability and raises a deeper problem that the authors confront directly. Aesthetic ideals vary across cultures, and an algorithm trained on particular standards may not translate to other populations. The field of medicine has already seen how algorithmic bias can produce disparities in care, and a scoring system for facial beauty is arguably even more sensitive to cultural context than a diagnostic tool. The authors also note that the algorithm is proprietary and commercially developed, and that one of the co-authors is the chief executive of the company behind it, which makes independent replication by investigators without financial ties an important next step.
Perhaps the most instructive finding is the exception that proves the rule. One patient in the series showed a postoperative decrease in her Facial Aesthetic Index of 2.2 percent, moving against the group trend. She had undergone the same combination of procedures as another patient who showed one of the largest gains, and her age fell within the range of the cohort, so no obvious factor distinguished her. The authors suggest the discordance may reflect the algorithm’s sensitivity to subtle differences in facial expression, head positioning, or lighting between the pre- and postoperative captures rather than any true clinical effect. Their conclusion is pointed: AI-derived attractiveness scores should not be treated as stand-alone measures of aesthetic success or failure, nor as medicolegal endpoints, but as adjunctive metrics considered alongside clinical judgment and patient-reported outcomes.
An intriguing puzzle remains in why the Skin Quality Index and its subcomponents improved at all, given that no dedicated skin-directed procedures such as laser resurfacing or chemical peels were performed. The authors offer plausible but unconfirmed mechanisms: the mechanical tension and redraping of the skin envelope during a deep-plane facelift may transiently smooth surface microtexture, the resolution of postoperative swelling and improved local perfusion may alter skin tone and light reflectance, and facial lipofilling may change subdermal support and the optical properties of the overlying skin. Part of the change may also simply reflect measurement variability related to positioning, lighting, or residual swelling. Distinguishing among these possibilities will require dedicated mechanistic studies.
What the study does establish is technical feasibility. An automated AI algorithm can be integrated with a standard clinical camera system to generate quantitative, reproducible scores after combined facial aesthetic surgery, with minimal manual input and no observer bias. The authors frame the work as a step toward greater consistency and transparency in outcome assessment rather than a means of fully standardizing beauty, a quality they acknowledge is shaped by individual, cultural, and contextual factors that no single metric can resolve. Future validation should correlate changes in the three indices with validated patient-reported instruments such as FACE-Q and with blinded surgeon panel ratings, determine minimum clinically important differences for each score, stratify analyses by procedure to isolate the contribution of individual operations, and extend follow-up to nine to twelve months. Until then, the authors recommend treating CAARISMA ARMM as an investigational research tool rather than a validated clinical instrument. The pilot dataset, they write, should be regarded as motivation for that validation program rather than its completion. It is a measured conclusion for a field where the temptation to declare a revolution is never far away, and it may prove to be the study’s most valuable contribution.
Subject of Research: Automated AI-based quantitative assessment of aesthetic outcomes after combined facial aesthetic surgery
Article Title: Automated assessment of aesthetic outcomes following combined facial aesthetic surgery – a pilot study
Article References: Sherwani, K., Knoedler, L., Schaschinger, T., Niederegger, T., Fenske, J., Sadati, K., Heiland, M., Pooth, R., Brown, T., Cetrulo, C. L., Jr., Aguglia, R., & Lellouch, A. G. (2026). Automated assessment of aesthetic outcomes following combined facial aesthetic surgery – a pilot study. BMC Plastic and Reconstructive Surgery, 2(1), Article 27. https://doi.org/10.1186/s44452-026-00040-w
Image Credits: AI Generated
DOI: 10.1186/s44452-026-00040-w
Keywords: artificial intelligence, facial aesthetic surgery, facelift, CAARISMA ARMM, Vectra imaging, outcome assessment, plastic surgery, machine learning, skin quality, facial rejuvenation, pilot study, quantitative metrics
Cite Scienmag News
Ophelia Keating. (September 30, 2026). AI Scoring System Puts Numbers on Facelift Results in Landmark Pilot Study. Scienmag. https://scienmag.com/ai-scoring-system-puts-numbers-on-facelift-results-in-landmark-pilot-study/
Ophelia Keating. "AI Scoring System Puts Numbers on Facelift Results in Landmark Pilot Study." Scienmag, 30 September 2026, https://scienmag.com/ai-scoring-system-puts-numbers-on-facelift-results-in-landmark-pilot-study/. Accessed 30 September 2026.
Ophelia Keating. "AI Scoring System Puts Numbers on Facelift Results in Landmark Pilot Study." Scienmag. September 30, 2026. https://scienmag.com/ai-scoring-system-puts-numbers-on-facelift-results-in-landmark-pilot-study/

