Tuesday, October 6, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Medicine

AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test

October 6, 2026
in Medicine
Dean Parker
By Dean Parker Scienmag Editorial Profile - Dermatology
Reading Time: 5 mins read
0
AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test

AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

When a patient presents with an unusual rash, the dermatologist across the table is drawing on years of training, thousands of clinical images, and the hard-won experience of board certification. Now a team at Rutgers Robert Wood Johnson Medical School has asked a provocative question: how would the newest generation of artificial intelligence chatbots fare against the same examination standards? In a research letter published in the Archives of Dermatological Research, Rose Rasty, Elyse Mackenzie, Shaunt Mehdikhani and Babar Rao report their evaluation of five prominent large language models on board-style dermatology questions drawn from the American Board of Dermatology’s testing tradition, offering one of the most direct head-to-head comparisons yet attempted in the specialty.

The five systems examined represent the current cutting edge of consumer-facing artificial intelligence. The study tested ChatGPT in its GPT-5.2 configuration, xAI’s Grok 4.1, Google’s Gemini 3, Perplexity running Sonar and GPT-5-series backends, and Microsoft Copilot built on GPT-5. This lineup matters because these are not obscure research prototypes buried in a laboratory; they are the tools that medical students, residents and practicing physicians already consult daily for quick answers. Understanding how they perform on rigorous, specialty-specific questions is therefore not an academic curiosity but a matter of practical patient safety, since patients too are increasingly turning to these chatbots with their own dermatological concerns.

Board-style questions are a particularly demanding benchmark for any question-answering system. The American Board of Dermatology’s examinations are designed to distinguish competent specialists from merely knowledgeable ones, and the questions typically present nuanced clinical vignettes in which several answer choices are plausible. A candidate must weigh the patient’s age, lesion morphology, distribution, history of prior treatments, and sometimes subtle histopathological findings before selecting the single best response. This format tests something closer to clinical reasoning than to factual recall, which is precisely why it has become a popular yardstick for evaluating medical artificial intelligence. A model can memorize that psoriasis affects roughly two to three percent of the population, but it must reason through a vignette to recognize a psoriasiform drug eruption masquerading as the idiopathic disease.

The technical architecture underlying these models helps explain both their promise and their pitfalls. Large language models are neural networks, generally built on the transformer architecture, trained on vast text corpora to predict the next token in a sequence. Through this deceptively simple objective, they absorb grammar, factual knowledge, and patterns of reasoning encoded in their training data. Medical knowledge enters the picture because textbooks, review articles, clinical guidelines and discussion forums form a substantial portion of the public internet. However, the models do not store facts as a database would; knowledge is distributed across billions of numerical parameters, which means retrieval is probabilistic rather than deterministic. This is why a model can answer an obscure question flawlessly one moment and hallucinate a nonexistent drug interaction the next, a behavior that examiners would immediately fail in a human candidate.

Dermatology poses distinctive challenges for text-based models, even in a question-answer format that strips away the visual element. The specialty sits at the intersection of internal medicine, immunology, infectious disease, oncology and pathology, and its vocabulary is notoriously dense. Distinguishing lichen planus from lichenoid drug eruption, or dermatitis herpetiformis from linear IgA bullous dermatosis, depends on pattern recognition honed through exposure to thousands of cases. Board-style questions compress that pattern recognition into prose, which arguably favors language models, yet the compression also removes the contextual cues a clinician would use at the bedside. The Rutgers team’s choice of this format follows a growing research tradition: earlier studies had already tested ChatGPT on board-style dermatology questions, including image-based items, and Liu and colleagues in 2025 assessed dermatological knowledge and image analysis using specialty certificate examinations as their benchmark.

What makes the new study notable is its breadth of comparison. Most prior evaluations examined a single model, usually a ChatGPT variant, against a question bank. By running the same board-style items through five different systems, each with distinct underlying architectures and training pipelines, the researchers could probe whether strong medical performance is a general property of frontier language models or something specific to particular products. The differences among these systems are not trivial. GPT-5.2, Grok 4.1, Gemini 3 and the GPT-5-based Copilot each reflect different training data mixtures, different reinforcement learning strategies, and different approaches to reasoning. Perplexity adds a further wrinkle: it couples a language model with live web search, meaning its answers may draw on retrieved documents rather than purely on parametric memory, a hybrid architecture that could behave very differently on questions with recently updated guidelines.

The distinction between parametric knowledge and retrieval-augmented answering is one of the most technically interesting aspects of this kind of evaluation. A pure language model answers from what is baked into its weights during training, frozen at a cutoff date. A retrieval-augmented system like Perplexity queries the live web, which brings freshness but also vulnerability: it can absorb misinformation, outdated guidelines or commercially biased content from the open internet. For dermatology, where treatment recommendations evolve and where the internet is saturated with anecdotal skin-care advice, that distinction could materially change performance. A benchmark like the one used in this study therefore measures not just a model’s medical knowledge but its entire answer-generation pipeline, including how it filters and prioritizes whatever sources it consults.

The authors of the research letter are appropriately measured about what such testing can and cannot establish. Answering multiple-choice questions correctly is not the same as managing a patient. A board examination candidate who selects the right option for a bullous disorder vignette may still struggle to perform a safe skin biopsy or to counsel a frightened patient about a new diagnosis of melanoma. Conversely, models can exploit statistical cues in question wording, patterns that human test-takers learn to recognize as well, without genuinely understanding the underlying medicine. There is also the ever-present risk of data contamination: if board-style questions or close paraphrases circulate online, a model trained on the internet may have effectively seen the answer key. Rigorous benchmarking in medical artificial intelligence must therefore always ask not only how well a model scored, but whether the score reflects reasoning or recall.

Nevertheless, the trajectory across successive studies is striking. When ChatGPT first burst into public awareness in late 2022, early medical evaluations produced mixed results, with the model often falling short of passing thresholds on professional examinations. Each subsequent generation has closed the gap, and by the mid-2020s frontier models were routinely performing at or near the level of human test-takers across a range of specialties. The Rutgers study extends that line of evidence into dermatology with the newest available models, and its very framing, published as a research letter rather than a lengthy original article, reflects how quickly the field is moving: the findings needed to reach the literature before the models they evaluated were superseded. The authors report that all data were generated using the five named systems, with the dataset available from the corresponding author on reasonable request, and the work received no external funding while the authors declared no conflicts of interest.

For clinicians and patients alike, the practical message is one of cautious engagement rather than either alarm or celebration. These systems are already embedded in clinical workflows, from drafting patient messages to summarizing literature, and their demonstrated competence on board-style questions suggests they can serve as powerful adjuncts for education and decision support. A resident reviewing for boards might legitimately use a chatbot to generate practice vignettes or explain the immunopathology of pemphigus vulgaris. But the same probabilistic architecture that enables fluent expertise also produces confident errors, and no benchmark, however demanding, substitutes for clinical supervision. The Rutgers team’s comparison of five frontier models gives the field a clearer map of where artificial intelligence stands in dermatology today, and a reminder that the map will need redrawing with every new model release.

Subject of Research: Evaluation of large language models on American Board of Dermatology-style examination questions

Article Title: Performance of five large language models on board-style dermatology questions from the American Board of Dermatology

Article References: Rasty, R., Mackenzie, E., Mehdikhani, S., & Rao, B. (2026). Performance of five large language models on board-style dermatology questions from the American Board of Dermatology. Archives of Dermatological Research, 318(1), Article 462. https://doi.org/10.1007/s00403-026-04894-z

Image Credits: AI Generated

DOI: 10.1007/s00403-026-04894-z

Keywords: large language models, dermatology, artificial intelligence, board examination, ChatGPT, Gemini, Grok, Microsoft Copilot, Perplexity, medical education, clinical reasoning, benchmarking

Cite Scienmag News

Dean Parker. (October 6, 2026). AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test. Scienmag. https://scienmag.com/ai-sits-a-dermatology-board-exam-five-chatbots-put-to-the-test/

Dean Parker. "AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test." Scienmag, 6 October 2026, https://scienmag.com/ai-sits-a-dermatology-board-exam-five-chatbots-put-to-the-test/. Accessed 6 October 2026.

Dean Parker. "AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test." Scienmag. October 6, 2026. https://scienmag.com/ai-sits-a-dermatology-board-exam-five-chatbots-put-to-the-test/

Tags: AI dermatology board examAI diagnostic accuracy in skin conditionsAI performance on specialty-specific medical questionsAI tools for medical education and practiceArtificial Intelligencebenchmarkingboard examinationchatbot performance in medical licensingChatGPTclinical reasoningcomparison of ChatGPT and other AI modelsconsumer-facing AI in healthcaredermatologyevaluation of GPT-5.2 in dermatologyGeminiGrokimpact of AI on dermatology certification examslarge language modelslarge language models in dermatologyMedical EducationMicrosoft Copilotperplexityrole of AI chatbots in dermatology diagnosticsRutgers study on AI in medical licensing
Share26Tweet16
Previous Post

One in Five Dutch Nursing Home Staff Faced Burnout or Depression Throughout COVID-19 Pandemic

Next Post

Three-Year Trial Shows Mediterranean Diet Eases Food Addiction Symptoms but Only for Some

Related Posts

Chemical Trick Gives MRI a Clear Window on Brain Molecules It Could Never See
Medicine

Chemical Trick Gives MRI a Clear Window on Brain Molecules It Could Never See

October 6, 2026
Fingertip Blood Sampling Device Outperforms Rival in HIV Monitoring Trial in Zambia
Medicine

Fingertip Blood Sampling Device Outperforms Rival in HIV Monitoring Trial in Zambia

October 6, 2026
When Bypass Grafts Slow Down: New Study Separates Competitive Flow from True Blockage
Medicine

When Bypass Grafts Slow Down: New Study Separates Competitive Flow from True Blockage

October 6, 2026
Longer Community Exercise Programs Sharpen Balance, Strength and Endurance in Older Adults
Medicine

Longer Community Exercise Programs Sharpen Balance, Strength and Endurance in Older Adults

October 6, 2026
Lifting Weights May Rejuvenate the Aging Brain’s Metabolism and Deepen Sleep
Medicine

Lifting Weights May Rejuvenate the Aging Brain’s Metabolism and Deepen Sleep

October 6, 2026
Three-Year Trial Shows Mediterranean Diet Eases Food Addiction Symptoms but Only for Some
Medicine

Three-Year Trial Shows Mediterranean Diet Eases Food Addiction Symptoms but Only for Some

October 6, 2026
Next Post
Three-Year Trial Shows Mediterranean Diet Eases Food Addiction Symptoms but Only for Some

Three-Year Trial Shows Mediterranean Diet Eases Food Addiction Symptoms but Only for Some

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Chemical Trick Gives MRI a Clear Window on Brain Molecules It Could Never See
  • Smart Vibration System Lines Up Maize Ears Without Breaking Them
  • Fingertip Blood Sampling Device Outperforms Rival in HIV Monitoring Trial in Zambia
  • Landmark Trial Pinpoints Which HER2-Positive Breast Cancers Respond to Shorter, Gentler Chemotherapy

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading