<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>impact of AI on medical licensing standards &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/impact-of-ai-on-medical-licensing-standards/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 10 Oct 2026 02:16:10 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>impact of AI on medical licensing standards &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Passes the Doctor&#8217;s Exam: How GPT-4o and DeepSeek Conquered China&#8217;s Toughest Medical Test</title>
		<link>https://scienmag.com/ai-passes-the-doctors-exam-how-gpt-4o-and-deepseek-conquered-chinas-toughest-medical-test/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 10 Oct 2026 02:16:10 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI and medical education reform]]></category>
		<category><![CDATA[AI benchmarks]]></category>
		<category><![CDATA[AI passing medical licensing examinations]]></category>
		<category><![CDATA[AI performance in Chinese medical licensing tests]]></category>
		<category><![CDATA[AI-driven advancements in medical diagnostics]]></category>
		<category><![CDATA[Artificial intelligence in medical licensing exams]]></category>
		<category><![CDATA[China]]></category>
		<category><![CDATA[Chinese-developed AI for medical testing]]></category>
		<category><![CDATA[clinical reasoning]]></category>
		<category><![CDATA[DeepSeek-R1]]></category>
		<category><![CDATA[ethical considerations of AI in healthcare licensing]]></category>
		<category><![CDATA[evaluation of AI capabilities in medical knowledge]]></category>
		<category><![CDATA[GPT-4o]]></category>
		<category><![CDATA[GPT-4o and DeepSeek healthcare applications]]></category>
		<category><![CDATA[impact of AI on medical licensing standards]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in clinical practice]]></category>
		<category><![CDATA[machine learning in healthcare certification]]></category>
		<category><![CDATA[Medical Education]]></category>
		<category><![CDATA[medical licensing examination]]></category>
		<category><![CDATA[model evaluation]]></category>
		<category><![CDATA[multimodal AI]]></category>
		<category><![CDATA[PLOS Digital Health]]></category>
		<category><![CDATA[scoping review]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=256982</guid>

					<description><![CDATA[A scoping review finds that GPT-4o and DeepSeek-R1 have surpassed the passing threshold on China's national medical licensing examination, while warning that inconsistent evaluation methods still cloud the field.]]></description>
										<content:encoded><![CDATA[<p>For decades, the Chinese National Medical Licensing Examination has stood as one of the most demanding gateways into clinical practice anywhere in the world. Every year, hundreds of thousands of aspiring physicians sit through its grueling written and practical components, and only those who clear its threshold are permitted to treat patients. Now, a new class of examinees has quietly joined the ranks: artificial intelligence. According to a scoping review published in PLOS Digital Health, the latest large language models, including OpenAI&#8217;s GPT-4o and the Chinese-developed DeepSeek-R1, have not merely dabbled at the edges of this examination but have decisively passed it, scoring above the official cutoff and, in doing so, raising a question that would have sounded like science fiction only a few years ago: can a machine now match the baseline knowledge of a licensed physician?</p>
<p>The review, conducted by Hui Zong, Jiao Wang, and colleagues including senior author Bairong Shen, set out to map how well large language models perform on Chinese medical licensing examinations and related specialty tests. The team followed the PRISMA-ScR guidelines, the accepted standard for scoping reviews, and searched PubMed and Web of Science on June 12, 2025, using search terms that combined references to large language models with Chinese medical examinations. To be included, a study had to evaluate at least one large language model on a national-level Chinese medical examination and report quantitative performance metrics. Two reviewers independently screened the candidate papers and extracted the data, a double-check designed to reduce the risk that a single reader&#8217;s judgment would skew the synthesis.</p>
<p>What emerged from that screening was a body of evidence both rapidly growing and strikingly uneven. The final analysis included 14 studies encompassing 51 separate evaluation records spread across 8 different types of Chinese medical examinations. Beyond the core licensing exam, the models were tested on specialty assessments spanning critical care medicine, radiation oncology, and ultrasound medicine, domains that demand not just recall of textbook facts but the kind of layered clinical reasoning that experienced specialists apply under pressure. The examination years covered by the studies stretched from 2017 to 2024, and the review noted a clear shift over time toward using more recent exams, a sign that researchers are increasingly trying to avoid the contamination problem that haunts this field: the possibility that a model has already memorized the very questions it is being tested on, because those questions circulated in the internet-scale text used during training.</p>
<p>Nine large language models have been put through this Chinese examination gauntlet so far. The roster reads like a tour of the global and domestic AI landscape: OpenAI&#8217;s GPT-3.5, GPT-4, and GPT-4o; Baidu&#8217;s ERNIE; DeepSeek-R1; Alibaba&#8217;s Qwen-72B; the Baichuan2 models at 7 billion and 13 billion parameters; and DISC-MedLLM, a model specifically tuned for Chinese medical dialogue. This diversity matters. Chinese medical examinations are administered in Chinese, rich in domain-specific terminology, clinical vignettes, and culturally specific standards of care, so models trained primarily on English text face an additional hurdle. The inclusion of homegrown Chinese models alongside their American counterparts reflects a growing recognition that medical AI benchmarks cannot simply be transplanted from one language to another and assumed to mean the same thing.</p>
<p>The headline results are dramatic. On the 2024 Chinese National Medical Licensing Examination, GPT-4o achieved the highest score recorded in the reviewed literature: 552 points, equivalent to 92.00 percent, while DeepSeek-R1 followed closely with 523 points, or 87.20 percent. Both scores comfortably surpass the passing threshold for the examination. To put those numbers in perspective, the licensing exam is designed so that a substantial fraction of human candidates fail it in any given sitting; a model answering more than nine out of ten questions correctly is performing at a level that would place it among the stronger human test-takers. That a general-purpose chatbot, built without any specific intention of practicing medicine, can clear a national physician licensing barrier is a milestone that even optimistic researchers did not widely anticipate when GPT-3.5 first appeared.</p>
<p>The trajectory of improvement is as important as the endpoint. In an exploratory comparison drawn from the included studies, GPT-4 achieved a higher pass rate than GPT-3.5, and the difference was statistically significant when assessed with Fisher&#8217;s exact test, a statistical method suited to comparing proportions in small samples. This generational leap suggests that the same scaling forces driving progress in general AI, larger models, better training data, and refined reasoning techniques, are translating directly into medical competence. DeepSeek-R1&#8217;s strong showing is particularly noteworthy because that model was explicitly engineered to work through problems step by step, generating extended chains of reasoning before committing to an answer. Its success hints that structured reasoning, rather than sheer memorization, may be the key ingredient that allows language models to handle multi-step clinical questions, where a single symptom must be traced through differential diagnoses, laboratory interpretation, and treatment selection.</p>
<p>Yet the review is emphatic that these eye-catching scores should be read with caution, because the studies beneath them are far from uniform. The authors assessed study quality using a 12-item checklist covering dataset characteristics, how each language model was set up, and the evaluation methods employed. What they found were persistent inconsistencies across the literature: different studies used different datasets, some models were evaluated under opaque conditions with little disclosure of how prompts were constructed or how answers were scored, and the evaluation criteria themselves varied from paper to paper. In practical terms, this means that a score of 87 percent reported in one study and a score of 60 percent reported in another may not be directly comparable, because the underlying tests, question formats, scoring rules, and even the number of times each model was queried may all differ. The review also flags the issue of model transparency: when a commercial model&#8217;s architecture and training data are proprietary, researchers cannot rule out that exam content leaked into training, inflating performance in ways that have nothing to do with genuine clinical reasoning.</p>
<p>These methodological gaps are more than academic quibbles. The stakes of medical AI evaluation are enormous, because licensing examinations are not just tests; they are the social contract that certifies who may be trusted with human lives. If a language model&#8217;s apparent competence rests on contaminated data or lenient scoring, the danger is not that the model will be overrated in a paper, but that hospitals, educators, and regulators might make real decisions based on inflated numbers. Conversely, if evaluation is too fragmented, genuinely capable systems may be underestimated, slowing beneficial applications such as clinical decision support, medical education tutoring, and assistance for physicians in underserved regions. The review&#8217;s authors argue that the field needs standardized benchmarks, shared protocols, and transparent reporting so that each new model&#8217;s medical performance can be measured against a common yardstick rather than a shifting patchwork of ad hoc tests.</p>
<p>The authors also point toward multimodality as the next frontier. Real medicine is not a text-only discipline: a radiologist reads images, a dermatologist inspects lesions, an ultrasonographer interprets moving scans in real time. The Chinese examinations reviewed here are largely text-based, which means the current results capture only one slice of what physician competence entails. Models that can jointly process clinical text, medical images, laboratory data, and even patient histories will be far better probes of genuine clinical ability, and the review identifies multimodal capabilities as a priority for future research. Combined with transparent evaluation and clinically meaningful benchmarks, this direction could transform these examinations from a publicity benchmark for AI companies into a rigorous scientific instrument for measuring how close machines are actually getting to the practice of medicine.</p>
<p>What the PLOS Digital Health review ultimately documents is a turning point. In the span of a few years, large language models have gone from failing medical exams to surpassing the passing threshold of one of the world&#8217;s most demanding licensing systems, with GPT-4o and DeepSeek-R1 leading a generation of models that improve with each release. The 14 studies and 51 evaluation records synthesized by Zong, Wang, Shen, and their colleagues tell a story of rapid ascent shadowed by unfinished methodological work: datasets that need standardizing, models that need transparency, and evaluations that need to reach beyond text into the multimodal reality of the clinic. The machines have passed the written test. Whether they can be trusted at the bedside is the question the next generation of research, and the next generation of benchmarks, must now answer.</p>
<p><strong>Subject of Research:</strong> Performance of large language models on the Chinese National Medical Licensing Examination</p>
<p><strong>Article Title:</strong> Performance of large language models in the Chinese National Medical Licensing Examinations and Beyond: Scoping review</p>
<p><strong>Article References:</strong> Zong, H., Wang, J., Cha, J., Wu, R., Hu, S., Zhao, Y., &amp; Shen, B. (2026). Performance of large language models in the Chinese National Medical Licensing Examinations and Beyond: Scoping review. <em>PLOS Digital Health, 5</em>(10), e0001773. <a href="https://doi.org/10.1371/journal.pdig.0001773" rel="noopener noreferrer">https://doi.org/10.1371/journal.pdig.0001773</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pdig.0001773" rel="noopener noreferrer">10.1371/journal.pdig.0001773</a></p>
<p><strong>Keywords:</strong> large language models, GPT-4o, DeepSeek-R1, medical licensing examination, China, medical education, scoping review, clinical reasoning, AI benchmarks, multimodal AI, PLOS Digital Health, model evaluation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">256982</post-id>	</item>
	</channel>
</rss>
