<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>medical AI safety study &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/medical-ai-safety-study/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 08:55:04 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>medical AI safety study &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Medical Scribes Still Make Dangerous Errors in Surgical Notes, Study Finds</title>
		<link>https://scienmag.com/ai-medical-scribes-still-make-dangerous-errors-in-surgical-notes-study-finds/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 08:55:04 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI in healthcare documentation]]></category>
		<category><![CDATA[AI medical scribes]]></category>
		<category><![CDATA[AI transcription error rates]]></category>
		<category><![CDATA[AI-assisted clinical documentation]]></category>
		<category><![CDATA[ambient voice technology]]></category>
		<category><![CDATA[ambient voice technology safety]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[clinical documentation]]></category>
		<category><![CDATA[clinical note accuracy]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[healthcare AI technology evaluation]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[medical AI safety study]]></category>
		<category><![CDATA[medical scribes]]></category>
		<category><![CDATA[NHS]]></category>
		<category><![CDATA[patient safety]]></category>
		<category><![CDATA[patient safety risks AI]]></category>
		<category><![CDATA[SNOMED CT]]></category>
		<category><![CDATA[speech recognition]]></category>
		<category><![CDATA[speech-to-text medical systems]]></category>
		<category><![CDATA[surgery]]></category>
		<category><![CDATA[surgical note transcription errors]]></category>
		<category><![CDATA[surgical record errors]]></category>
		<category><![CDATA[transcription errors]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=252877</guid>

					<description><![CDATA[A controlled evaluation of eight ambient voice technologies found that 30 to 68 percent of AI-transcribed surgical records contained clinically significant errors, with the most dangerous mistakes concentrated in medication doses and laboratory results.]]></description>
										<content:encoded><![CDATA[<p>Ambient voice technologies, the artificial intelligence tools that listen to clinicians and turn their spoken words into clinical notes, are being adopted across hospitals at a pace that has outstripped the evidence behind them. Now a controlled evaluation published in the Journal of Medical Systems has put eight of these systems through a rigorous safety-focused test, and the results are a sobering reminder that even the most fluent-sounding AI transcripts can hide errors capable of harming patients. In the study, a team of researchers from King&#8217;s College London and collaborating NHS trusts found that between 30 and 68 percent of transcribed surgical records contained at least one clinically significant error, depending on the system used.</p>
<p>The research team, led by Ruairi O&#8217;Kane and Martyn Cobourne, designed the experiment to isolate the speech-to-text component of ambient voice technologies under conditions that would be as controlled as possible. They compiled 100 written surgical case-reports from open-access journals indexed in PubMed Central, totalling 32,897 words, with individual reports ranging from 44 to 449 words. Each report was narrated aloud by a single speaker in a quiet, controlled environment at King&#8217;s College London, producing 100 audio recordings that served as identical inputs for every system tested. This single-speaker design deliberately removes the confounding effects of accent variation, background noise and multiple speakers, meaning the error rates observed likely represent a best-case scenario for these technologies rather than the messier reality of a busy hospital.</p>
<p>Eight systems were evaluated: three commercial transcription products, namely Nuance Dragon Medical One, Heidi Health and Tortus; four standalone automatic speech recognition systems accessed through application programming interfaces, including Speechmatics Enhanced, Amazon Medical Transcribe, OpenAI&#8217;s Whisper and GPT-4o Transcribe; and an experimental two-stage pipeline in which the raw GPT-4o Transcribe output was passed to the GPT-5 large language model for generative error correction. The result was 800 transcript outputs, each compared against the original written case-report that served as the ground truth. To ensure that any discrepancies reflected genuine transcription failures rather than mistakes in the audio recordings themselves, two independent reviewers audited a random sample of 20 recordings covering 5,932 words and found just two errors, an error rate of 0.034 percent, well below accepted validation thresholds.</p>
<p>The primary outcome was deliberately safety-oriented rather than purely technical. Rather than simply counting wrong words, the researchers applied a modified version of the Kanal error analysis framework to identify so-called Class 3 errors: mistranscriptions that change the meaning of a sentence in a way that is not obvious on immediate inspection. An error was deemed clinically significant if it could lead a reader to misunderstand the clinical picture and potentially affect patient care. Two reviewers independently annotated all errors, achieving 96.61 percent raw agreement on the binary primary outcome with a Cohen&#8217;s kappa of 0.923, indicating near-perfect reliability, with a third reviewer resolving any disagreements. Each significant error was then graded for potential harm on a four-level scale aligned with the NHS England Harm Framework, ranging from low harm to death.</p>
<p>The headline finding is stark. Across all 800 outputs, the proportion containing at least one clinically significant Class 3 error ranged from 30 percent for the experimental GPT-4o Transcribe-Corrected-5 pipeline to 68 percent for both Amazon Medical Transcribe and Dragon Medical One. Even the best-performing systems, which achieved impressively low overall word error rates, still produced transcripts that a clinician could not safely accept without careful review. In total, 53 errors across the systems carried the potential for severe harm or death, and these clustered in two particularly dangerous domains: medication type and dose, which accounted for 39.6 percent of severe errors, and investigations and laboratory results, which accounted for 30.2 percent.</p>
<p>The danger lies in the subtlety of these mistakes. Among the examples identified by the reviewers were reversed meanings of comorbidities, distorted laboratory values, incorrect procedures documented, and medication errors in which both the wrong and the right dose were clinically plausible. One striking case involved 300 milligrams of aspirin being transcribed as 75 milligrams; both are real doses used for different indications, making the error easy to miss and easy to act upon incorrectly. The researchers point out that numbers are especially vulnerable in speech recognition because of phonetic similarities, such as fifteen versus fifty, and because numerals carry little semantic context that would help a language model self-correct. Paradoxically, the more polished and coherent an AI-generated note appears, the lower a clinician&#8217;s guard may drop during review, a phenomenon well documented in the automation bias literature.</p>
<p>The experimental pipeline offered a glimpse of how large language models might mitigate, but not eliminate, these risks. When the raw GPT-4o Transcribe output was processed by GPT-5 for generative error correction, the proportion of transcripts containing a clinically significant error fell from 53 percent to 30 percent, a statistically significant reduction. Of the 89 clinically significant errors present in the raw output, 51 were resolved by the language model, including context-dependent mistakes such as correcting the misspelled antifungal flucocytosine to the antibiotic flucloxacillin, which was the appropriate treatment for methicillin-sensitive Staphylococcus aureus in that case. Yet 38 errors persisted, and crucially, the language model introduced four new clinically significant errors of its own, including replacing the word malformation with malignant. Numerical errors were particularly resistant to correction, with a transcribed 6 percent remaining wrong even when the reference said 60 percent, because the correct value could not be reliably inferred from the audio alone.</p>
<p>Accuracy metrics told a consistent story across the board. The researchers calculated a domain-specific word error rate focused on clinical terminology drawn from SNOMED CT, alongside general-language error rates, character error rates, and lexical and semantic similarity measures using ROUGE, BERT and BART scores. The GPT-4o Transcribe-Corrected-5 pipeline achieved the best performance with a domain word error rate of 3.60 percent, followed by Heidi Health at 5.67 percent, while Amazon Medical Transcribe fared worst at 24.03 percent. For context, professional human transcription of conversational speech has been benchmarked at around 5.9 percent, meaning the top systems approach human parity while others fall far short. For most systems, medical terminology was transcribed less accurately than general language, though Dragon Medical One, a product specifically trained for healthcare dictation, showed the reverse pattern, excelling at clinical terms at the cost of general-language performance. Hallucinations, in which content absent from the audio appears in the transcript, were rare but not absent, occurring in Whisper and Heidi Health outputs and including everything from stray phrases like thank you for watching to plausible clinical distortions.</p>
<p>The study also uncovered a practical risk factor for deployment: dictation length. For four of the eight systems, each additional 100 dictated words significantly increased the odds of a clinically significant error, with odds ratios ranging from 2.65 to 3.34. This means longer operative notes carry disproportionately greater risk on some platforms, a finding with direct implications for how clinicians use these tools and how post-deployment monitoring should be designed. The authors acknowledge limitations, including the single-speaker design, which likely underestimates error rates compared with real clinical environments full of accents, interruptions and noise, and the focus on transcription rather than the full end-to-end ambient scribing workflow in which summarisation adds another opportunity for error. They also note that human-written clinical notes contain errors too, and that reported hallucination and omission rates for AI summarisation can fall below human note-taking error rates.</p>
<p>The message for clinicians and health systems is not that ambient voice technology should be abandoned, but that it cannot yet be trusted blindly. Under NHS England guidance and the Royal College of Surgeons of England&#8217;s Good Surgical Practice standards, clinicians remain legally and professionally responsible for the accuracy of the records they sign off, and every output must be reviewed before entering the care pathway. The researchers argue that the next step is not simply asking tired clinicians to spot subtle errors in fluent text, but redesigning the electronic health record interface itself so that high-risk content such as medications, doses, laboratory values, laterality and procedures is automatically flagged for verification against existing structured patient data. Until then, the most dangerous transcript may be the one that reads perfectly.</p>
<p><strong>Subject of Research:</strong> Accuracy and patient safety of ambient voice technology speech-to-text systems for surgical clinical documentation</p>
<p><strong>Article Title:</strong> A Controlled Single-Speaker Evaluation of Ambient Voice Technologies’ Speech-to-Text Function Using Narrated Surgical Case-Reports</p>
<p><strong>Article References:</strong> O’Kane, R., Stonehouse-Smith, D., Shirazi, M., Hutchinson, R., Gregg, M., Patel, R., Seehra, J., Papageorgiou, S. N., &amp; Cobourne, M. T. (2026). A Controlled Single-Speaker Evaluation of Ambient Voice Technologies’ Speech-to-Text Function Using Narrated Surgical Case-Reports. <em>Journal of Medical Systems, 50</em>(1), Article 145. <a href="https://doi.org/10.1007/s10916-026-02471-5" rel="noopener noreferrer">https://doi.org/10.1007/s10916-026-02471-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10916-026-02471-5" rel="noopener noreferrer">10.1007/s10916-026-02471-5</a></p>
<p><strong>Keywords:</strong> ambient voice technology, speech recognition, artificial intelligence, clinical documentation, patient safety, surgery, large language models, transcription errors, electronic health records, medical scribes, SNOMED CT, NHS</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">252877</post-id>	</item>
	</channel>
</rss>
