Hearing tests produce some of medicine’s most deceptively simple images. An audiogram is a grid of symbols marking the faintest sounds a patient can detect at each frequency, in each ear, with and without masking noise. A tympanogram traces how the eardrum moves under changing pressure. Interpreting these charts requires more than reading numbers: it demands knowledge of specialty conventions, masking rules, and the subtle distinction between a threshold that was not measured and a sound so loud the patient still could not hear it. A new study published in the Journal of Medical Systems has now tested, with unusual methodological rigor, whether multimodal artificial intelligence can perform this interpretation, and where it still fails.
The research team, led by investigators at the University of Hong Kong and Ningbo Hospital of Integrated Traditional Chinese and Western Medicine in China, took a deliberately modular approach. Rather than asking an AI to produce an end-to-end diagnosis from an image, they split the problem into two separate modules. The first tested whether a large multimodal model could transcribe and interpret pure-tone audiometry and tympanometry images. The second tested whether AI workflows could calculate twenty-five prespecified clinical fields and draft professional and patient-facing reports in Chinese, given already-verified structured data. This separation matters because a single end-to-end score can hide whether errors come from reading the image, applying specialty rules, or communicating the result.
The study drew on 158 outpatient audiology encounters collected between December 2024 and March 2025, of which 155 records representing 151 unique patients and 302 ears were eligible. Reference standards were built painstakingly: two trained transcribers independently entered every threshold while masked to each other, and two audiologists with 13 and 16 years of clinical experience classified tympanogram curves, agreeing on 296 of 300 dually classified ears, a Cohen’s kappa of 0.979. Air-bone gaps, hearing-loss degrees, and loss types were derived under prespecified rules, including a local convention that a meaningful air-bone gap required at least two comparable frequencies with gaps of 15 dB or more.
The heart of the first module was the audiologist skill: a carefully engineered prompt, not a fine-tuned model, that encoded audiological practice. Early development errors were revealing. The model initially produced thresholds not on the standard 5-dB grid, misapplied degree boundaries, overcalled conductive components, treated insufficient bone-conduction evidence as a negative air-bone gap, and read cancelled 95-dB acoustic-reflex marks as present responses. The refined skill imposed a fixed sequence: verify the image and ear, inspect axes and legends, assign symbols, transcribe before calculating, preserve no-response entries, then derive and cross-check. When masking could not be assigned unambiguously, the prompt instructed the model to abstain rather than guess.
The results of the locked test were striking. In 50 development records, the skill raised hearing-loss-type agreement from 85.0 to 94.0 percent and acoustic-reflex agreement from 70.4 to 98.4 percent compared with a schema-only prompt. In the held-out evaluation of 101 independent patients using a Codex GPT5.5 agentic workflow, the system achieved 1809 of 1820 exact numeric thresholds, 99.4 percent, and 91 of 101 patients met every numeric and no-response criterion. Degree was correct in all 101 patients, hearing-loss type in 100, and tympanometry measurements in all 578 entries. Yet the Achilles’ heel persisted: only 9 of 15 no-response entries were correct, and 18 of 19 air-bone-gap mismatches occurred because the model forced a negative label when the evidence supported an indeterminate one.
A post hoc robustness analysis using DeepSeek-V4-Flash-Vision-Exp on the same patients showed how much performance depends on the implementation. With the same locked skill, hearing-loss-type agreement rose from 37.6 to 70.3 percent, a dramatic improvement, but exact numeric-threshold agreement reached only 72.1 percent, reflex agreement hovered near 69 percent, and no-response agreement was zero. The authors are careful to note this was a descriptive cross-model comparison, not a matched foundation-model experiment, since the execution environments differed. The lesson, however, is clear: specialty guidance can substantially improve rule-dependent interpretation, but raw visual accuracy remains tied to the underlying model and workflow.
The second module addressed reporting. Here the AI received adjudicated structured values rather than its own image predictions, calculated twenty-five prespecified fields, and drafted separate Chinese professional and patient-facing reports. After two audiologists reviewed 60 initial cases and identified overreliance on the speech-frequency average, omission of high-frequency losses, and weak integration of history with results, the prompts were refined and locked. In the formal evaluation of 91 independent patients, all twenty-five rule-derived fields were correct in 91 of 91 Codex-workflow cases and 87 of 91 DeepSeek-workflow cases. Both audiologists rated every single report from both workflows as accurate or basically accurate; no report received a rating of clear error or potentially misleading.
An exploratory lay evaluation added a human dimension. Five lay raters compared pre-refinement patient-facing reports from the two workflows in 48 patients. DeepSeek reports were preferred for explanations in 44 of 48 patients and for next steps in 30, but they were also significantly longer, with a median of 392 versus 235 Chinese characters. Overall preference and perceived ease did not differ significantly. The authors emphasize that preference and readability proxies do not establish comprehension, and that no lay evaluation of the refined reports was conducted. This matters because hearing-health materials often exceed recommended reading levels, and limited health literacy can coexist with hearing loss in older adults.
The study’s limitations are candidly enumerated. It came from a single hospital with two devices over four months. Only 15 no-response entries existed, from just three validation patients. The two audiologists who rated the final reports were the same ones who had refined the prompts, raising the possibility of incorporation bias. Even with zero unfavorable ratings among 91 patients, the statistical upper bound on the unfavorable-report rate remains roughly 4.1 percent. Reproducibility was constrained by reliance on proprietary services: the exact Codex snapshot and sampling settings were unavailable, and provider data retention could not be excluded. No end-to-end test connected the two modules, so error propagation from image to report was never measured.
What the study ultimately offers is an architecture rather than a product. The authors envision a safety-conscious pipeline in which image transcription, deterministic validation and calculation, report drafting, uncertainty flags, and clinician approval remain visible, auditable handoff points. This aligns with what Chinese audiologists themselves have said in qualitative work: AI may absorb repetitive technical work, but communication, judgment, and responsibility should remain clinician-led. The findings support prospective evaluation of modular, clinician-supervised assistance, not autonomous diagnosis. The next step, the authors argue, is a silent prospective deployment that connects the modules while preserving intermediate outputs, measuring abstention, correction burden, review time, patient comprehension, and downstream clinical decisions. Until then, the audiogram-reading AI remains a promising apprentice, one that can transcribe nearly every threshold perfectly yet still needs its audiologist to teach it what silence means.
Subject of Research: Audiologist-guided multimodal AI for interpreting pure-tone audiometry and tympanometry and generating clinical reports
Article Title: Audiologist-Guided Multimodal AI for Pure-Tone Audiometry and Tympanometry Interpretation and Reporting
Article References: Wu, X., Shen, X., Mo, C., Shao, S., Wang, J., & Wang, S. (2026). Audiologist-Guided Multimodal AI for Pure-Tone Audiometry and Tympanometry Interpretation and Reporting. Journal of Medical Systems, 50(1), Article 140. https://doi.org/10.1007/s10916-026-02463-5
Image Credits: AI Generated
DOI: 10.1007/s10916-026-02463-5
Keywords: artificial intelligence, audiology, pure-tone audiometry, tympanometry, large language models, multimodal AI, clinical decision support, hearing loss, medical imaging AI, patient communication, prompt engineering, Journal of Medical Systems
Cite Scienmag News
Ophelia Keating. (October 1, 2026). AI Learns to Read Hearing Tests Like an Audiologist, But Not Yet Like a Doctor. Scienmag. https://scienmag.com/ai-learns-to-read-hearing-tests-like-an-audiologist-but-not-yet-like-a-doctor/
Ophelia Keating. "AI Learns to Read Hearing Tests Like an Audiologist, But Not Yet Like a Doctor." Scienmag, 1 October 2026, https://scienmag.com/ai-learns-to-read-hearing-tests-like-an-audiologist-but-not-yet-like-a-doctor/. Accessed 1 October 2026.
Ophelia Keating. "AI Learns to Read Hearing Tests Like an Audiologist, But Not Yet Like a Doctor." Scienmag. October 1, 2026. https://scienmag.com/ai-learns-to-read-hearing-tests-like-an-audiologist-but-not-yet-like-a-doctor/

