Ophthalmic B-scan ultrasound is one of the most widely used diagnostic tools for examining the back of the eye. When cataracts, bleeding, or other obstructions prevent a clear view of the retina, ultrasound becomes the clinician’s window into the posterior segment, revealing vitreous opacities, retinal detachment, tumors, and other sight-threatening conditions. Yet for all its diagnostic value, the technology has long carried a hidden administrative burden: every scan must be accompanied by a written diagnostic report, drafted manually by a clinician, describing findings and impressions in precise medical language. Writing those reports is slow, requires substantial expertise, and varies considerably in quality between examiners. A study now published in the Journal of Big Data proposes a striking solution—a multimodal artificial intelligence system called OphthUS-GPT that can generate diagnostic reports from ultrasound images automatically, and then explain them to users through interactive, multi-turn dialogue.
The research was carried out at the Affiliated Eye Hospital of Jiangxi Medical College, Nanchang University, in China, and rests on one of the largest datasets ever assembled for this purpose. The retrospective analysis encompassed 103,237 ophthalmic ultrasound images and 51,618 diagnostic reports drawn from 51,618 patients. The scale matters. Teaching an artificial intelligence system to read an ultrasound image and describe it the way an experienced ophthalmologist would requires tens of thousands of paired examples of images and the corresponding clinical text. The ethics committee of the hospital approved the study, and because it was a retrospective analysis of de-identified existing data, the requirement for individual informed consent was waived.
At the technical heart of OphthUS-GPT is a two-stage framework built on the Bootstrapping Language-Image Pre-training model, widely known as BLIP, a vision-language architecture that learns to connect visual content with descriptive text. The researchers fine-tuned BLIP in two distinct stages. In the first stage, the model was trained to produce clinical findings—the objective descriptions of what appears in the ultrasound image, such as the presence of echogenic material in the vitreous cavity or an elevated retinal membrane. In the second stage, the model learned to integrate those findings with clinical impressions, producing the complete diagnostic report structure that ophthalmologists use in daily practice. This staged approach mirrors the way human clinicians actually reason: first observing, then interpreting.
But generating a report is only half the problem. A written description of an ultrasound scan is of limited value to a patient, a trainee, or even a non-specialist physician if it cannot be understood. To address this, the team incorporated a large language model—DeepSeek, deployed locally—into the system to serve as an interactive question-answering module. Users can ask follow-up questions about a generated report, request clarification of medical terminology, or explore the clinical implications of specific findings, and the system responds through multi-turn dialogue. The result is not merely an automated typist but an interpretive assistant designed to make ophthalmic ultrasound findings accessible and actionable.
Evaluating such a system demands rigorous, multi-dimensional testing, and the researchers deployed three complementary strategies. First, they measured text quality using standard natural language generation metrics: BLEU, which measures n-gram overlap with reference reports; ROUGE, which captures longer matching sequences; and CIDEr, which weighs consensus with human-written descriptions. OphthUS-GPT achieved a BLEU-1 score of 0.5739, a ROUGE-L score of 0.6131, and a CIDEr score of 0.9818 for report generation—figures indicating substantial agreement between machine-generated text and the reports written by clinicians.
Second, the team assessed diagnostic performance through disease classification metrics, measuring accuracy, sensitivity, specificity, and F1 score across the range of conditions that appear in posterior segment ultrasound. The results were strongest for common conditions. For the majority of frequently encountered findings—including vitreous opacities, retinal detachment, posterior scleral staphyloma, and cataracts—classification accuracy exceeded 90 percent. Across all evaluated disease categories, overall accuracy exceeded 80 percent, a level of performance that suggests the system could function reliably as a first-pass drafting and screening tool even for less common pathology.
Third, and perhaps most importantly for clinical credibility, expert ophthalmologists rated the generated reports for accuracy and completeness. More than 90 percent of the reports produced by the system received scores of 3 or higher—the evaluation scale’s threshold for acceptable—on both criteria. This expert-validated quality is crucial, because statistical similarity metrics alone cannot guarantee that a report is medically sound. A report can match the vocabulary of its references while missing a critical finding; human review remains the gold standard, and by this measure OphthUS-GPT performed well.
The interactive question-answering module was evaluated separately, on four dimensions: accuracy, completeness, security, and user satisfaction. Here the study produced one of its most interesting comparative findings. The locally deployed DeepSeek model demonstrated statistically superior accuracy compared with Claude, and higher user satisfaction compared with both Claude and GPT-4 Turbo. No significant differences emerged between the models in completeness or security. The comparison matters for real-world deployment: large language models differ in how faithfully they answer medical questions and how satisfied clinicians and patients are with their responses, and the finding suggests that a locally hosted open model can outperform commercial alternatives on the dimensions that matter most, while also addressing data privacy concerns that often constrain hospital AI adoption.
The implications extend well beyond a single hospital in Nanchang. Manual ultrasound report generation is time-consuming and heavily dependent on clinician expertise, and in many parts of the world—particularly in lower-resource settings and in clinics without an on-site ultrasound specialist—ophthalmic B-scan interpretation is a genuine bottleneck in eye care. An automated system that produces expert-quality reports and explains them conversationally could shorten reporting times, standardize documentation quality, support less experienced examiners, and free clinicians to spend more time with patients and less time at the keyboard. Because the system includes an interpretive dialogue layer, it could also serve as a teaching tool for ophthalmology trainees learning to read ultrasound images, offering immediate, interactive explanations of findings and terminology.
The work, led by Fan Gan, Lei Chen, Weiguo Qin, and colleagues including corresponding author Zhipeng You, with Fan Gan and Lei Chen contributing equally, was supported by the Jiangxi Provincial Health Commission and the Public Hospital High-Quality Development Research Public Welfare Project. It represents a broader shift in medical artificial intelligence: away from narrow, single-task classifiers and toward integrated multimodal systems that combine computer vision, natural language generation, and conversational reasoning in a single clinical workflow. Challenges remain before such systems become routine—the evaluation was retrospective, prospective clinical deployment would require further validation across diverse populations and equipment, and any AI-generated report would still need clinician review. But the study demonstrates that the pieces can fit together. A vision-language model trained on more than one hundred thousand real ultrasound images can produce reports that experts judge accurate and complete, and a large language model layered on top can turn those reports into an interactive clinical conversation. For a diagnostic field where the written report has long lagged behind the image itself, that is a meaningful step forward.
Subject of Research: A multimodal artificial intelligence system for generating and interpreting ophthalmic B-scan ultrasound diagnostic reports
Article Title: OphthUS-GPT: a multimodal AI system for ophthalmic B-scan ultrasound report generation and interpretation
Article References: OphthUS-GPT: a multimodal AI system for ophthalmic B-scan ultrasound report generation and interpretation. (n.d.). https://doi.org/10.1186/s40537-026-01563-w
Image Credits: AI Generated
DOI: 10.1186/s40537-026-01563-w
Keywords: artificial intelligence, multimodal learning, ophthalmic B-scan ultrasound, report generation, BLIP, DeepSeek, large language models, ophthalmology, medical imaging, clinical decision support, vision-language models, deep learning
Cite Scienmag News
Blake Davidson. (September 22, 2026). New AI System Writes Eye Ultrasound Reports and Explains Them. Scienmag. https://scienmag.com/new-ai-system-writes-eye-ultrasound-reports-and-explains-them/
Blake Davidson. "New AI System Writes Eye Ultrasound Reports and Explains Them." Scienmag, 22 September 2026, https://scienmag.com/new-ai-system-writes-eye-ultrasound-reports-and-explains-them/. Accessed 22 September 2026.
Blake Davidson. "New AI System Writes Eye Ultrasound Reports and Explains Them." Scienmag. September 22, 2026. https://scienmag.com/new-ai-system-writes-eye-ultrasound-reports-and-explains-them/

