<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>report generation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/report-generation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 22 Sep 2026 14:54:28 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>report generation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI System Writes Eye Ultrasound Reports and Explains Them</title>
		<link>https://scienmag.com/new-ai-system-writes-eye-ultrasound-reports-and-explains-them/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 14:54:28 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[BLIP]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[DeepSeek]]></category>
		<category><![CDATA[eye ultrasound images and generate accurate reports at this scale represents a significant advancement in ophthalmic diagnostics and medical AI applications]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Medical Imaging]]></category>
		<category><![CDATA[multimodal learning]]></category>
		<category><![CDATA[ophthalmic B-scan ultrasound]]></category>
		<category><![CDATA[ophthalmology]]></category>
		<category><![CDATA[report generation]]></category>
		<category><![CDATA[vision-language models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=206039</guid>

					<description><![CDATA[Researchers have developed OphthUS-GPT, a multimodal AI system that automatically generates and interprets diagnostic reports from ophthalmic B-scan ultrasound images with clinically strong accuracy.]]></description>
										<content:encoded><![CDATA[<p>Ophthalmic B-scan ultrasound is one of the most widely used diagnostic tools for examining the back of the eye. When cataracts, bleeding, or other obstructions prevent a clear view of the retina, ultrasound becomes the clinician&#8217;s window into the posterior segment, revealing vitreous opacities, retinal detachment, tumors, and other sight-threatening conditions. Yet for all its diagnostic value, the technology has long carried a hidden administrative burden: every scan must be accompanied by a written diagnostic report, drafted manually by a clinician, describing findings and impressions in precise medical language. Writing those reports is slow, requires substantial expertise, and varies considerably in quality between examiners. A study now published in the Journal of Big Data proposes a striking solution—a multimodal artificial intelligence system called OphthUS-GPT that can generate diagnostic reports from ultrasound images automatically, and then explain them to users through interactive, multi-turn dialogue.</p>
<p>The research was carried out at the Affiliated Eye Hospital of Jiangxi Medical College, Nanchang University, in China, and rests on one of the largest datasets ever assembled for this purpose. The retrospective analysis encompassed 103,237 ophthalmic ultrasound images and 51,618 diagnostic reports drawn from 51,618 patients. The scale matters. Teaching an artificial intelligence system to read an ultrasound image and describe it the way an experienced ophthalmologist would requires tens of thousands of paired examples of images and the corresponding clinical text. The ethics committee of the hospital approved the study, and because it was a retrospective analysis of de-identified existing data, the requirement for individual informed consent was waived.</p>
<p>At the technical heart of OphthUS-GPT is a two-stage framework built on the Bootstrapping Language-Image Pre-training model, widely known as BLIP, a vision-language architecture that learns to connect visual content with descriptive text. The researchers fine-tuned BLIP in two distinct stages. In the first stage, the model was trained to produce clinical findings—the objective descriptions of what appears in the ultrasound image, such as the presence of echogenic material in the vitreous cavity or an elevated retinal membrane. In the second stage, the model learned to integrate those findings with clinical impressions, producing the complete diagnostic report structure that ophthalmologists use in daily practice. This staged approach mirrors the way human clinicians actually reason: first observing, then interpreting.</p>
<p>But generating a report is only half the problem. A written description of an ultrasound scan is of limited value to a patient, a trainee, or even a non-specialist physician if it cannot be understood. To address this, the team incorporated a large language model—DeepSeek, deployed locally—into the system to serve as an interactive question-answering module. Users can ask follow-up questions about a generated report, request clarification of medical terminology, or explore the clinical implications of specific findings, and the system responds through multi-turn dialogue. The result is not merely an automated typist but an interpretive assistant designed to make ophthalmic ultrasound findings accessible and actionable.</p>
<p>Evaluating such a system demands rigorous, multi-dimensional testing, and the researchers deployed three complementary strategies. First, they measured text quality using standard natural language generation metrics: BLEU, which measures n-gram overlap with reference reports; ROUGE, which captures longer matching sequences; and CIDEr, which weighs consensus with human-written descriptions. OphthUS-GPT achieved a BLEU-1 score of 0.5739, a ROUGE-L score of 0.6131, and a CIDEr score of 0.9818 for report generation—figures indicating substantial agreement between machine-generated text and the reports written by clinicians.</p>
<p>Second, the team assessed diagnostic performance through disease classification metrics, measuring accuracy, sensitivity, specificity, and F1 score across the range of conditions that appear in posterior segment ultrasound. The results were strongest for common conditions. For the majority of frequently encountered findings—including vitreous opacities, retinal detachment, posterior scleral staphyloma, and cataracts—classification accuracy exceeded 90 percent. Across all evaluated disease categories, overall accuracy exceeded 80 percent, a level of performance that suggests the system could function reliably as a first-pass drafting and screening tool even for less common pathology.</p>
<p>Third, and perhaps most importantly for clinical credibility, expert ophthalmologists rated the generated reports for accuracy and completeness. More than 90 percent of the reports produced by the system received scores of 3 or higher—the evaluation scale&#8217;s threshold for acceptable—on both criteria. This expert-validated quality is crucial, because statistical similarity metrics alone cannot guarantee that a report is medically sound. A report can match the vocabulary of its references while missing a critical finding; human review remains the gold standard, and by this measure OphthUS-GPT performed well.</p>
<p>The interactive question-answering module was evaluated separately, on four dimensions: accuracy, completeness, security, and user satisfaction. Here the study produced one of its most interesting comparative findings. The locally deployed DeepSeek model demonstrated statistically superior accuracy compared with Claude, and higher user satisfaction compared with both Claude and GPT-4 Turbo. No significant differences emerged between the models in completeness or security. The comparison matters for real-world deployment: large language models differ in how faithfully they answer medical questions and how satisfied clinicians and patients are with their responses, and the finding suggests that a locally hosted open model can outperform commercial alternatives on the dimensions that matter most, while also addressing data privacy concerns that often constrain hospital AI adoption.</p>
<p>The implications extend well beyond a single hospital in Nanchang. Manual ultrasound report generation is time-consuming and heavily dependent on clinician expertise, and in many parts of the world—particularly in lower-resource settings and in clinics without an on-site ultrasound specialist—ophthalmic B-scan interpretation is a genuine bottleneck in eye care. An automated system that produces expert-quality reports and explains them conversationally could shorten reporting times, standardize documentation quality, support less experienced examiners, and free clinicians to spend more time with patients and less time at the keyboard. Because the system includes an interpretive dialogue layer, it could also serve as a teaching tool for ophthalmology trainees learning to read ultrasound images, offering immediate, interactive explanations of findings and terminology.</p>
<p>The work, led by Fan Gan, Lei Chen, Weiguo Qin, and colleagues including corresponding author Zhipeng You, with Fan Gan and Lei Chen contributing equally, was supported by the Jiangxi Provincial Health Commission and the Public Hospital High-Quality Development Research Public Welfare Project. It represents a broader shift in medical artificial intelligence: away from narrow, single-task classifiers and toward integrated multimodal systems that combine computer vision, natural language generation, and conversational reasoning in a single clinical workflow. Challenges remain before such systems become routine—the evaluation was retrospective, prospective clinical deployment would require further validation across diverse populations and equipment, and any AI-generated report would still need clinician review. But the study demonstrates that the pieces can fit together. A vision-language model trained on more than one hundred thousand real ultrasound images can produce reports that experts judge accurate and complete, and a large language model layered on top can turn those reports into an interactive clinical conversation. For a diagnostic field where the written report has long lagged behind the image itself, that is a meaningful step forward.</p>
<p><strong>Subject of Research:</strong> A multimodal artificial intelligence system for generating and interpreting ophthalmic B-scan ultrasound diagnostic reports</p>
<p><strong>Article Title:</strong> OphthUS-GPT: a multimodal AI system for ophthalmic B-scan ultrasound report generation and interpretation</p>
<p><strong>Article References:</strong> OphthUS-GPT: a multimodal AI system for ophthalmic B-scan ultrasound report generation and interpretation. (n.d.). <a href="https://doi.org/10.1186/s40537-026-01563-w" rel="noopener noreferrer">https://doi.org/10.1186/s40537-026-01563-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40537-026-01563-w" rel="noopener noreferrer">10.1186/s40537-026-01563-w</a></p>
<p><strong>Keywords:</strong> artificial intelligence, multimodal learning, ophthalmic B-scan ultrasound, report generation, BLIP, DeepSeek, large language models, ophthalmology, medical imaging, clinical decision support, vision-language models, deep learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">206039</post-id>	</item>
		<item>
		<title>New AI Assistant Reads Whole Pathology Slides and Answers Clinician Questions Across 31 Cancer Types</title>
		<link>https://scienmag.com/new-ai-assistant-reads-whole-pathology-slides-and-answers-clinician-questions-across-31-cancer-types/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 19:20:46 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[advanced diagnostic AI systems]]></category>
		<category><![CDATA[AI accuracy in cancer detection]]></category>
		<category><![CDATA[AI for multiple cancer types]]></category>
		<category><![CDATA[AI-assisted cancer reporting]]></category>
		<category><![CDATA[cancer diagnosis]]></category>
		<category><![CDATA[cancer diagnosis AI]]></category>
		<category><![CDATA[clinical question answering AI]]></category>
		<category><![CDATA[computational pathology]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[digital pathology]]></category>
		<category><![CDATA[digital pathology revolution]]></category>
		<category><![CDATA[gigapixel-scale medical imaging]]></category>
		<category><![CDATA[instruction tuning]]></category>
		<category><![CDATA[integration of AI and pathology]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[multimodal AI]]></category>
		<category><![CDATA[multimodal AI in pathology]]></category>
		<category><![CDATA[Nature Cancer]]></category>
		<category><![CDATA[pathology AI]]></category>
		<category><![CDATA[pathology slide interpretation]]></category>
		<category><![CDATA[report generation]]></category>
		<category><![CDATA[SlideChat]]></category>
		<category><![CDATA[whole-slide image analysis]]></category>
		<category><![CDATA[whole-slide imaging]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197812</guid>

					<description><![CDATA[Researchers have developed SlideChat, a multimodal generative AI assistant that interprets gigapixel whole-slide pathology images across 31 cancer types and outperforms leading baselines on diagnosis, question answering and report generation.]]></description>
										<content:encoded><![CDATA[<p>Pathology is undergoing a quiet revolution, and a new study published in Nature Cancer may accelerate it dramatically. A research team led by Ying Chen, Chenglong Ma and colleagues at the Shanghai Artificial Intelligence Laboratory, in collaboration with clinical partners including Shanghai General Hospital, the Eastern Hepatobiliary Surgery Hospital and Stanford University School of Medicine, has unveiled SlideChat, a multimodal generative artificial intelligence assistant designed to interpret gigapixel-scale whole-slide images, the enormous digital scans that form the backbone of modern cancer diagnosis. Unlike earlier AI systems that analyze only small crops of tissue, SlideChat is built to reason over the entire slide, answering clinical questions and generating diagnostic-style reports across 31 cancer types. Expert pathologists who reviewed the system&#8217;s answers rated them as accurate, clinically relevant and, in several dimensions, superior to those of every competing model tested.</p>
<p>The technical hurdle the team set out to overcome is deceptively simple to state and extraordinarily difficult to solve. A single whole-slide image can contain billions of pixels, far exceeding the input capacity of conventional vision-language models. Most existing multimodal AI assistants in pathology, including widely cited systems adapted from general-purpose models, operate at the patch level, examining isolated regions of tissue. That approach works for tasks such as detecting mitotic figures or classifying small lesions, but it fails when a diagnosis depends on global tissue architecture. Determining tumor staging, for example, requires appreciating whether invasive cells have breached anatomical boundaries visible only when the entire slide is considered. The researchers demonstrated this limitation directly: in a challenging bladder cancer case, patch-level models misidentified the invasive stage, likely because they could not integrate spatial context across the slide, while SlideChat correctly analyzed the global architecture and reported the accurate pT3 stage.</p>
<p>SlideChat&#8217;s architecture combines three essential ingredients. First, it integrates a patch-level pathology encoder, which captures fine-grained cellular and subcellular detail, with a slide-level pathology encoder that aggregates information across the full specimen into a compact representation. Second, it connects these visual encoders to a pretrained large language model, allowing the system to translate visual evidence into fluent natural-language answers. Third, and perhaps most importantly, it is trained on SlideInstruction, a new dataset of 274,233 multimodal instruction samples that pair whole-slide images with diagnostic reports and question-answer pairs. The instruction-tuning paradigm, which proved transformative in general-purpose chatbots, teaches the model not merely to classify images but to interpret complex, realistic clinical queries and respond in the structured language that pathologists actually use.</p>
<p>Building the training data required considerable curation. The team drew on publicly available resources including The Cancer Genome Atlas and the Clinical Proteomic Tumor Analysis Consortium, whose whole-slide images and associated pathology reports are accessible through the National Institutes of Health data commons. Additional data came from the BCNB breast cancer cohort and the HISTAI dataset, while a retrospectively collected hepatobiliary cohort, gathered under approval from the Eastern Hepatobiliary Surgery Hospital, provided controlled-access material. Raw reports were parsed and cleaned, concise captions were extracted to teach image-language alignment, and instruction-style question-answer pairs were generated from the reports and labels. A rigorous quality-control pipeline then applied large language model filtering and pathologist verification to ensure that questions genuinely could not be answered without visual inspection of the slide, guarding against shortcuts that would let the model succeed through text alone.</p>
<p>The evaluation was among the most comprehensive ever assembled for slide-level pathology AI. The team created SlideBench, a benchmark spanning five cohorts and 31 cancer types, comprising 8,836 closed-ended questions, 129 open-ended questions and 3,149 whole-slide diagnostic reports. On closed-ended questions, SlideChat outperformed leading baseline models by 19.1 percentage points in accuracy, a margin the authors validated with two-sided Wilcoxon signed-rank tests and Benjamini-Hochberg correction across 1,000 bootstrap replicates. On report generation, judged by the Metric for Evaluation of Translation with Explicit Ordering, a standard automatic measure of textual overlap with reference documents, SlideChat exceeded the best baselines by 7.7 points. For open-ended questions, expert pathologists scored SlideChat highest across five evaluation dimensions, noting that its answers were the most diagnostically accurate and case-specific. In one representative differential-diagnosis case, SlideChat integrated key morphological features to support a refined melanoma diagnosis, while a leading general-purpose model produced a broad but weakly prioritized differential and competing medical models returned generic, weakly reasoned answers.</p>
<p>The comparison with general-purpose and specialist baselines was particularly revealing. Models such as GPT-4o, LLaVA-Med, Quilt-LLaVA and dedicated slide-level systems including HistoGPT and PRISM all trailed SlideChat on most tasks. In report generation case studies, SlideChat consistently produced accurate, structured reports that captured tumor type, anatomical location, invasion status, lymphovascular features, TNM staging and relevant histological details, closely matching reference clinical reports. HistoGPT, by contrast, sometimes produced anatomically inconsistent or diagnostically mismatched descriptions, while PRISM returned brief diagnoses lacking contextual or morphological explanation. In renal and lung carcinoma examples from the CPTAC cohort, SlideChat correctly identified tumor subtype, Fuhrman grade, pT2a staging, the absence of vascular and perineural invasion and negative margins, and even contextualized incidental benign findings such as emphysematous changes in the lung without drifting into irrelevant speculation.</p>
<p>The study also probes how the model thinks, offering an unusually transparent window into its behavior. Question-guided attention heatmaps show that SlideChat reallocates its visual focus depending on the query: when asked about cytology in an adenocarcinoma case, attention concentrates on nuclear detail, while a question about differentiation shifts attention to glandular architecture; in a breast carcinoma case, the model attends to fibrous stroma when asked about stromal reaction but to tumor nests when asked about cellular arrangement. Ablation experiments confirmed that both the slide-level encoder and the two-stage training procedure are indispensable to performance. Sensitivity tests in which Gaussian noise was injected into patches with high or low attention scores showed a monotonic drop in accuracy when meaningful regions were corrupted, indicating that the model relies on genuine visual content rather than artifacts or dataset shortcuts. When adapted patch-level foundation models were given whole-slide inputs through strategies such as patch voting, thumbnail downsampling or joint multi-patch input, SlideChat still outperformed them, often by statistically decisive margins.</p>
<p>Not everything in the evaluation was flattering, and the authors are candid about the model&#8217;s weaknesses. In multiturn conversational settings, SlideChat exhibited cross-turn inconsistency in some cases, at one point initially identifying a node-negative tumor and later contradicting that finding during prognostic assessment. The model also occasionally produced self-contradictory hallucinations, describing a tumor as confined to the epithelium while simultaneously reporting vascular invasion. These failure modes mirror the reliability challenges that afflict large language models generally, and their documentation in a rigorous pathology benchmark provides a concrete target for future work. The researchers also note that performance did not correlate strongly with the number of training samples per organ, suggesting that data scale alone does not determine competence and that data quality and task diversity matter substantially.</p>
<p>The potential applications extend beyond diagnosis. The authors highlight medical education as a promising use case, since a system that can answer natural-language questions about whole slides could serve as an interactive tutor for pathology trainees. Clinical decision support is another frontier: by integrating slide-level reasoning with conversational interaction, the assistant could help oncologists and pathologists rapidly extract staging, grading and prognostic information from complex specimens. Notably, SlideChat achieves performance competitive with specialized pathology foundation models such as CONCH, TITON-class slide encoders, CHIEF and Prov-GigaPath on breast cancer classification tasks involving tumor status and hormone receptor and HER2 status, despite being a general-purpose assistant rather than a task-specific classifier. The team has released the training and evaluation data on Hugging Face, the model source code on GitHub and the model weights publicly, a level of openness that could seed a new generation of slide-level pathology AI research.</p>
<p>The work arrives at a moment when computational pathology is transitioning from narrow, single-task algorithms to foundation models and generative assistants, and it addresses what many in the field identify as the central bottleneck: the gap between patch-level perception and slide-level clinical reasoning. By demonstrating that a single model can answer closed-ended questions, engage in open-ended diagnostic dialogue and generate expert-quality reports across a broad range of cancers, and by subjecting that model to pathologist review, extensive ablations and honest failure analysis, the study sets a new reference point for what clinically useful pathology AI looks like. Substantial work remains before such systems can enter routine practice, including prospective validation, regulatory review and the resolution of conversational reliability issues. But the trajectory is clear. The microscope, long the pathologist&#8217;s solitary instrument, is gaining a conversational partner, and that partnership may reshape how cancer is diagnosed, taught and understood.</p>
<p><strong>Subject of Research:</strong> A multimodal generative AI assistant for whole-slide computational pathology across cancer types</p>
<p><strong>Article Title:</strong> SlideChat is a multimodal generative artificial intelligence assistant for whole-slide computational pathology across cancer types</p>
<p><strong>Article References:</strong> Chen, Y., Ma, C., Li, Q., Yan, F., Chen, Y., Li, T., Ye, J., Hu, M., Lin, Y., Li, Y., Wang, G., Xu, H., Dong, H., Wang, X., Xu, X., Zhou, Y., Zhu, X., Yang, S., Wang, X., &#8230; Ji, Y. (2026). SlideChat is a multimodal generative artificial intelligence assistant for whole-slide computational pathology across cancer types. <em>Nature Cancer</em>. <a href="https://doi.org/10.1038/s43018-026-01220-4" rel="noopener noreferrer">https://doi.org/10.1038/s43018-026-01220-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s43018-026-01220-4" rel="noopener noreferrer">10.1038/s43018-026-01220-4</a></p>
<p><strong>Keywords:</strong> SlideChat, computational pathology, whole-slide imaging, multimodal AI, large language models, cancer diagnosis, digital pathology, instruction tuning, Nature Cancer, report generation, pathology AI, deep learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197812</post-id>	</item>
	</channel>
</rss>
