<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>clinically meaningful performance evaluation of AI models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/clinically-meaningful-performance-evaluation-of-ai-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 23:15:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>clinically meaningful performance evaluation of AI models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Read X-rays Like a Radiologist to Spot Bone Tumors With Few False Alarms</title>
		<link>https://scienmag.com/ai-learns-to-read-x-rays-like-a-radiologist-to-spot-bone-tumors-with-few-false-alarms/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 23:15:11 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI system for rare primary bone tumors]]></category>
		<category><![CDATA[AI-based radiograph lesion classification]]></category>
		<category><![CDATA[AI-powered bone tumor detection]]></category>
		<category><![CDATA[bone tumor detection]]></category>
		<category><![CDATA[BTXRD]]></category>
		<category><![CDATA[clinically meaningful performance evaluation of AI models]]></category>
		<category><![CDATA[computer-aided detection]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for bone tumor diagnosis]]></category>
		<category><![CDATA[EfficientNetV2]]></category>
		<category><![CDATA[EfficientNetV2-S in medical imaging]]></category>
		<category><![CDATA[false alarm reduction in medical imaging]]></category>
		<category><![CDATA[false positives per image]]></category>
		<category><![CDATA[FROC analysis]]></category>
		<category><![CDATA[leave-one-center-out validation]]></category>
		<category><![CDATA[localizing bone tumors in X-rays]]></category>
		<category><![CDATA[MURA]]></category>
		<category><![CDATA[musculoskeletal radiology AI tools]]></category>
		<category><![CDATA[radiography]]></category>
		<category><![CDATA[radiologist-inspired AI detection pipelines]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[two-stage radiograph tumor localization]]></category>
		<category><![CDATA[X-ray image analysis for cancer detection]]></category>
		<category><![CDATA[YOLOv8]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=250329</guid>

					<description><![CDATA[Researchers built a two-stage, biomimetically inspired AI pipeline that detects primary bone tumors on X-rays while respecting strict false-alarm budgets, revealing both promising sensitivity and fragile cross-source generalization.]]></description>
										<content:encoded><![CDATA[<p>Primary bone tumors are among the most treacherous findings in musculoskeletal radiology. They are rare, their appearances on plain X-rays can be maddeningly subtle, and the consequences of a missed lesion can be devastating. A new study published in BMC Medical Imaging by Behnam Kiani Kalejahi, Sajid Khan and Mohammad Javad Rajabi takes aim at exactly this diagnostic blind spot, and it does so with an unusually honest evaluation strategy that could reshape how artificial intelligence tools for cancer detection are judged. Rather than chasing headline accuracy numbers, the researchers built a two-stage detection pipeline modeled loosely on how an expert radiologist actually reads a film, and then measured its performance under strict, clinically meaningful limits on how many false alarms it is allowed to raise.</p>
<p>The architecture is deliberately simple in concept. The system consists of two independently trained components: an image-level classifier that decides whether a radiograph contains a suspicious lesion at all, and a lesion detector that localizes and marks candidate tumors on the image. The classifier is built on an EfficientNetV2-S backbone, first pretrained on MURA, a large Stanford dataset of musculoskeletal radiographs, and then fine-tuned on BTXRD, a bone tumor X-ray repository, using a two-phase unfreezing schedule that gradually releases layers of the network for retraining. The detector is a YOLOv8 model initialized from weights learned on the COCO object-detection benchmark. Crucially, the classifier can act as a gate: when it judges an image normal, the detector need not run, a design that mirrors the way a human reader first forms a global impression before hunting for specific findings.</p>
<p>The evaluation framework is where the paper makes its most distinctive contribution. The authors argue that summary metrics such as the area under the receiver operating characteristic curve, or AUC, are fundamentally ill-suited to detection tasks because they say nothing about how many false marks a system places on each examination. A detector could post an impressive AUC while littering normal films with spurious annotations that would erode clinician trust in practice. Instead, the team used free-response receiver operating characteristic, or FROC, analysis, the established endpoint for computer-aided detection, and evaluated performance at explicit false-mark budgets that were fixed in advance on validation data before the held-out test set was ever touched.</p>
<p>The data behind the study come from the BTXRD collection, comprising 3,746 radiographs drawn from three sources, of which 1,867 show lesions and 1,879 are normal. These were partitioned into training, validation and test sets of 2,265, 738 and 743 images respectively, with the test set held out completely until all thresholds had been locked. This discipline matters. Many machine learning studies in medical imaging tune their operating points on the very data they then report, inflating performance. Here, every decision threshold was frozen on validation data, and the test results represent a genuine prediction on unseen material.</p>
<p>The headline results are solid if unspectacular, and the authors are refreshingly candid about that. On the test set, the classifier distinguished malignant from normal radiographs with an AUC of 0.933 when initialized with MURA pretraining, compared with 0.910 using standard ImageNet initialization. The benefit of musculoskeletal pretraining was consistent in direction across three random seeds, adding between 0.023 and 0.073 to the AUC, which the authors describe as suggestive of a modest advantage, while noting that the confidence interval at the pre-specified seed included zero. In other words, domain-specific pretraining probably helps, but the evidence stops short of certainty, and the paper says so rather than overclaiming.</p>
<p>Detection performance was assessed at an overlap criterion of IoU at or above 0.3, meaning a predicted mark counts as correct only if it overlaps the true lesion sufficiently. At an achieved rate of 0.027 false marks per normal radiograph, corresponding to a validation budget of 0.02, lesion-level sensitivity reached 0.646. Relaxing the budget to an achieved 0.055 false marks per normal film raised sensitivity to 0.698. At the primary operating point, the system produced 0.40 false marks per lesion-positive radiograph. Sensitivity differed by lesion biology: 0.736 for malignant lesions versus 0.631 for benign ones, a gap that makes intuitive sense given that malignant bone tumors often produce more dramatic radiographic changes.</p>
<p>The matched ablation experiments add further texture. By comparing MURA initialization against ImageNet initialization, and two-phase fine-tuning against full fine-tuning under otherwise identical conditions, the study isolates the contribution of each design choice rather than bundling them into an opaque final system. This kind of controlled comparison is still uncommon in applied medical AI, where papers frequently report only a single best configuration. The leave-one-center-out experiments, however, delivered the study&#8217;s most sobering finding. When the smallest of the three data sources was held out, sensitivity held at 0.669, but when the largest source was removed and only 808 training images remained, sensitivity collapsed to 0.287.</p>
<p>That collapse is not a footnote; it is arguably the central lesson of the paper. It demonstrates that cross-source transfer in bone tumor detection is fragile and heavily dependent on the volume and diversity of available training data. A model that performs admirably when trained on nearly 2,300 images from multiple sources can lose most of its diagnostic power when one source disappears from the training pool. For any hospital or research group hoping to deploy such a tool, this is exactly the kind of stress test that should precede real-world use, and the authors explicitly state that cross-source transfer and clinical utility at realistic disease prevalence remain to be established.</p>
<p>The biomimetic framing of the pipeline deserves attention as more than marketing language. Radiologists do not examine every pixel of an X-ray with equal scrutiny; they triage, forming an impression of whether anything is abnormal before committing attention to specific regions. By allowing the classifier to gate the detector, the system embodies a similar division of labor, which carries practical benefits beyond accuracy. Gating can reduce computational cost on normal films and, more importantly, can suppress false marks on images where no lesion exists, directly addressing the false-positive burden that FROC analysis is designed to quantify. The two-stage organization also means each component can be improved, retrained or audited independently, which matters for the regulatory pathways that medical AI systems must eventually navigate.</p>
<p>What makes this study resonate beyond its specific task is its methodological posture. The authors used only publicly available, fully anonymized datasets, received no external funding, and openly disclose the use of generative AI tools for language editing and script drafting while taking full responsibility for the underlying experiments. They resist the temptation to inflate a 0.646 sensitivity figure into a claim of clinical readiness, instead framing sensitivity at fixed false-mark budgets as information that curve-level summaries simply do not convey. In a field where detection systems are often promoted on glossy benchmark numbers, a paper that fixes its thresholds before touching the test set, reports confidence intervals throughout, and shows exactly where its model breaks down when a data source vanishes is a quiet but pointed argument for how medical AI should be evaluated. The bone tumor detector itself is a promising prototype; the evaluation culture it models may prove to be the more consequential contribution.</p>
<p><strong>Subject of Research:</strong> Machine learning detection of primary bone tumors on radiographs using FROC evaluation under false-positive constraints</p>
<p><strong>Article Title:</strong> A biomimetically inspired staged pipeline for primary bone tumor detection on radiographs: held-out FROC evaluation under explicit false-positive constraints</p>
<p><strong>Article References:</strong> Kalejahi, B. K., Khan, S., &amp; Rajabi, M. J. (2026). A biomimetically inspired staged pipeline for primary bone tumor detection on radiographs: held-out FROC evaluation under explicit false-positive constraints. <em>BMC Medical Imaging</em>. <a href="https://doi.org/10.1186/s12880-026-02915-x" rel="noopener noreferrer">https://doi.org/10.1186/s12880-026-02915-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12880-026-02915-x" rel="noopener noreferrer">10.1186/s12880-026-02915-x</a></p>
<p><strong>Keywords:</strong> bone tumor detection, radiography, FROC analysis, false positives per image, deep learning, EfficientNetV2, YOLOv8, transfer learning, MURA, BTXRD, leave-one-center-out validation, computer-aided detection</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">250329</post-id>	</item>
	</channel>
</rss>
