<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>model calibration &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/model-calibration/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 22 Sep 2026 15:39:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>model calibration &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Decision-Curve Analysis Is a Bridge, Not an Endpoint, for Clinical AI in Oncology</title>
		<link>https://scienmag.com/decision-curve-analysis-is-a-bridge-not-an-endpoint-for-clinical-ai-in-oncology/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 15:39:25 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[algorithmic bias and model validation in oncology]]></category>
		<category><![CDATA[bridging research and clinical practice in oncology AI]]></category>
		<category><![CDATA[challenges in translating AI to cancer care]]></category>
		<category><![CDATA[clinical artificial intelligence]]></category>
		<category><![CDATA[CONSORT-AI]]></category>
		<category><![CDATA[DECIDE-AI]]></category>
		<category><![CDATA[decision curve analysis]]></category>
		<category><![CDATA[Decision-curve analysis in clinical AI]]></category>
		<category><![CDATA[FUTURE-AI]]></category>
		<category><![CDATA[impact of AI decision tools on patient outcomes]]></category>
		<category><![CDATA[importance of clinical evaluation in AI deployment]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[model calibration]]></category>
		<category><![CDATA[net benefit]]></category>
		<category><![CDATA[oncology]]></category>
		<category><![CDATA[oncology prediction models]]></category>
		<category><![CDATA[pitfalls of over-relying on decision-curve analysis]]></category>
		<category><![CDATA[prospective validation]]></category>
		<category><![CDATA[regulatory considerations for AI in healthcare]]></category>
		<category><![CDATA[risk of premature implementation of AI models]]></category>
		<category><![CDATA[role of decision-curve analysis in clinical validation]]></category>
		<category><![CDATA[translational framework for cancer AI]]></category>
		<category><![CDATA[Translational Medicine]]></category>
		<category><![CDATA[TRIPOD+AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=206487</guid>

					<description><![CDATA[Researchers argue that a widely used statistical measure of clinical utility must sit within a staged evidence pathway before cancer AI models reach patients.]]></description>
										<content:encoded><![CDATA[<p>A statistical tool that has quietly become one of the most influential gatekeepers in clinical artificial intelligence is being asked to carry more weight than it can bear, according to a new letter published in the Journal of Translational Medicine. The correspondence, authored by Musfira Khalid of SUNY Upstate Medical University, Muzna Sarfraz of King Edward Medical University in Lahore, and Faheem Javad of Al Nafees Medical College, Isra University in Islamabad, takes aim at a growing tendency in the oncology AI literature: treating decision-curve analysis as definitive proof that a prediction model will help patients. Their argument, published as a letter responding to a recently proposed end-to-end framework for translating AI and machine-learning systems into cancer research and care, is deliberately provocative in its simplicity. Decision-curve analysis, they contend, is a bridge to clinical evaluation, not an endpoint in itself, and confusing the two risks sending poorly vetted algorithms toward hospital wards where the stakes include invasive biopsies, escalated treatment, and unnecessary surveillance.</p>
<p>The letter is a direct response to work by Saha and colleagues, who laid out a comprehensive translational framework covering data harmonization, explainability, federated learning, algorithmic bias, model drift, and regulatory readiness. The three correspondents are careful to praise much of that vision, and they specifically endorse the framework&#8217;s embrace of decision-curve analysis as a way to move model evaluation beyond the narrow metric of discrimination, which measures only how well a model separates patients who experience an outcome from those who do not. Discrimination, long the headline number in prediction studies, says nothing about what clinicians should actually do with a probability estimate. Decision-curve analysis was designed to fill that gap by estimating a quantity called net benefit across a range of threshold probabilities, effectively asking how many true positives a model captures at a given willingness to act, minus the harm caused by false positives weighted by their relative cost. It is an elegant piece of mathematics, and its adoption has been largely welcomed by methodologists who spent years warning that area under the receiver operating characteristic curve was being overinterpreted.</p>
<p>But elegance in a formula is not the same as validity at the bedside, and this is where the new letter sharpens its critique. Net benefit, the authors explain, is calculated across assumed threshold probabilities, which means it depends on three fragile conditions holding simultaneously. First, the model&#8217;s predicted risks must be sufficiently calibrated, meaning that when a model says a patient has a thirty percent chance of recurrence, roughly thirty percent of similar patients should actually recur. Second, the thresholds across which the curve is drawn must reflect real clinical choices, not arbitrary mathematical ranges. Third, the predictions must be available at the intended point of care, using variables that exist when the decision must actually be made. A model can display a favorable net-benefit curve despite poor calibration in a new population, unstable performance in clinically important subgroups, or reliance on variables that simply are not available when the oncologist needs an answer. The statistic, in other words, inherits every weakness of the model and every assumption of the analyst, and then presents the result as a single smooth curve that looks reassuringly quantitative.</p>
<p>The problem deepens when the threshold range itself is examined critically. A mathematically plausible sweep of threshold probabilities may include values that are clinically unrealistic, because they fail to reflect the genuine consequences of false-positive and false-negative decisions, the downstream capacity for confirmatory testing, patient preferences, or the existence of competing treatment options. In cancer care, where a flagged prediction might trigger a needle biopsy, additional imaging, treatment escalation, or years of intensified surveillance, each of which carries its own benefit and harm, the choice of threshold is a clinical and ethical judgment, not a plotting convention. The authors argue that decision-curve analysis should therefore be interpreted only after the target population, the prediction time point, the comparator strategies, the outcome horizon, and the clinically defensible thresholds have all been prespecified. Without that discipline, the curve becomes an exercise in retrospective storytelling, with analysts free to choose the window in which their model happens to look best.</p>
<p>What the correspondents propose in place of a single utility benchmark is a staged clinical-evidence pathway, a ladder that a cancer AI model must climb before its decision curve can be taken seriously. The first rung is analytical and predictive validity, established through transparent reporting, calibration assessment, internal validation, and external validation on datasets that are geographically and temporally distinct from the development data. This stage should follow TRIPOD + AI, the updated guidance for reporting clinical prediction models that use regression or machine-learning methods, published in the BMJ by Collins and colleagues in 2024. The second rung is where decision-curve analysis properly belongs: estimating potential net benefit against realistic alternatives, with uncertainty quantification and sensitivity analyses across prespecified thresholds rather than post hoc selections. Only after those two stages pass does the model earn the right to face the messier tests of real clinical environments.</p>
<p>The third rung takes the model out of the retrospective dataset entirely and into silent prospective validation, running within its intended clinical workflow without influencing care. This shadow phase exists to expose problems that historical data cannot reproduce: data latency, missing fields, distribution shift as the patient population changes, and plain implementation failures such as incompatible record systems or delayed result transmission. Many models that performed beautifully on archived data have stumbled here, and the letter&#8217;s authors argue that no net-benefit curve should be cited as evidence of clinical utility until these operational hazards have been observed and addressed. It is a humble but essential checkpoint, because the conditions under which a model was trained are never quite the conditions under which it must perform.</p>
<p>Fourth comes live early-stage evaluation, and this is where the correspondence widens its lens from the algorithm to the humans around it. What matters at this stage is not merely algorithmic accuracy but human-AI interaction: who receives the model&#8217;s output, how it changes decisions, whether clinicians can recognize incorrect recommendations, and whether automation bias, alert fatigue, or inequitable override patterns begin to emerge. Automation bias, the well-documented tendency of human experts to defer to machine suggestions even when the machine is wrong, is a particular hazard in high-volume oncology settings where time pressure makes a second opinion costly. The authors point to DECIDE-AI, the reporting guideline for early-stage clinical evaluation of AI-driven decision support systems, as the appropriate framework for this phase of testing. The fifth rung then asks the question that ultimately matters most: does AI-assisted care actually improve patient-relevant or process outcomes compared with current practice? Ideally this is tested with randomized or otherwise rigorous prospective designs, with CONSORT-AI, the extension of the CONSORT reporting guidelines for interventions involving artificial intelligence, guiding how such trials are reported.</p>
<p>The sixth and final rung extends beyond deployment day, which the authors insist is not the end of evaluation but the beginning of a monitoring obligation. Post-deployment surveillance should track subgroup performance, calibration drift, safety events, cybersecurity threats, and user behavior, against predefined criteria for recalibration, suspension, or retirement of the model. These lifecycle requirements align with the FUTURE-AI consensus, the international guideline for trustworthy and deployable healthcare AI published in the BMJ in 2025 by Lekadir and colleagues. The underlying philosophy is that clinical AI is not a static artifact that can be certified once and forgotten; it is a sociotechnical system whose performance erodes as populations, practices, and data streams evolve. A model that was safe and beneficial in 2026 may drift quietly into harm well before anyone notices, unless monitoring is built into the deployment from the start.</p>
<p>The oncology context gives this argument its urgency. Prediction models in cancer care may trigger invasive biopsies, additional imaging, treatment escalation, or surveillance regimens that create both benefit and harm, often in the same patient. Clinical utility cannot be inferred from a positive net-benefit curve alone, the authors write; it is demonstrated only when a well-calibrated, transportable model is usable by its intended clinicians, improves decisions relative to a credible comparator, benefits patients, and remains safe after deployment. Each clause in that sentence carries its own burden of proof, and none of them can be discharged by a statistical plot drawn from retrospective data. The letter does not diminish decision-curve analysis; rather, it relocates the tool to the rung where it belongs, as a valuable bridge between mathematical evaluation and clinical evaluation, and a strictly necessary but far from sufficient condition for trust.</p>
<p>The correspondence closes with an invitation rather than a rebuke. The framework proposed by Saha and colleagues, the authors note, is well positioned to incorporate this evidence ladder, and doing so would preserve decision-curve analysis as a valuable translational tool while clarifying that clinical utility is a property of the complete sociotechnical intervention, not the prediction model in isolation. As hospitals worldwide prepare to embed machine-learning systems into cancer diagnosis, staging, and treatment planning, the letter offers a sober benchmark for the field: the question is no longer whether a model produces an impressive curve, but whether the entire chain of evidence, from transparent reporting through silent validation, human interaction studies, comparative trials, and post-deployment monitoring, has been completed before the algorithm touches a patient&#8217;s care. In a discipline where the cost of a wrong prediction is measured in biopsies, toxic therapies, and missed cancers, the authors argue, that full chain is the only honest definition of utility.</p>
<p><strong>Subject of Research:</strong> The role of decision-curve analysis in evaluating clinical artificial intelligence for oncology</p>
<p><strong>Article Title:</strong> Decision-curve analysis should be a bridge, not an endpoint, for clinical artificial intelligence in oncology</p>
<p><strong>Article References:</strong> Khalid, M., Sarfraz, M., &amp; Javad, F. (2026). Decision-curve analysis should be a bridge, not an endpoint, for clinical artificial intelligence in oncology. <em>Journal of Translational Medicine, 24</em>(1), Article 1202. <a href="https://doi.org/10.1186/s12967-026-08775-x" rel="noopener noreferrer">https://doi.org/10.1186/s12967-026-08775-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12967-026-08775-x" rel="noopener noreferrer">10.1186/s12967-026-08775-x</a></p>
<p><strong>Keywords:</strong> decision-curve analysis, clinical artificial intelligence, oncology, net benefit, model calibration, TRIPOD+AI, DECIDE-AI, CONSORT-AI, FUTURE-AI, prospective validation, machine learning, translational medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">206487</post-id>	</item>
	</channel>
</rss>
