<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>decision thresholds &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/decision-thresholds/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 20:51:41 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>decision thresholds &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>GPT-4o Ranks Freight Cancellations Well but Fails the Cost Test</title>
		<link>https://scienmag.com/gpt-4o-ranks-freight-cancellations-well-but-fails-the-cost-test/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 20:51:41 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[API costs in predictive models]]></category>
		<category><![CDATA[architecture audit]]></category>
		<category><![CDATA[asymmetric error costs in freight logistics]]></category>
		<category><![CDATA[calibration]]></category>
		<category><![CDATA[cost-sensitive learning]]></category>
		<category><![CDATA[cost-sensitive machine learning in logistics]]></category>
		<category><![CDATA[decision thresholds]]></category>
		<category><![CDATA[freight cancellation]]></category>
		<category><![CDATA[freight cancellation risk prediction]]></category>
		<category><![CDATA[freight shipment intervention costs]]></category>
		<category><![CDATA[GPT-4o]]></category>
		<category><![CDATA[GPT-4o shipping risk ranking]]></category>
		<category><![CDATA[impact of false alarms in freight operations]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[logistic regression]]></category>
		<category><![CDATA[logistics optimization with AI]]></category>
		<category><![CDATA[machine learning deployment]]></category>
		<category><![CDATA[machine learning model evaluation in supply chain]]></category>
		<category><![CDATA[online retailer freight management]]></category>
		<category><![CDATA[real-world freight cancellation prediction]]></category>
		<category><![CDATA[shipping cancellation decision thresholds]]></category>
		<category><![CDATA[supply chain]]></category>
		<category><![CDATA[tabular prediction]]></category>
		<category><![CDATA[XGBoost]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198524</guid>

					<description><![CDATA[A new study of 354,565 freight shipments shows that prompted GPT-4o ranks cancellation risk as well as boosted trees yet never crosses the firm's cost-implied decision threshold, matching a free always-intervene baseline while incurring API costs.]]></description>
										<content:encoded><![CDATA[<p>A large language model can rank shipping-cancelation risks just as well as specialized statistical tools, yet still lose money the moment a company acts on its predictions. That paradox sits at the heart of a new study in Machine Learning with Applications by Wenyi Kuang, who builds a cost-sensitive architecture audit and deploys it against a real-world freight cancelation problem at a large U.S. online retailer. The research examines 354,565 truckload shipments and finds that prompted GPT-4o, despite achieving a competitive area under the ROC curve of 0.810, flags every single shipment for intervention because its lowest predicted probability never falls below the firm&#8217;s cost-implied decision threshold of 0.091. The result is an action policy identical to the crude strategy of always intervening, plus an extra $1.80 per thousand shipments in API costs.</p>
<p>The underlying business problem is one of asymmetric error costs. Every morning, freight planners must decide which scheduled shipments warrant pre-execution intervention, such as reconfirming the carrier, staging backup capacity, splitting the load, or warning downstream customers. A false alarm, intervening on a shipment that would never have been canceled, costs roughly $35 in coordination effort. A missed cancelation triggers last-minute re-tendering, delays, and service disruption at roughly $350 per shipment. That 10-to-1 asymmetry implies a mathematically optimal decision threshold of p* = 0.091: any shipment whose predicted cancelation probability exceeds 0.091 should be flagged, and anything below it should be released. Classical decision theory going back to Elkan&#8217;s foundational work on cost-sensitive learning prescribes exactly this kind of threshold, but the study shows that whether a model can actually use such a threshold depends on a property almost nobody measures: the range of probabilities the model is capable of producing.</p>
<p>Kuang formalizes this insight through four diagnostics that together decompose an architecture&#8217;s deployment value into four separable components: the information it consumes, the support and calibration of its predicted probabilities, the deployment policy, and the marginal inference cost. The first diagnostic asks whether competing architectures rank instances equally well when given identical features. The second, and the sharpest, detects a prediction-threshold mismatch: when a model&#8217;s score distribution lies entirely on one side of the decision threshold, the model degenerates into a constant-action rule no matter how well it ranks. The third evaluates routing architectures, testing whether selectively escalating hard cases to an expensive model beats a cheap deterministic fallback. The fourth probes behavior under a fixed intervention budget, where ties in the score distribution can silently destroy precision.</p>
<p>The empirical comparison was designed with unusual rigor. All architectures, including the language model, were restricted to the same 23-field audit record: ten categorical fields such as origin, destination, and carrier code, five numeric attributes such as mileage and planned transit hours, and eight engineered empirical-Bayes shrinkage priors estimated from historical cancelation rates by carrier and lane. The tabular learners, logistic regression, XGBoost, LightGBM, and an equal-weight ensemble, were trained on 266,915 shipments, calibrated on a separate 3,914-shipment fold, and scored once on a held-out set of 15,000 shipments. GPT-4o was never trained on any shipment; it was a frozen pretrained model queried once per shipment with the same features serialized into a prompt demanding a JSON cancelation probability. On ranking, the language model was statistically indistinguishable from the boosted trees, with paired bootstrap confidence intervals placing its AUC gaps to XGBoost (+0.002) and LightGBM (-0.002) inside a pre-specified equivalence band of plus or minus 0.02.</p>
<p>Ranking parity, however, masked what the author calls action collapse. GPT-4o&#8217;s predicted probabilities took only 96 distinct values across 15,000 shipments, and the lowest was 0.150, well above the operating threshold of 0.091. Consequently the threshold rule flagged 100 percent of shipments, matching the always-intervene baseline pointwise. Once the $1.80-per-thousand API cost was added, the language model architecture underperformed the free always-flag rule by about $2 per thousand shipments. The same-information learners avoided this trap because their continuous score distributions straddled the threshold. A raw-field logistic regression flagged only 91 percent of shipments and saved $394 per thousand against the baseline, the largest margin in the study, followed by the full-feature logistic model, the ensemble, the priors-only logistic model, and LightGBM.</p>
<p>Post-hoc calibration, the standard remedy for miscalibrated probabilities, did not rescue the language model. Isotonic regression improved the Brier score but left every probability above the threshold. Platt scaling and Beta calibration did push some scores below 0.091, yet they made operational cost worse, because the shipments they released included an outsized share of genuine cancelations whose $350 miss penalties overwhelmed the false-alarm savings. The study&#8217;s cost-ratio sweep shows the collapse is regime-dependent rather than universal: whenever the false-negative-to-false-positive ratio exceeds roughly 5.67 to 1, placing the threshold below the model&#8217;s 0.150 support floor, the always-flag degeneration is mechanically guaranteed. At lower ratios, such as 4:1, continuous-score learners still dominated, and across a 400-cell grid of cost pairs the ensemble and the raw-field logistic model split all winning cells, with the language model winning none.</p>
<p>An information ablation traced where the model&#8217;s apparent skill originated. An extraction regression of GPT-4o&#8217;s textual probabilities onto the prompt fields showed that raw numerics explained only 6.4 percent of output variance under cross-validation, while hashed categorical identifiers alone explained 87.5 percent and the eight engineered priors alone explained 88.2 percent, saturating at 89.2 percent with everything supplied. The language model was essentially re-expressing the firm&#8217;s own historical entity-level cancelation statistics back to it, not extracting novel signal from a pretrained world model. Stripping the priors from the prompt collapsed its discrimination to an AUC of 0.508, essentially chance.</p>
<p>The audit also exposed a subtle failure in popular routing architectures. A natural deployment routes sparse-history shipments, carriers appearing fewer than K times in the training data, to the language model while handling the rest with a calibrated tree. But because every routed shipment scored at or above 0.15, escalation and an unconditional flag-everything fallback produced identical actions on the routed subset at every value of K tested. Under any positive per-call API charge, the free deterministic rule strictly dominates, a result the paper proves formally and verifies empirically across the full routing grid. A capacity stress test added a further wrinkle: under a strict Top-500 intervention budget, the language model&#8217;s discrete score lattice forced 476 of the selected shipments to share the cutoff score, so precision was determined by the tie bin&#8217;s base cancellation rate of 97.1 percent rather than any within-bin ranking, while continuous-score architectures retained more decision-relevant resolution.</p>
<p>The central finding survived out-of-time validation. On a strictly chronological fold of 67,938 shipments from April to July 2021, split into four consecutive quartiles, the language model continued to rank competitively, trailing the raw-field logistic model by 0.006 to 0.019 AUC with all bootstrap intervals strictly positive, yet it again flagged 100 percent of shipments in every quartile and delivered no cost improvement over always-flag. A bounded robustness check with Anthropic&#8217;s Claude Sonnet under identical prompts reproduced the collapse, and a 28-cell prompt factorial across both providers found that no elicitation condition produced probabilities reaching zero, with 22 of 28 conditions flagging at least 99.4 percent of the cohort. The author is careful to scope the claim: the result does not show that language models cannot rank tabular data, only that at this cost ratio, with this probability elicitation interface and the observed score support, the model never changes the action the firm would take anyway. The ranking would likely flip under a lower cost ratio, a calibrated single-token log-probability interface, or richer unstructured inputs unavailable to the tabular baselines.</p>
<p>The transferable contribution is methodological rather than a verdict on any model class. Because deployment value depends jointly on information, score support, policy, and inference cost, a fair architecture comparison must audit whether a model&#8217;s scores actually cross the decision threshold and change actions, not merely whether it ranks well. In this freight application, a plain logistic regression on raw fields delivered the cheapest action policy, while an expensive frontier model reproduced a free baseline at a loss, a cautionary tale for any operations team tempted to equate benchmark scores with business value.</p>
<p><strong>Subject of Research:</strong> Cost-sensitive auditing of large language model architectures for tabular prediction in freight cancellation management.</p>
<p><strong>Article Title:</strong> COST-SENSITIVE ARCHITECTURE AUDITING FOR LLM-BASED TABULAR PREDICTION: EVIDENCE FROM FREIGHT CANCELLATION MANAGEMENT</p>
<p><strong>Article References:</strong> Kuang, W. (2026). Cost-sensitive architecture auditing for llm-based tabular prediction: Evidence from freight cancellation management. <em>Machine Learning with Applications, 25</em>, Article 100997. <a href="https://doi.org/10.1016/j.mlwa.2026.100997" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.100997</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.100997" rel="noopener noreferrer">10.1016/j.mlwa.2026.100997</a></p>
<p><strong>Keywords:</strong> large language models, GPT-4o, tabular prediction, cost-sensitive learning, freight cancellation, XGBoost, calibration, decision thresholds, logistic regression, architecture audit, machine learning deployment, supply chain</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198524</post-id>	</item>
		<item>
		<title>Entropy May Not Be the Fix Medicine Needs for Clinical Uncertainty</title>
		<link>https://scienmag.com/entropy-may-not-be-the-fix-medicine-needs-for-clinical-uncertainty/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 01:36:43 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[artificial intelligence and uncertainty quantification]]></category>
		<category><![CDATA[artificial intelligence in healthcare]]></category>
		<category><![CDATA[Bayesian inference]]></category>
		<category><![CDATA[challenges of applying thermodynamics to clinical practice]]></category>
		<category><![CDATA[clinical decision-making]]></category>
		<category><![CDATA[clinical judgment]]></category>
		<category><![CDATA[decision theory]]></category>
		<category><![CDATA[decision theory in medicine]]></category>
		<category><![CDATA[decision thresholds]]></category>
		<category><![CDATA[diagnostic uncertainty]]></category>
		<category><![CDATA[entropy]]></category>
		<category><![CDATA[entropy in medicine]]></category>
		<category><![CDATA[information theory in clinical reasoning]]></category>
		<category><![CDATA[internal medicine]]></category>
		<category><![CDATA[limitations of entropy for medical decisions]]></category>
		<category><![CDATA[medical decision-support tools]]></category>
		<category><![CDATA[Medical Education]]></category>
		<category><![CDATA[medical uncertainty]]></category>
		<category><![CDATA[quantitative measures of clinical uncertainty]]></category>
		<category><![CDATA[role of entropy in diagnosis]]></category>
		<category><![CDATA[uncertainty]]></category>
		<category><![CDATA[value of information]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193394</guid>

					<description><![CDATA[A letter in the Journal of General Internal Medicine warns that entropy-based measures of diagnostic uncertainty risk an illusion of precision and could undermine clinical judgment and medical education.]]></description>
										<content:encoded><![CDATA[<p>A concise letter published in the Journal of General Internal Medicine is igniting a debate that reaches far beyond its modest length. Written by Mucheli Sharavan Sadasiv and Minyang Chow of the Lee Kong Chian School of Medicine at Nanyang Technological University and the National Healthcare Group in Singapore, the correspondence takes aim at one of the more seductive ideas now circulating at the intersection of medicine, information theory, and artificial intelligence: the notion that entropy, a mathematical measure of uncertainty drawn from thermodynamics and information science, could serve as a unifying quantitative lens for clinical decision-making. The letter is a response to a narrative review by Rohlfsen and colleagues titled “Entropy in Clinical Decision-Making: A Narrative Review Through the Lens of Decision Theory,” and it argues that enthusiasm for the concept must be tempered by a fundamental mismatch between what entropy measures and what clinicians actually need in order to act.</p>
<p>The original review had presented entropy as a way to quantify uncertainty in medical reasoning, describing it as offering a concise summary of uncertainty that nonetheless lacks a built-in mechanism for action. That admission, the Singapore authors contend, is precisely where the trouble begins. In clinical practice, uncertainty is not merely a quantity to be measured; it is a condition to be navigated, weighed against risks, benefits, and patient values, and ultimately resolved into a decision: treat, test, observe, or reassure. A framework that summarizes uncertainty without specifying how to act on it, they argue, risks creating what they call an illusion of precision, presenting clinicians with a single descriptive number that feels rigorous but resists translation into a concrete clinical act.</p>
<p>The technical heart of the critique lies in a comparison between entropy and Bayesian inference, the dominant framework for reasoning under uncertainty in medicine and statistics. Bayesian models produce state-specific probabilities: the probability, for instance, that a patient with chest pain is having a myocardial infarction versus a benign cause. These actionable probabilities can then be compared against established decision thresholds, most famously formalized by Pauker and Kassirer in the New England Journal of Medicine in 1980. The threshold approach defines a testing threshold and a treatment threshold; if the probability of disease falls below the former, the clinician forgoes testing, and if it rises above the latter, treatment proceeds without further diagnostic workup. This architecture converts probability directly into action, providing a rational bridge between belief and behavior.</p>
<p>Entropy, by contrast, collapses an entire probability distribution into a single scalar. In information theory, the Shannon entropy of a diagnostic hypothesis set is maximal when all possibilities are equally likely and minimal when one diagnosis dominates. A high-entropy differential diagnosis tells the clinician that the situation is genuinely uncertain, but it does not say which diagnosis is most probable, what test would most efficiently reduce the uncertainty, or whether further investigation is even warranted given the stakes. Two patients could carry identical entropy values while demanding radically different management: one with a high-mortality condition hovering near a treatment threshold, the other with a trivial condition with little actionable consequence. The letter’s authors argue that this loss of state-specific information is not a minor technicality but an ontological mismatch between the descriptive reach of entropy and the prescriptive demands of clinical judgment.</p>
<p>The critique also engages with the literature on value of information, a family of methods for prioritizing research and testing by quantifying how much a new piece of information would be worth in terms of improved outcomes. Value of information analysis, as codified by Jackson and colleagues in Epidemiologic Methods in 2021, builds explicitly on decision-theoretic foundations, linking the acquisition of information to expected gains in health. Bayesian probability combined with threshold logic naturally accommodates these calculations: knowing a probability and the payoff matrix of actions allows one to compute the expected value of perfect or sample information. Entropy alone, stripped of state-specific probabilities and payoff structures, cannot perform this function. A clinician told that a case has an entropy of 1.7 bits has learned little about what to do next, whereas a clinician told that the probability of disease is 45 percent against a testing threshold of 30 percent knows immediately that more information is worth acquiring.</p>
<p>What makes the letter particularly provocative is its pivot from decision theory to pedagogy. The authors acknowledge that the original review rightly locates entropy’s true promise in standardization and scalability, especially for artificial intelligence systems trained on vast clinical datasets. In that context, entropy can serve as a useful computational statistic, a way for machine learning systems to flag cases of high diagnostic ambiguity, route them to specialists, or measure model confidence. But the authors warn that the vision of an “entropy-based medicine” must be weighed against its potential educational consequences. Medicine has long oscillated between the aspiration to quantify everything and the recognition that its core practice remains an interpretive, human activity. If trainees learn that good clinical reasoning means minimizing a calculated uncertainty value, the letter suggests, they may lose sight of a more important competency: the cultivated ability to tolerate uncertainty and still act responsibly.</p>
<p>That argument draws on a growing body of medical education scholarship, most prominently the 2016 New England Journal of Medicine perspective by Simpkin and Schwartzstein titled “Tolerating uncertainty — the next medical revolution?” That piece argued that discomfort with uncertainty drives a range of pathology in modern medicine, from excessive diagnostic testing and defensive medicine to communication failures and burnout. Uncertainty tolerance, far from being a soft skill, is framed as a professional capacity intimately linked to clinical judgment, effective patient communication, and patient safety. The Singapore authors build directly on this framing: an “entropy-minimization” mindset, they caution, could distract trainees from the deeper goal of becoming comfortable living with ambiguity. In a busy clinical environment, the temptation to chase a single number that promises clarity is strong, and a pedagogy built around minimizing entropy could reinforce precisely the reflexive, test-driven behavior that educators have spent years trying to moderate.</p>
<p>The debate also carries implications for how artificial intelligence tools will be explained and governed in medicine. As machine learning systems become embedded in triage, imaging interpretation, and predictive analytics, measures of model uncertainty such as entropy will increasingly be surfaced to clinicians, perhaps as confidence scores or risk flags. The letter’s warning suggests that how these numbers are taught, contextualized, and displayed will matter enormously. A confidence metric presented without a decision threshold or a treatment implication invites either blind deference or reflexive dismissal. Used well, however, uncertainty quantification can prompt exactly the right kind of reflection: a pause before acting on a low-confidence prediction, a request for a second opinion, or a conversation with the patient about the limits of what is known. The difference lies not in the mathematics but in the professional culture that surrounds it.</p>
<p>None of this amounts to a rejection of information theory in medicine. The letter is explicit in crediting the original review with a valuable service: introducing a complex concept to a general medical audience and sparking a necessary dialogue on the nature of clinical uncertainty. Its authors position their critique as a call for deeper conversation rather than a dismissal, insisting that before the profession embraces new quantitative tools, it must clarify their proper place in a practice that remains both a science and an art. The historical parallel is instructive. Bayesian reasoning took decades to move from statistical journals into bedside teaching, and only became genuinely useful to clinicians once it was paired with threshold frameworks, likelihood ratios, and pretest probability estimation. Entropy, if it follows a similar path, will need its own translation layer: ways of connecting a global uncertainty measure to the specific probabilities, stakes, and values that drive individual decisions.</p>
<p>For now, the Singapore letter stands as a compact but pointed intervention in one of the most consequential conversations in contemporary medicine: how a profession built on judgment should metabolize the quantitative machinery of the information age. Its message resonates well beyond internal medicine, touching any field wrestling with the promise of AI-assisted uncertainty quantification, from radiology to public health modeling. The core claim is deceptively simple. Measuring uncertainty is not the same as managing it, and a number that summarizes doubt without pointing toward action may, in the hands of an overburdened clinician or a trainee still forming professional habits, do more to obscure good judgment than to support it. As hospitals and developers race to embed uncertainty metrics in clinical workflows, this letter insists that the decisive questions are not computational but philosophical and pedagogical: what do we want clinicians to learn when we teach them to measure what they do not know?</p>
<p><strong>Subject of Research:</strong> The limitations of entropy as a quantitative measure of clinical uncertainty in medical decision-making, judgment, and education.</p>
<p><strong>Article Title:</strong> Beyond Entropy: Decision Thresholds, Judgment, and Pedagogy</p>
<p><strong>Article References:</strong> Sadasiv, M. S., &amp; Chow, M. (2026). Beyond Entropy: Decision Thresholds, Judgment, and Pedagogy. <em>Journal of General Internal Medicine</em>. <a href="https://doi.org/10.1007/s11606-026-10746-3" rel="noopener noreferrer">https://doi.org/10.1007/s11606-026-10746-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11606-026-10746-3" rel="noopener noreferrer">10.1007/s11606-026-10746-3</a></p>
<p><strong>Keywords:</strong> entropy, clinical decision-making, uncertainty, Bayesian inference, decision thresholds, medical education, clinical judgment, artificial intelligence, decision theory, value of information, diagnostic uncertainty, internal medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193394</post-id>	</item>
	</channel>
</rss>
