<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>uncertainty calibration &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/uncertainty-calibration/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 14:55:33 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>uncertainty calibration &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Check Itself: New Framework Makes Language Models Honest About Their Own Confidence</title>
		<link>https://scienmag.com/ai-learns-to-check-itself-new-framework-makes-language-models-honest-about-their-own-confidence/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 14:55:33 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[addressing AI hallucinations]]></category>
		<category><![CDATA[AI confidence calibration]]></category>
		<category><![CDATA[AI trustworthiness]]></category>
		<category><![CDATA[CLAIM-CAL model for fact verification]]></category>
		<category><![CDATA[confidence estimation]]></category>
		<category><![CDATA[decomposing AI responses into atomic claims]]></category>
		<category><![CDATA[enhancing reliability of AI-generated information]]></category>
		<category><![CDATA[Expected Calibration Error]]></category>
		<category><![CDATA[hallucination detection]]></category>
		<category><![CDATA[handling mixed accuracy in AI outputs]]></category>
		<category><![CDATA[improving AI honesty and transparency]]></category>
		<category><![CDATA[independent claim verification in AI systems]]></category>
		<category><![CDATA[isotonic regression]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[reliability estimation]]></category>
		<category><![CDATA[selective answering]]></category>
		<category><![CDATA[self-assessment mechanisms for AI accuracy]]></category>
		<category><![CDATA[self-checking AI frameworks]]></category>
		<category><![CDATA[self-verification]]></category>
		<category><![CDATA[truthfulness in large language models]]></category>
		<category><![CDATA[TruthfulQA]]></category>
		<category><![CDATA[uncertainty calibration]]></category>
		<category><![CDATA[verifying factual claims in language models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=195575</guid>

					<description><![CDATA[A new claim-level self-verification framework called CLAIM-CAL decomposes AI answers into atomic factual claims and applies isotonic calibration, cutting expected calibration error from 0.212 to 0.038 on TruthfulQA while enabling reliable selective answering.]]></description>
										<content:encoded><![CDATA[<p>Large language models have a confidence problem. They can produce fluent, polished, persuasive answers that are simply wrong, and they deliver those answers with the same commanding tone they use when they are right. For the millions of people now relying on AI systems for factual information, this mismatch between fluency and truthfulness has become one of the field&#8217;s most pressing unsolved challenges. A new study published in Discover Artificial Intelligence proposes a deceptively simple remedy: make the model break its own answer into tiny pieces, interrogate each piece separately, and then honestly recalculate how sure it should be.</p>
<p>The framework, called CLAIM-CAL, was developed by Abhigyan Pal, an independent researcher based in New Delhi, India. Rather than treating a model&#8217;s entire response as a single object to be trusted or doubted, CLAIM-CAL decomposes each answer into atomic factual claims, the smallest independently checkable statements an answer contains. A response such as &#8216;The Eiffel Tower is in Paris and was completed in 1889&#8217; is split into two distinct claims, each of which is then verified on its own merits. This granular approach reflects a core insight of the research: a single answer can be a mixture of accurate and inaccurate statements, and assigning one blanket confidence score to the whole response obscures exactly the information users need most.</p>
<p>The technical pipeline works in five stages. First, an answer model generates an initial response to a question without any external retrieval, deliberately isolating the contribution of self-verification. Second, gpt-4o-mini operating at temperature 0.0 extracts atomic claims using a structured JSON-format prompt, capping extraction at eight claims per answer with a sentence-splitting fallback if extraction fails. Third, each claim is interrogated by three differently framed verification probes: a balanced probe, a support-seeking probe, and a contradiction-seeking probe. Each probe labels the claim as Supported, Contradicted, or Unknown, and the deliberately adversarial framing of the contradiction probe is central to the design, because actively hunting for disconfirming evidence catches confidently wrong statements that a purely supportive check might wave through.</p>
<p>Fourth, the verdicts are converted into risk scores using a transparent hand-set heuristic. Contradiction carries the heaviest penalty because direct evidence against a claim is the strongest signal of unreliability; Unknown claims receive an intermediate penalty because they cannot be confidently verified; and lack of explicit support adds a smaller but nonzero penalty. The claim-level risks are averaged to produce an answer-level risk, which is subtracted from one to yield a raw confidence score. Crucially, the study does not stop there. The fifth stage applies isotonic regression, a non-parametric post-hoc calibration technique that learns a monotonic mapping from raw confidence scores to empirical correctness using a dedicated calibration split of sixty examples. Without this final step, even the sophisticated claim-level signal remains systematically overconfident.</p>
<p>The evaluation used the TruthfulQA generation benchmark, a dataset specifically engineered to expose cases where language models reproduce common human misconceptions. The experiment sampled 200 examples, reserving 60 for calibration and 140 for held-out testing. CLAIM-CAL was benchmarked against four alternatives: direct answering with assumed full confidence, verbal confidence where the model self-reports its certainty, self-consistency based on agreement across five independently sampled answers, and simple self-verification that judges the whole answer at once. On the held-out split, CLAIM-CAL achieved the highest observed accuracy of 0.757 among the evaluated methods, but the more striking finding concerned calibration quality rather than raw accuracy.</p>
<p>Raw CLAIM-CAL, like every other uncalibrated method, remained markedly overconfident, with an Expected Calibration Error of 0.212. ECE measures the bin-weighted gap between what a system claims to know and what it actually gets right. After isotonic calibration, that figure collapsed to 0.038, with a 95 percent bootstrap confidence interval of 0.015 to 0.106, while accuracy held steady. To rule out the possibility that calibration alone explained the advantage, the study then applied the same isotonic procedure to the variable-confidence baselines on the identical calibration split. In this fair head-to-head comparison, the calibrated claim-level method retained the strongest point estimates for both ECE and Brier Score, although the author is careful to note that paired bootstrap intervals overlapping zero mean the ECE advantage over the closest calibrated baselines should be interpreted cautiously rather than as statistically decisive superiority.</p>
<p>Where the approach truly shines is in selective answering, the practical scenario in which a deployed system must decide when to answer and when to defer. At a fixed 0.7 confidence threshold, CLAIM-CAL plus calibration achieved selective accuracy of 0.824 while retaining 0.893 coverage, meaning the system answered nearly 90 percent of questions while being right more than 82 percent of the time on the ones it chose to answer. Self-consistency, by contrast, reached comparable selective accuracy only at a coverage of 0.207, essentially answering so few questions that its apparent reliability becomes operationally useless. After calibration, self-consistency&#8217;s confidence values never exceeded 0.667, leaving it with zero coverage at the threshold entirely. This coverage-versus-accuracy trade-off is a point the study emphasizes repeatedly: a method can look trustworthy simply by refusing to engage.</p>
<p>The paper is notable as much for its methodological candor as for its results. Because the same model family handled answer generation, claim verification, and correctness judging, correlated errors could inflate apparent performance, a confound the author explicitly flags. A second-pass robustness check using an independent stricter judge prompt on 50 sampled evaluations achieved 96 percent agreement and a Cohen&#8217;s kappa of 0.896, which is strong consistency but not human validation. An ablation study confirmed that removing contradiction probing produced the worst calibration error among the variants, underscoring the value of adversarial verification, while removing claim decomposition lowered accuracy and weakened the reliability signal. The error analysis also showed the method&#8217;s honest limits: twenty-two incorrect answers still slipped through above the 0.7 confidence threshold, proving that CLAIM-CAL improves calibration and risk awareness without guaranteeing correctness.</p>
<p>The implications reach well beyond a single benchmark. The framework is designed to be complementary to prompt engineering and retrieval-augmented generation: better prompts can reduce initial errors, retrieved evidence can ground claims externally, and CLAIM-CAL can sit on top of either, estimating whether the final answer deserves user trust. A natural next step is verifying extracted claims directly against retrieved passages, converting the self-verification layer into a retrieval-grounded reliability check. Future work identified in the study includes scaling beyond 200 examples, testing multiple model families for generator, verifier, and judge roles, learning the risk weights from data rather than fixing them by hand, and integrating human review for low-confidence answers. The deeper conclusion is a reframing of what reliability means for artificial intelligence: a system should not be judged only by whether its answers are correct, but by whether its expressed confidence honestly reflects the probability of correctness. In an era when fluent language is too easily mistaken for truth, teaching models to know what they do not know may prove as important as teaching them what they do.</p>
<p><strong>Subject of Research:</strong> Claim-level self-verification and uncertainty calibration for improving the reliability of large language models</p>
<p><strong>Article Title:</strong> Improving reliability of large language models via claim-level self-verification and uncertainty calibration</p>
<p><strong>Article References:</strong> Pal, A. (2026). Improving reliability of large language models via claim-level self-verification and uncertainty calibration. <em>Discover Artificial Intelligence, 6</em>(1), Article 1132. <a href="https://doi.org/10.1007/s44163-026-02240-w" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02240-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02240-w" rel="noopener noreferrer">10.1007/s44163-026-02240-w</a></p>
<p><strong>Keywords:</strong> large language models, uncertainty calibration, self-verification, hallucination detection, TruthfulQA, isotonic regression, selective answering, Expected Calibration Error, reliability estimation, natural language processing, AI trustworthiness, confidence estimation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">195575</post-id>	</item>
		<item>
		<title>New AI Method Fuses Expert Opinions to Map Lung Vessels With Calibrated Confidence</title>
		<link>https://scienmag.com/new-ai-method-fuses-expert-opinions-to-map-lung-vessels-with-calibrated-confidence/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 04:42:50 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI methods for pulmonary research]]></category>
		<category><![CDATA[AI-driven medical image segmentation]]></category>
		<category><![CDATA[automated digital histology quantification]]></category>
		<category><![CDATA[computational pathology]]></category>
		<category><![CDATA[deep learning ensembles]]></category>
		<category><![CDATA[deep learning for lung tissue analysis]]></category>
		<category><![CDATA[ensemble neural network models in pathology]]></category>
		<category><![CDATA[ensemble segmentation]]></category>
		<category><![CDATA[expert opinion fusion in medical imaging]]></category>
		<category><![CDATA[histological vessel segmentation]]></category>
		<category><![CDATA[lung vessel segmentation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for vascular remodeling]]></category>
		<category><![CDATA[medical image analysis]]></category>
		<category><![CDATA[posterior fusion]]></category>
		<category><![CDATA[pulmonary hypertension]]></category>
		<category><![CDATA[pulmonary hypertension vessel analysis]]></category>
		<category><![CDATA[reliability calibration in medical AI]]></category>
		<category><![CDATA[ReliFuse]]></category>
		<category><![CDATA[scalable pulmonary disease assessment tools]]></category>
		<category><![CDATA[segmentation]]></category>
		<category><![CDATA[uncertainty calibration]]></category>
		<category><![CDATA[vessel mapping in diseased lungs]]></category>
		<category><![CDATA[vessel remodeling]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193766</guid>

					<description><![CDATA[Researchers have developed ReliFuse, a machine learning framework that fuses cached predictions from multiple segmentation experts to segment lung vessels in histology images with calibrated reliability and state-of-the-art overlap.]]></description>
										<content:encoded><![CDATA[<p>Quantifying how blood vessels remodel in diseased lungs has long been one of the most tedious bottlenecks in pulmonary research. Pathologists studying vascular changes associated with pulmonary hypertension must trace and outline vessel after vessel under a microscope, converting stained tissue sections into precise digital measurements. The work is slow, expert-dependent and difficult to scale, yet the numbers it produces underpin how researchers judge disease severity and treatment response. A team at the University of Science, Ho Chi Minh City, working with Vietnam National University, has now introduced a machine learning framework designed to automate this labor without sacrificing the reliability that clinical quantification demands.</p>
<p>The new method, called ReliFuse, is described in the journal Machine Learning and addresses a familiar irony in modern medical image analysis. Deep neural networks have become remarkably good at segmenting anatomical structures from histological images, producing masks that can rival human annotations. However, no single network is perfect, and the errors that individual models make are often complementary: one expert model may miss a faint peripheral vessel that another catches, while the second mislabels a fold of tissue that the first correctly ignores. Rather than treating these disagreements as noise, the Vietnamese team treats them as information, formulating the segmentation task as a problem of posterior fusion, in which multiple frozen expert models pool their predictions into a single, better-calibrated output.</p>
<p>What distinguishes ReliFuse from many ensemble techniques is a striking design constraint. At the fusion stage, the framework never looks at the underlying color image at all. Instead, it operates purely on cached probability maps produced beforehand by a bank of seven independently trained segmentation experts. These probability maps encode, for every pixel, how strongly each expert believes that the pixel belongs to a vessel. The fusion head then constructs so-called ensemble-state features from this stack of opinions, describing where the experts agree, where they diverge, and how their confidence is distributed. Working in logit space rather than raw probabilities, the method pools the evidence from all experts, estimates how trustworthy each local expert opinion is, and applies corrections only where the ambiguity is genuinely high.</p>
<p>Reliability estimation is the conceptual heart of the framework. For each expert model, the researchers compute validation-anchored priors from the model&#8217;s behavior on held-out validation data, giving the fusion head a sense of each expert&#8217;s typical strengths and weaknesses before it ever sees a test case. These priors are combined into a calibrated consensus opinion that serves as the starting point for the final segmentation. Crucially, ReliFuse does not rewrite the whole map. Its residual correction branch is bounded and gated by an ambiguity field, so that confident agreement among experts is preserved unchanged while corrections are concentrated exclusively in the contested regions where experts disagree or where boundary transitions are uncertain. This consensus-preservation principle ensures that the fusion step can refine the output without corrupting regions where the ensemble is already correct.</p>
<p>Training the fusion head is itself a multi-objective undertaking. The researchers combine an overlap loss with a boundary loss, a calibration loss, a consensus-preservation loss and a sparse-correction penalty. The boundary term compares gradient magnitudes between the predicted mask and the annotation, making contour errors visible even when vessels occupy few pixels. The consensus term is deliberately asymmetric, using stop-gradient operators to prevent the model from pulling its prior toward its own output or from simply lowering the ambiguity gate to dodge penalties. The sparse penalty is applied to the gated correction actually added to the logits, discouraging the network from making dense modifications everywhere rather than surgical fixes in ambiguous places. The calibration term supervises both the pooled prior and the final posterior with a Brier-style error, keeping the system&#8217;s confidence honest.</p>
<p>On a publicly available dataset of rat lung histology images with expert-annotated vessel masks, ReliFuse achieved the highest primary overlap among all methods in a matched comparison that gave every learned fusion head the same seven-expert posterior stack. The gains over the strongest competing learned fusion heads, which include ensemble-from-multiple-annotations approaches such as D-LEMA and locally calibrated federated methods such as LC-Fed, are modest in raw Dice and IoU terms. The authors are candid about this. In a paired statistical analysis across the held-out batches, the differences against these strongest references were small and not statistically significant, and the team treats those rows as evidence about effect direction and magnitude rather than proof of broad superiority.</p>
<p>Where ReliFuse genuinely pulls ahead is in the conditions that stress fusion methods hardest. In stress tests isolating batches with high expert disagreement and high vessel content, the improvements were clearest, consistent with the framework&#8217;s design focus on ambiguity and minority evidence. The method also held its own on boundary quality: while P-MoLE recorded the best boundary F1 scores and D-LEMA led on distance metrics such as HD95, ReliFuse remained close on these contour measures, indicating that its overlap gains did not come at the cost of degraded vessel geometry. A calibration and morphology analysis showed no single method dominating every diagnostic, with LC-Fed best on calibration error and D-LEMA best on centerline overlap, but ReliFuse remained competitive across morphology measures while producing the strongest primary Dice and IoU in the matched benchmark.</p>
<p>The practical economics of the approach are part of its appeal. Because the experts run only once and their probability maps are cached, the fusion stage is dramatically cheaper than re-running full segmentation networks. In the researchers&#8217; profiling experiments, recomputing the seven raw-image experts required roughly 9,357 milliseconds per batch of four images and more than 13.4 gigabytes of peak memory, far exceeding the cost of any cached-fusion pass. ReliFuse is slower than naive averaging because it must construct its diagnostic state, estimate calibrated opinions and apply gated corrections, but its parameter count remains modest relative to the base experts, and the framework is designed for settings where multiple frozen models are already available from prior development work.</p>
<p>The study is also notable for its methodological transparency. The authors report a full sensitivity analysis of how the expert bank is constructed, showing that a diversity-aware selection of experts improved every matched fusion rule compared with simply choosing the seven highest-scoring models. Ablation studies confirmed that the ambiguity gate, calibration supervision, boundary emphasis and consensus-preservation terms each contribute to the framework&#8217;s behavior, and the team documents a failure case in which the same gate that recovers a faint vessel can enlarge a false-positive region when the posterior evidence is misleading. By reporting the complete hard-subset stress matrices and labeling exploratory statistics as such, the researchers offer a template for honest evaluation in the crowded field of medical image segmentation.</p>
<p>For the researchers who need these measurements, the implications are concrete. The dataset underlying the work, published by Sinitca and colleagues in Scientific Data in 2024, contains 609 paired microphotographs and binary masks from rat models of pulmonary hypertension, split here into 517 development images and 92 held-out test images. The ReliFuse source code is publicly available on GitHub, and the framework requires no retraining of the underlying expert networks, only the lightweight fusion head. As quantitative histology moves from hand-tracing toward automated pipelines, ReliFuse suggests a pragmatic middle path: rather than chasing ever-larger single models, laboratories can combine the complementary strengths of the models they already have, and let a calibrated arbiter decide, pixel by pixel, whose opinion to trust.</p>
<p>The intellectual lineage of this approach stretches back several decades. Stacked generalization, introduced by Wolpert in 1992, established the idea of training a secondary learner to combine the outputs of base models, and Dietterich&#8217;s foundational work on ensemble methods later explained why combining diverse classifiers so often outperforms any single member. ReliFuse adapts this classical principle to dense prediction, where every pixel rather than every sample must receive a fused verdict, and where the cost of naively rerunning large networks makes caching an attractive design choice.</p>
<p>The framework also draws on a well-developed literature concerning the overconfidence of modern neural networks. Guo and colleagues demonstrated in 2017 that contemporary deep classifiers frequently produce probabilities that are poorly calibrated with respect to true correctness, a concern that is especially acute in medical settings where downstream decisions may hinge on a confidence value. Related work by Lakshminarayanan and colleagues on deep ensembles showed that simply averaging independently trained networks yields surprisingly strong uncertainty estimates, which helps explain why the reliability priors anchored on validation behavior prove so informative in the fusion stage.</p>
<p>Vessel segmentation itself has a long algorithmic history predating deep learning. Multiscale vessel-enhancement filtering, pioneered by Frangi and colleagues in 1998, remains influential, and subsequent surveys catalogued the breadth of methods, datasets and evaluation metrics used to judge tubular structure extraction. The pulmonary histology setting adds distinctive challenges, including stained tissue texture, irregular vessel branching and a pronounced class imbalance favoring background pixels, which is why overlap measures such as Dice and IoU are typically complemented by boundary and centerline diagnostics in this domain.</p>
<p>By situating posterior fusion within these established traditions, the study connects classical ensemble theory, calibration research and vascular image analysis into a single practical pipeline for quantitative histopathology.</p>
<p><strong>Subject of Research:</strong> A reliability-calibrated machine learning framework that fuses multiple segmentation experts&#x27; probability maps to automate histological pulmonary vessel segmentation.</p>
<p><strong>Article Title:</strong> ReliFuse: Reliability-Calibrated Posterior Fusion for Histological Vessel Segmentation</p>
<p><strong>Article References:</strong> Le, T. P., Nguyen, T. N., Tran, V. L. H., Doan, T. T., Nguyen, B. T., &amp; Huynh, S. T. (2026). ReliFuse: Reliability-Calibrated Posterior Fusion for Histological Vessel Segmentation. <em>Machine Learning, 115</em>(9), Article 218. <a href="https://doi.org/10.1007/s10994-026-07154-3" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07154-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07154-3" rel="noopener noreferrer">10.1007/s10994-026-07154-3</a></p>
<p><strong>Keywords:</strong> machine learning, ReliFuse, histological vessel segmentation, posterior fusion, pulmonary hypertension, medical image analysis, deep learning ensembles, uncertainty calibration, ensemble segmentation, computational pathology, vessel remodeling, segmentation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193766</post-id>	</item>
	</channel>
</rss>
