<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>posterior contraction &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/posterior-contraction/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 14:18:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>posterior contraction &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Bayesian Screening Method Hunts Hidden High-Risk Outliers in Count Data</title>
		<link>https://scienmag.com/new-bayesian-screening-method-hunts-hidden-high-risk-outliers-in-count-data/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 14:18:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Bayesian inference]]></category>
		<category><![CDATA[Bayesian screening for high-risk outliers in count data]]></category>
		<category><![CDATA[count data]]></category>
		<category><![CDATA[e-values]]></category>
		<category><![CDATA[exceedance probability]]></category>
		<category><![CDATA[exceedance probability estimation in statistical modeling]]></category>
		<category><![CDATA[false discovery rate]]></category>
		<category><![CDATA[health utilization]]></category>
		<category><![CDATA[Kullback–Leibler projection]]></category>
		<category><![CDATA[latent risk detection in healthcare and insurance claims]]></category>
		<category><![CDATA[multiple testing]]></category>
		<category><![CDATA[negative binomial regression]]></category>
		<category><![CDATA[open-access Bayesian approaches for data contamination detection]]></category>
		<category><![CDATA[outlier detection in count-based datasets]]></category>
		<category><![CDATA[posterior contraction]]></category>
		<category><![CDATA[posterior predictive distribution in Bayesian statistics]]></category>
		<category><![CDATA[probabilistic ranking of units based on Bayesian inference]]></category>
		<category><![CDATA[RAND Health Insurance Experiment]]></category>
		<category><![CDATA[robust methods for count data analysis]]></category>
		<category><![CDATA[robust statistics]]></category>
		<category><![CDATA[separating observed data from inferred risk]]></category>
		<category><![CDATA[statistical methods for identifying high-risk units]]></category>
		<category><![CDATA[tail functional analysis in count data]]></category>
		<category><![CDATA[threshold exceedance modeling in count data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223222</guid>

					<description><![CDATA[A new robust Bayesian framework separates latent tail-risk ranking from false-discovery-controlled screening in overdispersed count data, with applications to high-utilization patients in the RAND Health Insurance Experiment.]]></description>
										<content:encoded><![CDATA[<p>Statisticians have long faced a deceptively simple question: given a stream of counts — hospital visits, insurance claims, defect reports — which units truly carry elevated risk of exceeding a critical threshold, and which merely look extreme because of noise or contamination? A new study published in the International Journal of Data Science and Analytics by Abdolnasser Sadeghkhani of North Carolina Agricultural and Technical State University offers a carefully engineered answer. The work, an open-access regular paper published on 1 October 2026, develops a robust Bayesian procedure for what the author calls exceedance screening: ranking and selecting units according to their latent probability of producing a count above a scientifically meaningful cutoff, rather than according to the raw numbers they happen to display.</p>
<p>The central conceptual move is a sharp separation between what is observed and what is inferred. For each unit i, the method defines a parameter-conditional tail functional, the probability that a replicated count drawn from the fitted model would exceed a threshold, given that unit&#8217;s covariates. Averaging this functional over the posterior distribution of the model parameters yields the posterior predictive exceedance probability, the quantity used for ranking and calibration. A large realized count, the paper stresses, is informative but is not itself the inferential target: a person with moderate observed utilization may still carry high replicated-count risk once health status, insurance characteristics, and model uncertainty are accounted for. Conversely, an isolated spike in the data may reflect contamination rather than genuine tail risk.</p>
<p>The working model is a negative binomial regression, a standard choice for overdispersed counts because it relaxes the restrictive variance assumption of the Poisson model while retaining an interpretable conditional mean. The robustness, however, comes from an additional layer. Each observation receives a positive latent scale that locally inflates or deflates its mean, drawn from a two-component mixture: a point mass at one, which preserves the ordinary negative binomial baseline, and a Beta-prime slab with substantial mass near zero and a polynomially decaying right tail. This construction, adapted from recent robust Bayesian count modeling, allows excess zeros and isolated large counts to be absorbed by the local scale rather than forcing them to distort the regression coefficients or the fitted tail probabilities. In effect, unusual observations are quarantined instead of being allowed to hijack the analysis.</p>
<p>The theory is deliberately framed for the realistic case where the model is wrong. Rather than assuming the data follow the negative binomial family exactly, the contraction results are stated around the Kullback–Leibler projection of the true law onto the working family — a pseudo-true parameter that represents the best available approximation. Following the classical misspecification perspective of Berk and of Kleijn and van der Vaart, the posterior is shown to concentrate near this projection, and a Lipschitz transfer argument shows that concentration in parameters carries over to uniform concentration of the exceedance functionals themselves. Pinsker&#8217;s inequality then bounds the error between true tail probabilities and the working-model targets by the Kullback–Leibler discrepancy, giving a quantitative account of how much misspecification can distort screening decisions.</p>
<p>Discovery, as opposed to ranking, is handled by a separate layer built on e-values. Posterior thresholding of the evidence probability — the posterior probability that a unit&#8217;s exceedance probability exceeds a chosen level gamma — is a natural local Bayes rule under a loss function penalizing false positives and false negatives, but it does not by itself control the false discovery rate across many simultaneous tests. The paper therefore converts posterior odds into Bayes factors and submits them to the e-BH procedure of Wang and Ramdas, which controls the false discovery rate under arbitrary dependence among the test statistics. The validity of this step is stated with unusual care: under the Bayes-marginal null induced by the same hierarchical model, posterior odds are Bayes factors and hence valid e-values with null expectation at most one. Under a mere working-model interpretation, the same quantities are evidence scores unless calibrated, and the author includes an approximate-validity result in which the FDR bound is inflated by a calibration error term. This honesty avoids the common pitfall of treating a parametric Bayesian fit as automatically producing model-free frequentist guarantees.</p>
<p>Computation is provided at three levels of fidelity. A reference Markov chain Monte Carlo sampler uses Pólya–Gamma augmentation to update the regression coefficients in a single Gaussian block, with slice sampling and adaptive random-walk Metropolis steps for the latent scales, dispersion parameter, and slab hyperparameters; convergence is monitored with rank-normalized split R-hat statistics and effective sample size thresholds. For large datasets and repeated simulations, a bounded-influence iteratively reweighted least squares Laplace approximation downweights extreme Pearson residuals and high-leverage observations, producing a Gaussian approximate posterior from which the exceedance summaries can be drawn directly. A mean-field variational option offers fast screening. A theorem bounds the error of any such approximation: the discrepancy between exact and approximate posterior exceedance probabilities is controlled by the total variation distance between the two posteriors, which in turn is bounded by a square root of the Kullback–Leibler divergence between them.</p>
<p>The simulation evidence demonstrates a clear and quantified robustness–efficiency tradeoff. Under clean data, the ordinary negative binomial fit is slightly more efficient, with marginally higher true positive rates and smaller calibration error. Under contamination — engineered as a mixture of near-zero structural zeros and heavy upper-tail multiplicative contamination at rates of ten and twenty percent — the picture reverses dramatically. At ten percent contamination, the robust procedure cuts the false discovery proportion from 0.029 to 0.002 and halves the calibration error metrics. At twenty percent contamination and a sample size of one thousand, the naive method&#8217;s false discovery proportion climbs to roughly 0.099 while the robust method holds near 0.004, with expected calibration error of about 0.090 versus 0.040. A more aggressive variant using Tukey&#8217;s biweight weights proves most conservative of all, illustrating that the degree of robustness tunes a power–conservatism dial rather than dominating uniformly. Reliability plots confirm that the robust fit tracks the clean oracle tail probabilities across the range, while the naive fit inflates the upper bins under contamination.</p>
<p>The real-data application turns to the RAND Health Insurance Experiment, a classic dataset of 20,190 observations on outpatient physician visits, adjusted for insurance plan, income, physical limitation, chronic disease burden, and self-rated health. Screening for individuals whose replicated visit count exceeds eight visits, with a discovery level of 0.3 and e-BH applied at a false discovery rate of 0.1, both fits select subgroups whose observed exceedance rates — 0.339 for the naive fit and 0.315 for the robust fit — tower over the full-sample rate of 0.071. The robust fit selects a somewhat larger set while shrinking tail probabilities more aggressively over most of the sample, yet still assigns high risk to individuals with severe health profiles. A training–holdout split with 2,000 posterior predictive replicates showed the observed zero rate falling at the edge of its predictive interval, while the mean count and exceedance rate exceeded theirs — a candid indication that the working model remains conservative in the upper tail and not fully distributionally adequate.</p>
<p>Perhaps the most practically important feature of the paper is its insistence on reporting three distinct outputs separately: calibrated tail-risk rankings, posterior evidence scores, and multiplicity-controlled discovery sets. Some individuals selected by the procedure did not actually exceed eight visits in the observed record, which the author presents as expected behavior rather than a failure — the method targets replicated-count risk conditional on covariates and model uncertainty, not the realized indicator alone. The selected individuals are characterized by physical limitation and high chronic-disease burden, exactly the covariate profiles one would associate with latent high utilization. The paper also warns against a subtle computational trap: e-value caps that are too small can mechanically prevent the e-BH procedure from making any rejections, so decision steps must use uncapped Bayes factors or caps compatible with the rejection thresholds.</p>
<p>Taken together, the study offers a transparent template for combining Bayesian tail modeling, robust count regression, and e-value-based multiple testing in settings where the data are messy, overdispersed, and possibly contaminated. The author suggests that future work may extend the framework to richer dependence structures, more flexible count families, and sequential e-value methods for online monitoring — a natural next step for applications ranging from epidemiological surveillance to industrial quality control, where the question is not merely how many events occur, but who is genuinely at risk of crossing the line.</p>
<p><strong>Subject of Research:</strong> Robust Bayesian exceedance screening and false discovery rate control for overdispersed count data</p>
<p><strong>Article Title:</strong> Robust Bayesian exceedance screening for count data</p>
<p><strong>Article References:</strong> Sadeghkhani, A. (2026). Robust Bayesian exceedance screening for count data. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 323. <a href="https://doi.org/10.1007/s41060-026-01312-5" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01312-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01312-5" rel="noopener noreferrer">10.1007/s41060-026-01312-5</a></p>
<p><strong>Keywords:</strong> Bayesian inference, exceedance probability, e-values, negative binomial regression, robust statistics, false discovery rate, count data, Kullback–Leibler projection, RAND Health Insurance Experiment, multiple testing, posterior contraction, health utilization</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223222</post-id>	</item>
	</channel>
</rss>
