<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Bayesian models in machine learning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/bayesian-models-in-machine-learning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 06:43:55 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Bayesian models in machine learning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>One Outlier Can Break Gaussian Process Training, New Study Shows</title>
		<link>https://scienmag.com/one-outlier-can-break-gaussian-process-training-new-study-shows/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 06:43:55 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Bayesian inference]]></category>
		<category><![CDATA[Bayesian models in machine learning]]></category>
		<category><![CDATA[effects of data anomalies on Bayesian models]]></category>
		<category><![CDATA[Gaussian process fragility]]></category>
		<category><![CDATA[Gaussian process training vulnerabilities]]></category>
		<category><![CDATA[Gaussian processes]]></category>
		<category><![CDATA[Huber loss]]></category>
		<category><![CDATA[hyperparameter estimation]]></category>
		<category><![CDATA[hyperparameter sensitivity in Gaussian processes]]></category>
		<category><![CDATA[influence functions]]></category>
		<category><![CDATA[kernel function impact on model stability]]></category>
		<category><![CDATA[kernel methods]]></category>
		<category><![CDATA[M-estimation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning model reliability]]></category>
		<category><![CDATA[marginal likelihood]]></category>
		<category><![CDATA[marginal likelihood optimization issues]]></category>
		<category><![CDATA[mathematical limitations of Gaussian processes]]></category>
		<category><![CDATA[outlier influence on Gaussian process training]]></category>
		<category><![CDATA[outliers]]></category>
		<category><![CDATA[robust statistics]]></category>
		<category><![CDATA[robustness of Gaussian process models]]></category>
		<category><![CDATA[Student-t likelihood]]></category>
		<category><![CDATA[uncertainty quantification in Gaussian processes]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226238</guid>

					<description><![CDATA[A new theoretical study proves that standard Gaussian process hyperparameter estimation has an unbounded influence function, meaning a single outlier can arbitrarily distort kernel parameters, and identifies heavy-tailed likelihoods and curvature-preserving robust losses as the fix.]]></description>
										<content:encoded><![CDATA[<p>Gaussian processes have quietly become one of the most trusted tools in modern machine learning. From robotics to geostatistics to the emulation of expensive computer simulations, these Bayesian models deliver not just predictions but honest measures of uncertainty, which is precisely why engineers and scientists rely on them when the cost of being confidently wrong is high. Yet a new theoretical study published in the International Journal of Data Science and Analytics by Akash Sedai and Francesca Medda of University College London delivers an uncomfortable finding: the standard way of training Gaussian processes is fundamentally fragile, and the fragility is not a glitch that better software can fix. It is baked into the mathematics itself.</p>
<p>The heart of Gaussian process training lies in its hyperparameters, the knobs that define how the model believes data points relate to one another. A kernel function encodes assumptions about smoothness, correlation length, signal strength and noise level, and these quantities are typically learned by maximizing the marginal likelihood, a quantity that balances how well the model fits the data against how complex it becomes. For decades this procedure has been regarded as statistically principled and, under ideal Gaussian assumptions, it genuinely is. What Sedai and Medda show, however, is that the estimators produced by this procedure possess what statisticians call an unbounded influence function, a formal certificate that a single extreme observation can exert arbitrarily large leverage on the fitted kernel parameters.</p>
<p>The influence function is a classic concept from robust statistics, introduced by Frank Hampel in the 1970s as a diagnostic for measuring how much an infinitesimal speck of contamination at any point in the data distribution shifts an estimator. If the influence function is bounded, no single observation can dominate the result no matter how extreme it is. If it is unbounded, a lone outlier can drag the estimate anywhere. The authors adapt this framework to the Gaussian process setting with a subtle twist: because the joint log marginal likelihood couples all observations through the kernel covariance matrix, it does not decompose into a sum of independent per-point terms the way classical M-estimators do. Their solution is a leave-one-in contamination construction, in which the perturbation is represented by the predictive log-likelihood of an added data point conditional on the observed sample.</p>
<p>From this construction they derive a closed-form expression for the influence function of the hyperparameter estimator, factoring it into two interpretable pieces: the global curvature of the objective, captured by the Hessian of the log marginal likelihood, and a local score contribution from the contaminated point. That local score depends on the residual between the contaminated response and the model&#8217;s predictive mean, on the predictive variance at the contaminated input, and on where the input sits relative to the observed design. When the input lies far from the training data, the kernel coupling weakens and the geometric influence attenuates, though variance-related parameters such as signal and noise variance remain affected even then.</p>
<p>The central result is stark. Under the standard Gaussian likelihood, the contamination score grows quadratically in the magnitude of the response, so the influence function diverges as the outlier&#8217;s value tends toward infinity. The authors prove this formally in a proposition requiring only mild regularity conditions, and they show that the conclusion is automatic whenever the model includes a signal variance or noise variance parameter, which virtually every practical Gaussian process does. They further demonstrate that the result survives log-parameterizations and automatic relevance determination kernels, closing the escape hatches that practitioners might hope would soften the blow. In Hampel&#8217;s terminology, maximum likelihood training of Gaussian processes is nonrobust by construction.</p>
<p>Crucially, the paper does not stop at diagnosis. The authors extend the influence function framework to the robust alternatives that have long circulated in the Gaussian process literature, including heavy-tailed likelihoods such as the Student-t and Laplace distributions, and robustified objectives built on the Huber loss and its smooth cousin, the Pseudo-Huber or Charbonnier loss. For each case they establish the conditions under which the influence function becomes bounded. The unifying principle is elegant: if the derivative of the pointwise loss with respect to the mean prediction remains bounded over all residuals, then the contamination score is bounded, and so is the influence function, provided the local curvature of the objective stays finite and nonsingular. This connects Gaussian process training directly to the classical theory of M-estimation, recasting it as a kernelized, probabilistic analogue of robust regression.</p>
<p>The empirical validation is unusually thorough. The authors first verify that their analytic influence function matches an independent automatic-differentiation implementation to machine precision, that it is insensitive to the small ridge regularization used for numerical conditioning, and that it reproduces the true first-order response of the optimizer to small data perturbations as measured by finite differences. On a synthetic one-dimensional benchmark with two hundred points, they then compare Gaussian, Student-t, Huber and Pseudo-Huber training under single-point stress tests and controlled contamination. The scatter of per-point influence norms against whitened residuals tells the story visually: the Gaussian fit reaches large influence values at modest residual sizes, while Huber and Pseudo-Huber accommodate residuals several times larger within a comparable influence range, with Student-t sitting in between.</p>
<p>Perhaps the most surprising empirical finding concerns curvature. Clipping the score function, as Huber does, turns out to be necessary but not sufficient for stable hyperparameter behaviour. Because Huber&#8217;s score is flat in the tail, contaminated points contribute no data curvature to the Hessian, and with enough contamination the smallest eigenvalue of the curvature matrix can collapse toward zero or below, destabilizing the local solve and sometimes amplifying influence beyond even the Gaussian fit. The Student-t likelihood, whose score redescends toward zero while retaining curvature, preserved positive curvature in every contaminated run and produced the smallest hyperparameter drift, whereas Huber and Pseudo-Huber lost positive curvature in most runs. The lesson is that robustness depends on both suppressing the direct pull of outliers and maintaining a well-conditioned training objective.</p>
<p>The practical implications ripple outward. In experiments on real datasets including Kin8nm and superconductivity data, the authors show that robust objectives cost essentially nothing when the Gaussian model is appropriate, matching point accuracy and calibrated coverage after a single variance calibration. When contamination or heavy tails appear, however, the Gaussian fit degrades fastest, compensating by inflating its noise estimate and widening its intervals, while Student-t and Pseudo-Huber keep hyperparameters stable and uncertainty honest. The authors also verify that their influence analysis carries over to sparse variational Gaussian processes, the scalable variant used with large datasets, finding that the analytic influence function closely predicts the optimizer&#8217;s actual movement.</p>
<p>For a field that increasingly deploys Gaussian processes in safety-critical emulation, environmental modelling and Bayesian optimization, the message is hard to ignore. The fragility of marginal-likelihood training is a structural property of the estimator, not an artefact of finite precision or poor implementation, and it can be measured, predicted and repaired. Sedai and Medda&#8217;s work gives practitioners both the warning and the toolkit: check the influence structure of your fitted model, prefer heavy-tailed likelihoods when outliers are plausible, and remember that a bounded score alone will not save you if the curvature of your objective quietly disappears. Robustness, it turns out, is a two-part bargain, and Gaussian process training has only recently learned to sign it.</p>
<p><strong>Subject of Research:</strong> Robustness analysis of Gaussian process hyperparameter estimation using influence functions</p>
<p><strong>Article Title:</strong> Influence functions for Gaussian process hyperparameter estimation</p>
<p><strong>Article References:</strong> Sedai, A., &amp; Medda, F. (2026). Influence functions for Gaussian process hyperparameter estimation. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 316. <a href="https://doi.org/10.1007/s41060-026-01285-5" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01285-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01285-5" rel="noopener noreferrer">10.1007/s41060-026-01285-5</a></p>
<p><strong>Keywords:</strong> Gaussian processes, influence functions, hyperparameter estimation, robust statistics, marginal likelihood, outliers, Student-t likelihood, Huber loss, M-estimation, Bayesian inference, kernel methods, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226238</post-id>	</item>
	</channel>
</rss>
