<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>hyperparameter optimization impact &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/hyperparameter-optimization-impact/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 12:08:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>hyperparameter optimization impact &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Stop trusting single accuracy scores: why ML models need distributional evaluation</title>
		<link>https://scienmag.com/stop-trusting-single-accuracy-scores-why-ml-models-need-distributional-evaluation/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 12:08:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[best practices for model validation]]></category>
		<category><![CDATA[confidence intervals]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[distributional performance metrics]]></category>
		<category><![CDATA[experimental design]]></category>
		<category><![CDATA[hyperparameter optimization]]></category>
		<category><![CDATA[hyperparameter optimization impact]]></category>
		<category><![CDATA[importance of multiple model assessments]]></category>
		<category><![CDATA[interpreting model performance distributions]]></category>
		<category><![CDATA[limitations of single accuracy scores]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning evaluation]]></category>
		<category><![CDATA[model evaluation]]></category>
		<category><![CDATA[model robustness and variability]]></category>
		<category><![CDATA[quantiles]]></category>
		<category><![CDATA[randomized data augmentation]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[risk assessment]]></category>
		<category><![CDATA[statistical analysis in ML]]></category>
		<category><![CDATA[statistics]]></category>
		<category><![CDATA[stochastic nature of ML training]]></category>
		<category><![CDATA[train-test split effects]]></category>
		<category><![CDATA[uncertainty quantification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=262126</guid>

					<description><![CDATA[A large-scale study argues that machine learning performance should be reported as distributions with confidence intervals on quantiles rather than single point estimates, revealing model weaknesses that average metrics conceal.]]></description>
										<content:encoded><![CDATA[<p>Every machine learning practitioner knows the ritual: train a model, run it on a held-out test set, and report a single number — an accuracy of 0.91, a root mean squared error of 10.4. That number then travels through papers, dashboards, and boardroom slides as if it were a definitive statement about the model. A new study argues that this ritual is quietly misleading us. Machine learning training is a stochastic process, shaped by random weight initializations, random data splits, random augmentation, and randomized hyperparameter search, and a single performance figure is just one draw from an unknown underlying distribution. Treating it as the whole story, the authors contend, is like judging a drug trial from a single patient.</p>
<p>The research, published in Machine Learning with Applications by Christoph Lehmann and Yahor Paromau, proposes a deceptively simple reframing: view ML training as a controlled scientific experiment with repeated measurements. By varying one confounding factor at a time — the train–test split, the initial weights, the data augmentation scheme, the dropout configuration, or the hyperparameter optimization procedure — while holding everything else fixed via random seed control, each training run becomes a sample from the distribution of a target metric of interest. Repeating the run with different seeds yields repeated measurements, and with them the tools of classical inferential statistics become applicable: quantiles, confidence intervals, and explicit uncertainty quantification.</p>
<p>The scale of the empirical work is unusual for this kind of methodological study. The authors ran roughly 500 to 1,000 seed-controlled training repetitions for each source of variation across three real-world use cases: image classification of Simpsons characters with convolutional neural networks, the standard CIFAR10 benchmark with a VGG11 architecture, and a regression task predicting the critical temperature of superconductors using deep neural networks and gradient boosting trees. The effort consumed approximately 19,000 GPU hours on NVIDIA A100 and H100 accelerators, producing empirical performance distributions dense enough to serve as near-ground-truth references against which small-sample statistical methods could be validated.</p>
<p>The core statistical move is a shift from means to quantiles. A quantile answers a risk-oriented question that a mean cannot: the 25% quantile of accuracy tells you the accuracy falls below that threshold in only a quarter of runs, while the 90% quantile of RMSE tells you the error exceeds that value in just one run out of ten. The authors draw an explicit parallel to Value-at-Risk in financial regulation, where a quantile of the loss distribution defines the threshold exceeded only with a small pre-specified probability. For high-stakes applications — medicine, banking, criminal justice — this threshold language maps naturally onto acceptance criteria: a model can be required, in advance, to keep its error below a bound with a stated probability, something a mean-based evaluation cannot certify.</p>
<p>Quantiles, however, are harder to estimate than means. They are sensitive to the shape of the distribution, and the more extreme the quantile level, the more data is needed. The study characterizes this trade-off precisely. Exact nonparametric confidence intervals, built from order statistics and the binomial distribution, exist only when the sample size clears a minimum that grows steeply toward the distribution&#8217;s tails: estimating a 10% quantile with 90% confidence requires at least 22 runs, while a 1% quantile demands 230. Asymptotic intervals based on the normal approximation of sample quantiles have their own requirements, and the authors show a clever sign transformation — analyzing the negated metric — that converts a costly lower-tail quantile into a cheaper upper-tail one, cutting the required sample size for a 10% quantile from 42 to 25 runs.</p>
<p>For situations where even those minimums cannot be met, the authors evaluate a semiparametric bootstrap that extrapolates into the distribution&#8217;s tails using the spacing of the outermost order statistics. This method remains applicable where the closed-form intervals are undefined, at the price of extrapolation that may not match the true tail behavior. Across 42 experimental settings and a simulation study spanning skewed Beta distributions, normal mixtures, and Gumbel-distributed errors, the asymptotic nonparametric interval emerged as the practical default, offering a favorable balance of coverage, interval length, and simplicity. Bootstrap intervals remained a usable fallback for very small samples, though with coverage dropping to around 0.85 and noticeably wider intervals.</p>
<p>The practical payoff is demonstrated with two worked examples. In the Simpsons classification task, two training configurations — one varying hyperparameter optimization, one varying data augmentation — had nearly identical mean accuracies of 0.874 and 0.872, indistinguishable by any mean-based comparison. But their 10% quantile confidence intervals, estimated from just 25 runs each, did not overlap at all: [0.870, 0.873] for hyperparameter optimization versus [0.860, 0.868] for data augmentation. The hyperparameter-optimized configuration was substantially more resistant to producing poor runs, a fact invisible in the means. Against a deployment requirement that accuracy not drop below 0.87 in more than 10% of cases, only one configuration qualified — a statistically grounded model selection decision that point estimates could never support.</p>
<p>The regression example on superconductors showed the complementary value of the framework. A deep neural network and a gradient boosting tree differed clearly in mean RMSE, 10.47 versus 9.59, and their confidence intervals for the mean, the 10% quantile, and the 90% quantile were all cleanly separated with similar widths. That uniform pattern, the authors note, is the empirical signature of a simple location shift between the two distributions — and its absence of tail-specific effects is itself informative, justifying a mean-based comparison in this case. Across 2,000 repeated paired draws, the quantile intervals separated the two models in 92 to 99 percent of draws at n=25, while the mean separated them essentially always.</p>
<p>The study also distills concrete guidance for practitioners working under realistic compute budgets. With fewer than 10 repetitions, only a t-interval for the mean is advisable. Between 10 and 15 runs, exact nonparametric intervals for the quartiles become feasible, with a maximum recommended confidence level of 0.9. At 25 runs, asymptotic intervals with the sign transformation extend coverage to the 10% and 90% quantile levels, and at 50 or more runs, extreme quantiles such as the 5% and 95% levels come within reach. Interval lengths themselves carry diagnostic value: a median interval markedly longer than the mean&#8217;s t-interval signals low density near the center and hints at bimodality, while asymmetric tail intervals reveal skewness. The authors also caution that metrics like accuracy are formally discrete, computed on a grid set by test-set size, and recommend inspecting the fraction of distinct values in a sample before trusting quantile estimates.</p>
<p>What emerges is a bridge between two communities that rarely speak. Statisticians have long treated experimental outcomes as random variables deserving of distributional analysis; machine learning practice, optimized for leaderboard positions and single benchmark numbers, has largely resisted that view. This work shows that the statistical machinery is not merely compatible with modern deep learning — it is affordable. Even 10 to 25 seed-controlled runs, a modest addition to any training budget, unlock tail behavior, stability assessment, and requirement-driven evaluation that no point estimate can provide. As models move into domains where a single bad run can mean a misdiagnosis or a denied loan, the authors&#8217; message lands with force: the average is not the answer, and the distribution is the evidence.</p>
<p><strong>Subject of Research:</strong> Distributional uncertainty quantification in machine learning performance evaluation using quantile estimation and confidence intervals</p>
<p><strong>Article Title:</strong> Beyond point estimates: Distributional uncertainty in machine learning performance evaluation</p>
<p><strong>Article References:</strong> Lehmann, C., &amp; Paromau, Y. (2026). Beyond point estimates: Distributional uncertainty in machine learning performance evaluation. <em>Machine Learning with Applications, 26</em>, Article 101033. <a href="https://doi.org/10.1016/j.mlwa.2026.101033" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101033</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> machine learning, model evaluation, uncertainty quantification, confidence intervals, quantiles, statistics, reproducibility, deep learning, risk assessment, benchmarking, hyperparameter optimization, experimental design</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">262126</post-id>	</item>
	</channel>
</rss>
