<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>statistical inference &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/statistical-inference/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 16:56:35 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>statistical inference &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New 7S Framework Aims to Unify How Science Judges Data Credibility</title>
		<link>https://scienmag.com/new-7s-framework-aims-to-unify-how-science-judges-data-credibility/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:56:35 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[7S Framework]]></category>
		<category><![CDATA[7S framework for data validation]]></category>
		<category><![CDATA[biomedical engineering]]></category>
		<category><![CDATA[biomedical engineering data verification]]></category>
		<category><![CDATA[clinical decision-making]]></category>
		<category><![CDATA[computational model validation]]></category>
		<category><![CDATA[credibility assessment]]></category>
		<category><![CDATA[credibility of measurement instruments]]></category>
		<category><![CDATA[in silico medicine]]></category>
		<category><![CDATA[interdisciplinary data credibility standards]]></category>
		<category><![CDATA[machine learning predictors]]></category>
		<category><![CDATA[metrology]]></category>
		<category><![CDATA[predictive models]]></category>
		<category><![CDATA[predictive simulation verification]]></category>
		<category><![CDATA[quantitative information]]></category>
		<category><![CDATA[regulatory decision-making in in silico medicine]]></category>
		<category><![CDATA[scientific data credibility assessment]]></category>
		<category><![CDATA[sensor data reliability assessment]]></category>
		<category><![CDATA[statistical inference]]></category>
		<category><![CDATA[statistical inference validation]]></category>
		<category><![CDATA[synthetic data]]></category>
		<category><![CDATA[uncertainty quantification in scientific models]]></category>
		<category><![CDATA[unified approach to data evaluation]]></category>
		<category><![CDATA[verification validation and uncertainty quantification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196671</guid>

					<description><![CDATA[A University of Bologna researcher has proposed a seven-step framework that unifies how science assesses the credibility of measured, inferred, and predicted quantitative information.]]></description>
										<content:encoded><![CDATA[<p>Every number that enters a scientific argument arrives by one of three routes. It is either measured directly with an instrument, inferred statistically from other data, or predicted using a model built from prior knowledge. For decades, each route has carried its own separate machinery for deciding whether the resulting figure deserves to be trusted: metrology governs measurements, statistics governs inference, and the computational science and engineering community relies on Verification, Validation, and Uncertainty Quantification, known as VVUQ, to police predictions. A new letter published in the Annals of Biomedical Engineering argues that this tidy separation is breaking down, and it proposes a single, unified recipe for credibility assessment designed to work across all three sources of quantitative information.</p>
<p>The paper, written by Marco Viceconti of the Department of Industrial Engineering at the University of Bologna, introduces what the author calls the 7S Framework, a seven-step general procedure for evaluating the credibility of any quantitative estimate, whether it originates from a sensor, a statistical model, or a predictive simulation. The motivation is practical rather than purely philosophical. In fields such as in silico medicine, where computational models increasingly inform clinical decisions and regulatory submissions, a new generation of tools refuses to sit neatly within any one of the traditional categories. In silico-augmented clinical trials, physics-informed machine learning predictors, and machine learning models trained on synthetic datasets all blend measured data, statistical inference, and causal prediction into single estimators, leaving established credibility frameworks unable to cover them cleanly.</p>
<p>To build the framework, Viceconti begins with a conceptual scaffolding often described as the pyramid of knowledge. In this picture, observation lifts raw signals produced by a system of interest into data; annotation with metadata about who, what, where, and when lifts data into information; modelling correlations lifts information into tentative causal beliefs; and subjecting those beliefs to falsification experiments lifts them into actionable knowledge. Quantitative information, in this scheme, is an annotated set of values whose metadata specifies the domain of information, everything sender and receiver must know for the values to be meaningful, but without the causal &#8220;why&#8221; that would elevate it to knowledge. New information can be created by measurement, which converts signals into data; by inference, which derives new information from existing information; or by prediction, which uses causal knowledge to generate estimates of quantities that were never observed.</p>
<p>From these foundations, the author generalises a vocabulary that statistics normally reserves for inference. The target quantity is the estimand, the value produced is the estimate, and whatever produces it, be it a thermometer, a regression, or a finite element model, is an estimator. Credibility is then defined with a deliberately demanding definition: the minimum accuracy with which an estimator recovers the true value of the estimand across the entire information space, the bounded multidimensional region defined by all the observable quantities on which the quantity of interest depends. Accuracy itself borrows from metrology, where trueness captures systematic error and precision captures random error, combined into a normalised class of accuracy averaged over repeated estimations.</p>
<p>A crucial insight of the paper is that credibility expectations come in three levels, and that these levels are properties of the intended use rather than of the type of estimator. Level 1 credibility demands only that an estimate fall within a predefined uncertainty band around the true value, appropriate when knowing the order of magnitude suffices. Level 2 requires that the average of repeated estimates match the true value in a statistical sense, as when comparing central properties of populations. Level 3 demands local accuracy at every validation point, the standard for subject-specific models intended to predict individual outcomes. A predictive model can therefore be L1, L2, or L3 credible depending on whether it is meant to capture a scale, a population mean, or a person-specific value, and the same hierarchy applies to measurement and inference.</p>
<p>Because brute-force induction, measuring the error at an effectively infinite number of points, is practically impossible and not even theoretically guaranteed for all estimators, the framework follows the strategy shared by metrology, statistics, and VVUQ: decompose the estimation error into its sources and check that each component behaves as theory predicts for a well-behaved estimator. The seven steps formalise this logic. Step S1 defines the context of use and the acceptable error threshold, the maximum error that still leaves the information useful for the decision it must support, and sets this against the limits of validity imposed by the physics of the phenomenon. Step S2 establishes the source of true values, insisting on measurement chains at least an order of magnitude more accurate than the threshold. Step S3 quantifies estimation error through controlled experiments sampled across the solution space. Step S4 identifies the sources of error, which the paper groups into approximation, aleatoric, and epistemic contributions. Step S5 decomposes the overall error among these sources, sometimes requiring special experiments in which all but one error source is excluded. Step S6 critically reviews whether each error component is distributed as expected. Step S7 examines robustness to biases that could emerge in routine use, including applicability, the guarantee that real-world inputs never stray beyond the validity limits explored during assessment. Transparency throughout, particularly about which error sources are considered and how they are separated, is flagged as essential.</p>
<p>The paper demonstrates the framework on seven use cases drawn from the author&#8217;s research programme, three of which are summarised in detail. The first concerns strain gauge measurements of bone tissue deformation, used to validate finite element models that predict fracture. The context of use fixes an error threshold derived from the strain difference used to determine fracture, attenuated by two orders of magnitude to account for the chain of inference. True values come from beam-theory calculations on machined aluminium alloy specimens corrected for curvature error; trueness is computed as a root-mean-square error; normality tests confirm the expected distribution of random and systematic errors; and repeated tests on bone specimens establish applicability.</p>
<p>The second case applies the framework to the BBCT-Hip predictor, a biophysical model that estimates mechanical strains in a patient&#8217;s bone from a calibrated computed tomography scan. Here the error threshold is set at two percent of the cortical bone failure strain in compression, about 146 microstrain. True values come from strain gauge measurements on carefully preserved cadaveric femurs. The model&#8217;s predictions carry numerical, aleatoric, and epistemic errors, and the VVUQ procedure separates them, with the numerical component required to be negligible, the aleatoric component normally distributed with a near-zero mean, and the epistemic component showing a root-mean-square error close to zero. Applicability is probed by exhaustive experiments spanning inter-subject variability and all relevant loading conditions.</p>
<p>The third case is the most forward-looking: assessing a synthetic cohort inferred from a real clinical cohort of elderly patients at risk of hip fracture. The goal is to run in silico trials on virtual populations far larger than any experimentally collected cohort could be, comparing central properties such as means and medians of feature distributions and model predictions. The error thresholds are tied to the measurement and prediction accuracy of each quantity; epistemic error vanishes because the synthetic data are generated by inference, leaving aleatoric error from measurement uncertainty and numerical error from the interpolation functions, which must be shown negligible by sensitivity analysis. Applicability restricts use of the synthetic cohort to the portion of the information space actually sampled by the clinical data.</p>
<p>Viceconi is careful about scope. For estimators that fall squarely into the classical categories, he recommends continuing to use metrology, statistics, or VVUQ, which are more mature and widely accepted within their communities. The 7S Framework is positioned as a supplement for the growing class of hybrid estimators that do not fit anywhere: in silico-augmented trials that inject model predictions into Bayesian device trials, physics-informed neural networks that encode biomechanical law inside learned predictors, synthetic datasets generated to sidestep privacy constraints, and machine learning surrogates trained to replace expensive biophysical simulations. The framework was also applied to cases covering fracture-risk prediction and machine learning surrogates of computational models, and the author reports that it proved effective, sufficiently general, and sensitive to the subtle differences in what credibility means for each information type. Limitations acknowledged include the restriction of the exposition to single scalar quantities, although extension to vectors and time-dependent quantities requires only adding a norm, and the framework&#8217;s status as a generalisation rather than a replacement of existing practice. As computational medicine pushes further into regulatory territory, the stakes of getting credibility assessment right rise accordingly, and a shared epistemological vocabulary spanning measurement, inference, and prediction may prove to be exactly what regulators, developers, and clinicians need.</p>
<p><strong>Subject of Research:</strong> A general seven-step framework for assessing the credibility of measured, inferred, and predicted quantitative information in computational medicine</p>
<p><strong>Article Title:</strong> Assessing the Credibility of Quantitative Information: A General Framework</p>
<p><strong>Article References:</strong> Viceconti, M. (2026). Assessing the Credibility of Quantitative Information: A General Framework. <em>Annals of Biomedical Engineering</em>. <a href="https://doi.org/10.1007/s10439-026-04367-4" rel="noopener noreferrer">https://doi.org/10.1007/s10439-026-04367-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10439-026-04367-4" rel="noopener noreferrer">10.1007/s10439-026-04367-4</a></p>
<p><strong>Keywords:</strong> credibility assessment, quantitative information, metrology, statistical inference, verification validation and uncertainty quantification, in silico medicine, machine learning predictors, synthetic data, 7S Framework, biomedical engineering, clinical decision-making, predictive models</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196671</post-id>	</item>
		<item>
		<title>Hidden Sampling Gaps Skew Plankton Models, Study Warns</title>
		<link>https://scienmag.com/hidden-sampling-gaps-skew-plankton-models-study-warns/</link>
		
		<dc:creator><![CDATA[Gavin Prescott]]></dc:creator>
		<pubDate>Thu, 03 Sep 2026 17:42:35 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[biological pump and carbon export]]></category>
		<category><![CDATA[calanoid copepod abundance estimation]]></category>
		<category><![CDATA[challenges in marine biodiversity assessment]]></category>
		<category><![CDATA[copepods]]></category>
		<category><![CDATA[dispersion modelling]]></category>
		<category><![CDATA[effects of vertical sampling intervals on marine data]]></category>
		<category><![CDATA[generalized linear models]]></category>
		<category><![CDATA[heteroscedasticity]]></category>
		<category><![CDATA[impact of sampling gaps on plankton models]]></category>
		<category><![CDATA[implications for marine conservation and climate studies]]></category>
		<category><![CDATA[importance of accurate plankton data]]></category>
		<category><![CDATA[inverse Gaussian distribution]]></category>
		<category><![CDATA[Kerguelen Islands]]></category>
		<category><![CDATA[log-transformation]]></category>
		<category><![CDATA[marine ecology]]></category>
		<category><![CDATA[marine food web dynamics]]></category>
		<category><![CDATA[ocean ecosystem response to climate change]]></category>
		<category><![CDATA[oceanographic sampling techniques]]></category>
		<category><![CDATA[sampling bin width]]></category>
		<category><![CDATA[Southern Ocean]]></category>
		<category><![CDATA[statistical inference]]></category>
		<category><![CDATA[statistical modeling in marine ecology]]></category>
		<category><![CDATA[zooplankton]]></category>
		<category><![CDATA[Zooplankton sampling biases]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=186519</guid>

					<description><![CDATA[A new study of Southern Ocean copepods shows that ignoring differences in sampling depth intervals can distort statistical inference in zooplankton abundance models.]]></description>
										<content:encoded><![CDATA[<p>Zooplankton may be small, but they carry the weight of the ocean&#8217;s food web on their translucent shoulders. These drifting animals form the critical trophic bridge between the microscopic phytoplankton that fuel marine primary production and the fish, seabirds, and whales that depend on them, while also playing a central role in the biological pump that exports carbon from the surface ocean to the deep sea. Getting their abundance estimates right is therefore not a niche statistical concern but a foundational requirement for understanding how marine ecosystems respond to a changing climate. A new study published in Discover Oceans argues that one of the most common tools in the plankton ecologist&#8217;s statistical toolbox may be quietly distorting the very patterns scientists are trying to detect.</p>
<p>The research, led by Yulia Egorova of the University of Miami&#8217;s Rosenstiel School of Marine, Atmospheric, and Earth Science together with statisticians and oceanographers at the University of British Columbia, focuses on a deceptively simple question: how should scientists model the abundance of calanoid copepods, one of the most widespread groups of zooplankton in the world&#8217;s oceans? The team&#8217;s answer carries a warning for the field. When datasets combine samples collected over different vertical depth intervals, the resulting differences in measurement precision can quietly reshape which environmental relationships appear statistically significant, and which fade into uncertainty.</p>
<p>The problem begins with how zooplankton are actually collected. Research vessels tow nets through the water column, and the vertical extent covered by each tow, known as the sampling bin width, varies widely between surveys and even between stations on the same cruise. A net hauled through a 400 to 600 meter layer integrates a much smaller volume of water than one dragged from 300 to 700 meters. Although abundance is routinely standardized to volumetric units of individuals per cubic meter, this standardization does not eliminate the underlying difference in precision. A wide bin blends several distinct vertical habitats, each with its own temperature, oxygen regime, and copepod concentration, into a single averaged number, making that number inherently less consistent than an estimate from a narrow, focused interval.</p>
<p>Compounding this design issue is a long-standing habit in biological oceanography: log-transforming abundance data before fitting standard Gaussian linear models. Because volumetric abundance is strictly positive and strongly right-skewed, with most values clustering below one individual per cubic meter and a few extreme values reaching four, researchers have long reached for the logarithm to make the data look more normally distributed. Statisticians have cautioned for years that this transformation can bias inference and obscure how results depend on distributional assumptions, yet the practice remains entrenched in the zooplankton literature.</p>
<p>To test whether these conventions hold up under scrutiny, the team assembled 867 sampling records from 1987 in the waters surrounding the Kerguelen Islands in the Southern Ocean, drawn from the Mesopelagic Mesozooplankton and Micronekton Database. The records covered five dominant calanoid species, including Rhincalanus gigas, Metridia lucens, Calanus simillimus, Pleuromamma robusta, and Ctenocalanus vanus, matched to environmental conditions from the World Ocean Atlas 2018. After screening for multicollinearity among candidate predictors, temperature and average sampling depth, along with their interaction, were retained as the key covariates, with salinity and dissolved oxygen excluded due to strong correlations with depth and with each other.</p>
<p>The researchers then compared four statistical frameworks: a Gaussian model fitted to log-transformed abundance, reflecting conventional practice, and three generalized linear models fitted directly to the original abundance scale assuming log-normal, Gamma, and inverse Gaussian error distributions. Model selection using the generalized Akaike information criterion, with a Jacobian correction to place all models on a common response scale, identified the inverse Gaussian model as the clear winner. Its variance structure, in which variability grows with the cube of the mean, proved best suited to data where dense copepod aggregations are far less predictable than sparse observations. The inverse Gaussian model outperformed the alternatives by a substantial margin, and the log-normal and Gamma models failed to improve on the transformed Gaussian baseline.</p>
<p>But choosing the right error distribution was only half the story. The team extended the best-performing model using generalized additive modeling for location, scale and shape functionality, allowing the dispersion parameter to vary as a function of normalized sampling bin width. This single change improved model fit further and, crucially, altered the ecological conclusions. The temperature-by-depth interaction, which was statistically significant in the constant-dispersion model with a p-value of 0.0023, became non-significant once bin-width-dependent precision was included, with the p-value rising to 0.0657. In contrast, the negative relationship between copepod abundance and depth remained robust throughout, with coefficient estimates changing only marginally and standard errors staying small.</p>
<p>The quantification of the bin-width effect is striking: for every additional 100 meters of sampling depth interval, the residual spread around predicted abundance increased by approximately 15.4 percent. In the study dataset, bin widths ranged across 31 unique intervals from 195 to 537 meters, a level of heterogeneity that is far from unusual in compiled zooplankton databases. The authors suggest that wider bins integrate multiple vertical habitats containing different environmental conditions and copepod concentrations, reducing vertical resolution and producing less consistent abundance estimates. Ignoring this heterogeneity, they argue, can overstate confidence in weaker covariate effects, potentially leading researchers to report environmental relationships that are artifacts of unequal sampling resolution rather than robust biological patterns.</p>
<p>A leave-one-out sensitivity analysis, refitting the final model 867 times after deleting each observation in turn, strengthened confidence in the core findings. The negative depth effect and the positive bin-width effect in the dispersion model were never rendered non-significant by the removal of any single observation, and the temperature main effect remained non-significant throughout. The temperature-by-depth interaction proved more fragile: deleting seven of the 867 observations pushed its p-value below 0.05, although the interaction coefficient stayed positive in every refit. The most influential cases were observations with large positive residuals, primarily from Rhincalanus gigas and Calanus simillimus. The authors recommend treating the interaction as a tentative ecological hypothesis rather than a firm conclusion.</p>
<p>The implications extend well beyond Southern Ocean copepods. The authors suggest that the approach is likely applicable to any ecological dataset with a positive, right-skewed response collected under unequal sampling effort or integration scales, including benthic abundance indexed by sampled area and environmental DNA concentrations indexed by processed volume. They emphasize that allowing dispersion to depend on covariates is established statistical functionality rather than a new model class; the novelty lies in using sampling bin width, a feature of survey design, as a predictor of precision. The team cautions that their conclusions are limited to the four candidate frameworks and the single-year dataset evaluated, and that validation across other regions, years, taxa, and comparisons with weighting schemes, measurement-error models, and hierarchical formulations remain important future directions. For now, the message to marine ecologists is clear: the depth intervals printed in the methods section of a survey report may matter as much as the environmental variables in the analysis itself.</p>
<p>The setting of the study itself adds ecological weight to its methodological message. The Kerguelen Islands sit within the circumpolar Southern Ocean, where the surrounding waters are among the most productive in the region, supporting food webs that include krill, seabirds, and marine mammals. Copepods such as Calanus simillimus and Rhincalanus gigas dominate the mesozooplankton there, and their vertical distributions shift seasonally and with life stage, which is precisely why the depth interval covered by a net tow shapes both what is captured and how precise the resulting estimate can be. In waters where abundance declines steeply with depth, averaging across a broad vertical slab blurs real ecological structure into a single number.</p>
<p>The inverse Gaussian distribution, the best-supported error structure in the comparison, has a long history in statistics. It describes positive, right-skewed data in which the variance grows steeply with the mean, a pattern familiar from physics and reliability engineering before its adoption in ecology. Its arrival as the top candidate for copepod abundance is biologically intuitive: sparse plankton samples are relatively predictable, while dense aggregations, which arise from swarming behavior and patchy advection, are far more variable. Distributions such as the Gamma allow variance to scale linearly with the mean, which evidently understates this heterogeneity in the Kerguelen data.</p>
<p>The environmental covariates came from the World Ocean Atlas 2018, a gridded climatological product built from decades of ship-based measurements. Matching atlas values to individual plankton records by location and mean sampling depth is standard practice, but it introduces its own smoothing, since the atlas represents long-term average conditions rather than the water properties a net actually encountered on a given day. The authors noted that temperature and depth were retained after screening out salinity and dissolved oxygen, which were strongly correlated with depth and with each other, a common multicollinearity problem in oceanographic datasets.</p>
<p>The reliance on a single year, 1987, deserves emphasis. That year supplied the largest eligible sample in the source database by a wide margin, which made it attractive for a methodological comparison but leaves open whether the same error structure and dispersion behavior hold in other years or across seasonal cycles. Interannual variability in Southern Ocean zooplankton is substantial, and the authors themselves frame validation across regions, years, and taxa as the necessary next step before their recommendations become general guidance.</p>
<p><strong>Subject of Research:</strong> Statistical modelling of zooplankton abundance accounting for error distribution choice and sampling depth bin width</p>
<p><strong>Article Title:</strong> How error distribution and sampling depth strata affect plankton abundance modelling</p>
<p><strong>Article References:</strong> Egorova, Y., Tamvada, N., Forrest, D., Pakhomov, E. A., &amp; Auger-Méthé, M. (2026). How error distribution and sampling depth strata affect plankton abundance modelling. <em>Discover Oceans, 3</em>(1), Article 53. <a href="https://doi.org/10.1007/s44289-026-00166-w" rel="noopener noreferrer">https://doi.org/10.1007/s44289-026-00166-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44289-026-00166-w" rel="noopener noreferrer">10.1007/s44289-026-00166-w</a></p>
<p><strong>Keywords:</strong> zooplankton, copepods, generalized linear models, inverse Gaussian distribution, heteroscedasticity, sampling bin width, Kerguelen Islands, Southern Ocean, log-transformation, dispersion modelling, marine ecology, statistical inference</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">186519</post-id>	</item>
	</channel>
</rss>
