A new mathematical result promises to change how scientists revisit old analyses, how hospitals and research consortiums share statistical power, and how machine learning systems avoid expensive retraining. In a paper published in the International Journal of Data Science and Analytics, Alessandro Maria Selvitella of Purdue University Fort Wayne and the University of Washington’s eScience Institute derives an exact algebraic bridge between two of the most common forms of regression: models that include an intercept term and models that force the fitted line through the origin. The striking part is that the bridge can be crossed using nothing but summary statistics, with no access to the underlying raw data at all.
The problem the paper addresses is deceptively simple. Statisticians routinely fit regression models with an intercept, allowing the line to pass through the centroid of the data cloud. But many scientific questions demand a no-intercept, or homogeneous, model, in which the response must vanish when the predictor vanishes. Metabolic rate, for instance, should arguably fall to zero as body mass falls to zero. The default strategy for switching between these specifications has always been to refit the model from scratch, which implicitly assumes the analyst can return to the full design matrix and response vector. In legacy studies, privacy-restricted medical research, and federated learning systems, that assumption often fails.
Selvitella’s central theorem shows that for any quadratic-loss regression, including ordinary least squares, ridge regression, and weighted least squares, the no-intercept estimator can be written in closed form as a function of the with-intercept coefficients, the empirical means of the variables, and the Gram matrices that summarize the data’s geometry. The key identity decomposes the uncentered cross-products used by the no-intercept model into centered cross-products plus a mean-product term. Because the with-intercept solution already satisfies a centered normal equation, the no-intercept solution can be recovered by a single matrix inversion that depends only on the design, the weighting, and any penalty structure, never on the response itself.
Just as importantly, uncertainty propagates through the same bridge. The no-intercept estimator turns out to be an affine function of the with-intercept estimator, so standard covariance calculus delivers its variance directly. Under Gaussian errors, the transformed estimator exactly reproduces the finite-sample distribution of the classical no-intercept solution, and under weak moment conditions it inherits the usual asymptotic normality through Slutsky’s theorem. In other words, confidence intervals and hypothesis tests for the through-the-origin model can be reconstructed from published summaries alone, an algebraically exact result rather than an approximation.
The first major application is federated learning, where data are distributed across many clients, such as hospitals or mobile devices, that never share raw observations. If a consortium has already trained a with-intercept model through the popular FedAvg protocol and then decides it needs the no-intercept version, the conventional answer is a second full federated run, costing on the order of the number of rounds times the sample size times the predictor dimension in local computation, plus repeated communication rounds. The new theorem shows the server can instead aggregate one round of local sufficient statistics, each client transmitting only its sample count, moment vectors, and Gram matrix, and then compute the no-intercept estimator with a cost that scales with the predictor dimension alone and is independent of the number of clients, observations, or communication rounds.
Numerical experiments in the paper back the theory with striking precision. In simulations with ten clients and five thousand observations, the transformed estimator agreed with the centralized no-intercept ordinary least squares solution at floating-point precision, around ten to the minus fifteen. The direct federated no-intercept baseline, by contrast, drifted from the centralized answer as client data became heterogeneous, with relative errors growing from roughly four ten-thousandths in the balanced case to a full ten percent under severe non-IID conditions, a manifestation of the well-known client-drift problem. Timing measurements showed the transformation running between one thousand and twenty-three thousand times faster than a second federated training run, and the agreement held at machine precision even in high-dimensional settings with up to five hundred predictors.
The paper is careful about scope. The savings are relative to a second iterative federated optimization, not to one-shot protocols in which clients transmit exact Gram matrices from the outset, and the method is not by itself a formal privacy guarantee. Aggregate summaries can still leak information when a client holds few observations or unusual covariate patterns, so the author notes that secure aggregation or calibrated noise would need to be layered on top where formal differential privacy is required. There is also a communication trade-off: the one-shot summary protocol costs on the order of the predictor dimension squared per client, so it beats a second federated run only when the number of additional rounds exceeds roughly a quarter of the predictor dimension, a break-even the paper quantifies explicitly.
The second implication is more conceptually surprising: a regression reversal phenomenon. In simple linear regression, the slope estimated with an intercept and the slope estimated through the origin can have strictly opposite signs. Selvitella proves an exact necessary and sufficient condition, namely that the centered cross-product and the uncentered cross-product have opposite signs, which happens precisely when the mean-product term dominates the centered signal. The reversal is driven by the intercept constraint and by leverage, not by classical confounding, yet it is structurally analogous to Simpson’s paradox: in both cases an unadjusted association can point opposite to an adjusted one, with the omitted constant term playing the role of a lurking group-level offset. Because the condition can be checked from summary statistics alone, practitioners can detect potential sign flips before committing to a refit.
The third implication reaches into biology, where only aggregate results are often available and the intercept carries real scientific weight. Reanalyzing a published study of basal metabolic rate in sixty-eight adult men, the paper reconstructs the through-the-origin fit entirely from reported means, standard deviations, and the correlation coefficient. The no-intercept slope comes out near 0.0975, nearly double the with-intercept slope of about 0.054, illustrating how a nonzero intercept can attenuate an estimated scaling relationship. Goodness-of-fit measures can also be reconstructed, though the paper cautions that the uncentered coefficient of determination for no-intercept models is not directly comparable to the conventional centered one; on a common centered scale, the no-intercept model explains less variance than the with-intercept fit, consistent with the original authors’ conclusion that simple linear models inadequately capture allometric scaling.
The broader message is that a piece of linear algebra most statisticians thought they already understood still holds surprises with practical teeth. By placing the affine-to-linear map within a unified framework covering penalized and weighted losses, and by propagating both point estimates and their distributions, the result turns intercept-specification changes from an expensive refitting problem into a cheap post-processing step. For meta-analysts working with decades of published summaries, for consortia bound by data-use agreements, and for engineers designing federated systems, the ability to move between model specifications using sufficient statistics alone offers a rare combination of mathematical exactness and immediate utility, and it suggests that other transformations across the regression landscape may be waiting for similar treatment.
Subject of Research: Summary-statistics-only affine-to-linear transformations linking with-intercept and no-intercept quadratic-loss regression, with applications to federated learning, regression reversal, and biological scaling
Article Title: Affine-to-linear transformations for quadratic-loss regression: computational implications to federated learning, reversal paradox, and applications
Article References: Selvitella, A. M. (2026). Affine-to-linear transformations for quadratic-loss regression: computational implications to federated learning, reversal paradox, and applications. International Journal of Data Science and Analytics, 22(1), Article 289. https://doi.org/10.1007/s41060-026-01268-6
Image Credits: AI Generated
DOI: 10.1007/s41060-026-01268-6
Keywords: regression, federated learning, sufficient statistics, ordinary least squares, ridge regression, weighted least squares, intercept, Simpson's paradox, metabolic scaling, covariance propagation, communication cost, data privacy
Cite Scienmag News
Blake Davidson. (October 2, 2026). Simple Algebra Lets Regression Models Switch Intercept Rules Without Raw Data. Scienmag. https://scienmag.com/simple-algebra-lets-regression-models-switch-intercept-rules-without-raw-data/
Blake Davidson. "Simple Algebra Lets Regression Models Switch Intercept Rules Without Raw Data." Scienmag, 2 October 2026, https://scienmag.com/simple-algebra-lets-regression-models-switch-intercept-rules-without-raw-data/. Accessed 2 October 2026.
Blake Davidson. "Simple Algebra Lets Regression Models Switch Intercept Rules Without Raw Data." Scienmag. October 2, 2026. https://scienmag.com/simple-algebra-lets-regression-models-switch-intercept-rules-without-raw-data/








