<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Sparse regression for large-scale genetic data &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/sparse-regression-for-large-scale-genetic-data/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 08:12:01 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Sparse regression for large-scale genetic data &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Sparse Regression Method Tames Biobank-Scale Genetic Data for Risk Prediction</title>
		<link>https://scienmag.com/new-sparse-regression-method-tames-biobank-scale-genetic-data-for-risk-prediction/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 08:12:01 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[Biotechnology]]></category>
		<category><![CDATA[biobank data analysis techniques]]></category>
		<category><![CDATA[biobank-scale genetic data modeling]]></category>
		<category><![CDATA[clinical application of genetic risk scores]]></category>
		<category><![CDATA[cost-effective genetic variant selection]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[genome-wide association studies]]></category>
		<category><![CDATA[genome-wide association studies with sparse models]]></category>
		<category><![CDATA[genomics]]></category>
		<category><![CDATA[high-dimensional genetic analysis]]></category>
		<category><![CDATA[high-dimensional statistics]]></category>
		<category><![CDATA[interpretable genetic risk prediction]]></category>
		<category><![CDATA[LASSO]]></category>
		<category><![CDATA[penalized regression in genomics]]></category>
		<category><![CDATA[PLOS Genetics]]></category>
		<category><![CDATA[polygenic risk score computation]]></category>
		<category><![CDATA[polygenic risk scores]]></category>
		<category><![CDATA[predictive modeling]]></category>
		<category><![CDATA[scalable genetic risk prediction algorithms]]></category>
		<category><![CDATA[sparse regression]]></category>
		<category><![CDATA[Sparse regression for large-scale genetic data]]></category>
		<category><![CDATA[statistical methods for big genomic datasets]]></category>
		<category><![CDATA[summary statistics]]></category>
		<category><![CDATA[UK Biobank]]></category>
		<category><![CDATA[uniLasso]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=252749</guid>

					<description><![CDATA[Researchers have adapted a two-stage sparse regression method called uniLasso to the UK Biobank, producing polygenic risk scores that match the accuracy of standard approaches while using far fewer genetic variants.]]></description>
										<content:encoded><![CDATA[<p>Geneticists have long dreamed of distilling the vast complexity of the human genome into compact, interpretable models that can predict an individual&#8217;s risk of disease. A new study published in PLOS Genetics brings that dream closer to reality. Researchers led by Joshua Richland, Tuomo Kiiskinen, and colleagues, working with statisticians Robert Tibshirani and Trevor Hastie of Stanford University, have adapted a recently introduced statistical technique called Univariate-Guided Sparse Regression, or uniLasso, to the scale of the UK Biobank, one of the largest biomedical research resources in the world. The method promises polygenic risk scores that are as accurate as those produced by established approaches, yet built from far fewer genetic variants, making them cheaper to compute, easier to interpret, and more practical for clinical deployment.</p>
<p>Polygenic risk scores, commonly abbreviated as PRS, summarize the combined influence of thousands or millions of genetic variants on a person&#8217;s likelihood of developing a particular condition. Constructing them is a formidable statistical challenge. Modern biobanks measure hundreds of thousands of individuals at millions of genetic positions, and the number of potential predictors vastly exceeds the number of study participants. In this high-dimensional regime, ordinary regression breaks down, and researchers must rely on penalized regression methods that shrink or eliminate weak signals. The most famous of these, the Lasso introduced by Tibshirani in 1996, adds a penalty proportional to the sum of absolute coefficient values, driving the coefficients of uninformative variants all the way to zero and thereby selecting a sparse subset of predictors.</p>
<p>Despite its elegance, the Lasso faces difficulties when the number of variables dwarfs the number of observations, as it does in genomic data. Feature selection can be unstable: small perturbations in the data may cause the method to pick different variants, and correlated variants can compete with one another in ways that obscure the true signal. UniLasso addresses these problems with a two-stage strategy. In the first stage, the method performs simple univariate regressions, examining each genetic variant one at a time for its individual association with the trait of interest. The signs and magnitudes of these univariate coefficients are then used to guide the second stage, a multivariate penalized regression fitted on the full set of variants simultaneously. Variants with strong univariate evidence receive favorable treatment in the penalty, while those with weak or contradictory evidence are discouraged from entering the model.</p>
<p>The intuition behind this design is that univariate screening provides a stable, data-driven prior about which variants matter. By anchoring the multivariate fit to this prior, uniLasso stabilizes feature selection across datasets and reduces the variance of the resulting coefficient estimates. The authors emphasize that the univariate stage does not replace the multivariate fit; it merely informs it. The final model is still learned from individual-level data in the target population, which means the method can adapt to the specific ancestry, cohort composition, and phenotype definitions of the dataset at hand. This two-stage architecture also confers a computational advantage: the initial univariate screen is embarrassingly parallel and can be distributed across many processors, making the approach feasible for datasets with more than one million genetic variants measured on hundreds of thousands of people.</p>
<p>To test the method at true biobank scale, the team applied uniLasso to the UK Biobank, a population-based repository containing genetic, health, and lifestyle data from roughly half a million participants across the United Kingdom. The researchers adapted the algorithm to handle the sheer dimensionality of the resource, where more than a million variants must be evaluated simultaneously. They also introduced an extension called uniLasso ES, short for external scores, which incorporates summary statistics from pre-existing genome-wide association studies. These external signals, drawn from large meta-analyses that may involve millions of additional participants, guide the regression toward variants with prior evidence of association by informing penalty weights and imposing sign constraints that keep coefficient directions consistent with earlier findings.</p>
<p>The uniLasso ES framework is particularly significant because it bridges two traditionally separate paradigms in polygenic risk prediction. Summary-statistic methods, such as PRS-CS and lassosum2, rely exclusively on aggregated association results and linkage disequilibrium reference panels, avoiding the privacy and logistical hurdles of individual-level data but sacrificing some flexibility. Individual-level methods, by contrast, can model the full joint structure of the data but require access to raw genotypes. UniLasso ES occupies a middle ground: external summary statistics act as a soft guide, shaping the penalty landscape, while the definitive fitting happens on individual-level target data. The external evidence informs the model without dictating it, allowing the method to correct for differences between the discovery population and the target population.</p>
<p>The empirical results reported in the study are striking. Across a range of traits and disease outcomes in the UK Biobank, uniLasso attained predictive performance comparable to the standard Lasso while selecting substantially fewer variants. Sparser models offer tangible benefits beyond aesthetics. A risk score built from a few thousand variants rather than hundreds of thousands costs less to genotype in a clinical setting, reduces the storage and computational burden of scoring large patient cohorts, and, crucially, is easier to interrogate biologically. When a compact set of variants drives a prediction, researchers can more readily trace which genes and pathways contribute to the score, potentially generating new hypotheses about disease mechanisms.</p>
<p>Interpretability has been a persistent criticism of polygenic risk scores. Critics note that scores built from millions of tiny effects are statistical black boxes whose individual components rarely correspond to established biology. By producing models that are sparse by construction and whose selected variants carry coherent univariate evidence, uniLasso nudges the field toward scores that a geneticist can actually read. The method&#8217;s reliance on univariate coefficient signs also guards against a known pathology of penalized regression in the presence of correlated predictors, where a variant may enter the model with a counterintuitive sign simply because it is absorbing the effect of a nearby variant. Sign constraints informed by univariate evidence keep the fitted coefficients aligned with the marginal signal.</p>
<p>Benchmarking against competing approaches reinforced the method&#8217;s promise. The authors report that both uniLasso and uniLasso ES remained competitive with PRS-CS and lassosum2, two widely used PRS estimation methods that represent the current state of the art. Achieving parity with these established tools while producing markedly sparser models positions uniLasso as a practical alternative for research groups that need both accuracy and parsimony. The scalability of the implementation matters as much as its statistical properties; as biobanks grow toward cohorts of millions and sequencing expands the variant catalogue into the billions, methods that cannot be distributed efficiently will simply fall out of use.</p>
<p>The study arrives at a moment when polygenic risk scores are edging toward clinical application, with health systems beginning to explore their use in screening programs for conditions such as cardiovascular disease, diabetes, and several cancers. The promise of personalized prevention depends on scores that are accurate, portable across ancestries, and affordable to deploy. A method that delivers competitive accuracy with a fraction of the variants addresses all three concerns at once. The work also exemplifies a broader trend in statistics: the creative combination of classical ideas, in this case univariate screening and penalized regression, to meet the demands of modern data scale. As the authors and their collaborators continue to refine the framework, uniLasso may well become a standard tool in the geneticist&#8217;s arsenal, helping to convert the torrent of biobank data into models that clinicians and patients can understand and act upon.</p>
<p><strong>Subject of Research:</strong> Scalable sparse regression methods for polygenic risk score computation in biobank-scale genomic data</p>
<p><strong>Article Title:</strong> Univariate-guided sparse regression for Biobank-scale high-dimensional omics data</p>
<p><strong>Article References:</strong> Univariate-guided sparse regression for Biobank-scale high-dimensional omics data. (n.d.). <a href="https://doi.org/10.1371/journal.pgen.1012314" rel="noopener noreferrer">https://doi.org/10.1371/journal.pgen.1012314</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pgen.1012314" rel="noopener noreferrer">10.1371/journal.pgen.1012314</a></p>
<p><strong>Keywords:</strong> polygenic risk scores, uniLasso, UK Biobank, sparse regression, Lasso, genome-wide association studies, high-dimensional statistics, summary statistics, predictive modeling, genomics, feature selection, PLOS Genetics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">252749</post-id>	</item>
	</channel>
</rss>
