<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>cis-regulatory variants &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/cis-regulatory-variants/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 16:53:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>cis-regulatory variants &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>xDecoder Pushes Genomic Foundation Models Toward Personal Gene Expression Prediction</title>
		<link>https://scienmag.com/xdecoder-pushes-genomic-foundation-models-toward-personal-gene-expression-prediction/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:53:00 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[ATAC-seq]]></category>
		<category><![CDATA[challenges in personalized genomics]]></category>
		<category><![CDATA[Chromatin Accessibility]]></category>
		<category><![CDATA[cis-regulatory variants]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for gene regulation]]></category>
		<category><![CDATA[DNA-only AI models]]></category>
		<category><![CDATA[eQTL]]></category>
		<category><![CDATA[eQTL studies limitations]]></category>
		<category><![CDATA[gene expression prediction]]></category>
		<category><![CDATA[gene expression prediction in individuals]]></category>
		<category><![CDATA[genomic foundation models]]></category>
		<category><![CDATA[genomic language models]]></category>
		<category><![CDATA[GEUVADIS]]></category>
		<category><![CDATA[improving gene expression prediction accuracy]]></category>
		<category><![CDATA[personal gene expression prediction]]></category>
		<category><![CDATA[personalized genomics]]></category>
		<category><![CDATA[personalized regulatory genomics]]></category>
		<category><![CDATA[rare and novel genetic variants]]></category>
		<category><![CDATA[sequence-to-function models]]></category>
		<category><![CDATA[sequence-to-function models in genomics]]></category>
		<category><![CDATA[xDecoder]]></category>
		<category><![CDATA[xDecoder computational framework]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196611</guid>

					<description><![CDATA[A new decoding framework called xDecoder more than triples few-shot accuracy of genomic foundation models for personal gene expression prediction while exposing a persistent cross-locus generalization bottleneck in DNA-only models.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers at the University of Hong Kong and Shenzhen University has developed a new computational framework that dramatically improves how well genomic foundation models can predict gene expression in individual people, while simultaneously exposing a fundamental weakness in today&#8217;s DNA-only artificial intelligence models. The tool, called xDecoder, is described in a study published in Molecular Systems Biology and offers both a practical advance and a candid diagnosis of where the field of personalized regulatory genomics currently falls short.</p>
<p>Predicting how a person&#8217;s unique genetic makeup shapes their gene expression has long been a central goal of modern biology. Expression quantitative trait locus, or eQTL, studies have linked thousands of genetic locations to gene activity, but these statistical approaches depend on massive population datasets and are therefore largely restricted to common genetic variants. Rare and novel variants, which are often critical for understanding disease mechanisms, remain stubbornly difficult to interpret. Sequence-to-function models such as Enformer and Borzoi were designed to close this gap by using deep learning to map DNA sequence directly to regulatory outputs, and they perform well when predicting average expression across genes and tissues. Yet a series of benchmarking studies published in recent years showed that these models capture very little of the transcriptomic variation between individuals, prompting researchers to ask whether a fundamentally different approach was needed.</p>
<p>The Hong Kong team, led by Shumin Li, Ruibang Luo, and Yuanhua Huang, saw promise in a newer class of models: genomic language models, or gLMs. Unlike sequence-to-function models, which are trained on curated functional annotations across many cellular contexts, gLMs such as Evo2-7B, the Nucleotide Transformer, and Caduceus learn generalizable representations of DNA through self-supervised pretraining on raw sequence alone. Whether these models could complement or extend existing sequence-to-function approaches for personal expression prediction was an open question. xDecoder was built to answer it. The framework is a lightweight convolutional neural network decoding tower that consumes the full token-level embeddings produced by six different foundation models, three genomic language models and three sequence-to-function models, and translates them into personal gene expression predictions.</p>
<p>The researchers trained and evaluated xDecoder using paired genome and transcriptome data from the GEUVADIS consortium, which includes RNA sequencing and high-quality genotypes from lymphoblastoid cell lines of 1000 Genomes Project participants. Individualized DNA sequences were reconstructed by substituting single-nucleotide variants from phased genotypes into the reference genome. The team focused on genes with significant cis-eQTLs, reserving 200 genes for personalized training and holding out 287 genes on chromosomes 5 and 10 for stringent, locus-disjoint evaluation, with 50 individuals for training and 100 for testing. Importantly, the xDecoder architecture deliberately avoids compressing embeddings through global pooling. Because only about 0.01 percent of positions differ between any two people across a gene&#8217;s sequence, preserving positional resolution is essential, so the model aggregates local neighborhoods with small stacked depthwise-separable convolutional layers before pooling and regression, and averages predictions across the two haplotypes.</p>
<p>The study&#8217;s central finding is a tale of two learning regimes. In the few-shot setting, where models are fine-tuned on paired personal genome-expression data for genes seen during training, xDecoder delivered consistent and substantial improvements. Cross-individual prediction correlations more than tripled relative to pretrained Enformer and Borzoi and to reference-trained xDecoder variants, rising from 0.089 to 0.194 on average. For specific genes the gains were striking: for HLA-DQA1, pretrained Enformer and AlphaGenome produced negative correlations, while individually fine-tuned xDecoder models exceeded 0.3, with a Caduceus-based variant reaching 0.74. The best genomic language model embeddings even outperformed the best sequence-to-function embeddings, and rp-Caduceus successfully modeled 46 genes for which the widely used linear method PrediXcan produced identical, uninformative predictions across all individuals. Exploratory variant-level perturbation analysis further showed that high-scoring xDecoder perturbations overlapped known eQTL variants, hinting that the model can nominate candidate cis-regulatory variants in loci where fixed genotype weights fail.</p>
<p>The zero-shot picture was far less encouraging. When the models were asked to predict expression across unseen individuals for genes never encountered during training, none performed reliably. Zero-shot performance remained close to zero even as the number of training individuals was scaled from 50 up to 250, demonstrating that the limitation is not simply a matter of insufficient data. Performance did improve modestly for genes with stronger cis-genetic signal, particularly for AlphaGenome, but even in the most favorable strata, zero-shot results fell well below gene-specific supervised models such as PrediXcan, which achieved a mean Spearman correlation of 0.216. The authors characterize this as a persistent cross-locus transfer bottleneck: current sequence models fail to learn a transferable grammar of cis-regulation that generalizes from one genomic neighborhood to another. Cross-population experiments added a further caution, showing that benefits of individual-level training diminished when models trained on European individuals were tested on Yoruba individuals from Ibadan, with genes carrying higher novel variant rates and less similar linkage disequilibrium structure showing larger performance drops.</p>
<p>To diagnose what information DNA-only models are missing, the team conducted a revealing experiment with chromatin accessibility. By adding a base-resolution ATAC-seq channel to Caduceus embeddings, they significantly improved prediction for unseen genes, lifting the mean cross-individual correlation from -0.003 to 0.129 and flipping correlation direction to positive for many genes; for the gene CHST3, adding the ATAC track transformed a wrongly signed prediction into a correlation of 0.755. However, control analyses complicated the interpretation. In the unseen-gene setting, a simple unsupervised baseline that just summed ATAC signal across the same genomic window was competitive with the full model, suggesting that measured chromatin accessibility largely explains the gain. Meanwhile, AlphaGenome-predicted ATAC signals showed only modest concordance with observed accessibility across individuals, indicating that current DNA-only models cannot yet reliably infer a person&#8217;s chromatin state from sequence alone. In the seen-gene setting, by contrast, the full DNA-plus-ATAC model outperformed an ATAC-only linear baseline, showing that sequence-derived information adds genuine value when paired training data exist.</p>
<p>The authors are careful about the limits of their advance. Absolute prediction accuracy, while improved, remains modest, and the framework should not yet be interpreted as reliable enough for gene-level prediction from personal genomes in any clinical sense. The evaluation was restricted to lymphoblastoid cell lines, and extending the approach to larger cohorts, multiple tissues, and context-specific regulation remains future work. The researchers also stress that variant prioritization applications require systematic benchmarking against eQTL-supported variants, and that clinical interpretation would demand far stronger calibration and independent validation. Rather than presenting DNA-plus-ATAC inputs as a practical replacement for RNA sequencing, they frame the ATAC experiments as a diagnostic that pinpoints the key bottleneck of current DNA-only models.</p>
<p>Even so, the study sketches a concrete roadmap for the next generation of personalized regulatory models. Combining pretrained sequence encoders with variant-aware modules, using individual-level multi-omic profiles such as ATAC-seq as auxiliary supervision or intermediate regulatory-state targets, and exploiting the few-shot regime where sequence-based models already outperform fixed-variant statistical tools together form a coherent research agenda. Because sequence-based models operate directly on DNA, they can in principle evaluate perturbations outside the fixed variant-weight space of linear predictors, offering potential utility wherever rare or novel variants lack population-derived weights. The xDecoder code has been released openly on GitHub, and the authors suggest that bridging the cross-locus generalization gap, likely through multi-omic, variant-aware training, is the defining challenge that will determine whether genomic foundation models can ultimately deliver truly personalized expression prediction from DNA alone.</p>
<p><strong>Subject of Research:</strong> Few-shot personal gene expression prediction using genomic foundation models</p>
<p><strong>Article Title:</strong> xDecoder unlocks the potential of genomic foundation models for few-shot personal gene expression prediction</p>
<p><strong>Article References:</strong> Li, S., Luo, R., &amp; Huang, Y. (2026). xDecoder unlocks the potential of genomic foundation models for few-shot personal gene expression prediction. <em>Molecular Systems Biology</em>. <a href="https://doi.org/10.1038/s44320-026-00238-1" rel="noopener noreferrer">https://doi.org/10.1038/s44320-026-00238-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s44320-026-00238-1" rel="noopener noreferrer">10.1038/s44320-026-00238-1</a></p>
<p><strong>Keywords:</strong> genomic foundation models, gene expression prediction, personalized genomics, xDecoder, genomic language models, sequence-to-function models, eQTL, chromatin accessibility, ATAC-seq, deep learning, GEUVADIS, cis-regulatory variants</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196611</post-id>	</item>
	</channel>
</rss>
