<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>genomic language models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/genomic-language-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 23:40:32 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>genomic language models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Codon Optimality Predicts mRNA Lifespan but Fails for Noncoding RNAs</title>
		<link>https://scienmag.com/codon-optimality-predicts-mrna-lifespan-but-fails-for-noncoding-rnas/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:40:32 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[codon optimality]]></category>
		<category><![CDATA[comparative analysis of RNA half-life]]></category>
		<category><![CDATA[cross-species RNA stability]]></category>
		<category><![CDATA[cross-validation]]></category>
		<category><![CDATA[Gene regulation]]></category>
		<category><![CDATA[genomic language models]]></category>
		<category><![CDATA[limitations of codon optimality in noncoding RNAs]]></category>
		<category><![CDATA[Long non-coding RNA]]></category>
		<category><![CDATA[long non-coding RNAs]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[metabolic labeling sequencing assays]]></category>
		<category><![CDATA[molecular determinants of RNA longevity]]></category>
		<category><![CDATA[mRNA stability]]></category>
		<category><![CDATA[noncoding RNA lifespan]]></category>
		<category><![CDATA[nonsense-mediated decay]]></category>
		<category><![CDATA[RNA decay]]></category>
		<category><![CDATA[RNA half-life]]></category>
		<category><![CDATA[RNA-binding proteins]]></category>
		<category><![CDATA[sequence-based transcript decay prediction]]></category>
		<category><![CDATA[transcriptome]]></category>
		<category><![CDATA[transcriptome regulation in mammals]]></category>
		<category><![CDATA[translation efficiency and RNA decay]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204012</guid>

					<description><![CDATA[A new computational study shows that codon optimality, one of the strongest sequence determinants of messenger RNA half-life, does not predict the lifespan of long non-coding RNAs.]]></description>
										<content:encoded><![CDATA[<p>For more than a decade, molecular biologists have been captivated by a remarkably elegant idea: that the very spelling of a gene, codon by codon, can determine how long its messenger RNA survives inside a cell. Codon optimality, the principle that ribosomes translate optimal codons more rapidly and that slow translation recruits mRNA decay machinery, has emerged as one of the strongest sequence-based predictors of messenger RNA half-life across species as distant as yeast, zebrafish and humans. But a provocative new study published in Molecular Genetics and Genomics asks a question that cuts to the heart of this framework: does the same coding-sequence logic govern the lifespans of long non-coding RNAs, the vast family of transcripts that make up much of the mammalian transcriptome but are, by definition, largely untranslated? The answer, according to a rigorous computational analysis, is an emphatic no.</p>
<p>The study, conducted by Hidenori Tani of Yokohama University of Pharmacy, takes an unusually careful comparative approach. Rather than comparing mRNA and lncRNA stability data from different experiments, different laboratories or different measurement platforms, Tani assembled half-life measurements derived from the same metabolic labeling and sequencing assays, allowing messenger RNAs and long non-coding RNAs to be evaluated on a genuinely level playing field. This design matters, because the field has been plagued by comparisons in which differences in assay chemistry, cellular context or transcript abundance could masquerade as biological differences in decay regulation. By holding the measurement method constant, the analysis isolates the one variable of interest: whether the sequence of a transcript encodes its own stability.</p>
<p>The findings for messenger RNAs confirm the power of codon-mediated regulation. When Tani re-estimated codon stabilization coefficients, quantitative measures of how strongly each codon is associated with transcript longevity, inside every cross-validation fold to eliminate information leakage between training and test data, the resulting model predicted mRNA half-life with a cross-validated coefficient of determination of 0.174 and a Spearman rank correlation of 0.44. Critically, the codon stabilization coefficient on its own, yielding an R-squared of 0.084, outperformed a six-member set of translation-independent features, which managed only 0.008. That ordering, with codon optimality beating a battery of sequence characteristics such as GC content, transcript length and known destabilizing elements, held consistently across three additional datasets, two distinct measurement technologies and two different species.</p>
<p>The picture for long non-coding RNAs could hardly be more different. In a carefully matched set of 364 lncRNAs measured in the same HeLa cell assay as the messenger RNAs, sequence-based models of half-life performed no better than chance, yielding an R-squared of minus 0.050, a value meaning the predictions were actually worse than simply guessing the average half-life for every transcript. The result was not a quirk of a small or unrepresentative sample. When Tani turned to the largest published lncRNA stability dataset, encompassing 33,285 transcripts, k-mer composition features achieved an R-squared of essentially zero, at 0.0007. Even the cross-validated rank correlation reported in that dataset&#8217;s original study, 0.091, is vanishingly small in absolute terms.</p>
<p>Here the study makes a subtle but important statistical argument. A Spearman correlation of 0.091 in a dataset of more than thirty thousand transcripts does clear the permutation null, meaning it is statistically detectable rather than pure noise. But detectability is not the same as explanatory power, and the contrast between the two RNA classes is one of magnitude, not of presence or absence. Messenger RNA half-life is meaningfully sequence-encoded; long non-coding RNA half-life, by every measure Tani applied, is not. The distinction has practical consequences for anyone attempting to model RNA decay: a signal that is real but negligible cannot support the kind of predictive machinery that works for coding transcripts.</p>
<p>The analysis also confronts one of the most fashionable tools in modern genomics: pretrained genomic language models. These deep neural networks, trained on enormous corpora of genomic sequence, have been touted as universal feature extractors capable of discovering biological signals without explicit programming. Tani tested models including HyenaDNA, which can process nearly complete transcripts at single-nucleotide resolution, and DNABERT-2. Even with one model covering 99.6 percent of transcripts in their entirety, the language models left messenger RNA prediction at an R-squared of just 0.042, and lncRNA prediction at values between 0.0002 and 0.0036. The implication is sobering: whatever the language models learned about genomic sequence, they did not uncover a hidden stability code in non-coding transcripts that simpler approaches had missed.</p>
<p>To rule out confounding, Tani stratified transcripts by predicted coding potential, testing whether a subset of lncRNAs with translated open reading frames might behave like messenger RNAs after all. They did not. Nor did differences in transcript length, GC content or sample size explain the gap: when lncRNAs and mRNAs were matched on all three characteristics and analyzed with an identical feature set, the messenger RNAs yielded an R-squared of 0.033 while the lncRNAs yielded 0.0002. Perhaps most convincingly, a signal-injection experiment established that the analytical pipeline could reliably detect effects explaining as little as 0.5 percent of the variance in half-life, a sensitivity threshold comfortably above every single lncRNA result reported in the study. If a sequence-based stability signal existed in these transcripts at even a modest level, the method would have found it.</p>
<p>The biological interpretation is as interesting as the statistical one. The mechanistic chain linking codon optimality to decay runs through translation itself: slow-moving ribosomes on unoptimized codons recruit decay factors, and the DEAD-box helicase Dhh1p in yeast monitors codon optimality directly to couple translation to destruction. Long non-coding RNAs, which are largely untranslated, simply cannot participate in this feedback loop. Their lifespans are instead governed by other forces, including RNA-binding proteins, nuclear retention mechanisms, structural elements and, in many organisms, the nonsense-mediated decay pathway, which can eliminate lncRNAs carrying premature stop codons. A growing literature documents hundreds of short-lived non-coding transcripts in mammalian cells and shows that nonsense-mediated decay restricts lncRNA expression even in RNAi-capable budding yeasts, underscoring that non-coding transcript turnover has its own logic, written in a different biochemical language.</p>
<p>The study&#8217;s broader lesson extends beyond RNA biology into the methodology of machine-learning-driven science. By re-estimating features within each cross-validation fold, Tani guarded against data leakage, the insidious practice pattern that has inflated reported performance across computational biology, and the work aligns with recent guidance on reproducibility in machine-learning applications. The bottom-line recommendation is blunt: mRNA-derived decay models should not be transferred to long non-coding RNAs without re-validation. For researchers designing RNA therapeutics, engineering synthetic transcripts or interpreting lncRNA dysregulation in disease, the message is clear. The stability of the non-coding transcriptome cannot be read from codon usage tables, and understanding how these molecules are scheduled for destruction will require looking elsewhere, at the proteins and structures that shepherd them through the cell.</p>
<p><strong>Subject of Research:</strong> Codon optimality as a predictor of mRNA half-life and its failure to predict long non-coding RNA stability</p>
<p><strong>Article Title:</strong> Codon optimality predicts mRNA half-life but does not transfer to lncRNAs</p>
<p><strong>Article References:</strong> Tani, H. (2026). Codon optimality predicts mRNA half-life but does not transfer to lncRNAs. <em>Molecular Genetics and Genomics, 301</em>(1), Article 196. <a href="https://doi.org/10.1007/s00438-026-02520-1" rel="noopener noreferrer">https://doi.org/10.1007/s00438-026-02520-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00438-026-02520-1" rel="noopener noreferrer">10.1007/s00438-026-02520-1</a></p>
<p><strong>Keywords:</strong> codon optimality, mRNA stability, long non-coding RNA, RNA decay, RNA half-life, machine learning, cross-validation, genomic language models, nonsense-mediated decay, transcriptome, RNA-binding proteins, gene regulation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204012</post-id>	</item>
		<item>
		<title>xDecoder Pushes Genomic Foundation Models Toward Personal Gene Expression Prediction</title>
		<link>https://scienmag.com/xdecoder-pushes-genomic-foundation-models-toward-personal-gene-expression-prediction/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:53:00 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[ATAC-seq]]></category>
		<category><![CDATA[challenges in personalized genomics]]></category>
		<category><![CDATA[Chromatin Accessibility]]></category>
		<category><![CDATA[cis-regulatory variants]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for gene regulation]]></category>
		<category><![CDATA[DNA-only AI models]]></category>
		<category><![CDATA[eQTL]]></category>
		<category><![CDATA[eQTL studies limitations]]></category>
		<category><![CDATA[gene expression prediction]]></category>
		<category><![CDATA[gene expression prediction in individuals]]></category>
		<category><![CDATA[genomic foundation models]]></category>
		<category><![CDATA[genomic language models]]></category>
		<category><![CDATA[GEUVADIS]]></category>
		<category><![CDATA[improving gene expression prediction accuracy]]></category>
		<category><![CDATA[personal gene expression prediction]]></category>
		<category><![CDATA[personalized genomics]]></category>
		<category><![CDATA[personalized regulatory genomics]]></category>
		<category><![CDATA[rare and novel genetic variants]]></category>
		<category><![CDATA[sequence-to-function models]]></category>
		<category><![CDATA[sequence-to-function models in genomics]]></category>
		<category><![CDATA[xDecoder]]></category>
		<category><![CDATA[xDecoder computational framework]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196611</guid>

					<description><![CDATA[A new decoding framework called xDecoder more than triples few-shot accuracy of genomic foundation models for personal gene expression prediction while exposing a persistent cross-locus generalization bottleneck in DNA-only models.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers at the University of Hong Kong and Shenzhen University has developed a new computational framework that dramatically improves how well genomic foundation models can predict gene expression in individual people, while simultaneously exposing a fundamental weakness in today&#8217;s DNA-only artificial intelligence models. The tool, called xDecoder, is described in a study published in Molecular Systems Biology and offers both a practical advance and a candid diagnosis of where the field of personalized regulatory genomics currently falls short.</p>
<p>Predicting how a person&#8217;s unique genetic makeup shapes their gene expression has long been a central goal of modern biology. Expression quantitative trait locus, or eQTL, studies have linked thousands of genetic locations to gene activity, but these statistical approaches depend on massive population datasets and are therefore largely restricted to common genetic variants. Rare and novel variants, which are often critical for understanding disease mechanisms, remain stubbornly difficult to interpret. Sequence-to-function models such as Enformer and Borzoi were designed to close this gap by using deep learning to map DNA sequence directly to regulatory outputs, and they perform well when predicting average expression across genes and tissues. Yet a series of benchmarking studies published in recent years showed that these models capture very little of the transcriptomic variation between individuals, prompting researchers to ask whether a fundamentally different approach was needed.</p>
<p>The Hong Kong team, led by Shumin Li, Ruibang Luo, and Yuanhua Huang, saw promise in a newer class of models: genomic language models, or gLMs. Unlike sequence-to-function models, which are trained on curated functional annotations across many cellular contexts, gLMs such as Evo2-7B, the Nucleotide Transformer, and Caduceus learn generalizable representations of DNA through self-supervised pretraining on raw sequence alone. Whether these models could complement or extend existing sequence-to-function approaches for personal expression prediction was an open question. xDecoder was built to answer it. The framework is a lightweight convolutional neural network decoding tower that consumes the full token-level embeddings produced by six different foundation models, three genomic language models and three sequence-to-function models, and translates them into personal gene expression predictions.</p>
<p>The researchers trained and evaluated xDecoder using paired genome and transcriptome data from the GEUVADIS consortium, which includes RNA sequencing and high-quality genotypes from lymphoblastoid cell lines of 1000 Genomes Project participants. Individualized DNA sequences were reconstructed by substituting single-nucleotide variants from phased genotypes into the reference genome. The team focused on genes with significant cis-eQTLs, reserving 200 genes for personalized training and holding out 287 genes on chromosomes 5 and 10 for stringent, locus-disjoint evaluation, with 50 individuals for training and 100 for testing. Importantly, the xDecoder architecture deliberately avoids compressing embeddings through global pooling. Because only about 0.01 percent of positions differ between any two people across a gene&#8217;s sequence, preserving positional resolution is essential, so the model aggregates local neighborhoods with small stacked depthwise-separable convolutional layers before pooling and regression, and averages predictions across the two haplotypes.</p>
<p>The study&#8217;s central finding is a tale of two learning regimes. In the few-shot setting, where models are fine-tuned on paired personal genome-expression data for genes seen during training, xDecoder delivered consistent and substantial improvements. Cross-individual prediction correlations more than tripled relative to pretrained Enformer and Borzoi and to reference-trained xDecoder variants, rising from 0.089 to 0.194 on average. For specific genes the gains were striking: for HLA-DQA1, pretrained Enformer and AlphaGenome produced negative correlations, while individually fine-tuned xDecoder models exceeded 0.3, with a Caduceus-based variant reaching 0.74. The best genomic language model embeddings even outperformed the best sequence-to-function embeddings, and rp-Caduceus successfully modeled 46 genes for which the widely used linear method PrediXcan produced identical, uninformative predictions across all individuals. Exploratory variant-level perturbation analysis further showed that high-scoring xDecoder perturbations overlapped known eQTL variants, hinting that the model can nominate candidate cis-regulatory variants in loci where fixed genotype weights fail.</p>
<p>The zero-shot picture was far less encouraging. When the models were asked to predict expression across unseen individuals for genes never encountered during training, none performed reliably. Zero-shot performance remained close to zero even as the number of training individuals was scaled from 50 up to 250, demonstrating that the limitation is not simply a matter of insufficient data. Performance did improve modestly for genes with stronger cis-genetic signal, particularly for AlphaGenome, but even in the most favorable strata, zero-shot results fell well below gene-specific supervised models such as PrediXcan, which achieved a mean Spearman correlation of 0.216. The authors characterize this as a persistent cross-locus transfer bottleneck: current sequence models fail to learn a transferable grammar of cis-regulation that generalizes from one genomic neighborhood to another. Cross-population experiments added a further caution, showing that benefits of individual-level training diminished when models trained on European individuals were tested on Yoruba individuals from Ibadan, with genes carrying higher novel variant rates and less similar linkage disequilibrium structure showing larger performance drops.</p>
<p>To diagnose what information DNA-only models are missing, the team conducted a revealing experiment with chromatin accessibility. By adding a base-resolution ATAC-seq channel to Caduceus embeddings, they significantly improved prediction for unseen genes, lifting the mean cross-individual correlation from -0.003 to 0.129 and flipping correlation direction to positive for many genes; for the gene CHST3, adding the ATAC track transformed a wrongly signed prediction into a correlation of 0.755. However, control analyses complicated the interpretation. In the unseen-gene setting, a simple unsupervised baseline that just summed ATAC signal across the same genomic window was competitive with the full model, suggesting that measured chromatin accessibility largely explains the gain. Meanwhile, AlphaGenome-predicted ATAC signals showed only modest concordance with observed accessibility across individuals, indicating that current DNA-only models cannot yet reliably infer a person&#8217;s chromatin state from sequence alone. In the seen-gene setting, by contrast, the full DNA-plus-ATAC model outperformed an ATAC-only linear baseline, showing that sequence-derived information adds genuine value when paired training data exist.</p>
<p>The authors are careful about the limits of their advance. Absolute prediction accuracy, while improved, remains modest, and the framework should not yet be interpreted as reliable enough for gene-level prediction from personal genomes in any clinical sense. The evaluation was restricted to lymphoblastoid cell lines, and extending the approach to larger cohorts, multiple tissues, and context-specific regulation remains future work. The researchers also stress that variant prioritization applications require systematic benchmarking against eQTL-supported variants, and that clinical interpretation would demand far stronger calibration and independent validation. Rather than presenting DNA-plus-ATAC inputs as a practical replacement for RNA sequencing, they frame the ATAC experiments as a diagnostic that pinpoints the key bottleneck of current DNA-only models.</p>
<p>Even so, the study sketches a concrete roadmap for the next generation of personalized regulatory models. Combining pretrained sequence encoders with variant-aware modules, using individual-level multi-omic profiles such as ATAC-seq as auxiliary supervision or intermediate regulatory-state targets, and exploiting the few-shot regime where sequence-based models already outperform fixed-variant statistical tools together form a coherent research agenda. Because sequence-based models operate directly on DNA, they can in principle evaluate perturbations outside the fixed variant-weight space of linear predictors, offering potential utility wherever rare or novel variants lack population-derived weights. The xDecoder code has been released openly on GitHub, and the authors suggest that bridging the cross-locus generalization gap, likely through multi-omic, variant-aware training, is the defining challenge that will determine whether genomic foundation models can ultimately deliver truly personalized expression prediction from DNA alone.</p>
<p><strong>Subject of Research:</strong> Few-shot personal gene expression prediction using genomic foundation models</p>
<p><strong>Article Title:</strong> xDecoder unlocks the potential of genomic foundation models for few-shot personal gene expression prediction</p>
<p><strong>Article References:</strong> Li, S., Luo, R., &amp; Huang, Y. (2026). xDecoder unlocks the potential of genomic foundation models for few-shot personal gene expression prediction. <em>Molecular Systems Biology</em>. <a href="https://doi.org/10.1038/s44320-026-00238-1" rel="noopener noreferrer">https://doi.org/10.1038/s44320-026-00238-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s44320-026-00238-1" rel="noopener noreferrer">10.1038/s44320-026-00238-1</a></p>
<p><strong>Keywords:</strong> genomic foundation models, gene expression prediction, personalized genomics, xDecoder, genomic language models, sequence-to-function models, eQTL, chromatin accessibility, ATAC-seq, deep learning, GEUVADIS, cis-regulatory variants</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196611</post-id>	</item>
	</channel>
</rss>
