<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>DNA language models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/dna-language-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 23:45:19 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>DNA language models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Large Language Models Are Learning to Read the Language of Life</title>
		<link>https://scienmag.com/large-language-models-are-learning-to-read-the-language-of-life/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 23:45:19 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI in genomics research]]></category>
		<category><![CDATA[AlphaFold]]></category>
		<category><![CDATA[bioinformatics and deep learning]]></category>
		<category><![CDATA[bioinformatics education]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[digital transformation in molecular biology]]></category>
		<category><![CDATA[DNA and protein sequence interpretation]]></category>
		<category><![CDATA[DNA language models]]></category>
		<category><![CDATA[DNA sequence analysis]]></category>
		<category><![CDATA[foundation models]]></category>
		<category><![CDATA[genomic language models]]></category>
		<category><![CDATA[genomics]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in molecular biology]]></category>
		<category><![CDATA[machine learning for life sciences]]></category>
		<category><![CDATA[natural language processing for biological data]]></category>
		<category><![CDATA[protein structure prediction]]></category>
		<category><![CDATA[RNA and protein decoding]]></category>
		<category><![CDATA[RNA secondary structure]]></category>
		<category><![CDATA[sequence-to-function prediction in biology]]></category>
		<category><![CDATA[single-cell analysis]]></category>
		<category><![CDATA[transformer architecture]]></category>
		<category><![CDATA[transformer models in genomics]]></category>
		<category><![CDATA[variant effect prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=232502</guid>

					<description><![CDATA[A new review from Wuhan University details how large language models are rapidly transforming protein, DNA, and RNA research, from structure prediction to single-cell analysis and genomic education.]]></description>
										<content:encoded><![CDATA[<p>When ChatGPT stunned the world with its command of human language, few predicted that the same underlying technology would soon be trained on DNA, RNA, and proteins. Yet a comprehensive review published in Frontiers of Digital Education by Shaopeng Li, Weiliang Fan, and Yu Zhou of Wuhan University maps out exactly how large language models, or LLMs, have swept into genomics in just four years, transforming how scientists interpret the molecular instructions that govern life. The review, published on 28 May 2025, argues that the sequential nature of biological data makes it strikingly similar to human text, and that the architectures built to master one can be repurposed to decode the other.</p>
<p>The analogy is more than a metaphor. DNA is written in an alphabet of four nucleotide letters, proteins in an alphabet of twenty amino acids, and RNA in a four-letter code of its own. Just as words derive meaning from their context in a sentence, the function of a nucleotide or amino acid depends on its position within a sequence and on the sequences that surround it. The transformer architecture, introduced in 2017 in the landmark paper Attention Is All You Need, excels precisely at capturing such long-range contextual dependencies. By applying self-attention mechanisms across millions of biological sequences, models can learn statistical patterns that reflect evolutionary constraints, structural motifs, and functional sites without any explicit programming.</p>
<p>The review organizes the field into biological foundation models, general-purpose systems pretrained on massive unlabeled sequence databases, and specialized models tailored to particular problems. On the protein side, models such as ProtTrans and the Evolutionary Scale Modeling family demonstrated that unsupervised learning on hundreds of millions of sequences yields internal representations from which structure and function can be read off. This line of work culminated in AlphaFold and its successors, which achieved atomic-level accuracy in protein structure prediction and, with AlphaFold 3, extended to predicting interactions between proteins, nucleic acids, and small molecules. Generative protein LLMs such as ProGen and ProGen2 have gone further, producing entirely novel amino acid sequences that fold into functional proteins, a capability the review highlights as one of the most consequential advances of the past decade.</p>
<p>DNA language models have followed a parallel trajectory. The Nucleotide Transformer, built and evaluated on human and multi-species genomes, showed that robust foundation models could be trained at nucleotide resolution. Subsequent work demonstrated that such models are powerful predictors of genome-wide variant effects, meaning they can estimate whether a mutation in a patient&#8217;s genome is likely to be harmful, a task central to clinical genetics. The review also points to Evo and its successor Evo 2, genome-scale models capable of modeling and designing sequences across all domains of life, from bacteria to humans. These systems operate at a scale that was unthinkable for genomics only a few years ago, processing contexts long enough to encompass entire genes and their regulatory neighborhoods.</p>
<p>RNA has proven a particularly fertile ground because its structure and function are notoriously difficult to predict from sequence alone. Classical approaches relied on thermodynamic folding algorithms, but deep learning models such as UFold and RNAformer have shown that learned representations can match or exceed physics-based methods for secondary structure prediction. More recently, language model-based deep learning approaches have achieved accurate RNA 3D structure prediction, and motif-aware pretraining strategies have produced multi-purpose RNA models that handle splicing prediction, binding site identification, and function annotation within a single framework. The review notes that self-supervised learning on millions of primary RNA sequences from dozens of vertebrate species has substantially improved sequence-based splicing prediction, with direct implications for interpreting disease-causing mutations in untranslated regions and splice sites.</p>
<p>Beyond the three canonical sequence types, the review surveys a rapidly growing ecosystem of specialized applications. Interaction prediction models forecast which proteins bind which RNAs, which RNAs pair with other RNAs, and how mutations perturb these contacts. Structure prediction systems now cover proteins, RNA, and biomolecular complexes at near-experimental accuracy. Perhaps most striking is the single-cell revolution: foundation models such as scGPT, scBERT, and Geneformer treat gene expression profiles across thousands of individual cells as a kind of language, learning universal representations that support cell type annotation, perturbation response prediction, and the integration of massive transcriptomic datasets. Studies assessing GPT-4 for cell type annotation suggest that general-purpose chatbots can also contribute, though specialized models currently hold the edge on most benchmarks.</p>
<p>The authors are candid about the obstacles. Hallucination, the tendency of generative models to produce plausible but false outputs, is a documented failure mode of natural language generation that carries serious risks when outputs inform drug design or clinical interpretation. Benchmarking remains immature: evaluation suites such as DART-Eval for regulatory DNA and PertEval-scFM for perturbation prediction reveal that performance gains do not always transfer across datasets and cell types. Computational cost is another barrier, since genome-scale attention models demand enormous memory, motivating architectural innovations such as the Hyena hierarchy and Mamba-based state space models that the review identifies as promising directions for long-context biological modeling. Data quality, batch effects in single-cell experiments, and the interpretability of learned representations round out the open challenges.</p>
<p>What distinguishes this review from many surveys of artificial intelligence in biology is its educational dimension. Because it appears in a journal focused on digital education, the authors devote substantial attention to integrating LLMs into genomics teaching and learning. They propose practical projects in which students fine-tune pretrained models on small datasets, use conversational agents to simplify bioinformatics workflows, and learn to critically evaluate model outputs. Tools such as BioMANIA, which allows researchers to interrogate bioinformatics pipelines through natural language conversation, and BioCoder, a benchmark for bioinformatics code generation, illustrate how LLMs can lower the technical barriers that traditionally separate biologists from computational analysis. The authors argue that fluency in these tools is becoming as essential to the modern life scientist as statistical literacy.</p>
<p>The trajectory the review describes is remarkably compressed. In roughly four years, language models have progressed from proof-of-concept embeddings to systems that simulate 500 million years of evolutionary history, design functional proteins, predict the pathogenicity of genetic variants, and annotate cell atlases containing millions of cells. The Wuhan University team frames the current moment as an inflection point: the same scaling laws and reasoning enhancements that drive progress in general-purpose AI, including reinforcement learning approaches that sharpen chain-of-thought reasoning, are poised to accelerate biological discovery further. If the challenges of reliability, evaluation, and interpretability can be met, the authors suggest, the fusion of language models and genomics may redefine both how biology is researched and how the next generation of scientists is trained, turning the genome from an inscrutable code into a text that machines can read, annotate, and even rewrite.</p>
<p><strong>Subject of Research:</strong> Applications of large language models in genomics, including biological foundation models for protein, DNA, and RNA sequence analysis</p>
<p><strong>Article Title:</strong> AI-Empowered Genome Decoding: Applications of Large Language Models in Genomics</p>
<p><strong>Article References:</strong> Li, S., Fan, W., &amp; Zhou, Y. (2025). AI-Empowered Genome Decoding: Applications of Large Language Models in Genomics. <em>Frontiers of Digital Education, 2</em>(1), Article 14. <a href="https://doi.org/10.1007/s44366-025-0051-1" rel="noopener noreferrer">https://doi.org/10.1007/s44366-025-0051-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44366-025-0051-1" rel="noopener noreferrer">10.1007/s44366-025-0051-1</a></p>
<p><strong>Keywords:</strong> large language models, genomics, deep learning, protein structure prediction, DNA language models, RNA secondary structure, single-cell analysis, foundation models, AlphaFold, transformer architecture, bioinformatics education, variant effect prediction</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">232502</post-id>	</item>
		<item>
		<title>Interpretable DNABERT Models Predict DNA Replication Origins in S. cerevisiae</title>
		<link>https://scienmag.com/interpretable-dnabert-models-predict-dna-replication-origins-in-s-cerevisiae/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 29 Aug 2026 06:37:27 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI in genomics]]></category>
		<category><![CDATA[AI model explainability]]></category>
		<category><![CDATA[AI model interpretability]]></category>
		<category><![CDATA[analysis of replication origin signals]]></category>
		<category><![CDATA[artificial intelligence in biology]]></category>
		<category><![CDATA[biological signal extraction]]></category>
		<category><![CDATA[biological signal recognition by AI models]]></category>
		<category><![CDATA[DNA language models]]></category>
		<category><![CDATA[DNA replication origins]]></category>
		<category><![CDATA[DNA replication origins in yeast]]></category>
		<category><![CDATA[DNABERT]]></category>
		<category><![CDATA[DNABERT for genomic sequence analysis]]></category>
		<category><![CDATA[genomic feature extraction using language models]]></category>
		<category><![CDATA[genomic sequence analysis]]></category>
		<category><![CDATA[interpretability of machine learning]]></category>
		<category><![CDATA[machine learning in DNA replication]]></category>
		<category><![CDATA[open interpretability of AI in biology]]></category>
		<category><![CDATA[prediction of replication initiation sites]]></category>
		<category><![CDATA[replication origin prediction]]></category>
		<category><![CDATA[Saccharomyces cerevisiae]]></category>
		<category><![CDATA[transformer-based language models in genomics]]></category>
		<category><![CDATA[transformer-based models]]></category>
		<category><![CDATA[understanding DNA regulatory elements]]></category>
		<category><![CDATA[yeast genome replication mechanisms]]></category>
		<guid isPermaLink="false">https://scienmag.com/interpretable-dnabert-models-predict-dna-replication-origins-in-s-cerevisiae/</guid>

					<description><![CDATA[DNA’s replication machinery may be ancient, but one of the newest tools for studying it borrows its logic from artificial intelligence trained to process language. In a study of budding yeast, researchers fine-tuned two DNA “language models”—DNABERT and DNABERT-2—to predict where chromosomes begin copying themselves. More importantly, they opened the models’ black boxes to determine [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>DNA’s replication machinery may be ancient, but one of the newest tools for studying it borrows its logic from artificial intelligence trained to process language. In a study of budding yeast, researchers fine-tuned two DNA “language models”—DNABERT and DNABERT-2—to predict where chromosomes begin copying themselves. More importantly, they opened the models’ black boxes to determine whether the patterns guiding their predictions corresponded to real biological signals. The results suggest that transformer-based AI can recognize replication origins while also revealing which stretches of DNA influence its decisions, although the way those signals emerge depends strongly on how the models divide DNA into computational “words.”</p>
<p>DNA replication begins at specific genomic locations known as replication origins. From each origin, molecular machines assemble and copy the surrounding chromosome in both directions. Bacteria often use a single origin, but eukaryotic chromosomes are much larger and typically require many starting points. The researchers focused on <em>Saccharomyces cerevisiae</em>, or budding yeast, whose replication origins are unusually well characterized. In this organism, origins occur within autonomous replication sequences, or ARS regions, generally about 100 to 200 base pairs long. These regions contain several functional elements, including the A element and B1, which together help form the principal binding site for the origin recognition complex, or ORC. The B2 element may provide an additional ORC-binding site or a platform for components of the replicative helicase.</p>
<p>A central sequence signal in yeast origins is the 11-base-pair ARS consensus sequence, or ACS. It is commonly represented as WTTTAYRTTTW, where W means either adenine or thymine and Y means cytosine or thymine. An extended version spans 17 base pairs and is written WWW-WTTTAYRTTTW-GTT. Yet the ACS alone cannot explain which origins actually function. The yeast genome contains roughly 12,000 matches to ACS-like motifs, but only about 500 are functional under normal conditions. Origin activity is therefore shaped by additional factors, including chromatin accessibility, transcription that can interfere with initiation, and secondary DNA features. This mismatch between abundant sequence motifs and comparatively rare active origins makes replication-origin prediction a demanding test for machine learning.</p>
<p>Earlier computational approaches often depended on labor-intensive feature engineering. Researchers had to convert sequences into numerical descriptions of nucleotide composition, DNA shape, physical properties, or sequence order before training support-vector machines, random forests, or other classifiers. Deep-learning methods reduced some of that manual work but frequently required models to be trained from scratch, and their learned features could be difficult to connect to recognizable biological mechanisms. The new study instead used pretrained transformer models. Transformers process sequences through self-attention, a mechanism that allows each input segment to weigh information from other segments while constructing a contextual representation. In principle, this enables a model to detect both short motifs and relationships between separated regions without being told in advance which biological features to search for.</p>
<p>DNABERT and DNABERT-2 share the basic BERT architecture, but they read DNA differently. DNABERT was pretrained on the human genome and converts sequences into overlapping four-base units, or 4-mers. Because adjacent tokens overlap, a DNA sequence is represented through a dense series of partially shared local windows. DNABERT-2 was pretrained on genomes from multiple species, including yeast, and uses byte-pair encoding, or BPE. BPE builds a vocabulary by repeatedly merging frequently occurring sequence segments, producing tokens of variable length. The newer model has about 117 million parameters and a vocabulary of 4,096 tokens, compared with approximately 86 million parameters for DNABERT. The difference in size comes mainly from the larger embedding matrix needed for BPE, not from deeper or wider transformer layers.</p>
<p>To train and test the models, the team assembled balanced datasets using 325 experimentally confirmed yeast origins from the curated OriDB resource. Each origin sequence was placed into a standardized 500-base-pair window. Shorter origins were extended with genuine neighboring genomic DNA, with the extra sequence distributed randomly between the two sides. This prevented the model from learning a trivial rule based on where the origin appeared within the window. The investigators created two contrasting negative datasets. In the Random-Neg set, non-origin sequences were sampled from genomic regions that did not overlap known origins. In the more difficult ACS-Neg set, negative sequences were selected from approximately 12,000 ACS matches that do not function as origins. Positive and negative sequences in the latter set therefore contained ACS-like signals of comparable strength, forcing the models to search for additional distinctions.</p>
<p>Both models were fine-tuned for binary classification: origin-containing sequence versus non-origin sequence. The DNA entered each model with special beginning and end tokens, while DNABERT-2 sequences were padded when its variable-length tokenization required equal input lengths. A classification layer attached to the representation of the initial classification token produced the probability that a sequence belonged to the origin class. The researchers used seven independent 70/10/20 training, validation, and test splits, selecting checkpoints based on validation accuracy and area under the receiver operating characteristic curve rather than simply taking the final training epoch. This design was intended to test robustness rather than maximize a single headline score. They also performed chromosome-based splitting to check that random partitioning had not artificially inflated performance.</p>
<p>DNABERT achieved an average test accuracy of 0.83 and an area under the curve of 0.90 on the easier Random-Neg dataset. DNABERT-2 reached 0.81 accuracy and 0.82 area under the curve. When ACS-rich non-origins were used as negatives, both models achieved approximately 0.72 accuracy, demonstrating that the task became substantially harder once the most obvious motif-based distinction was removed. The results indicate that the models were not merely detecting the presence of an ACS match. They retained some ability to distinguish functional origins from nonfunctional ACS sites, although the moderate accuracy also shows that sequence alone does not fully determine origin activity. Chromatin state, transcriptional context, and other cellular features remain outside the information available in the sequence windows.</p>
<p>The most revealing differences appeared when the researchers examined how the models arrived at their decisions. For DNABERT, they extracted attention scores associated with the classification token and projected those token-level values back onto individual nucleotides. Sharp attention peaks repeatedly fell within experimentally annotated origin regions rather than being distributed randomly across the 500-base-pair windows. The team collected short, 20-base-pair fragments around the strongest peaks and analyzed them with MEME, a probabilistic motif-discovery program that uses expectation-maximization to align variable motif instances. In test sequences from the Random-Neg dataset, the resulting motif was TTTTTWTTTATRTTT, with an E-value of 2.6 × 10−6, closely matching the known ACS pattern. A related training-set motif, TATATTTATRTWTWT, had an E-value of 2.3 × 10−32. In the ACS-Neg condition, where ACS motifs appeared in both classes, no comparably significant test-set motif emerged, consistent with the greater difficulty of the classification problem.</p>
<p>Attention maps from DNABERT-2 were less straightforward to interpret, so the researchers used perturbation experiments and Shapley additive explanations, or SHAP. Perturbation analysis asks what happens when a portion of the input is removed, altered, or rearranged. Eliminating ACS motifs from origin sequences changed DNABERT-2’s predictions, confirming that the motif contributed to classification. Randomly shuffling the model’s BPE tokens produced another important result: token identity appeared to matter strongly, while the precise order of some tokens mattered less than expected. SHAP analysis, repeated 30 times to reduce the randomness of its estimates, identified tokens that consistently supported origin or non-origin predictions. The team combined these findings into an AT-index incorporating overall adenine-thymine content, frequencies of AT-rich motifs, the longest uninterrupted AT run, and alternating AT runs. The index provided a compact way to quantify the AT-rich character of sequences associated with the model’s decisions, but the authors caution that SHAP values describe model behavior relative to its training background, not universal causal rules governing DNA replication.</p>
<p>The study also used shuffled sequences to test whether the models depended mainly on nucleotide composition or on the arrangement of bases. In one control, nucleotides within each positive sequence were randomly rearranged, preserving overall composition but destroying sequence order. In another, the sequence was divided into five-base-pair blocks and the blocks were shuffled, preserving local patterns while disrupting larger-scale organization. These experiments were designed to reveal whether the classifiers had learned broad compositional differences between origins and non-origins or more structured sequence information. The models’ near-random performance on shuffled data before task-specific fine-tuning showed that pretraining alone did not automatically confer the ability to distinguish intact from rearranged origin sequences. That capability emerged during adaptation to the replication-origin task.</p>
<p>Training from scratch produced weaker results than fine-tuning the pretrained systems. On the Random-Neg dataset, DNABERT trained from scratch averaged 0.71 accuracy, while DNABERT-2 averaged 0.60. The pretrained models used without fine-tuning performed at 0.37 and 0.55 accuracy, respectively, showing that general DNA representations alone were insufficient. Fine-tuning was therefore essential, but pretraining still supplied a useful starting point. The researchers emphasize that their goal was not to declare one architecture the universal winner. Instead, they wanted to determine whether genomic language models could reduce manual feature engineering while producing biologically interpretable signals. On that measure, the two systems behaved differently: overlapping k-mers gave DNABERT attention maps that more visibly reflected known motifs, whereas BPE tokenization in DNABERT-2 produced less transparent attention but supported alternative attribution analyses.</p>
<p>The findings matter because replication origins are not simply strings containing one magic sequence. In yeast, the ACS is an important anchor for ORC recognition, yet thousands of similar matches fail to initiate replication. A useful predictive model must therefore identify combinations of sequence properties while avoiding shortcuts created by biased negative examples. By deliberately challenging the models with ACS-containing non-origins, the researchers showed that transformer classifiers can recover signals beyond the canonical motif, though not perfectly. The work also illustrates a broader principle for biological AI: predictive accuracy is only part of the result. Tokenization, dataset construction, and explanation method can substantially influence what a model appears to have learned. As genomic language models are applied to human and other eukaryotic genomes—where replication origins are less sharply defined—the ability to distinguish meaningful biological signals from statistical artifacts will be crucial.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Prediction and interpretability of DNA replication origins in <em>Saccharomyces cerevisiae</em> using DNABERT and DNABERT-2</p>
<p><strong>Article Title:</strong> Interpretable prediction of DNA replication origins in <em>S. cerevisiae</em> using DNABERT and DNABERT-2</p>
<p><strong>Article References:</strong> Piroozeh, Z., Akerman, I., Kalinina, O. V., Kesselheim, S., &amp; Bazarova, A. (2026). Interpretable prediction of DNA replication origins in S. cerevisiae using DNABERT and DNABERT-2. <em>BMC Bioinformatics, 27</em>(1), Article 157. <a href="https://doi.org/10.1186/s12859-026-06562-5" target="_blank" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06562-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06562-5" target="_blank" rel="noopener noreferrer">10.1186/s12859-026-06562-5</a></p>
<p><strong>Keywords:</strong> DNA replication, replication origins, <em>Saccharomyces cerevisiae</em>, DNABERT, DNABERT-2, transformer models, genomic language models, explainable artificial intelligence, SHAP, ACS motif</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">184506</post-id>	</item>
		<item>
		<title>Dynamic Cross-Strand Interactions Boost DNA Language Models</title>
		<link>https://scienmag.com/dynamic-cross-strand-interactions-boost-dna-language-models/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 04 Jun 2026 11:40:26 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced genomic AI techniques]]></category>
		<category><![CDATA[AI in genomics]]></category>
		<category><![CDATA[complementary DNA strand encoding]]></category>
		<category><![CDATA[CrossDNA language model]]></category>
		<category><![CDATA[DNA double helix communication]]></category>
		<category><![CDATA[DNA language models]]></category>
		<category><![CDATA[DNA repair mechanisms]]></category>
		<category><![CDATA[DNA replication dynamics]]></category>
		<category><![CDATA[dynamic cross-strand DNA interactions]]></category>
		<category><![CDATA[genome function decoding]]></category>
		<category><![CDATA[genomic sequence modeling]]></category>
		<category><![CDATA[transcription regulation modeling]]></category>
		<guid isPermaLink="false">https://scienmag.com/dynamic-cross-strand-interactions-boost-dna-language-models/</guid>

					<description><![CDATA[In a groundbreaking development at the intersection of genomics and artificial intelligence, a team of scientists has unveiled a novel approach to DNA sequence language modeling that promises to revolutionize how we interpret the human genome. Traditional DNA sequence models have typically either analyzed genomic data directionally—as if reading through a text—or applied static, approximative [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking development at the intersection of genomics and artificial intelligence, a team of scientists has unveiled a novel approach to DNA sequence language modeling that promises to revolutionize how we interpret the human genome. Traditional DNA sequence models have typically either analyzed genomic data directionally—as if reading through a text—or applied static, approximative methods to simulate the interactions between the two complementary strands of the DNA double helix. However, these strategies fall short in capturing the rich, dynamic exchanges that naturally occur between DNA strands in living cells. Addressing this critical gap, researchers have introduced CrossDNA, an innovative language model designed explicitly to encode and learn from the dynamic interplay between both strands of DNA.</p>
<p>In biological systems, the information encoded within the DNA duplex is not a mere linear script but a complex network of interactions where each strand influences and coordinates with its complement. This physical and functional coupling is essential, orchestrating genomic processes such as transcription regulation, replication, and DNA repair. Hence, the ability to model these cross-strand relationships dynamically offers a powerful and nuanced way to decode genomic function with unprecedented fidelity. CrossDNA takes a bold step forward by explicitly modeling these relationships rather than relying on implicit or static approximations.</p>
<p>The architecture of CrossDNA is notably distinctive. It employs a dual-branch framework in which the model alternates between processing forward and reverse-complement segments of DNA sequences. This design emulates the natural duplex structure of DNA, providing the model with both “views” of the genomic code and forcing it to learn the interplay across strands actively. Moreover, CrossDNA facilitates explicit interstrand communication through a lightweight but highly effective cross-strand communication module. This feature enables real-time information sharing between the forward and reverse branches during the learning process, ensuring that context-dependent interactions are captured dynamically rather than treating strands as isolated entities.</p>
<p>A significant technical challenge when working with genomic sequences is accommodating their length and contextual dependencies. Genomic regulatory elements can span thousands of base pairs and require models to attend to vast, complex sequence contexts. To address this, the developers of CrossDNA ingeniously combined a recurrent long-context backbone with sliding-window attention mechanisms. This hybrid approach allows the model to maintain an extensive memory of the sequence context while efficiently focusing on local relevant regions. As a result, CrossDNA achieves a new level of long-range genomic understanding that surpasses existing models’ capabilities.</p>
<p>The performance gains offered by CrossDNA are not merely theoretical. When benchmarked across a variety of genomics prediction tasks—including enhancer element identification, transcription factor binding prediction, and non-coding variant prioritization—the model consistently outperformed traditional DNA language models. Particularly notable is its performance on enhancer prediction, where the ability to model cross-strand interactions directly correlates with improved robustness to sequence orientation changes. These findings underscore the functional relevance of explicitly modeling DNA as a duplex rather than as a one-dimensional sequence or a simplistic double-complement symmetry.</p>
<p>One of the most striking aspects of CrossDNA is its parameter efficiency. While many state-of-the-art DNA foundation models contain hundreds of millions of parameters, CrossDNA achieves comparable—and often superior—predictive performance with only a fraction of the parameter count, in the million-parameter scale. This streamlined design not only accelerates training and inference but also enhances the model&#8217;s accessibility for broader scientific use, especially where computational resources might be constrained. It represents a paradigm shift towards more biologically-grounded and computationally sustainable genomic AI models.</p>
<p>Beyond improving model metrics, CrossDNA opens new avenues for interpretation and discovery within genomics. By capturing explicit dynamic cross-strand interactions, it provides a framework to better understand the regulatory logic underlying gene expression and chromatin organization. This could lead to the identification of novel regulatory elements that have been elusive to previous models and experimental assays. Additionally, CrossDNA’s capability to prioritize disease-associated non-coding variants promises to accelerate the interpretation of human genetic variation, facilitating advances in personalized medicine and genomic diagnostics.</p>
<p>The design principles behind CrossDNA also highlight the importance of mimicking biological reality within computational models. Many earlier efforts tried to impose reverse-complement symmetry or strand equivalence through data augmentation or static equivariant transformations. While useful, these approaches inevitably gloss over the dynamic, context-dependent nuance of real DNA strand interactions. CrossDNA’s approach to explicitly and iteratively learning cross-strand dependencies reflects an important conceptual leap, treating DNA as a fundamentally duplex molecular entity rather than as two separate strands.</p>
<p>In terms of technical implementation, CrossDNA’s cross-strand communication module is a lightweight yet powerful component that acts as a bridge transmitting information between the dual branches. This module dynamically integrates contextual signals during training, allowing each branch to incorporate what the other learns in a way that mirrors physical strand interactions. The synergy derived from this interbranch communication is essential for the model&#8217;s superior performance and ability to understand complex genomic structures.</p>
<p>The recurrent long-context architecture embedded into CrossDNA deserves special mention as well. Long-range dependencies in DNA sequences pose profound challenges due to the sheer length and complexity of genetic material. The combination of a recurrent backbone with sliding-window attention ensures that the model can both hold onto historical context and prioritize immediate, biologically relevant sequence patterns. This architecture mitigates the memory bottlenecks and computational inefficiencies that typically plague large sequence models, charting a path forward in genomic deep learning.</p>
<p>Perhaps most excitingly, CrossDNA transforms the concept of DNA language modeling into a more faithful analog of biological reality. By explicitly modeling cross-strand interactions dynamically, it transcends previous approximations that reduced the duplex DNA to unidirectional strings or symmetrical pairs. This leap forwards will not only enhance computational genomics but may also deepen our fundamental understanding of DNA’s role as an information carrier within the cell.</p>
<p>In practical terms, the advent of CrossDNA paves the way for more reliable and interpretable genomic prediction tools. Researchers investigating regulatory element functions, epigenetic markers, and mutation impacts will benefit from this improved modeling fidelity. Clinical geneticists tasked with identifying pathogenic variants in non-coding regions—an area historically challenging due to data complexity—now have a powerful computational ally that integrates contextual nuances from both DNA strands simultaneously.</p>
<p>Looking ahead, the interdisciplinary team behind CrossDNA has laid a foundation that could extend beyond human genomics. The principles underpinning cross-strand modeling may find applications in broader biological sequence analysis, including RNA duplexes, protein-DNA interactions, and even synthetic biology. This creates exciting possibilities for AI-driven innovation rooted in biomolecular structure and function.</p>
<p>Moreover, CrossDNA exemplifies a successful marriage of biological insight with AI techniques, showcasing how domain expertise can guide architectural decisions for transformative results. The model&#8217;s ability to efficiently leverage cross-strand information without ballooning parameter counts sets a precedent for future genomics models that balance complexity and interpretability with resource constraints.</p>
<p>In sum, CrossDNA represents a paradigm shift in the functional interpretation of genomic sequences by embracing the inherently duplex nature of DNA. Its explicit, dynamic modeling of cross-strand interactions, combined with innovative architecture for long-context handling and parameter efficiency, establishes new benchmarks in the field. This breakthrough has profound implications for genetics, molecular biology, and precision medicine, positioning CrossDNA as a pioneering tool in the new age of genome interpretation fueled by artificial intelligence.</p>
<hr />
<p><strong>Subject of Research:</strong><br />
Innovative DNA sequence language modeling focusing on explicit, dynamic cross-strand interactions within the DNA duplex to enhance genomic function interpretation and prediction accuracy.</p>
<p><strong>Article Title:</strong><br />
Explicit dynamic cross-strand interactions for DNA sequence language modelling.</p>
<p><strong>Article References:</strong><br />
Yang, C., Liu, Y., Ling, L. et al. Explicit dynamic cross-strand interactions for DNA sequence language modelling. Nat Mach Intell (2026). <a href="https://doi.org/10.1038/s42256-026-01249-1">https://doi.org/10.1038/s42256-026-01249-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s42256-026-01249-1">https://doi.org/10.1038/s42256-026-01249-1</a></p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">163809</post-id>	</item>
	</channel>
</rss>
