<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>deep learning in genomics &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/deep-learning-in-genomics/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 13 Sep 2026 00:52:46 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>deep learning in genomics &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Language Model Learns the Grammar of RNA Sequences</title>
		<link>https://scienmag.com/ai-language-model-learns-the-grammar-of-rna-sequences/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 00:52:46 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI language models for genetic sequences]]></category>
		<category><![CDATA[AI-driven understanding of ribonucleic acid]]></category>
		<category><![CDATA[computational biology]]></category>
		<category><![CDATA[deep learning in genomics]]></category>
		<category><![CDATA[embeddings]]></category>
		<category><![CDATA[language models]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for RNA annotation]]></category>
		<category><![CDATA[natural language processing for RNA sequences]]></category>
		<category><![CDATA[non-coding RNA]]></category>
		<category><![CDATA[NucleicBERT]]></category>
		<category><![CDATA[NucleicBERT transformer model]]></category>
		<category><![CDATA[predicting RNA roles with neural networks]]></category>
		<category><![CDATA[RNA]]></category>
		<category><![CDATA[RNA sequence analysis]]></category>
		<category><![CDATA[RNA structure]]></category>
		<category><![CDATA[RNA structure-function prediction]]></category>
		<category><![CDATA[RNA therapeutics]]></category>
		<category><![CDATA[self-supervised learning]]></category>
		<category><![CDATA[self-supervised learning in molecular biology]]></category>
		<category><![CDATA[sequence biology]]></category>
		<category><![CDATA[sequence-structure relationship in RNA]]></category>
		<category><![CDATA[transformers]]></category>
		<category><![CDATA[unsupervised learning in bioinformatics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200260</guid>

					<description><![CDATA[A self-supervised language model called NucleicBERT offers researchers a new computational lens on the vast and poorly charted space of RNA sequences.]]></description>
										<content:encoded><![CDATA[<p>Ribonucleic acid has spent decades in the shadow of DNA and proteins, treated by many molecular biologists as a humble courier, a disposable intermediate in the flow of genetic information from gene to protein. That view has collapsed under the weight of discovery. RNA is now known to catalyse chemical reactions, silence genes, scaffold molecular machines, tune translation and orchestrating development, and every one of those functions is written in the language of its sequence. Yet compared with proteins, where decades of structural and evolutionary data have taught researchers to read amino-acid patterns, the sequence-structure-function logic of RNA remains largely opaque. A new study published in Nature Machine Intelligence argues that the fastest route to fluency in this language may come from an unlikely teacher: the same family of self-supervised neural networks that learned to write prose.</p>
<p>The system, called NucleicBERT, applies a transformer-based language model to ribonucleic acid sequences, training it to predict masked positions in nucleotide strings drawn from large public databases. The approach deliberately avoids labels. Instead of being told which sequences are ribozymes, which are microRNAs, or which bind particular proteins, the model is simply asked to fill in the blanks across millions of natural sequences. In doing so, it is forced to internalise the statistical regularities of real RNA: which nucleotides tend to co-occur, which motifs recur across distant branches of life, and which combinations essentially never appear. Those patterns, the authors contend, encode a compressed representation of the physical and evolutionary constraints that shape functional RNA.</p>
<p>The technical foundation is the bidirectional encoder architecture popularised by models such as BERT. In natural language, such models read text in both directions and learn contextual embeddings, so that the meaning of a word depends on its neighbours. NucleicBERT imports that idea wholesale into molecular biology. Each nucleotide in an RNA sequence is treated as a token, and the encoder produces a vector for every position that reflects its biological context. A cytosine embedded in a stem-loop of a transfer RNA acquires a different representation from the same cytosine sitting in the loop of a riboswitch, even though the raw letter is identical. This context sensitivity is precisely what hand-crafted features, position-weight matrices and simple motif searches have historically lacked.</p>
<p>Pretraining proceeds with a masked-language objective. Random positions in each training sequence are hidden, and the model must reconstruct them from surrounding context. Because the training corpus spans diverse RNA families and organisms, the network cannot succeed by memorising shallow patterns; it must capture deeper regularities such as compensatory mutations in paired regions, conserved loops, and the compositional biases of different RNA classes. The resulting embeddings can then be transferred downstream: a relatively small amount of labelled data is sufficient to fine-tune the pretrained network for specific prediction tasks, a strategy that has transformed fields from computer vision to protein biochemistry.</p>
<p>The practical payoff comes in the form of benchmark performance on tasks that matter to RNA biologists. According to the paper, NucleicBERT embeddings improve predictive accuracy on problems including the classification of non-coding RNA families, the identification of RNA-binding protein sites, and the assessment of sequence variants that disrupt splicing or translation. In each case the pretrained model outperforms baselines trained from scratch on the same labelled data, and the advantage is largest precisely where labelled examples are scarcest. That pattern is the classic signature of useful pretraining: the model arrives at a task already fluent in the underlying vocabulary, so supervision only needs to teach the final grammar.</p>
<p>What makes the work conceptually significant is not merely the benchmark numbers but the interpretability experiments layered on top of them. The authors probe what the model has learned by examining attention patterns and embedding geometry. Sequences with related structures and functions cluster together in the embedding space even when their nucleotide identities differ substantially, suggesting the model has discovered homology that raw sequence comparison misses. Attention heads, the internal components that let a transformer weigh relationships between positions, turn out to concentrate on regions that biologists recognise as structurally or functionally meaningful, such as paired stems and conserved catalytic motifs. In effect, the network rediscovers, from raw data alone, some of the hard-won knowledge that RNA biochemists assembled over half a century.</p>
<p>The study also confronts one of the central puzzles of RNA biology: the sheer size of sequence space. An RNA molecule of only 100 nucleotides has 4 to the power of 100 possible sequences, a number that dwarfs the number of atoms in the observable universe. Natural RNA occupies a vanishingly sparse subset of that space, organised into families shaped by common ancestry and common physics. Language models are, in a formal sense, tools for modelling the distribution of data, and NucleicBERT can therefore be read as a statistical map of where functional RNA lives within the vast combinatorial wilderness. Sequences the model assigns high likelihood are, heuristically, sequences that look like biology; sequences it assigns low likelihood are candidates for exotic synthetic designs, or for failure.</p>
<p>That map has immediate applications in engineering. RNA therapeutics, from messenger RNA vaccines to small interfering RNAs and antisense oligonucleotides, all depend on the properties of sequence: how stably a molecule folds, how efficiently it is translated, how recognisable it is to the innate immune system, and how long it survives in the cell. The authors report that NucleicBERT representations correlate with measurable properties such as secondary-structure stability and expression level, offering drug developers a way to screen and optimise candidate sequences in silico before expensive synthesis and testing begin. The same representations can guide the design of synthetic riboswitches and regulatory elements for synthetic biology, where designers currently iterate through costly build-and-test cycles.</p>
<p>The researchers are candid about limitations. RNA databases are biased towards well-studied model organisms and abundant RNA classes, so the model&#8217;s fluency is strongest where data are richest and weakest for rare transcripts and poorly characterised clades. The masked-language objective captures linear sequence context directly and higher-order structure only indirectly, so tasks that hinge on detailed three-dimensional folding may still require complementary physics-based or structure-specific models. And like all deep networks, NucleicBERT offers correlations rather than mechanisms: its embeddings are a powerful substrate for prediction, but turning them into causal explanations of why a particular fold catalyses a particular reaction remains future work. The authors frame the model not as a replacement for biochemical experiment but as a hypothesis engine that tells experimentalists where to look.</p>
<p>Even with those caveats, the arrival of a mature nucleic-acid language model marks a turning point in how the life sciences approach sequence data. For twenty years, genome annotation has leaned on alignment-based tools that compare new sequences against known ones, a strategy that fails for molecules with no recognisable relatives. Self-supervised models offer a different epistemology: knowledge distilled from the totality of observed sequences, applicable even to orphans with no evolutionary cousins. As sequencing technologies continue to generate data far faster than any human can annotate them, systems like NucleicBERT are likely to become standard equipment in the computational biology toolkit, reading the genome&#8217;s least understood language at a pace no human reader could match and pointing the way to RNA molecules that biology has not yet invented.</p>
<p><strong>Subject of Research:</strong> Self-supervised language modelling of RNA sequence space with the NucleicBERT neural network</p>
<p><strong>Article Title:</strong> NucleicBERT interprets RNA sequence space through self-supervised language modelling</p>
<p><strong>Article References:</strong> Upadhyay, U., Herold, J., Götz, M., &amp; Schug, A. (2026). NucleicBERT interprets RNA sequence space through self-supervised language modelling. <em>Nature Machine Intelligence</em>. <a href="https://doi.org/10.1038/s42256-026-01295-9" rel="noopener noreferrer">https://doi.org/10.1038/s42256-026-01295-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s42256-026-01295-9" rel="noopener noreferrer">10.1038/s42256-026-01295-9</a></p>
<p><strong>Keywords:</strong> NucleicBERT, RNA, self-supervised learning, language models, machine learning, transformers, non-coding RNA, RNA therapeutics, sequence biology, computational biology, embeddings, RNA structure</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200260</post-id>	</item>
		<item>
		<title>Ancient DNA Word Vocabularies Govern How Cell Types Evolve Across Species</title>
		<link>https://scienmag.com/ancient-dna-word-vocabularies-govern-how-cell-types-evolve-across-species/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 12:54:02 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[ancestral DNA regulatory elements]]></category>
		<category><![CDATA[cell type evolution]]></category>
		<category><![CDATA[cellular diversity across species]]></category>
		<category><![CDATA[Chromatin Accessibility]]></category>
		<category><![CDATA[conserved DNA sequences]]></category>
		<category><![CDATA[cross-species cell type comparison]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning in genomics]]></category>
		<category><![CDATA[developmental homology]]></category>
		<category><![CDATA[evolutionary biology of cell types]]></category>
		<category><![CDATA[flatworms]]></category>
		<category><![CDATA[Gene regulation]]></category>
		<category><![CDATA[gene regulatory networks]]></category>
		<category><![CDATA[homology in cell types]]></category>
		<category><![CDATA[motif vocabularies]]></category>
		<category><![CDATA[Nature Ecology & Evolution]]></category>
		<category><![CDATA[regulatory genome evolution]]></category>
		<category><![CDATA[regulatory syntax]]></category>
		<category><![CDATA[single-cell multi-omics]]></category>
		<category><![CDATA[single-nucleus multi-omics]]></category>
		<category><![CDATA[transcription factors]]></category>
		<category><![CDATA[vertebrate and invertebrate genome regulation]]></category>
		<category><![CDATA[vertebrates]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194519</guid>

					<description><![CDATA[Deep-learning analysis of single-nucleus multi-omic data from flatworms and vertebrates reveals that conserved motif vocabularies maintain cell type family identities across hundreds of millions of years while individual cell type regulatory programmes evolve rapidly.]]></description>
										<content:encoded><![CDATA[<p>DNA word vocabularies conserved across half a billion years of animal evolution are revealed to act as the controlling factors that govern which parts of the genome are opened up for reading in each distinct cell type. This finding, published in Nature Ecology &amp; Evolution, emerges from a sophisticated combination of single-nucleus multi-omic sequencing and deep-learning models applied to flatworms and vertebrates, offering an unprecedented view into the regulatory logic that shapes cellular diversity across vastly divergent species. The study suggests that while individual cell types evolve their own regulatory programmes at a rapid rate, the family-level identity of cells is maintained collectively through large pools of conserved regulatory factors, drawing a parallel to the developmental principle of homology.</p>
<p>The research team, led by investigators at Stanford University including Chew Chai, Jesse Gibson, Pengyang Li, Brennan D. McDonald, Anusri Pampari, Aman Patel, Anshul Kundaje, and Bo Wang, set out to address a fundamental question in evolutionary biology: what mechanisms define and maintain families of related cell types across deep evolutionary time? Cell types can be organized into related families based on their functional and molecular properties, yet the regulatory underpinnings that sustain these families across hundreds of millions of years of divergence have remained largely unknown. By integrating single-nucleus multi-omic sequencing data from three species of flatworms and comparing it with vertebrate data, the researchers were able to identify hundreds of sequence motifs that dictate chromatin accessibility and partition into distinct, conserved sets referred to as vocabularies.</p>
<p>The concept of motif vocabularies represents a significant conceptual advance in the field. Each vocabulary is associated with a specific cell type family, meaning that the short DNA sequences recognized by transcription factors are not randomly distributed across the genome but instead cluster into coherent sets that define broad categories of cellular identity. When the researchers examined the combinatorial relationships among these motifs, they found that the particular combinations preferred by individual cell types are largely species-specific. This means that while the building blocks, the individual motifs themselves, remain stable across vast evolutionary distances, the ways in which those blocks are assembled into functional regulatory programmes evolve rapidly and independently in each lineage.</p>
<p>To dissect this layered organization, the team employed ChromBPNet, a deep-learning architecture designed to model chromatin accessibility at base resolution while factoring out technical biases introduced by the Tn5 transposase used in ATAC-seq library preparation. Models trained on chromatin accessibility data from one species accurately predicted family-level chromatin accessibility in distantly related species, demonstrating that the vocabulary-level information encoded in DNA sequences is conserved in a functionally meaningful way. However, interpretability analyses of the model predictions revealed a striking pattern: the deep-learning models frequently relied on different motifs from the shared vocabularies to arrive at convergent predictions. In other words, two species might achieve the same regulatory outcome for a given cell type family, but they do so by drawing on different members of the same motif vocabulary rather than by using identical regulatory elements.</p>
<p>The picture changes dramatically when the resolution of analysis shifts from the cell type family level to the individual cell type level. Models trained on chromatin accessibility data from a specific cell type within one species lost their predictive power when applied to the corresponding cell type in a distantly related species. This loss of cross-species transferability indicates that the regulatory syntax governing cell type-level identity, the precise arrangements and combinations of motifs that specify an individual cell type, evolves much more rapidly than the vocabulary-level constraints that define broader cell type families. The researchers refer to this hierarchical organization as a collective maintenance model, in which the identity of a cell type family is preserved not by any single conserved regulatory element but by the collective stability of a large pool of conserved regulatory factors.</p>
<p>This collective maintenance framework draws a compelling parallel to the concept of developmental homology in evolutionary biology. In homology, a character identity persists across species through conservation at the network level, even as the individual components of the network undergo extensive rewiring. The flatworm and vertebrate data suggest that cell type family identity operates under a similar logic: the vocabulary of sequence motifs defining a family is evolutionarily stable, yet the recombination of these motifs generates cell type-specific regulatory programmes that can differ substantially between species. This decoupling of family-level conservation from cell type-level innovation provides a mechanistic explanation for how new cell types can arise during evolution without disrupting the fundamental identities of existing cell type families.</p>
<p>The technical rigor underlying these conclusions is substantial. The researchers generated single-nucleus multi-omic sequencing data, capturing both gene expression and chromatin accessibility from the same individual nuclei, across three flatworm species: Schmidtea mediterranea, Schistosoma mansoni, and Macrostomum lignano. These species span a considerable range of evolutionary divergence within the flatworm phylum, providing a robust framework for comparative analysis. The team also extended their analysis to vertebrate systems by leveraging existing single-cell multi-omic data from mouse and zebrafish, two species separated by approximately 450 million years of evolution. The cross-species comparison between flatworms and vertebrates is particularly informative because these lineages diverged over 550 million years ago, representing one of the deepest evolutionary comparisons feasible with single-cell genomics.</p>
<p>Among the findings that emerge from this analysis is the observation that chromatin accessibility is conserved within cell type families even when the individual regulatory elements driving that accessibility are not. This means that the overall pattern of which parts of the genome are open and accessible in a given cell type family remains similar across species, but the specific DNA sequences responsible for opening those regions differ. The deep-learning models captured this distinction with remarkable fidelity, as they were able to predict family-level accessibility patterns across species using different combinations of motifs from the same conserved vocabulary. When the researchers examined neural cell types specifically, they documented extensive turnover in combinatorial motif usage, with individual neural cell types in different species employing different sets of motif pairs from the shared neural vocabulary to achieve their specific regulatory identities.</p>
<p>The implications of this work extend beyond basic evolutionary biology into the realm of biomedical research. Understanding that the vocabulary-level organization of regulatory motifs is conserved across species suggests that insights gained from model organisms about cell type family regulation may be more broadly transferable than previously appreciated, provided the analysis is conducted at the appropriate level of biological organization. Conversely, the rapid evolution of cell type-specific regulatory syntax means that extrapolating detailed regulatory mechanisms from one species to another requires caution, particularly for cell types that have undergone significant diversification. The collective maintenance model also raises intriguing questions about the evolutionary dynamics that maintain vocabulary stability while permitting combinatorial flexibility, and whether disruptions to vocabulary-level conservation might underlie certain developmental disorders or diseases.</p>
<p>As the field of single-cell genomics continues to expand the catalog of cell types across the tree of life, the framework developed in this study provides a conceptual scaffold for interpreting cross-species comparisons at multiple levels of resolution. The finding that conserved motif vocabularies constrain genome access while their flexible recombination drives cell type innovation bridges a persistent gap between the stability of cellular identities and the remarkable diversity of cell types observed across the animal kingdom. The research team has made all data and computational tools publicly available, including the single-cell multi-ome datasets deposited in the Sequence Read Archive, the ChromBPNet models shared through figshare, and the complete analysis code archived on GitHub and Zenodo, ensuring that the broader scientific community can build upon these findings to further unravel the regulatory architecture of cell type evolution.</p>
<p><strong>Subject of Research:</strong> Evolutionary conservation of regulatory motif vocabularies governing chromatin accessibility and cell type family identity across flatworms and vertebrates</p>
<p><strong>Article Title:</strong> Flexible use of conserved motifs constrains genome access in cell type evolution</p>
<p><strong>Article References:</strong> Flexible use of conserved motifs constrains genome access in cell type evolution. (n.d.). <a href="https://doi.org/10.1038/s41559-026-03164-5" rel="noopener noreferrer">https://doi.org/10.1038/s41559-026-03164-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41559-026-03164-5" rel="noopener noreferrer">10.1038/s41559-026-03164-5</a></p>
<p><strong>Keywords:</strong> cell type evolution, chromatin accessibility, deep learning, single-cell multi-omics, motif vocabularies, regulatory syntax, flatworms, vertebrates, developmental homology, gene regulation, Nature Ecology &amp; Evolution, transcription factors</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194519</post-id>	</item>
		<item>
		<title>High-Resolution Mapping of Cell-Specific Gene Regulation from Bulk Sequencing</title>
		<link>https://scienmag.com/high-resolution-mapping-of-cell-specific-gene-regulation-from-bulk-sequencing/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Mon, 13 Jul 2026 12:38:23 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[bulk sequencing data analysis]]></category>
		<category><![CDATA[cell-type-specific gene regulation]]></category>
		<category><![CDATA[cell-type-specific regulatory mechanisms]]></category>
		<category><![CDATA[chromatin immunoprecipitation sequencing (ChIP-seq)]]></category>
		<category><![CDATA[cross-modality deconvolution]]></category>
		<category><![CDATA[deep learning in genomics]]></category>
		<category><![CDATA[genomic signal dissection]]></category>
		<category><![CDATA[high-resolution genomic mapping]]></category>
		<category><![CDATA[multi-omics integration]]></category>
		<category><![CDATA[nascent transcription profiling]]></category>
		<category><![CDATA[single-cell chromatin accessibility]]></category>
		<category><![CDATA[tissue heterogeneity analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/high-resolution-mapping-of-cell-specific-gene-regulation-from-bulk-sequencing/</guid>

					<description><![CDATA[In a groundbreaking advancement for genomic research, scientists have unveiled DeepDETAILS, a novel deep-learning framework that dramatically enhances our ability to dissect cell-type-specific regulatory mechanisms from complex tissue samples. Traditional single-cell sequencing methods, such as scRNA-seq and scATAC-seq, have revolutionized cellular biology by profiling regulatory landscapes at the individual cell level. However, adapting these techniques [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking advancement for genomic research, scientists have unveiled DeepDETAILS, a novel deep-learning framework that dramatically enhances our ability to dissect cell-type-specific regulatory mechanisms from complex tissue samples. Traditional single-cell sequencing methods, such as scRNA-seq and scATAC-seq, have revolutionized cellular biology by profiling regulatory landscapes at the individual cell level. However, adapting these techniques for other genome-wide assays—especially those that measure diverse chromatin features and transcriptional activity—has remained a formidable challenge.</p>
<p>DeepDETAILS addresses this gap by performing cross-modality deconvolution, integrating high-resolution single-cell open chromatin references with bulk sequencing data from complementary assays. This quasisupervised algorithm enables researchers to dissect locus-specific genomic signals at base-pair resolution, resolving the contributions of different cell types within heterogeneous tissue samples. Impressively, the method is compatible with various genomic layers including nascent transcription measurements from PRO-cap and PRO-seq, as well as chromatin immunoprecipitation sequencing (ChIP-seq) for histone modifications.</p>
<p>By leveraging single-cell chromatin accessibility as a reference, DeepDETAILS constructs precise, cell-type-resolved maps of transcriptional regulatory processes, overcoming technical barriers that have limited the study of complex tissues. In applied analyses spanning 39 human tissues and 86 cell types, the team compiled a comprehensive atlas of cell-type-specific nascent RNA synthesis and epigenetic modifications, providing an unparalleled resource for the community.</p>
<p>Beyond resource generation, the utility of DeepDETAILS was showcased in fine-mapping genetic risk variants associated with primary sclerosing cholangitis (PSC), a devastating liver disorder marked by progressive bile duct inflammation. The framework pinpointed specific cell types and regulatory elements implicated in disease etiology, opening new avenues for understanding pathogenic mechanisms and identifying potential therapeutic targets.</p>
<p>This deep-learning driven approach fundamentally transforms how bulk sequencing data can be harnessed to infer cell-type-specific regulatory activity with unprecedented resolution. It circumvents the cost and technical complexity of generating single-cell data for every assay type by computationally integrating modalities, significantly broadening the scope of genomic interrogation.</p>
<p>As the biomedical field continues to grapple with the complexity of cellular heterogeneity in tissues, DeepDETAILS sets a new standard for multi-omic deconvolution. Its adaptable framework promises to accelerate discovery in developmental biology, disease pathogenesis, and therapeutic intervention by providing accurate, interpretable, and scalable solutions for dissecting regulatory dynamics within diverse biological contexts.</p>
<p>With this tool, researchers are now equipped to unlock hidden layers of transcriptional regulation from existing bulk datasets, greatly expanding the potential for insights into the molecular architecture of human health and disease.</p>
<p>Subject of Research: Single-cell resolution deconvolution of bulk genomic sequencing data for transcriptional regulatory analysis</p>
<p>Article Title: High-resolution reconstruction of cell-type-specific transcriptional regulatory processes from bulk sequencing samples</p>
<p>Article References:<br />
Yao, L., Shah, S.R., Ozer, A. et al. <em>High-resolution reconstruction of cell-type-specific transcriptional regulatory processes from bulk sequencing samples.</em> Nat Biotechnol (2026). <a href="https://doi.org/10.1038/s41587-026-03218-w">https://doi.org/10.1038/s41587-026-03218-w</a></p>
<p>DOI: <a href="https://doi.org/10.1038/s41587-026-03218-w">https://doi.org/10.1038/s41587-026-03218-w</a></p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">172038</post-id>	</item>
		<item>
		<title>Deep Learning Uncovers Multiomic Data Integration Insights</title>
		<link>https://scienmag.com/deep-learning-uncovers-multiomic-data-integration-insights/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 26 Jan 2026 22:40:28 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[advancements in cellular dynamics understanding]]></category>
		<category><![CDATA[cellular heterogeneity analysis]]></category>
		<category><![CDATA[challenges in multiomics analysis]]></category>
		<category><![CDATA[contrastive learning in bioinformatics]]></category>
		<category><![CDATA[cross-modal integration of omics layers]]></category>
		<category><![CDATA[deep learning in genomics]]></category>
		<category><![CDATA[extracting insights from biological data]]></category>
		<category><![CDATA[high-dimensional biological datasets]]></category>
		<category><![CDATA[Innovative approaches in genomics]]></category>
		<category><![CDATA[machine learning for biological data]]></category>
		<category><![CDATA[regulatory mechanisms in cell function]]></category>
		<category><![CDATA[single-cell multiomic data integration]]></category>
		<guid isPermaLink="false">https://scienmag.com/deep-learning-uncovers-multiomic-data-integration-insights/</guid>

					<description><![CDATA[In the ever-evolving landscape of genomics and bioinformatics, the need for innovative approaches to analyze complex biological data is paramount. Recent research led by Cheng et al. introduces a groundbreaking method for analyzing single-cell multiomic data through deep contrastive learning, paving the way for advancements in our understanding of cellular heterogeneity and functional integration across [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In the ever-evolving landscape of genomics and bioinformatics, the need for innovative approaches to analyze complex biological data is paramount. Recent research led by Cheng et al. introduces a groundbreaking method for analyzing single-cell multiomic data through deep contrastive learning, paving the way for advancements in our understanding of cellular heterogeneity and functional integration across different biological modalities. This pioneering study primarily focuses on the aligned cross-modal integration of various omics layers, which can unveil the intricate regulatory mechanisms governing cell function and identity.</p>
<p>The study emphasizes the versatility and effectiveness of deep learning techniques in extracting meaningful insights from high-dimensional biological datasets. Single-cell multiomics, which combines genomic, transcriptomic, and epigenomic data at the single-cell level, presents a formidable challenge due to its inherent complexity. Traditional analytical methods often struggle to capture the multifaceted relationships among different omics layers. However, this new approach adeptly bridges the gap between disparate data modalities, leading to a deeper understanding of cellular dynamics.</p>
<p>One of the core innovations detailed in the study is the application of contrastive learning principles to the realm of genomics. In typical machine learning tasks, contrastive learning assists in distinguishing between similar and dissimilar instances by training models to maximize agreement between positive pairs while minimizing it for negative pairs. Cheng and colleagues adapted these principles to the analysis of multiomic datasets, effectively generating robust representations that incorporate both common and unique features of different omic layers.</p>
<p>The research presents a detailed methodology that integrates deep contrastive learning with single-cell multiomics, providing a systematic framework for analyzing heterogeneous cellular populations. Esto enables researchers to tackle key biological questions regarding cell-type identification, cellular states, and regulatory networks with unprecedented accuracy and sensitivity. The authors highlight that this method not only enhances performance in clustering and classification tasks but also provides significant insights into the functional implications of cellular diversity.</p>
<p>Moreover, the study highlights the importance of considering the interactions among various molecular layers. By aligning omics data through deep contrastive representations, the research underscores the significance of cross-modal relationships that contribute to cellular identity and function. This holistic view of molecular data allows for a more nuanced interpretation and understanding of cellular behavior in health and disease.</p>
<p>Furthermore, the implications of this research extend beyond basic biology into potential clinical applications. Understanding cell-specific regulatory mechanisms can inform therapeutic strategies for diseases characterized by cellular dysregulation, including cancer and autoimmune disorders. By providing a clearer picture of the cellular landscape and its influences, this study opens avenues for targeted interventions and precision medicine.</p>
<p>Additionally, the authors discuss the computational efficiency of their approach. While traditional methods may require extensive preprocessing and manual integration of datasets, the deep learning-based framework significantly reduces the overhead associated with these steps. This not only expedites the analysis process but also minimizes the introduction of biases that can arise during data integration.</p>
<p>The research findings are showcased through various case studies, demonstrating the method’s capability to uncover biologically relevant signals and regulatory pathways. These examples illustrate how aligned cross-modal integration can lead to discoveries of novel cell types and states that were previously obscured in the noise of high-dimensional data.</p>
<p>As the field continues to progress towards personalized medicine, methodologies like the one proposed by Cheng et al. are crucial. The ability to conduct integrated analyses of single-cell multiomics will empower researchers to decipher the underlying genetic and epigenetic mechanisms of complex diseases, ultimately guiding the development of more effective treatment strategies.</p>
<p>In conclusion, the work presented by Cheng, Su, Fan, and their team marks a significant advancement in the field of multiomics analysis. By leveraging the power of deep contrastive learning, this research provides a novel lens through which the multifaceted nature of single-cell data can be explored and understood. As the scientific community continues to harness the potential of AI and machine learning in biology, studies like this will undoubtedly shape the future of genomic research and its applications in healthcare.</p>
<p>The study’s results not only demonstrate the feasibility of applying advanced machine learning techniques to biological data but also emphasize the importance of integrative approaches that can capture the complexity of living systems. As more researchers adopt these cutting-edge methodologies, we can anticipate a sharper understanding of the biological underpinnings of health and disease.</p>
<p>By continuously pushing the boundaries of what is possible in genomics, researchers are setting the stage for transformative breakthroughs that could redefine how we approach the complexities of life itself.</p>
<p>This research serves as a reminder that the journey into the cellular world, now facilitated by deep learning and sophisticated analytic techniques, is only just beginning. The tools and insights generated through these studies will forge new paths in our quest to unravel the molecular intricacies of life, ultimately enhancing our understanding of ourselves and the biological universe around us.</p>
<p><strong>Subject of Research</strong>: Integrated analysis of single-cell multiomic data using deep contrastive learning.</p>
<p><strong>Article Title</strong>: Aligned cross-modal integration and regulatory heterogeneity characterization of single-cell multiomic data with deep contrastive learning.</p>
<p><strong>Article References</strong>: Cheng, Y., Su, Y., Fan, Y. <i>et al.</i> Aligned cross-modal integration and regulatory heterogeneity characterization of single-cell multiomic data with deep contrastive learning. <i>Genome Med</i> <b>18</b>, 10 (2026). https://doi.org/10.1186/s13073-025-01586-7</p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: https://doi.org/10.1186/s13073-025-01586-7</p>
<p><strong>Keywords</strong>: single-cell multiomics, deep contrastive learning, cross-modal integration, genomic data, machine learning, cellular heterogeneity, regulatory networks, precision medicine.</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">131334</post-id>	</item>
		<item>
		<title>Ancient Recombination Desert Drives Mammal Speciation</title>
		<link>https://scienmag.com/ancient-recombination-desert-drives-mammal-speciation/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 13 Nov 2025 02:07:43 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[ancient recombination desert]]></category>
		<category><![CDATA[deep learning in genomics]]></category>
		<category><![CDATA[evolutionary biology advancements]]></category>
		<category><![CDATA[gene flow in evolution]]></category>
		<category><![CDATA[genetic introgression effects]]></category>
		<category><![CDATA[mammal speciation mechanisms]]></category>
		<category><![CDATA[phylogenetic tree reconstruction]]></category>
		<category><![CDATA[placental mammal phylogenetics]]></category>
		<category><![CDATA[recombination rate dynamics]]></category>
		<category><![CDATA[species formation processes]]></category>
		<category><![CDATA[supergene evolution in mammals]]></category>
		<category><![CDATA[X chromosome genetics]]></category>
		<guid isPermaLink="false">https://scienmag.com/ancient-recombination-desert-drives-mammal-speciation/</guid>

					<description><![CDATA[In a groundbreaking study published in Nature, researchers have unveiled the existence of an ancient recombination desert on the X chromosome that acts as a formidable speciation supergene across placental mammals. This discovery not only sheds new light on the genetic underpinnings of species formation but also offers a powerful new tool for resolving challenging [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking study published in <em>Nature</em>, researchers have unveiled the existence of an ancient recombination desert on the X chromosome that acts as a formidable speciation supergene across placental mammals. This discovery not only sheds new light on the genetic underpinnings of species formation but also offers a powerful new tool for resolving challenging evolutionary relationships that have eluded scientists for decades.</p>
<p>Gene flow—the interbreeding and genetic exchange between different species—is a widespread phenomenon across the tree of life. It plays a crucial role in adaptation and evolution but also complicates our ability to decipher true species relationships. While such genetic introgression can generate new genetic combinations and fuel diversity, it simultaneously blurs historical signals essential for reconstructing phylogenetic trees. The interplay between gene flow and recombination—a biological process that shuffles genetic material during meiosis—has remained a complex and not fully understood aspect of evolutionary biology.</p>
<p>The study tackled this longstanding challenge by leveraging cutting-edge deep learning algorithms applied to comprehensive genome alignments spanning 22 distinct placental mammal species. Researchers trained their models to infer the evolutionary dynamics of the recombination landscape—specifically, how recombination rates have changed and been maintained across tens of millions of years. Their analyses pinpointed a remarkable feature: a large recombination desert occupying roughly 30% of the X chromosome. This region exhibits drastically reduced recombination rates compared to the rest of the genome.</p>
<p>What makes this recombination desert remarkable is its evolutionary conservation. Despite the immense diversity and divergence among placental mammals, the recombination desert on the X chromosome has remained intact for millions of years. This suggests it performs a vital biological function, beyond being a mere genomic quirk. Indeed, further phylogenomic analyses incorporating data from 94 species revealed that the X-linked recombination desert serves as a longstanding barrier to gene flow. In scenarios where introgression dominates genome-wide ancestry, the recombination desert faithfully retains the true species history.</p>
<p>This genetic stronghold operates as a speciation supergene—a cluster of tightly linked genes that collectively underpin reproductive isolation. The researchers discovered that this locus is enriched with genes involved in sex chromosome silencing and key reproductive traits. Such genetic architecture supports the idea that suppressed recombination in this region protects co-adapted gene complexes critical for species integrity. By guarding against the homogenizing effects of hybridization, the recombination desert forms a genomic firewall preserving species boundaries.</p>
<p>The concept of a speciation supergene on a sex chromosome isn’t entirely new, but the scale and evolutionary longevity documented here are unprecedented. Unlike smaller supergenes often identified in insects or plants, this X-linked recombination desert spans nearly a third of the chromosome and remains conserved across multiple mammalian orders. This points to a generalized role in maintaining reproductive isolation in a broad array of placental mammals—a phenomenon not previously appreciated at this magnitude.</p>
<p>From a methodological perspective, the use of deep learning to infer recombination landscapes from genome alignments represents a significant advance. Traditional methods rely heavily on experimentally derived recombination maps, which are rare and challenging to obtain across many species. The AI-driven approach circumvents these limitations by detecting subtle genomic signatures indicative of recombination suppression. This opens the door for large-scale comparative analyses that were previously unfeasible.</p>
<p>Perhaps one of the most exciting implications of this work lies in its application to phylogenetics—the science of reconstructing species evolutionary histories. The study shows that incorporating recombination-aware models dramatically improves the resolution of phylogenetic trees, particularly when gene flow confounds conventional approaches. By focusing on the genomic region resistant to introgression, researchers obtain a clearer signal of species relationships, overcoming one of the most persistent obstacles in evolutionary biology.</p>
<p>The supergene’s enrichment for genes mediating sex chromosome inactivation dovetails with our understanding of hybrid incompatibilities. The process of X chromosome silencing during meiosis—critical for normal gamete development—is highly sensitive to disturbances, often underlying hybrid sterility in mammals. The recombination desert’s maintenance may thus reflect selective pressures to preserve crucial meiotic mechanisms and fertility barriers that reinforce speciation.</p>
<p>Furthermore, the identification of this ancient recombination desert provides novel insights into the evolutionary forces shaping sex chromosomes. Sex chromosomes are known for their unique dynamics, including suppressed recombination and accumulation of reproductive genes. This study elegantly illustrates how these genomic peculiarities integrate with macroevolutionary patterns, linking chromosome biology to the broader speciation landscape in mammals.</p>
<p>Overall, the findings carry profound implications for understanding how complex genomes navigate the tension between gene flow and species divergence. The recombination desert emerges as a pivotal evolutionary feature that secures species boundaries and preserves the authenticity of evolutionary histories amidst pervasive hybridization. As such, it stands as a cornerstone for future investigations into mammalian speciation, genome evolution, and the genetic architecture of reproductive isolation.</p>
<p>In an era where genomic data are accumulating at unprecedented rates, this study exemplifies the power of integrating advanced computational methods with evolutionary theory to uncover hidden genomic phenomena. It also underscores the necessity of considering recombination landscapes when interpreting genome-wide data, especially in systems characterized by extensive gene flow.</p>
<p>Beyond its scientific impact, the discovery holds potential applied relevance. The recombination desert locus could serve as a molecular marker in conservation genetics, systematics, and breeding programs, helping identify cryptic species boundaries and maintain biodiversity. Moreover, understanding the genetic basis of reproductive isolation may inform medical research on fertility and chromosome biology.</p>
<p>As this pioneering research gains traction, it invites the scientific community to revisit long-held assumptions about genomic recombination and speciation. The ancient recombination desert on the X chromosome is not merely a passive genomic feature but an active architect of mammalian biodiversity—a supergene standing guard over the essence of species identity.</p>
<hr />
<p><strong>Subject of Research</strong>: Evolutionary genetics of placental mammals, focusing on recombination landscapes and speciation mechanisms.</p>
<p><strong>Article Title</strong>: An ancient recombination desert is a speciation supergene in placental mammals.</p>
<p><strong>Article References</strong>:<br />
Foley, N.M., Rasulis, R.G., Wani, Z. <em>et al.</em> An ancient recombination desert is a speciation supergene in placental mammals. <em>Nature</em> (2025). <a href="https://doi.org/10.1038/s41586-025-09740-2">https://doi.org/10.1038/s41586-025-09740-2</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: <a href="https://doi.org/10.1038/s41586-025-09740-2">https://doi.org/10.1038/s41586-025-09740-2</a></p>
<p><strong>Keywords</strong>: recombination desert, speciation supergene, placental mammals, gene flow, introgression, phylogenomics, sex chromosome silencing, reproductive isolation, X chromosome, deep learning, evolutionary genetics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">104967</post-id>	</item>
		<item>
		<title>Unlocking Noncoding Variants&#8217; Influence on Gene Expression</title>
		<link>https://scienmag.com/unlocking-noncoding-variants-influence-on-gene-expression/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 02 Oct 2025 00:24:25 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Assay for Transposase-Accessible Chromatin]]></category>
		<category><![CDATA[challenges in gene regulatory prediction]]></category>
		<category><![CDATA[chromatin accessibility and gene regulation]]></category>
		<category><![CDATA[computational approaches in genetics]]></category>
		<category><![CDATA[deep learning in genomics]]></category>
		<category><![CDATA[EMO model for epigenomic modeling]]></category>
		<category><![CDATA[genomic science advancements]]></category>
		<category><![CDATA[integrating sequencing and chromatin data]]></category>
		<category><![CDATA[noncoding variants and gene expression]]></category>
		<category><![CDATA[predicting noncoding mutation effects]]></category>
		<category><![CDATA[regulatory impact of noncoding SNPs]]></category>
		<category><![CDATA[tissue-specific gene regulation]]></category>
		<guid isPermaLink="false">https://scienmag.com/unlocking-noncoding-variants-influence-on-gene-expression/</guid>

					<description><![CDATA[In the rapidly evolving field of genomic science, the ability to predict how noncoding mutations influence gene expression has increasingly become a frontier of investigation. Scientists have long recognized the importance of noncoding regions of DNA, which make up a substantial portion of the human genome and play critical roles in regulatory mechanisms. However, accurately [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In the rapidly evolving field of genomic science, the ability to predict how noncoding mutations influence gene expression has increasingly become a frontier of investigation. Scientists have long recognized the importance of noncoding regions of DNA, which make up a substantial portion of the human genome and play critical roles in regulatory mechanisms. However, accurately assessing the regulatory impact of noncoding single nucleotide polymorphisms (SNPs) remains a formidable challenge, particularly due to their tissue-specific and cell-type-specific effects. Recent advancements have paved the way for novel computational approaches that harness the power of deep learning to better decipher these complex relationships.</p>
<p>Introducing the EMO model, researchers have taken a significant leap forward in the computation and prediction of the regulatory influences exerted by noncoding variants. EMO, which stands for Epigenomic Modelling for Omics, employs a transformer-based architecture designed to integrate DNA sequencing with chromatin accessibility data. Specifically, it utilizes Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) data to highlight regions of the genome that are epigenetically active and potentially influential in gene regulation. This symbiosis between sequence data and chromatin state data forms a robust foundation for exploring the functional consequences of genetic variation.</p>
<p>One of EMO&#8217;s standout features is its capacity to integrate personalized functional genomic profiles. This unique capability allows the model to not only generate generalizable predictions across various tissues and cell types but also to tailor its predictions to individual genomic contexts. This personalization addresses a critical limitation often seen in conventional models that lack the granularity needed for precise predictions tied to specific genetic backgrounds or disease states.</p>
<p>Incorporating both short- and long-range regulatory interactions enables EMO to capture the dynamic regulatory landscape that influences gene expression. This dynamic approach is particularly crucial when considering the progression of diseases, as gene expression patterns can shift substantially in response to pathological changes. By modeling these interactions with a deep learning framework, EMO stands apart from other predictive models in its ability to adapt to and analyze changes in gene expression tied to specific conditions.</p>
<p>Moreover, benchmark evaluations have demonstrated EMO&#8217;s superiority over existing predictive frameworks in the domain of noncoding variant impacts. Through a process of pretraining, the model has developed strong baseline capabilities that are further enhanced when fine-tuning is performed on smaller, specific samples. This method of transfer learning allows EMO to refine its predictive performance in target tissue types, showcasing the flexibility and power of this computational tool.</p>
<p>In single-cell contexts, which have emerged as vital for understanding cellular heterogeneity and specialized gene expression, EMO showcases remarkable performance. The model adeptly identifies regulatory patterns specific to various cell types, detecting nuanced differences that could be pivotal in elucidating disease mechanisms. For instance, the ability to pinpoint how adhesion molecules or transcription factors are regulated differently in immune cells as compared to neuronal cells can lead to profound insights into diseases that manifest in specific tissues.</p>
<p>Various studies have highlighted the association of SNPs with disease susceptibility, yet the pathways through which these genetic variants exert their influence on gene expression remained largely obscure. EMO addresses this knowledge gap by linking genetic variation not only to gene expression changes but also to disease-relevant pathways. This pathway-centric approach opens new avenues for therapeutic interventions, as understanding which genetic variants are functionally impactful allows for more targeted strategies in managing diseases.</p>
<p>While the advances presented by EMO are promising, there is also an intrinsic complexity within the integration of genomic data and epigenomic features. Deciphering the effects of noncoding mutations involves navigating intricate regulatory networks, and thus the challenge resides in the multifaceted nature of these interactions. The transformer architecture employed by EMO is adept at managing such complexities, enabling it to discern patterns within vast datasets.</p>
<p>The implications of this research extend beyond mere academic interest; they pose transformative potential for personalized medicine. As we inch closer toward understanding individual genetic architectures, the ability to predict how specific noncoding variants will affect gene expression could translate into actionable insights for tailored treatments. This precision in medicine relies heavily on the functional understanding gained through advanced computational models like EMO.</p>
<p>The future of genomic research demands interdisciplinary approaches, where biology and computational science converge. The development of models like EMO highlights the necessity for innovative tools that can not only improve predictive accuracy but also facilitate collaborative efforts across research fields. As the relationship between genetic variation and phenotypic expression becomes clearer, it promises to propel advancements across varied scientific domains, including development, evolution, and disease mitigation.</p>
<p>To summarize, EMO represents a crucial step forward in our understanding of noncoding variants and their regulatory roles. By effectively integrating multiple layers of genomic data, it enhances the predictive capabilities essential for dissecting the complexities of gene regulation. As experts continue to unravel the intricate threads of the human genome, tools like EMO will be indispensable in paving the way toward breakthroughs in genetic research, disease understanding, and ultimately, personalized medicine.</p>
<p>The importance of the studies surrounding gene expression regulation cannot yet be overstated. Each discovery not only solidifies foundational knowledge but also catalyzes the emergence of novel research directions. Given the breadth of applications stemming from this work, EMO and similar models are set to become central players in the genomic landscape, resulting in enriched insights that forge new pathways in human health and disease.</p>
<p>As the realm of functional genomics continues to evolve, the collaborative intersections between computational tools and biological inquiry will only deepen. With models like EMO leading the charge, there is a growing anticipation for what the next frontier in genomic research will entail, along with its implications for health, disease, and the future of medical science.</p>
<p>The launch of EMO marks a pivotal moment that could redefine how scientists approach the complexities of gene regulation. By addressing the challenges presented by noncoding mutations, EMO not only elevates predictive accuracy but also enriches our understanding of the underlying biological phenomena. This endeavor embodies a crucial step toward merging computational prowess with biological specificity, setting the stage for a new era in understanding the human genome.</p>
<p>The excitement surrounding this research is palpable within the scientific community as individuals grapple with the potential it holds. The implications of uncovering the functional roles of noncoding variants extend far beyond theoretical exploration—they could redefine therapeutic approaches and improve individualized care strategies significantly. As we eagerly await further developments stemming from EMO&#8217;s capabilities, the anticipation for groundbreaking discoveries accompanying its implementation continues to grow.</p>
<p>In summation, EMO exemplifies the convergence of genomic science and computational innovation, heralding a new age for functional genomics. As researchers navigate the intricacies of noncoding mutations and their regulatory impacts, the tools developed through such work promise to enhance our understanding of gene expression, paving the way for tailored therapies and improved health outcomes.</p>
<hr />
<p><strong>Subject of Research</strong>: Predicting the regulatory impacts of noncoding variants on gene expression through epigenomic integration.</p>
<p><strong>Article Title</strong>: Predicting the regulatory impacts of noncoding variants on gene expression through epigenomic integration across tissues and single-cell landscapes.</p>
<p><strong>Article References</strong>:<br />
Liu, Z., Bao, Y., Gu, A. <em>et al.</em> Predicting the regulatory impacts of noncoding variants on gene expression through epigenomic integration across tissues and single-cell landscapes.<br />
<em>Nat Comput Sci</em> (2025). <a href="https://doi.org/10.1038/s43588-025-00878-7">https://doi.org/10.1038/s43588-025-00878-7</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: 10.1038/s43588-025-00878-7</p>
<p><strong>Keywords</strong>: Noncoding mutations, gene expression, EMO model, chromatin accessibility, SNPs, personalized medicine, regulatory patterns, disease progression, computational genomics, transformer-based models.</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">84995</post-id>	</item>
		<item>
		<title>Deep Learning Enhances Polygenic Score Accuracy</title>
		<link>https://scienmag.com/deep-learning-enhances-polygenic-score-accuracy/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Mon, 02 Jun 2025 19:12:47 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[advancements in public health genomics]]></category>
		<category><![CDATA[AI in genetic analysis]]></category>
		<category><![CDATA[complex trait prediction]]></category>
		<category><![CDATA[computational techniques in medicine]]></category>
		<category><![CDATA[deep learning frameworks for genetics]]></category>
		<category><![CDATA[deep learning in genomics]]></category>
		<category><![CDATA[genetic variant interactions]]></category>
		<category><![CDATA[high-dimensional genomic data analysis]]></category>
		<category><![CDATA[nonlinear effects in genetics]]></category>
		<category><![CDATA[personalized medicine with AI]]></category>
		<category><![CDATA[polygenic risk prediction accuracy]]></category>
		<category><![CDATA[polygenic score enhancement]]></category>
		<guid isPermaLink="false">https://scienmag.com/deep-learning-enhances-polygenic-score-accuracy/</guid>

					<description><![CDATA[In recent years, the field of genomics has witnessed remarkable advancements driven by the integration of artificial intelligence, particularly deep learning, with traditional genetic analyses. Among the most promising applications is the enhancement of polygenic scores, which estimate an individual&#8217;s susceptibility to various diseases by aggregating the effects of numerous genetic variants. A groundbreaking study [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In recent years, the field of genomics has witnessed remarkable advancements driven by the integration of artificial intelligence, particularly deep learning, with traditional genetic analyses. Among the most promising applications is the enhancement of polygenic scores, which estimate an individual&#8217;s susceptibility to various diseases by aggregating the effects of numerous genetic variants. A groundbreaking study published in <em>Nature Communications</em> by Kelemen, Xu, Jiang, and colleagues in 2025 has pushed the boundaries of this approach by rigorously evaluating the performance of deep-learning-based methods for improving polygenic risk prediction. Their work offers novel insights into how cutting-edge computational techniques can transform personalized medicine and public health genomics.</p>
<p>Polygenic scores have revolutionized our ability to quantify inherited risk for complex traits and diseases by synthesizing information across millions of common genetic variants. Traditional methods to compute these scores often rely on linear models, which may fail to capture intricate genetic architectures involving interactions among variants or nonlinear effects. The research team sought to address these inherent limitations by employing various deep-learning frameworks capable of modeling complex patterns within high-dimensional genomic data. Their objective was to determine whether these sophisticated models could outperform existing approaches and deliver more accurate and clinically actionable polygenic scores.</p>
<p>The core of their investigation involved designing and training multiple deep neural network architectures on extensive genome-wide association study (GWAS) datasets. These networks included convolutional layers to identify local sequence patterns, fully connected layers to integrate signals across the genome, and attention mechanisms to prioritize relevant genomic regions. The models were rigorously validated using independent cohorts, ensuring robustness against overfitting and generalizability across populations. By comparing performance metrics such as predictive accuracy, area under the receiver operating characteristic curve (AUC), and calibration scores, the authors provided a comprehensive benchmark of state-of-the-art methods.</p>
<p>One of the most striking findings was that deep-learning-based models demonstrated consistent improvements in predictive accuracy relative to canonical polygenic scoring techniques, especially for traits with complex genetic underpinnings. Diseases such as type 2 diabetes, coronary artery disease, and various psychiatric disorders exhibited enhanced risk stratification when analyzed through these neural networks. The study underscored that the capacity of deep learning to capture nonlinear relationships and higher-order interactions among variants was key to this superior performance. Moreover, the interpretability modules integrated within the models enabled the identification of biologically meaningful variant clusters, providing mechanistic insights that were previously elusive.</p>
<p>The researchers also tackled the challenge of computational efficiency and scalability, which are critical for clinical adoption. Training deep neural networks on genomic-scale data is notoriously resource-intensive, but through innovative algorithmic optimizations and parallel computing techniques, the team was able to reduce training times significantly. This optimization enables the potential deployment of deep-learning-enhanced polygenic scoring in routine medical settings, where timely risk assessments could inform prevention strategies and tailored therapeutic interventions.</p>
<p>An important aspect of the study was the exploration of transfer learning approaches, wherein neural networks pre-trained on one trait or population were fine-tuned for another. This methodology demonstrated promising results, allowing models to leverage shared genetic architectures across phenotypes and ancestries. Transfer learning thus offers a pathway to mitigate disparities in genomic research, where underrepresented populations suffer from a lack of well-powered GWAS datasets. By enhancing prediction accuracy across diverse cohorts, deep learning can contribute to more equitable healthcare outcomes in genomic medicine.</p>
<p>Despite these advancements, the authors acknowledge several limitations and areas for future research. For instance, while deep learning models improve polygenic score accuracy, they still depend on the quality and diversity of the underlying GWAS data. Phenotypic heterogeneity, environmental confounders, and gene-environment interactions remain challenging to incorporate fully. The study calls for integrating multi-omics data, longitudinal health records, and environmental metrics to construct more holistic risk models. Additionally, the interpretability of deep learning remains an ongoing technical and ethical concern, necessitating transparent model designs and validation protocols to foster clinical trust.</p>
<p>The societal implications of improved polygenic scoring using deep learning are profound. By enabling earlier and more precise identification of individuals at heightened genetic risk, healthcare systems can implement targeted screening programs and preventive lifestyle modifications. Such proactive approaches could reduce the burden of chronic diseases and improve population health outcomes. Moreover, these models can aid drug discovery by pinpointing genetic pathways most strongly linked to disease risk, accelerating the development of novel therapeutics. However, ethical considerations surrounding genetic privacy, data security, and potential discrimination must keep pace with these technological innovations.</p>
<p>Kelemen and colleagues’ work exemplifies the power of interdisciplinary collaboration, blending genomics, machine learning, and biomedical science to address one of the most complex challenges in human health. Their rigorous benchmarking framework sets a new standard for evaluating polygenic score methodologies, encouraging transparency and reproducibility in this fast-evolving research domain. Their public release of trained models and code repositories further democratizes access to these tools, facilitating broader adoption and iterative improvements by the scientific community.</p>
<p>Looking ahead, the integration of deep learning into polygenic risk modeling opens new horizons for precision medicine. The ability to untangle the multifaceted genetic basis of disease at scale promises unprecedented insights into pathogenesis and individual variability. As biobank-linked cohorts expand globally and computational resources continue to grow, we anticipate a proliferation of ever more nuanced and powerful models that transcend current limitations. The synergy of genomic data and artificial intelligence heralds a transformative era, wherein preventive healthcare is predictive, personalized, and participatory.</p>
<p>In summary, the 2025 study by Kelemen, Xu, Jiang, and colleagues represents a landmark contribution to the genomics community and beyond. By rigorously demonstrating the tangible benefits of deep-learning techniques for polygenic score enhancement, this research paves the way for more accurate genetic risk prediction tools. These advancements not only deepen our biological understanding but fundamentally reshape how we approach disease prevention, diagnosis, and treatment in the 21st century. As we stand on the cusp of widespread clinical translation, this work exemplifies the profound impact that AI can have when thoughtfully harnessed in biomedicine.</p>
<p>The continued convergence of machine learning innovation and genomic science, as epitomized by this study, ensures that the future of predictive health is both data-driven and deeply human-centric. Ultimately, empowering individuals and clinicians with precise genetic insights derived from sophisticated deep-learning models will be a cornerstone of next-generation healthcare systems worldwide. The implications for extending healthy lifespans and alleviating disease burdens are immense. This pioneering research marks a critical milestone on that transformative journey.</p>
<hr />
<p><strong>Subject of Research</strong>: Deep-learning approaches to enhance polygenic risk scores for disease prediction.</p>
<p><strong>Article Title</strong>: Performance of deep-learning-based approaches to improve polygenic scores.</p>
<p><strong>Article References</strong>:<br />
Kelemen, M., Xu, Y., Jiang, T. <em>et al.</em> Performance of deep-learning-based approaches to improve polygenic scores. <em>Nat Commun</em> <strong>16</strong>, 5122 (2025). <a href="https://doi.org/10.1038/s41467-025-60056-1">https://doi.org/10.1038/s41467-025-60056-1</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">50628</post-id>	</item>
		<item>
		<title>AI Model Decodes the Universal &#8216;Language&#8217; of Regulatory Genomics, Unveiling Cellular Narratives</title>
		<link>https://scienmag.com/ai-model-decodes-the-universal-language-of-regulatory-genomics-unveiling-cellular-narratives/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Wed, 29 Jan 2025 20:02:34 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[AI in regulatory genomics]]></category>
		<category><![CDATA[artificial intelligence in healthcare]]></category>
		<category><![CDATA[chromatin accessibility maps]]></category>
		<category><![CDATA[collaboration in genomic research]]></category>
		<category><![CDATA[Dana-Farber Cancer Institute research]]></category>
		<category><![CDATA[decoding cellular landscapes]]></category>
		<category><![CDATA[deep learning in genomics]]></category>
		<category><![CDATA[EpiBERT model for gene expression]]></category>
		<category><![CDATA[gene regulation across cell types]]></category>
		<category><![CDATA[innovative genomic models]]></category>
		<category><![CDATA[predictive analytics in biology]]></category>
		<category><![CDATA[understanding gene expression mechanisms]]></category>
		<guid isPermaLink="false">https://scienmag.com/ai-model-decodes-the-universal-language-of-regulatory-genomics-unveiling-cellular-narratives/</guid>

					<description><![CDATA[Researchers at Dana-Farber Cancer Institute, in collaboration with renowned institutions such as The Broad Institute of MIT and Harvard, Google, and Columbia University, have unveiled an innovative artificial intelligence model named EpiBERT. This groundbreaking model has been specifically engineered to predict gene expression across diverse human cell types, ultimately advancing our understanding of regulatory genomics. [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Researchers at Dana-Farber Cancer Institute, in collaboration with renowned institutions such as The Broad Institute of MIT and Harvard, Google, and Columbia University, have unveiled an innovative artificial intelligence model named EpiBERT. This groundbreaking model has been specifically engineered to predict gene expression across diverse human cell types, ultimately advancing our understanding of regulatory genomics. By decoding the intricate cellular landscapes, this study offers significant insight into how genes are expressed, regulated, and influenced within various contexts.</p>
<p>The allure of EpiBERT arises from its deep learning foundation, drawing inspiration from BERT—a model originally developed for natural language processing. Just as BERT learns from vast textual data to create coherent sentences, EpiBERT has been trained on a substantial genomic dataset encompassing hundreds of human cell types. The underlying mechanics involve feeding the model with whole genomic sequences, which span approximately 3 billion base pairs, along with intricate maps of chromatin accessibility. Such maps reveal which sections of the DNA are unwound and transliterated into biological function by the cell.</p>
<p>The initial training phase of EpiBERT focused on establishing the relationship between DNA sequences and chromatin accessibility within specific cell types. This foundational learning plays a critical role in the model&#8217;s subsequent ability to predict the activation of particular genes, providing valuable insights into cellular behavior. By accurately identifying regulatory elements—segments of the genome acknowledged by transcription factors—EpiBERT develops a generalized predictive framework, or &quot;grammar,&quot; for gene regulation across various cell types.</p>
<p>This regulatory framework operates similarly to how a language model like ChatGPT constructs meaningful linguistic patterns by sifting through extensive text examples. The capacity of EpiBERT to comprehend chromatin accessibility not only enables it to predict functional bases but also allows it to estimate RNA expression levels for previously unobserved cell types. This capability opens new avenues for exploring how different cells respond to internal and external stimuli in a very nuanced manner. </p>
<p>EpiBERT further enriches the field of regulatory genomics by addressing an elementary yet profound question: What distinguishes one cell type from another if all cells contain the same genome sequence? The answer lies predominantly in the regulation of gene expression—the timing, extent, and specificity with which genes are activated. Approximately 20% of the human genome codes for various regulatory elements that orchestrate these expression patterns; however, the precise locations and functionalities of these regulatory codes remain largely unexplored. By leveraging EpiBERT&#8217;s predictive power, researchers can shed light on these crucial elements that govern cellular identity and function.</p>
<p>The implications of this research extend beyond basic biological insights, potentially paving the way for breakthroughs in our understanding of human diseases. The understanding gained from the EpiBERT model may illuminate how mutations in regulatory elements disrupt cellular function, contributing to pathological conditions such as cancer. By elucidating the underlying mechanisms that dictate gene regulation, EpiBERT may help identify novel therapeutic strategies for tackling diverse cancers and other genetic disorders.</p>
<p>EpiBERT&#8217;s development was made possible through an impressive collaboration backed by significant funding sources. Organizations such as the Broad Institute, the Novo Nordisk Foundation, and the National Genome Research Institute offered their financial support, while Google provided vital computational resources with its Tensor Processing Unit (TPU) technology. This collaborative effort underscores the necessity of combining expertise from multiple disciplines to tackle the complex challenges within modern genomics.</p>
<p>As we delve into this study, the methodology employed in constructing and validating the EpiBERT model becomes clear. By utilizing a multi-modal approach, researchers harness various types of data—genomic sequences, chromatin state information, and expression profiles—enabling the model to perform cell type-agnostic predictions. This methodology not only ensures versatility in its applications but also enhances its relevant predictive accuracy across a broad spectrum of biological contexts.</p>
<p>Additionally, the cutting-edge nature of EpiBERT highlights the transformative impact of artificial intelligence in life sciences. Models like EpiBERT not only exemplify the potential for AI to revolutionize our insights into biological systems but also facilitate new research methodologies that can be readily adapted across various scientific disciplines. The insights gleaned from EpiBERT could benefit fields extending well beyond genomics alone.</p>
<p>In conclusion, the unveiling of EpiBERT marks a significant milestone in the field of molecular biology, merging the realms of artificial intelligence and genomics. With its promise to enhance our comprehension of gene regulation, EpiBERT is poised to contribute to ongoing research efforts that unravel the complex narrative of human health and disease. As researchers continue to parse the vast complexities of regulatory networks, EpiBERT stands as a testament to human ingenuity, bridging the gap between computational prowess and biomolecular understanding.</p>
<p>The successful application of EpiBERT paves the way for future research endeavors that aim to decode the complexities of gene regulation further. By building on the foundations laid by this model, subsequent studies might refine our knowledge of how regulatory elements function in health and disease. The ongoing collaboration between research institutions is crucial in ensuring that the powerful tools developed will be readily accessible to a global scientific community, thereby fostering innovation across the life sciences.</p>
<p>As EpiBERT continues to be utilized in various research projects, the anticipated discoveries will not only deepen our understanding of cellular mechanisms but will also hold the potential to transform clinical practices by providing a more comprehensive approach to genetic and epigenetic research. These advancements reinforce the significance of interdisciplinary collaboration in propelling the boundaries of current scientific understanding.</p>
<p>In summation, EpiBERT exemplifies the convergence of artificial intelligence and genomic research, fostering hope for groundbreaking developments in medical science. With ongoing research initiatives, the insights gained from EpiBERT will be instrumental in tackling future challenges posed by genetic diseases, further establishing its role as a key player in the evolving landscape of genomic research.</p>
<p><strong>Subject of Research</strong>: Gene expression and regulatory genomics<br />
<strong>Article Title</strong>: A multi-modal transformer for cell type agnostic regulatory predictions<br />
<strong>News Publication Date</strong>: January 29, 2025<br />
<strong>Web References</strong>: <a href="https://www.cell.com/cell-genomics/fulltext/S2666-979X(25)00018-7">Cell Genomics</a><br />
<strong>References</strong>: N/A<br />
<strong>Image Credits</strong>: Courtesy of Dana-Farber Cancer Institute<br />
<strong>Keywords</strong>: EpiBERT, regulatory genomics, gene expression, artificial intelligence, transcription factors, cancer research, chromatin accessibility, deep learning, molecular biology, genomics.</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">24845</post-id>	</item>
	</channel>
</rss>
