<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>application of deep learning in bioinformatics &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/application-of-deep-learning-in-bioinformatics/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 01:30:04 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>application of deep learning in bioinformatics &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Self-Adaptive Cell Graphs Push Single-Cell Clustering Into Sharper Focus</title>
		<link>https://scienmag.com/self-adaptive-cell-graphs-push-single-cell-clustering-into-sharper-focus/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 01:30:04 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced methods for single-cell data interpretation]]></category>
		<category><![CDATA[application of deep learning in bioinformatics]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[biological diversity in tissues]]></category>
		<category><![CDATA[cell graph]]></category>
		<category><![CDATA[cell-type clustering in single-cell data]]></category>
		<category><![CDATA[cell-type discovery]]></category>
		<category><![CDATA[challenges in single-cell gene expression data]]></category>
		<category><![CDATA[clustering]]></category>
		<category><![CDATA[clustering algorithms for single-cell genomics]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[dimensionality reduction]]></category>
		<category><![CDATA[dropout]]></category>
		<category><![CDATA[graph autoencoder]]></category>
		<category><![CDATA[high-dimensional single-cell data analysis]]></category>
		<category><![CDATA[improving accuracy of cell type identification]]></category>
		<category><![CDATA[SaDGAE deep learning framework]]></category>
		<category><![CDATA[self-supervised learning]]></category>
		<category><![CDATA[single-cell RNA sequencing analysis]]></category>
		<category><![CDATA[single-cell RNA-seq]]></category>
		<category><![CDATA[Transcriptomics]]></category>
		<category><![CDATA[unsupervised deep learning for cell clustering]]></category>
		<category><![CDATA[zero-inflated gene expression matrices]]></category>
		<category><![CDATA[ZINB]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=230019</guid>

					<description><![CDATA[A new unsupervised deep learning framework called SaDGAE uses a self-adaptive cell graph and a doubly enhanced graph autoencoder to achieve strong clustering performance on single-cell RNA sequencing data.]]></description>
										<content:encoded><![CDATA[<p>Every cell in the human body carries the same genome, yet a neuron, an immune cell, and a pancreatic beta cell behave in radically different ways because of which genes they switch on. Single-cell RNA sequencing, or scRNA-seq, has become one of the most powerful tools in modern biology precisely because it can read out that gene activity cell by cell, revealing the hidden diversity within tissues that bulk sequencing averages away. But the technology produces data that are notoriously difficult to analyze: tens of thousands of genes measured across thousands or millions of cells, with vast stretches of the expression matrix reduced to zeros by technical artifacts rather than true biology. A new unsupervised deep learning framework called SaDGAE, published in the International Journal of Data Science and Analytics, tackles the clustering problem at the heart of single-cell analysis and reports strong, competitive performance across a broad panel of benchmark datasets.</p>
<p>The central task that SaDGAE addresses is cell-type clustering. Before biologists can ask what a tissue is doing, they need to know which cells belong together, and in an unbiased experiment nobody can label the cells in advance. The algorithm must therefore group cells purely from their expression profiles, an unsupervised problem that has resisted clean solutions for years. The difficulty stems from three intertwined properties of scRNA-seq data. First, the data are extremely high-dimensional, with each cell described by a vector of gene expression values spanning the entire transcriptome. Second, the data are sparse, because most genes are not expressed in most cells. Third, and most troublesome, the data suffer from dropout events, in which a gene that is genuinely active in a cell registers as zero simply because the sequencing reaction failed to capture its messenger RNA molecules. These artifacts can make two cells of the same type look more different than two cells of entirely different lineages.</p>
<p>SaDGAE, developed by Xianghui Liu, Yanmei Hu, Shenglin Yang, Jie Wei, and Xiangtao Li at Chengdu University of Technology and Jilin University, is built as a graph autoencoder framework that jointly models two complementary views of the data: the gene expression patterns of individual cells and the relationships between cells. The architecture unfolds in a sequence of cooperating modules, each designed to correct a specific weakness of the raw data. The first stage is a denoising autoencoder based on the zero-inflated negative binomial, or ZINB, distribution, a statistical model that has become a standard way of describing scRNA-seq count data because it explicitly accounts for both the overdispersion of gene expression and the excess of zeros produced by dropout. By training an autoencoder to reconstruct the data under this model, SaDGAE generates dimensionally reduced representations that suppress technical noise while preserving the biological variation that actually distinguishes cell types.</p>
<p>The most distinctive innovation lies in what the authors call a self-adaptive cell graph construction mechanism. Graph-based methods represent each cell as a node and connect similar cells with edges, allowing a neural network to propagate information along those connections so that neighboring cells reinforce one another&#8217;s identities. The trouble is that the graph is usually built once, at the start of the analysis, from noisy or poorly normalized data. If the initial connections are wrong, every downstream step inherits those mistakes. SaDGAE instead treats the graph as a living structure: the cell-to-cell relationships are dynamically refined throughout training, optimized jointly with the embedding learning and the clustering objective itself. As the model learns better representations of the cells, it rewrites the graph; as the graph improves, it feeds better structural information back into the representations. This feedback loop is designed to capture subtle, context-specific connectivity that a static graph built from raw expression data would miss.</p>
<p>On top of this adaptive graph sits the doubly enhanced graph autoencoder, abbreviated DeGAE, which gives the framework its name. Conventional graph autoencoders encode node features through the graph and then attempt to reconstruct the original adjacency matrix from the learned embeddings. SaDGAE strengthens this reconstruction signal in two ways at once. It reconstructs the cell graph not only from the latent embeddings produced by the encoder but also from decoder-recovered features, effectively asking the network to agree on the cell-to-cell structure from two independent directions. This dual reconstruction provides a more informative training signal, pushing the model toward embeddings that are simultaneously faithful to the expression data and to the topology of the cell neighborhood, which in turn yields more discriminative representations for separating cell populations.</p>
<p>The final stage of the pipeline is a self-optimizing clustering module that works in a self-supervised manner. Rather than assigning cells to clusters in a single pass, the module constructs a target distribution over cluster assignments and iteratively aligns the learned embeddings with that target, sharpening confident assignments and correcting ambiguous ones. This strategy, common in deep clustering research, allows the clustering result and the representation learning to improve each other over the course of training, so that by the end of the optimization the cluster boundaries reflect both the denoised expression profiles and the refined cell graph.</p>
<p>To test the framework, the team evaluated SaDGAE on thirteen publicly available scRNA-seq datasets spanning a remarkable range of tissues and technologies. The benchmark collection includes the Pollen dataset (SRP041736) of neural cells, the Klein dataset (GSE65525), the Muraro pancreatic islet data (GSE85241), the Romanov hypothalamus dataset (GSE74672), the Wang lung dataset (GSE106960), the Xin pancreatic islet data (GSE81608), the Yan preimplantation embryo dataset (GSE109555), the Zeisel brain dataset (GSE60361), the Goolam embryo data (E-MTAB-3321), and the Quake 10x Bladder and Quake Smart-seq2 Diaphragm datasets (GSE109774), along with the widely used 10x Genomics PBMC-3k dataset of peripheral blood mononuclear cells. These datasets vary enormously in size, sequencing platform, and biological complexity, from a few hundred cells to tens of thousands, and from full-length Smart-seq2 protocols to droplet-based 10x chemistry, making them a demanding test of generality.</p>
<p>According to the authors, the extensive experiments, together with case studies and an additional large-scale scalability study, demonstrate that SaDGAE achieves strong and competitive clustering performance across this diverse collection. Just as important for working biologists, the clusters the method produces are biologically interpretable: the framework accurately recovers known marker gene patterns, meaning that the groups it discovers can be validated against established knowledge of which genes define which cell types. The scalability study is a notable addition, since single-cell experiments are growing explosively in size and algorithms that work beautifully on a thousand cells can grind to a halt on a million. The authors have also released their processed data and dataset metadata in a curated public repository on GitHub, supporting reproducibility for other groups.</p>
<p>The significance of this work extends beyond one algorithm&#8217;s benchmark scores. The field of single-cell computational biology has been converging on graph neural networks as a natural fit for the problem, since cells in a tissue genuinely form neighborhoods and communication networks, and several recent methods have combined graph convolutions, contrastive learning, and autoencoders for scRNA-seq analysis. SaDGAE&#8217;s contribution to this crowded landscape is the insistence that the graph itself should not be fixed. By making the cell graph self-adaptive and by reconstructing it through two parallel pathways, the method addresses one of the quiet failure modes of graph-based single-cell analysis: the propagation of early, noise-driven mistakes in cell-to-cell connectivity through every subsequent stage of the pipeline.</p>
<p>For the broader scientific community, better unsupervised clustering is a gateway capability. Cell-type discovery underpins atlas projects that aim to catalog every cell type in the human body, disease studies that search for rare pathogenic cell states hidden within tumors or inflamed tissue, and drug screens that monitor how treatments reshape cellular populations. Each of these applications begins with the same question SaDGAE was built to answer: given thousands of unlabeled cells, which ones are truly alike? If the self-adaptive graph philosophy proves robust across laboratories and platforms, it could become a standard component of the single-cell analysis toolbox, helping researchers turn the noisy, sparse, dropout-riddled output of a sequencing machine into clean, biologically meaningful maps of cellular identity. The study, received in January 2026 and published in August 2026, was supported by the National Natural Science Foundation of China, and its code and curated datasets are available to the community for further exploration.</p>
<p><strong>Subject of Research:</strong> Unsupervised deep graph autoencoder clustering of single-cell RNA-seq data</p>
<p><strong>Article Title:</strong> Sadgae: doubly enhanced graph autoencoder with self-adaptive cell graph for single-cell RNA-Seq clustering</p>
<p><strong>Article References:</strong> Liu, X., Hu, Y., Yang, S., Wei, J., &amp; Li, X. (2026). Sadgae: doubly enhanced graph autoencoder with self-adaptive cell graph for single-cell RNA-Seq clustering. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 287. <a href="https://doi.org/10.1007/s41060-026-01252-0" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01252-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01252-0" rel="noopener noreferrer">10.1007/s41060-026-01252-0</a></p>
<p><strong>Keywords:</strong> single-cell RNA-seq, clustering, graph autoencoder, deep learning, cell graph, dropout, ZINB, dimensionality reduction, self-supervised learning, bioinformatics, transcriptomics, cell-type discovery</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">230019</post-id>	</item>
	</channel>
</rss>
