<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>deep learning frameworks for biological data &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/deep-learning-frameworks-for-biological-data/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 06:51:49 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>deep learning frameworks for biological data &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Hydra: An Interpretable AI Ensemble Tames the Chaos of Single-Cell Data</title>
		<link>https://scienmag.com/hydra-an-interpretable-ai-ensemble-tames-the-chaos-of-single-cell-data/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 06:51:49 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[advancements in single-cell sequencing analysis]]></category>
		<category><![CDATA[Alzheimer's disease]]></category>
		<category><![CDATA[cell type annotation]]></category>
		<category><![CDATA[challenges in identifying rare cell types]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[computational methods for noisy biological data]]></category>
		<category><![CDATA[data sparsity and noise in single-cell datasets]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning frameworks for biological data]]></category>
		<category><![CDATA[ensemble learning]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[feature selection in high-dimensional genomics]]></category>
		<category><![CDATA[Integrated Gradients]]></category>
		<category><![CDATA[interpretable deep learning in biology]]></category>
		<category><![CDATA[molecular signatures in single-cell sequencing]]></category>
		<category><![CDATA[multimodal integration]]></category>
		<category><![CDATA[multimodal single-cell data integration]]></category>
		<category><![CDATA[multiomics]]></category>
		<category><![CDATA[order and interpretability in single-cell analysis]]></category>
		<category><![CDATA[rare cell populations]]></category>
		<category><![CDATA[single-cell data analysis]]></category>
		<category><![CDATA[single-cell omics]]></category>
		<category><![CDATA[transparent AI models for cellular identity]]></category>
		<category><![CDATA[variational autoencoder]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226266</guid>

					<description><![CDATA[Researchers have unveiled Hydra, an interpretable ensemble deep learning framework that jointly performs feature selection and cell type annotation across unimodal and multimodal single-cell omics data, with particular strength in identifying rare cell populations in health and Alzheimer's disease.]]></description>
										<content:encoded><![CDATA[<p>Every cell in the human body carries a molecular autobiography, and modern single-cell sequencing technologies have become extraordinarily adept at reading it. Yet the data these technologies produce are notoriously unruly: thousands of genes measured per cell, vast numbers of missing measurements, and a biological reality in which the most interesting cell types are often the rarest. A new deep learning framework called Hydra, described in Molecular Systems Biology by Manoj M. Wagle of the University of Sydney and colleagues, including collaborators at the Massachusetts Institute of Technology, promises to bring order to this chaos while remaining unusually transparent about how it reaches its conclusions.</p>
<p>The core problem Hydra tackles is one that has haunted computational biologists for years. Single-cell datasets are sparse and noisy, with each cell described by thousands of molecular features, most of which are irrelevant to distinguishing one cell type from another. This high dimensionality buries the key molecular signatures that define cellular identity. Feature selection, the process of identifying which genes or genomic regions truly matter, has therefore become a critical preprocessing step. But existing tools struggle on three fronts: most are built exclusively for transcriptomic data and cannot handle newer multimodal assays; they systematically overlook small cell populations; and the deep learning methods that do perform well tend to operate as inscrutable black boxes.</p>
<p>Hydra&#8217;s architecture is built around an ensemble of variational autoencoders, or VAEs, a class of neural networks that learn compressed probabilistic representations of complex data. Each VAE in the ensemble is paired with a cell type classification head, and the whole system is trained jointly with a loss function that balances reconstructing the input data against correctly classifying cell types. The framework operates through two connected modules. The first performs ensemble feature ranking, producing a consensus list of cell-type-specific markers. The second is an annotation module that deploys an ensemble of simple neural network classifiers, trained on the selected features, to automatically assign cell type labels to new query datasets.</p>
<p>The most ingenious element of the design is how Hydra confronts class imbalance, the persistent bias that causes algorithms to favor abundant cell types while ignoring rare ones. Because a VAE learns the probability distribution underlying the data, it can generate synthetic cells that faithfully mimic underrepresented populations. Hydra combines this generative augmentation with random downsampling of dominant cell types, producing balanced training sets for each member of the ensemble. Each refined model is then interrogated using Integrated Gradients, a post hoc attribution technique that traces each prediction back to the individual input features that drove it, accumulating gradients along a path from a baseline input to the actual data point.</p>
<p>The ensemble approach proved decisive for reliability. When the researchers perturbed lung transcriptomic datasets through stratified subsampling and measured the consistency of feature importance scores using Pearson correlations, ensemble models were markedly more stable than single models. Stability improved with ensemble size up to 25 members, after which gains plateaued, so the team adopted 25 as the default configuration. Integrated Gradients also outperformed three alternative attribution methods, Saliency, GradientSHAP, and DeepLIFT, delivering higher feature stability and lower variability across all cell types, a property the authors argue is essential for identifying reproducible markers that generalize across studies and sequencing platforms.</p>
<p>Benchmarking was extensive. The team evaluated Hydra on 21 datasets spanning unimodal and multimodal single-cell technologies, comparing it against 13 state-of-the-art methods. On a subsampled Mouse Cell Atlas containing 20 cell types with a severe 100-to-2 imbalance between major and minor populations, Hydra&#8217;s selected features clustered biologically related cell types together, correctly grouping naive B cells with late pro-B cells and classical monocytes with promonocytes, distinctions that statistical methods such as Welch&#8217;s t test, the Wilcoxon rank-sum test, and Limma-Voom failed to make. For kidney proximal convoluted tubule epithelial cells, the top five genes Hydra identified, including GPX3, TIMP3, and FTH1, were all highly and specifically expressed in that population.</p>
<p>The functional relevance of these selections was confirmed through gene ontology enrichment analysis. The top features for naive B cells were enriched for B-cell activation and B-cell receptor signaling, neutrophil features pointed to migration, chemotaxis, and inflammatory response programs, and mesenchymal cell features captured extracellular matrix organization and skeletal system development. In cell type prediction tasks across 13 transcriptomic datasets, Hydra achieved the highest balanced accuracy in both intra-dataset validation on prostate urethra and colon data, at 68.86 percent and 86.90 percent respectively, and in inter-dataset benchmarking across 22 train-test pairs spanning kidney, lung, peripheral blood mononuclear cells, and retina, where it reached 73.81 percent balanced accuracy and a 66.77 percent macro F1-score.</p>
<p>Hydra&#8217;s multimodal capabilities may prove even more consequential. Modern assays can simultaneously measure gene expression, chromatin accessibility, and surface protein abundance in the same cell, and Hydra&#8217;s architecture processes each modality through dedicated encoder layers before merging them in a shared latent space of 100 neurons. Across seven multiome technologies, including SHARE-seq, SNARE-seq, sciCAR, CITE-seq, and the trimodal TEA-seq, Hydra outperformed seven competing integration methods, among them MOFA+, totalVI, MultiVI, scGLUE, scJoint, UMINT, and scMoMaT, achieving the highest balanced accuracy and macro F1-score in both intra-dataset and inter-dataset evaluations across 28 train-test splits.</p>
<p>The framework&#8217;s most striking demonstration came in Alzheimer&#8217;s disease. Using a previously published dataset profiling the transcriptome and epigenome of the medial frontal cortex, the team trained Hydra on healthy brain tissue containing 27 distinct cell populations and asked whether it could transfer those annotations to diseased tissue, where molecular changes can obscure cellular identity. Hydra maintained balanced accuracy of roughly 85 percent on healthy held-out samples, about 85 percent on early-stage Alzheimer&#8217;s samples, and 84 percent on late-stage samples. Critically, it preserved disease-relevant signals: its predictions captured cell type proportion shifts between early and late disease stages and maintained strong Spearman correlations with differential gene expression signatures derived from expert annotations, including for rare populations that competing methods failed to resolve.</p>
<p>Practicality matters too, and Hydra completes training in under ten minutes even on datasets of ten thousand cells, though its ensemble design does demand higher peak GPU memory than baseline methods. The authors are candid about limitations: as a supervised framework, Hydra can only predict cell types present in its training reference and cannot discover novel populations, and its performance depends on the quality of reference labels. They also note that multiome benchmarks relying on RNA-derived ground truth may understate the value of methods that genuinely integrate additional modalities. Even so, with code and documentation publicly available through the Sydney BioX repository, Hydra arrives at a moment when single-cell multiomics is becoming routine, offering researchers a tool that is simultaneously powerful, balanced toward the rare cells that often matter most, and interpretable enough to trust.</p>
<p><strong>Subject of Research:</strong> Interpretable deep generative ensemble learning for feature selection and cell type annotation in single-cell omics</p>
<p><strong>Article Title:</strong> Interpretable deep generative ensemble learning for single-cell omics with Hydra</p>
<p><strong>Article References:</strong> Wagle, M. M., Liu, C., Liu, Z., Wang, Y., Kellis, M., Patrick, E., &amp; Yang, P. (2026). Interpretable deep generative ensemble learning for single-cell omics with Hydra. <em>Molecular Systems Biology, 22</em>(7), 1161-1179. <a href="https://doi.org/10.1038/s44320-026-00208-7" rel="noopener noreferrer">https://doi.org/10.1038/s44320-026-00208-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s44320-026-00208-7" rel="noopener noreferrer">10.1038/s44320-026-00208-7</a></p>
<p><strong>Keywords:</strong> single-cell omics, deep learning, variational autoencoder, cell type annotation, feature selection, multimodal integration, Integrated Gradients, class imbalance, Alzheimer&#x27;s disease, rare cell populations, multiomics, ensemble learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226266</post-id>	</item>
	</channel>
</rss>
