<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>gene selection &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/gene-selection/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 13:49:47 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>gene selection &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Federated AI Framework Reaches Near-Perfect Accuracy in Cancer Gene Selection</title>
		<link>https://scienmag.com/federated-ai-framework-reaches-near-perfect-accuracy-in-cancer-gene-selection/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 13:49:47 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[bio-inspired optimization algorithms]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[cancer classification]]></category>
		<category><![CDATA[cancer gene expression analysis]]></category>
		<category><![CDATA[challenges in analyzing high-dimensional genetic data]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep neural networks for gene selection]]></category>
		<category><![CDATA[DNA microarray]]></category>
		<category><![CDATA[DNA microarray data analysis]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[federated AI frameworks in computational biology]]></category>
		<category><![CDATA[federated learning]]></category>
		<category><![CDATA[federated learning in biomedical data]]></category>
		<category><![CDATA[gene selection]]></category>
		<category><![CDATA[gene subset selection in cancer diagnosis]]></category>
		<category><![CDATA[high-accuracy cancer classification]]></category>
		<category><![CDATA[interdisciplinary approaches in bioinformatics]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[machine learning for tumor subtype identification]]></category>
		<category><![CDATA[manta ray foraging optimization]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[privacy-preserving machine learning]]></category>
		<category><![CDATA[privacy-preserving machine learning in healthcare]]></category>
		<category><![CDATA[stacked ensemble]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=228075</guid>

					<description><![CDATA[Researchers have developed a vertical federated deep learning framework using a binary manta ray optimization algorithm that selects cancer-related genes with near-perfect accuracy across five microarray datasets without centralizing sensitive patient data.]]></description>
										<content:encoded><![CDATA[<p>A team of computer engineering researchers in Iran has unveiled a new machine learning framework that combines federated learning, bio-inspired optimization, and deep neural networks to pick out the most informative genes from DNA microarray data, and the results are striking. Writing in the journal Cluster Computing, Seyedeh Mina Salimi, Babak Nouri-Moghaddam, and Abbas Mirzaei report that their approach achieved accuracies approaching 100 percent across five widely used cancer gene expression datasets, including Leukemia, Lymphoma, MLL, Ovarian, and SRBCT. The work addresses one of the most persistent headaches in computational biology: how to sift through thousands of gene measurements to find the small handful that actually matter for diagnosing disease, without pooling sensitive patient data in one place.</p>
<p>DNA microarray technology has transformed modern medicine by allowing scientists to measure the activity levels of thousands of genes simultaneously. In cancer research, this means clinicians can potentially distinguish tumor subtypes by their molecular fingerprints rather than by appearance under a microscope alone. Yet the very richness that makes microarrays powerful also makes them difficult to analyze. A typical dataset may contain tens of thousands of gene expression values for only a few dozen or a few hundred patients, a classic case of high dimensionality paired with small sample size. The authors note that large-scale data, a lack of reliable instances, and complex correlations among genes have all complicated the analysis process, making it easy for learning algorithms to latch onto noise instead of genuine biological signal.</p>
<p>The standard remedy is gene selection, a form of feature selection in which an algorithm searches for the subset of genes that best supports accurate classification. Over the past two decades, researchers have deployed filters, wrappers, and embedded methods, often hybridized with metaheuristic optimizers such as genetic algorithms, moth flame optimization, and intelligent water drop algorithms. These techniques have produced steady gains, but they typically assume that all the data sits in a single repository. In real-world medicine, that assumption is increasingly untenable. Hospitals, research centers, and laboratories hold their own patient records, and privacy regulations, institutional policies, and ethical concerns often prevent them from sharing raw genomic data with outside parties.</p>
<p>Federated learning has emerged as a compelling answer to this dilemma. In a federated setup, multiple participants, called federates, train models collaboratively while keeping their raw data local; only model updates or intermediate results travel across the network. The approach has proven its worth in mobile keyboard prediction and medical imaging, but the authors of the new study argue that many existing federated gene selection methods lack an optimal, hierarchical mechanism for feature reduction. In other words, they may coordinate model training across institutions but do not systematically narrow down which genes should be considered at each level of the process, leaving accuracy on the table when datasets are as extreme as microarrays.</p>
<p>The new framework, described as a vertical federated stacking deep learning system, tackles that gap with a two-stage distributed feature selection strategy that operates at both the local and central levels. The word vertical refers to how the data is partitioned: rather than splitting patients among federates, the framework splits the genes. Using a weight-based distribution, the full set of gene features is divided among scalable federates, so each participant holds expression values for a different slice of the genome while all of them see the same patients. This design keeps the collaborative process scalable, because adding more federates distributes the gene space more thinly rather than demanding more data from any single institution.</p>
<p>At the first stage, each federate independently searches its own slice of genes using a wrapper-based method built on the Binary Manta Ray Foraging Optimization algorithm, abbreviated BMRFO. The manta ray foraging optimization algorithm is a relatively recent nature-inspired optimizer that mimics the foraging behavior of manta rays, including their characteristic cyclone feeding patterns, to explore a search space efficiently. In its binary form, each candidate solution is a string of on-off decisions indicating whether a given gene should be included or excluded. A wrapper approach means the optimizer evaluates each candidate gene subset by actually training and testing a classifier on it, so the selection is directly tied to predictive performance rather than to statistical correlations alone. Previous studies have shown the manta ray optimizer to be competitive in feature selection tasks ranging from network intrusion detection to medical diagnosis, which motivated the team to adapt it for the federated setting.</p>
<p>Once the local searches are complete, the genes selected by every federate are merged and sent to a central server, which performs a final refinement step to identify the near-optimal gene set. This hierarchical structure is the core novelty of the method: a coarse-grained distributed search that trims the gene space in parallel, followed by a centralized fine-grained selection that reconciles the candidates and eliminates redundancy. The design mirrors the way a large research consortium might operate, with each member institution pre-screening its own portion of the genome before a coordinating body consolidates the findings. Crucially, because the raw expression data never leaves the federates, the framework preserves the privacy benefits that make federated learning attractive in the first place.</p>
<p>The final selected genes then feed into a robust deep stacking classification model that combines four different neural architectures: a multilayer perceptron, a convolutional neural network, a recurrent neural network, and a long short-term memory network, known as LSTM. Stacking, or stacked generalization, is an ensemble technique in which the outputs of multiple base learners are combined, often by a higher-level model, so that the strengths of one architecture can compensate for the weaknesses of another. The multilayer perceptron captures nonlinear relationships among the selected genes, the convolutional network detects local patterns in the expression profiles, and the recurrent and LSTM networks, which are designed to handle sequential data, can model dependencies across the ordered gene features. By fusing these complementary perspectives, the stacked model is better equipped to generalize from the small, high-dimensional datasets that dominate microarray research.</p>
<p>The evaluation spanned five benchmark datasets that have become standard proving grounds for gene selection algorithms. On the Leukemia dataset, the framework achieved an accuracy of 97.78 percent in predicting test samples, while on the Lymphoma, MLL, Ovarian, and SRBCT datasets it reached 99.99 percent. The authors report that these figures represent a significant improvement over previous methods, a notable claim in a field where benchmark accuracies have been climbing for years and where differences of even a fraction of a percent can be hard to earn. The datasets themselves are publicly available through repositories such as Kaggle and the biolab cancer projection pages, and the authors state that raw processed data supporting the findings are available from the corresponding author upon reasonable request from editors or reviewers.</p>
<p>Beyond the headline numbers, the study carries broader implications for how machine learning and genomics intersect. As personalized medicine pushes toward models trained on data scattered across hospitals and countries, frameworks that can perform sophisticated feature selection without centralizing sensitive records become increasingly valuable. The combination of a nature-inspired optimizer with a hierarchical federated architecture suggests a template that could extend beyond microarrays to other high-dimensional biomedical data, from single-cell RNA sequencing to proteomics. The work also adds to a growing literature on hybrid metaheuristic-deep learning systems, in which optimizers inspired by animal behavior, swarm intelligence, or physical processes are used to tame search spaces that would overwhelm brute-force methods. For a discipline where the goal is to find a handful of diagnostic genes among tens of thousands of candidates, and to do so across institutional boundaries, the Iranian team&#8217;s results indicate that distributing the problem, both the data and the search, may be exactly the right way to solve it.</p>
<p><strong>Subject of Research:</strong> Vertical federated learning with bio-inspired optimization for gene selection in DNA microarray cancer classification</p>
<p><strong>Article Title:</strong> Vertical federated deep stacking learning gene selection based on BMRFO algorithm</p>
<p><strong>Article References:</strong> Vertical federated deep stacking learning gene selection based on BMRFO algorithm. (n.d.). <a href="https://doi.org/10.1007/s10586-026-06626-4" rel="noopener noreferrer">https://doi.org/10.1007/s10586-026-06626-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10586-026-06626-4" rel="noopener noreferrer">10.1007/s10586-026-06626-4</a></p>
<p><strong>Keywords:</strong> federated learning, gene selection, DNA microarray, cancer classification, manta ray foraging optimization, deep learning, stacked ensemble, feature selection, bioinformatics, neural networks, LSTM, privacy-preserving machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">228075</post-id>	</item>
	</channel>
</rss>
