<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>biomarker discovery &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/biomarker-discovery/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 25 Sep 2026 00:41:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>biomarker discovery &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Network Tool Sheds Light on Metabolomics Dark Matter for Biomarker Discovery</title>
		<link>https://scienmag.com/new-network-tool-sheds-light-on-metabolomics-dark-matter-for-biomarker-discovery/</link>
		
		<dc:creator><![CDATA[Alexandra Wallace]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 00:41:25 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advances in metabolomics detection methods]]></category>
		<category><![CDATA[annotation propagation]]></category>
		<category><![CDATA[bile acids]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[biomarker discovery in metabolomics]]></category>
		<category><![CDATA[Cell Reports Methods]]></category>
		<category><![CDATA[cellular metabolite analysis]]></category>
		<category><![CDATA[chemical chatter of metabolites]]></category>
		<category><![CDATA[dark metabolome]]></category>
		<category><![CDATA[Gut microbiome]]></category>
		<category><![CDATA[mass spectrometry]]></category>
		<category><![CDATA[mass spectrometry in metabolite identification]]></category>
		<category><![CDATA[metabolite identification challenges]]></category>
		<category><![CDATA[metabolite structure elucidation]]></category>
		<category><![CDATA[Metabolomics]]></category>
		<category><![CDATA[metabolomics and disease biomarkers]]></category>
		<category><![CDATA[metabolomics dark matter]]></category>
		<category><![CDATA[metabolomics research techniques]]></category>
		<category><![CDATA[molecular community network]]></category>
		<category><![CDATA[network analysis]]></category>
		<category><![CDATA[new network tools for metabolomics]]></category>
		<category><![CDATA[unidentified metabolites in metabolomics]]></category>
		<category><![CDATA[University of Central Florida]]></category>
		<category><![CDATA[Vladimir Boginski]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213671</guid>

					<description><![CDATA[A University of Central Florida team has developed the Molecular Community Network, an algorithmic tool that connects nearly 95 percent of mass spectrometry data into molecular communities, enabling the identification of unknown metabolites and already yielding a new class of gut microbe bile acids.]]></description>
										<content:encoded><![CDATA[<p>Every time a cell breaks down food, processes a drug, or responds to a chemical, it produces a cascade of small molecules known as metabolites. Reading that chemical chatter is the business of metabolomics, a field that promises new signs of disease, new ways to measure whether a treatment is working, and new insight into how diet and nutrition shape the body. Yet for all its promise, metabolomics has been haunted by an embarrassing problem: most of the molecules it detects cannot be identified. Researchers at the University of Central Florida have now built a tool designed to change that, and early results suggest the hidden fraction of the molecular universe may finally be within reach.</p>
<p>The vast majority of metabolites are detected with mass spectrometry, a laboratory technique that measures the mass and charge of tiny particles. When a sample is run through a mass spectrometer, each molecule produces a characteristic pattern called a mass spectrum, which in principle can be matched against libraries of known compounds. In practice, the matching works only for a minority of spectra. The rest belong to molecules that are clearly present and clearly measurable, but whose structures remain unknown. Researchers have come to call this inaccessible territory the dark matter of metabolomics, or the dark metabolome, a fitting name for a substance that can be detected but not identified.</p>
<p>The scale of the problem is enormous. According to Professor Vladimir Boginski of UCF, who led the new work, public repositories now hold roughly 8.4 million observed mass spectra of known and unknown molecules combined, and by current estimates the dark matter accounts for up to 90 percent of the entire observed molecular space. Every one of those unidentified spectra represents a potential biomarker, a molecular signature that could distinguish sick tissue from healthy, flag the early stages of disease, or reveal how a patient is responding to therapy. Instead, the data sit in databases, detectable but biologically silent, leaving a gap in what scientists can say about the genetic biome, cellular health, disease, and drug response.</p>
<p>Boginski and an interdisciplinary team of co-authors describe their solution in the most recent issue of Cell Reports Methods, in a paper titled Ordering molecular diversity in untargeted metabolomics via molecular community networking. Their tool, the Molecular Community Network, or MCN, takes a fundamentally different approach to organizing spectral data than the molecular networking methods that have dominated the field. Those traditional methods connect molecules only when their similarity scores, calculated from observed mass spectra, exceed a predetermined threshold. If two biologically related molecules fall just below the cutoff, the link between them is severed, and molecular families fragment into disconnected pieces.</p>
<p>That fragmentation is more than a cosmetic inconvenience. When a molecular family is broken apart, the known members lose touch with their unknown relatives, and the opportunity to infer the identity of the unknowns from their neighbors disappears. Boginski argues that this is precisely where biomarker discovery tends to stall. A molecule that reliably differentiates sick patients from healthy ones may show up clearly in a mass spectrum, yet if its structure cannot be assigned, the finding becomes a dead end. The spectrum is real, the signal is reproducible, but the molecule behind it remains a question mark, and a biomarker that cannot be identified cannot be developed into a clinical test.</p>
<p>The MCN replaces the rigid threshold with an algorithm drawn from the mathematics of community detection in large networks. Rather than asking whether each pair of molecules exceeds a fixed similarity score, the method divides the entire molecular network into its natural communities, groups of nodes that share a vast number of strong links internally while remaining only weakly connected to other groups. Once those communities are established, the strongest connections are retained to keep each community intact. The result is not a new structure imposed on the data, but a revelation of structure that was already present, hidden beneath the noise of near-threshold scores and fragmented families.</p>
<p>The practical consequence is striking. In the traditional approach, molecules that fall below the similarity cutoff can end up isolated, with no links at all. Under the MCN, almost every molecule in the network is linked to at least one neighbor, and those links typically connect molecules from similar molecular families. Boginski reports that nearly 95 percent of molecules are now connected and assigned to network communities, a dramatic expansion of the searchable molecular space. With nearly every unknown anchored inside a community alongside known compounds, the tool enables annotation propagation, the process of predicting the identity of an unknown molecule from the known identities of its neighbors.</p>
<p>Annotation propagation is where the dark matter begins to give up its secrets. If an unidentified spectrum sits inside a community dominated by, say, a particular family of lipids or bile acids, the algorithm can propose that the unknown belongs to that family, narrowing the search from millions of possibilities to a manageable set of candidates. Boginski notes that the team has shown this approach can indeed find previously unknown molecules based on their positions in the molecular community network, converting spectral orphans into annotated candidates that chemists can then pursue in the laboratory.</p>
<p>The method has already produced a discovery of its own. Using the MCN, Boginski and his co-authors identified a new class of bile acids, molecules produced by gut microbes that play important roles in digestion and metabolism. Among the newly recognized compounds is one that appears only in early infants, a finding that carries obvious implications for understanding the developing infant microbiome and the chemical dialogue between microbes and their youngest hosts. Crucially, the discovery followed the tool&#8217;s characteristic workflow: the new bile acids were first predicted computationally by their positions in the molecular community network, and then confirmed experimentally by the researchers in the lab, a sequence that demonstrates the pipeline from prediction to validation.</p>
<p>Boginski describes that first discovery as just the tip of the iceberg, and the reasoning behind his confidence is straightforward. Because the MCN runs on mass spectrometry data that has already been collected, the roughly 8.4 million spectra sitting in public repositories are now mapped and open to reanalysis. No new experiments are required to begin mining the dark metabolome; the raw material has existed for years, waiting for a method capable of organizing it. Each reanalysis pass has the potential to convert dead-end spectra into identifiable molecules, and each identified molecule becomes a candidate biomarker for disease, a measure of treatment efficiency, or a window into how diet and nutrition act on the body. For a field in which up to 90 percent of what can be measured has remained uninterpretable, the arrival of a tool that connects nearly every molecule to a molecular family may mark the moment the dark matter starts to give way.</p>
<p><strong>Subject of Research:</strong> Computational molecular networking for identifying unknown metabolites in untargeted metabolomics</p>
<p><strong>Article Title:</strong> UCF researcher develops new resource to identify unknown molecules for biomarker discovery</p>
<p><strong>Article References:</strong> UCF researcher develops new resource to identify unknown molecules for biomarker discovery. (n.d.). <a href="https://www.eurekalert.org/news-releases/1145440" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> metabolomics, dark metabolome, mass spectrometry, molecular community network, biomarker discovery, bile acids, gut microbiome, annotation propagation, network analysis, University of Central Florida, Cell Reports Methods, Vladimir Boginski</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213671</post-id>	</item>
		<item>
		<title>Budget Mass Spectrometers Can Match the Big Machines in Plasma Proteomics, Study Finds</title>
		<link>https://scienmag.com/budget-mass-spectrometers-can-match-the-big-machines-in-plasma-proteomics-study-finds/</link>
		
		<dc:creator><![CDATA[Kenneth Gardner]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 08:50:06 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[affordable mass spectrometry equipment]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[biomarker discovery in plasma]]></category>
		<category><![CDATA[blood plasma analysis]]></category>
		<category><![CDATA[clinical proteomics]]></category>
		<category><![CDATA[clinical proteomics advancements]]></category>
		<category><![CDATA[cost-effective proteomics tools]]></category>
		<category><![CDATA[data-independent acquisition]]></category>
		<category><![CDATA[impact of budget mass spectrometers]]></category>
		<category><![CDATA[LC-MS]]></category>
		<category><![CDATA[liquid chromatography-mass spectrometry (LC-MS)]]></category>
		<category><![CDATA[low-abundance biomarker detection]]></category>
		<category><![CDATA[MARS14]]></category>
		<category><![CDATA[mass spectrometry]]></category>
		<category><![CDATA[mass spectrometry in clinical research]]></category>
		<category><![CDATA[nanoparticle enrichment]]></category>
		<category><![CDATA[overcoming high-abundance protein interference]]></category>
		<category><![CDATA[plasma proteomics]]></category>
		<category><![CDATA[protein depletion]]></category>
		<category><![CDATA[protein identification in plasma]]></category>
		<category><![CDATA[Proteograph XT]]></category>
		<category><![CDATA[proteome coverage]]></category>
		<category><![CDATA[quantitative precision]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212282</guid>

					<description><![CDATA[A new study shows that widely accessible mass spectrometry instruments can reliably distinguish the performance of seven plasma preparation platforms, extending deep proteome analysis to laboratories without next-generation equipment.]]></description>
										<content:encoded><![CDATA[<p>Blood plasma is often described as the most information-rich liquid in the human body, a fluid that carries molecular echoes of nearly every organ and every disease process unfolding somewhere in the circulatory system. It is also, from an analytical chemist&#8217;s point of view, one of the most frustrating samples imaginable. A handful of abundant proteins, chiefly albumin and immunoglobulins, account for the overwhelming majority of the protein mass in plasma, while the low-abundance molecules that most interest biomarker hunters, such as signaling proteins, tissue leakage products and disease-specific fragments, are buried beneath them at concentrations many orders of magnitude lower. A new study published in Clinical Proteomics now shows that laboratories working with standard, widely available mass spectrometry equipment can navigate this problem far more effectively than many had assumed, opening the door to plasma proteomics for research groups that cannot afford the newest generation of instruments.</p>
<p>The research, led by Yeongshin Kim, Junho Park, Dongyoon Shin and Youngsoo Kim of CHA University in Seongnam, South Korea, set out to answer a deceptively simple question: do the elaborate plasma preparation platforms that have been benchmarked on cutting-edge mass spectrometers behave the same way when run on more accessible liquid chromatography-mass spectrometry, or LC-MS, instrumentation? The answer matters because most published evaluations of these platforms have relied on next-generation instruments with exceptional sensitivity and speed, leaving a genuine gap in knowledge for the many laboratories around the world that operate older or more modest equipment. If platform performance were fundamentally tied to top-tier hardware, the democratization of plasma biomarker discovery would stall; if not, the field could expand dramatically.</p>
<p>To address the question, the team processed a commercial pooled plasma sample through seven different preparation strategies. These included neat plasma with no treatment at all, serving as the baseline; the Multiple Affinity Removal System Human 14, known as MARS14, an antibody-based column that strips out the fourteen most abundant plasma proteins; and perchloric acid precipitation with neutralization, a chemical approach that preferentially precipitates abundant proteins while leaving many low-abundance species in solution. Alongside these depletion methods, the researchers tested four enrichment platforms: ENRICH-iST, Mag-Net, Proteonano and Proteograph XT. Each of these uses a different physical or chemical principle, from nanoparticle surfaces to specialized capture chemistries, to concentrate the dilute population of low-abundance proteins that standard workflows miss.</p>
<p>All samples were then analyzed in data-independent acquisition, or DIA, mode, a mass spectrometry strategy that systematically fragments all ions within defined mass windows rather than selecting individual precursors. DIA has become the workhorse of modern quantitative proteomics because it produces highly reproducible measurements across large sample cohorts, reducing the stochastic missing values that plague older data-dependent approaches. The choice of DIA was itself significant: it is the acquisition mode most commonly implemented on accessible instrumentation, so any conclusions drawn from the study would apply directly to the laboratories the researchers hoped to reach.</p>
<p>The results were striking. Untreated neat plasma yielded an average of just 777 identified proteins, a figure that captures the brutal reality of plasma&#8217;s dynamic range problem. The two depletion platforms performed substantially better, identifying between 1,278 and 1,468 proteins on average. But the enrichment platforms left the depletion approaches behind, with coverage ranging from 1,891 to a remarkable 6,060 proteins. Proteograph XT, the nanoparticle-based platform, delivered the deepest coverage of all and achieved the lowest missing value rate in the study, at just 1.9 percent, meaning that nearly every protein it detected was quantified consistently across the analysis rather than appearing and disappearing between runs.</p>
<p>Yet the study&#8217;s most important finding may be the one that complicates the simple narrative that more proteins equals better science. The researchers observed that greater proteome coverage did not always correlate with better quantitative precision. A platform that identifies thousands of additional proteins may do so at the cost of noisier measurements, and a coefficient of variation that looks acceptable for one platform may be unacceptable for another. Because the ultimate goal of plasma proteomics is reliable quantification, particularly when comparing patient cohorts to find proteins that distinguish disease from health, precision matters as much as depth. The study makes clear that platform choice significantly affected both the abundance profiles of the resulting datasets and the coverage of clinically relevant proteins, meaning that laboratories must match their preparation strategy to their biological question rather than simply chasing the largest protein count.</p>
<p>This platform-specific reshaping of the plasma proteome is a phenomenon that has been noted in previous work but never before systematically confirmed on accessible instrumentation. Each preparation method imposes its own bias: antibody-based depletion removes not only its targets but also proteins that travel bound to them, such as carrier proteins shuttling hormones and metabolites; chemical precipitation can lose proteins that co-precipitate with the abundant fraction; and nanoparticle enrichment selects for proteins with affinity for particular surface chemistries, so different particles fish out different, partially overlapping slices of the proteome. The consequence is that two laboratories studying the same plasma pool with different platforms may report substantially different protein abundance profiles, a fact that has major implications for reproducibility and for the design of multi-center biomarker studies.</p>
<p>Crucially, when the researchers compared their results with recent benchmarks generated on next-generation mass spectrometers, the platform-specific differences they observed reproduced those earlier findings. This is the study&#8217;s central vindication: the accessible LC-MS setups captured the same key distinctions among preparation platforms that the flagship instruments had revealed. In other words, the relative ranking of platforms, the patterns of proteome reshaping and the qualitative conclusions about which strategies best expose the low-abundance proteome all held true on standard equipment. What the newest instruments add is primarily depth and throughput, not a fundamentally different picture of how the platforms behave.</p>
<p>The practical implications extend well beyond methodological housekeeping. Plasma biomarker discovery has accelerated enormously in recent years, fueled by large-scale efforts such as the Human Plasma Proteome Project and by growing interest in early cancer detection, neurodegenerative disease monitoring and cardiovascular risk prediction. But much of that progress has been concentrated in well-funded centers equipped with the latest triple-TOF or Orbitrap Astral class instruments. If the platform evaluation framework demonstrated here holds, hospital-affiliated laboratories, research institutes in lower-resource settings and clinical translation teams can now participate meaningfully in plasma proteomics using instrumentation they already own. The study&#8217;s authors frame this as supporting the broader adoption of plasma proteomics in settings where next-generation instrumentation is not routinely available, and the data appear to justify that framing.</p>
<p>There remain caveats that the field will need to keep in view. The study used a single commercial pooled plasma sample, an excellent control for technical comparison but not a substitute for testing platforms across real patient cohorts with their biological variability, pre-analytical noise and disease-specific matrix effects. The evaluation also reflects one DIA acquisition strategy on one class of accessible instrument, and other configurations could shift the absolute numbers, even if the relative platform behavior is likely to persist. And the finding that coverage and precision can trade off against each other means that no single platform emerges as a universal winner; instead, the study provides something arguably more useful, a rigorous, reproducible map of what each of seven platforms delivers on hardware that most laboratories can access. As plasma proteomics moves from discovery science toward clinical application, that kind of practical, instrument-agnostic evidence may prove to be exactly what the field needs to turn an information-rich body fluid into a routine diagnostic resource.</p>
<p><strong>Subject of Research:</strong> Comparative evaluation of plasma preparation platforms for proteomics using accessible liquid chromatography-mass spectrometry instrumentation</p>
<p><strong>Article Title:</strong> Comprehensive evaluation of plasma proteomics platforms toward practical applications using accessible mass spectrometry instrumentation</p>
<p><strong>Article References:</strong> Kim, Y., Park, J., Shin, D., &amp; Kim, Y. (2026). Comprehensive evaluation of plasma proteomics platforms toward practical applications using accessible mass spectrometry instrumentation. <em>Clinical Proteomics</em>. <a href="https://doi.org/10.1186/s12014-026-09631-2" rel="noopener noreferrer">https://doi.org/10.1186/s12014-026-09631-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12014-026-09631-2" rel="noopener noreferrer">10.1186/s12014-026-09631-2</a></p>
<p><strong>Keywords:</strong> plasma proteomics, mass spectrometry, biomarker discovery, data-independent acquisition, protein depletion, nanoparticle enrichment, Proteograph XT, MARS14, LC-MS, proteome coverage, quantitative precision, Clinical Proteomics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212282</post-id>	</item>
		<item>
		<title>AI Reads Three Layers of Breast Cancer Biology at Once to Sharpen Subtype Diagnosis</title>
		<link>https://scienmag.com/ai-reads-three-layers-of-breast-cancer-biology-at-once-to-sharpen-subtype-diagnosis/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 21:33:55 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[BMC Bioinformatics]]></category>
		<category><![CDATA[breast cancer]]></category>
		<category><![CDATA[breast cancer subtypes]]></category>
		<category><![CDATA[computational methods in cancer subtype diagnosis]]></category>
		<category><![CDATA[cross-attention]]></category>
		<category><![CDATA[DNA Methylation]]></category>
		<category><![CDATA[gene expression profiling in breast cancer]]></category>
		<category><![CDATA[gene regulation in cancer]]></category>
		<category><![CDATA[Hierarchical Cross-Attention Multi-omics framework]]></category>
		<category><![CDATA[improving breast cancer treatment strategies]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[molecular biomarkers for cancer prognosis]]></category>
		<category><![CDATA[molecular layers in cancer diagnosis]]></category>
		<category><![CDATA[multi-omics data integration in oncology]]></category>
		<category><![CDATA[multi-omics integration]]></category>
		<category><![CDATA[PAM50]]></category>
		<category><![CDATA[PAM50 gene signature limitations]]></category>
		<category><![CDATA[precision oncology]]></category>
		<category><![CDATA[prognostic genes]]></category>
		<category><![CDATA[subtype classification]]></category>
		<category><![CDATA[TCGA]]></category>
		<category><![CDATA[transparency in cancer classification]]></category>
		<category><![CDATA[tumor classification accuracy]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=210445</guid>

					<description><![CDATA[A new hierarchical cross-attention framework integrates mRNA, microRNA, and DNA methylation data to classify breast cancer subtypes more accurately and identify subtype-specific prognostic genes.]]></description>
										<content:encoded><![CDATA[<p>Breast cancer is not one disease. Under the microscope, tumors may look similar, but at the molecular level they split into distinct subtypes that respond differently to treatment and carry very different prognoses. Clinicians have long relied on a 50-gene signature known as PAM50 to sort breast tumors into categories such as Luminal A, Luminal B, Her2-enriched, and Basal-like. Yet even this standard tool struggles with certain boundaries, particularly the line separating Luminal B from Her2-enriched tumors, where misclassification can change the course of therapy. A new computational framework published in BMC Bioinformatics by Md. Neaz Ali and Suman Biswas of the Department of Statistics and Data Science at Islamic University in Kushtia, Bangladesh, aims to make that sorting both more accurate and more transparent, while pointing clinicians toward genes that may matter for survival.</p>
<p>The framework, called HCAM-BRCA, short for Hierarchical Cross-Attention Multi-omics, tackles a problem that has dogged computational oncology for years: how to combine different kinds of molecular data without losing the biological relationships between them. Modern cancer research generates information from multiple molecular layers simultaneously. Messenger RNA profiles reveal which genes are being actively transcribed into protein-building instructions. MicroRNA profiles capture short regulatory molecules that silence gene expression after transcription. DNA methylation data, measured at cytosine-phosphate-guanine sites across the genome, show chemical tags that can switch genes on or off without altering the underlying sequence. Each layer tells part of the story, and each layer regulates the others in a web of interactions that single-omics analyses simply cannot see.</p>
<p>Many existing integration approaches handle this complexity in limited ways. Some methods analyze one data type at a time, while others combine only two layers in pairwise fashion, for instance joining RNA and microRNA data or RNA and methylation data but never all three at once. More importantly, most approaches do not explicitly model the hierarchical regulatory structure that connects the layers: methylation influences transcription, microRNAs modulate messenger RNA, and feedback loops run in both directions. HCAM-BRCA was designed to capture exactly these relationships. The architecture applies modality-specific self-attention to each data type, allowing the model to learn which features within a layer depend on one another, and then deploys cross-attention mechanisms between layers so the model can learn how, for example, methylation patterns inform the interpretation of gene expression.</p>
<p>Attention mechanisms are the same mathematical machinery that powers modern large language models, and their great advantage here is interpretability. Rather than acting as an inscrutable black box, an attention-based model assigns weights that reveal which features it considered most important when making a decision. In HCAM-BRCA, those weights translate directly into a ranked list of genes and regulatory elements that drove each subtype classification. The learned low-dimensional embeddings, compact numerical representations of each tumor&#8217;s multi-omics profile, were then fed into a battery of conventional machine learning classifiers, allowing the researchers to test how well the attention-derived representation supported downstream prediction.</p>
<p>The results were evaluated on data from The Cancer Genome Atlas, the large public repository of de-identified tumor profiles that has become the workhorse of computational cancer research. Performance was measured with rigorous statistics, including the Matthews correlation coefficient, a demanding metric that accounts for class imbalance, and the area under the precision-recall curve, which is particularly informative when one subtype is rare. All metrics were reported as means with standard deviations across five-fold cross-validation, meaning the data were repeatedly split so that every model was tested on samples it had never seen during training. The researchers also addressed a common pitfall in cancer genomics: class imbalance, where some subtypes have far fewer samples than others. Using an adaptive synthetic sampling technique known as ADASYN, they generated balanced training sets to prevent the models from simply defaulting to the most common categories.</p>
<p>Across every classifier tested, HCAM-BRCA&#8217;s three-layer integration outperformed the pairwise RNA-microRNA and RNA-CpG strategies, confirming that the full hierarchical picture carries information that two-layer views miss. Among the twelve machine learning models evaluated, CatBoost, a gradient-boosting method, achieved the highest macro-average Matthews correlation coefficient at 0.7628 with a standard deviation of 0.0130, while Extra Trees attained the highest area under the precision-recall curve at 0.8925 with a standard deviation of 0.0032. Those numbers represent strong and stable discrimination across subtypes, and notably the framework performed well precisely where existing tools struggle most, in separating the clinically challenging Luminal B and Her2-enriched categories.</p>
<p>Accuracy alone, however, was only half the story. Because the model&#8217;s attention weights expose which features influenced each decision, the researchers could interrogate the biology behind the predictions. Attention-based feature prioritization surfaced genes associated with transcriptional regulation, chromatin organization, immune signaling, and established cancer-related pathways, a pattern consistent with what independent functional analyses using Gene Ontology and KEGG pathway annotations would expect from genuinely relevant candidates. In other words, the model was not latching onto statistical noise; it was converging on genes that biologists already recognize as players in tumor behavior, along with others that merit new investigation.</p>
<p>The most clinically provocative findings came from survival analysis. The team examined whether expression of the prioritized genes was associated with patient outcomes within specific subtypes, and several strong subtype-specific prognostic links emerged. In Basal-like tumors, the aggressive category that largely overlaps with triple-negative breast cancer, the genes CUX1, MYZAP, and CEBPA showed significant survival associations. In Luminal B tumors, HOXA10, a developmental regulator frequently implicated in cancer, and HLA-DQA2, an immune-related gene, carried prognostic weight. In Her2-enriched tumors, GDF10 and ZNF879 were associated with survival outcomes. Each of these associations suggests a potential biomarker that could, with further validation, help stratify patients within a subtype for more tailored follow-up or therapy.</p>
<p>The significance of such subtype-specific markers is easy to underestimate. A gene that predicts poor survival in Basal-like tumors may be irrelevant in Luminal B, and pooling all subtypes together in a single analysis can wash out these signals entirely. By classifying first and then probing survival within each molecular category, HCAM-BRCA mirrors the way precision oncology actually operates: treatment decisions are made for a specific patient with a specific tumor subtype, not for an average cancer. A framework that simultaneously classifies accurately and nominates candidate biomarkers within each class therefore addresses two bottlenecks in the same pipeline.</p>
<p>The study, published open access on 23 September 2026 with no external funding, arrives amid a broader wave of multi-omics integration methods, including graph convolutional approaches such as MOGONET and neural frameworks like moBRCA-net. What distinguishes HCAM-BRCA is its explicit attention to hierarchical cross-layer regulation and the interpretability that attention weights provide. The authors acknowledge that the work rests on publicly available TCGA data and used no new human subjects, which means the biomarker candidates remain hypotheses awaiting validation in independent cohorts and, eventually, clinical settings. Still, the combination of robust subtype classification, biologically coherent feature prioritization, and subtype-specific prognostic signals makes a compelling case that teaching artificial intelligence to read all three molecular layers of a tumor at once, and to explain what it read, could move computational oncology closer to the clinic. For patients facing the diagnostic gray zones that complicate breast cancer care today, tools that sharpen those boundaries while revealing the genes behind them are exactly the kind of advance the field has been waiting for.</p>
<p><strong>Subject of Research:</strong> Interpretable multi-omics integration using hierarchical cross-attention for breast cancer subtype classification and biomarker discovery</p>
<p><strong>Article Title:</strong> HCAM-BRCA: an interpretable multi-omics framework for breast cancer subtype classification and biomarker discovery</p>
<p><strong>Article References:</strong> Ali, M. N., &amp; Biswas, S. (2026). HCAM-BRCA: an interpretable multi-omics framework for breast cancer subtype classification and biomarker discovery. <em>BMC Bioinformatics</em>. <a href="https://doi.org/10.1186/s12859-026-06607-9" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06607-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06607-9" rel="noopener noreferrer">10.1186/s12859-026-06607-9</a></p>
<p><strong>Keywords:</strong> breast cancer, multi-omics integration, cross-attention, machine learning, biomarker discovery, TCGA, subtype classification, PAM50, prognostic genes, precision oncology, DNA methylation, BMC Bioinformatics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">210445</post-id>	</item>
		<item>
		<title>New SPAID Database Maps Hidden Autoantigens Behind Autoimmune Diseases</title>
		<link>https://scienmag.com/new-spaid-database-maps-hidden-autoantigens-behind-autoimmune-diseases/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:03:51 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[autoantigens]]></category>
		<category><![CDATA[autoantigens in autoimmune diseases]]></category>
		<category><![CDATA[autoimmune disease biomarkers]]></category>
		<category><![CDATA[autoimmune disease diagnostics]]></category>
		<category><![CDATA[autoimmune diseases]]></category>
		<category><![CDATA[bioinformatics in immunology]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[comprehensive autoantigen mapping]]></category>
		<category><![CDATA[epitopes]]></category>
		<category><![CDATA[HLA class I]]></category>
		<category><![CDATA[immune epitope validation]]></category>
		<category><![CDATA[Immunogenicity prediction]]></category>
		<category><![CDATA[mass spectrometry]]></category>
		<category><![CDATA[mass spectrometry in autoantigen discovery]]></category>
		<category><![CDATA[non-canonical proteins]]></category>
		<category><![CDATA[non-canonical proteins in autoimmunity]]></category>
		<category><![CDATA[non-coding genome translation]]></category>
		<category><![CDATA[novel autoantigen identification]]></category>
		<category><![CDATA[protein databases]]></category>
		<category><![CDATA[Proteomics]]></category>
		<category><![CDATA[rheumatoid arthritis]]></category>
		<category><![CDATA[SPAID database]]></category>
		<category><![CDATA[T-cell and MHC ligand assays]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204364</guid>

					<description><![CDATA[Researchers have launched SPAID, a comprehensive database that maps both canonical and non-canonical candidate autoantigens across 14 autoimmune diseases by integrating validated epitope evidence with large-scale proteomic analysis.]]></description>
										<content:encoded><![CDATA[<p>Autoimmune diseases, in which the immune system turns against the body&#8217;s own tissues, affect hundreds of millions of people worldwide and remain notoriously difficult to diagnose early and precisely. At the heart of every autoimmune response lies a molecular trigger: an autoantigen, a self-protein or peptide that the immune system mistakenly recognizes as foreign. Yet despite decades of research, the full landscape of these triggers remains incomplete, in part because scientists have traditionally focused only on canonical, well-annotated protein-coding genes. Now, a research team led by scientists at Sun Yat-sen University and collaborating institutions in China has unveiled SPAID, a comprehensive database designed to systematically catalog candidate autoantigens across 14 autoimmune disorders, including both canonical proteins and a vast, largely unexplored universe of non-canonical proteins translated from non-coding regions of the genome.</p>
<p>SPAID, which is freely accessible online at spaid.renlab.cn, organizes its evidence into two distinct levels. The first, a validated level, contains proteins carrying experimentally confirmed epitopes drawn from the Immune Epitope Database, supported by positive T-cell assays and major histocompatibility complex (MHC) ligand assays. The second, a proteomics-based level, aggregates disease-associated peptides and proteins identified through mass spectrometry from human patient samples, annotated with differential expression patterns, predicted immunogenicity scores, and functional features. This two-tier architecture allows researchers to distinguish between candidates backed by direct immunological experimentation and those flagged through high-throughput proteomic discovery that await laboratory validation.</p>
<p>The technical ambition behind SPAID is considerable. To capture non-canonical proteins, the team assembled candidate sequences from more than 660,000 non-coding RNA entries in RNAcentral and over 332,000 intronic sequences from the IntroVerse database. Each candidate was evaluated for coding potential using two independent algorithms, CPAT and CNCI, and only sequences passing both thresholds were retained. Open reading frames were then predicted with NCBI&#8217;s ORFfinder and translated into amino acid sequences, which were de-duplicated against the UniProt reference set. The result is a unified protein sequence space of 576,516 sequences, combining 42,444 canonical UniProt proteins with 534,072 non-canonical proteins, including 447,445 intron-derived and 86,627 ncRNA-derived candidates.</p>
<p>Onto this reference framework, the researchers mapped a wealth of experimental data. From 292 publications, they integrated T-cell and MHC ligand assay records, ultimately identifying 1,141 unique validated epitopes from 10 autoimmune diseases supported by 2,750 positive T-cell assay records, alongside 20,424 distinct epitopes from 21,566 positive MHC ligand assays across five diseases. In total, these experimentally supported epitopes mapped to 21,349 unique proteins, spanning 16,966 canonical, 1,681 intron-derived, and 2,702 ncRNA-derived proteins. The inclusion of non-canonical proteins at this level is particularly striking, as it suggests that proteins translated from non-coding RNAs and introns can serve as genuine immune targets in human autoimmunity.</p>
<p>The proteomics-based level is equally extensive. Drawing on public repositories including PRIDE, MassIVE.quant, jPOST, PeptideAtlas, and iProX, the team curated 675 human proteomic samples spanning 14 autoimmune diseases, stratified into 51 disease- and tissue-specific cohorts. Peptides were identified by searching tandem mass spectra against the unified sequence space using DIA-NN for data-independent acquisition datasets and MaxQuant for data-dependent acquisition, with stringent false discovery rate control of 1 percent at the peptide-spectrum match, peptide, and protein-group levels. This rigorous filtering was essential because non-canonical peptides carry a heightened risk of false-positive identification. The search yielded 176,363 mass spectrometry-identified peptides assigned to 26,085 disease-associated proteins, including 927 ncRNA-derived and 531 intron-derived proteins.</p>
<p>To transform raw protein identifications into biologically meaningful signals, SPAID performs differential expression analysis for each cohort, comparing disease samples against matched controls. Proteins were classified as disease-only detected, upregulated, downregulated, or other, and results across multiple cohorts were integrated using Robust Rank Aggregation to assess cross-study consistency. Across the 14 diseases, the platform identified 4,577 disease-only detected proteins, 2,571 significantly upregulated proteins, and 757 significantly downregulated proteins. Notably, the disease-only category included 193 ncRNA-derived and 28 intron-derived proteins, demonstrating that non-canonical translation products participate in disease-specific proteomic signatures rather than representing background noise.</p>
<p>One of the most consequential questions the team addressed was whether these non-canonical proteins are reproducible. In diseases supported by at least three independent proteomic cohorts, more than 40 percent of non-canonical proteins were repeatedly detected: 61.51 percent in psoriasis, 54.58 percent in systemic lupus erythematosus, 51.41 percent in Crohn&#8217;s disease, and 40.32 percent in rheumatoid arthritis. Even more striking, among the repeatedly detected proteins, expression patterns showed remarkable concordance, with 99.33 percent consistency in rheumatoid arthritis, 93.93 percent in lupus, and 87.72 percent in psoriasis. Cross-referencing with published literature revealed that only 1.10 percent of the 1,458 non-canonical proteins identified across the diseases had prior experimental support, meaning the overwhelming majority represent previously unrecognized translation products now documented at scale for the first time.</p>
<p>To pinpoint which of these proteins might actually provoke immune responses, SPAID incorporates an immunogenicity prediction pipeline. Every mass spectrometry-detected peptide was segmented into overlapping 8- to 14-mer fragments and evaluated for HLA class I presentation across 12 functional supertypes, integrating MHC binding affinity, peptide-MHC stability, and T-cell recognition probability. Candidates were then refined using PanPep, a machine-learning tool that estimates T-cell receptor interaction probabilities against a panel of 419 CDR3 sequences. Overall, 14.70 percent of the 176,363 detected peptides were classified as putatively immunogenic, mapping to 15,558 immunogenic proteins, or 59.64 percent of all disease-associated proteins in the database. Immunogenic candidates were strongly enriched among disease-only detected proteins, representing 69.24 percent of that subset, consistent with the idea that proteins elevated under inflammatory conditions feed the antigen-processing machinery that can expose sequestered self-determinants and cryptic epitopes.</p>
<p>By intersecting three features, disease-specific proteomic detection, predicted immunogenicity, and experimental epitope support, the team defined a high-confidence core set of 2,023 candidate autoantigens. The platform&#8217;s practical utility was then demonstrated in an independent rheumatoid arthritis cohort, where serum proteomics of five patients and five healthy controls identified 1,559 proteins, 132 of which were RA-associated. A two-step validation pipeline using SPAID recovered clinically established biomarkers such as gamma-interferon-inducible protein 16 (IFI16) and immunoglobulin mu heavy chain, while also flagging novel candidates. Nine proteins, including myosin-9 (MYH9), glutathione S-transferase P (GSTP1), and hemoglobin subunit gamma-1/2 (HBG1/2), harbored experimentally validated epitopes. APOA4, apolipoprotein A-IV, emerged as an entirely novel candidate with highly specific enrichment in RA samples and strong predicted immunogenicity but no prior literature link to the disease, illustrating how the database can surface unexpected therapeutic leads.</p>
<p>Beyond its scientific content, SPAID offers a polished web interface built on a MySQL backend with a Java-based server and interactive ECharts visualizations. Users can search by disease, tissue, or protein attributes, run BLAST searches against transcript, protein, and peptide datasets, and explore hierarchical gene, protein, and peptide pages featuring expression boxplots, volcano plots, 3D structural models from the Protein Data Bank or ColabFold, predicted post-translational modification sites generated with PTM-Mamba, and an MS/MS spectrum annotator. The authors are candid about limitations: immunogenicity predictions currently cover only HLA class I, omitting CD4 T-cell biology tied to HLA class II, and the underlying proteomic data skew toward accessible tissues such as blood and skin. Mass spectrometry, however stringent, cannot on its own prove functional translation or physiological epitope presentation. Positioned as a candidate discovery resource rather than a definitive catalog, SPAID nonetheless represents a foundational shift in how autoantigen research can be conducted, and its developers plan future expansions to include HLA class II data and broader tissue proteomics, potentially accelerating diagnostics and targeted therapies for millions of autoimmune patients.</p>
<p><strong>Subject of Research:</strong> A comprehensive database for disease-specific autoantigen discovery in autoimmune disorders</p>
<p><strong>Article Title:</strong> SPAID: a comprehensive database for disease-specific autoantigens in autoimmune disorders</p>
<p><strong>Article References:</strong> Deng, S., Wei, F., Pang, Y., Zhang, L., Zhi, S., Chen, T., Zuo, Z., Ren, J., Xie, Y., &amp; Luo, X. (2026). SPAID: a comprehensive database for disease-specific autoantigens in autoimmune disorders. <em>Advanced Biotechnology, 4</em>(2), Article 23. <a href="https://doi.org/10.1007/s44307-026-00117-8" rel="noopener noreferrer">https://doi.org/10.1007/s44307-026-00117-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44307-026-00117-8" rel="noopener noreferrer">10.1007/s44307-026-00117-8</a></p>
<p><strong>Keywords:</strong> autoimmune diseases, autoantigens, SPAID database, proteomics, non-canonical proteins, mass spectrometry, epitopes, HLA class I, immunogenicity prediction, rheumatoid arthritis, biomarker discovery, protein databases</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204364</post-id>	</item>
		<item>
		<title>Urine Chemistry Reveals Hidden Signatures of Parkinson&#8217;s Disease</title>
		<link>https://scienmag.com/urine-chemistry-reveals-hidden-signatures-of-parkinsons-disease/</link>
		
		<dc:creator><![CDATA[Diana Fleming]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:50:38 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[advances in metabolomics for]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[biomarker discovery for neurodegeneration]]></category>
		<category><![CDATA[chemical signatures of Parkinson's in bodily fluids]]></category>
		<category><![CDATA[dansylation]]></category>
		<category><![CDATA[Decoding]]></category>
		<category><![CDATA[dopamine metabolism]]></category>
		<category><![CDATA[early detection of Parkinson's disease through urine analysis]]></category>
		<category><![CDATA[Gut microbiome]]></category>
		<category><![CDATA[gut-brain axis]]></category>
		<category><![CDATA[mass spectrometry]]></category>
		<category><![CDATA[metabolic]]></category>
		<category><![CDATA[metabolic disturbances in Parkinson's disease]]></category>
		<category><![CDATA[metabolomic fingerprint of Parkinson's]]></category>
		<category><![CDATA[Metabolomics]]></category>
		<category><![CDATA[neurodegeneration]]></category>
		<category><![CDATA[non-invasive urine test for neurodegenerative disorders]]></category>
		<category><![CDATA[Parkinson's disease]]></category>
		<category><![CDATA[Parkinson's disease urinary biomarkers]]></category>
		<category><![CDATA[potential for urine-based Parkinson's disease monitoring]]></category>
		<category><![CDATA[role of urine in diagnosing movement disorders]]></category>
		<category><![CDATA[submetabolome mapping in Parkinson's research]]></category>
		<category><![CDATA[urinary amines and phenols in Parkinson's diagnosis]]></category>
		<category><![CDATA[urinary biomarkers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204144</guid>

					<description><![CDATA[Researchers used dansylated urinary metabolomics to reveal a distinctive amine and phenol chemical signature of Parkinson's disease and identify candidate biomarkers for earlier diagnosis.]]></description>
										<content:encoded><![CDATA[<p>Scientists have uncovered a detailed chemical fingerprint of Parkinson&#8217;s disease hidden in one of the most routinely collected and least invasive fluids in medicine: urine. In a new study published in npj Parkinson&#8217;s Disease, researchers mapped the submetabolome of dansylated urinary amines and phenols, showing that the small nitrogen- and phenol-containing molecules excreted by patients with Parkinson&#8217;s disease form a distinctive pattern that can separate them from healthy individuals with striking clarity. The work, which appeared online in November 2026, offers a fresh window into the metabolic upheaval that accompanies the neurodegenerative disorder and points toward a practical route to biomarkers that could one day support earlier diagnosis and better monitoring of disease progression.</p>
<p>Parkinson&#8217;s disease affects more than ten million people worldwide, and its numbers continue to climb as populations age. Yet the diagnosis remains stubbornly clinical, resting on the observation of motor symptoms such as tremor, rigidity, and slowness of movement. By the time those symptoms become obvious, a substantial fraction of the dopamine-producing neurons in the substantia nigra has already been lost, and no available therapy can restore them. Decades of research have made clear that Parkinson&#8217;s begins long before tremors appear, with disturbances in protein handling, mitochondrial function, inflammation, and metabolism unfolding across years or even decades. A reliable molecular readout of that process, drawn from an accessible body fluid, has been a long-standing goal of the field.</p>
<p>The new study addresses that goal through a targeted lens on the urinary metabolome. Rather than attempting to measure every small molecule in urine at once, the researchers focused on amines and phenols, two chemically related classes of compounds that include neurotransmitter breakdown products, microbial metabolites, and products of amino acid metabolism. To capture these molecules comprehensively, they used dansylation chemistry, a labeling technique in which dansyl chloride reacts with compounds bearing an amine or phenol group, attaching a fluorescent and easily ionizable tag to each one. This derivatization dramatically enhances the detectability of these compounds in liquid chromatography–mass spectrometry, boosting sensitivity, improving chromatographic separation, and suppressing interference from salts and other matrix components that normally complicate urine analysis.</p>
<p>The strategy allowed the team to profile thousands of tagged metabolite features in each urine sample with high reproducibility. Urine was collected from patients with Parkinson&#8217;s disease and from matched healthy controls, and the dansylated extracts were analyzed under standardized conditions. After rigorous preprocessing to align chromatographic peaks, remove noise, and normalize signal intensities across batches, the resulting data matrix captured the amine and phenol submetabolome of each participant in exquisite detail. Statistical and machine-learning approaches were then applied to identify the metabolite features that best discriminated patients from controls and to build predictive models capable of classifying new samples.</p>
<p>The analysis revealed a coherent disease signature rather than a scattering of random chemical differences. Among the compounds that shifted most consistently were metabolites tied to neurotransmitter metabolism, including products of the catecholamine pathways that are directly affected by the degeneration of dopaminergic circuits. Other discriminating features pointed to alterations in phenolic compounds, many of which originate in the gut, where microbial enzymes transform dietary constituents into phenols that are absorbed into the bloodstream and excreted by the kidneys. The involvement of these gut-derived molecules fits squarely within a growing body of evidence linking the intestinal microbiome to Parkinson&#8217;s disease, from changes in microbial composition reported in patient cohorts to the observation that gastrointestinal symptoms often precede motor onset by many years.</p>
<p>Beyond individual metabolites, the investigators examined the pathways in which the altered compounds participate. The results implicate disturbances in the metabolism of tyrosine and phenylalanine, the aromatic amino acids that serve as precursors to dopamine and to numerous phenolic products, as well as in tryptophan catabolism, which feeds both the serotonin and the kynurenine pathways and has been repeatedly connected to neurodegeneration and neuroinflammation. Shifts in these interconnected routes suggest that Parkinson&#8217;s disease is accompanied not by a single metabolic lesion but by a coordinated remodeling of how the body processes aromatic compounds, a remodeling that reflects the combined influence of the brain, the periphery, and the resident microbiota.</p>
<p>The translational payoff of the study lies in its biomarker candidates. Using feature-selection algorithms, the researchers distilled the thousands of measured variables down to a compact panel of metabolites that together classify samples with high accuracy in the discovery data and hold up under cross-validation. The panel&#8217;s performance was evaluated using standard metrics, including the area under the receiver operating characteristic curve, and the selected markers retained discriminative power when tested on independent sample sets. Enrichment analyses confirmed that the chosen compounds were not statistical artifacts but chemically meaningful indicators, clustering in the same metabolic pathways implicated by the broader dataset. A urine test built on such a panel could be repeated easily, costs little compared with imaging or cerebrospinal fluid analysis, and could in principle be deployed in clinics and community settings far beyond specialized movement disorder centers.</p>
<p>Methodological rigor underpins the credibility of these findings. Dansylated metabolomics is technically demanding, and the authors took extensive precautions to ensure that the observed differences reflected genuine biology rather than analytical drift. Internal standards were used to monitor derivatization efficiency, quality-control samples were interspersed throughout the analytical runs to track instrument stability, and batch effects were corrected statistically before group comparisons were made. Putative metabolite identifications were assigned with appropriate levels of confidence based on accurate mass, retention time, and comparison with labeled standards where available, following the community conventions for reporting metabolomics data. This attention to annotation standards matters, because it allows other laboratories to reproduce the measurements and to build on the reported signatures.</p>
<p>The study also carries implications for how Parkinson&#8217;s disease is understood at a systems level. Metabolomics sits at the downstream end of the biological information flow, integrating changes in genes, transcripts, proteins, and environment into a chemical readout of physiology. The urinary amine and phenol submetabolome, in particular, sits at the convergence of central neurotransmitter metabolism, peripheral amino acid handling, and gut microbial activity. Its alteration in Parkinson&#8217;s disease reinforces the view of the disorder as a multisystem condition in which the gut-brain axis and peripheral metabolism are active participants rather than bystanders. That perspective is already reshaping therapeutic thinking, with interventions targeting the microbiome, the enteric nervous system, and systemic metabolism joining the traditional focus on neurons of the substantia nigra.</p>
<p>Important caveats remain. The metabolic signature reported here was established in specific cohorts, and its generalizability across populations, disease stages, medications, diets, and comorbidities will need confirmation in large, prospective, multicenter studies. Levodopa therapy, which virtually all patients eventually receive, is itself a rich source of dopamine metabolites and must be carefully accounted for in any diagnostic application. Longitudinal data will be essential to determine whether the biomarker panel tracks disease progression, predicts conversion from prodromal states such as REM sleep behavior disorder, or responds to disease-modifying treatments once such treatments become available. Standardization of sample collection, storage, and processing across sites will likewise be critical before a urine-based test can enter routine practice.</p>
<p>Even so, the study represents a substantial step toward a long-elusive goal. It demonstrates that a chemically defined slice of the urinary metabolome, accessed through a well-established derivatization technique and interrogated with modern mass spectrometry and machine learning, carries enough disease-specific information to distinguish Parkinson&#8217;s patients from healthy controls with confidence. If validated at scale, the approach could complement emerging tools such as alpha-synuclein seed amplification assays and advanced imaging, offering a complementary, low-cost, and patient-friendly measure of the disease&#8217;s systemic chemistry. In a condition whose diagnosis currently depends on the arrival of irreversible motor damage, a simple urine test that reflects the underlying biology earlier would be a genuinely transformative addition to the clinical arsenal, and the present work provides a detailed molecular roadmap for building one.</p>
<p><strong>Subject of Research:</strong> Urinary amine and phenol submetabolome profiling for Parkinson&#x27;s disease biomarker discovery</p>
<p><strong>Article Title:</strong> Decoding the metabolic landscape of Parkinson’s disease: dansylated urinary amines and phenols submetabolomes for signature profiling and biomarker discovery</p>
<p><strong>Article References:</strong> Li, Z., Cui, P., Zhang, L., Huang, X., Zhou, Y., Xu, S., Mao, Y., Wang, Y., Liu, L., &amp; Zhang, Y. (2026). Decoding the metabolic landscape of Parkinson’s disease: dansylated urinary amines and phenols submetabolomes for signature profiling and biomarker discovery. <em>npj Parkinson&#x27;s Disease</em>. <a href="https://doi.org/10.1038/s41531-026-01572-9" rel="noopener noreferrer">https://doi.org/10.1038/s41531-026-01572-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41531-026-01572-9" rel="noopener noreferrer">10.1038/s41531-026-01572-9</a></p>
<p><strong>Keywords:</strong> Parkinson&#x27;s disease, metabolomics, urinary biomarkers, dansylation, mass spectrometry, biomarker discovery, gut microbiome, neurodegeneration, dopamine metabolism, gut-brain axis, Decoding, metabolic</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204144</post-id>	</item>
		<item>
		<title>How AI and Multi-Omics Are Unlocking the Hidden Microbial World Inside Tumors</title>
		<link>https://scienmag.com/how-ai-and-multi-omics-are-unlocking-the-hidden-microbial-world-inside-tumors/</link>
		
		<dc:creator><![CDATA[Morgan Morrow]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 22:23:59 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI-driven cancer microbiome analysis]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[artificial intelligence in oncology]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[cancer microbiome]]></category>
		<category><![CDATA[clinical translation]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Fusobacterium nucleatum]]></category>
		<category><![CDATA[Gut microbiome]]></category>
		<category><![CDATA[gut microbiome influence on cancer]]></category>
		<category><![CDATA[immunotherapy response]]></category>
		<category><![CDATA[integrating multi-omics for cancer diagnosis]]></category>
		<category><![CDATA[metagenomics]]></category>
		<category><![CDATA[microbial impact on cancer treatment response]]></category>
		<category><![CDATA[multi-omics]]></category>
		<category><![CDATA[multi-omics technologies in cancer research]]></category>
		<category><![CDATA[precision oncology]]></category>
		<category><![CDATA[role of microbiome in cancer progression]]></category>
		<category><![CDATA[systemic immune modulation by microbes]]></category>
		<category><![CDATA[tumor ecosystem and microbiota interactions]]></category>
		<category><![CDATA[tumor microenvironment]]></category>
		<category><![CDATA[tumor-associated bacteria]]></category>
		<category><![CDATA[tumor-associated microbiome]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203500</guid>

					<description><![CDATA[A Genome Biology review charts how multi-omics technologies and artificial intelligence are transforming cancer microbiome research from observational associations toward clinically actionable precision oncology tools.]]></description>
										<content:encoded><![CDATA[<p>Cancer has long been understood as a disease of corrupted genes and rogue cells, but a quieter story has been unfolding in laboratories around the world. Tumors are not just masses of malignant tissue; they are ecosystems, populated by bacteria and other microbes that appear to shape how cancers begin, how they grow, and how they respond to treatment. At the same time, the trillions of microbes living in the gut send systemic signals that influence immunity and metabolism far beyond the digestive tract. A comprehensive review published in Genome Biology now maps out how researchers are combining multi-omics technologies with artificial intelligence to decode this hidden biology, and what it will take to turn those discoveries into real clinical tools.</p>
<p>The stakes are enormous. Cancer affects roughly 20 million people each year, and projections suggest the annual burden could climb to 30.5 million by 2050. While tumor-intrinsic factors such as mutations and dysregulated signaling pathways remain central to oncology, researchers increasingly recognize that tumor-extrinsic factors, including everything non-cancerous within the tumor microenvironment and the broader tumor macroenvironment, powerfully influence disease trajectories. The cancer microbiome, spanning both the gut microbiome and the tumor-associated microbiome, has emerged as one of the most intriguing of these factors. Early studies focused on colorectal cancer simply because of anatomical proximity, but evidence now shows the gut microbiome acts systemically, modulating host immunity and metabolism across the body, and has been implicated in tumorigenesis, progression, treatment response, and immune-related adverse events.</p>
<p>The tumor-associated microbiome tells a different story. Unlike the gut&#8217;s rich microbial communities, intratumoral microbes exist in low abundance, sparse populations that vary dramatically by cancer type. Tumors exposed to the external environment, such as colorectal, gastric, and oral cancers, harbor relatively more microbial biomass, while pancreatic, liver, lung, and breast tumors are considered low-biomass settings. Evidence from experimental models and human datasets has linked these intratumoral communities to cancer progression, prognosis, and treatment response, and the microbes can localize both outside and inside cancer and immune cells, sometimes with distinct spatial organization. They influence their surroundings through infection, inflammation, and the production of metabolites, yet separating genuine microbial signals from laboratory contaminants remains one of the field&#8217;s hardest problems.</p>
<p>Computationally, the field has traveled a long road. Early studies relied on classical statistical tools such as differential abundance methods, including LEfSe and metagenomeSeq, to identify taxa associated with cancer risk, survival, and treatment response. But microbiome data are notoriously difficult: high-dimensional, sparse, zero-inflated, and compositional, properties that can reduce reproducibility. Network-based approaches like SparCC, CoNet, and SPIEC-EASI extended the toolkit by modeling microbial interactions and identifying community structures linked to cancer processes, yet they struggle with complex nonlinear relationships. Deep learning has now entered the picture, enabling integration of microbiome and multi-omics data to model higher-order interactions and improve biomarker discovery and clinical outcome prediction. The catch is that most AI models must work with high-dimensional but small-sample datasets, raising overfitting risks and threatening biomarker stability, especially without strong external or prospective validation.</p>
<p>Generating reliable data is the first battleground. Cancer microbiome studies produce diverse data types, each capturing different aspects of microbial composition, function, and host interaction. Partial 16S rRNA sequencing is affordable but usually resolves taxonomy only to the genus level; full-length 16S improves resolution to species; shotgun metagenomics captures all genes in a sample, enabling both taxonomy and functional prediction. Metatranscriptomics captures real-time gene expression but is technically demanding, limited by RNA instability and stringent handling requirements. Metaproteomics, which profiles expressed proteins, offers a more direct functional view but has barely touched cancer: a PubMed search as of June 2026 identified only nine cancer microbiome metaproteomics studies. Metabolomics rounds out the picture, measuring the small molecules microbes produce, with databases such as MiMeDB, the Natural Products Atlas, and MASST helping to attribute metabolites to microbial origins, while spatial metabolomics now maps region-specific metabolic changes within tumors.</p>
<p>Detecting intratumoral microbes demands special tools. Researchers have reanalyzed bulk RNA-seq and whole-genome sequencing data to infer microbial signals from non-human reads, but this approach is vulnerable to contamination, and a recent large-scale tumor whole-genome analysis found that after host subtraction and decontamination, detectable microbiome signals were largely restricted to orodigestive cancers. Specialized pipelines have emerged to help: CSI_Microbe extracts microbial reads from The Cancer Genome Atlas sequencing data, SAHMI denoises microbial signals from single-cell RNA sequencing, and INVADEseq adds a primer targeting the conserved 16S region to map microbes within individual human cells. Spatial technologies such as imaging mass cytometry with mass-tagged antibodies and desorption electrospray ionization mass spectrometry imaging can visualize microbial presence alongside host immune and tumor cells. The review also lays out practical standards for credible signals: negative controls, conservative host-read subtraction, evaluation of batch structure, and orthogonal validation through qPCR, culture, in situ hybridization, or spatial imaging.</p>
<p>Once data are trustworthy, AI modeling begins in earnest, and the review offers a sobering lesson: bigger is not always better. Benchmarking studies show that classical, regularized models remain strong baselines. In a 16S rRNA benchmark with 490 subjects and 6,920 features, L2-regularized logistic regression matched random forest performance while training faster and remaining more interpretable. A larger benchmark across 83 gut microbiome cohorts and 20 diseases found ridge regression and random forest among the best performers, with neural networks and gradient boosting not consistently outperforming them. Deep learning architectures, including multilayer perceptrons, transformers, graph neural networks, and autoencoders, expand modeling capacity for nonlinear and structure-aware analysis, and foundation models pretrained on large-scale microbiome data, such as MGM, GenomeOcean, and Evo2, promise transferable representations, but fine-tuning on small cancer cohorts still risks overfitting, and interpretability remains limited.</p>
<p>The applications are already impressive. Multi-view deep learning frameworks distinguish metastatic from non-metastatic colorectal cancer using gut microbial features, while methods like GDmicro combine graph convolutional networks with domain adaptation to improve cross-cohort robustness against differences in region, diet, and sequencing protocols. Integration frameworks such as VTrans use large-scale pretraining and selective co-attention to combine microbiome features with host transcriptomics and copy-number profiles, enhancing survival risk stratification in small cohorts. Interpretability methods, from SHAP values and integrated gradients to graph-based community explanations like Micah, are evolving from simple feature ranking toward direction-aware, network-level insights. Looking ahead, causal AI frameworks such as DAG-deepVASE, which combines deep networks with knockoff features to identify nonlinear causal relationships, could move the field beyond pure association, though the review stresses that such findings remain hypothesis-generating until validated.</p>
<p>Mechanistic evidence is accumulating for real biological effects. Intratumoral Fusobacterium nucleatum has been implicated in promoting tumor progression, metastasis, and chemoresistance through immune modulation, autophagy activation, and oncogenic signaling, while enterotoxigenic Bacteroides fragilis promoted tumorigenesis and metastasis in breast cancer models. Gut microbes directly metabolize therapeutic drugs: bacterial β-glucuronidase reactivates the inactive metabolite of irinotecan in the gut, causing diarrhea, and inhibiting this enzyme can preserve drug efficacy while limiting toxicity. In immunotherapy, antibiotic use before immune checkpoint inhibitor treatment is associated with poorer response and survival, fecal microbiota transplantation from responders enhances anti-PD-1 responses in preclinical models, and findings from the CIAO clinical trial showed that total intratumoral bacterial abundance was the only microbiome-related feature predicting immune checkpoint blockade response in head and neck cancer, with higher abundance linked to an immunosuppressive microenvironment.</p>
<p>Translating these discoveries into the clinic is the final and steepest climb. The review emphasizes a three-stage evidentiary hierarchy, analytical validation, clinical validation, and demonstrated clinical utility, that most cancer microbiome biomarkers have not yet climbed. Fecal metagenomic classifiers for colorectal cancer are the most mature, retaining accuracy around an AUC of 0.8 in independent cohorts, while tumor-intrinsic signatures have faced serious challenges over host-read misclassification and normalization artifacts. Interventional strategies show encouraging but preliminary signals, with responder-derived fecal transplants reinstating anti-PD-1 responses in some ICI-refractory melanoma patients, yet a fatal transmission of a drug-resistant bacterium during fecal transplantation underscores the safety stakes. The path forward, the authors argue, requires standardized reporting under frameworks like STORMS, contamination-aware pipelines, prospective multicenter validation, and AI systems treated as prioritization tools rather than oracles. Emerging paradigms, including agentic AI systems like Eubiota, human-in-the-loop frameworks, and digital twin approaches, may eventually knit microbiome data into iterative clinical translation, but only if every model is paired with uncertainty estimation and external validation. The era of the cancer microbiome is no longer a question of whether microbes matter in oncology, but of whether the field can prove it rigorously enough for patients to benefit.</p>
<p><strong>Subject of Research:</strong> Integration of multi-omics technologies and artificial intelligence for decoding the cancer microbiome and translating discoveries into clinical oncology applications</p>
<p><strong>Article Title:</strong> Decoding the cancer microbiome: multi-omics, AI, and translational opportunities</p>
<p><strong>Article References:</strong> Decoding the cancer microbiome: multi-omics, AI, and translational opportunities. (n.d.). <a href="https://doi.org/10.1186/s13059-026-04284-8" rel="noopener noreferrer">https://doi.org/10.1186/s13059-026-04284-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13059-026-04284-8" rel="noopener noreferrer">10.1186/s13059-026-04284-8</a></p>
<p><strong>Keywords:</strong> cancer microbiome, tumor-associated microbiome, gut microbiome, multi-omics, artificial intelligence, deep learning, biomarker discovery, immunotherapy response, Fusobacterium nucleatum, metagenomics, precision oncology, clinical translation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203500</post-id>	</item>
		<item>
		<title>Gene Selection Gets Smarter: Co-expression Networks Meet Genetic Algorithms</title>
		<link>https://scienmag.com/gene-selection-gets-smarter-co-expression-networks-meet-genetic-algorithms/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:18:41 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[bioinformatics feature selection methods]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[cancer classification]]></category>
		<category><![CDATA[co-expression networks]]></category>
		<category><![CDATA[computational biology data challenges]]></category>
		<category><![CDATA[dimensionality reduction in genomics]]></category>
		<category><![CDATA[disease classification gene markers]]></category>
		<category><![CDATA[gene co-expression network analysis]]></category>
		<category><![CDATA[gene feature selection]]></category>
		<category><![CDATA[genetic algorithms]]></category>
		<category><![CDATA[genetic algorithms for feature selection]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[high-throughput sequencing data analysis]]></category>
		<category><![CDATA[information-theoretic genetic operators]]></category>
		<category><![CDATA[integrating biology and evolutionary mathematics]]></category>
		<category><![CDATA[machine learning in biomedical data]]></category>
		<category><![CDATA[multi-objective optimization]]></category>
		<category><![CDATA[mutual information]]></category>
		<category><![CDATA[noise reduction in genetic datasets]]></category>
		<category><![CDATA[NSGA-II]]></category>
		<category><![CDATA[Precision medicine]]></category>
		<category><![CDATA[Weighted Non-dominated Sorting Genetic Algorithm]]></category>
		<category><![CDATA[WGCNA]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196239</guid>

					<description><![CDATA[A new hybrid algorithm called CJWGA combines gene co-expression networks with enhanced genetic operators to select small, accurate gene subsets from high-dimensional medical data.]]></description>
										<content:encoded><![CDATA[<p>Modern medicine is drowning in data, and a new study argues that the way out is not more computing power but a smarter partnership between biology and evolutionary mathematics. In research published in the Journal of Advanced Research, a team led by Zhilin Wang, Weiping Ding, Jinquan Zhang, Ali Asghar Heidari, Mingjing Wang and Huiling Chen introduces a feature selection framework called CJWGA, a Weighted Non-dominated Sorting Genetic Algorithm that combines gene co-expression networks with information-theoretic genetic operators. The method is designed to tackle one of the most stubborn problems in computational biology: how to find the handful of genes that truly matter for disease classification inside datasets containing thousands of candidate features, most of which are noise, redundancy, or statistical distraction.</p>
<p>The scale of the problem is easy to underestimate. High-throughput sequencing and mass spectrometry now allow laboratories to measure the expression of every gene in the human genome across hundreds of samples at once. A dataset might record 10,000 genes while including fewer than a hundred patients. This imbalance creates what statisticians call the curse of dimensionality: the number of possible feature subsets grows as two to the power of n, so for a dataset with 10,000 genes the search space is astronomically larger than anything a brute-force enumeration could ever cover. Worse, adding features does not reliably improve a model. Extra genes can introduce redundancy and noise, causing classifiers to overfit the training data while performing poorly on patients they have never seen. Running times grow as well, because computational complexity rises steadily with the number of features examined.</p>
<p>Existing feature selection strategies fall into three broad families, each with well-known trade-offs. Filtering methods, which rank genes using statistical measures such as mutual information, are fast and scalable but blind to the interactions between features. Wrapper methods, which evaluate subsets by feeding them to a classifier, capture those nonlinear relationships but at a punishing computational cost. Embedded methods such as LASSO regression and tree-based models select features during training, but none of these approaches ask the deeper biological question: which genes actually work together, and which modules of co-regulated genes drive the disease being studied? The new framework was built precisely to fill that gap, treating the biology of gene cooperation as the starting point rather than an afterthought.</p>
<p>The first stage of CJWGA relies on Weighted Gene Co-expression Network Analysis, or WGCNA, a technique originally proposed by Zhang and Horvath that constructs a weighted network linking genes whose expression levels rise and fall together across samples. Genes are not loners; they participate in biological processes through intricate webs of interaction, and WGCNA captures those relationships from a systems perspective. The pipeline begins with Z-score normalization of expression values, followed by a Pearson correlation matrix that is then transformed into a weighted adjacency matrix using a soft thresholding exponent chosen so the network follows a scale-free topology, in which a few highly connected hub genes dominate while most genes have few connections. A Topological Overlap Measure, which accounts for shared neighbors, is then fed into hierarchical clustering to identify modules of functionally related genes, with module eigengenes derived by principal component analysis.</p>
<p>But the authors recognized that relying on a single eigengene per module throws away too much information. A lone principal component cannot reflect the diversity of functions within a module, and it can be biased by outlier expression patterns. Their answer is a preprocessing step called IMGCNet, which uses conditional mutual information to rank genes within each module by how much extra information they carry about the disease label, given the eigengene is already known. A higher conditional mutual information value means a gene retains a strong dependency on the phenotype even after controlling for what the module representative already explains. Larger modules are allowed to retain more genes and smaller modules fewer, through a descending allocation rule that preserves the biological representativeness of each module without letting small, specialized groups flood the analysis.</p>
<p>The second stage hands the modules to an enhanced version of NSGA-II, the classic multi-objective genetic algorithm that balances competing goals by evolving a population of candidate solutions toward a Pareto front. Here the two objectives are minimizing the number of selected genes and maximizing classification accuracy, formalized with a binary decision vector over features and evaluated with a K-Nearest Neighbor classifier on a 70-30 train-test split. Crucially, the researchers designed a hierarchical encoding scheme: the first layer of each chromosome encodes which modules are selected, and the second layer encodes which genes within each chosen module survive. This two-layer structure preserves the biological meaning of the modularization rather than flattening it back into a flat string of bits.</p>
<p>The heart of the contribution lies in two new operators. The Combined Information Entropy Crossover Operator, or CIECO, computes a joint mutual information score across the genes selected in both parents, those selected in neither, and those selected in only one. The resulting value, transformed through a probabilistic function, decides whether crossover should prune doubly-selected genes, promote single-selected ones, or hold steady. When the score is positive, unselected genes carry little information and conservative trimming is favored; when it is negative, redundant double selections are removed and a few unselected genes are introduced to seek greater information content. The Joint Adaptive Mutation Operator, or JAMO, then fine-tunes individual genes using an adaptive rate that depends on iteration progress, the proportion of genes already selected in the module, and the ratio of joint mutual information between selected and unselected genes, with an exponent parameter that keeps the balance under control.</p>
<p>The experimental evaluation covered eight publicly available gene expression datasets, including Brain_Tumor1, Brain_Tumor2, CNS, Leukemia, Leukemia1, Leukemia2, Lung_Cancer and Prostate_Tumor, all with more than 5,000 features and sample sizes between 50 and 203. Against three specialist algorithms, FQEISS, WMOSS and WQEISS, CJWGA achieved the lowest classification error on the CNS, Leukemia, Leukemia1, Leukemia2 and Prostate_Tumor datasets, while selecting the smallest feature subsets on six of the eight datasets. On Inverted Generational Distance, a standard measure of how well a computed Pareto front approximates the true optimum, CJWGA scored zero, meaning perfect overlap with the reference front, on six datasets. Ablation experiments confirmed that both new operators contribute measurably: removing the crossover operator or the mutation operator individually degraded either accuracy or subset compactness. Parameter sweeps established that a crossover proportion of 0.2 and a mutation exponent of 3 offered the most robust results. Because joint mutual information is computed only within compact modules rather than across the entire feature space, the framework retains reasonable scalability even as dataset dimensionality grows.</p>
<p>The implications reach beyond benchmark tables. A feature selection method that respects gene co-expression relationships can point clinicians toward biologically meaningful biomarkers rather than statistical artifacts, a prerequisite for precision medicine where a compact, interpretable gene panel must support diagnosis and treatment decisions. The authors caution, however, that systematic biological interpretation of the selected genes remains future work, and they note that the framework could be extended to dimensionality reduction problems well outside genomics. As high-throughput biology continues to generate data faster than medicine can absorb it, tools like CJWGA suggest that the path forward lies in algorithms that speak both languages fluently: the language of information theory and the language of biological networks. The study is available as open access, supported by the National Natural Science Foundation of China and several provincial research programs.</p>
<p><strong>Subject of Research:</strong> Gene feature selection in high-dimensional medical gene expression data using co-expression networks and genetic algorithms</p>
<p><strong>Article Title:</strong> Advancing Gene Feature Selection: A Synergistic Approach with Co-expression Networks and Genetic Algorithms</p>
<p><strong>Article References:</strong> Wangy, Z., Ding, W., Zhang, J., Heidari, A. A., Wang, M., &amp; Chen, H. (2026). Advancing Gene Feature Selection: A Synergistic Approach with Co-expression Networks and Genetic Algorithms. <em>Journal of Advanced Research</em>. <a href="https://doi.org/10.1016/j.jare.2026.08.064" rel="noopener noreferrer">https://doi.org/10.1016/j.jare.2026.08.064</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.jare.2026.08.064" rel="noopener noreferrer">10.1016/j.jare.2026.08.064</a></p>
<p><strong>Keywords:</strong> gene feature selection, co-expression networks, WGCNA, genetic algorithms, multi-objective optimization, NSGA-II, mutual information, bioinformatics, cancer classification, precision medicine, high-dimensional data, biomarker discovery</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196239</post-id>	</item>
		<item>
		<title>Hidden messengers: matrix-bound vesicles rewrite the rules of tissue signalling</title>
		<link>https://scienmag.com/hidden-messengers-matrix-bound-vesicles-rewrite-the-rules-of-tissue-signalling/</link>
		
		<dc:creator><![CDATA[Gregory Coleman]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 12:38:35 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[Biomarkers]]></category>
		<category><![CDATA[cell-to-cell communication]]></category>
		<category><![CDATA[decellularization]]></category>
		<category><![CDATA[Drug delivery]]></category>
		<category><![CDATA[extracellular matrix]]></category>
		<category><![CDATA[extracellular vesicles]]></category>
		<category><![CDATA[immunomodulation]]></category>
		<category><![CDATA[local tissue communication]]></category>
		<category><![CDATA[matrix-bound vesicles]]></category>
		<category><![CDATA[microRNA]]></category>
		<category><![CDATA[molecular cargo transfer]]></category>
		<category><![CDATA[nanoscale membrane sacs]]></category>
		<category><![CDATA[Nature Reviews Bioengineering]]></category>
		<category><![CDATA[Regenerative Medicine]]></category>
		<category><![CDATA[spatially confined reservoirs]]></category>
		<category><![CDATA[therapeutic potential]]></category>
		<category><![CDATA[tissue engineering]]></category>
		<category><![CDATA[Tissue microenvironment]]></category>
		<category><![CDATA[tissue signaling]]></category>
		<category><![CDATA[tumour microenvironment]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194259</guid>

					<description><![CDATA[A new review argues that matrix-bound extracellular vesicles retained within the extracellular matrix form a distinct, tissue-specific signalling population with major implications for regenerative medicine and disease diagnostics.]]></description>
										<content:encoded><![CDATA[<p>Deep inside every tissue, beyond the reach of blood and lymph, a quiet postal service has been operating undetected for decades. A sweeping new review published in Nature Reviews Bioengineering argues that extracellular vesicles—the nanoscale membrane sacs that cells use to exchange molecular messages—come in two fundamentally different flavours, and that science has spent most of its attention on the wrong half. While liquid-phase vesicles drift through blood and other biofluids, carrying signals systemically, a second population remains anchored within the extracellular matrix itself, acting as spatially confined reservoirs of molecular cargo that encode the local state of a tissue. These matrix-bound extracellular vesicles, the authors contend, deserve to be treated as a distinct biological entity with their own rules, functions and translational promise.</p>
<p>The distinction is more than taxonomic housekeeping. Liquid-phase vesicles, which have fuelled a decade of biomarker discovery and therapeutic development, are subject to dilution, clearance and non-specific biodistribution the moment they enter circulation. Matrix-bound vesicles, by contrast, never leave home. They stay tethered to the fibrous network of collagen, glycoproteins and other macromolecules that gives tissue its structure, delivering their cargo of proteins, lipids and microRNAs to neighbouring cells in a strictly local, context-dependent manner. In doing so, they function not merely as passengers within the matrix but as functional components of it, shaping tissue development, regeneration and day-to-day homeostasis.</p>
<p>The technical case for treating matrix-bound vesicles as a separate population rests on molecular evidence. Lipidomic and RNA sequencing studies of vesicles extracted from decellularized extracellular matrix bioscaffolds have revealed cargo profiles that differ measurably from those of vesicles harvested from the surrounding fluid. Proteomic comparisons of liquid-phase and matrix-bound vesicles grown in both two-dimensional and three-dimensional cell cultures reinforce the picture of two biochemically distinct populations. Matrix-bound vesicles carry tissue-specific protein and microRNA signatures, suggesting that they act as regulatory components of the matrix rather than as incidental debris trapped in its fibres.</p>
<p>That tissue specificity is one of the most striking features of the new framework. Vesicles isolated from decellularized cardiac, skeletal muscle, bone and other tissues recapitulate the angiogenic and immunomodulatory properties of their parent matrices, and their microRNA profiles differ from tissue to tissue. In bone, matrix vesicle cargo such as the microRNA miR-125b accumulates within the mineralized matrix and inhibits bone resorption in mouse models. In skeletal muscle, vesicle-associated interleukin-33 has been shown to initiate a pro-regenerative shift in macrophage phenotype after injury, supporting functional recovery. The vesicles, in effect, carry a molecular record of the tissue they came from—and a set of instructions appropriate to that tissue.</p>
<p>The review also documents how this record changes with age and disease, and the implications are unsettling. Studies of aged breast tissue show that matrix-bound vesicles from older matrices carry cargo that promotes invasiveness in breast epithelial and cancer cells, suggesting that the aged microenvironment itself can actively contribute to tumour progression. Cardiac tissue-resident vesicles regulate fibroblast activation in an age- and sex-dependent manner, and multi-omics profiling of young and aged plasma and matrix-bound vesicles has identified anti-fibrotic microRNAs enriched in the young versions, with therapeutic activity validated in a heart-on-a-chip model. In colorectal cancer, vesicles trapped in decellularized tumour matrix preserve disease-associated signatures of the tumour microenvironment, while cancer-associated fibroblasts have been shown to produce matrix-bound vesicles that influence endothelial cell function. The matrix, in other words, is not a passive scaffold but an active archive of pathological state.</p>
<p>Therapeutically, the localized nature of matrix-bound vesicles is both their greatest asset and their central challenge. Because they act where they are placed, they are natural candidates for integration into engineered tissues and biomaterial delivery platforms. Matrix-bound vesicles embedded in cartilaginous extracellular matrix have enabled functional reconstruction of tracheal defects, vesicles from decellularized tumours have been used as platforms for targeting parent tumour cells and tumour-associated stromal cells, and injectable microsphere systems are being developed for sustained delivery in adipose tissue engineering. Immunomodulatory matrix-bound vesicles derived from urinary bladder matrix have mitigated influenza-mediated lung inflammation while preserving antiviral responses, eased rheumatoid arthritis in preclinical models, and alleviated particulate-induced periprosthetic osteolysis. Recent work even suggests these vesicles can accumulate in bone marrow and induce durable epigenetic changes in myeloid progenitors and macrophages, hinting at effects that outlast the vesicles themselves.</p>
<p>Compared with their liquid-phase cousins, matrix-bound vesicles may also sidestep some of the biodistribution problems that have plagued systemic vesicle therapies. Circulating vesicles must survive the bloodstream, cross vascular barriers and find the right tissue before releasing their cargo, and much of the dose is lost along the way. A vesicle pre-positioned within an implanted scaffold or hydrogel faces no such gauntlet. The trade-off is that delivery becomes inseparable from the biomaterial itself: the scaffold must retain the vesicles, present them to infiltrating host cells and release them, if release is desired, on a controlled schedule. This couples vesicle therapy directly to the maturing field of engineered extracellular matrices, decellularized scaffolds and bioprinted tissues.</p>
<p>Getting there, the authors caution, requires solving problems that the liquid-phase vesicle field has only partially addressed. Isolation of matrix-bound vesicles depends on decellularization protocols whose harshness can alter yield, purity and function, and different isolation methods produce vesicles with different biological behaviour. Characterization remains hampered by the heterogeneity of vesicle populations and by the lack of standardized reporting, although community frameworks such as the MISEV guidelines are pushing the field toward rigor. Downstream, translating cargo profiles into diagnostics or therapeutics will demand bioinformatics integration across proteomics, lipidomics and transcriptomics, and manufacturing clinical-grade material will require scalable, validated processes. The review singles out innovations in isolation, characterization, bioinformatics and bioengineered delivery as the four pillars on which translational success will rest.</p>
<p>What emerges from the analysis is a reframing of how biologists should think about the space between cells. The extracellular matrix has long been appreciated as a mechanical and structural environment that influences stem cell fate, angiogenesis and fibrosis. The new work positions matrix-bound vesicles as the signalling layer of that environment—a distributed, tissue-encoded communication network that operates alongside, and distinct from, the systemic vesicle traffic carried in blood, urine, saliva and other biofluids. If the framework holds, diagnostics could one day read the vesicular archive embedded in a biopsy or decellularized scaffold to reconstruct a tissue&#8217;s recent history, and regenerative therapies could seed engineered implants with vesicles pre-loaded with the molecular instructions a healing tissue needs. The quiet postal service in the matrix, long overlooked, may prove to be one of the most consequential mail routes in the body.</p>
<p><strong>Subject of Research:</strong> Matrix-bound extracellular vesicles as tissue-specific, matrix-anchored mediators of local intercellular signalling in health, ageing and disease</p>
<p><strong>Article Title:</strong> Matrix-bound extracellular vesicles</p>
<p><strong>Article References:</strong> Matrix-bound extracellular vesicles. (n.d.). <a href="https://doi.org/10.1038/s44222-026-00477-9" rel="noopener noreferrer">https://doi.org/10.1038/s44222-026-00477-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s44222-026-00477-9" rel="noopener noreferrer">10.1038/s44222-026-00477-9</a></p>
<p><strong>Keywords:</strong> extracellular vesicles, matrix-bound vesicles, extracellular matrix, tissue engineering, regenerative medicine, biomarkers, microRNA, decellularization, immunomodulation, tumour microenvironment, drug delivery, Nature Reviews Bioengineering</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194259</post-id>	</item>
		<item>
		<title>AI Learns to Predict Patient Outcomes Even When Key Omics Data Are Missing</title>
		<link>https://scienmag.com/ai-learns-to-predict-patient-outcomes-even-when-key-omics-data-are-missing/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 04:41:58 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI techniques for missing data]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[block-wise modality missingness in clinical research]]></category>
		<category><![CDATA[challenges in multi-omics data collection]]></category>
		<category><![CDATA[clinical outcome prediction]]></category>
		<category><![CDATA[cost-effective multi-omics analysis]]></category>
		<category><![CDATA[data integration]]></category>
		<category><![CDATA[genomic and proteomic data prediction]]></category>
		<category><![CDATA[handling incomplete multi-omics datasets]]></category>
		<category><![CDATA[imputation]]></category>
		<category><![CDATA[incomplete data]]></category>
		<category><![CDATA[latent representations]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for incomplete biological datasets]]></category>
		<category><![CDATA[missing modality learning]]></category>
		<category><![CDATA[missing omics data in precision medicine]]></category>
		<category><![CDATA[multi-omics]]></category>
		<category><![CDATA[multi-omics data integration]]></category>
		<category><![CDATA[personalized treatment prediction algorithms]]></category>
		<category><![CDATA[Precision medicine]]></category>
		<category><![CDATA[predictive medicine]]></category>
		<category><![CDATA[predictive modeling with missing omics layers]]></category>
		<category><![CDATA[robustness of AI models in healthcare]]></category>
		<category><![CDATA[Single-Cell Genomics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193762</guid>

					<description><![CDATA[A new review maps the machine learning methods that allow multi-omics models to predict clinical outcomes even when entire layers of patient data are missing.]]></description>
										<content:encoded><![CDATA[<p>Precision medicine has long promised a future in which a patient&#8217;s treatment is tailored to the molecular signatures written into their genome, transcriptome, epigenome, proteome and metabolome. In theory, combining these layers of biological information—the field known as multi-omics integration—should let algorithms forecast how a disease will progress, which drugs will work and which patients are at highest risk. In practice, however, real-world clinical cohorts almost never deliver the complete, neatly paired datasets that many machine learning models quietly assume. A new review published in Artificial Intelligence Review by Ricky Nguyen and Fatemeh Vafaee of the University of New South Wales in Sydney examines the growing family of techniques designed to keep multi-omics prediction working when entire layers of data are simply missing.</p>
<p>The problem the researchers describe is known as block-wise modality missingness, and it is endemic to clinical research. Sequencing a genome, profiling the methylome or running a mass-spectrometry-based proteomic assay each carries its own costs, technical demands and failure rates. A hospital may afford whole-exome sequencing for every patient in a cancer cohort but collect RNA-sequencing data for only a subset. Assays fail. Study designs evolve mid-project, adding omics layers that earlier patients never received. The result is a data matrix riddled with entire missing blocks rather than scattered gaps—and this pattern is far more damaging to conventional integrative pipelines than ordinary single-cell missingness.</p>
<p>Most existing multi-omics methods were built with fully paired data in mind. When confronted with incomplete cohorts, practitioners typically fall back on one of two workarounds: complete-case filtering, in which every patient lacking any omics layer is discarded, or point-wise imputation, in which missing values are filled in one element at a time. Both strategies carry serious drawbacks. Complete-case analysis can shrink a cohort so drastically that statistical power collapses, and it systematically biases the remaining sample toward patients who received the most thorough work-up—often those with better access to care or more advanced disease at diagnosis. Point-wise imputation, meanwhile, treats block-wise absence as if it were random noise, which it emphatically is not, and can manufacture false confidence in downstream predictions while masking genuine biological signal.</p>
<p>The review&#8217;s central contribution is a methodological taxonomy that organises the emerging solutions into coherent families. One prominent family comprises missingness-aware fusion architectures: models that explicitly encode which modalities are present for each patient and adapt their internal computations accordingly. Rather than demanding a full complement of inputs, these networks learn fusion functions that can operate on whatever subset of omics layers happens to be available, weighting contributions in a way that accounts for both the information content and the absence of particular views. The absence of a modality becomes a structured condition the model reasons about, rather than a defect it must repair.</p>
<p>A second family relies on shared latent representations with subset-conditioned inference. Here, the idea is to project each available omics layer into a common latent space where modalities become comparable and combinable. Because the encoding is learned jointly across patients with different patterns of availability, the model can capture the correlations that link, say, methylation patterns to transcriptomic states, and exploit those correlations when one view is missing. At inference time, the model conditions on the observed subset for a given patient and produces outcome predictions from that partial evidence. The approach borrows conceptually from multi-view learning and from variational frameworks in which each modality is treated as a partial observation of a single underlying biological state.</p>
<p>The third major category in the taxonomy is modality-completion: frameworks that attempt to synthesise the missing layer itself before integration proceeds. Generative models, including adversarial and autoencoder-based designs, learn the cross-modal relationships in the complete subset of the cohort and then produce plausible surrogates for missing omics profiles. Crucially, the review stresses that the goal is not to conjure the true molecular measurements of a patient who was never assayed, but to supply the downstream predictor with an estimate that preserves the predictive information the missing layer would have contributed. Done well, completion can recover much of the discriminative power lost to missingness; done poorly, it can inject hallucinated structure that inflates apparent accuracy without reflecting real biology.</p>
<p>Nguyen and Vafaee pay particular attention to cross-pollination from an unexpected corner of computational biology: single-cell research. Single-cell multi-omics experiments frequently produce mosaic datasets in which each cell is profiled for only a subset of modalities—RNA in one cell, chromatin accessibility in another—and an entire literature has arisen on integrating such fragmentary data. The review asks when the architectural tricks developed for that setting, such as modality dropout during training, cross-modality translation and shared embedding spaces, transfer to cohort-level supervised prediction of clinical outcomes. The authors conclude that the underlying mechanisms are often architecturally transferable, but that the statistical regimes differ: single-cell datasets contain thousands to millions of sparse observations, whereas clinical cohorts are typically small, heterogeneous and confounded by treatment and demographics, demanding greater caution and stronger regularisation.</p>
<p>Throughout, the review contrasts the design philosophies, inference mechanisms and robustness properties of competing approaches, and it makes clear that no single strategy dominates. Missingness-aware fusion tends to be the most conservative, never inventing data but sometimes sacrificing performance when a highly informative modality is absent. Latent-space methods offer flexibility and elegant handling of heterogeneous subsets but can be sensitive to how well the shared space is learned from limited samples. Completion frameworks can be the most powerful when cross-modal correlations are strong, yet they carry the greatest risk of propagating fabricated signal into clinical decisions. The right choice, the authors argue, depends on the missingness pattern itself—how it arises, whether it is informative, and which modalities it affects.</p>
<p>The practical stakes are considerable. As multi-omics assays move from research laboratories into routine oncology, immunology and rare-disease care, the models that guide treatment will inevitably be deployed on patients whose molecular work-ups are incomplete. A clinical prediction system that silently fails, or silently biases its estimates, whenever a modality is missing is not safe for bedside use. By mapping the assumptions each method makes about missingness—whether it treats absence as random, informative or structural—the review offers clinicians and bioinformaticians a principled framework for matching model class to data reality, and for recognising when a published benchmark built on artificially deleted data says little about performance in a genuinely incomplete cohort.</p>
<p>The work, which was supported by Australia&#8217;s CSIRO Next-Generation Graduate Program and the National Health and Medical Research Council, arrives as the field confronts a widening gap between the tidy datasets of methodological papers and the messy matrices of real hospitals. By clarifying when conventional integration breaks down, how missingness-aware designs hold together, and which lessons from single-cell genomics carry over to patient-level prediction, Nguyen and Vafaee provide a roadmap for building predictive models that meet clinical data as it actually exists: partial, uneven and imperfect, but still rich enough, if handled with the right mathematics, to improve the odds for the patients behind the numbers.</p>
<p>The review appears as an open-access publication, meaning its taxonomy and comparative analyses are freely available to researchers in low-resource settings who often face the very data limitations the paper addresses. The article was received in March 2026 and accepted in September 2026, placing it among the first comprehensive treatments of block-wise missingness in patient-level multi-omics prediction, a topic that has previously been scattered across methodological papers in machine learning venues and bioinformatics journals without a unifying framework.</p>
<p>One useful lens for understanding the review&#8217;s scope comes from its positioning within predictive medicine and biostatistics. Classical statistical approaches to incomplete data, such as likelihood-based methods that ignore missingness under certain assumptions, were developed for low-dimensional settings where a handful of covariates might be unobserved. Multi-omics data break these assumptions in two directions at once: the dimensionality is enormous, with tens of thousands of features per modality, and the missingness operates at the level of whole data layers rather than individual entries. This means that techniques from the missing-data literature in statistics cannot simply be imported wholesale, and the machine learning architectures surveyed in the review represent a genuinely new methodological territory rather than an incremental extension of older tools.</p>
<p>The authors&#8217; institutional context is also relevant to the review&#8217;s perspective. Nguyen and Vafaee are based at UNSW Sydney&#8217;s School of Biotechnology and Biomolecular Sciences, with Vafaee additionally affiliated with the UNSW AI Institute, a setting that bridges experimental molecular biology and artificial intelligence research. The work was funded through CSIRO&#8217;s Next-Generation Graduate Program and the National Health and Medical Research Council, reflecting Australian investment in translational computational health research.</p>
<p>Accompanying the article are supplementary data files in spreadsheet format, which likely catalogue the methods included in the taxonomy and their characteristics, offering readers a practical reference tool when selecting approaches for their own incomplete cohorts. The article&#8217;s keyword set, spanning multimodal machine learning, missing-modality learning, incomplete multi-omics data integration and clinical outcome prediction, signals its intended audience across both computer science and clinical informatics communities, and the authors declare no competing interests.</p>
<p><strong>Subject of Research:</strong> Machine learning methods for clinical outcome prediction from incomplete multi-omics datasets with missing modalities</p>
<p><strong>Article Title:</strong> Modelling missing modalities in multi-omics clinical outcome prediction</p>
<p><strong>Article References:</strong> Nguyen, R., &amp; Vafaee, F. (2026). Modelling missing modalities in multi-omics clinical outcome prediction. <em>Artificial Intelligence Review</em>. <a href="https://doi.org/10.1007/s10462-026-11702-7" rel="noopener noreferrer">https://doi.org/10.1007/s10462-026-11702-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10462-026-11702-7" rel="noopener noreferrer">10.1007/s10462-026-11702-7</a></p>
<p><strong>Keywords:</strong> multi-omics, missing modality learning, clinical outcome prediction, machine learning, precision medicine, data integration, imputation, latent representations, single-cell genomics, biomarker discovery, incomplete data, predictive medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193762</post-id>	</item>
	</channel>
</rss>
