Monday, October 5, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Biology

AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data

October 5, 2026
in Biology
Nathaniel Bowman
By Nathaniel Bowman Scienmag Editorial Profile - Precision Oncology
Reading Time: 5 mins read
0
AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data

AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Cancer is never a single disease, and even within one tumor type, molecular subtypes can behave in radically different ways, respond to different therapies, and carry different prognoses. Distinguishing those subtypes accurately from molecular data is one of the central tasks of modern precision oncology. Now, a team of researchers at Hunan City University in China has developed a new artificial intelligence framework designed to do exactly that, while confronting two stubborn problems that have long hampered computational approaches: the sheer complexity of integrating multiple layers of molecular information, and the fact that some cancer subtypes are far rarer than others, leaving machine learning models with too few examples to learn from. The framework, called CALT-GNN, is described in an open-access paper published in BMC Bioinformatics.

The core idea behind CALT-GNN, which stands for Cross-Attention and Long-Tail Expert Graph Neural Network, is to represent patients not as isolated rows of numbers but as nodes in a network. The method begins by constructing separate patient-similarity graphs for each omics layer, meaning that patients are linked to one another when their copy number alteration profiles, DNA methylation patterns, or messenger RNA expression signatures resemble each other. Graph convolutional networks then learn latent representations of each patient within these graphs, allowing information to flow between molecularly similar individuals. Finally, a technique known as similarity network fusion merges the separate graphs into a single unified network that captures relationships spanning all the molecular layers at once.

This graph-based strategy addresses the first major challenge in multi-omics cancer classification: high dimensionality. Each patient in a typical multi-omics dataset carries tens of thousands of molecular measurements, spanning genomic, epigenomic, transcriptomic, and sometimes proteomic levels. Feeding such high-dimensional vectors directly into a classifier invites overfitting and obscures the biological relationships between data types. By converting patients into nodes embedded in similarity graphs, the framework reduces the effective complexity of the problem and lets the learning algorithm exploit the structure of the data, namely the fact that patients with similar molecular profiles tend to belong to the same subtype.

The second innovation lies in how the model integrates different omics types with one another. Rather than simply concatenating measurements from copy number alteration, DNA methylation, and mRNA expression into one long vector, CALT-GNN employs a cross-attention branch that explicitly models the complementary relationships between modalities. In this design, copy number alteration and DNA methylation serve as source modalities, while mRNA expression acts as the target modality. Cross-attention, a mechanism borrowed from modern deep learning architectures, allows the model to learn which features in the source modalities are most informative for interpreting each feature in the target modality. Biologically, this mirrors real regulatory logic: DNA copy number changes and methylation patterns both influence gene expression, and the model can, in principle, learn those influences directly from the data.

The third component tackles the long-tail problem, which is arguably the most underappreciated obstacle in cancer subtype classification. Real cancer cohorts are almost never balanced. A common subtype may account for the majority of patients in a dataset, while rare but clinically important subtypes may be represented by only a handful of samples. Standard classifiers, optimized for overall accuracy, tend to become experts on the common subtypes and largely ignore the rare ones, which is precisely backwards from a clinical standpoint, since correctly identifying a rare subtype may be the most consequential decision for an individual patient. CALT-GNN addresses this with a long-tail expert branch built around two specialized components: a Major Expert and a Minor Expert, combined with prototype-guided routing that directs each sample toward the expert best suited to its position in the class distribution.

The routing mechanism works by comparing a patient’s learned representation against class prototypes, essentially reference points that summarize what each subtype looks like in the model’s internal feature space. Samples that resemble well-populated classes are handled by the Major Expert, which is tuned to the dense regions of the data, while samples from sparse, rare subtypes are routed to the Minor Expert, which specializes in the long tail of the distribution. This class-distribution-aware strategy means the model does not force a single classifier to serve both the abundant and the scarce subtypes simultaneously, a compromise that typically degrades performance on the rare end of the spectrum.

To combine the insights from the cross-attention branch and the long-tail expert branch, the framework uses a learnable global weight that determines how much each branch contributes to the final prediction. Rather than fixing this balance in advance, the model learns it during training, allowing the optimal mixture to differ across cancer types and datasets. The authors report that ablation studies, in which individual components were removed to test their contribution, supported the complementary roles of cross-omics interaction modeling and adaptive long-tail expert routing, although the magnitude of the improvements varied across cohorts and subtypes, a candid acknowledgment that no single mechanism dominates in every setting.

The evaluation was conducted on eight multi-omics cohorts drawn from The Cancer Genome Atlas, one of the largest and most widely used public resources in cancer genomics. The datasets were obtained in processed form from the MO-GCAN repository on Figshare, which derives from TCGA PanCancer Atlas data accessed through cBioPortal. Because the data are publicly available and de-identified, the study required no additional ethics approval. Across the eight cohorts, CALT-GNN achieved competitive or comparable classification performance relative to representative baseline methods, with particularly stable results on two metrics that are sensitive to class imbalance: Macro-F1, which averages the F1 score across classes so that rare subtypes count as much as common ones, and the Matthews Correlation Coefficient, which provides a balanced measure of classification quality even when class sizes differ sharply.

The choice of these two metrics is significant. A model can post an impressive overall accuracy simply by predicting the most common subtype for nearly every patient, while failing almost completely on rare ones. Macro-F1 and MCC expose such failures, and CALT-GNN’s stability on these measures suggests that its long-tail machinery is doing real work rather than merely inflating headline numbers. The authors also performed subtype-level analyses to examine how the model behaved on individual classes, providing a more granular picture than aggregate scores alone. For a field where the clinically hardest cases are often the rarest, this emphasis on imbalance-sensitive evaluation is a methodological point worth emphasizing.

None of this means that CALT-GNN is ready to guide treatment decisions in a clinic tomorrow. The study is a methodological contribution, demonstrating a framework and benchmarking it against existing approaches on public data, and the authors themselves note that improvements varied in magnitude across cohorts and subtypes. But the work illustrates a broader and important trend in computational oncology: the recognition that the hardest problems in cancer subtyping are not just about bigger models or more data, but about the structure of the data itself, including the tangled regulatory relationships among genomic layers and the skewed distributions of disease subtypes. By building those two realities directly into the architecture of a graph neural network, the Hunan City University team has offered a template that other researchers working on multi-omics integration, from rare disease classification to drug response prediction, may well find worth adapting. The paper is open access, and the underlying data are publicly available, lowering the barrier for the community to test, refine, and extend the approach.

Subject of Research: A graph neural network framework for multi-omics cancer subtype classification addressing cross-omics relationships and imbalanced subtype distributions

Article Title: CALT-GNN: a graph neural network with cross-attention and long-tail experts for multi-omics cancer subtype classification

Article References: Wang, K., Zheng, J., Zhao, L., Xiao, W., Li, Z., & He, Q. (2026). CALT-GNN: a graph neural network with cross-attention and long-tail experts for multi-omics cancer subtype classification. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06683-x

Image Credits: AI Generated

DOI: 10.1186/s12859-026-06683-x

Keywords: multi-omics integration, cancer subtype classification, graph neural network, cross-attention, long-tail learning, TCGA, copy number alteration, DNA methylation, mRNA expression, class imbalance, precision oncology, bioinformatics

Cite Scienmag News

Nathaniel Bowman. (October 5, 2026). AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data. Scienmag. https://scienmag.com/ai-model-tackles-rare-cancer-subtypes-by-learning-from-imbalanced-molecular-data/

Nathaniel Bowman. "AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data." Scienmag, 5 October 2026, https://scienmag.com/ai-model-tackles-rare-cancer-subtypes-by-learning-from-imbalanced-molecular-data/. Accessed 5 October 2026.

Nathaniel Bowman. "AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data." Scienmag. October 5, 2026. https://scienmag.com/ai-model-tackles-rare-cancer-subtypes-by-learning-from-imbalanced-molecular-data/

Tags: bioinformaticsCALT-GNN frameworkcancer subtype classificationCancer subtypesclass imbalancecomputational approaches for rare cancerscopy number alterationcross-attentioncross-attention in AI modelsDNA MethylationGraph neural networkgraph neural networks in bioinformaticsimbalanced data in cancer researchlong-tail learningmachine learning in cancermolecular data integrationmRNA expressionmulti-omics data analysismulti-omics integrationprecision oncologyrare cancer subtype detectionTCGAtumor heterogeneity and molecular profiling
Share26Tweet16
Previous Post

New Graph Alignment Method Matches Users Across Networks Without Any Labelled Data

Next Post

Two new US patents widen Ultrasound AI’s reach from pregnancy imaging to drug detection

Related Posts

Human GLUD2 Gene Variant Worsens Parkinson’s Disease by Igniting an Astrocyte–Microglia Complement Cascade
Biology

Human GLUD2 Gene Variant Worsens Parkinson’s Disease by Igniting an Astrocyte–Microglia Complement Cascade

October 5, 2026
Dental Students Say a Three-Phase Case Method Bridges the Preclinical-Clinical Divide
Biology

Dental Students Say a Three-Phase Case Method Bridges the Preclinical-Clinical Divide

October 5, 2026
Gut Microbes Turn Amino Acid Into Toxin That Worsens Pancreatitis in the Elderly
Biology

Gut Microbes Turn Amino Acid Into Toxin That Worsens Pancreatitis in the Elderly

October 5, 2026
Ribosomal RNA Passes Its Test as a Steady Reference in High-Throughput Plasma qPCR
Biology

Ribosomal RNA Passes Its Test as a Steady Reference in High-Throughput Plasma qPCR

October 5, 2026
GRETA: New Database Turns Mountains of Genome Sequencing Data into Searchable Science
Biology

GRETA: New Database Turns Mountains of Genome Sequencing Data into Searchable Science

October 5, 2026
Deep Learning Tool Reads RNA Tails Straight From Nanopore Signals
Biology

Deep Learning Tool Reads RNA Tails Straight From Nanopore Signals

October 5, 2026
Next Post
Two new US patents widen Ultrasound AI’s reach from pregnancy imaging to drug detection

Two new US patents widen Ultrasound AI's reach from pregnancy imaging to drug detection

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Two new US patents widen Ultrasound AI’s reach from pregnancy imaging to drug detection
  • AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data
  • New Graph Alignment Method Matches Users Across Networks Without Any Labelled Data
  • Worm-Compost and Bacteria Team Up to Cut Maize Fertilizer Use by a Quarter

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading