Cancer is never a single disease, and even within one tumor type, molecular subtypes can behave in radically different ways, respond to different therapies, and carry different prognoses. Distinguishing those subtypes accurately from molecular data is one of the central tasks of modern precision oncology. Now, a team of researchers at Hunan City University in China has developed a new artificial intelligence framework designed to do exactly that, while confronting two stubborn problems that have long hampered computational approaches: the sheer complexity of integrating multiple layers of molecular information, and the fact that some cancer subtypes are far rarer than others, leaving machine learning models with too few examples to learn from. The framework, called CALT-GNN, is described in an open-access paper published in BMC Bioinformatics.
The core idea behind CALT-GNN, which stands for Cross-Attention and Long-Tail Expert Graph Neural Network, is to represent patients not as isolated rows of numbers but as nodes in a network. The method begins by constructing separate patient-similarity graphs for each omics layer, meaning that patients are linked to one another when their copy number alteration profiles, DNA methylation patterns, or messenger RNA expression signatures resemble each other. Graph convolutional networks then learn latent representations of each patient within these graphs, allowing information to flow between molecularly similar individuals. Finally, a technique known as similarity network fusion merges the separate graphs into a single unified network that captures relationships spanning all the molecular layers at once.
This graph-based strategy addresses the first major challenge in multi-omics cancer classification: high dimensionality. Each patient in a typical multi-omics dataset carries tens of thousands of molecular measurements, spanning genomic, epigenomic, transcriptomic, and sometimes proteomic levels. Feeding such high-dimensional vectors directly into a classifier invites overfitting and obscures the biological relationships between data types. By converting patients into nodes embedded in similarity graphs, the framework reduces the effective complexity of the problem and lets the learning algorithm exploit the structure of the data, namely the fact that patients with similar molecular profiles tend to belong to the same subtype.
The second innovation lies in how the model integrates different omics types with one another. Rather than simply concatenating measurements from copy number alteration, DNA methylation, and mRNA expression into one long vector, CALT-GNN employs a cross-attention branch that explicitly models the complementary relationships between modalities. In this design, copy number alteration and DNA methylation serve as source modalities, while mRNA expression acts as the target modality. Cross-attention, a mechanism borrowed from modern deep learning architectures, allows the model to learn which features in the source modalities are most informative for interpreting each feature in the target modality. Biologically, this mirrors real regulatory logic: DNA copy number changes and methylation patterns both influence gene expression, and the model can, in principle, learn those influences directly from the data.
The third component tackles the long-tail problem, which is arguably the most underappreciated obstacle in cancer subtype classification. Real cancer cohorts are almost never balanced. A common subtype may account for the majority of patients in a dataset, while rare but clinically important subtypes may be represented by only a handful of samples. Standard classifiers, optimized for overall accuracy, tend to become experts on the common subtypes and largely ignore the rare ones, which is precisely backwards from a clinical standpoint, since correctly identifying a rare subtype may be the most consequential decision for an individual patient. CALT-GNN addresses this with a long-tail expert branch built around two specialized components: a Major Expert and a Minor Expert, combined with prototype-guided routing that directs each sample toward the expert best suited to its position in the class distribution.
The routing mechanism works by comparing a patient’s learned representation against class prototypes, essentially reference points that summarize what each subtype looks like in the model’s internal feature space. Samples that resemble well-populated classes are handled by the Major Expert, which is tuned to the dense regions of the data, while samples from sparse, rare subtypes are routed to the Minor Expert, which specializes in the long tail of the distribution. This class-distribution-aware strategy means the model does not force a single classifier to serve both the abundant and the scarce subtypes simultaneously, a compromise that typically degrades performance on the rare end of the spectrum.
To combine the insights from the cross-attention branch and the long-tail expert branch, the framework uses a learnable global weight that determines how much each branch contributes to the final prediction. Rather than fixing this balance in advance, the model learns it during training, allowing the optimal mixture to differ across cancer types and datasets. The authors report that ablation studies, in which individual components were removed to test their contribution, supported the complementary roles of cross-omics interaction modeling and adaptive long-tail expert routing, although the magnitude of the improvements varied across cohorts and subtypes, a candid acknowledgment that no single mechanism dominates in every setting.
The evaluation was conducted on eight multi-omics cohorts drawn from The Cancer Genome Atlas, one of the largest and most widely used public resources in cancer genomics. The datasets were obtained in processed form from the MO-GCAN repository on Figshare, which derives from TCGA PanCancer Atlas data accessed through cBioPortal. Because the data are publicly available and de-identified, the study required no additional ethics approval. Across the eight cohorts, CALT-GNN achieved competitive or comparable classification performance relative to representative baseline methods, with particularly stable results on two metrics that are sensitive to class imbalance: Macro-F1, which averages the F1 score across classes so that rare subtypes count as much as common ones, and the Matthews Correlation Coefficient, which provides a balanced measure of classification quality even when class sizes differ sharply.
The choice of these two metrics is significant. A model can post an impressive overall accuracy simply by predicting the most common subtype for nearly every patient, while failing almost completely on rare ones. Macro-F1 and MCC expose such failures, and CALT-GNN’s stability on these measures suggests that its long-tail machinery is doing real work rather than merely inflating headline numbers. The authors also performed subtype-level analyses to examine how the model behaved on individual classes, providing a more granular picture than aggregate scores alone. For a field where the clinically hardest cases are often the rarest, this emphasis on imbalance-sensitive evaluation is a methodological point worth emphasizing.
None of this means that CALT-GNN is ready to guide treatment decisions in a clinic tomorrow. The study is a methodological contribution, demonstrating a framework and benchmarking it against existing approaches on public data, and the authors themselves note that improvements varied in magnitude across cohorts and subtypes. But the work illustrates a broader and important trend in computational oncology: the recognition that the hardest problems in cancer subtyping are not just about bigger models or more data, but about the structure of the data itself, including the tangled regulatory relationships among genomic layers and the skewed distributions of disease subtypes. By building those two realities directly into the architecture of a graph neural network, the Hunan City University team has offered a template that other researchers working on multi-omics integration, from rare disease classification to drug response prediction, may well find worth adapting. The paper is open access, and the underlying data are publicly available, lowering the barrier for the community to test, refine, and extend the approach.
Subject of Research: A graph neural network framework for multi-omics cancer subtype classification addressing cross-omics relationships and imbalanced subtype distributions
Article Title: CALT-GNN: a graph neural network with cross-attention and long-tail experts for multi-omics cancer subtype classification
Article References: Wang, K., Zheng, J., Zhao, L., Xiao, W., Li, Z., & He, Q. (2026). CALT-GNN: a graph neural network with cross-attention and long-tail experts for multi-omics cancer subtype classification. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06683-x
Image Credits: AI Generated
DOI: 10.1186/s12859-026-06683-x
Keywords: multi-omics integration, cancer subtype classification, graph neural network, cross-attention, long-tail learning, TCGA, copy number alteration, DNA methylation, mRNA expression, class imbalance, precision oncology, bioinformatics
Cite Scienmag News
Nathaniel Bowman. (October 5, 2026). AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data. Scienmag. https://scienmag.com/ai-model-tackles-rare-cancer-subtypes-by-learning-from-imbalanced-molecular-data/
Nathaniel Bowman. "AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data." Scienmag, 5 October 2026, https://scienmag.com/ai-model-tackles-rare-cancer-subtypes-by-learning-from-imbalanced-molecular-data/. Accessed 5 October 2026.
Nathaniel Bowman. "AI Model Tackles Rare Cancer Subtypes by Learning From Imbalanced Molecular Data." Scienmag. October 5, 2026. https://scienmag.com/ai-model-tackles-rare-cancer-subtypes-by-learning-from-imbalanced-molecular-data/

