Predicting how a patient’s tumor will respond to a particular drug remains one of the most stubborn problems in precision oncology. Two patients can carry the same driver mutation, receive the same targeted therapy, and experience dramatically different outcomes, a phenomenon rooted in the genomic and transcriptomic heterogeneity that defines cancer at the single-tumor and even single-cell level. A new computational method described in Genome Medicine aims to close that gap by borrowing a strategy from graph-based machine learning: instead of treating each tumor sample as an isolated data point, the model aggregates information from the sample and its most molecularly similar counterparts before making a prediction. The tool, called DrugSAGE, was developed by Peilin Jia of the China National Center for Bioinformation and the Beijing Institute of Genomics, Chinese Academy of Sciences, together with Zhongming Zhao of Vanderbilt University Medical Center, and it offers an interpretable route to imputing drug responses from gene expression data.
The core insight behind DrugSAGE is that no single tumor profile exists in a vacuum. Cell line repositories such as the Cancer Cell Line Encyclopedia (CCLE) and drug sensitivity resources such as the Genomics of Drug Sensitivity in Cancer (GDSC) database contain rich measurements of how hundreds of cancer cell lines respond to hundreds of compounds, paired with transcriptomic profiles of those lines. The challenge has always been transferring that knowledge to patient tumors, whose expression patterns differ systematically from cell lines grown in culture. DrugSAGE addresses the transfer problem by constructing a graph in which each sample, whether a cell line or a patient tumor, is connected to its nearest neighbors in transcriptomic space. A graph neural network then learns to aggregate features across these connections, so that a patient tumor’s predicted response to a drug is informed not only by its own gene expression signature but also by the responses of the cell lines that most closely resemble it.
Technically, DrugSAGE builds on the GraphSAGE architecture, a framework originally designed to generate embeddings for nodes in large graphs by sampling and aggregating features from local neighborhoods. In the drug response setting, the nodes are biological samples and the node features are genome-wide gene expression values, typically derived from RNA-sequencing data. The aggregation mechanism allows the model to smooth and enrich each sample’s representation with information from similar samples, which is particularly valuable when the target sample’s own data are noisy or incomplete. This is precisely the situation in clinical imputation tasks: a patient tumor has a measured transcriptome but no measured drug response, and the model must infer that response from the patterns learned across the cell line graph.
One of the most distinctive features of the method is its commitment to biological interpretability, an attribute often sacrificed in deep learning approaches. Rather than relying solely on opaque layers of artificial neurons, DrugSAGE incorporates a customized linear layer that encodes gene-pathway annotations directly into the model’s architecture. Genes are grouped according to known pathway memberships, and the model’s first stage of processing respects this biological organization. As a result, the learned weights can be traced back to specific pathways, giving researchers a window into why the model made a particular prediction. The authors complement this architectural interpretability with SHAP analysis, a game-theoretic technique for attributing a model’s output to its input features, which allows them to identify the key genes and pathways driving each drug response prediction.
The benchmarking strategy was deliberately broad. The authors evaluated DrugSAGE across independent bulk RNA-seq datasets and single-cell datasets, testing whether the model’s predictions were associated with known drug targets and with patient groups stratified by treatment outcome. In the bulk setting, the model was trained on cell line data from CCLE and GDSC and then applied to patient cohorts from The Cancer Genome Atlas (TCGA) and datasets deposited in the Gene Expression Omnibus (GEO), spanning cancer types that include breast cancer, lung adenocarcinoma, colon adenocarcinoma, head and neck squamous cell carcinoma, skin cutaneous melanoma, stomach adenocarcinoma, and thyroid carcinoma. The reported results showed significant associations between predicted responses and established pharmacology, suggesting that the model captures biologically meaningful signal rather than dataset-specific artifacts.
Perhaps the most forward-looking aspect of the study is the extension to single-cell drug response prediction. Bulk transcriptomics averages gene expression across millions of cells in a tumor, masking the subclonal diversity that increasingly appears to govern therapeutic resistance. Single-cell RNA sequencing resolves that diversity cell by cell, but drug response measurements at single-cell resolution are scarce and technically demanding to generate. DrugSAGE’s graph-based design makes it naturally suited to this sparse-data regime, because each single cell’s predicted response can be aggregated from the responses of transcriptionally similar cells and samples. The authors report that DrugSAGE effectively predicts single-cell drug responses, opening the possibility of dissecting how different subpopulations within the same tumor might respond differently to the same therapy.
The clinical validation extended to treatment-stratified patient groups, where the model’s predicted drug responses were compared against actual pathological outcomes. In breast cancer cohorts, for example, the framework was applied to distinguish patients who achieved pathological complete response from those with residual invasive disease after neoadjuvant treatment. The ability of predicted response scores to separate these groups provides evidence that the computational imputation reflects clinically relevant biology, not merely statistical correlation. Similar analyses across other cancer types reinforced the pattern: samples predicted to be sensitive to a drug were enriched among patients known to benefit from it, while predicted resistance aligned with poor outcomes such as shortened progression-free survival.
The interpretability layer yielded biologically coherent findings as well. When the authors interrogated which genes and pathways drove predictions for specific drugs, the model recovered known mechanisms of action and resistance. For targeted agents, key driver genes such as EGFR in lung cancers and VEGFR-related signaling in angiogenesis-directed therapies emerged as influential features, consistent with the established pharmacology. Pathway-level analysis highlighted processes such as epithelial-mesenchymal transition, a well-characterized program associated with invasion and drug resistance, as contributors to predicted response patterns. This convergence between model attribution and prior biological knowledge is a critical benchmark for any machine learning tool intended for clinical translation, because it suggests the model has learned mechanisms rather than memorizing confounds.
The broader context makes the contribution timely. Drug response prediction has been pursued with a range of computational strategies, including elastic net regression, deep neural networks, and variational autoencoder-based approaches such as VAEN and multi-omic synthetic augmentation methods. Each approach trades off accuracy, interpretability, and data requirements in different ways. DrugSAGE’s benchmarking showed superior or comparable performance relative to existing methods, but its distinguishing advantages are the graph-based aggregation of neighbor information and the pathway-informed architecture that makes predictions traceable. In a field where black-box models have struggled to earn clinical trust, a framework that can explain itself in the language of genes and pathways may prove more actionable for oncologists and translational researchers alike.
Limitations and next steps remain, as with any computational method trained primarily on cell line data. Cell lines capture only a fraction of tumor biology, and the fidelity of any imputation ultimately depends on how well the nearest-neighbor structure of the graph bridges the gap between culture dish and clinic. Nevertheless, the open-access publication of DrugSAGE, with its reliance on publicly available resources from GDSC, TCGA, and GEO, means that the broader research community can immediately test, extend, and refine the approach. As single-cell drug screening technologies mature and more treatment-stratified patient cohorts become available, graph-based transcriptome aggregation of the kind embodied in DrugSAGE could become a standard layer in the pipeline that translates tumor gene expression into concrete treatment decisions, moving precision oncology one step closer to predicting, rather than observing, each tumor’s response to therapy.
Subject of Research: Graph neural network-based prediction of cancer drug response from transcriptomic data
Article Title: DrugSAGE: a transcriptome aggregation approach using cell lines for drug response imputation
Article References: Jia, P., & Zhao, Z. (2026). DrugSAGE: a transcriptome aggregation approach using cell lines for drug response imputation. Genome Medicine. https://doi.org/10.1186/s13073-026-01781-0
Image Credits: AI Generated
DOI: 10.1186/s13073-026-01781-0
Keywords: DrugSAGE, graph neural network, drug response prediction, transcriptomics, precision oncology, cancer cell lines, GDSC, CCLE, single-cell RNA-seq, gene-pathway interpretability, pharmacogenomics, machine learning
Cite Scienmag News
Nathaniel Bowman. (September 30, 2026). AI Model Borrows From Lookalike Cells to Predict How Tumors Respond to Drugs. Scienmag. https://scienmag.com/ai-model-borrows-from-lookalike-cells-to-predict-how-tumors-respond-to-drugs/
Nathaniel Bowman. "AI Model Borrows From Lookalike Cells to Predict How Tumors Respond to Drugs." Scienmag, 30 September 2026, https://scienmag.com/ai-model-borrows-from-lookalike-cells-to-predict-how-tumors-respond-to-drugs/. Accessed 30 September 2026.
Nathaniel Bowman. "AI Model Borrows From Lookalike Cells to Predict How Tumors Respond to Drugs." Scienmag. September 30, 2026. https://scienmag.com/ai-model-borrows-from-lookalike-cells-to-predict-how-tumors-respond-to-drugs/

