Friday, October 2, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Medicine

Massive Proteomics Dataset Powers a Virtual Cell Model for Drug Discovery

October 2, 2026
in Medicine, Technology and Engineering
Kenneth Gardner
By Kenneth Gardner Scienmag Editorial Profile - Proteomics
Reading Time: 5 mins read
0
Massive Proteomics Dataset Powers a Virtual Cell Model for Drug Discovery

Massive Proteomics Dataset Powers a Virtual Cell Model for Drug Discovery

Massive Proteomics Dataset Powers a Virtual Cell Model for Drug Discovery

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

A team of researchers in China has unveiled one of the most ambitious attempts yet to build a working virtual cell: an artificial intelligence system trained on tens of millions of protein measurements that can predict how cancer cells respond to drugs, suggest new drug combinations, and even anticipate how patients will fare on therapy. The model, called ProteinTalks, is described in a study published in Nature on 9 September 2026, and its creators argue that it marks a turning point for in silico drug discovery, the practice of screening and prioritizing candidate treatments entirely inside a computer before any experiment is run.

The central obstacle that the team set out to overcome is a familiar one in the emerging field of AI virtual cells. Most existing models are trained on gene expression data, particularly single-cell transcriptomics, which captures the messenger RNA molecules inside cells. But RNA abundance is an imperfect proxy for what actually happens in a cell, because proteins, not transcripts, carry out most biological functions and are the direct targets of nearly every drug. Large-scale, time-resolved proteomics data, which would allow a model to watch the protein machinery of a cell change hour by hour after a drug perturbation, simply did not exist at the scale needed for machine learning. The new study addresses that gap head-on by generating the data itself.

Using systematically perturbed breast cancer cell lines, the researchers generated more than 38 million temporal protein-abundance measurements, assembled into what they call a perturbation proteomics dataset. The experiments covered 63 FDA-approved drugs spanning a wide range of mechanisms of action, including alkylating agents, HDAC inhibitors, topoisomerase inhibitors, CDK inhibitors, hormonal agents and kinase inhibitors. Each cell line, drug and time point was measured with multiple biological replicates, and extensive quality control showed high reproducibility, with pooled samples, technical replicates and biological replicates all showing strong Pearson correlations and low coefficients of variation. Unsupervised visualization of the whole-proteome profiles revealed clean separation by cell line, treatment duration and drug class, with minimal batch effects from the mass spectrometry instruments, a critical prerequisite for training a model that learns biology rather than laboratory noise.

On top of this resource, the team built the ProteinTalks architecture, whose key innovation is a pretraining framework that learns transferable dynamical latent representations from temporal proteome trajectories. Rather than treating each protein measurement as a static snapshot, the model is trained to represent how the entire proteome evolves conditionally in response to different perturbations over time. The approach draws on ideas from dynamical systems modeling and neural ordinary differential equations, a line of thinking championed by applied mathematician Weinan E, who is a co-author and provided guidance on the modeling. By learning the dynamics of protein responses rather than just their endpoints, the model acquires representations that can be transferred to new drugs, new cell types and, ultimately, new patients.

The practical payoff is that ProteinTalks functions as what the authors call an operational tool, meaning it can be deployed for a diverse set of concrete drug discovery tasks rather than serving only as a scientific demonstration. In benchmarks under the evaluated protocols, the model predicted drug efficacy and drug synergy, discovered new drug combinations, identified proteins associated with drug resistance, stratified patient responses, and prioritized drug candidates for testing in patient-derived organoids. The authors report that it generally achieved higher performance than selected benchmark implementations, which included deep learning approaches for predicting anti-cancer drug synergy and other perturbation-response models. This matters because recent evaluations of single-cell foundation models have been sobering, with some studies finding that deep-learning-based gene perturbation predictions do not yet outperform simple linear baselines, a criticism the proteomics-based approach appears better positioned to answer.

Interpretability was a deliberate design goal rather than an afterthought. Using SHAP values, a technique that attributes a model’s output to its individual inputs, the researchers could trace which proteins drove predictions of drug efficacy for each drug class, and even examine dynamic SHAP values that reveal how one protein’s influence on another changes across the response trajectory. The team validated some of these mechanistic insights experimentally. For example, the model highlighted the protein TYMS in the response to the drug capecitabine, and the researchers confirmed with siRNA knockdown experiments that depleting TYMS in HCC1143 breast cancer cells enhanced sensitivity to the drug, exactly as the model’s perturbation scores implied. Another example involved AKR1C3 in antimitotic drug responses, again validated by knockdown. These experiments suggest the model is not merely a black-box predictor but can generate testable hypotheses about drug mechanism.

Perhaps the most consequential claim in the study is transferability beyond the cell lines on which the model was trained. ProteinTalks extended to patient-derived organoids, three-dimensional cultures grown from individual patients’ tumors, and to clinical biopsy samples, generally maintaining strong performance. The team also fine-tuned the model with transcriptomic data from patient-derived tumor xenografts, showing that pretraining on perturbation proteomics improved predictions of drug efficacy compared with a model trained from scratch on the same transcriptomic subset. In a clinical cohort of triple-negative breast cancer patients, the model’s predicted risk scores stratified patients treated with three- or four-drug combinations into groups with significantly different relapse-free and overall survival, as assessed by Kaplan-Meier analysis. For a field where most virtual cell models remain confined to cell line benchmarks, that bridge to clinical data is a notable step.

The study is also notable for its scale of collaboration and its openness. The work was led by Tiannan Guo of Westlake University together with Yi Zhu, Han Wen and Peijie Zhou, and involved researchers from DP Technology, Peking University, Westlake Omics, the AI for Science Institute in Beijing and Harbin Medical University Cancer Hospital, the latter contributing tumor samples and clinical information from the triple-negative breast cancer cohort. The raw mass spectrometry data have been deposited in the ProteomeXchange Consortium via iProX, and the protein matrix underlying the dataset is available for academic and non-commercial use through an open-access platform at db.prottalks.com, with analysis code released on GitHub. That openness could accelerate a race now clearly underway, as multiple groups worldwide pursue virtual cell models trained on transcriptomic, proteomic and multimodal data.

Challenges remain before such models change clinical practice. The training data, however vast, come from a limited set of breast cancer cell lines and a defined panel of approved drugs, and the authors themselves frame the model’s performance claims within the specific protocols they evaluated. Proteomics experiments remain more expensive and slower than transcriptomic ones, which may slow the community-wide accumulation of the kind of data this approach requires. Yet the study makes a forceful argument that dynamics-aware, proteomics-based pretraining is a viable path to operational virtual cells, and the demonstration that one model can move from cell lines to organoids to patient biopsies while remaining interpretable will likely set a benchmark for what the next generation of AI cell simulators must achieve.

Subject of Research: An AI virtual cell model trained on large-scale temporal perturbation proteomics data for in silico drug discovery in breast cancer

Article Title: An operational perturbation proteomics-based virtual cell model

Article References: Sun, R., Qian, L., Li, Y., Liu, T., Cheng, H., Zhang, X., Zhou, X., Zhan, Y., Zhang, G., Luo, Z., Ma, K., Wu, C., Ji, D., Xue, Z., Meng, H., Xiang, Y., Lei, D., Zhou, Q., Hu, W., … Guo, T. (2026). An operational perturbation proteomics-based virtual cell model. Nature. https://doi.org/10.1038/s41586-026-11001-9

Image Credits: AI Generated

DOI: 10.1038/s41586-026-11001-9

Keywords: virtual cell model, ProteinTalks, perturbation proteomics, drug discovery, breast cancer, machine learning, mass spectrometry, drug synergy, patient-derived organoids, transfer learning, SHAP interpretability, drug resistance

Cite Scienmag News

Kenneth Gardner. (October 2, 2026). Massive Proteomics Dataset Powers a Virtual Cell Model for Drug Discovery. Scienmag. https://scienmag.com/massive-proteomics-dataset-powers-a-virtual-cell-model-for-drug-discovery/

Kenneth Gardner. "Massive Proteomics Dataset Powers a Virtual Cell Model for Drug Discovery." Scienmag, 2 October 2026, https://scienmag.com/massive-proteomics-dataset-powers-a-virtual-cell-model-for-drug-discovery/. Accessed 2 October 2026.

Kenneth Gardner. "Massive Proteomics Dataset Powers a Virtual Cell Model for Drug Discovery." Scienmag. October 2, 2026. https://scienmag.com/massive-proteomics-dataset-powers-a-virtual-cell-model-for-drug-discovery/

Tags: AI in personalized medicineAI-driven drug discoverybreast cancercancer cell response predictiondrug combination suggestionsdrug discoverydrug resistancedrug synergyin silico drug screeninglarge-scale proteomics dataMachine learningmass spectrometrypatient therapy outcome predictionpatient-derived organoidsperturbation proteomicsprotein measurement analysisprotein-based cellular function modelingProteinTalksproteomics datasetSHAP interpretabilitytime-resolved proteomics analysistransfer learningvirtual cell modelvirtual cell modeling
Share26Tweet16
Previous Post

How Policy Documents Shape Student Evaluations in Nursing Education

Next Post

Surfing Protons Hit Record 132 MeV Thanks to Graphene and Long-Pulse Lasers

Related Posts

AI Learns to Hunt Radio Spectrum: War-Strategy Boost for Cognitive Networks
Technology and Engineering

AI Learns to Hunt Radio Spectrum: War-Strategy Boost for Cognitive Networks

October 2, 2026
Surfing Protons Hit Record 132 MeV Thanks to Graphene and Long-Pulse Lasers
Technology and Engineering

Surfing Protons Hit Record 132 MeV Thanks to Graphene and Long-Pulse Lasers

October 2, 2026
How Policy Documents Shape Student Evaluations in Nursing Education
Medicine

How Policy Documents Shape Student Evaluations in Nursing Education

October 2, 2026
AI Reads Breast MRI to Predict Gene-Based Cancer Risk Without a Biopsy
Medicine

AI Reads Breast MRI to Predict Gene-Based Cancer Risk Without a Biopsy

October 2, 2026
New AI Network Learns How Relationships Evolve to Predict Future Links
Technology and Engineering

New AI Network Learns How Relationships Evolve to Predict Future Links

October 2, 2026
New Hormone Drugs Transform Prostate Cancer Outcomes in Japan, Landmark Real-World Study Shows
Medicine

New Hormone Drugs Transform Prostate Cancer Outcomes in Japan, Landmark Real-World Study Shows

October 2, 2026
Next Post
Surfing Protons Hit Record 132 MeV Thanks to Graphene and Long-Pulse Lasers

Surfing Protons Hit Record 132 MeV Thanks to Graphene and Long-Pulse Lasers

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • AI Learns to Hunt Radio Spectrum: War-Strategy Boost for Cognitive Networks
  • Surfing Protons Hit Record 132 MeV Thanks to Graphene and Long-Pulse Lasers
  • Massive Proteomics Dataset Powers a Virtual Cell Model for Drug Discovery
  • How Policy Documents Shape Student Evaluations in Nursing Education

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading