Tuesday, September 22, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Chemistry

Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors

September 22, 2026
in Chemistry
Nathaniel Bowman
By Nathaniel Bowman Scienmag Editorial Profile - Precision Oncology
Reading Time: 6 mins read
0
Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors

Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors

Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Immunotherapy has transformed the treatment of many cancers, and few targets have proven as consequential as the interaction between the programmed cell death protein 1 receptor, known as PD-1, and its ligand PD-L1. When PD-L1 on tumor or immune cells engages PD-1 on effector T cells, the resulting signal dampens T-cell activation and allows tumors to slip past immune surveillance. Blocking this axis with monoclonal antibodies has produced durable responses in melanoma, non-small cell lung cancer, and renal cell carcinoma, reshaping standards of care. Yet antibody drugs carry practical drawbacks: they are expensive to manufacture, must be given by injection, and can trigger immune-related adverse events. These limitations have fueled a sustained push toward small-molecule alternatives that could be taken orally, titrated more easily, produced at scale, and potentially better tolerated during long-term therapy.

Small molecules face a formidable structural challenge, however. The PD-1/PD-L1 interface is a broad, shallow protein-protein contact surface, long considered a difficult target for drug-like compounds. Progress has come from chemotypes such as the biphenyl-based inhibitors developed by Bristol-Myers Squibb, including BMS-202, which bind PD-L1 and induce dimerization that blocks PD-1 recognition. Compounds such as INCB086550 and the oral modulator CA-170 have entered early clinical evaluation, proving that small-molecule checkpoint inhibition is feasible. Still, only a small fraction of explored scaffolds has advanced to the clinic, and screening the relevant chemical space experimentally is costly and slow. Computational triage has therefore become essential for deciding which candidates deserve synthesis and bioassay follow-up.

A team led by Monsin Sangsawat and Pornchai Rojsitthisak at Chulalongkorn University, with collaborators including Worathat Thitikornpong, Cong Feng, and Ming Chen, has now published an integrated solution in the journal Results in Chemistry. Their study reports an open-source, end-to-end cheminformatics framework for predicting the potency of small-molecule PD-L1 inhibitors, built entirely in the KNIME Analytics Platform. Every processing step, from data retrieval to model deployment, is encoded as an explicit node sequence, and the executable workflow file, curated datasets, and fixed split identifiers have been deposited in a public GitHub repository. The design directly confronts a persistent problem in the field: many published machine-learning models for PD-L1 are difficult to reproduce because training data, preprocessing details, or workflow code are inaccessible, and few are evaluated across independent data sources.

The foundation of the framework is a rigorously curated dataset. The researchers queried ChEMBL, targeting the human PD-L1 entry CHEMBL4523993, and retrieved more than 15,000 bioactivity records. To reduce label noise, they applied assay-level filters restricting the modeling set to homogeneous Alpha and HTRF assay formats, the two readouts with the most consistent annotation and the broadest activity coverage. After removing ambiguous qualifiers, duplicates, and non-numeric annotations, and after converting IC50 values to logarithmic pIC50 units, 2413 unique compounds remained. These were stratified into three potency tiers: active compounds with IC50 below 10 nanomolar, moderate inhibitors between 10 and 100 nanomolar, and weak or inactive compounds above 100 nanomolar. The curated set spans 953 unique Bemis-Murcko scaffolds, preserving the chemical diversity reported for PD-L1 inhibitors, from biphenyls and combretastatin analogs to cyclic peptides.

Molecular representation was treated as a central experimental variable. Each compound was encoded with 119 RDKit physicochemical descriptors, which were filtered for zero variance, pruned for high inter-correlation, and refined using recursive feature elimination with cross-validation, ultimately yielding 23 descriptors covering lipophilicity, ring content, van der Waals surface areas, and molecular quantum numbers. In parallel, two fingerprint families were generated: MACCS structural keys and Morgan circular fingerprints with a radius of two, together with a hybrid concatenation of the two. Four regression algorithms were benchmarked across the three representations, producing twelve configurations: Random Forest, a Keras deep neural network, Gradient Boosting, and XGBoost. Each configuration was trained under four fixed random seeds, and results were reported as means with standard deviations to capture stochastic sensitivity.

The Random Forest consistently delivered the most reliable balance between fit and generalization. Using the hybrid MACCS-Morgan representation, it achieved a coefficient of determination of 0.871 on the held-out test set and 0.863 on a strictly external validation set of 403 compounds that were never touched during model selection. Cross-validated Q2 values ranged from 0.888 to 0.902 across descriptor sets, and external Qext2 values reached as high as 0.919 with MACCS fingerprints alone. The deep neural network, although it attained higher training performance for some representations, showed a larger decline on held-out data, suggesting a greater tendency toward overfitting at the current sample size. Importantly, the team verified that no test or external compound shared an identical canonical SMILES with any training compound, and performance was unchanged after excluding high-similarity nearest neighbors, ruling out information leakage through duplicates.

Generalization was then stress-tested with a scaffold-disjoint partition, in which entire Bemis-Murcko scaffold families were assigned exclusively to training or held-out sets. Under this stricter regime, the hybrid Random Forest model achieved a mean test R2 of 0.858 and external validation R2 of 0.851, closely matching the random-split results. This indicates that the reported accuracy was not inflated by scaffold sharing and that the model genuinely extends to structural families absent from training. Prediction reliability was strongest in the moderate-to-potent activity range most relevant to virtual screening, while weak inhibitors with pIC50 below 7 showed larger errors and a modest tendency toward over-prediction, a pattern attributed to the relative scarcity of weak compounds in the curated data.

A distinctive strength of the study is its emphasis on knowing when to trust the model. Two applicability-domain diagnostics were implemented: a distance-based criterion derived from the mean and standard deviation of minimum training-set distances, and a leverage-based Williams plot with the classical threshold of three times the number of descriptors divided by the number of training compounds. Nearly all test and external-validation compounds fell inside the domain, and out-of-domain compounds in the independent BindingDB application set showed markedly larger errors than in-domain ones, with mean absolute errors of 1.884 versus 0.809 under the distance criterion. The researchers applied the model to 124 unique BindingDB compounds after removing 444 overlapping entries. Compounds sharing biphenyl scaffolds or cyclic peptide chemotypes with the training distribution showed small residuals, while structurally distinct pyrazolone derivatives and arylindanyloxypyridine frameworks showed substantially larger errors, pinpointing precisely where future data expansion would pay off.

The framework also translates statistical features into medicinal chemistry insight. SHAP analysis of the Random Forest identified the number of aromatic heterocycles, molar refractivity, and lipophilicity as the most influential descriptors, with threshold-like dependencies. The strongest potency-associated MACCS key simply records the presence of chlorine, but the corresponding Morgan features resolve it more sharply as a chlorinated aryl ring system and a chloro-substituted biaryl, discriminating potency with correlations up to 0.75. Conversely, a potency-decreasing CH2-O linkage characteristic of dibenzyl ether motifs is captured by MACCS but not the leading Morgan features, explaining why the two fingerprint blocks carry non-redundant, oppositely signed information. Scaffold analysis showed simple dibenzyl-ether and stilbene-like cores averaging pIC50 values of 6.3 to 7.3, while phenylpyridine and phenylpyrazine cores bearing basic amines averaged 8.4 to 10.2. These trends map directly onto the predominantly hydrophobic PD-L1 binding cleft, defined by residues such as Tyr56, Met115, and Tyr123, where halogenated aromatic extension and nitrogen heterocycles favor productive binding, and outlier compounds flagged by the model were examined through GNINA docking at this cleft against the reference inhibitor BMS-202.

The authors are candid about limits. Assay-format annotations come from ChEMBL descriptions rather than independently verified conditions, which may bias the set toward chemotypes preferentially tested in Alpha and HTRF formats, and a comparison against the published PDL1-inhi.predictor showed that an external model captured the BindingDB chemical space somewhat better. Yet the study’s contribution extends beyond any single metric: it delivers a transparent, inspectable, and adaptable platform that couples potency prediction with uncertainty quantification, supports pre-docking prioritization, and can be extended with uncertainty-guided active learning. The team suggests that high-ranking compounds within the applicability domain, especially those representing novel chemotypes, could be selected for prospective testing in the same Alpha or HTRF assays, followed by model retraining. For a field where reproducibility is often the missing ingredient, an open workflow anyone can audit and rerun on an ordinary laptop may prove as valuable as the model itself.

Subject of Research: Machine learning-guided discovery and prioritization of small-molecule PD-L1 inhibitors using an open-source KNIME workflow with curated bioactivity data and applicability-domain assessment

Article Title: Machine-learning-guided discovery and prioritization of PD-L1 inhibitors: An open-source KNIME workflow with curated bioactivity data, molecular fingerprints, and applicability-domain assessment

Article References: Sangsawat, M., Watpuang, B., Nathapinthu, T., Srisiriroj, P., Nalinratana, N., Taechawattananant, P., Thitikornpong, W., Feng, C., Chen, M., & Rojsitthisak, P. (2026). Machine-learning-guided discovery and prioritization of PD-L1 inhibitors: An open-source KNIME workflow with curated bioactivity data, molecular fingerprints, and applicability-domain assessment. Results in Chemistry, 30, Article 103877. https://doi.org/10.1016/j.rechem.2026.103877

Image Credits: AI Generated

DOI: 10.1016/j.rechem.2026.103877

Keywords: PD-L1, PD-1/PD-L1 immune checkpoint, machine learning, KNIME workflow, Random Forest, molecular fingerprints, MACCS, Morgan fingerprints, applicability domain, ChEMBL, BindingDB, cancer immunotherapy

Cite Scienmag News

Nathaniel Bowman. (September 22, 2026). Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors. Scienmag. https://scienmag.com/open-source-machine-learning-workflow-accelerates-discovery-of-small-molecule-pd-l1-cancer-inhibitors/

Nathaniel Bowman. "Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors." Scienmag, 22 September 2026, https://scienmag.com/open-source-machine-learning-workflow-accelerates-discovery-of-small-molecule-pd-l1-cancer-inhibitors/. Accessed 22 September 2026.

Nathaniel Bowman. "Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors." Scienmag. September 22, 2026. https://scienmag.com/open-source-machine-learning-workflow-accelerates-discovery-of-small-molecule-pd-l1-cancer-inhibitors/

Tags: applicability domainBindingDBcancer immunotherapyCancer Immunotherapy Targeting PD-1/PD-L1 AxisChallenges in Targeting Protein-Protein InteractionsChEMBLChemotype-Based PD-L1 InhibitorsClinical Progress of PD-LDeep Learning Approaches in Cancer Drug DesignKNIME workflowMACCSMachine learningMachine Learning Workflow in Drug Discoverymolecular fingerprintsMorgan fingerprintsOpen-Source Machine Learning for Small-Molecule PD-L1 Inhibitor DiscoveryOral Small-Molecule Immune Checkpoint BlockersPD-1/PD-L1 immune checkpointPD-L1Random ForestSmall-Molecule Cancer Immunotherapy DevelopmentStructural Challenges of PD-L1 Targeting
Share26Tweet16
Previous Post

Urban Trees Ease Negative Emotions on Social Media, Especially During COVID-19

Next Post

Real-World Liraglutide Study Shows Meaningful Weight Loss in Taiwanese Adults at Sub-Maximal Doses

Related Posts

Coal Experiment Reveals Carbon Dioxide Diffusion Can Rise and Then Fall With Pressure
Chemistry

Coal Experiment Reveals Carbon Dioxide Diffusion Can Rise and Then Fall With Pressure

September 22, 2026
Glass Waste Turns Into Buoyant Ceramic Beads That Could End Styrofoam Pollution at Sea
Chemistry

Glass Waste Turns Into Buoyant Ceramic Beads That Could End Styrofoam Pollution at Sea

September 22, 2026
Friction Stir Processing Repairs Fusion Welded Aluminium Alloy Joints
Chemistry

Friction Stir Processing Repairs Fusion Welded Aluminium Alloy Joints

September 22, 2026
Machine Learning Joins Forces With Quantum Physics to Accelerate Clean Energy Materials Discovery
Chemistry

Machine Learning Joins Forces With Quantum Physics to Accelerate Clean Energy Materials Discovery

September 22, 2026
Chemists Thread Eight Molecular Helices Into Single Nanographenes
Chemistry

Chemists Thread Eight Molecular Helices Into Single Nanographenes

September 22, 2026
Gamma-Grafted Polymer Separates Zirconium and Yttrium With Impressive Selectivity
Chemistry

Gamma-Grafted Polymer Separates Zirconium and Yttrium With Impressive Selectivity

September 22, 2026
Next Post
Real-World Liraglutide Study Shows Meaningful Weight Loss in Taiwanese Adults at Sub-Maximal Doses

Real-World Liraglutide Study Shows Meaningful Weight Loss in Taiwanese Adults at Sub-Maximal Doses

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Probiotics, Zinc and Copper Shield Broilers From Clostridium perfringens Damage
  • Danish Therapy Questionnaires Reveal Gaps Between What Clients Prefer and Therapists Deliver
  • Real-World Liraglutide Study Shows Meaningful Weight Loss in Taiwanese Adults at Sub-Maximal Doses
  • Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading