<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>KNIME workflow &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/knime-workflow/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 22 Sep 2026 15:24:01 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>KNIME workflow &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Open-Source Machine Learning Workflow Accelerates Discovery of Small-Molecule PD-L1 Cancer Inhibitors</title>
		<link>https://scienmag.com/open-source-machine-learning-workflow-accelerates-discovery-of-small-molecule-pd-l1-cancer-inhibitors/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 15:24:01 +0000</pubDate>
				<category><![CDATA[Chemistry]]></category>
		<category><![CDATA[applicability domain]]></category>
		<category><![CDATA[BindingDB]]></category>
		<category><![CDATA[cancer immunotherapy]]></category>
		<category><![CDATA[Cancer Immunotherapy Targeting PD-1/PD-L1 Axis]]></category>
		<category><![CDATA[Challenges in Targeting Protein-Protein Interactions]]></category>
		<category><![CDATA[ChEMBL]]></category>
		<category><![CDATA[Chemotype-Based PD-L1 Inhibitors]]></category>
		<category><![CDATA[Clinical Progress of PD-L]]></category>
		<category><![CDATA[Deep Learning Approaches in Cancer Drug Design]]></category>
		<category><![CDATA[KNIME workflow]]></category>
		<category><![CDATA[MACCS]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Machine Learning Workflow in Drug Discovery]]></category>
		<category><![CDATA[molecular fingerprints]]></category>
		<category><![CDATA[Morgan fingerprints]]></category>
		<category><![CDATA[Open-Source Machine Learning for Small-Molecule PD-L1 Inhibitor Discovery]]></category>
		<category><![CDATA[Oral Small-Molecule Immune Checkpoint Blockers]]></category>
		<category><![CDATA[PD-1/PD-L1 immune checkpoint]]></category>
		<category><![CDATA[PD-L1]]></category>
		<category><![CDATA[Random Forest]]></category>
		<category><![CDATA[Small-Molecule Cancer Immunotherapy Development]]></category>
		<category><![CDATA[Structural Challenges of PD-L1 Targeting]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=206351</guid>

					<description><![CDATA[Researchers have built a fully open-source KNIME machine-learning workflow that curates thousands of PD-L1 bioactivity records and reliably predicts inhibitor potency while flagging when predictions can be trusted.]]></description>
										<content:encoded><![CDATA[<p>Immunotherapy has transformed the treatment of many cancers, and few targets have proven as consequential as the interaction between the programmed cell death protein 1 receptor, known as PD-1, and its ligand PD-L1. When PD-L1 on tumor or immune cells engages PD-1 on effector T cells, the resulting signal dampens T-cell activation and allows tumors to slip past immune surveillance. Blocking this axis with monoclonal antibodies has produced durable responses in melanoma, non-small cell lung cancer, and renal cell carcinoma, reshaping standards of care. Yet antibody drugs carry practical drawbacks: they are expensive to manufacture, must be given by injection, and can trigger immune-related adverse events. These limitations have fueled a sustained push toward small-molecule alternatives that could be taken orally, titrated more easily, produced at scale, and potentially better tolerated during long-term therapy.</p>
<p>Small molecules face a formidable structural challenge, however. The PD-1/PD-L1 interface is a broad, shallow protein-protein contact surface, long considered a difficult target for drug-like compounds. Progress has come from chemotypes such as the biphenyl-based inhibitors developed by Bristol-Myers Squibb, including BMS-202, which bind PD-L1 and induce dimerization that blocks PD-1 recognition. Compounds such as INCB086550 and the oral modulator CA-170 have entered early clinical evaluation, proving that small-molecule checkpoint inhibition is feasible. Still, only a small fraction of explored scaffolds has advanced to the clinic, and screening the relevant chemical space experimentally is costly and slow. Computational triage has therefore become essential for deciding which candidates deserve synthesis and bioassay follow-up.</p>
<p>A team led by Monsin Sangsawat and Pornchai Rojsitthisak at Chulalongkorn University, with collaborators including Worathat Thitikornpong, Cong Feng, and Ming Chen, has now published an integrated solution in the journal Results in Chemistry. Their study reports an open-source, end-to-end cheminformatics framework for predicting the potency of small-molecule PD-L1 inhibitors, built entirely in the KNIME Analytics Platform. Every processing step, from data retrieval to model deployment, is encoded as an explicit node sequence, and the executable workflow file, curated datasets, and fixed split identifiers have been deposited in a public GitHub repository. The design directly confronts a persistent problem in the field: many published machine-learning models for PD-L1 are difficult to reproduce because training data, preprocessing details, or workflow code are inaccessible, and few are evaluated across independent data sources.</p>
<p>The foundation of the framework is a rigorously curated dataset. The researchers queried ChEMBL, targeting the human PD-L1 entry CHEMBL4523993, and retrieved more than 15,000 bioactivity records. To reduce label noise, they applied assay-level filters restricting the modeling set to homogeneous Alpha and HTRF assay formats, the two readouts with the most consistent annotation and the broadest activity coverage. After removing ambiguous qualifiers, duplicates, and non-numeric annotations, and after converting IC50 values to logarithmic pIC50 units, 2413 unique compounds remained. These were stratified into three potency tiers: active compounds with IC50 below 10 nanomolar, moderate inhibitors between 10 and 100 nanomolar, and weak or inactive compounds above 100 nanomolar. The curated set spans 953 unique Bemis-Murcko scaffolds, preserving the chemical diversity reported for PD-L1 inhibitors, from biphenyls and combretastatin analogs to cyclic peptides.</p>
<p>Molecular representation was treated as a central experimental variable. Each compound was encoded with 119 RDKit physicochemical descriptors, which were filtered for zero variance, pruned for high inter-correlation, and refined using recursive feature elimination with cross-validation, ultimately yielding 23 descriptors covering lipophilicity, ring content, van der Waals surface areas, and molecular quantum numbers. In parallel, two fingerprint families were generated: MACCS structural keys and Morgan circular fingerprints with a radius of two, together with a hybrid concatenation of the two. Four regression algorithms were benchmarked across the three representations, producing twelve configurations: Random Forest, a Keras deep neural network, Gradient Boosting, and XGBoost. Each configuration was trained under four fixed random seeds, and results were reported as means with standard deviations to capture stochastic sensitivity.</p>
<p>The Random Forest consistently delivered the most reliable balance between fit and generalization. Using the hybrid MACCS-Morgan representation, it achieved a coefficient of determination of 0.871 on the held-out test set and 0.863 on a strictly external validation set of 403 compounds that were never touched during model selection. Cross-validated Q2 values ranged from 0.888 to 0.902 across descriptor sets, and external Qext2 values reached as high as 0.919 with MACCS fingerprints alone. The deep neural network, although it attained higher training performance for some representations, showed a larger decline on held-out data, suggesting a greater tendency toward overfitting at the current sample size. Importantly, the team verified that no test or external compound shared an identical canonical SMILES with any training compound, and performance was unchanged after excluding high-similarity nearest neighbors, ruling out information leakage through duplicates.</p>
<p>Generalization was then stress-tested with a scaffold-disjoint partition, in which entire Bemis-Murcko scaffold families were assigned exclusively to training or held-out sets. Under this stricter regime, the hybrid Random Forest model achieved a mean test R2 of 0.858 and external validation R2 of 0.851, closely matching the random-split results. This indicates that the reported accuracy was not inflated by scaffold sharing and that the model genuinely extends to structural families absent from training. Prediction reliability was strongest in the moderate-to-potent activity range most relevant to virtual screening, while weak inhibitors with pIC50 below 7 showed larger errors and a modest tendency toward over-prediction, a pattern attributed to the relative scarcity of weak compounds in the curated data.</p>
<p>A distinctive strength of the study is its emphasis on knowing when to trust the model. Two applicability-domain diagnostics were implemented: a distance-based criterion derived from the mean and standard deviation of minimum training-set distances, and a leverage-based Williams plot with the classical threshold of three times the number of descriptors divided by the number of training compounds. Nearly all test and external-validation compounds fell inside the domain, and out-of-domain compounds in the independent BindingDB application set showed markedly larger errors than in-domain ones, with mean absolute errors of 1.884 versus 0.809 under the distance criterion. The researchers applied the model to 124 unique BindingDB compounds after removing 444 overlapping entries. Compounds sharing biphenyl scaffolds or cyclic peptide chemotypes with the training distribution showed small residuals, while structurally distinct pyrazolone derivatives and arylindanyloxypyridine frameworks showed substantially larger errors, pinpointing precisely where future data expansion would pay off.</p>
<p>The framework also translates statistical features into medicinal chemistry insight. SHAP analysis of the Random Forest identified the number of aromatic heterocycles, molar refractivity, and lipophilicity as the most influential descriptors, with threshold-like dependencies. The strongest potency-associated MACCS key simply records the presence of chlorine, but the corresponding Morgan features resolve it more sharply as a chlorinated aryl ring system and a chloro-substituted biaryl, discriminating potency with correlations up to 0.75. Conversely, a potency-decreasing CH2-O linkage characteristic of dibenzyl ether motifs is captured by MACCS but not the leading Morgan features, explaining why the two fingerprint blocks carry non-redundant, oppositely signed information. Scaffold analysis showed simple dibenzyl-ether and stilbene-like cores averaging pIC50 values of 6.3 to 7.3, while phenylpyridine and phenylpyrazine cores bearing basic amines averaged 8.4 to 10.2. These trends map directly onto the predominantly hydrophobic PD-L1 binding cleft, defined by residues such as Tyr56, Met115, and Tyr123, where halogenated aromatic extension and nitrogen heterocycles favor productive binding, and outlier compounds flagged by the model were examined through GNINA docking at this cleft against the reference inhibitor BMS-202.</p>
<p>The authors are candid about limits. Assay-format annotations come from ChEMBL descriptions rather than independently verified conditions, which may bias the set toward chemotypes preferentially tested in Alpha and HTRF formats, and a comparison against the published PDL1-inhi.predictor showed that an external model captured the BindingDB chemical space somewhat better. Yet the study&#8217;s contribution extends beyond any single metric: it delivers a transparent, inspectable, and adaptable platform that couples potency prediction with uncertainty quantification, supports pre-docking prioritization, and can be extended with uncertainty-guided active learning. The team suggests that high-ranking compounds within the applicability domain, especially those representing novel chemotypes, could be selected for prospective testing in the same Alpha or HTRF assays, followed by model retraining. For a field where reproducibility is often the missing ingredient, an open workflow anyone can audit and rerun on an ordinary laptop may prove as valuable as the model itself.</p>
<p><strong>Subject of Research:</strong> Machine learning-guided discovery and prioritization of small-molecule PD-L1 inhibitors using an open-source KNIME workflow with curated bioactivity data and applicability-domain assessment</p>
<p><strong>Article Title:</strong> Machine-learning-guided discovery and prioritization of PD-L1 inhibitors: An open-source KNIME workflow with curated bioactivity data, molecular fingerprints, and applicability-domain assessment</p>
<p><strong>Article References:</strong> Sangsawat, M., Watpuang, B., Nathapinthu, T., Srisiriroj, P., Nalinratana, N., Taechawattananant, P., Thitikornpong, W., Feng, C., Chen, M., &amp; Rojsitthisak, P. (2026). Machine-learning-guided discovery and prioritization of PD-L1 inhibitors: An open-source KNIME workflow with curated bioactivity data, molecular fingerprints, and applicability-domain assessment. <em>Results in Chemistry, 30</em>, Article 103877. <a href="https://doi.org/10.1016/j.rechem.2026.103877" rel="noopener noreferrer">https://doi.org/10.1016/j.rechem.2026.103877</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.rechem.2026.103877" rel="noopener noreferrer">10.1016/j.rechem.2026.103877</a></p>
<p><strong>Keywords:</strong> PD-L1, PD-1/PD-L1 immune checkpoint, machine learning, KNIME workflow, Random Forest, molecular fingerprints, MACCS, Morgan fingerprints, applicability domain, ChEMBL, BindingDB, cancer immunotherapy</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">206351</post-id>	</item>
	</channel>
</rss>
