<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>data leakage in drug synergy models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/data-leakage-in-drug-synergy-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 08:38:37 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>data leakage in drug synergy models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Machine Learning Predicts Antibacterial Drug Synergy Without Data Leakage</title>
		<link>https://scienmag.com/machine-learning-predicts-antibacterial-drug-synergy-without-data-leakage/</link>
		
		<dc:creator><![CDATA[Teresa Odom]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 08:38:37 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[antibacterial compounds]]></category>
		<category><![CDATA[antibacterial drug synergy prediction]]></category>
		<category><![CDATA[Antibiotic resistance]]></category>
		<category><![CDATA[avoiding data leakage in machine learning]]></category>
		<category><![CDATA[bioactivity signatures]]></category>
		<category><![CDATA[challenges in antibacterial drug discovery]]></category>
		<category><![CDATA[Chemical Checker]]></category>
		<category><![CDATA[clinical applications of drug synergy prediction]]></category>
		<category><![CDATA[combination therapy]]></category>
		<category><![CDATA[combination therapy optimization]]></category>
		<category><![CDATA[computational modeling of antibiotic interactions]]></category>
		<category><![CDATA[cross-validation]]></category>
		<category><![CDATA[data leakage]]></category>
		<category><![CDATA[data leakage in drug synergy models]]></category>
		<category><![CDATA[drug synergy]]></category>
		<category><![CDATA[evaluation methods for synergy prediction models]]></category>
		<category><![CDATA[HALO]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in antimicrobial research]]></category>
		<category><![CDATA[PLOS Computational Biology]]></category>
		<category><![CDATA[PLOS Computational Biology antibiotic research]]></category>
		<category><![CDATA[predictive modeling]]></category>
		<category><![CDATA[predictive modeling for antibiotic resistance]]></category>
		<category><![CDATA[small datasets in machine learning for medicine]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=261722</guid>

					<description><![CDATA[Researchers developed HALO, a machine-learning framework that predicts antibacterial drug synergy from multi-level bioactivity features while exposing how data leakage has inflated performance in earlier models.]]></description>
										<content:encoded><![CDATA[<p>Antibiotic resistance is one of the defining medical challenges of our time, and one of the most promising responses to it is deceptively simple: combine existing drugs so that they attack bacteria in ways a single compound cannot. Combination therapies have repeatedly shown improved efficacy across diverse clinical settings, from tuberculosis to difficult hospital-acquired infections. Yet finding pairs of antibacterial compounds that genuinely reinforce each other has remained a slow, expensive, and largely empirical process. The search space is enormous, the biology is messy, and the experimental data available to train computational models is small by the standards of modern machine learning. A new study published in PLOS Computational Biology by Hannie Yousefabadi and Mahya Mehrmohamadi confronts this problem directly, and in doing so exposes an uncomfortable truth about how synergy-prediction models have been evaluated in the past.</p>
<p>The core difficulty is not merely predicting synergy; it is predicting it honestly. The researchers point out that many existing machine-learning approaches for antibacterial synergy prediction rely on permissive cross-validation schemes in which the same drug pair can appear in both the training and the test portions of the data. When that happens, the model is not really learning general rules about which combinations work. Instead, it is partially memorizing specific pairs it has already seen, and the performance metrics that result are inflated. In a field where the ultimate goal is to prioritize untested combinations for laboratory validation, such leakage can turn a modest predictor into an apparently spectacular one, and can send experimentalists chasing combinations that a properly evaluated model would never have flagged.</p>
<p>To build a foundation for rigorous modeling, the team first assembled a curated dataset of 3,160 drug–pair–strain interactions, spanning 97 antibacterial compounds tested against 10 bacterial strains. That scale is meaningful: it captures variability both across drugs and across bacterial strains, which is precisely where earlier efforts have struggled. Strain variability means that a combination synergistic in one genetic background may be neutral or even antagonistic in another, so any model that ignores this dimension risks producing predictions that do not transfer beyond the strains it was trained on. By encoding interactions at the level of drug pairs within specific strains, the dataset allows a model to learn how context shapes the outcome of combining two compounds.</p>
<p>The centerpiece of the study is HALO, which stands for Held-out Antibacterial interaction Learning from latent bioactivity Observations. HALO represents each drug pair using multi-level Chemical Checker similarity features, a framework that describes compounds not by a single fingerprint but across five nested domains: chemical structure, protein targets, biological networks, cellular phenotypes, and clinical bioactivity. This layered representation is the conceptual heart of the work. Two antibiotics may look dissimilar as molecules yet converge on overlapping targets or perturb similar cellular processes, and it is often this deeper functional resemblance, rather than surface chemistry, that determines whether their combined effect is synergistic, additive, or antagonistic. By encoding drugs at multiple biological levels simultaneously, HALO gives the learning algorithm access to the kinds of information that medicinal chemists and pharmacologists would weigh when reasoning about a combination.</p>
<p>Evaluation is where HALO sets itself apart. The researchers used strictly nested, drug-pair held-out cross-validation, meaning that entire drug pairs were excluded from training when they appeared in the test fold, and feature selection was performed inside each fold so that no information from the test set could influence which features the model relied on. Under these conservative conditions, HALO achieved a ROC–AUC of 0.78 for generalizing to unseen combinations. That number may look unremarkable next to the near-perfect scores sometimes reported in the literature, but the authors show why the comparison is misleading. As the held-out schemes became increasingly stringent, performance dropped notably, and under random data splits, where leakage is permitted, metrics inflated substantially. The study thereby clarifies the realistic performance limits of current models under truly leakage-free evaluation, a contribution that may prove as influential as the model itself.</p>
<p>The distinction between random splits and pair-level held-out evaluation deserves emphasis, because it gets at a subtle but consequential flaw in how machine learning results are often reported. In a random split, fragments of information about a given drug pair, and even about individual drugs, are scattered across training and test sets. The model can interpolate from near-neighbors rather than extrapolate to genuinely novel territory. Pair-level held-out validation forbids this shortcut: the model must predict the behavior of a combination it has never encountered in any form. For drug discovery, this is the only evaluation that mirrors reality, because the whole point of a predictor is to rank untested pairs. A model that scores brilliantly on random splits but collapses under pair-level validation is, in practical terms, a model that cannot yet do its job.</p>
<p>Despite these deliberately conservative conditions, HALO demonstrated genuine transferability. When applied to two independent external datasets, the framework achieved ROC–AUC values of 0.90 and 0.71, and average precision scores of 0.76 and 0.94, for the task of distinguishing synergistic interactions from antagonistic ones. The variability between the two datasets is itself informative. Transfer performance depends on how similar the new data are to the training distribution, including the compounds, strains, and assay conditions involved. A ROC–AUC of 0.90 on one independent dataset suggests that the multi-level bioactivity signatures capture something real and portable about antibacterial interaction biology, while the more modest 0.71 on the second tempers expectations and underscores that no representation, however rich, fully insulates a model from distribution shift.</p>
<p>One of the quieter strengths of the approach is interpretability. Because the Chemical Checker features are organized into recognizable biological levels, from chemistry to targets to networks to cellular and clinical effects, the model&#8217;s reliance on particular feature families can be inspected rather than treated as an opaque black box. This matters for a field in which predictions ultimately need to persuade experimentalists to commit laboratory resources. A synergy prediction backed by evidence that two compounds share target-level or network-level signatures is more actionable, and more falsifiable, than a bare probability score emerging from an inscrutable architecture. Scalability is the other practical advantage: latent bioactivity features can in principle be computed for large compound libraries, allowing the framework to screen far more candidate pairs than any feasible wet-lab campaign.</p>
<p>The implications extend beyond antibacterial research. The methodological lesson, that permissive cross-validation can manufacture optimism where little exists, applies to synergy prediction in cancer drug combinations, drug–drug interaction prediction, and indeed much of computational pharmacology. The study provides a template: curate data at the level of the biological unit that matters, here the drug pair within a strain; encode that unit with features that span multiple levels of biological organization; and evaluate with held-out schemes that forbid the model from seeing anything resembling the test case during training. Teams adopting this template may report lower headline numbers than their predecessors, but those numbers will describe what models can actually be expected to do when confronted with genuinely novel combinations.</p>
<p>For clinicians and microbiologists, the near-term significance is a more honest map of what computational synergy prediction can and cannot yet deliver. HALO&#8217;s results suggest that machine learning, fed with multi-level bioactivity representations, can meaningfully prioritize which untested antibiotic pairs deserve experimental attention, even under evaluation conditions that strip away the advantages of memorization. At the same time, the study is a caution against reading inflated benchmark scores as evidence of clinical readiness. Progress against resistant bacteria will depend on pairing ambitious models with evaluation standards rigorous enough to reveal their true limits, and on datasets large and diverse enough to train them. This work moves both dials forward, offering the field a framework that is at once more demanding and more trustworthy, and a demonstration that when the evaluation is honest, multi-level bioactivity signatures still carry real predictive signal for one of medicine&#8217;s most urgent searches.</p>
<p><strong>Subject of Research:</strong> Machine learning prediction of antibacterial drug synergy using multi-level bioactivity features</p>
<p><strong>Article Title:</strong> Bioactivity-driven prediction of antibacterial synergy using machine learning models</p>
<p><strong>Article References:</strong> Yousefabadi, H., &amp; Mehrmohamadi, M. (2026). Bioactivity-driven prediction of antibacterial synergy using machine learning models. <em>PLOS Computational Biology, 22</em>(10), e1014849. <a href="https://doi.org/10.1371/journal.pcbi.1014849" rel="noopener noreferrer">https://doi.org/10.1371/journal.pcbi.1014849</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pcbi.1014849" rel="noopener noreferrer">10.1371/journal.pcbi.1014849</a></p>
<p><strong>Keywords:</strong> antibiotic resistance, drug synergy, machine learning, HALO, Chemical Checker, cross-validation, bioactivity signatures, PLOS Computational Biology, combination therapy, antibacterial compounds, data leakage, predictive modeling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">261722</post-id>	</item>
	</channel>
</rss>
