<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>global aggregation pooling in drug-target models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/global-aggregation-pooling-in-drug-target-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 09:03:21 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>global aggregation pooling in drug-target models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Smarter Classifiers, Not Flashier Attention, Drive Gains in Drug-Target AI</title>
		<link>https://scienmag.com/smarter-classifiers-not-flashier-attention-drive-gains-in-drug-target-ai/</link>
		
		<dc:creator><![CDATA[Louis Brooks]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 09:03:21 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[ablation study]]></category>
		<category><![CDATA[advances in computational drug discovery]]></category>
		<category><![CDATA[AUROC]]></category>
		<category><![CDATA[BCAG-DTI model for drug discovery]]></category>
		<category><![CDATA[bidirectional cross-attention in neural networks]]></category>
		<category><![CDATA[BindingDB]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[BIOSNAP]]></category>
		<category><![CDATA[classifier design]]></category>
		<category><![CDATA[controlled model comparison in bioinformatics]]></category>
		<category><![CDATA[cross-attention]]></category>
		<category><![CDATA[Davis]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning model evaluation]]></category>
		<category><![CDATA[distinguishing architecture from training improvements]]></category>
		<category><![CDATA[drug-target interaction]]></category>
		<category><![CDATA[drug-target interaction prediction]]></category>
		<category><![CDATA[global aggregation pooling in drug-target models]]></category>
		<category><![CDATA[impact of training strategies on model performance]]></category>
		<category><![CDATA[interpretability of AI in pharmacology]]></category>
		<category><![CDATA[model performance assessment in bioinformatics]]></category>
		<category><![CDATA[MolTrans]]></category>
		<category><![CDATA[MolTrans baseline in drug-target prediction]]></category>
		<category><![CDATA[optimisation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=226730</guid>

					<description><![CDATA[A controlled ablation study shows that classifier design and optimisation, not cross-attention alone, account for most of the improvement a new MolTrans-based model achieves in drug-target interaction prediction.]]></description>
										<content:encoded><![CDATA[<p>Predicting whether a drug molecule will bind to a particular protein target is one of the most consequential questions in computational drug discovery, and deep learning models have promised to answer it at scale. Yet a persistent problem has haunted the field: when a new model reports better numbers than its predecessor, it is often unclear whether the improvement comes from a genuinely smarter architecture or from quieter changes to how the model is trained and how its outputs are classified. A team at the University of Nottingham Ningbo China has now tackled that ambiguity head-on, publishing a controlled evaluation in BMC Bioinformatics that dissects exactly where the performance gains of a MolTrans-based drug-target interaction predictor actually come from.</p>
<p>The study, led by Hao Pang, Fiseha Berhanu Tesema, Tianxiang Cui, Yuan Cheng, and Yanwen Mao, introduces a model called BCAG-DTI, short for Bidirectional Cross-Attention and Global Aggregation drug-target interaction model. Rather than simply claiming superiority over the widely used MolTrans baseline, the researchers ran eight carefully controlled configurations that separately toggled bidirectional cross-attention, a global average-and-max pooling layer, the design of the classification head, and the optimisation strategy. By changing one factor at a time across fixed training, validation, and test partitions with five matched random seeds, they could attribute each gain to its true source, a level of experimental hygiene that remains rare in a literature crowded with simultaneous architectural and training tweaks.</p>
<p>The headline results are striking. On the BindingDB benchmark, the complete BCAG-DTI configuration lifted the mean area under the receiver operating characteristic curve, or AUROC, from 0.8815 for MolTrans to 0.9063. On BIOSNAP, the corresponding figure rose from 0.8631 to 0.8891, and on DAVIS from 0.8808 to 0.8955. Because AUROC alone can flatter models on imbalanced datasets, the team also measured the area under the precision-recall curve, AUPRC, which is often the more honest metric for interaction prediction. There the improvements were even more pronounced: gains of 0.0831 on BindingDB, 0.0346 on BIOSNAP, and 0.0619 on DAVIS, representing substantial advances in the model&#8217;s ability to rank true interactions above false ones.</p>
<p>The most provocative finding, however, lies in the ablation analysis that followed. When each component was tested in isolation, the enhanced classifier design emerged as the single most influential element, achieving the highest mean AUROC on BindingDB and BIOSNAP and the highest mean AUPRC and F1-score on all three datasets. The enhanced training strategy also delivered substantial improvements on its own. By contrast, bidirectional cross-attention alone actually reduced performance when paired with the baseline training configuration. In other words, the glamorous cross-modal attention machinery that gives the model its name contributes value only under the right optimisation conditions, and its benefit is dataset-dependent rather than universal. On DAVIS, it was the combination of cross-attention, global pooling, and enhanced training that achieved the highest mean AUROC, underscoring how sensitive these components are to context.</p>
<p>This pattern carries a lesson that extends well beyond one model family. Transformer-style cross-attention between molecular and protein representations has become a fashionable design choice in bioinformatics, frequently credited with allowing models to learn which parts of a ligand speak to which residues of a binding pocket. The Nottingham results suggest that such credit may sometimes be misplaced, or at least overstated, because classifier head design and optimisation schedules can account for a large share of the reported improvement. For a field racing to publish ever-larger architectures, the message is that rigorous component-wise evaluation is not optional bookkeeping but the difference between genuine scientific progress and an artefact of training procedure.</p>
<p>Robustness testing added another layer of nuance. In experiments on BIOSNAP designed to simulate real-world sparsity, BCAG-DTI outperformed MolTrans when predicting interactions for unseen drugs, for unseen proteins, and in settings where 70 to 90 percent of interaction labels were missing. These scenarios matter enormously in practice, because the vast majority of possible drug-protein pairs have never been measured, and any screening tool must generalise beyond the sparse matrix of known interactions. The one blemish appeared at 95 percent missing data, where BCAG-DTI&#8217;s AUROC and F1-score dipped slightly below the baseline, a reminder that even improved models have limits when supervision becomes extremely thin.</p>
<p>To situate their results against the broader literature, the authors also evaluated an external reference model, CPI-GGS, on the same fixed BIOSNAP partitions. CPI-GGS achieved 0.8619 plus or minus 0.0023 AUROC, 0.8645 plus or minus 0.0042 AUPRC, and 0.7938 plus or minus 0.0028 F1-score, compared with 0.8891 plus or minus 0.0078, 0.8992 plus or minus 0.0067, and 0.8178 plus or minus 0.0094 for BCAG-DTI. Crucially, the team interpreted this comparison with appropriate caution, noting that the two systems use different input preprocessing pipelines and differ substantially in model capacity. Rather than declaring outright victory, they framed the numbers as a same-split reference point, a refreshing contrast to the practice of comparing scores across incompatible datasets and splits that still plagues the field.</p>
<p>Perhaps the most intellectually honest portion of the paper concerns attention interpretability. Cross-attention weights are often visualised and marketed as evidence that a model has learned genuine binding contacts between a ligand and a protein pocket. The authors&#8217; case analysis pushes back on this convention, indicating that cross-attention weights should be treated as model-internal allocation patterns rather than validated binding contacts. This distinction matters for downstream users: a researcher who trusts an attention heatmap as a structural hypothesis could waste laboratory resources chasing a pattern that reflects the model&#8217;s internal bookkeeping rather than physical chemistry. It is a caution that resonates across the growing literature on attention-based models in the life sciences.</p>
<p>The study&#8217;s methodology deserves emphasis as a model for the field. Fixing the data partitions across all principal experiments and running five matched random seeds allowed the authors to report means with a meaningful sense of variance, guarding against the lucky-seed effect that can inflate single-run results. The three benchmarks themselves span different regimes: BindingDB offers a large, chemically diverse collection of measured affinities, BIOSNAP provides a network-derived interaction dataset well suited to cold-start and missing-data experiments, and DAVIS presents a smaller, kinase-focused panel. Showing consistent gains across all three, while honestly reporting where components underperform, lends the conclusions a credibility that single-benchmark papers rarely achieve.</p>
<p>For drug discovery pipelines, the practical implications are immediate. Teams building interaction predictors should scrutinise their classification heads and optimisation strategies before investing in architectural complexity, since the cheapest improvements may lie in the least glamorous parts of the stack. Cross-attention and pooling modules should be adopted with the understanding that their value is conditional, emerging only in combination with appropriate training regimes and on certain data distributions. And anyone publishing comparisons against established baselines should take a page from this study&#8217;s playbook: hold the data fixed, vary one thing at a time, and report what actually moved the needle. As machine learning continues to reshape how candidate drugs are prioritised for synthesis and testing, evaluations of this rigour will be essential to ensure that the field&#8217;s progress is real rather than an artefact of experimental design.</p>
<p><strong>Subject of Research:</strong> Controlled ablation evaluation of architectural, classifier, and training refinements in deep learning-based drug-target interaction prediction</p>
<p><strong>Article Title:</strong> Controlled evaluation of architectural, classifier, and training refinements in MolTrans-based drug-target interaction prediction</p>
<p><strong>Article References:</strong> Controlled evaluation of architectural, classifier, and training refinements in MolTrans-based drug-target interaction prediction. (n.d.). <a href="https://doi.org/10.1186/s12859-026-06634-6" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06634-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06634-6" rel="noopener noreferrer">10.1186/s12859-026-06634-6</a></p>
<p><strong>Keywords:</strong> drug-target interaction, MolTrans, cross-attention, ablation study, deep learning, BindingDB, BIOSNAP, DAVIS, AUROC, classifier design, optimisation, bioinformatics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">226730</post-id>	</item>
	</channel>
</rss>
