Every new medicine begins with a question of geometry: how exactly does a small molecule nestle into the pocket of a disease-related protein? Getting that three-dimensional arrangement, known as the ligand pose, right is one of the foundational tasks of structure-based drug discovery, because the position, orientation, and conformation of a ligand inside a binding site determine everything that follows, from virtual screening of compound libraries to the fine-tuning of lead candidates. A team of researchers from Bangladesh University of Engineering and Technology, Rajshahi University of Engineering and Technology, the University of Newcastle, and Monash University now reports a deep learning framework, called KTransPose, that measurably improves the accuracy of these poses, and their results, published in Applied Intelligence, suggest that a careful combination of architectural design and physically informed training signals can push graph neural networks beyond their current limits.
The new work builds on a wave of geometric learning methods that have transformed computational docking in recent years. Classical tools such as AutoDock, AutoDock Vina, GOLD, and MedusaDock combine a conformational search over the ligand’s translational, rotational, and torsional degrees of freedom with scoring functions that approximate the energetics of binding. Machine learning scoring functions like NNScore and RF-Score later learned nonlinear relationships between structural features and binding properties, and deep models such as AtomNet and GNINA learned interaction representations directly from three-dimensional structures. More recently, methods like EquiBind, TANKBind, E3Bind, FABind, and DiffDock have moved beyond scoring toward direct prediction or refinement of ligand geometry, with DiffDock famously reformulating pose prediction as a diffusion process over translational, rotational, and torsional variables. Systems such as DynamicBind, NeuralPLexer, AlphaFold 3, FABFlex, and FlowDock have extended these ideas to settings involving protein flexibility and joint modeling of entire protein-ligand complexes.
Yet the authors of the new study identified a persistent gap in how refinement methods handle coordinate errors. A docked ligand may need a common shift in position and orientation, which is a rigid transformation applied to the whole molecule, while at the same time individual atoms may require their own local adjustments to match the true bound conformation. A global transformation alone cannot represent atom-specific corrections, and an atom-wise displacement parameterization, the approach used by the earlier graph-based method MedusaGraph, never explicitly parameterizes a shared rotation and translation for the complete ligand. Pose accuracy must also be considered alongside molecular geometry: a low coordinate error does not by itself guarantee that the ligand’s internal structure is preserved or that it avoids physically implausible overlap with the protein. These observations motivated a design that combines complementary coordinate predictions with pose supervision and additional geometric regularization.
KTransPose answers that challenge with a dual-branch architecture built on a graph neural network encoder. The protein-ligand complex is represented as a graph whose nodes are the ligand atoms and the selected protein-pocket atoms, with edges encoding molecular connectivity and interatomic distances. The encoder, which uses a hidden dimension of 256 and residual connections with GELU activations and dropout, produces a shared representation of the graph. From this representation, a global pose transformation branch pools the ligand and protein node features separately, concatenates them, and passes them through a small multilayer perceptron that outputs an axis-angle rotation vector and a translation vector. The rotation is converted to a valid rotation matrix using Rodrigues’ formula, and the transformation is applied around the ligand centroid. Meanwhile, a local coordinate correction branch applies an additional graph layer with three output channels, producing an independent three-dimensional adjustment for every ligand atom. The final displacement for each atom is simply the sum of the global and local contributions.
The training objective is where the framework earns its docking-oriented character. Rather than supervising displacements directly, the loss operates on the predicted absolute coordinates and combines four components. The first is the root mean squared distance computed in the original receptor coordinate frame, without any alignment, which directly supervises where the ligand sits in the binding site. The second is a mean squared error computed after optimally aligning the prediction to the reference using the classical Kabsch algorithm, implemented differentiably so that gradients flow through the singular value decomposition; this auxiliary term evaluates structural agreement after removing the best rigid transformation. The third term regularizes the complete pairwise distance matrices of the predicted and reference ligand structures, preserving internal ligand geometry independently of absolute position. The fourth is a clash penalty that quadratically punishes protein-ligand atom pairs closer than a cutoff of 2.0 in the normalized coordinate representation. A sensitivity analysis settled on weights of 1.0, 0.10, 0.05, and 0.02 for these four terms, respectively.
That sensitivity analysis produced one of the study’s most instructive findings. With only the receptor-frame RMSD term active, the dual-branch model achieved a mean RMSD of 5.08 angstroms on the PDBbind-2020 benchmark. Adding the auxiliary terms lowered this to 4.67 angstroms, a clear demonstration that alignment-aware, intraligand, and clash supervision genuinely help. But the reverse experiment was even more striking: when the weight of the raw RMSD term was reduced while the auxiliary terms stayed fixed, performance collapsed catastrophically, from 4.67 angstroms at full weight to 27.96 angstroms when the term was removed entirely. The auxiliary geometric terms complement rather than replace the supervision needed for correct absolute placement, a nuance the authors argue is essential for anyone designing docking losses.
The team evaluated seven graph neural network configurations within the framework, including TransformerConv, GraphSAGE, Graph Attention Networks, ARMAConv, and three hybrid architectures combining GIN, GAT, and TransformerConv layers. TransformerConv, an attention-based message-passing operator that can incorporate edge attributes, delivered the strongest overall performance, followed by GraphSAGE and the G3 hybrid, though it also required the most training time per epoch. All experiments were run on modest hardware, an Intel Core i7-7700 desktop with a GeForce GTX 1080 GPU, and five training runs with different random seeds showed a standard deviation of only about 0.01 angstroms, indicating that the results are highly reproducible.
The headline comparison, against MedusaGraph under strictly identical conditions, same initial MedusaDock poses, same preprocessing, same test complexes, atom ordering, graph construction, and RMSD calculation, showed consistent gains. On PDBbind-2020, a dataset of 3,956 preprocessed protein-ligand complexes split by CD-HIT sequence clustering at 90 percent identity, MedusaGraph achieved a mean RMSD of 5.08 angstroms, while KTransPose reached 4.67, 4.51, and 4.47 angstroms across three successive refinement iterations. On the independent CASF-2016 benchmark of 278 evaluated complexes, MedusaGraph’s 4.81 angstroms fell to 4.41, 4.27, and 4.24 angstroms for KTransPose, an improvement of roughly 12 percent on both datasets. Paired two-sided Wilcoxon signed-rank tests confirmed the differences were statistically significant at the matched iterations. Notably, iterative refinement for KTransPose reduced RMSD monotonically, whereas MedusaGraph actually degraded at its third iteration on both datasets, rising from 4.68 to 4.90 angstroms on PDBbind-2020.
The refinement dynamics themselves carry a practical lesson. Most of the total improvement over the initial MedusaDock pose is achieved in the very first iteration, with the second and third iterations contributing progressively smaller gains, just 0.16 and then 0.04 angstroms for the TransformerConv model on PDBbind-2020. Because each additional iteration requires applying the preceding frozen models, updating coordinates, and reconstructing the ligand-dependent graph edges within a 6 angstrom distance threshold, the authors stopped at three iterations, judging the diminishing returns not worth the extra computational cost. Success-rate analyses at fixed RMSD thresholds added further texture: KTransPose showed its clearest advantages at the 3 and 5 angstrom cutoffs, while the strict 2 angstrom success rate remained low for both methods and was not uniformly improved. A supplementary analysis using the DAAP affinity predictor also found that KTransPose-refined complexes yielded better affinity-prediction metrics across Pearson correlation, error, and concordance measures than MedusaGraph outputs.
The authors are candid about the boundaries of their results. KTransPose refines an existing docked pose rather than performing blind or de novo docking, so its performance depends on the quality of the supplied MedusaDock starting point, and protein coordinates are held fixed throughout, meaning receptor flexibility and conformational change are not modeled. The comparison with MedusaGraph is a method-level comparison, so the gains cannot be attributed to the global branch or the loss function alone, and the evaluation rests primarily on PDBbind-2020 with CASF-2016 as an additional test. Even so, the framework’s consistent improvements, its reproducibility, and its publicly released codebase mark it as a meaningful step for computational drug discovery, and the authors point toward equivariant architectures, flexible-receptor settings, and more diverse docking initializations as the next frontiers. For a field where fractions of an angstrom can separate a plausible drug candidate from a dead end, a reliable 12 percent reduction in pose error is no small matter.
Subject of Research: Deep learning-based refinement of protein-ligand binding pose prediction in structure-based drug discovery
Article Title: Enhancing protein–ligand pose prediction via dual-branch deep learning with docking-oriented multi-component loss and iterative refinement
Article References: Alam, M. K., Rahman, J., Newton, M. A. H., & Ali, M. E. (2026). Enhancing protein–ligand pose prediction via dual-branch deep learning with docking-oriented multi-component loss and iterative refinement. Applied Intelligence, 56(15), Article 458. https://doi.org/10.1007/s10489-026-07475-9
Image Credits: AI Generated
DOI: 10.1007/s10489-026-07475-9
Keywords: protein-ligand pose prediction, molecular docking, graph neural networks, drug discovery, KTransPose, MedusaGraph, PDBbind-2020, CASF-2016, Kabsch alignment, TransformerConv, iterative refinement, virtual screening
Cite Scienmag News
Louis Brooks. (September 30, 2026). Dual-Branch AI Sharpens Molecular Docking Poses for Drug Discovery. Scienmag. https://scienmag.com/dual-branch-ai-sharpens-molecular-docking-poses-for-drug-discovery/
Louis Brooks. "Dual-Branch AI Sharpens Molecular Docking Poses for Drug Discovery." Scienmag, 30 September 2026, https://scienmag.com/dual-branch-ai-sharpens-molecular-docking-poses-for-drug-discovery/. Accessed 30 September 2026.
Louis Brooks. "Dual-Branch AI Sharpens Molecular Docking Poses for Drug Discovery." Scienmag. September 30, 2026. https://scienmag.com/dual-branch-ai-sharpens-molecular-docking-poses-for-drug-discovery/

