One of the most stubborn bottlenecks in modern drug discovery is not imagining new molecules but actually making them. Generative algorithms can propose millions of candidate drugs in an afternoon, yet each suggestion is worthless if a chemist cannot synthesize it in the laboratory. A team of researchers from ETH Zurich and Roche Pharma Research and Early Development has now unveiled a closed-loop workflow that tackles this problem head-on, combining active learning, geometric deep learning and self-supervised training to predict chemical reactions with striking accuracy even when experimental data are scarce. The work, published in Nature Computational Science, demonstrates that an intelligent feedback loop between algorithm and experiment can transform how medicinal chemists decide which reactions to run next.
The central insight of the study is that the limiting resource in early drug discovery has shifted. Algorithms for molecular design have become powerful, but the cost and speed of generating experimental data now constrain how quickly pharmacologically active molecules can be identified and optimized. High-throughput experimentation (HTE) partially relieves this pressure by rapidly producing findable, accessible, interoperable and reusable reaction data, but even automated laboratories have finite bandwidth. Chemists must still decide, under stringent time and material constraints, which of the many conceivable reactions, substrates and conditions deserve to be explored. The researchers argue that this allocation problem is precisely where machine intelligence can pay the greatest dividends.
Their solution begins with active learning, a strategy in which a model itself selects the next most informative experiments rather than passively waiting for data. In chemical reaction optimization, the goal differs from typical drug screening: instead of chasing immediate hits, the aim is to map the reaction landscape, including negative outcomes, so that models generalize to unseen substrates. To serve as the decision-making oracle inside this loop, the team benchmarked ensembles of Random Forest, CatBoost and XGBoost classifiers. Deep neural ensembles are strong uncertainty estimators, but retraining them after every iteration and scoring millions of candidate molecules would be computationally prohibitive. Tree-based methods proved competitive at a fraction of the cost, and the XGBoost ensemble delivered the best combination of classification performance and uncertainty calibration, scoring the full pool of 22,253 drug-like aromatic compounds in roughly 0.1 seconds.
With the oracle chosen, the team benchmarked ten acquisition functions, the mathematical rules that determine which molecules an active learner requests next. The Bayesian Active Learning by Disagreement (BALD) function consistently emerged as the top performer. Its explorative character proved crucial: rather than repeatedly sampling molecules similar to those already tested, BALD prioritized structurally diverse substrates, broadening scaffold coverage in the training data. Over three prospective rounds, the workflow tested approximately 96 reaction configurations on 50 previously unreported substrates, generating 4,821 new reactions and expanding the dataset to 6,865 reactions spanning 568 unique substrates. Quantitative analysis confirmed the explorative behavior: the mean Tanimoto distance between selected substrates and their nearest training-set neighbors was 0.69, and the campaign added 42 previously unseen Bemis–Murcko scaffolds, expanding scaffold diversity from 105 to 147 unique frameworks, a 40 percent increase.
The reaction chosen to showcase the workflow was iridium-catalyzed C–H borylation, a late-stage functionalization technique with a special place in medicinal chemistry. Late-stage functionalization generates analogs by installing functional groups on fully elaborated scaffolds without designing new synthetic routes from scratch, allowing chemists to probe structure–activity relationships and fine-tune absorption, distribution, metabolism, excretion and toxicity profiles on accelerated timelines. Among such methods, C–H borylation is particularly attractive because the resulting boronate esters serve as versatile handles for a broad range of downstream cross-coupling transformations, dramatically expanding the chemical space accessible from a single intermediate. The catch is regioselectivity: drug-like molecules often carry multiple similarly reactive C–H bonds, and predicting which one will be borylated has long been a formidable challenge.
Once the active learning campaign had enriched the dataset, the computational calculus changed. With sufficient data in hand, the focus shifted from rapid iteration to maximizing predictive performance, and the researchers transitioned to geometric graph neural networks (GNNs). These models operate directly on three-dimensional molecular graphs, encoding translational, rotational and permutational symmetries through architectural constraints. The team benchmarked a panel of ten GNN architectures, including invariant networks such as SchNet, equivariant networks such as PaiNN, Ponita and EquiformerV2, and unconstrained designs such as FAENet, against several XGBoost baselines across three increasingly challenging data splits: random, Butina clustering and Bemis–Murcko scaffold. On random splits, the best GNN and a condition-aware XGBoost model performed comparably, but on the harder scaffold split, which tests generalization to unseen molecular series, the GNNs pulled decisively ahead, achieving a mean Matthews correlation coefficient between 0.43 and 0.50 compared with 0.35 for the best XGBoost model.
A particularly elegant contribution lies in the self-supervised auxiliary tasks the researchers grafted onto their networks. Rather than relying on offline pretraining with large external molecular corpora, they incorporated two online objectives: coordinate denoising, in which Gaussian noise is added to atomic positions and the model must predict the noise vector, and node masking, in which atomic identities are hidden and must be reconstructed, following a scheme reminiscent of masked language modeling. These tasks require no additional unlabeled molecules and no separate pretraining stage, yet they consistently improved performance across most architectures, especially on the challenging scaffold splits. The finding suggests that robust chemical priors can be extracted from limited labeled examples, and that data quality and training strategy may matter more than architectural expressiveness once a baseline level of geometric awareness is in place.
The regioselectivity results are the study’s most striking demonstration. Framed as node-level binary classification of whether each sp2 and sp3 carbon is borylated, the task was first benchmarked against T5Chem, a prominent reaction-aware chemical language model. The condition-aware GNNs outperformed T5Chem on the harder Butina and scaffold splits, with Ponita Fiber reaching an MCC of 0.59 on the scaffold split compared with 0.38 for T5Chem. For prospective validation, the team selected EquiformerV2 with coordinate denoising, retrained it on the complete dataset of 812 starting materials and 920 borylated products, and asked it to predict the borylation sites of ten previously unreported substrates bearing challenging N-heteroaryl motifs. In every single case, the most probable predicted site matched the experiment. For compounds with a single dominant predicted site, the top-ranked position was confirmed, with high-confidence predictions of 0.98 and 0.83 for two substrates, and for symmetric molecules the model sensibly assigned near-equal probabilities to equivalent positions.
To show the practical payoff, the researchers conducted an in silico late-stage functionalization exercise, using the top-predicted borylation sites of ten substrates to enumerate virtual products via Suzuki–Miyaura coupling against a curated library of aryl-halide building blocks. Visualization of the embedded products confirmed that accurate regioselectivity prediction opens access to a wide range of chemical diversity, allowing medicinal chemists to map the synthetically accessible space around a scaffold and concentrate experimental effort on the most promising analogs. The generality of the approach was further tested on Minisci C–H alkylation, a mechanistically distinct radical-mediated reaction comprising 13,490 reactions across 80 heterocyclic substrates and 59 alkyl radical sources. The same pattern held: GNNs and XGBoost performed comparably on random splits, while GNNs were markedly superior on scaffold splits for both classification and yield prediction, with auxiliary tasks again providing broad benefit.
The authors are candid about limitations. The XGBoost choice reflects a practical trade-off favoring simplicity rather than a claim of universal optimality, and Gaussian processes or deep neural ensembles would be tractable at this dataset scale. The EquiformerV2 model showed reduced accuracy on the most structurally intricate substrates, and the field would benefit from larger, more diverse public reaction datasets with better class balance and a higher proportion of drug-like molecules. Still, the study delivers a validated template for data-efficient chemistry: an explorative active learning loop that simultaneously produced a larger, publicly available dataset and steadily improved predictive models, coupled with symmetry-aware networks whose self-supervised regularization squeezes more insight from every labeled experiment. As 3D foundation models and adaptive multitask strategies mature, this closed-loop paradigm points toward a future in which the question is no longer whether a molecule can be made, but which of thousands of possible analogs are worth making first.
Subject of Research: Machine learning-guided prediction of chemical reaction outcomes and regioselectivity for late-stage functionalization in drug discovery
Article Title: Advancing chemical reaction prediction in data-scarce drug discovery with active and geometric deep learning
Article References: Minot, M., Stenzhorn, Y., Wolfard, J., Strobel, S., Jablonski, P., Zimmerli, D., Binder, M., Grether, U., Martin, R. E., Müller, A. T., Nippa, D. F., Atz, K., & Schneider, G. (2026). Advancing chemical reaction prediction in data-scarce drug discovery with active and geometric deep learning. Nature Computational Science. https://doi.org/10.1038/s43588-026-01063-0
Image Credits: AI Generated
DOI: 10.1038/s43588-026-01063-0
Keywords: active learning, geometric deep learning, graph neural networks, C–H borylation, regioselectivity, late-stage functionalization, drug discovery, medicinal chemistry, high-throughput experimentation, self-supervised learning, reaction prediction, XGBoost
Cite Scienmag News
Louis Brooks. (October 9, 2026). AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions. Scienmag. https://scienmag.com/ai-chemists-learn-where-to-break-molecules-active-learning-steers-drug-discovery-reactions/
Louis Brooks. "AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions." Scienmag, 9 October 2026, https://scienmag.com/ai-chemists-learn-where-to-break-molecules-active-learning-steers-drug-discovery-reactions/. Accessed 9 October 2026.
Louis Brooks. "AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions." Scienmag. October 9, 2026. https://scienmag.com/ai-chemists-learn-where-to-break-molecules-active-learning-steers-drug-discovery-reactions/

