Friday, October 9, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions

October 9, 2026
in Technology and Engineering
Louis Brooks
By Louis Brooks Scienmag Editorial Profile - Medicinal Chemistry
Reading Time: 6 mins read
0
AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions

AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

One of the most stubborn bottlenecks in modern drug discovery is not imagining new molecules but actually making them. Generative algorithms can propose millions of candidate drugs in an afternoon, yet each suggestion is worthless if a chemist cannot synthesize it in the laboratory. A team of researchers from ETH Zurich and Roche Pharma Research and Early Development has now unveiled a closed-loop workflow that tackles this problem head-on, combining active learning, geometric deep learning and self-supervised training to predict chemical reactions with striking accuracy even when experimental data are scarce. The work, published in Nature Computational Science, demonstrates that an intelligent feedback loop between algorithm and experiment can transform how medicinal chemists decide which reactions to run next.

The central insight of the study is that the limiting resource in early drug discovery has shifted. Algorithms for molecular design have become powerful, but the cost and speed of generating experimental data now constrain how quickly pharmacologically active molecules can be identified and optimized. High-throughput experimentation (HTE) partially relieves this pressure by rapidly producing findable, accessible, interoperable and reusable reaction data, but even automated laboratories have finite bandwidth. Chemists must still decide, under stringent time and material constraints, which of the many conceivable reactions, substrates and conditions deserve to be explored. The researchers argue that this allocation problem is precisely where machine intelligence can pay the greatest dividends.

Their solution begins with active learning, a strategy in which a model itself selects the next most informative experiments rather than passively waiting for data. In chemical reaction optimization, the goal differs from typical drug screening: instead of chasing immediate hits, the aim is to map the reaction landscape, including negative outcomes, so that models generalize to unseen substrates. To serve as the decision-making oracle inside this loop, the team benchmarked ensembles of Random Forest, CatBoost and XGBoost classifiers. Deep neural ensembles are strong uncertainty estimators, but retraining them after every iteration and scoring millions of candidate molecules would be computationally prohibitive. Tree-based methods proved competitive at a fraction of the cost, and the XGBoost ensemble delivered the best combination of classification performance and uncertainty calibration, scoring the full pool of 22,253 drug-like aromatic compounds in roughly 0.1 seconds.

With the oracle chosen, the team benchmarked ten acquisition functions, the mathematical rules that determine which molecules an active learner requests next. The Bayesian Active Learning by Disagreement (BALD) function consistently emerged as the top performer. Its explorative character proved crucial: rather than repeatedly sampling molecules similar to those already tested, BALD prioritized structurally diverse substrates, broadening scaffold coverage in the training data. Over three prospective rounds, the workflow tested approximately 96 reaction configurations on 50 previously unreported substrates, generating 4,821 new reactions and expanding the dataset to 6,865 reactions spanning 568 unique substrates. Quantitative analysis confirmed the explorative behavior: the mean Tanimoto distance between selected substrates and their nearest training-set neighbors was 0.69, and the campaign added 42 previously unseen Bemis–Murcko scaffolds, expanding scaffold diversity from 105 to 147 unique frameworks, a 40 percent increase.

The reaction chosen to showcase the workflow was iridium-catalyzed C–H borylation, a late-stage functionalization technique with a special place in medicinal chemistry. Late-stage functionalization generates analogs by installing functional groups on fully elaborated scaffolds without designing new synthetic routes from scratch, allowing chemists to probe structure–activity relationships and fine-tune absorption, distribution, metabolism, excretion and toxicity profiles on accelerated timelines. Among such methods, C–H borylation is particularly attractive because the resulting boronate esters serve as versatile handles for a broad range of downstream cross-coupling transformations, dramatically expanding the chemical space accessible from a single intermediate. The catch is regioselectivity: drug-like molecules often carry multiple similarly reactive C–H bonds, and predicting which one will be borylated has long been a formidable challenge.

Once the active learning campaign had enriched the dataset, the computational calculus changed. With sufficient data in hand, the focus shifted from rapid iteration to maximizing predictive performance, and the researchers transitioned to geometric graph neural networks (GNNs). These models operate directly on three-dimensional molecular graphs, encoding translational, rotational and permutational symmetries through architectural constraints. The team benchmarked a panel of ten GNN architectures, including invariant networks such as SchNet, equivariant networks such as PaiNN, Ponita and EquiformerV2, and unconstrained designs such as FAENet, against several XGBoost baselines across three increasingly challenging data splits: random, Butina clustering and Bemis–Murcko scaffold. On random splits, the best GNN and a condition-aware XGBoost model performed comparably, but on the harder scaffold split, which tests generalization to unseen molecular series, the GNNs pulled decisively ahead, achieving a mean Matthews correlation coefficient between 0.43 and 0.50 compared with 0.35 for the best XGBoost model.

A particularly elegant contribution lies in the self-supervised auxiliary tasks the researchers grafted onto their networks. Rather than relying on offline pretraining with large external molecular corpora, they incorporated two online objectives: coordinate denoising, in which Gaussian noise is added to atomic positions and the model must predict the noise vector, and node masking, in which atomic identities are hidden and must be reconstructed, following a scheme reminiscent of masked language modeling. These tasks require no additional unlabeled molecules and no separate pretraining stage, yet they consistently improved performance across most architectures, especially on the challenging scaffold splits. The finding suggests that robust chemical priors can be extracted from limited labeled examples, and that data quality and training strategy may matter more than architectural expressiveness once a baseline level of geometric awareness is in place.

The regioselectivity results are the study’s most striking demonstration. Framed as node-level binary classification of whether each sp2 and sp3 carbon is borylated, the task was first benchmarked against T5Chem, a prominent reaction-aware chemical language model. The condition-aware GNNs outperformed T5Chem on the harder Butina and scaffold splits, with Ponita Fiber reaching an MCC of 0.59 on the scaffold split compared with 0.38 for T5Chem. For prospective validation, the team selected EquiformerV2 with coordinate denoising, retrained it on the complete dataset of 812 starting materials and 920 borylated products, and asked it to predict the borylation sites of ten previously unreported substrates bearing challenging N-heteroaryl motifs. In every single case, the most probable predicted site matched the experiment. For compounds with a single dominant predicted site, the top-ranked position was confirmed, with high-confidence predictions of 0.98 and 0.83 for two substrates, and for symmetric molecules the model sensibly assigned near-equal probabilities to equivalent positions.

To show the practical payoff, the researchers conducted an in silico late-stage functionalization exercise, using the top-predicted borylation sites of ten substrates to enumerate virtual products via Suzuki–Miyaura coupling against a curated library of aryl-halide building blocks. Visualization of the embedded products confirmed that accurate regioselectivity prediction opens access to a wide range of chemical diversity, allowing medicinal chemists to map the synthetically accessible space around a scaffold and concentrate experimental effort on the most promising analogs. The generality of the approach was further tested on Minisci C–H alkylation, a mechanistically distinct radical-mediated reaction comprising 13,490 reactions across 80 heterocyclic substrates and 59 alkyl radical sources. The same pattern held: GNNs and XGBoost performed comparably on random splits, while GNNs were markedly superior on scaffold splits for both classification and yield prediction, with auxiliary tasks again providing broad benefit.

The authors are candid about limitations. The XGBoost choice reflects a practical trade-off favoring simplicity rather than a claim of universal optimality, and Gaussian processes or deep neural ensembles would be tractable at this dataset scale. The EquiformerV2 model showed reduced accuracy on the most structurally intricate substrates, and the field would benefit from larger, more diverse public reaction datasets with better class balance and a higher proportion of drug-like molecules. Still, the study delivers a validated template for data-efficient chemistry: an explorative active learning loop that simultaneously produced a larger, publicly available dataset and steadily improved predictive models, coupled with symmetry-aware networks whose self-supervised regularization squeezes more insight from every labeled experiment. As 3D foundation models and adaptive multitask strategies mature, this closed-loop paradigm points toward a future in which the question is no longer whether a molecule can be made, but which of thousands of possible analogs are worth making first.

Subject of Research: Machine learning-guided prediction of chemical reaction outcomes and regioselectivity for late-stage functionalization in drug discovery

Article Title: Advancing chemical reaction prediction in data-scarce drug discovery with active and geometric deep learning

Article References: Minot, M., Stenzhorn, Y., Wolfard, J., Strobel, S., Jablonski, P., Zimmerli, D., Binder, M., Grether, U., Martin, R. E., Müller, A. T., Nippa, D. F., Atz, K., & Schneider, G. (2026). Advancing chemical reaction prediction in data-scarce drug discovery with active and geometric deep learning. Nature Computational Science. https://doi.org/10.1038/s43588-026-01063-0

Image Credits: AI Generated

DOI: 10.1038/s43588-026-01063-0

Keywords: active learning, geometric deep learning, graph neural networks, C–H borylation, regioselectivity, late-stage functionalization, drug discovery, medicinal chemistry, high-throughput experimentation, self-supervised learning, reaction prediction, XGBoost

Cite Scienmag News

Louis Brooks. (October 9, 2026). AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions. Scienmag. https://scienmag.com/ai-chemists-learn-where-to-break-molecules-active-learning-steers-drug-discovery-reactions/

Louis Brooks. "AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions." Scienmag, 9 October 2026, https://scienmag.com/ai-chemists-learn-where-to-break-molecules-active-learning-steers-drug-discovery-reactions/. Accessed 9 October 2026.

Louis Brooks. "AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions." Scienmag. October 9, 2026. https://scienmag.com/ai-chemists-learn-where-to-break-molecules-active-learning-steers-drug-discovery-reactions/

Tags: active learningC–H borylationdrug discoverydrug synthesis reactions to prioritize. The researchers' innovative approach uses active learning to guide these decisionsgeometric deep learningGraph Neural Networkshigh-throughput experimentationlate-stage functionalizationmedicinal chemistryoptimizing the reaction selection process and accelerating drug development timelines.reaction predictionregioselectivityself-supervised learningXGBoost
Share26Tweet16
Previous Post

Chemotherapy Drug Speeds Ovarian Aging by Breaking the Body’s Cellular Clock, Study Finds

Next Post

Vitamin C and Exercise Leave New Glucose Sensor Unshaken, Trial Finds

Related Posts

Where mRNA Vaccines Really Go: New Review Maps the Journey of Lipid Nanoparticles Through the Body
Technology and Engineering

Where mRNA Vaccines Really Go: New Review Maps the Journey of Lipid Nanoparticles Through the Body

October 9, 2026
Alveolar Stem Cells Cross the Lung’s Divide to Rebuild Damaged Airways
Medicine

Alveolar Stem Cells Cross the Lung’s Divide to Rebuild Damaged Airways

October 9, 2026
Atomic-Scale Simulations Reveal How Palladium and Grain Size Shape the Strength of High-Entropy Alloys
Technology and Engineering

Atomic-Scale Simulations Reveal How Palladium and Grain Size Shape the Strength of High-Entropy Alloys

October 9, 2026
Performance Bonds Could Make Satellite Operators Clean Up Their Own Orbital Mess
Technology and Engineering

Performance Bonds Could Make Satellite Operators Clean Up Their Own Orbital Mess

October 9, 2026
AI-Modulated Urns Unify Reinforcement Learning, Risk Control and Portfolio Design
Technology and Engineering

AI-Modulated Urns Unify Reinforcement Learning, Risk Control and Portfolio Design

October 9, 2026
Hidden Rules Behind the Mirror Symmetry of NMR Spectra Revealed
Chemistry

Hidden Rules Behind the Mirror Symmetry of NMR Spectra Revealed

October 9, 2026
Next Post
Vitamin C and Exercise Leave New Glucose Sensor Unshaken, Trial Finds

Vitamin C and Exercise Leave New Glucose Sensor Unshaken, Trial Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Vitamin C and Exercise Leave New Glucose Sensor Unshaken, Trial Finds
  • AI Chemists Learn Where to Break Molecules: Active Learning Steers Drug Discovery Reactions
  • Chemotherapy Drug Speeds Ovarian Aging by Breaking the Body’s Cellular Clock, Study Finds
  • Moon Cities May Run Dry: Polar Water Too Scarce for Large Lunar Settlements

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading