For more than half a century, one of the most demanding intellectual exercises in chemistry has remained largely unchanged: retrosynthesis, the art of working backward from a desired molecule to the simpler building blocks and reactions that could realistically produce it. Now, a team of researchers from Microsoft Research AI for Science, Novartis Biomedical Research, the University of Cambridge, Jagiellonian University, GSK, and Bergische Universität Wuppertal reports a new artificial intelligence system, called RetroChimera, that appears to close a stubborn gap between what machine learning models can predict and what experienced chemists actually consider plausible. The study, published in Nature, describes both a detailed diagnosis of why existing AI synthesis planners fail and a novel architecture that addresses those failures head-on.
The motivation for the work stems from a simple but uncomfortable observation. Although AI-assisted synthesis planning has proliferated across the chemical and pharmaceutical industries in recent years, a detailed understanding of how and why these systems fail has never been achieved. According to the authors, current models still struggle with predicting less frequent yet strategically critical reactions, and they continue to produce hallucinated, incorrect predictions that are fundamentally misaligned with chemists’ expectations. For a discovery chemist weighing whether to trust a model’s suggestion, such failures are not merely inconvenient; they can derail weeks of planning and undermine confidence in computational tools altogether.
The core of the new system rests on two newly developed model components with what the researchers describe as complementary inductive biases. Inductive biases are the built-in assumptions a model makes about the structure of the problem it is learning, and in chemistry these assumptions matter enormously. One component may excel at capturing the local, rule-like transformations that dominate reaction databases, while another is better positioned to generalize to sparse, unusual reaction classes where such rules break down. By pairing architectures whose strengths and blind spots differ in a structured way, the team ensured that the ensemble’s errors would be less correlated than those of any single model.
Crucially, however, simply averaging predictions from different architectures is not enough, because the outputs of retrosynthesis models are discrete chemical transformations rather than continuous scores. The researchers therefore introduced a novel, learning-based ensembling strategy that integrates the diverse components. Rather than treating the ensemble as a post-hoc投票 mechanism, the approach learns how to weigh and combine candidate predictions in a way that reflects chemical validity and likelihood. This design allows RetroChimera to draw on the complementary knowledge embedded in each constituent model without letting the weaknesses of one corrupt the strengths of the other.
The experimental evaluation was deliberately broad, spanning several orders of magnitude in data scale, and the results point to robustness that previous systems lacked. RetroChimera outperformed leading baselines across the benchmark settings tested, and it demonstrated the ability to learn from very small numbers of examples per reaction class. This latter capability addresses one of the most persistent pain points in machine learning for chemistry: the long tail of the reaction literature. Highly unusual transformations, such as those used to forge complex molecular frameworks in natural product synthesis or to install challenging stereochemical motifs, are chronically underrepresented in public databases, yet they are often the exact steps that determine whether an ambitious synthesis route is feasible at all.
The team also subjected the model to tests of generalization outside its training data, a regime in which many deep learning systems collapse into confident nonsense. RetroChimera’s predictions remained reliable under distribution shift, suggesting that the combination of diverse inductive biases and learned ensembling confers a form of robustness that single-architecture models have not matched. The researchers frame this as a demonstration of the viability of deep learning for accurate synthesis prediction in increasingly challenging regimes, where the easy, well-documented reactions that dominate benchmarks are no longer representative of the problems chemists face.
Perhaps the most consequential validation, however, came not from automated benchmarks but from human experts. Using both pairwise and pointwise evaluation setups, the researchers asked organic chemists to compare RetroChimera’s predictions against published reference reactions and against the outputs of other AI models. In these blind assessments, the chemists systematically preferred the suggestions generated by RetroChimera. That preference for model output over the actual published record of real reactions is a striking result, because it indicates the system can propose transformations that experts judge more sensible or more practical than the routes that chemists historically chose and journals published.
To test whether these gains survive contact with industrial reality, the team performed zero-shot transfer and fine-tuning experiments on internal datasets from two major pharmaceutical companies. These proprietary datasets differ from public reaction corpora in both scale and character, reflecting the molecule types, reaction conditions, and strategic priorities of real drug discovery programs. RetroChimera showed robust generalization under this distribution shift, performing well without any additional training in the zero-shot setting and improving further when fine-tuned on internal data. For pharmaceutical companies that have long been skeptical of benchmarks built on public data alone, this industrial validation may prove as important as any academic result.
The collaboration itself reflects the interdisciplinary demands of the problem. The author list spans machine learning researchers specializing in molecular AI, computational chemists embedded in industrial research organizations, and practicing medicinal and synthetic chemists at GSK, Novartis, and other partner organizations. Equal contributions are credited to Krzysztof Maziarz, Guoqing Liu, and Felix Pultar of Microsoft Research AI for Science, with corresponding authorship shared with Marwin H. S. Segler, who has worked on AI-driven synthesis planning for nearly a decade. Representatives from GSK’s discovery sciences group in Stevenage and Wuppertal-based chemist Mario P. Wiesenfeldt round out a team designed to ensure that evaluation criteria mirrored the judgments of working chemists rather than proxy metrics alone.
Looking ahead, the work suggests a broader lesson for applying AI in the natural sciences: progress may depend less on scaling a single architecture than on deliberately designing and combining models with different, complementary biases. As chemical synthesis remains a critical bottleneck in the discovery and manufacture of functional small molecules, tools that align with expert intuition while generalizing beyond the training literature could reshape how new medicines, materials, and agrochemicals are planned. If systems like RetroChimera continue to earn the trust of the chemists who use them, the long-promised collaboration between artificial intelligence and synthetic chemistry may finally be moving from demonstration to daily practice.
Subject of Research: Chemist-aligned AI retrosynthesis planning using ensembles of diverse inductive bias models
Article Title: Chemist-aligned retrosynthesis by ensembling diverse inductive bias models
Article References: Maziarz, K., Liu, G., Pultar, F., Gardner, J., Gensch, T., Helie, J., Misztela, H., Tripp, A., Li, J., Kornev, A., Gaiński, P., Hoefling, H., Fortunato, M., Gupta, R., Baxter, A., Poole, D. L., Elward, J. M., Krzyzanowski, A., Pogány, P., … Segler, M. H. S. (2026). Chemist-aligned retrosynthesis by ensembling diverse inductive bias models. Nature. https://doi.org/10.1038/s41586-026-11160-9
Image Credits: AI Generated
DOI: 10.1038/s41586-026-11160-9
Keywords: retrosynthesis, artificial intelligence, synthesis planning, machine learning, ensembles, inductive bias, chemical synthesis, pharmaceutical discovery, cheminformatics, deep learning, distribution shift, RetroChimera
Cite Scienmag News
Bethany Barker. (September 22, 2026). AI Retrosynthesis Model RetroChimera Learns Chemistry the Way Chemists Do. Scienmag. https://scienmag.com/ai-retrosynthesis-model-retrochimera-learns-chemistry-the-way-chemists-do/
Bethany Barker. "AI Retrosynthesis Model RetroChimera Learns Chemistry the Way Chemists Do." Scienmag, 22 September 2026, https://scienmag.com/ai-retrosynthesis-model-retrochimera-learns-chemistry-the-way-chemists-do/. Accessed 22 September 2026.
Bethany Barker. "AI Retrosynthesis Model RetroChimera Learns Chemistry the Way Chemists Do." Scienmag. September 22, 2026. https://scienmag.com/ai-retrosynthesis-model-retrochimera-learns-chemistry-the-way-chemists-do/

