Skip to content
Tech News
← Back to articles

Chemist-aligned retrosynthesis by ensembling diverse inductive bias models

read original more articles
Why This Matters

Chemical synthesis planning is a major bottleneck in drug and materials discovery, and existing AI retrosynthesis tools often fail on rare but strategically important reactions or produce chemically implausible predictions. RetroChimera addresses these gaps by combining models with complementary strengths through an ensembling approach, offering more reliable, chemist-preferred predictions. This matters because better AI-driven synthesis planning could accelerate drug discovery pipelines and reduce costly trial-and-error in pharmaceutical and chemical manufacturing.

Key Takeaways

Chemical synthesis remains a critical bottleneck in the discovery and manufacture of functional small molecules1−3. While AI-assisted synthesis planning has proliferated in recent years, a detailed understanding of its failure modes has not been achieved, and models still struggle with predicting less frequent, yet strategically critical reactions, as well as hallucinated, incorrect predictions misaligned with chemists’ expectations4−12. In this work, we analyze the failure modes of current AI models and propose RetroChimera: a frontier retrosynthesis model, built upon two newly developed components with complementary inductive biases, integrated via a novel, learning-based ensembling strategy. Through experiments across several orders of magnitude in data scale, we show RetroChimera outperforms leading baselines, demonstrating robustness outside the training data, as well as the ability to learn from very small numbers of examples per reaction class. Using both pairwise and pointwise setups, we find that organic chemists prefer predictions from RetroChimera over published reference reactions and over other AI models. Finally, we demonstrate zero-shot transfer and fine-tuning on internal datasets from two major pharmaceutical companies, showing robust generalization under distribution shift. Our work demonstrates the viability of deep learning for accurate synthesis prediction in increasingly challenging regimes.