The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan many steps ahead for how to assemble simple building blocks into an intricate target, devise backup strategies, and anticipate procedural challenges. It is also a profoundly creative activity. For half a century, efforts to automate the retrosynthetic design of natural products and other complex molecules have drawn on catalogued reactions, and the resulting tools now report near-complete success on benchmarks built from that same source. But these tools were shaped to fit benchmarked chemistry, and they falter on many natural products, the frontier of the field, whose densely functionalized, polycyclic architectures demand precisely the inventive chemistry the record contains least. Whether a machine could reasonably design such syntheses like an expert chemist does has remained unclear. Here, we show that SynthEx, an agentic framework built on large language models, plans routes to complex natural products that lie beyond the reach of conventional design algorithms. SynthEx proposes competing strategies, assembles a sequence of routine and key steps into a cohesive route, and critiques and improves its own design; the chemistry it favours is more convergent than existing tools produce, and spans a region of reaction space that catalogue-based tools cannot match. Most notably, in blinded assessments, expert chemists judged its key steps comparable to those of published human syntheses and engaged with them as genuine synthesis plans, a response algorithmic route prediction has not previously accomplished. We release routes to more than a thousand natural products as SynthAtlas, an open, interactive database, and anticipate it will become a shared resource for a collection of complex target molecules that lack existing literature routes.
Accurate generation of intricate 3D molecular structures is a fundamental challenge in computational chemistry. Transition states (TS), the transient structures that dictate reaction kinetics, exemplify this challenge, as their calculation remains a major bottleneck in elucidating reaction mechanisms. While generative AI has shown promise in automating TS generation, existing methods are largely confined to simple systems, struggling with the complex structures like those in transition metal catalysis. Here we show a unified framework for general-purpose TS generation, comprising the UniTS-Lib library of 4,391 high-quality structures spanning 42 elements and diverse chemical transformations, coupled with the UniTS-Gen diffusion model that generates 3D TS configurations from 2D reactant graphs using a custom-designed higher-degree equivariant network. Validation demonstrates UniTS-Gen’s superior accuracy and robust generalization to unseen chemical systems. We show that UniTS-Gen provides reliable initial guesses and accelerates discovery by locating kinetically favored conformations. This work provides a scalable and transferable solution for automating mechanistic studies in organic synthesis and beyond.
Li-Cheng Xu, Junyi An, Weiqi Liu et al.· Nature Communications· 0 citations
Developing reliable synthesis routes for complex materials remains a major bottleneck in accelerating materials discovery. This study establishes a large language model-based framework for predicting and optimizing synthesis conditions directly from the literature data. Key synthesis information, including target compounds, precursors, and processing parameters, was systematically extracted from 4407 open-access solid-state synthesis papers and organized into a structured recipe dataset. Using a retrieval-augmented generation (RAG) approach, the system first retrieves similar recipes from the corpus and then generates a new candidate recipe conditioned on those exemplars. The generated recipes were benchmarked against literature data using quantitative scoring metrics, achieving strong agreement with experimentally reported conditions. To validate the predictive capability, the framework was applied to unreported solid-state electrolyte candidates identified through first-principles screening, and multiple oxy-selenide compounds were successfully synthesized through iterative feedback between the model and experiment. The recipe generator accurately refined synthesis parameters over successive trials, demonstrating its ability to reproduce phase-pure products while minimizing trial-and-error. This approach establishes a data-driven, feedback-optimized route to accelerate synthesis design, offering a generalizable paradigm for integrating language models into experimental materials research.
Dong Won Jeon, Dong Hwi Kim, Taeyang Jeon et al.· Advances in Materials· 0 citations
RetroAgent is introduced, an LLM agent that bridges symbolic search and neural reasoning through a harness with structured memory, enabling informed decisions grounded in both global progress and domain knowledge in multi-step retrosynthesis planning.
Yanqiao Zhu, Jingru Gan, Xiaoqi Sun et al.· arXiv.org· 1 citation
This work introduces a domain specific language (DSL)-guided strategy to improve the reasoning and design capability of LLM agents by translating natural language design rules into symbolic predicates encoded in a predefined chemistry DSL, and developed a multi-agent materials design framework.
Dong Hyeon Mok, Seoin Back, Victor Fung et al.· 0 citations
ConspectusThe efficient exploration of natural product-like chemical space remains a central challenge in synthetic chemistry and chemical biology. Despite major advances, many synthetic strategies remain organized around individual targets or predefined compound collections, which can limit the continuity and cumulative expansion of chemical space exploration across successive synthetic campaigns. Developing approaches that enable sustained, scalable, and unbiased expansion of biologically relevant molecular diversity is therefore a key objective for modern synthesis.Here, we formalize Synthetic Framework Evolution (SFE) as a scaffold-centric conceptual framework that organizes synthetic planning in terms of persistence, connectivity, and growth. Building upon emerging concepts in bioinspired chemical divergence and scaffold evolution, SFE directs retrosynthetic analysis toward the identification of information-rich intermediates that function as reusable nodes within an evolving synthetic network. Rather than converging on isolated end points, synthesis is organized around persistent scaffolds that enable iterative divergence and cumulative expansion of interconnected molecular architectures.We demonstrate the implementation of SFE in sesquiterpenoid synthesis, a natural product family whose biosynthetic diversity provides an ideal platform for framework-based design. Central to this approach is the hierarchical organization of conformationally flexible, low-oxidation-state scaffolds that encode latent reactivity and stereochemical information. These intermediates enable access to multiple carbocyclic frameworks through controlled rearrangements, cyclizations, and functional group interconversions, allowing structurally distinct natural product families to emerge from unified and continuously expandable synthetic pathways.Beyond scaffold-level divergence, SFE incorporates oxidative transformations as a second, amplifying dimension of chemical space generation. In this context, oxidation is not treated as a terminal functionalization step, but as a generative process that unlocks new topologies through rearrangements, cascade reactions, and bond reorganization, closely reflecting biosynthetic oxidative phases. In this sense, SFE places oxidative diversification alongside scaffold rearrangement as a complementary and generative driver of chemical space expansion. The integration of these two layers, scaffold persistence and oxidative amplification, enables the efficient construction of structurally complex and densely connected regions of chemical space from a limited set of intermediates.Cheminformatic analysis reveals that SFE-derived compounds occupy broad and biologically relevant regions of chemical space, exhibiting high natural product-likeness, shape diversity, and molecular complexity relative to their size. In contrast to established paradigms such as BIOS, DOS, and CtD, which often focus on the generation of discrete collections or predefined diversification campaigns, SFE emphasizes the cumulative expansion of a unified synthetic framework, enabling diversity to emerge as a function of system evolution rather than target enumeration.By shifting the focus of synthesis from molecule production to framework evolution, SFE establishes a conceptual bridge between synthetic chemistry and biosynthetic organization. This perspective enables a more dynamic and scalable approach to chemical space exploration and provides a general framework for the discovery of new molecular architectures and functions.
Kalliopi Mazaraki, George Karageorgis, A. Zografos· Accounts of Chemical Researc...· 0 citations
This work introduces Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions and establishes Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.
B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.