Skip to content

Toward Synthesizability-Aware Multi-Step Retrosynthetic Planning

· 0 citations · 40 references

TL;DR

This work proposes GuideRetro, a synthesizability-aware framework for multi-step retrosynthetic planning that integrates global syn-thesizability knowledge into step-wise retrosyn-thetic prediction and improves planning accuracy and search efficiency under realistic retrosynthetic settings.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

This work introduces Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions and establishes Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.

B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al. · 1 citation
Sep 2026

SynOmega: Simplifying Retrosynthesis for Efficient Synthesizability Scoring

Retrosynthesis-based synthesizability scoring triages molecules from generative design but is expensive: every score requires a multi-step search. We present SynOmega, an open-source toolkit that couples a single-step template model, an AND–OR route search, and a route-based synthesizability score (SynScore). Its single-step model can be restricted, at the reaction-template level, to simplifying disconnections that split the target into smaller precursors. On 1000 ChEMBL drug molecules this yields two findings. (1) The simplifying constraint cuts node expansions by about 30% on jointly solved targets and lowers the median search time by about a third. This reduction in search effort is the robust, budget-independent result; the small accompanying rise in solved rate is a secondary effect between the two separately trained models, not the isolated result of toggling one model’s action space. (2) As a complete system, under matched search depth, width and iteration budget, SynOmega reaches about 1.8× the solved rate of the open-source planner AiZynthFinder while searching about 13× faster. SynOmega thus offers a cheap, data-level action-space constraint that makes route-based synthesizability scoring more efficient without sacrificing solvability.

Bai-Cheng Zhang, Guo-Qing Zhang, Jun Jiang et al. · 0 citations
Aug 2026

Leveraging the Condensed Graph of Reaction for Clustering Retrosynthetic Pathways.

Modern retrosynthetic tools can propose hundreds of alternative pathways for a single target, making it challenging to effectively explore and navigate the resulting route space. We present a CGR-based framework for the analysis and clustering of synthetic routes that integrates both target-centered and all-species-centered perspectives. Entire reaction pathways are encoded as single-molecule graphs (RouteCGR) or reduced representations retaining only target atoms (SB-CGR), enabling automatic identification of strategic bond patterns (SBPs). These representations can be transformed into Morgan fingerprints for a quantitative route similarity assessment. We propose a two-level clustering strategy in which routes are first grouped by shared SBPs, ensuring high interpretability based on key retrosynthetic disconnections, and then further differentiated using RouteCGR similarity to capture variations in starting materials and auxiliary transformations. The method demonstrates near-linear scalability and computational efficiency for large data sets. Application to synthetic routes for apatinib generated by multiple planning tools reveals tool-dependent diversity in strategic disconnections and highlights the benefit of combining tools to expand route space. The framework also supports cross-target analysis, enabling the identification of reusable route families that share common strategic disconnections and building blocks across related molecules. Overall, the SBP-based approach provides an interpretable and scalable solution for automated synthesis route analysis and informed decision-making. The proposed approach is implemented in SynPlanner retrosynthesis planning software and is available as a standalone module at https://github.com/Laboratoire-de-Chemoinformatique/SynPlanner.

Almaz Gilmullin, T. Akhmetshin, D. Zankov et al. · 0 citations
Preprint Aug 2026

Synthesizing like a chemist: an iterative, feedback-driven loop for materials discovery

Most computationally predicted materials are never synthesized because conventional synthesis optimization is slow, expertise-dependent, and iterative. Here we present a closed-loop framework that automates this expert workflow by placing human tacit knowledge in the loop through a large language model (LLM) that distills synthesis knowledge from the literature, high-throughput hyperspectral imaging for rapid film evaluation, and multi-objective Bayesian optimization guided by experimental feedback. In a paired optimization campaign, LLM-assisted initialization produced more Pareto-optimal samples and higher hypervolume than a Latin hypercube sampling baseline at matched trial counts, and this advantage persisted throughout iterative optimization. We demonstrate the framework by synthesizing the previously unreported perovskite-inspired compound Rb3BiI6 as thin films and validating the optimized films by optical bandgap analysis and X-ray diffraction. The framework transforms synthesis prediction from single-shot recommendation to iterative learning, providing a generalizable strategy to accelerate automated and fully autonomous experimental materials discovery.

Fang Sheng, Steven B. Torrisi, Amanda A. Volk et al. · 0 citations
Open access Aug 2026

Closed-Loop Solid-State Synthesis Planning for Materials Discovery With Large Language Models.

Developing reliable synthesis routes for complex materials remains a major bottleneck in accelerating materials discovery. This study establishes a large language model-based framework for predicting and optimizing synthesis conditions directly from the literature data. Key synthesis information, including target compounds, precursors, and processing parameters, was systematically extracted from 4407 open-access solid-state synthesis papers and organized into a structured recipe dataset. Using a retrieval-augmented generation (RAG) approach, the system first retrieves similar recipes from the corpus and then generates a new candidate recipe conditioned on those exemplars. The generated recipes were benchmarked against literature data using quantitative scoring metrics, achieving strong agreement with experimentally reported conditions. To validate the predictive capability, the framework was applied to unreported solid-state electrolyte candidates identified through first-principles screening, and multiple oxy-selenide compounds were successfully synthesized through iterative feedback between the model and experiment. The recipe generator accurately refined synthesis parameters over successive trials, demonstrating its ability to reproduce phase-pure products while minimizing trial-and-error. This approach establishes a data-driven, feedback-optimized route to accelerate synthesis design, offering a generalizable paradigm for integrating language models into experimental materials research.

Dong Won Jeon, Dong Hwi Kim, Taeyang Jeon et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.