Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 10681-10692· 0 citations· 26 references
Abstract
Retrosynthesis, the process of predicting reactants from products, remains a critical challenge in computational chemistry and drug discovery. While recent deep learning methods have shown strong performance, they remain overly reliant on reaction datasets, which are limited in availability and quality. Large-scale unlabeled molecular data encode rich structural patterns that can be leveraged to learn transferable chemical knowledge, but remain largely unexplored. In this work, we propose KnowRetro (Knowledge-Guided Retrosynthesis Prediction), a chemically-aware framework that learns chemical knowledge from large-scale unlabeled molecules to enhance the accuracy and diversity of retrosynthesis prediction. Specifically, KnowRetro first builds a hierarchical knowledge graph from millions of unlabeled molecules, which captures transformation-relevant relationships among molecules, substructures, and functional groups. It then employs chemically guided pre-training based on substructure decomposition to encourage the model to capture fundamental reaction patterns, followed by fine-tuning with an adapter designed to inject task-relevant knowledge into reactant generation. Extensive experiments demonstrate that KnowRetro achieves high accuracy with improved robustness and diversity in reactant generation. Our code is available at https://github.com/chenyujie1127/KnowRetro.
The results highlight the promise of CRG-based RxnCLF as a scalable reaction foundation model, with the potential to generalize across broader reaction spaces and support diverse downstream reaction informatics tasks, including regioselectivity prediction, enantioselectivity prediction, and reaction condition optimization.
Yiting Zheng, Cheng Fang, Anthony Donofrio et al.· 0 citations
This work introduces Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions and establishes Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.
B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al.· 1 citation
A framework that systematically standardizes and integrates multiple reaction datasets into a high-quality, unique-structure-per-entity dataset, coupled with active learning to strategically expand chemical space is introduced, establishing a blueprint for robust machine learning in synthetic chemistry.
Paulo Neves, Bo Hao, Santeri Aikonen et al.· Nature Computational Science· 1 citation
A mechanism-driven approach to alleviate data dependence and develop a multisource transfer learning (MS-TL) framework that leverages the knowledge embedded in abundant adsorption data sets while accurately capturing local structural dependence, enabling a deep fusion of multidimensional thermodynamic knowledge while preserving local structural information.
Wangqiang Lin, Huiyan Zhang, Jinxin Sun et al.· Journal of the American Chem...· 0 citations
CRISP is introduced, a large language model-assisted framework that treats representation construction as a rule-space exploration and compilation problem: it repeatedly samples target-relevant chemical rules without access to structures, labels or data splits, consolidates related concepts, and compiles each into an executable scalar descriptor supplied to a conventional learner.
Jaehwan Choi, Kunik Jang, Seongmin Kim et al.· 0 citations
Cocrystal formation is a widely used strategy in solid-state chemistry and pharmaceutical development to improve the solubility, stability, and bioavailability of molecules with otherwise poor physicochemical properties. Identifying viable coformer combinations remains laborious and uncertain. A key but underappreciated challenge is that experimental databases overwhelmingly report successful cocrystals, while unsuccessful attempts are rarely documented, creating biased data sets that cause many machine-learning models to make overly optimistic and unreliable predictions when applied to new chemical systems. Here, we address this limitation by reframing cocrystal prediction as a learning problem with missing negative information and by adopting a conservative strategy that focuses on identifying molecular pairs that are very unlikely to form cocrystals. We leverage multiple, independent molecular descriptions─including structural, electronic, and physicochemical characteristics─that provide complementary views for identifying reliable negatives, and use their agreement to exclude implausible combinations from large sets of untested pairs. These highly confident pseudonegative examples are then used to mitigate data imbalance and to fine-tune a pretrained graph attention network for cocrystal prediction. Across large and chemically diverse data sets, this data-centric strategy significantly improves the reliability and generalization of cocrystal prediction models compared with existing deep-learning approaches, demonstrating that carefully correcting for missing negative information is critical for making computational screening more realistic and more useful for guiding future experimental discovery.
Mohammad Amin Ghanavati, S. Moosavi, Sohrab Rohani· Journal of Chemical Informat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.