Jul 2026· The journal of physical chemistry. C, Nanomaterials and interfaces· Vol 130, pp. 10487 - 10503· 0 citations· 71 references
Medicine
TL;DR
This perspective focuses on three opportunities where LLMs can significantly contribute to catalysis: (1) text to properties; (2) text to structure; and (3) text to mechanistic models.
Abstract
Heterogeneous catalysis presents a distinct challenge for artificial intelligence (AI). Data sets are often small and inconsistently reported, catalyst representations are not standardized, and extracting fundamental knowledge requires integrating performance data, spectroscopic characterizations, and mechanistic models across multiple scales. Language offers a unifying representation across these modalities, making catalysis well suited for leveraging large language models (LLMs). By standardizing how catalytic data is represented, LLMs make dispersed experimental results more accessible to downstream statistical modeling. In this perspective, we focus our discussion around three opportunities where LLMs can significantly contribute to catalysis: (1) text to properties; (2) text to structure; and (3) text to mechanistic models. The discussion is followed by a perspective section on LLM-readiness of data, aligning LLM outputs with scientific correctness, and bridging lab-scale discovery to industrial deployment. Across each area, the most productive applications couple dispersed chemical knowledge with physics-grounded validation to produce verifiable hypotheses and actionable representations.
Modern chemistry is pushing the limits of traditional Artificial Intelligence (AI) models, placing unprecedented demands on data availability to address humanity's most pressing challenges. One particular concern is AI's dependence on large, curated data and its tendency to deviate from or misrepresent fundamental chemistry principles. Nonetheless, this concern is often overshadowed by the urgent demand for emergent solutions to real‐world problems. This perspective describes the incorporation of a domain‐specific knowledge representation & reasoning (KR&R) framework with machine learning (ML) for predictive chemistry. KR&R is presented as a framework to represent chemical knowledge, making a formal connection between inductive hypothesis generation and deductive reasoning. By integrating scientific rules into data‐driven processes, upholding a “chemist in the loop” approach, KR&R ensures that ML models are understandable and consistent with existing chemical theory. These concepts are illustrated by case studies where KR&R improves the interpretability of ML predictive models targeting thermodynamic properties (Δ
G
sol
, Δ
vap
H
m
°), reaction yields, and catalytic performance. These examples also show KR&R's importance in managing the complexity of modern computational chemistry, establishing it as a key component of explainable AI in the field.
José Ferraz-Caetano, Filipe Teixeira, M. N. D. S. Cordeiro· WIREs Computational Molecula...· 0 citations
CRISP is introduced, a large language model-assisted framework that treats representation construction as a rule-space exploration and compilation problem: it repeatedly samples target-relevant chemical rules without access to structures, labels or data splits, consolidates related concepts, and compiles each into an executable scalar descriptor supplied to a conventional learner.
Jaehwan Choi, Kunik Jang, Seongmin Kim et al.· 0 citations
Despite the potential of Large Language Models (LLMs) in chemical discovery, current LLMs still lack fundamental chemical domain knowledge, produce incoherent reasoning trajectories, and exhibit suboptimal performance across diverse chemical tasks. To address these challenges, we propose Chem-R, a general Chemical Reasoning model designed to emulate the deliberative processes of chemists. To build advanced reasoning capabilities of Chem-R, we design a three-phase training framework, including: 1) Chemical Foundation Training (CFT), which establishes core chemical knowledge. 2) Chemical Reasoning Protocol (CRP) Distillation, incorporating structured, expert-like reasoning traces to guide systematic and reliable problem solving. 3) Chemical Multi-Task Optimization (CMO) that optimizes the model for generalizable capabilities across diverse molecular- and reaction-level tasks. This structured pipeline enables Chem-R to achieve state-of-the-art performance on comprehensive benchmarks, surpassing leading LLMs, including Gemini-3-Pro and Kimi-k2.5, by up to 19% on molecular tasks and 40% on reaction tasks. Meanwhile, Chem-R also consistently outperforms existing chemical foundation models across both molecular and reaction level tasks. These results demonstrate Chem-R's superior generalization, interpretability, and potential as a foundation for next-generation AI-driven chemical discovery. The code and model are available at https://github.com/davidweidawang/Chem-R.
Weida Wang, Benteng Chen, Di Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
Onepot-Bench 0 is introduced, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution and probes basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.
Brandon Wang, Andrei S. Tyrin, Daniil A. Boiko· 0 citations
An LLM-guided SR framework is introduced, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery, demonstrating that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.
Roberto Aliaga Medina, Paulina Quintanilla, Antonio del Río Chanona· 1 citation
This work introduces Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions and establishes Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.
B. Zagribelnyy, Ivan D. Ilin, N. Bondarev et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.