Organoboron compounds are widely used across pharmaceuticals and materials science, where 11B NMR spectroscopy serves as a valuable tool for structural characterization. However, severe spectral line broadening induced by the quadrupolar nature of the boron nucleus often causes signal overlap, making it exceptionally difficult to experimentally resolve chemically inequivalent sites in complex multiboron architectures. While traditional density functional theory can resolve these ambiguities, it faces prohibitive computational bottlenecks, whereas data-driven alternatives remain constrained by the scarcity of high-quality data sets. Herein, we report a manually verified, solvent-annotated 11B NMR data set constructed via a large language model (LLM)-assisted workflow. Interpretable machine learning identifies a strong correlation between the BCUT2D_MRLOW descriptor and the boron hybridization. Integrating these ML-derived features as prior knowledge, we developed a prior-guided Graph Transformer for accurate atom-level chemical shift prediction. Notably, the model provides a form of virtual spectral resolution, enabling the discrimination of chemically inequivalent boron sites that are difficult to resolve experimentally. We further deploy the framework as an open-access Web tool to support the rapid structural analysis of organoboron compounds.
Penghui Li, Ben Gao, Shiyang Wang et al.· JACS Au· 0 citations
Despite the potential of Large Language Models (LLMs) in chemical discovery, current LLMs still lack fundamental chemical domain knowledge, produce incoherent reasoning trajectories, and exhibit suboptimal performance across diverse chemical tasks. To address these challenges, we propose Chem-R, a general Chemical Reasoning model designed to emulate the deliberative processes of chemists. To build advanced reasoning capabilities of Chem-R, we design a three-phase training framework, including: 1) Chemical Foundation Training (CFT), which establishes core chemical knowledge. 2) Chemical Reasoning Protocol (CRP) Distillation, incorporating structured, expert-like reasoning traces to guide systematic and reliable problem solving. 3) Chemical Multi-Task Optimization (CMO) that optimizes the model for generalizable capabilities across diverse molecular- and reaction-level tasks. This structured pipeline enables Chem-R to achieve state-of-the-art performance on comprehensive benchmarks, surpassing leading LLMs, including Gemini-3-Pro and Kimi-k2.5, by up to 19% on molecular tasks and 40% on reaction tasks. Meanwhile, Chem-R also consistently outperforms existing chemical foundation models across both molecular and reaction level tasks. These results demonstrate Chem-R's superior generalization, interpretability, and potential as a foundation for next-generation AI-driven chemical discovery. The code and model are available at https://github.com/davidweidawang/Chem-R.
Weida Wang, Benteng Chen, Di Zhang et al.· Proceedings of the 32nd ACM...· 0 citations