Aug 2026· Analytical Chemistry· Vol 98 32, pp.
23717-23725
· 0 citations· 17 references
Medicine
TL;DR
PINS (Physics-Informed NMR Structure elucidation model), a generative framework that explicitly bridges the gap between spectral data and molecular topology by enforcing multiphysical priors, provides a trustworthy, automated strategy for decoding novel chemical structures in data-scarce regimes.
Abstract
1D Nuclear Magnetic Resonance (NMR) spectra offer a rapid and accessible alternative to time-consuming 2D experiments, making them ideal for high-throughput structure elucidation. However, reconstructing molecular topology solely from 1D NMR spectral data remains a formidable combinatorial challenge. This difficulty arises from the loss of explicit atomic connectivity information, and existing data-driven models generate chemically invalid or hallucinated structures. Here, we present PINS (Physics-Informed NMR Structure elucidation model), a generative framework that explicitly bridges the gap between spectral data and molecular topology by enforcing multiphysical priors. By constraining the generative search space within strict physical laws, PINS effectively mitigates structural hallucinations and helps resolve structural ambiguities compared with end-to-end deep learning models. PINS achieves 100% chemical validity and surpasses the state-of-the-art (SOTA) baseline by a substantial margin of 21.6 percentage points (a 68.9% relative improvement). Furthermore, PINS demonstrates robust extrapolation capabilities on novel molecular scaffolds within the specific domain of challenging new psychoactive substances (NPS), achieving an 89.7% identification rate in forensic scenarios. Our findings underscore the necessity of physical constraints in generative AI, providing a trustworthy, automated strategy for decoding novel chemical structures in data-scarce regimes.
Accurate interpretation of one-dimensional proton nuclear magnetic resonance (1H NMR) spectra remains a rate-limiting step in molecular structure elucidation, particularly when signal overlap, strong spin coupling, and instrumental distortions mask key features. Existing automated approaches depend on computationally intensive and sensitive iterative quantum-mechanical fitting and still require expert oversight. Here we introduce MolDeTr, a chemistry-informed deep-learning framework derived from the detection-transformer architecture that unifies peak picking, multiplet identification, and extraction of chemical shifts, scalar coupling constants, relaxation-dependent decay times, and proton counts in a single-network pass. The method targets prototypical spin systems of single-component small molecules in 1D 1H NMR, with up to ten distinct groups of chemically equivalent spins (multiplets). MolDeTr is trained exclusively on synthetic spectra generated by spin-dynamics simulations and augmented with realistic experimental artifacts, enabling it to generalize to unseen compounds─including experimental spectra with overlapping and strongly coupled multiplets─without reference standards or prior spin-system knowledge. Unlike structure-conditioned shift-prediction or calculation models, e.g., density functional theory (DFT), that assume the molecular structure is known, MolDeTr addresses the spectrum-conditioned inverse problem and extracts spin-system parameters directly from measured 1D 1H NMR spectra, thereby substantially improving chemical-shift prediction precision by one to 2 orders of magnitude compared to existing structure-conditioned approaches. Benchmarking against a diverse experimental set of 1H NMR spectra of modestly sized small molecules, spanning 80 to 600 MHz base frequency, shows median absolute errors of 0.89 Hz for chemical shifts and 0.20 Hz for coupling constants, while absolute proton counts are predicted with 93.5% accuracy, outperforming state-of-the-art spectrum analysis software and experienced spectroscopists. By eliminating iterative fitting and expert intervention, MolDeTr offers a scalable route to fully automated spectral analysis, accelerating molecular discovery across the chemical sciences.
N. Schmid, Marc Wanner, G. Fischetti et al.· Analytical Chemistry· 1 citation
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.
Chengchun Liu, Zhiyuan Yan, Li Yuan et al.· 0 citations
X-ray absorption spectroscopy (XAS) is a key technique for probing local atomic environments, yet learning based modeling must bridge two heterogeneous modalities: 1D continuous spectra and 3D atomic structures. Existing approaches typically decouple forward spectrum prediction and inverse structure inference into separate regression tasks, hindering shared representation learning. Moreover, severe permutation ambiguity among identical atoms often limits inverse modeling to coarse structure descriptors rather than explicit 3D structure generation. In this work, we present Uni-XAS, a unified benchmark and learning framework that reframes bidirectional XAS modeling as a cross-modal alignment and conditional generation problem. We first propose XASLip, an alignment recipe coupling a physics-aware spectral encoder with an absorberaware manifold optimization strategy to resolve fine-grained intra-element coordination variations. Building upon this shared latent space, we formulate forward prediction as anchored absolute-spectrum generation via retrieval-augmented decoding, effectively preventing physical scale collapse and energy drift. For the inherently ill-posed inverse problem, we introduce Permutation-Rectified Flow Matching, which integrates type-wise optimal transport into a continuous generative flow to provide a principled solution to ligand permutation ambiguity without relying on heavy high-order equivariant architectures. Evaluated on a largescale standardized benchmark of 328,839 structure-spectrum pairs, Uni-XAS demonstrates strong performance in cross-modal retrieval, accurate absolute-spectrum prediction, and composition-conditional 3D structure generation, establishing a scalable, reproducible, and protocol-consistent foundation for multimodal learning and standardized evaluation in scientific spectroscopy.
Suyang Zhong, Yuhao Zhao, Boying Huang et al.· 0 citations
Biomolecular binder design for peptides and antibodies requires generating diverse candidates that satisfy stringent three-dimensional geometric constraints while enabling affinity-oriented exploration under strong structural priors. In current generative models, the effective search space for structurally feasible binders is severely constrained, as the complexity of biochemical interactions is not explicitly encoded into a semantically grounded representation of viable molecular manifolds. To address this challenge, we propose Pretrained Representation Induced Molecular gEneration (PRIME), a unified generative framework for three-dimensional binder design across peptides and antibodies. PRIME grounds stochastic generation on frozen large-scale pretrained structural representations, inheriting robust physical priors to ensure structural feasibility without training a manifold from scratch. However, defining a feasible space alone is insufficient for effective exploration. Under commonly used isotropic perturbations, chain topology is ignored, allowing local noise to propagate into global structural distortions. To enable controlled exploration within the feasible space, we introduce Semantics-Preserving Exploratory Sampling (SPES), which integrates Graph Laplacian Spectral Noise to respect chain connectivity and Conditional Freedom Modulation to dynamically balance exploration with fidelity. By aligning stochastic exploration with structural semantics, PRIME enables diversity-enhanced generation without sacrificing geometric validity under the reported structural metrics, improving the empirical exploration--fidelity trade-off. PRIME achieves state-of-the-art performance on unified peptide and antibody benchmarks, effectively reconciling geometric validity with functional optimization under computational proxy metrics. The source code is available at https://github.com/simplaj/PRIME.
Zhihua Tian, Jiale Zhou, Rubo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
Nuclear Magnetic Resonance (NMR) spectroscopy is the gold standard for molecular structure elucidation, yet interpreting complex spectra for unknown molecules remains a bottleneck reliant on human expertise. While artificial intelligence has advanced this field, current methods face a critical trade-off: database retrieval cannot identify novel scaffolds, while de novo molecular structure elucidation models operate as black boxes, lacking the atom-level interpretability required for rigorous scientific validation. Here, we present NMRAgent, an evidential reasoning agent powered by large language models (LLMs) that bridges this gap by integrating specialized spectral analysis tools with chemical knowledge graphs. Unlike previous approaches, NMRAgent mimics the deductive reasoning of human experts: it takes experimental NMR spectra and molecular formula as input, plans the elucidation process, proposes candidate structures, verifies peak-atom consistency, and refines misaligned substructure through formula-aware fragment optimization. Enabled by its evidential reasoning, NMRAgent outperforms state-of-the-art methods, improving top-1 accuracy by 46.5% and Tanimoto similarity by 0.502 on a scaffold-split benchmark with novel scaffolds in the test set. Besides, we demonstrate the agent's practical utility by elucidating the structures of two previously unknown natural products isolated from Hydrangea davidii and Vitex trifolia, and by correcting structural misassignments in established literature. By combining high-accuracy prediction with transparent and evidence-based reasoning, NMRAgent establishes a new paradigm for interpretable AI in analytical chemistry.
Zheng Fang, Yang Chen, Yusen Tan et al.· arXiv.org· 0 citations
Organoboron compounds are widely used across pharmaceuticals and materials science, where 11B NMR spectroscopy serves as a valuable tool for structural characterization. However, severe spectral line broadening induced by the quadrupolar nature of the boron nucleus often causes signal overlap, making it exceptionally difficult to experimentally resolve chemically inequivalent sites in complex multiboron architectures. While traditional density functional theory can resolve these ambiguities, it faces prohibitive computational bottlenecks, whereas data-driven alternatives remain constrained by the scarcity of high-quality data sets. Herein, we report a manually verified, solvent-annotated 11B NMR data set constructed via a large language model (LLM)-assisted workflow. Interpretable machine learning identifies a strong correlation between the BCUT2D_MRLOW descriptor and the boron hybridization. Integrating these ML-derived features as prior knowledge, we developed a prior-guided Graph Transformer for accurate atom-level chemical shift prediction. Notably, the model provides a form of virtual spectral resolution, enabling the discrimination of chemically inequivalent boron sites that are difficult to resolve experimentally. We further deploy the framework as an open-access Web tool to support the rapid structural analysis of organoboron compounds.
Penghui Li, Ben Gao, Shiyang Wang et al.· JACS Au· 0 citations