Skip to content

Trustworthy Molecular AI: Multi-dimensional Validation of LLM Generated Chemical Descriptions

2026 · International Conference on Conceptual Structures · pp. 337-345 · 0 citations · 19 references
Computer Science

TL;DR

This work sets foundations for trustworthy AI in chemistry, providing quality assurance infrastructure necessary for deploying LLMs in high-stakes applications that carry material consequences, by integrating grammatical fluency and lexical validity checking, RDKit -based computational chemistry validation, and DrugBank verification.

View source

Similar papers

Preprint Aug 2026

Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models

Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity. Existing validation approaches can identify non-computation-ready structures, but they often rely on heuristic rules, license requirement, or offer limited interpretability. Here, we show that large language models (LLMs) can serve as interpretable validators of MOF structures when crystallographic information is transformed into chemically meaningful text. By benchmarking nine descriptors, we find that successful LLM-based validation depends not on the amount of structural information alone, but on whether local coordination, framework connectivity, and chemical context are organized into a linguistically learnable representation. Fine-tuned LLMs using specialized descriptors (mof2text) achieve performance comparable to graph-based models in identifying unreasonable MOFs. Importantly, these models extend beyond black-box classification by generating diagnostic rationales for likely error sources, including abnormal bonding, connectivity, and charge states, as well as error-category predictions for annotated datasets. This work establishes chemically informed textualization as the key step that transforms LLMs from generic text models into practical and explainable tools for curating MOF databases.

G. Zhao, Xiao-Yan Li · 0 citations
Preprint Aug 2026

Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

This is the first method to expose GNN-derived attributions to an LLM as evidence for property prediction, and achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task.

Junwoo Park, Minyoung Shin, C. Lee et al. · 0 citations
Preprint Aug 2026

Compiling Chemical Knowledge into Executable Descriptors for Materials Prediction

CRISP is introduced, a large language model-assisted framework that treats representation construction as a rule-space exploration and compilation problem: it repeatedly samples target-relevant chemical rules without access to structures, labels or data splits, consolidates related concepts, and compiles each into an executable scalar descriptor supplied to a conventional learner.

Jaehwan Choi, Kunik Jang, Seongmin Kim et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model Synergy

CoMPASS is presented, a retrieval-calibrated framework for small-large model collaboration that retains a graph attention network as the predictive anchor, retrieves locally relevant training molecules, provides attention-grounded evidence to an LLM, and converts its proposal into a bounded correction through an agreement-aware gate.

Wen-Tao Li, Jiang-Jie Qiu, Yi-Jun Li et al. · 0 citations
Preprint Aug 2026

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

Onepot-Bench 0 is introduced, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution and probes basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.

Brandon Wang, Andrei S. Tyrin, Daniil A. Boiko · 0 citations
Jul 2026

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).

Chuhan Shi, Xiaoquan Ren, Sicheng Song et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.