The results suggest that coupling generation with executable verification and feedback-guided refinement is an effective way to improve text-to-molecule generation.
Abstract
Text-to-molecule generation is typically formulated as a one-shot sequence generation problem, where a model directly maps target descriptions to molecular representations. However, molecular descriptions often contain informative structural constraints, and violating such constraints can change the molecular identity. This makes chemical verification and error correction important but underexplored. To fill this gap, we propose MolGVR, a chemistry-grounded Generator--Verifier--Refiner framework. The Generator infers structural evidence and generates candidate molecules. The Verifier addresses the lack of chemical validation by converting descriptions into chemical constraints and checking candidates against them. The Refiner addresses generation failures by revising candidates rejected by the Verifier. Experiments on ChEBI-20 and PCDes show that MolGVR improves exact-match performance. These results suggest that coupling generation with executable verification and feedback-guided refinement is an effective way to improve text-to-molecule generation.
Text-guided molecule generation enables controlled molecular design from natural language descriptions and has broad applications in areas such as drug discovery. While recent methods have demonstrated promising capability in generating molecules that align well with textual descriptions, they often overlook the structural properties of the generated graphs. As a result, these approaches struggle to simultaneously ensure consistency with the input text and high structural quality of the generated molecules. In this paper, we propose a text-guided molecular graph generation framework that leverages the structural modeling power of graph diffusion models to achieve both strong alignment with textual descriptions and high-quality molecular structures. However, accomplishing this goal involves several key challenges: 1) how to align graph diffusion models with natural language instructions in order to generate molecular graphs with expected relational semantics from text, 2) how to directly optimize the quality of the generated molecular graphs without sacrificing fine-grained alignment with text-specific details. To tackle these challenges, we introduce Text-guided Conditional Discrete Graph Diffusion (TDGD), a discrete diffusion-based framework for generating molecular graphs from natural language descriptions. Our model incorporates a structure-aware cross-attention mechanism that aligns textual semantics with molecular structures by capturing relational semantics between textual descriptions and molecular structures. In addition, we propose a molecule structure consistency loss that explicitly enforces structural coherence during generation, leading to higher-quality and more consistent molecular graphs. Extensive experiments on ChEBI-20 and L+M-24 datasets demonstrate the effectiveness of our proposed TDGD model.
Yang Yao, Xin Wang, Yaofei Wu et al.· Proceedings of the 32nd ACM...· 0 citations
This work takes PROTAC linker design as a representative case where data scarcity, multi-constraint satisfaction, and interpretability requirements simultaneously hold, and provides quantitative evidence of distributional mismatch between PROTAC linkers and general small-molecule linkers.
Mol-CADiff is introduced, a diffusion-based framework that uses causal attention mechanisms for text-conditional molecular generation and enhances dependency modeling both within and across modalities, enabling precise control over the generation process.
Crystal generators can now propose periodic structures, but their control interfaces remain poorly matched to the mixed descriptors used in materials design. Text provides a compact way to combine composition, symmetry, prototype and property cues, yet it has not been clear whether such information can steer flow-based crystal generation. Here we introduce TFMat, a text-conditioned flow-matching framework that uses structured materials language as a semantic prior for a CrystalFlow generator. Across Perov-5, Carbon-24 and MP-20 crystal structure prediction benchmarks, TFMat improves one-candidate match rates over CrystalFlow and reaches a 92.04% MP-20 match rate with 20 candidates; in de novo generation, it improves element-count and density distribution alignment while retaining coarse property consistency in composition-selected outputs. These results position structured text as an inspectable control layer for translating human-readable materials intent into candidate crystals for downstream simulation and validation.
Recently, Large Language Models (LLMs) have become the dominant paradigm for molecular editing due to their strong generalization capabilities across diverse tasks. However, treating molecules as 1D text strings (SMILES) introduces significant challenges in controllability and structural validity. In this paper, we question whether sequence generation is truly optimal for this topological task. We propose UniEdit, a Unified Graph-based Mixture-of-Experts (MoE) Molecular Editing model that offers a robust alternative to LLMs. Diverging from the generative approach, UniEdit reformulates molecular editing as a hierarchical node-level classification task. By predicting discrete edit actions (e.g., Add, Remove, Replace) directly on the graph, our model ensures topological precision by design. To handle conflicting multi-objective constraints within a single framework, we incorporate a Mixture-of-Experts architecture that dynamically routes tasks to specialized components. Extensive experiments across 28 diverse tasks demonstrate that UniEdit significantly outperforms sequence-based baselines. Furthermore, a preliminary scaling study reveals that our graph-based approach benefits consistently from increased model capacity, suggesting a scalable path toward general-purpose molecular editing. Our code is available at https://github.com/jiajunyu1999/GraphEditing.
Jiajun Yu, Zhihao Wu, Yizhen Zheng et al.· Proceedings of the 32nd ACM...· 0 citations
Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity. Existing validation approaches can identify non-computation-ready structures, but they often rely on heuristic rules, license requirement, or offer limited interpretability. Here, we show that large language models (LLMs) can serve as interpretable validators of MOF structures when crystallographic information is transformed into chemically meaningful text. By benchmarking nine descriptors, we find that successful LLM-based validation depends not on the amount of structural information alone, but on whether local coordination, framework connectivity, and chemical context are organized into a linguistically learnable representation. Fine-tuned LLMs using specialized descriptors (mof2text) achieve performance comparable to graph-based models in identifying unreasonable MOFs. Importantly, these models extend beyond black-box classification by generating diagnostic rationales for likely error sources, including abnormal bonding, connectivity, and charge states, as well as error-category predictions for annotated datasets. This work establishes chemically informed textualization as the key step that transforms LLMs from generic text models into practical and explainable tools for curating MOF databases.
G. Zhao, Xiao-Yan Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.