Skip to content
Open access

StructureSAFE: A structure-aware chemical language model for unified hit identification and lead optimization

Jul 2026 · bioRxiv · 0 citations · 16 references
Biology

Abstract

Structure-based generative models (SBGMs) hold great promises for accelerating drug discovery by enabling target-aware molecular design. However, existing approaches face fundamental challenges: three-dimensional graph-based models can explicitly incorporate protein structural information but often generate chemically implausible molecules due to limited training data, while chemical language models (CLMs) produce chemically plausible molecules but struggle to effectively leverage three-dimensional structural information for structure-conditioned generation and hard to incorporate lead optimization functionality due to the nature of SMILES string. Here, we present StructureSAFE, a structure-aware chemical language model that resolves this trade-off by integrating protein structural and evolutionary encoders with the SAFE molecular representation via pretraining and finetuning training scheme, enabling both de novo hit identification and a comprehensive suite of lead optimization subtasks within a unified framework. Comprehensive benchmarking on the MolGenBench dataset demonstrates that StructureSAFE achieves state-of-the-art (SOTA) performance across multiple metrics, with particularly pronounced improvements in chemical plausibility relative to graph-based models lacking pretraining. Evaluation on a rigorously constructed held-out test set further confirms its ability to generate drug-like, synthetically accessible molecules with competitive predicted binding affinities for previously unseen targets on both hit identification and lead optimization setting. In silico case studies across four therapeutically relevant targets validate its capacity to generate chemically plausible molecules that recapitulate key binding interactions of known high-affinity ligands while proposing novel interactions for potential better affinity and exploring previously unknown regions of chemical space. Taking together, StructureSAFE represents a versatile and practical tool to provide high-quality candidate molecules for augmenting medicinal chemistry workflows in both hit identification and lead optimization campaigns.

Read PDF

Similar papers

Preprint Jul 2026

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

A clear pattern is revealed in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

Thomas MacDougall, Maksim Kuznetsov, Roman Schutski et al. · 0 citations
Jul 2026

A Scalable Structure-Aware Multimodal Architecture for Accurate Drug-Target Affinity Prediction.

Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.

Junlin Xu, Ye Yuan, Menglong Hu et al. · 0 citations
Open access Jul 2026

ScrambleBench: a workflow for comparative assessment of structure-based de novo generative models.

ScrambleBench provides a holistic medicinal chemistry-oriented framework that identifies methodological strengths, limitations, and opportunities for future model development and highlights the importance of evaluating chemical diversity explicitly and using the recently proposed metrics such as Hamiltonian Diversity (HamDiv) which assess both quantity and dissimilarity of a molecular set.

Veincent Yap, Pan Xu, Frankie S. Mak et al. · 0 citations
Jun 2025

READ: A Retrieval-Alignment Diffusion Framework for Structure-based Drug Design.

Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities among protein-ligand complexes. This fragmented perspective constrains their ability to capture the fundamental principles governing molecular recognition and binding specificity. Moreover, the limited availability of high-quality experimental data further hampers model generalization and real-world applicability. To address these challenges, we present READ, a retrieval-alignment molecular generation framework that conditions the generative process on small molecules targeting homologous proteins. Retrieved ligands are aligned with a diffusion model across multiple representational spaces and integrated as conditional guidance throughout successive stages of generation. Under a standardized docking-based evaluation protocol, READ achieves consistently strong performance against state-of-the-art SBDD methods. More importantly, it introduces a retrieval-alignment paradigm for structure-based molecular generation, offering a practical framework for early-stage computational hit generation while leaving prospective experimental validation as future work.

Dong Xu, Zhangfan Yang, Junchuang Cai et al. · 1 citation
Open access Jul 2026

FragBERTa: a fragment-aware molecular representation model with sequential attachment-based fragment embeddings

FragBERTa is introduced, a molecular fragment-aware transformer-based representation language model pretrained using masked language modeling on Sequential Attachment-based Fragment Embedding (SAFE) representations, suggesting that fragment-based string representations offer advantages over atom-level representations for scaffold-sensitive and interaction-driven tasks.

Neerav Kaushal, Ajay Mnv Penmatsa · 0 citations
Preprint Jul 2026

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

Macrocyclic peptides are an increasingly important therapeutic modality, but existing computational methods for modeling their structures and properties are limited in scope and do not generalize well across the synthetically accessible chemical space. In this work, we introduce Vilya-1, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability. Vilya-1 operates on a uniform all-atom representation and is trained on heterogeneous structural datasets spanning diverse topologies and chemical classes. Across a broad set of macrocycles composed of canonical and non-canonical residues, Vilya-1 substantially improves geometric accuracy relative to physics-based methods, co-folding networks, and deep-learning conformer generators, while maintaining broad chemical coverage that extends to small molecules. Vilya-1 also supports generative applications, enabling the design of novel macrocycles with tailored chemical, structural, and property profiles. Together, these capabilities establish Vilya-1 as a foundation model for accelerating the development of next-generation macrocycle therapeutics.

Vilya Research Pascal Sturmfels, M. Salem, Naozumi Hiranuma et al. · 1 citation