Skip to content
Open access

ConfRetro: a 3D-aware template-free method for enhancing retrosynthesis via molecular conformer information

Jan 2025 · Bioinform. · Vol 42 · 1 citation · 52 references
Computer Science Medicine

Abstract

Abstract Motivation Retrosynthesis plays a crucial role in organic synthesis and drug discovery, focusing on identifying a set of reactants capable of synthesizing a target product molecule. Although the existing approaches have shown promising results, they do not fully exploit 3D conformer information and molecular spatial structure, which can hinder stereochemically consistent and chemically plausible predictions. Results To tackle this problem, we propose ConfRetro, a Transformer-based template-free method that integrates molecular conformer information and spatial structure. We devise an Atom-align Fusion module to combine 3D positional information at the model input stage, ensuring alignment between atom tokens and corresponding 3D representations. Furthermore, we design a Distance-weighted Attention mechanism to guide self-attention, constraining the receptive field of model and emphasizing chemically relevant atom pairs in 3D space. Experiments conducted on the USPTO-50K and USPTO-FULL datasets demonstrate that ConfRetro significantly outperforms existing template-free approaches, achieving a new state-of-the-art performance. Case studies further highlight its capability to predict accurate and chemically plausible reactants, even for target molecules with intricate structures. Moreover, when plugged into a standard retrosynthetic search, ConfRetro recovers feasible synthetic routes for multiple representative drug molecules (e.g. Camptothecin). Availability and implementation ConfRetro is available at https://github.com/Jesse-zjx/ConfRetro. Archival snapshot is available at https://doi.org/10.5281/zenodo.20785018.

Read PDF

Similar papers

Open access Jul 2026

RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction.

Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progress, they are designed to autonomously learn reaction patterns from extensive retrosynthesis data sets with limited explicit integration of established chemical knowledge as priors. To address this limitation, we introduce RetroMPA, a molecular property-aware, posthoc enhancement module that injects chemical knowledge into the retrosynthesis pipeline. Rather than functioning as an independent, standalone SMILES sequence generator from scratch, RetroMPA is conceptualized as a broadly applicable, model-agnostic chemical filter designed to recalibrate and optimize the predictive pathways of various existing algorithms. This plug-and-play framework can be seamlessly integrated with a range of existing data-driven retrosynthesis methods, enhancing model outputs without necessitating any modifications to the original model architecture or requiring resource-intensive, model-specific retraining procedures. By operating at the molecular level and leveraging a property-aware latent embedding space, RetroMPA consistently improves top-1 accuracy across eight representative retrosynthesis models by an average of 5.50% on USPTO-50K. Furthermore, we demonstrate its scalability by validating its performance on the large-scale USPTO-Full data set, achieving an average improvement of about 2.03% across both template-based and template-free architectures. In addition, wet-lab experiments provide preliminary support for the practical utility of the framework. These syntheses confirmed viable, previously unreported substrate combinations for established, classic reaction paradigms─specifically, the Suzuki-Miyaura coupling, the Bucherer reaction, and the Friedel-Crafts acylation, thereby suggesting that RetroMPA can operate beyond mere data fitting. The code is open-sourced at https://github.com/MengzhouLu/RetroMPA.

Mianzhi Liu, Fan Xiao, Zhi-Qiang Yu et al. · 0 citations

Bridging the Biophysical Gap: Holistic Environmental Awareness for 3D Linker Design

LinkerBridge is proposed, an equivariant framework that unifies biochemical semantics with physical constraints through two innovations: a Contextual Interaction-Aware Representation module that internalizes pre-existing bio-chemical semantics and a Differentiable Physical Guidance mechanism derived from Van der Waals potentials to steer generation away from collision zones.

Mengwei Sun, Chengwei Ai, Xiaoyi Liu et al. · 0 citations

An Explicit Interaction-Prompted Diffusion Framework for High-Fidelity 3D Molecular Generation.

Current structure-based drug design generative models often struggle to faithfully recapitulate genuine ligand-protein binding interactions. Instead, under the coupling of implicit learning architectures and biased training data, they tend to learn spurious statistical correlations. To address this, we propose EIP-Diff (Explicit Interaction-Prompted Diffusion), an architecture featuring a novel explicit interaction-prompt embedding mechanism that is better suited for real-world target-specific drug design. This architecture replaces biased implicit learning with explicit, residue-level biological guidance, thereby promoting more fine-grained geometric fidelity and more precise interaction-aware conditioning. To fully realize the capabilities of EIP-Diff and provide a reliable basis for performance evaluation, we further constructed CrystalData set, which provides higher-fidelity and less-biased structural supervision than existing data sets. This explicit architecture markedly improves distribution consistency: even when trained on the crossdocked data set, EIP-Diff achieves the highest alignment with authentic pharmacological distributions among evaluated models. Training on CrystalData set further enhances this alignment and improves 3D geometric accuracy, while retaining strong controllability, high chemical space coverage, and near-perfect uniqueness. In addition, target-based validation on KAT6A and YTHDC1 confirmed that EIP-Diff accurately recapitulates native-like binding modes. Furthermore, in a real-world drug design task against IDO1, we successfully designed a novel lead compound with nanomolar potency (IC50 = 0.31 nM). These results demonstrate that the EIP-Diff architecture can explicitly leverage experimentally derived structural data and biologically meaningful interaction information for target-specific molecular generation, thereby enabling its effective application to real-world structure-based drug design.

Huabin Du, Mingyang Wang, M. Luo et al. · 0 citations
#machine learning Preprint Aug 2026

Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation

LiFT, a language-informed cross-modal framework built on Flow Matching for trend-guided 3D molecular generation across both de novo design and scaffold hopping, and suggests that language-derived chemical priors provide effective trend-level guidance for 3D molecular generation.

Tian-Yu Gao, Zhi-Kai Su, Jia-Shu Li et al. · 0 citations
Jul 2026

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

Vilya-1 is introduced, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability.

Vilya Research Pascal Sturmfels, M. Salem, Naozumi Hiranuma et al. · 1 citation
Book Open access Aug 2026

PRIME: A Pretrained Representation-Induced Model for 3D Molecules in De Novo Binder Design

Biomolecular binder design for peptides and antibodies requires generating diverse candidates that satisfy stringent three-dimensional geometric constraints while enabling affinity-oriented exploration under strong structural priors. In current generative models, the effective search space for structurally feasible binders is severely constrained, as the complexity of biochemical interactions is not explicitly encoded into a semantically grounded representation of viable molecular manifolds. To address this challenge, we propose Pretrained Representation Induced Molecular gEneration (PRIME), a unified generative framework for three-dimensional binder design across peptides and antibodies. PRIME grounds stochastic generation on frozen large-scale pretrained structural representations, inheriting robust physical priors to ensure structural feasibility without training a manifold from scratch. However, defining a feasible space alone is insufficient for effective exploration. Under commonly used isotropic perturbations, chain topology is ignored, allowing local noise to propagate into global structural distortions. To enable controlled exploration within the feasible space, we introduce Semantics-Preserving Exploratory Sampling (SPES), which integrates Graph Laplacian Spectral Noise to respect chain connectivity and Conditional Freedom Modulation to dynamically balance exploration with fidelity. By aligning stochastic exploration with structural semantics, PRIME enables diversity-enhanced generation without sacrificing geometric validity under the reported structural metrics, improving the empirical exploration--fidelity trade-off. PRIME achieves state-of-the-art performance on unified peptide and antibody benchmarks, effectively reconciling geometric validity with functional optimization under computational proxy metrics. The source code is available at https://github.com/simplaj/PRIME.

Zhihua Tian, Jiale Zhou, Rubo Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.