Skip to content
Open access

LEN-Seek: Fast and scalable ligand binding-site similarity search in the latent space of an SE(3)-invariant graph VAE

Aug 2026 · bioRxiv · 0 citations · 63 references
Biology

Abstract

Motivation Ligand binding-site similarity search is a crucial step in drug discovery that reduces the conformational search space for docking and other downstream tasks by comparing a target protein against experimentally identified binding sites. Existing methods rely on either direct structural alignment or lossy compression of structural information, producing a trade-off between scalability and precision. Results We propose LEN-Seek, a ligand binding-site search method based on a graph neural network (GNN)-driven variational autoencoder (VAE) that encodes the 3D structural and physicochemical context of a binding site into a probabilistic latent space, enabling similarity search within a low-dimensional vector space. A binding site is modeled as a graph of amino acid residues, with node features adopted from the protein language model, Ankh, and edges encoded as SE(3)-invariant (roto-translational invariant) geometric relationships, thereby avoiding expensive data augmentation or SE(3)-equivariant models. Compared to ProBiS, the purely geometric graph-clique based method, LEN-Seek successfully retrieves a substantial portion of similar binding sites with a roughly 3,400-fold lower per-comparison cost, demonstrating its potential as a scalable approach to template-based ligand binding-site search in large-scale protein structure databases. Supplementary information Supplementary data are available at Bioinformatics online.

Read PDF

Similar papers

Book Open access Jun 2026

DeepMetal: A Hierarchical Coarse-to-Fine Framework for Metal-Binding Site Prediction via Protein Language Models and SE(3)-Equivariant Graph Neural Networks

Metal ions serve as essential cofactors in approximately 30%–40% of proteins, and accurate recognition of their binding sites is central to function annotation, drug discovery, and metalloenzyme design. Existing predictors often operate at residue level, generate many false positives, or depend strongly on high-quality bound structures. We present DeepMetal, a hierarchical coarse-to-fine framework that combines ESM-2 residue screening, biophysics-constrained Dynamic Center-Iterative Clustering (DCIC), and a site-level SE(3)-equivariant graph neural network for candidate-site validation and metal typing. On a non-redundant BioLiP2-derived benchmark, DeepMetal achieves an AUROC of 0.775 and an F2 score of 0.533 for transition-metal site localization, outperforming representative baselines MetalNet2 and PinMyMetal under the same intersectional evaluation setting. These results show that sequence-driven screening, geometry-aware assembly, and equivariant validation can jointly improve practical metal-binding site prediction from predicted protein structures.

Bing Liu, Yangfan Xu, Yunpeng Wang et al. · 0 citations
Jul 2026

Native Contact Ratio as a Topological Metric for Machine Learning Based Molecular Docking

Accurate identification of near-native ligand binding poses is a central challenge in structure-based drug design. From a physical point of view, the successful construction of a protein–ligand complex structure is dependent on whether protein and ligand can form enough atomically pairwise interactions that result in a global energy minimum. In this work, we report a machine learning scoring strategy for protein–ligand screening which explicitly considers the Native Contact Ratio (NCR), a topology inspired metric that quantifies the preservation of protein–ligand interfacial contacts as well as interaction energy. This physics-awared supervision strategy provides a simple but efficient gradient field that faithfully reflects the complicated protein energy landscape than conventional 3D coordinate-based objectives. Building on this principle, we present DeepNCR, an energy-informed Transformer framework that encodes approximate Coulombic and dispersive interaction potentials across the protein–ligand binding interface. Furthermore, we introduce a feature pruning step that compresses the interaction tensor from 1470 to 868 dimensions, further improving signal-to-noise ratio and directing model attention toward the interaction motifs critical for binding specificity. The model optimizes topological objectives and at inference drives pose refinement through a differentiable hybrid gradient field integrating predicted NCR and AutoDock Vina energetics. Extensive evaluation on the CASF-2016 benchmark and the 3D-DISCO cross-docking data set demonstrates consistently high performance: a Top-1 docking success rate of 94.7%, a 1% Enrichment Factor of 21.21 in virtual screening, and a Top-1 cross-docking success rate of 34.8%. Mechanistic analysis reveals that NCR-guided optimization enables decoy escaping from local energy minima and drives the recovery of disrupted native interactions, confirming that NCR captures the physical determinants of binding rather than mere geometric proximity.

Zhenqiang Zhang, Zhihao Wang, Yang Liu et al. · 0 citations
Aug 2026

NextTopDocker: A Large-Scale Docking-Power Benchmark Reveals Limitations of Current End-to-End Machine-Learning Docking and the Strength of Hybrid Rescoring

Predicting three-dimensional binding orientations of drug-like molecules remains challenging in structure-based drug design. Despite methodological advances, docking performance is often assessed on small and outdated benchmarks. We present “NextTopDocker,” a large, up-to-date, open-access data set for docking-power assessment comprising 14,038 training and 5201 test entries across 3173 unique protein targets, constructed from the Protein Data Bank. Developed with open-source tools, it includes crystallographic structures, Smina-generated docking poses, and ligand-similarity-aware training subsets. We benchmarked four state-of-the-art machine-learning (ML) docking frameworks (DeepDock, Interformer, SurfDock, and Uni-Mol Docking v.2) against classical (Smina) and hybrid baselines (GNINA 1.3 and logistic regression using Smina and GNINA 1.3 scores). Interformer alone matched the docking power of logistic regression on Smina poses, while the others showed dependence on downstream physics-based correction. Most raw ML-generated poses displayed steric clashes and/or implausible geometries, highlighting the need for physics-informed constraints in autonomous docking. “NextTopDocker” is available at https://github.com/caominhtr/NextTopDocker and https://zenodo.org/records/17492994.

Cao-Minh Truong, Pedro J. Ballester, O. Taboureau et al. · 0 citations
Preprint Aug 2026

RAVEN: Frozen Random Graph Reservoirs with Physics-Informed Interaction Fingerprints for Protein-Ligand Binding Affinity Prediction

The results indicate that frozen multi-view graph representations, explicit physicochemical statistics, and heterogeneous model fusion provide a robust and flexible framework for protein-ligand binding-affinity prediction.

Qingyan Zou, Jiaye Huang, Hangbo Xie et al. · 0 citations
Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations