Back to feed
Open access

Application of vision transformers to protein-ligand affinity prediction

Jul 2026 · Scientific Reports · 0 citations

TL;DR

Despite challenges related to data sparsity and conformational variability, ViTs show strong performance and high robustness in structure-based affinity prediction tasks, underscore their effectiveness in learning spatial patterns and suggest broader applicability to related tasks, such as protein-protein or protein-nucleic acid interaction modeling.

Abstract

Predicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations. In this study, we investigate the application of vision transformers (ViTs) to the problem of affinity prediction. Unlike other neural networks used in this problem, the ViT framework can capture global, long-range interactions across the entire protein-ligand complex via self-attention, without relying on local receptive fields or predefined interaction cutoffs. We evaluate this advantage in representation of spatial information across two benchmark datasets, demonstrating competitive performance and, in some cases, surpassing state-of-the-art models. We study the model’s behavior using explainable AI (XAI) techniques, revealing that spatially proximal patches with similar attention scores cluster around biologically relevant regions, confirming the model’s ability to capture key interaction features. Furthermore, we show that data augmentation strategies can yield performance improvements, highlighting the potential for further enhancement. Despite challenges related to data sparsity and conformational variability, ViTs show strong performance and high robustness in structure-based affinity prediction tasks. Our findings underscore their effectiveness in learning spatial patterns and suggest broader applicability to related tasks, such as protein-protein or protein-nucleic acid interaction modeling.

Read PDF

Similar papers

Open access Jul 2026

A Preparation-Free Mixture-of-Experts Framework for Protein-Ligand Affinity Prediction

The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.

Huiming Bao, Shouliang Dong · 0 citations
Jul 2026

Hierarchical Graph Representation Learning From a Statistical Perspective for Generalizable and Interpretable Protein-Ligand Binding Affinity Prediction.

Protein-ligand binding affinity (PLA) prediction aims to guide rational drug design by estimating the strength of interaction. The effectiveness of the representation learning of protein and ligand is key to successful PLA prediction. To this end, attention mechanism, as a powerful architectural paradigm, has been introduced and gradually emerged as the prevailing approach. However, intuitively, the classical attention paradigm based on similarity does not fit the biological mechanisms relevant for binding. Worse still, the cooperative and antagonistic effects among multiple atoms are deliberately disregarded in the classical formulation of attention mechanisms. Consequently, the rigid transplantation of classical architectures substantially undermines the PLA prediction performance. To address these challenges, we employ a hierarchical statistical attention model (HISA). Specifically, HISA employs a statistical attention mechanism (SAM) based on non-similarity computation to fit the biological prior and perceive the relationship of multiple atoms. In addition, we optimize HISA by employing clustering, enabling hierarchical representations of biomolecules. Extensive experiments demonstrate that HISA achieves state-of-the-art performance on multiple PLA benchmarks while simultaneously exhibiting generalizability and interpretability.

Changming Yao, Shunfanyi Li, Shanghui Deng et al. · 0 citations
Jul 2026

A Scalable Structure-Aware Multimodal Architecture for Accurate Drug-Target Affinity Prediction.

Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.

Junlin Xu, Ye Yuan, Menglong Hu et al. · 0 citations
Open access Aug 2026

On the generalization and usability of cofolding models for GPCR drug discovery

The generalizability of co-folding models for protein–ligand structure prediction remains unclear. Here, we benchmark Boltz, a state-of-the-art co-folding model, using a curated set of ligand-bound human G protein-coupled receptors (GPCRs) from families unseen during training. We show that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data when tested with FEP +. We further show that physics‑based refinement of Boltz models can correct ligand poses to near‑experimental accuracy and rescue FEP+ performance to that of the native structure. These results highlight the strengths and limitations of co-folding methods and motivate a workflow that pairs them with physics-based refinement and validation before high-stakes decisions in drug discovery.

Lichirui Zhang, R. Friesner, Edward B. Miller et al. · 0 citations
Jan 2026

An algebraic graph neural network model for protein-ligand binding affinity prediction

An Algebraic Graph Neural Network model designed to encode molecular structures into a low-dimensional graph representation while preserving critical biochemical interactions is introduced, demonstrating superior performance in binding affinity prediction compared to state-of-the-art scoring functions.

Augustine Ouru, Xi Chen, Cameron Yeagle et al. · 0 citations
Jul 2026

Bridging between Structure-Based and Data-Driven Affinity Prediction.

Protein-ligand binding affinity prediction is central in computer-aided drug design, and a wide range of physics-based, empirical, and machine learning (ML) tools have been developed for this purpose. Scientists are often tasked with choosing an optimal tool for a given discovery problem, which can be difficult. Physical models typically perform best in the absence of target-specific data but are often outperformed by ML models as the amount of data on a given target grows. Here, we introduce a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model. We apply this framework to combine docking scores with predictions from a Gaussian Process model trained on binding affinities from two industrial data sets. We show that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired. We also show that structure-based methods, being insensitive to the distribution of binders across chemical space, can improve generalizability to new chemistries and increase the hit rate in active learning screens where the binder density varies among scaffolds in a virtual library. Thus, integrating predictions from multiple tools not only optimizes the use of limited experimental data but also ensures more robust performance compared to reliance on a single model.

Ažbeta Kubincová, David L. Mobley · 1 citation