Skip to content
Open access

On the generalization and usability of cofolding models for GPCR drug discovery

Aug 2026 · npj Drug Discovery · Vol 3 · 0 citations · 57 references
Medicine

TL;DR

Boltz is benchmarked using a curated set of ligand-bound human G protein-coupled receptors from families unseen during training, showing that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data when tested with FEP+.

Abstract

The generalizability of co-folding models for protein–ligand structure prediction remains unclear. Here, we benchmark Boltz, a state-of-the-art co-folding model, using a curated set of ligand-bound human G protein-coupled receptors (GPCRs) from families unseen during training. We show that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data when tested with FEP +. We further show that physics‑based refinement of Boltz models can correct ligand poses to near‑experimental accuracy and rescue FEP+ performance to that of the native structure. These results highlight the strengths and limitations of co-folding methods and motivate a workflow that pairs them with physics-based refinement and validation before high-stakes decisions in drug discovery.

Read PDF

Similar papers

Open access Aug 2026

Boltz-Perturb: Improving Diversity and Accuracy in Protein-Ligand Co-Folding through Training-Free Conditioning Perturbation

Protein-ligand co-folding models hold promise in structure-based drug discovery and small molecule interaction prediction, but often fail in predicting correct small molecule binding poses. We present Boltz-Perturb, a framework for addressing this through perturbing model conditioning signals during model inference, and show that such perturbations improve correct ligand binding mode predictions. We first show with true-coordinate injection experiments that the model’s learned energy landscape contains correct binding-mode basins, allowing us to reframe the problem as one of sampling deficiency. We then introduce two inference-time perturbation strategies, Token Bias Perturbation (TBP) and Token Conditioning Perturbation (TCP), which increase exploration of alternative binding poses. Across diverse protein–ligand systems, TCP improves top-20 oracle success rates by 2.6 to 7.8 fold. Boltz-Perturb attains higher oracle success rates compared to the Boltz-2 high diffusion temperature variant while requiring over 75% less compute. To our knowledge, this is the first systematic perturbation analysis of a co-folding architecture for small-molecule binding mode diversity. We demonstrate that inference-time perturbations can unlock latent structural diversity in generative co-folding models and improve protein-ligand predictions without costly retraining.

Hyeyun Jung, BoRam Lee, Alan C. Cheng · 0 citations
Open access Jul 2026

A Preparation-Free Mixture-of-Experts Framework for Protein-Ligand Affinity Prediction

The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.

Huiming Bao, Shouliang Dong · 0 citations
Preprint Jul 2026

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

Structure-prediction networks built on co-evolutionary statistics have transformed protein-based drug discovery, yet their accuracy does not extend to peptide therapeutics--an increasingly important modality defined by non-canonical residues, macrocyclization, and complex topologies. We introduce Vilya-2, a diffusion transformer that extends the all-atom representation of Vilya-1 from modeling individual molecules to modeling their interactions with protein targets. This all-atom representation enables transfer learning between different molecular types, and delivers highly accurate structural modeling of peptides across sizes, classes, and compositions bound to therapeutically relevant targets. By generating diverse structural ensembles and ranking them with calibrated confidence, Vilya-2 recovers 59.1% of peptide interfaces to sub-2 {\AA} backbone RMSD, far exceeding the performance of a representative co-folding model even when that model is given the bound receptor as a template. In addition, Vilya-2 is state-of-the-art at small-molecule docking, and generalizes to novel protein-small molecule complexes unlike those seen in training. It also generalizes to modeling molecular conformations of diverse macrocycles and disulfide-stapled miniproteins several-fold larger than any molecule seen in training. Finally, Vilya-2 can be used as a foundation model, and fine-tuned to enrich for active compounds in hit-to-lead campaigns. By unifying predictive accuracy with broad generalizability across chemical space, Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Vilya Research Pascal Sturmfels, Naozumi Hiranuma, M. Salem et al. · 0 citations
Jul 2026

Bridging between Structure-Based and Data-Driven Affinity Prediction.

This work introduces a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model and shows that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired.

Ažbeta Kubincová, David L. Mobley · 1 citation
Open access Jul 2026

Capabilities, specificity gaps and training-data dependence of AlphaFold3 across diverse application areas

It is found that, while AF3 can perform well in favourable settings, this performance is uneven across applications and its predictions and use of confidence metrics will depend strongly on the specific application area and must be interpreted with respect to training-set overlap.

O. Follonier, Yan Liu, Pablo Campomanes et al. · 1 citation
Open access Jul 2026

Application of vision transformers to protein-ligand affinity prediction

Despite challenges related to data sparsity and conformational variability, ViTs show strong performance and high robustness in structure-based affinity prediction tasks, underscore their effectiveness in learning spatial patterns and suggest broader applicability to related tasks, such as protein-protein or protein-nucleic acid interaction modeling.

Jakub Poziemski, Paweł Siedlecki · 0 citations