On the generalization and usability of cofolding models for GPCR drug discovery
Abstract
The generalizability of co-folding models for protein–ligand structure prediction remains unclear. Here, we benchmark Boltz, a state-of-the-art co-folding model, using a curated set of ligand-bound human G protein-coupled receptors (GPCRs) from families unseen during training. We show that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data when tested with FEP +. We further show that physics‑based refinement of Boltz models can correct ligand poses to near‑experimental accuracy and rescue FEP+ performance to that of the native structure. These results highlight the strengths and limitations of co-folding methods and motivate a workflow that pairs them with physics-based refinement and validation before high-stakes decisions in drug discovery.