Skip to content
Open access

Structure-free, site-resolved contrastive learning extends small-molecule discovery beyond the reach of structure-based modeling

Aug 2026 · bioRxiv · 1 citation · 69 references
Biology

TL;DR

Ptarmigan-1 is presented, a contrastive model that co-embeds the residues of a protein with candidate small molecules in a shared latent space, from sequence and two-dimensional chemistry alone, and without ever constructing a pose.

Abstract

Virtual screening asks which molecules, among an enormous space of drug-like chemistry, are worth synthesizing and testing against a protein target. Most modern methods answer this question by building and scoring an explicit three-dimensional pose through molecular docking, or the co-folding models that now approach experimental accuracy. Building these poses presumes a well-defined pocket. However, the non-orthosteric, cryptic, and intrinsically disordered sites where unexplored ligandability lies offer no such pocket to explore. Here we present Ptarmigan-1, a contrastive model that co-embeds the residues of a protein with candidate small molecules in a shared latent space, from sequence and two-dimensional chemistry alone, and without ever constructing a pose. Freed from the requirement for protein structures, Ptarmigan-1 trains directly on chemoproteomic and bioactivity data of mixed resolution, scores a compound in ten milliseconds rather than the tens of seconds a co-folding model demands, and resolves each prediction to the residues a compound engages. On well-folded, orthosteric targets it performs comparably to a collection of co-folding and docking models, and on covalent, cryptic, and disordered sites it matches or exceeds them. Ptarmigan-1 localizes reversible and covalent inhibitors to the pockets they engage, even for targets withheld from training, and screens the entire human proteome against a library of 3.4 billion compounds in under a day. By decoupling molecular recognition from structure, Ptarmigan-1 recasts virtual screening as a nearest-neighbor query in a rich latent space shared by protein residues and the compounds that bind them.

Read PDF

Similar papers

Jul 2026

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Pascal Sturmfels, Naozumi Hiranuma, Milad Salem et al. · 0 citations
Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations
Jul 2026

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

Vilya-1 is introduced, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability.

Vilya Research Pascal Sturmfels, M. Salem, Naozumi Hiranuma et al. · 1 citation
Aug 2026

Evaluating BioEmu-Generated Kinase Ensembles Reveals Structure Selection as the Virtual Screening Bottleneck.

It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.

Jaeoh Shin, K. Joo, Jejoong Yoo · 0 citations
2026

Open-Sourced In Silico Drug Screening.

This chapter describes a structure-based computational approach to perform high-throughput ligand screens of chemical libraries using open-source software programs and illustrates this workflow with the enzymatic molecular target NAD(P)H:quinone oxidoreductase1 (NQO1), which is overexpressed in a number of human solid tumors.

Audrey G. Fikes, Melissa C. Srougi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.