Skip to content
#data science Dataset Open access

A selectivity-direction benchmark for co-folding models

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

A matched-pair benchmark that asks whether a co-folding model can tell which of two related targets a substitution moves toward. Each item is two close analogues measured against two related targets in one report with one readout; only the sign of the change in selectivity is scored. The primary set holds 7,656 decidable changes over 5,784 molecule pairs from 335 reports across seven target pairs (thrombin/factor Xa, JAK1/JAK2, PI3Kalpha/PI3Kdelta, COX-1/COX-2, MAO-A/MAO-B, AChE/BuChE, CDK2/CDK9), published 1992 to 2025. The archive contains the pairs, the pre-registered scorer, six parameter-free baselines, the permutation nulls, the tier assignment, the figure scripts and the Boltz-2 predictions reported in the accompanying manuscript. python sel_bench_all.py re-scores the deposited predictions and reproduces the main result, 186 of 304 reports (61.2 per cent, 95% CI 55.5-66.7), and the floor sweep; it downloads nothing and fits nothing. sel_bench_collect.py rebuilds the pair table from the ChEMBL records in _chembl_cache/ without network access, and three further scripts reproduce the time split. README.md lists each command and what it regenerates. A MANIFEST.sha256 lists every shipped file. The activity values are derived from ChEMBL and the receptor constructs from UniProt; the source report of every value is named in the data. ChEMBL data are from https://www.ebi.ac.uk/chembl/ - the version of ChEMBL is ChEMBL_37. The records were retrieved from the ChEMBL web services between 4 and 7 September 2026, when ChEMBL_37 was the latest release; the release number was not recorded at the time, and README.md describes how it was identified and checked. Licence: the authors distribute every .tsv and .json file under CC BY-SA 3.0, the licence under which ChEMBL is made available: the ChEMBL records in _chembl_cache/, the tables built from them and the results computed from them. The .py files are MIT. The documents (README.md, LICENSE.md, MANIFEST.sha256 and the two design documents) are CC BY 4.0. LICENSE.md lists every file under its licence, with the reason. This work used compute credits provided by the Anthropic AI for Science program. Version 1.1 adds the ChEMBL records the pairs were built from, the script that builds the pairs from them (sel_bench_collect.py), the census of censored activities, the time-split scripts and the modules the figure scripts import. selectivity_pairs.tsv now holds every pair whose four values come from one report, decidable or not, as the collector writes it; the decidable pairs that v1.0 shipped under that name are in selectivity_pairs_decidable.tsv. The changes are listed in README.md. No scored number changes. Version 1.1 also corrects the licence statement. v1.0 declared CC BY 4.0 for every .tsv, .json and .md file; for the ChEMBL data in those files this was wrong, because ChEMBL is made available under CC BY-SA 3.0. The correction does not withdraw any licence that v1.0 validly granted; such a licence remains in effect for the v1.0 files. In v1.1 the authors distribute every .tsv and .json file under CC BY-SA 3.0 as a distribution policy, and LICENSE.md lists every file under its licence. The related works of v1.0 named ChEMBL release 35, which was wrong; the records are inferred to come from ChEMBL_37, as README.md describes. SHA-256 of selectivity_benchmark_v1.1.zip: 57c51ad7571a5936cb620c095304ff5626598409d2353c291b0f81f887df92a3

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Trajectory Balance: Improved Credit Assignment in GFlowNets

It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequenc...

Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al. · 302 citations · ⚡60

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.