This work presents CROWN (Curated Repository Of Well-resolved Non-covalent interactions), a machine-learning–ready dataset that reconciles scale and rigor through a fully automated preprocessing pipeline and can in principle support any model that learns from or is evaluated against protein–ligand complex structures.
Abstract
The development of machine-learning models for protein–ligand interactions is constrained by the quality and diversity of the available structural data. Existing resources force researchers into a trade-off: carefully curated collections such as PDBBind and HiQBind offer high structural reliability but cover only a narrow slice of the Protein Data Bank (PDB), whereas large-scale resources such as PLINDER provide broad coverage with minimal quality control. We present CROWN (Curated Repository Of Well-resolved Non-covalent interactions), a machine-learning–ready dataset that reconciles scale and rigor through a fully automated preprocessing pipeline. Starting from the PDB database, CROWN applies a series of interleaved quality filters and processing stages that address crystallographic resolution, ligand identity, pocket completeness, structural repair, interaction quality, and protonation at physiological pH. The pipeline finishes with a constrained energy-minimization step built on custom flat-bottomed restraints — a step absent from all existing protein–ligand datasets — that balances crystallographic evidence against the relaxation of intramolecular strain. By reconciling the heterogeneous refinement practices of different depositions without distorting the experimentally observed binding geometry, this step yields a structurally uniform collection of 178,263 complexes, representing a roughly four-fold increase in protein diversity over PDBBind and HiQBind. Rather than organizing the data around sparsely available, bias-prone binding affinities, CROWN adopts a geometry-centric design philosophy that treats the three-dimensional arrangement of atoms at the binding interface as a self-consistent source of information. To demonstrate its value as a training resource, we trained two knowledge-based scoring functions on CROWN and benchmarked them on CASF-2016: relative to HiQBind-trained counterparts, CROWN-trained models showed markedly improved ranking power (mean Spearman correlation rising from 0.509 to 0.637) and docking power (top-1 near-native pose recovery of 0.785 versus 0.724). Because CROWN imposes no requirement for affinity labels, it can in principle support any model that learns from or is evaluated against protein–ligand complex structures. We anticipate that it will serve as a broadly useful resource for tasks such as the training of binder generation, protein design or protein folding models conditioned on bound ligands, the development of scoring functions or benchmarking of interaction-prediction methods.
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.
Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al.· Journal of Systems and Softw...· 78 citations· ⚡6
This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.
D. Graziotin, Xiaofeng Wang, P. Abrahamsson· SSE@SIGSOFT FSE· 56 citations· ⚡4
This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.
Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 44 citations· ⚡5
It is demonstrated that linker-free PROTACs can outperform traditional designs, marking a paradigm shift in PROTAC development for targeted protein degradation.
Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or sequence constraints.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.