Skip to content
Open access

A comprehensive dataset of 32 million pentapeptide structures for high-throughput virtual screening.

Jul 2026 · Scientific Data · 0 citations
Medicine

Abstract

Small peptides are widely used as binders, modulators, and structural motifs, but their conformational flexibility complicates structure-based analysis and high-throughput screening. We present an open dataset of three-dimensional structures for the complete space of canonical amino-acid pentapeptides: 3,200,000 unique sequences with up to 10 conformers per sequence, for a total of 32,000,000 peptide conformers. Structures were generated directly from sequence using an automated workflow built on UCSF ChimeraX for model construction, Reduce for hydrogen placement, and RDKit for conformer generation and optimization. The dataset is distributed as compressed archives with an accompanying index that maps each sequence and conformer identifier to its coordinate record, enabling efficient download, subset selection, and programmatic access. Technical validation includes symmetry-aware inter-conformer RMSD analysis, Ramachandran quality assessment, and benchmarking against experimentally observed pentapeptide fragments from the Protein Data Bank. Although we focus here on pentapeptides to enable exhaustive sequence coverage, the publicly released workflow is solely based on open-source software and can be applied to other short peptides to generate comparable conformer libraries. This resource supports virtual screening with pre-generated peptide conformer ensembles, method benchmarking, and machine-learning applications in peptide design and protein engineering by removing the need for researchers to repeatedly generate large conformer ensembles from scratch.

Read PDF

Similar papers

Open access Aug 2026

PepXPro: a framework for curating, generating, and optimizing structure-affinity protein-peptide datasets

P PepXPro is presented, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria that provides an extensible foundation for reproducible protein- peptide benchmark construction.

L. A. Chi, F. M. Ytreberg · 0 citations
Jul 2026

HighDB: a structure-annotated cyclic peptide database for comparative analysis, template retrieval, and design-oriented applications

A unified annotation framework for topological and conformational descriptors is provided and HighDB is compared with representative cyclic peptide resources, showing broader annotation coverage and the largest collection of experimentally resolved cyclic peptide structures among the databases examined.

Yi-Qi Xu, Ning Zhu, Tianfeng Shang et al. · 0 citations
Aug 2026

Evaluating BioEmu-Generated Kinase Ensembles Reveals Structure Selection as the Virtual Screening Bottleneck.

It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.

Jaeoh Shin, K. Joo, Jejoong Yoo · 0 citations
Open access Jul 2026

Rich structure alphabets enable highest accuracy protein search

Protein structure databases have grown from thousands of experimentally determined structures to hundreds of millions of AI-predicted models, creating an urgent need for search methods that combine high accuracy with practical scalability. Here, I present the third generation of Reseek, a protein structure search algorithm achieving the highest overall accuracy (median rank 1) according to diverse metrics among tested methods including DALI, Foldseek and TM-align. Improved accuracy is obtained by parallel sequence alignment of many discrete alphabets capturing primary, secondary and tertiary features, giving a combined space of ∼1022 possible states. Separate statistical models are optimized for family, superfamily and fold discrimination, respectively, revealing distinct combinations of features that characterize each level. Hundreds of query structures can be searched a against a multimillion-structure database on a server computer in minutes, making large-scale structure search at state-of-the-art accuracy practical on commodity hardware.

Robert C. Edgar · 0 citations
Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.