A unified annotation framework for topological and conformational descriptors is provided and HighDB is compared with representative cyclic peptide resources, showing broader annotation coverage and the largest collection of experimentally resolved cyclic peptide structures among the databases examined.
P PepXPro is presented, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria that provides an extensible foundation for reproducible protein- peptide benchmark construction.
The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.
Yana Rose, Ronald S. Brown, Maria Voigt et al.· Protein Science· 0 citations
Peptide-protein interactions are fundamental to many biological processes, and peptide design is gaining interest due to its therapeutic potential. This has led to the emergence of various structural databases, such as PepBDB and Peptipedia, that classify peptides as polypeptides with fewer than 50 amino acids. These databases provide valuable starting points for studying peptide recognition by protein partners and are widely used for machine learning applications, docking, and scoring functions. However, under such a length definition for peptides, we have very different cases that are likely to confound the analysis of peptide binding, including miniproteins, peptide-peptide complexes, intramolecular peptide disulfide bonds, intermolecular disulfide bridges, and proteins undergoing internal cleavage, such as serpins, as well as non-natural amino acids and covalently bound cofactors. Here, we present a rigorous classification of peptide-protein complexes to generate datasets suitable for comparative energetic analysis with a focus on peptides that are unstructured in the absence of their target protein. The analysis of this dataset shows that peptide binding is typically driven by a small number of hotspot residues mainly enriched in aromatic and bulky hydrophobic side chains. Their number of hotspots and their spatial organization depend on peptide length, secondary structure, and covalent constraints. Short peptides rely on central anchor regions, whereas longer peptides distribute hotspots more broadly, with helices showing periodic spacing and β-strands relying more on backbone-mediated stabilization. Disulfide bonds further decrease the number of hotspots per peptide length by either pre-organizing the peptide or acting as covalent anchors. This work provides a curated resource and general principles for peptide recognition. It highlights the importance of structurally classifying peptide-protein complexes to avoid bias in downstream computational and machine-learning applications.
Rahma Hamdani, Javier Delgado, Luis Serrano· Protein Science· 0 citations
HighMorph is presented, an interaction-guided framework that combines protein–protein interaction information with artificial intelligence for rational cyclic peptide design and provides insights for developing therapeutics targeting challenging protein interfaces.
M. Lan, Chengyun Zhang, Wentong Wang et al.· Journal of Medicinal Chemist...· 0 citations
Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities among protein-ligand complexes. This fragmented perspective constrains their ability to capture the fundamental principles governing molecular recognition and binding specificity. Moreover, the limited availability of high-quality experimental data further hampers model generalization and real-world applicability. To address these challenges, we present READ, a retrieval-alignment molecular generation framework that conditions the generative process on small molecules targeting homologous proteins. Retrieved ligands are aligned with a diffusion model across multiple representational spaces and integrated as conditional guidance throughout successive stages of generation. Under a standardized docking-based evaluation protocol, READ achieves consistently strong performance against state-of-the-art SBDD methods. More importantly, it introduces a retrieval-alignment paradigm for structure-based molecular generation, offering a practical framework for early-stage computational hit generation while leaving prospective experimental validation as future work.
Dong Xu, Zhangfan Yang, Junchuang Cai et al.· IEEE transactions on computa...· 1 citation
Post-translational modifications (PTMs) and genetic variants regulate protein function, signalling, and disease, but their interpretation requires integration of sequence annotations with structural, interaction, and biophysical context. Although resources such as Scop3P, UniProt, the Protein Data Bank, and AlphaFold provide extensive annotations and structural information, integrating these data into reproducible structure-aware analyses still requires custom scripting and manual coordination between multiple independent tools. To address this challenge, we developed Scop3P-Toolkit, an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context. The toolkit integrates protein annotation retrieval with structural mapping, residue interaction network analysis, comparative structural analysis, and residue-level biophysical profiling within a unified framework. Experimentally supported phosphosites, phosphopeptides, and phosphoproteomics evidence are provided for human proteins through Scop3P, with optional integration of curated UniProt PTM annotations. UniProt-derived PTMs, sequence features, and genetic variants are available for proteins from any species, extending the framework beyond the human phosphoproteome. Scop3P-Toolkit supports structure-centric analyses including interpretation of PTMs and disease-associated variants, analysis of residue interaction networks and their rewiring across alternative conformations, structural localisation of peptides, and exploration of protein–protein, protein–ligand, and host–pathogen interfaces. Interactive visualisation links sequence annotations, three-dimensional structures, residue interaction networks, and biophysical profiles, enabling coordinated exploration across multiple molecular representations. The toolkit is distributed as Jupyter notebooks, browser-based Voilà applications, and a Galaxy interactive tool, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers. By integrating biological annotation resources into executable, structure-aware workflows, Scop3P-Toolkit enables reproducible interpretation of PTMs, mutations, and proteomics data.
Adrián Díaz, Natalia Tichshenko, Boris Depoortere et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.