Skip to content
#protein folding Dataset Open access

BOLTRA functional pilot dataset: seven-mode biomolecular design and post-design analysis

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This dataset contains the complete computational records from seven 20-design pilot runs performed during the development and validation of BOLTRA v1.0.0 (BoltzGen Orchestration Layer for Targeted design and Results Analysis). BOLTRA is a guided and resumable workflow layer for BoltzGen that supports design-specification generation, PDB/mmCIF residue-number conversion, input validation, interruption recovery, provenance capture, resource monitoring, candidate screening, and publication-oriented reporting. The seven pilot runs represent all design modes supported by BOLTRA: Peptide: Design of 12-15-residue head-to-tail cyclic peptides against chain A of PDB 1T2P. Cyclotide: Design of 29-residue cyclic cystine-knot peptides using the compact sequence specification C3C4C4C1C4C7 and disulfide linkages 1-15, 5-17, and 10-22 against chain A of PDB 1T2P. De novo protein binder: Design of 100-120-residue protein binders against a selected binding site on chain A of PDB 1T2P. Protein redesign: Redesign of 47 selected residues in chain A of PDB 4J9A while retaining the input protein structure as the design context. Small-molecule binder: Design of 123-residue protein binders against the Chemical Component Dictionary ligand D0G. Nanobody: Nanobody design against chain A of PDB 1T2P using the dynamically discovered BoltzGen 7eow scaffold. Antibody: Antibody design against chain A of PDB 1T2P using the dynamically discovered BoltzGen adalimumab.6cr1 scaffold. Each pilot requested 20 designs with a final quality-and-diversity budget of 20, yielding 140 generated structures across the seven design modes. The peptide, cyclotide, de novo protein-binder, nanobody, and antibody pilots used a common binding-site selection on chain A of PDB 1T2P. BOLTRA converted the supplied PDB author residue numbers into the sequential resolved-chain indices required by BoltzGen and retained the resulting residue-mapping tables for auditability. The deposit contains 136 metric-bearing design records. All 20 generated, inverse-folded, and refolded structures are retained for every pilot; however, one generated design did not produce a consolidated analysis-table record in each of four projects: peptide design 00, cyclotide design 01, de novo protein-binder design 01, and small-molecule-binder design 01. The protein-redesign, nanobody, and antibody projects each contain metric records for all 20 designs. This distinction is documented in the dataset README and machine-readable inventory. Design identifiers are zero-based, meaning that identifiers 00-19 represent 20 generated structures. The archive includes the source PDB structures, validated BoltzGen design specifications, copied project inputs, project settings, residue-mapping tables, execution logs, workflow manifests, pipeline configurations, generated structures, inverse-folded and refolded structures, native BoltzGen metric tables, ranked designs, overview reports, BOLTRA screening tables, custom and native candidate classifications, publication-quality figures, figure captions, scientific reports, analysis summaries, and SHA-256 integrity checksums. The post-design analyses retain BOLTRA’s user-defined multi-metric screening results separately from BoltzGen’s native filtering decisions. The screening thresholds were selected for these individual pilot datasets after examining their observed metric distributions and should not be interpreted as universal biological acceptance criteria. The pilot calculations were performed with BoltzGen 0.3.2. The dataset demonstrates BOLTRA workflow execution, input traceability, residue-number conversion, scaffold discovery, interruption recovery, output auditing, candidate selection, and scientific reporting across all seven supported design modes. These results are computational predictions produced for workflow demonstration and software validation. They do not establish binding affinity, specificity, structural correctness, expression, stability, biological activity, therapeutic suitability, or experimental success. All generated and prioritized candidates require appropriate independent experimental validation. BOLTRA source code, installation instructions, and a comprehensive beginner-oriented tutorial are available at: https://github.com/Olanrewaju-Durojaye/BOLTRA

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.