Skip to content

Similar papers

Open access Jul 2026

Muon Reduces the Training Cost of Regulatory DNA Transformers

Analysis of Transformer models on ENCODE cis-regulatory sequences with Adam and Muon indicates that optimizer update structure and norm-control choices are practical levers for reducing the training resources required to reach matched perplexity targets in regulatory DNA pretraining.

Viraj Doshi, Malhar Bhide, Akshat Singh et al. · 0 citations
Open access Jul 2026

Prediction-Guided Design of a More Developable FGF21 Construct

For structural-biology and protein-production pipelines, the hardest part of a difficult protein is not the biology — it is obtaining a well-behaved sample for functional studies. Programs routinely stall at construct design, expression, and purification: deciding where to truncate, which tags to use, how to express, and how to purify so the protein survives concentration and handling. These decisions are still made largely by literature precedent and experimental experience, and they require trial-and-error before arriving at a functional construct for hard targets. We present a prospective, single-pair wet-lab case study testing whether an integrated computational platform can improve these decisions. For human fibroblast growth factor 21 (FGF21) — a clinically important and stability-challenged metabolic hormone — we compared two expression constructs produced side by side under the same experimental workflow, using two different design strategies: one designed by a scientist from the literature (reproducing the published core-domain construct, PDB 6M6E), and one designed by the Orbion platform — an AI, prediction-guided protein-design system (orbion.life) — which additionally generated the expression and purification protocols (executed scientist-in-the-loop). The platform’s construct used an unconventional, longer C-terminal boundary not found in public sequence databases. Since the two constructs differ in more than one feature, we treat them as workflow-level designs throughout. The scientist construct gave a higher initial yield (∼2.4 ×more protein recovered at affinity capture). The platform-designed construct, however, showed a more favourable downstream developability profile: it concentrated higher (1.4 vs 0.7 mg/mL) while remaining more monodisperse by dynamic light scattering (DLS). The scientist construct, in contrast, aggregated on concentration, so its initial-yield advantage did not survive: in the final concentrated sample the Orbion construct provided the more usable material for downstream studies. Computed for the mammalian host used, the platform had prospectively scored its own design higher (composite 68.7 vs 59.0 for the scientist-designed construct), and its predictions of yield, solubility, and disorder matched the wet-lab outcome. This is a single, deliberately scoped case study, not a population-level benchmark; the two constructs differ in more than one feature, and biological activity was not assayed. Alongside the bottlenecks of this approach discussed here, used as a decision aid, prediction-guided construct and protocol design has the potential to remove costly iteration cycles of protein production campaigns.

Çağlar Bozkurt, Evangelia Nathanail, Aniruddh Goteti · 0 citations
Review Jul 2026

Toward generalizable and interpretable AI in regulatory genomics

It is suggested that progress requires reframing seq2func models as continually refined systems, in which targeted perturbation experiments, systematic evaluation and iterative model updates are tightly coupled through artificial intelligence-experiment feedback loops, enabling self-improving models that progressively deepen mechanistic understanding and more reliably support biological discovery.

Masayuki Nagai, A. E. Murphy, Kaeli Rizzo et al. · 2 citations
Open access Jul 2026

Deep Learning Predicts Dissimilar DNA-DNA Binding and Engineers Hyperconnected Networks

BINND is developed, a binding and interaction neural network to predict non-orthogonal DNA interactions, which could aid diagnostic, bioengineering, and DNA origami design, and supporting a shift toward exploiting the full sequence space.

Karishma Matange, Gunavaran Brihadiswaran, Kyle J. Tomek et al. · 0 citations
Review Open access Aug 2026

How to Build Machine-Learning Models for Molecular Science: A Step-by-Step, Annotated Tutorial

This tutorial provides a comprehensive, end-to-end workflow from raw data to deployed models,icitly designed for environmental chemists with limited prior experience in ML modeling while also providing practical guidance for other users seeking to strengthen their modeling workflows.

Kai Zhang, Yushu Cheng, Hai-Ping Ai et al. · 0 citations
Open access Nov 2025

seqme: a Python library for evaluating biological sequence design from generative models

This work introduces seqme, a modular and highly extendable open-source Python library, containing model-agnostic metrics for evaluating computational methods for biological sequence design, and can be used to evaluate both one-shot generation and iterative optimization.

Rasmus Møller-Larsen, Adam Izdebski, Jan Olszewski et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.