Reimagining the Synthetic Biology DBTL Cycle with Machine Learning.
TL;DR
Generalizable principles that can be applied to data-driven circuit design projects are highlighted and defined in this work.
AI Networking Cookbook: Practical recipes for AI-assisted network automation and development
TL;DR
Generalizable principles that can be applied to data-driven circuit design projects are highlighted and defined in this work.
Analysis of Transformer models on ENCODE cis-regulatory sequences with Adam and Muon indicates that optimizer update structure and norm-control choices are practical levers for reducing the training resources required to reach matched perplexity targets in regulatory DNA pretraining.
For structural-biology and protein-production pipelines, the hardest part of a difficult protein is not the biology — it is obtaining a well-behaved sample for functional studies. Programs routinely stall at construct design, expression, and purification: deciding where to truncate, which tags to use, how to express, and how to purify so the protein survives concentration and handling. These decisions are still made largely by literature precedent and experimental experience, and they require trial-and-error before arriving at a functional construct for hard targets. We present a prospective, single-pair wet-lab case study testing whether an integrated computational platform can improve these decisions. For human fibroblast growth factor 21 (FGF21) — a clinically important and stability-challenged metabolic hormone — we compared two expression constructs produced side by side under the same experimental workflow, using two different design strategies: one designed by a scientist from the literature (reproducing the published core-domain construct, PDB 6M6E), and one designed by the Orbion platform — an AI, prediction-guided protein-design system (orbion.life) — which additionally generated the expression and purification protocols (executed scientist-in-the-loop). The platform’s construct used an unconventional, longer C-terminal boundary not found in public sequence databases. Since the two constructs differ in more than one feature, we treat them as workflow-level designs throughout. The scientist construct gave a higher initial yield (∼2.4 ×more protein recovered at affinity capture). The platform-designed construct, however, showed a more favourable downstream developability profile: it concentrated higher (1.4 vs 0.7 mg/mL) while remaining more monodisperse by dynamic light scattering (DLS). The scientist construct, in contrast, aggregated on concentration, so its initial-yield advantage did not survive: in the final concentrated sample the Orbion construct provided the more usable material for downstream studies. Computed for the mammalian host used, the platform had prospectively scored its own design higher (composite 68.7 vs 59.0 for the scientist-designed construct), and its predictions of yield, solubility, and disorder matched the wet-lab outcome. This is a single, deliberately scoped case study, not a population-level benchmark; the two constructs differ in more than one feature, and biological activity was not assayed. Alongside the bottlenecks of this approach discussed here, used as a decision aid, prediction-guided construct and protocol design has the potential to remove costly iteration cycles of protein production campaigns.
It is suggested that progress requires reframing seq2func models as continually refined systems, in which targeted perturbation experiments, systematic evaluation and iterative model updates are tightly coupled through artificial intelligence-experiment feedback loops, enabling self-improving models that progressively deepen mechanistic understanding and more reliably support biological discovery.
BINND is developed, a binding and interaction neural network to predict non-orthogonal DNA interactions, which could aid diagnostic, bioengineering, and DNA origami design, and supporting a shift toward exploiting the full sequence space.
This tutorial provides a comprehensive, end-to-end workflow from raw data to deployed models,icitly designed for environmental chemists with limited prior experience in ML modeling while also providing practical guidance for other users seeking to strengthen their modeling workflows.
This work introduces seqme, a modular and highly extendable open-source Python library, containing model-agnostic metrics for evaluating computational methods for biological sequence design, and can be used to evaluate both one-shot generation and iterative optimization.
We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.