An unsupervised inference model that integrates directed-evolution sequencing time series with natural homologs is presented and the inferred couplings improve structural contact prediction by combining global evolutionary constraints from nature with local, experiment-specific signals.
Abstract
High-throughput directed evolution produces longitudinal sequence libraries that are ideal for probing local fitness neighborhoods but often underpowered for global inference tasks such as contact prediction. We present an unsupervised inference model that integrates directed-evolution sequencing time series with natural homologs. We project sequences into a low-dimensional latent space [1] learned from the natural multiple sequence alignment and model the experimental process as an Ornstein–Uhlenbeck dynamics in that space. Maximum-likelihood estimation of the latent drift and noise parameters determines a stationary Gaussian distribution, which induces an effective Potts model in sequence space. The inferred couplings improve structural contact prediction by combining global evolutionary constraints from nature with local, experiment-specific signals. Experiments on PSE1 β-lactamase and dihydrofolate reductase demonstrate the ability to identify correct complementary contacts not recovered by methods using either natural or experimental data alone, with gains concentrated in intermediate- and long-range contacts.
This work uses a general-purpose "post-training" algorithm grounded in statistical physics that employs quantitative experimental rankings to directly produce a sampler for diverse, high fitness sequences with fewer data points than competing methods.
Sebastian Ibarraran, Shriram Chennakesavalu, Frank Hu et al.· Journal of Chemical Informat...· 0 citations
This work introduces ALSEBO (Active Learning Sequence Exploration via Bayesian Optimization), which couples a generative latent sequence landscape to Bayesian optimization and featurizes candidates with direct-coupling-analysis (DCA) coevolutionary statistics.
D. P. Kulathunga, Divyanshu Shukla, D. Potoyan· bioRxiv· 0 citations
Protein language models organize sequence and structure at scale; here we add a global coordinate system for how proteins respond to mutation.1–12 We present RegimeFormer, a large protein perturbation model coupled to RegimeAtlas, which we constructed by harmonizing, representing and indexing 202,556,313 non-redundant...
Simkins Ma, Yinxiang Chai, Yi Wu et al.· bioRxiv· 0 citations
This work demonstrates how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences.
Talal Widatalla, Ashir Borah, Samuel H. King et al.· Nature Methods· 3 citations
Persistence is a useful and auditable edge target, but the present mapping recovers only a modest, protein-dependent improvement rather than a new state of the art.
Minghan Lyu· Theoretical and Natural Scie...· 0 citations
A cheap density prior is introduced over natural protein activations and keeps only the candidates that remain typical under it, a training-free post-hoc step the authors call Mahalanobis filtering that improves both the property score and the structural plausibility of the sequences it keeps at negligible cost, withou...
Shuibai Zhang, Xin-Chi Liu, Fred Zhangzhi Peng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.