Protein language models (PLMs) such as ESM-2 encode protein sequences as embeddings for downstream tasks. PLMs are trained on a masked learning objective that leverages evolutionary constraints. While interpretability studies of ESM-2 have focused on folded proteins, their behavior on intrinsically disordered proteins (IDPs), which constitute a substantial fraction of the human proteome and are implicated in numerous diseases, remains understudied. Because IDPs experience different types of evolutionary constraints on their amino acid sequences, we hypothesized that PLMs would behave differently on disordered versus folded regions. Here we show that ESM-2 exhibits reduced attention on disordered regions, yet still encodes meaningful biological signals. The model assigns heightened attention to disease-relevant residues even at high levels of disorder. Moreover, we show that both the radius of gyration and individual dynamic contact maps, key characteristics of IDPs, can be obtained from the model logits and embeddings. These findings suggest PLMs capture valuable information relevant to IDP biology despite their bias toward structured residues.
Lauren Naworski, Lydia L. Good, Rob M. Scrutton et al.· bioRxiv· 0 citations
It is shown that excipient-mediated solubilization of therapeutic antibodies is markedly molecularly specific, and the integration of high-throughput experimentation with molecular feature analysis offers a foundation for improving the understanding and prediction of antibody-specific formulation behavior.
Zexiang Han, Nadia A. Erkamp, Rob M. Scrutton et al.· mAbs· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.