A three-stage investigation framework is established to empirically study the inference design space of multimodal pLMs across three representative pLMs and four fundamental tasks, and to derive conclusions about base models that differ from prior consensus.
Abstract
Multimodal protein language models (pLMs) learn joint protein sequence-structure distributions, and their generation performance should also depend critically on inference-time sampling strategies. Yet prior work has focused more on model training than on how inference-time strategies behave. In this paper, we establish a three-stage investigation framework to empirically study the inference design space of multimodal pLMs across three representative pLMs and four fundamental tasks. We evaluate vanilla sampling, task-specific classifier-free guidance, and reward-guided beam search on multimodal pLMs, corresponding to controls over sampling distributions, per-step logits, and parallel trajectories. Throughout the complementary advancements centered on exploration-exploitation trade-off, we (1) reveal the suboptimality of default inference protocols and identify task-oriented sampling preferences; (2) observe substantial quantitative gains across tasks, consistently boosting the upper bound performance of multimodal pLMs without updating model parameters; (3) derive conclusions about base models that differ from prior consensus.
It is shown that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification, and MPEP conditioning is established as a lightweight strategy for low-data peptide classification.
Joshua Almonte, Minh N. Vu, Andrew Ahn et al.· bioRxiv· 0 citations
Diagnostics reveal that RL on PLMs is governed by two reward properties: verifiability, whether the reward is a fixed environment or a learned surrogate vulnerable to distribution shift, and coverage, the fraction of sequence space giving an informative gradient.
Hanqun Cao, Hongrui Zhang, Junde Xu et al.· Proceedings of the 32nd ACM...· 0 citations
TTS-Design is proposed, a test-time compute scaling framework that enhances protein sequence design without retraining models or relying on larger training datasets, and can consistently improve sequence recovery and structural reliability across different backbone models, without retraining or increasing model size.
Zizhe Jin, Yi Zheng, Huan Yee Koh et al.· 0 citations
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs). By consensus, embeddings from the models'last layers are used, while the models'internal behavior remains poorly understood. We analyzed 13 PLMs across 15 DTs and 9 datasets to assess the value of embeddings from intermediate PLM layers. We trained probe models on embeddings from each layer, compared their performance, and showed that the last layers of PLMs rarely produced embeddings that led to the best results on downstream tasks. Furthermore, we identified a connection between how models learn a certain DT and the similarity between that DT and the pre-training objective. For example, for residue-level downstream tasks, we observed a steady increase in performance across almost all PLM layers, which we attributed to their similarity to most PLMs'pre-training objectives. To allow the community to capitalize on our findings, we provide PLMSommelier, a Python package that automatically identifies the best PLM layer for a given DT with ~98% accuracy and creates a truncated model using only the early layers up to the best-performing layer. This will help users save time and memory during inference and yield better predictive performance.
R. Joeres, Ilya S. Senatorov, A. Kolchina et al.· 0 citations
This work demonstrates how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences.
Talal Widatalla, Ashir Borah, Samuel H. King et al.· Nature Methods· 1 citation
GenomeHarness, an agentic harness for adapting genome language models through controlled search over fine-tuning recipes, improves mean test MCC in 47 settings, and shows gains on Genomic Benchmarks and on tasks where the root recipe is unstable or poorly matched.
Weicai Long, Yusen Hou, Houcheng Su et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.