Book
Open access
Aug 2026
The Forgetting-Learning Trade-off: Making Reinforcement Learning Work for Protein Language Models
Diagnostics reveal that RL on PLMs is governed by two reward properties: verifiability, whether the reward is a fixed environment or a learned surrogate vulnerable to distribution shift, and coverage, the fraction of sequence space giving an informative gradient.
Hanqun Cao, Hongrui Zhang, Junde Xu et al.
· Proceedings of the 32nd ACM... · 0 citations