Protein language models learn underlying mutation biases alongside fitness landscapes
A parameter sweep is implemented to explicitly couple empirical nucleotide mutational supply from PLM-assessed amino acid substitution pseudo-probabilities across evolutionary forecasting tasks and finds that base PLMs implicitly learn generic nucleotide-level mutational constraints, an effect strongly amplified by virus-specific fine-tuning.