Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction
It is found that, at variance with networks not based on the attention mechanism, the lack of a first--order--like transition in the loss of the transformer produces a range of intermediate temperatures with good learning properties; this is true both for synthetic and natural protein sequences.