In LLM Reasoning, there is Irrationality on top of Value Misalignment
It is argued that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning, and the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart whose responses maximise utility in the steepest direction is formalised.