Reinforcement Learning Can Amplify Emergent Misalignment from Harmless Rewards
It is shown that rewarding narrow, overtly misaligned behavior produces substantially higher general-domain misalignment than sample-matched SFT, and that EM from RL can be induced by reward signals that could plausibly arise naturally, such as unpopular aesthetic preferences or poor rhetorical appeals.