Skip to content

Author

Marvin Gülhan

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint May 2026

Reinforcement Learning Can Amplify Emergent Misalignment from Harmless Rewards

It is shown that rewarding narrow, overtly misaligned behavior produces substantially higher general-domain misalignment than sample-matched SFT, and that EM from RL can be induced by reward signals that could plausibly arise naturally, such as unpopular aesthetic preferences or poor rhetorical appeals.

Magnus Jørgenvåg, David Kaczér, Lasse Ruttert et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.