Large Reasoning Models Learn Better Alignment from Flawed Thinking
RECAP (Robust Safety Alignment via Counter-Aligned Prefilling), a principled reinforcement learning (RL) method for post-training that explicitly teaches models to override flawed reasoning trajectories and reroute to safe and helpful responses, substantially improves safety and jailbreak robustness, reduces overrefusal, and preserves core reasoning capability.