Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning
Switch-Reasoner is proposed, a GRPO-based framework that learns to adaptively select reasoning modes for MLLMs and introduces a dual-level regulation mechanism that balances the overall use of Thinking Mode and Direct Mode while providing sample-level supervision based on the relative benefit of the two choices.