Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data, achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute.