Offline-Online Curriculum RL for Multimodal Reasoning
This work proposes $O^2-CritiCuRL, a novel curriculum reinforcement learning framework that introduces critical-step awareness through an iterative offline-online paradigm, and employs a progressive step-level reinforcement learning strategy, where truncated chains guide the model to infer missing steps and refine its reasoning.