TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing local...