Preprint
Aug 2026
From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection
This work proposes Golden-GRPO Injection (GRIN), a three-stage self-learning framework for continual knowledge injection that substantially outperforms SFT and mixed-policy RL baselines on the harder question types while matching them on basic fact recall.
Zhibo Hou, Fan Zhao, Zhiyu An et al.
· 0 citations