#artificial intelligence
Dec 2025
Better World Models Can Lead to Better Post-Training Performance
It is found that explicit world-modeling yields better representations in terms of higher probing accuracy and steerability of the model, and that better representations yield larger gains from GRPO, especially on harder cube states.
Prakhar Gupta, Henry Conklin, Sarah-Jane Leslie et al.
· arXiv.org · 3 citations