Deep Research Pretraining (DRP), an offline framework that derives predictive navigation supervision from naturally occurring evidence structures, is introduced, an offline framework that derives predictive navigation supervision from naturally occurring evidence structures.
Jiang-Nan Zhou, Zhi-Yuan Fan, Xing Wu et al.· 1 citation
PolicyLong is proposed, shifting data construction towards a dynamic on-policy paradigm, by iteratively re-executing data screening (entropy computation, retrieval, and verification) using the current model, which ensures the training distribution tracks evolving capabilities, yielding an emergent self-curriculum.