Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-genera...
Jin-Yuan Li, Chengsong Huang, Lang-Lin Huang et al.· 0 citations
VisPlay is introduced, a self-evolving RL framework that enables VLMs to autonomously improve their reasoning capabilities from massive unlabeled image data and establishes a scalable path toward self-evolving multimodal intelligence.
Yicheng He, Chengsong Huang, Zongxia Li et al.· 0 citations
Environment Harness is proposed, a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic, enabling continuous, targeted co-evolution of the policy and its environment.
Chengsong Huang, Zifeng Wang, Ru-Jun Han et al.· 11 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.