Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works ref...
Qi Yu, Rui-Zhong Qiu, Zhichen Zeng et al.· 0 citations
PolicyMem is introduced, a geometric policy memory that externalizes natural-language policies as reusable geometric memory objects represented by low-rank subspaces in a shared representation space that achieves state-of-the-art unsafe behavior detection while enabling effective policy attribution, rewriting, and post...
Yuan-Chen Bei, Zheng-Zhang Chen, Yan-Jun Zhao et al.· 0 citations
This formulation enables a systematic study of key self-improvement factors through the proposed Evo-Harness, and provides a principled understanding of how LLM agents can effectively learn on the fly.
Tian-Xin Wei, Zhan Shi, Min-hua Lin et al.· 12 citations
This work proposes method that unifies textual reasoning and graph message passing within a masked diffusion language model, a language model with bidirectional attention and generative decoding that outperforms graph neural networks, graph transformers, and LLM-based baselines on all three TAG benchmarks across two ta...
This work proposes a principled VLM TTA method called \algname, and theoretically reveals that the InfoNCE loss can be neatly reformulated as a Wasserstein OT formulation, thereby unifying the objectives of the inference and adaptation of VLMs to achieve their mutual benefits.
Qi Yu, Zhichen Zeng, Katherine Tieu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.