Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works ref...
Qi Yu, Rui-Zhong Qiu, Zhichen Zeng et al.· 0 citations
This survey synthesizes agentic reasoning methods into a unified roadmap bridging thought and action, and outlines open challenges and future directions, including personalization, long-horizon interaction, world modeling, scalable multi-agent training, and governance for real-world deployment.
Tian-Xin Wei, Ting-Wei Li, Zhining Liu et al.· 39 citations· ⚡6
PolicyMem is introduced, a geometric policy memory that externalizes natural-language policies as reusable geometric memory objects represented by low-rank subspaces in a shared representation space that achieves state-of-the-art unsafe behavior detection while enabling effective policy attribution, rewriting, and post...
Yuan-Chen Bei, Zheng-Zhang Chen, Yan-Jun Zhao et al.· 0 citations
The proposed Mixture of Roles (MoRe), which adaptively composes multiple specializations into a single steering vector for single-turn inference, enables multi-perspective specialization in a single-agent, single-turn inference process.
Zhichen Zeng, Hui-Yuan Chen, Jingru Cheng et al.· 2 citations
PILL (Probing-based InfiLling with preset-Length-free decoding), an efficient infilling method for DLMs that requires no preset initial length and adds far fewer extra forward passes than baselines, substantially reducing inference time is proposed.
Hao-Bo Xu, Si-Rui Chen, Yuanchen Bei et al.· 3 citations
AFANet is introduced, a lightweight graph-based framework that models interaction trajectories through step-level semantic signals and agent-level relationships and suggests that effective agent failure attribution does not require heavy LLM reasoning and a lightweight, structured approach can achieve strong performanc...
Ting-Wei Li, Yuan-Chen Bei, Xiao Lin et al.· 1 citation
This work proposes method that unifies textual reasoning and graph message passing within a masked diffusion language model, a language model with bidirectional attention and generative decoding that outperforms graph neural networks, graph transformers, and LLM-based baselines on all three TAG benchmarks across two ta...
EvoHarness-RL is introduced, which exposes Belief, Progress, and Experience (BPE) as policy-facing harness state and reveals two key dynamics: harness annealing, where training internalizes recurring harness-use patterns into the model policy and shifts the agent from frequent harness calls toward selective external-st...
Xuying Ning, Dongqi Fu, Tianxin Wei et al.· 0 citations
This work proposes a principled VLM TTA method called \algname, and theoretically reveals that the InfoNCE loss can be neatly reformulated as a Wasserstein OT formulation, thereby unifying the objectives of the inference and adaptation of VLMs to achieve their mutual benefits.
Qi Yu, Zhichen Zeng, Katherine Tieu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.