Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on s...
X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al.· 0 citations
The results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system.
X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al.· 1 citation
This work proposes MemLoc, a unified Retrieve-Localize-Generate framework for long-term conversational memory QA, and introduces a reasoning-based evidence locator trained with Self-reflective Hint Policy Optimization, which performs progressive refinement by extracting query-relevant fragments within memory units to s...
Yi-Fan Wang, Xin-Kui Lin, Yong-Xiu Xu et al.· 0 citations
gui-PRA is introduced, a Process Reward Agent that transforms GUI process evaluation from passive scoring into active investigation, and demonstrates strong competitiveness against fully trained critic models.
Tao Xiong, Xavier Hu, Yurun Chen et al.· arXiv.org· 5 citations
Results show that TRACE converts high model potential into stable, consistent performance gain, and bridge the gap between potential and reliable performance to just 4.0 points.
Wen-Hao Wu, Meng-Hao Zhang, X. Wang et al.· 2 citations
GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization, is introduced.
Long Zhang, Yuhan Chen, Chaoran Zhang et al.· 0 citations
This work studies reference-free post-training for multilingual machine translation with open large language models and finds that on-policy distillation reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation.
Chris Han, Pengzhi Gao, Pei Fu et al.· 0 citations
See, a two-stage data synthesis framework consisting of an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling, yields reproducible and explainable data gen...
Zhuohang Fan, Beichen Zhang, Yuanfa Li et al.· 0 citations
G-ReAct is a reasoning framework for deep search that organizes reasoning as state evolution over a fixed-topology query graph, transforming exploratory search driven by textual history into graph-guided reasoning under explicit constraints.
Shaoxiong Yang, Mengyuan Zhang, Shao-Jun Lin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.