Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on s...
X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al.· 0 citations
Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides eff...
Changhao Xiang, Shi-Lin Zhang, Zheng Ma et al.· 1 citation
The results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system.
X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al.· 1 citation
Value Flattening is identified as an important yet overlooked failure mode of critic learning in standard PPO and a simple sparse supervision strategy can mitigate it; SParse Proximal Policy Optimization is introduced, which applies the value loss to only a few well-separated states in each response to mitigate both ef...
Yi-Zhuo Li, Jian-Hao Yan, Yun Luo et al.· 1 citation
Results show that TRACE converts high model potential into stable, consistent performance gain, and bridge the gap between potential and reliable performance to just 4.0 points.
Wen-Hao Wu, Meng-Hao Zhang, X. Wang et al.· 2 citations
FailForge is proposed, an agentic framework that converts failed rollouts into training signal, and recovers over 26% of previously failed instances at marginal additional cost, and training Qwen3.5-4B on the augmented corpus improves the SWE-bench Verified resolve rate by 6.6 points over a strong RFT baseline.
Dongyi Lv, E. Fushun, Aichen Cai et al.· 1 citation
By decoupling action proposal from consequence evaluation, SVA preserves the generalization capacity of the VLA backbone while substantially improving task success rates, and shows that SVA consistently improves generalization on unseen tasks and exhibits strong test-time scaling behavior.
Xinyi Xie, Zican Hu, Zhanyun Liu et al.· 0 citations
This work uses a full $2^3$ factorial design to decompose three recurring interventions in formalization pipelines: parametric expert drafting, Mathlib/context search, and Lean elaboration feedback, suggesting that formal validity, proof-oriented Lean competence, and faithful statement generation should be reported sep...
Ke Zhang, P. Gallardo, S. Murthy et al.· arXiv.org· 1 citation
Experiments on MemFuseBench show that MemFuse achieves the best overall performance among the evaluated memory systems under all three LLM settings and consistently improves performance on questions requiring cross-source evidence fusion.
Chao Li, Yuan-Fang Li, Wenhao Wu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.