Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Un...
Zi-Hao Chen, Fan-Xiang Xiong, Hong-Ran Ren et al.· 0 citations
Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) f...
Shao-Qing Zhang, Ke-Hai Chen, Xue-Feng Bai et al.· 0 citations
Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with pseudo-labels constructed through majority voting. While effective, the reward signal assigned from majority voting is highly sensitive to consensus strength, defined as...
Kunbin Xu, Xingzuo Li, Xue-Feng Bai et al.· 1 citation
DM-Align is introduced, which derives a complementary gradient direction to guide the model toward human-preferred samples, and eliminates the need for multi-step reward evaluation and complex ODE-SDE conversions inherent in traditional RL.
Jiu-Zhou Lin, Jun-Long Wu, Feilong Zuo et al.· 1 citation
DualAnchor is proposed, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation that achieves strong overall performance on both PHOENIX-2014T and CSL-Daily.
Hongbin Zhang, Jun-Hao Liu, Xue-Feng Bai et al.· arXiv.org· 0 citations
The proposed DR 2 identifies and localizes non-deterministic reasoning behaviors, uncovering the underlying semantic representation deficiencies in LLMs, and designs abductive reasoning-based preference learning, which promotes fine-grained semantic discrimination and mitigates non-deterministic reasoning errors.
Ge Liang, Mufan Xu, Kehai Chen et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.