Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the corre...
Zhong-An Bi, Ke-Peng Lin, Xuan-Ang Gao et al.· 0 citations
Quality-Gated Length Advantage Shaping (QGLAS), which first computes advantages from quality rewards alone, then adds bounded bonuses only to shorter positive-advantage responses, leaving all other advantages unchanged, consistently achieves a stronger quality--length trade-off than representative baselines.
Zi-Jun Weng, Zhong-An Bi, Xuan-Ang Gao et al.· 0 citations
HAE-GEO is introduced, a benchmark that tracks the full trajectory from exposure to recovery under progressively more persuasive Web poisoning and finds three recurring patterns: evidence recognition degrades under the corroboration trap, agentic search improves final resistance without improving evidence recognition o...
Zhong-An Bi, Qi-Wen Wang, Jian-Rong Jiang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.