This work proposes AIPO, an enhanced reinforcement learning framework that improves LLM reasoning through active multi-agent interaction during exploration and consistently improves reasoning performance and generalizes across different policy models, collaborator backbones, and RLVR algorithms.
Junnan Liu, Linhao Luo, Thuy-Trang Vu et al.· arXiv.org· 0 citations
Results suggest that different edits are sparsely encoded within dominant singular components and can be separable when the number of edits is moderate, making selective spectral reversal a promising direction for locating edit-specific components and repairing edited language models.
Wei-Feng Jiang, Rui-Rui Chen, Qian-Ren Mao et al.· 0 citations
EASy is proposed, a trainable agentic framework that jointly optimizes task performance and computational efficiency through reinforcement learning and consistently achieves stronger performance-efficiency trade-offs than strong agentic baselines.
Junnan Liu, Linhao Luo, Thuy-Trang Vu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.