This work proposes AIPO, an enhanced reinforcement learning framework that improves LLM reasoning through active multi-agent interaction during exploration and consistently improves reasoning performance and generalizes across different policy models, collaborator backbones, and RLVR algorithms.
Junnan Liu, Linhao Luo, Thuy-Trang Vu et al.· arXiv.org· 0 citations
JRDB-AVR is introduced, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap between answer accuracy and evidence accuracy in current baselines into an explicit evaluation.
Zhixi Cai, Fu-Cai Ke, Sukai Huang et al.· 0 citations
Conformal Privacy Auditing is introduced, a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries and enables audits of open-source models and proprietary API models in a unified framework.
Shuo Huang, G. Haffari, Xing-Liang Yuan et al.· 0 citations
Compilable Academic Document Parsing (CADP) is proposed, a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page.
Rihui Jin, Jun Wang, Chen Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.