Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between thes...
Zi-Hang Shen, Qi Liu, Zi-Xuan Yang et al.· 0 citations
Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuff...
Zi-Xuan Yang, Yi-Qun Chen, Qi Liu et al.· 0 citations
Fetch-then-Explore is proposed, which separates page selection from evidence extraction and keeps what it selects: pages are recorded in a per-question workspace on the filesystem rather than the context window or a transient session, and evidence is pulled from them on demand later.
Qi Liu, Yi-Qun Chen, Zidan Chen et al.· 0 citations
CoSkill is a unified multi-agent RL framework that recasts the static meta-skill workflow as a learnable Meta-Skill Agent and jointly trains it with a Reasoning Agent over a hierarchical skill library and achieves superior early-stage sample efficiency, asymptotic performance, and wall-clock efficiency.
Jin-Yuan Feng, Dong-Min Li, Yi-Qun Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.