Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works ref...
Qi Yu, Rui-Zhong Qiu, Zhichen Zeng et al.· 0 citations
This work synthesizes a task-conditioned temporal workflow graph that jointly specifies agent connectivity and edge-level communication semantics, and introduces ReActNet, a training-free framework that compiles a query and a set of role-specialized agents into a sequence of directed communication graphs.
Scientific datasets, such as materials and molecular datasets, are often large, complex, and open-ended, posing a core challenge for data efficiency and model training. While data attribution (DA) offers a principled way to score and select samples for efficient learning, we identify a fundamental misalignment between...
Jianpeng Chen, Wangzhi Zhan, Haohui Wang et al.· Proceedings of the 32nd ACM...· 0 citations
EvoHarness-RL is introduced, which exposes Belief, Progress, and Experience (BPE) as policy-facing harness state and reveals two key dynamics: harness annealing, where training internalizes recurring harness-use patterns into the model policy and shifts the agent from frequent harness calls toward selective external-st...
Xuying Ning, Dongqi Fu, Tianxin Wei et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.