Jul 2026
MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
A controlled study isolates the source of MADA-RL's gains: the counterfactual advantage produces the highest critic improvement rate of any model evaluated, indicating that trained critics learn to correct generator errors rather than to imitate them.
Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov et al.
· arXiv.org · 0 citations