A controlled study isolates the source of MADA-RL's gains: the counterfactual advantage produces the highest critic improvement rate of any model evaluated, indicating that trained critics learn to correct generator errors rather than to imitate them.
Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov et al.· arXiv.org· 0 citations
Experiments show that SearchEyes achieves state-of-the-art performance among open-source multimodal search agents, with SearchEyes-27B improving over the strongest open-source baseline by 6.2 points on average.
Zhengbo Jiao, YiMing Cheng, Yilei Jiang et al.· 1 citation
According to Mendelian principles of controlled inheritance, Mendel G\"odel Machine (MGM) is introduced, which includes two new types of self-modification that better utilizes evidences accumulated and facilitates a faster and better convergence over single-trajectory baselines.
Changzhi Liu, Yilun Liu, Sikuan Yan et al.· 0 citations
OPD-V is introduced, a visual OPSD paradigm that instantiates privileged information through the Positive Teacher and Negative Teacher that consistently improves reasoning performance while reducing training cost.
MetaSkill-Evolve is introduced, a two-timescale framework that makes agentic skill improvement recursive and outperforms no-skill, static-skill, and single-level evolution baselines on three agentic benchmarks, improving held-out test accuracy over the raw backbone by +23.54, +16.09, and +1.92 points respectively.
Zefeng Wang, Minxi Yan, Jinhe Bi et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.