Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an an...
Jia-Rui Liu, Ren-Jie Tao, Yi-Wei Liao et al.· 0 citations
This work introduces RAPTOR - a Role-Aware Private Training framework, which alternates shared and expert optimization and targets each failure directly, using expert-specific clipping and noise together with a public expected-owner denominator and a count-independent update schedule that avoids conditioning on private...
Duc Dm, Khai Le-Duc, D. Nguyen et al.· 0 citations
It is shown that LLM-as-judge RL induces reward-hacking patterns, and that LLM-as-judge RL detectors can mitigate them during post-training, suggesting that behavioral foundation models require rethinking the LLM training paradigm.
Xuhui Zhou, Weiwei Sun, Weihua Du et al.· arXiv.org· 8 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.