Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy d...
Yi-Tong Qiao, Tian-Tian He, Lei Liu et al.· 0 citations
In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We intro...
Yi-Tong Qiao, Yan-Cheng Jin, Lei Liu et al.· 0 citations
A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Misalignment (EVM). To disentangle evidential grounding indepen...
Ke-Hua Feng, Yun-Sheng Lu, Yi-Tong Qiao et al.· 0 citations
ConRub-Med is introduced to preserve useful distinctions as rubric feedback moves from construction to policy optimization, and ranks first on six of nine benchmarks and achieves the highest medical and generalization averages.
Taojie Zhu, Yuan Xia, Tao Sun et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.