Skip to content

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy d...

Yi-Tong Qiao, Tian-Tian He, Lei Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning

In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We intro...

Yi-Tong Qiao, Yan-Cheng Jin, Lei Liu et al. · 0 citations
#natural language process... Preprint Sep 2026

Can LLMs Value the Right Evidence? Evidence-Value Misalignment in Dynamic Medical Diagnosis

A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Misalignment (EVM). To disentangle evidential grounding indepen...

Ke-Hua Feng, Yun-Sheng Lu, Yi-Tong Qiao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EOPSA: Efficient On-Policy Self-Distilled Safety Alignment

On-Policy Self-Distillation (OPSD) has emerged as a promising paradigm for safety alignment, delivering dense, token-level supervision by distilling from a teacher conditioned on refusal-oriented privileged prompts. However, we reveal that this paradigm suffers from critical inefficiencies that degrade both training ef...

Qi-Rui Liu, Yi-Chen Sun, Yan Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.