Skip to content

Author

H. Bai

We have 6 of 21 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Oct 2026

HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation

Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO inco...

Tie-Zheng Yu, Yu-Xin Jiang, Jin-Peng Li et al. · 0 citations

OVD: On-policy Verbal Distillation

On-policy Verbal Distillation is introduced, a framework that uses verbal scores from black-box teachers to rank student-generated sub-trajectories, retaining high-scoring ones and replacing low-scoring ones with teacher-generated continuations and suggests that retaining student-generated prefixes helps preserve explo...

Jing Xiong, Hui Shen, Shansan Gong et al. · 8 citations
#machine learning Preprint Sep 2026

EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the...

Kun-Ming Shao, Jie-Run Chen, Jiang-Nan Yu et al. · 0 citations
Review Jul 2026

SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review

Experiments show that agentic review continuously improves PRs through a generate-review-revise loop, outperforms single-turn fixed-context review in both decision accuracy and resolve rate after revision, transfers beyond review to improve issue-resolution models, and enables effective and efficient test-time scaling.

Ruoyu Wang, Jierun Chen, Shaowei Wang et al. · 4 citations
Preprint Aug 2026

OptiMAS: Automatically Optimize Multi-Agent System

This work presents OptiMAS, a task-agnostic agentic optimizer that leverages textual interaction trajectories and task feedback as loss signals for end-to-end MAS evolution and sustains performance improvement over extended optimization horizons.

Yuxin Cheng, Chang Liu, Hanxin Yu et al. · 0 citations
Preprint Aug 2026

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward ha...

Yi-Ming Du, Yu-Xin Jiang, Tao Yuan et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.