Skip to content

Author

Kewei Fu

We have 3 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a \textit{reward fairness problem} and propose \textbf{ARC} (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core $\tau/\tau^2$ tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.

Yongqi Tong, Tan Li Hui Faith, C. Marcus et al. · 0 citations
Preprint Aug 2026

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

ACA-RL supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.

Yong-Qi Tong, Zhenyu Zhang, Zimou Liu et al. · 0 citations
Preprint Aug 2026

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

This work proposes \methodname, a stability-guided active-set controller for controlled objective admission, a stability-guided active-set controller for controlled objective admission in reward-vector RLHF, which positions objective-entry timing as a concrete control variable in reward-vector RLHF.

Yong-Qi Tong, Z. Zhang, Ruirui Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.