Skip to content

Author

Kehai Chen

We have 6 of 92 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Un...

Zi-Hao Chen, Fan-Xiang Xiong, Hong-Ran Ren et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) f...

Shao-Qing Zhang, Ke-Hai Chen, Xue-Feng Bai et al. · 0 citations
Preprint Aug 2026

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with pseudo-labels constructed through majority voting. While effective, the reward signal assigned from majority voting is highly sensitive to consensus strength, defined as...

Kunbin Xu, Xingzuo Li, Xue-Feng Bai et al. · 1 citation
Preprint Sep 2026

Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

DM-Align is introduced, which derives a complementary gradient direction to guide the model toward human-preferred samples, and eliminates the need for multi-step reward evaluation and complex ODE-SDE conversions inherent in traditional RL.

Jiu-Zhou Lin, Jun-Long Wu, Feilong Zuo et al. · 1 citation
Jul 2026

DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation

DualAnchor is proposed, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation that achieves strong overall performance on both PHOENIX-2014T and CSL-Daily.

Hongbin Zhang, Jun-Hao Liu, Xue-Feng Bai et al. · 0 citations
Conference Open access 2026

Diagnosing and Remedying Representation Deficiencies for Deterministic Reasoning in KGQA

The proposed DR 2 identifies and localizes non-deterministic reasoning behaviors, uncovering the underlying semantic representation deficiencies in LLMs, and designs abductive reasoning-based preference learning, which promotes fine-grained semantic discrimination and mitigates non-deterministic reasoning errors.

Ge Liang, Mufan Xu, Kehai Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.