Skip to content

Author

Huihao Jing

We have 6 of 19 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation....

Wen-Bin Hu, Hui-Hao Jing, Hao-Chen Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SoFT: Soft Targets for Generalizable LLM Fine-Tuning

This work proposes soft-target fine-tuning (SoFT) to balance learning from teacher demonstrations with retaining the Base model's existing capabilities, with improvements in both in-distribution capability acquisition and out-of-distribution generalization.

Hui-Hao Jing, Wen-Bin Hu, Shao-Jin Chen et al. · 0 citations
Preprint Jul 2026

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

Overall, persistent skill self-evolution is better understood as sparse, validation-filtered search with model- and benchmark-dependent returns, rather than steady improvement from additional rounds.

Yuxuan Liu, Zhaochen Su, Yuhao Zhang et al. · 2 citations
Review Jul 2026

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

This survey treats isolation as a first-class principle for LLM-agent system safety, and organizes the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment.

Huihao Jing, Wenbin Hu, Shaojin Chen et al. · 0 citations
Jul 2026

RLPF: Reinforcement Learning from Performance Feedback for Code Generation

This work proposes RLPF, reinforcement learning from performance feedback, which turns execution outcomes into a staged reward, and suggests that code agents can be trained not only to pass tests, but also to optimize the programs they write.

Huihao Jing, Hao-Zhe Cui, Wenbin Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.