Skip to content

Author

Lei Feng

We have 9 of 19 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Preprint Oct 2026

Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions

OrthoPurify is proposed, a more efficient method to purify backdoored model weights via one-step orthogonal projection, which reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead.

Bo-Jun Yang, Hao-Chen Zhou, Zhi-Fang Zhang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex...

Yu Li, Guang-Feng Cai, Long-Fei Li et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more...

Yu Li, Zheng Zhang, Xin Liu et al. · 0 citations
Preprint Sep 2026

SRPO: Setwise Relative Policy Optimization for Multi-Agent Systems

Multi-agent systems enable complex reasoning and tool use by coordinating agents that divide roles and refine candidate solutions. Existing methods typically update individual agent responses or treat a complete trajectory as one training example. However, these methods may produce misleading policy updates because the...

Sheng-Tian Yang, Zi-Yun Xiong, Yu Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs

Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when several outputs jointly cause one state transition. Consequently, the update unit differs fr...

Sheng-Tian Yang, Zi-Yun Xiong, Yu Li et al. · 4 citations
#artificial intelligence Preprint Sep 2026

AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

AgentBrew is proposed, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts, and demonstrates that fine-grained offline learning can recover useful supervision from raw trajectories that filterin...

Zhiyi Lyu, Ye-Wen Li, Long-Tao Zheng et al. · 2 citations
#machine learning Preprint Sep 2026

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, ofte...

Shen-Zhi Yang, Guang-Cheng Zhu, Kai Tang et al. · 0 citations
Preprint Aug 2026

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

AntiSkillBench is introduced, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline, and experiments show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits.

Yongli Xiang, Zhi-Fang Zhang, Bojun Yang et al. · 2 citations
Jul 2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Progress-conditioned Group Policy Optimization is proposed, which uses first-visit observation coverage only when all samples in a group receive zero outcome reward, and consistently improves over group-based baselines, with particularly large gains on hard tasks.

Kaibing Yang, Guangfeng Cai, Sheng-Tian Yang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.