Skip to content

Author

Jicheng Wang

We have 5 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review Oct 2026

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...

Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation

ForkLeft, a distillation framework that resolves a fundamental mismatch between autoregressive teacher predicts from a left prefix and NTP teacher under the same context, is introduced, showing that DLMs can learn NTP-style reasoning without sacrificing native parallel generation.

Jun-Ming Liu, Ji-Cheng Wang, Yifeng He et al. · 0 citations
#artificial intelligence Preprint Sep 2026

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

BenchShield is presented, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation that grounds detection in a finite lifecycle model of an evaluation's reward-relevant events and achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.

Sheng-Han Zheng, Zong-Lin Di, Yimin Liu et al. · 0 citations
Jul 2026

Is Progressive Disclosure All You Need for Long-Context Agents?

The first controlled study of the pattern is run, comparing raw-document navigation and several designs of Agent Skills packs against a classical hybrid retriever across three agent harnesses and three model families on InfiniteBench.

Yifeng He, Yin Zhao, Jicheng Wang et al. · 2 citations
Preprint Jul 2026

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that...

Hanwen Xing, Pengyu Wang, Bingxu Meng et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.