Skip to content

Author

Xiangchi Yuan

We have 2 of 11 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and steering of agent honesty.

Zheyuan Liu, Weiliang Zhao, Xiangchi Yuan et al. · 0 citations

Steering Multimodal Large Language Models Decoding for Context-Aware Safety

Safety-aware Contrastive Decoding (SafeCoDe) is introduced, a lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context that consistently improves context-sensitive refusal behaviors while preserving model helpfulness.

Zheyuan Liu, Zhangchen Xu, Guangyao Dou et al. · 6 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.